Qwen-Image-2.1 · MNN (int4) for Android

Qwen/Qwen-Image-2.1 converted to MNN for on-device text-to-image and image editing on Android, with the OpenCL GPU running the DiT. Any size with sides a multiple of 32 works, from 256×256 up; the app offers 7 aspect ratios at three pixel budgets (~512², ~384², ~320²), e.g. 512×512, 576×448, 672×384, 480×320, 384×288. An optional Turbo LoRA (dit_turbo.mnn) cuts a run from 20–40 steps to a fixed 6. The VAE is TAEQI2.1 by default — a distilled VAE ~1/20th the size of the real one and ~200x faster to decode, on the GPU instead of a forced CPU fallback — see Tiny VAE below.

Runtime, Android library and demo app: github.com/scsonic/libQwenImage21

Left two: base model, 20 steps. Right three: Turbo, 6 steps — a text-to-image portrait, then two edits of it (same face, new outfit and background). All generated on a phone; see Turbo (6-step) below for the rest of the set and per-run timings.

Tested on Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13
Text to image, 448×576, 20 steps 451 s total (DiT 19.1 s/step on OpenCL fp16)
Image edit, 352×448, 20 steps 348 s total (DiT 12.6 s/step)
Text to image, 512×512, Turbo 6 steps ~216–235 s total (DiT ~22 s/step)
Image edit, 352×448, Turbo 6 steps ~196–200 s total

Files

Path What Size
dit.mnn + .weight 7B single-stream DiT (32 blocks + norm_out/proj_out). Block linears int4 (block 32), small layers int8 4.5 GB
dit_turbo.mnn + .weight Same DiT with the Viggle-turbo LoRA applied unmerged (a small extra fp16 branch per targeted layer); fixed 6-step schedule. Optional — see below 5.2 GB
dit_2bit.mnn + .weight Same DiT from the GGUF repo's Q2_K build instead of Q4_K — smaller, not faster (see below). Optional 3.6 GB
dit_2bit_turbo.mnn + .weight dit_2bit.mnn with Turbo applied the same unmerged way. Optional 4.3 GB
txt_in.mnn, img_in.mnn text (int8) / latent (fp16) input projections 36 MB
vae_decoder_tiny.mnn, vae_encoder_tiny.mnn TAEQI2.1, a distilled VAE matching the real one's latent directly. Default — see below 31 MB
vae_decoder.mnn, vae_encoder.mnn The real Qwen-Image-2.1 VAE, 64-ch latent ↔ RGBA, fp16 weights, dynamic size. Optional — swap in for max fidelity 0.66 GB
text_encoder/ Qwen3-VL-8B-Instruct, MNN int4 (from taobao-mnn/Qwen3-VL-8B-Instruct-MNN). te_config.json runs it text-only; te_vl_config.json adds the vision tower (visual.mnn) for image editing. Both return the last decoder layer before the final norm 5.4 GB

Default download (tiny VAE, skips the real VAE's 0.66 GB) — add --exclude "dit_turbo.mnn*" to also skip Turbo, or drop the --exclude on the VAE files for the real one instead/as well:

hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21 --exclude "vae_decoder.mnn" --exclude "vae_encoder.mnn"

Turbo (6-step)

Viggle-turbo distills Qwen-Image-2.1 to a fixed 6-step schedule with no CFG. dit_turbo.mnn applies it unmerged, the way diffusers and the LoRA's own ComfyUI node do: the int4 base weights are untouched (identical to dit.mnn's), and the LoRA's rank-128 correction is added as a small extra fp16 branch per targeted layer, exported as its own MNN model file that happens to carry a second copy of the base weights (MNN has no format for patching an already-compiled graph). Merging the correction into the weights instead — especially into int4 — is what the LoRA's own README specifically measures as lossy; unmerged is the accurate path. txt_in.mnn, img_in.mnn, both VAE models and the text encoder are unaffected and shared with the base pipeline.

Text-to-image (studio portrait, kimono, Harajuku street fashion) and two edits of the studio portrait — same face, new outfit and setting. All on a Snapdragon 8 Gen 2, OpenCL, 6 steps.

text encoder (+ prefix) DiT (6 steps) VAE total
Text to image, 448×576–512×512 ~15 s 6 × ~22 s ≈ 134 s ~18 s 216–235 s
Image edit → 352×448 ~20–80 s (with vision; P≈680) 6 × ~18 s ≈ 108 s ~13 s 196–200 s

A Turbo DiT step (22 s at ~512²) is slower than a base-model step at the same size (19 s) — the extra fp16 branch costs roughly the 10–25% diffusers/ComfyUI themselves measure — but 6 steps instead of 20 still roughly halves the total time for both modes. Base model at 6 steps without the LoRA is visibly worse (soft, muddy) — the schedule alone isn't what's doing the work.

Load it like the base model, just with dit_turbo.mnn instead of dit.mnn; the demo app has a Turbo LoRA checkbox that fixes the step count to 6. See the runtime repo for the CLI/library API.

Tiny VAE (default)

TAEQI2.1 by madebyollin is a tiny distilled autoencoder — a handful of convolution layers trained to match Qwen-Image-2.1's real VAE at its own "latent API" directly (16x spatial compression, 64 latent channels, RGBA), so it's a drop-in replacement: same tensor names, shapes and value ranges as vae_decoder.mnn / vae_encoder.mnn. The real VAE decoder needed ~4.3 GB of runtime memory at 512² — enough that it exhausted GPU memory and rebooted the test phone, so it always ran on CPU, at ~19.5 s/decode. TAEQI2.1 is ~15 MB per direction, runs on the GPU with no memory concerns, and decodes in well under a tenth of a second once the shader cache is warm.

Four prompts, Turbo 6 steps, decoded with TAEQI2.1. Measured against the real VAE on the same seeds/prompts (Snapdragon 8 Gen 2, vae_on_cpu forced for the real VAE since GPU isn't safe for it):

Prompt Real VAE decode / RAM Tiny VAE decode / RAM Total time (real → tiny)
"A red apple on a wooden table..." 19.63 s / 4326 MB 0.57 s¹ / 187 MB 220.9 s → 202.8 s
"A cat sitting on a windowsill at sunset" 19.54 s / 4326 MB 0.09 s / 187 MB 226.0 s → 202.1 s
"A futuristic city skyline at night..." 19.54 s / 4326 MB 0.09 s / 187 MB 225.9 s → 203.5 s
"A bowl of ramen noodles..., steam rising" 19.54 s / 4326 MB 0.09 s / 187 MB 229.4 s → 207.5 s

¹ First run after install pays a one-time OpenCL shader-compile cost; every run after that was 0.09 s.

export/export_vae_tiny.py in the runtime repo rebuilds TAEQI2.1's exact architecture (from taesd.py, MIT) in plain PyTorch and exports it with the same wrapper conventions as the real VAE's own export script, so the two are interchangeable at the runtime level. Verified two ways before touching the phone: the published weights load strict=True into the reimplemented architecture, and decoding the real VAE encoder's actual latents with the tiny decoder reproduces the source photo essentially unchanged. Full write-up: docs/TINY_VAE.md.

It's the demo app's default (a Tiny VAE checkbox, checked) and the library's default (Options.tinyVae = true); uncheck it / set it false for the original real VAE.

2-bit (optional)

dit_2bit.mnn is the same lossless-repack approach as the Q4_K dit.mnn, applied to the GGUF repo's Q2_K quantization instead: every Q2_K sub-block (w = d·sc·q − dmin·m, q in [0,3], 16 weights per sub-block) maps exactly onto MNN's asymmetric int2, copied without re-quantization. Combines with Turbo the same unmerged way (dit_2bit_turbo.mnn).

The honest result: smaller, not faster. 3.6 GB instead of 4.5 GB (20%, not 50% — Q2_K's native block size is 16 instead of 32, so per-block overhead eats into the saving), but each DiT step measured slightly slower than int4's on this runtime, since MNN's OpenCL low-memory int2 kernel takes a less-optimized path than int4's. If you want speed, use Turbo; if you want a smaller download, 2-bit delivers that.

DiT step Total (512×512) File size
int4 ~19.1 s ~451 s (20 steps, 448×576) 4.5 GB
int2 ~20.5 s 495.9 s (20 steps) 3.6 GB
int4 + Turbo ~22 s ~216–235 s (6 steps) 5.2 GB
int2 + Turbo ~24.3 s 242.1 s (6 steps) 4.3 GB

Getting this working also surfaced and fixed two bugs in MNN's own OpenCL low-memory int2/int3 path (a use-after-free on load, and a missing dequant-offset correction that produced NaN) — not exercised before since GGUF Q2_K's block size of 16 is new territory for that code path. Full write-up, images and the MNN fix details: docs/TWOBIT.md.

How it was made

  • DiT: taken from the GGUF Q4_K build (leejet/Qwen-Image-2.1-GGUF), or optionally its Q2_K build (dit_2bit.mnn — see 2-bit above). Every Q4_K/Q2_K sub-block (w = d·sc·q − dmin·m) maps exactly onto MNN's asymmetric int4/int2, so the weights are copied without re-quantization (scales stored as fp16).
  • Text encoder: Qwen-Image-2.1's text_encoder is byte-identical to Qwen3-VL-8B-Instruct, so the existing MNN export is reused unchanged.
  • VAE: the residual stream is divided by 256 (exact, power of two) and RMSNorm pre-divides by max|x| so the decoder fits fp16 (it peaks at ~3.5e5 otherwise). The decoded image is unchanged.
  • The pipeline caches the text K/V once per prompt (Qwen-Image-2.1's block-causal attention), so each denoising step only runs the image tokens. The cache is one tensor per layer (past_kv_0…past_kv_31): a single [32, 2, P, 32, 128] tensor is exactly 1 MiB per prefix token, and an image-edit prefix (P > 1000) would exceed OpenCL's 1 GiB maximum buffer size on an Adreno 740.

2026-09-23: dit.mnn / dit.mnn.weight were re-exported for that per-layer K/V cache. Older copies do not load with the current runtime — re-download both files. 2026-09-25: added dit_turbo.mnn / .weight (optional, see Turbo above). 2026-10-01: added vae_decoder_tiny.mnn / vae_encoder_tiny.mnn (default now — see Tiny VAE above); the real vae_decoder.mnn / vae_encoder.mnn are still here, now opt-in. 2026-10-01: added dit_2bit.mnn / dit_2bit_turbo.mnn (optional — see 2-bit above).

Conversion scripts: export/ in the GitHub repo.

License

Derived from Qwen-Image-2.1 and released under the Qwen Research License Agreement (see LICENSE), i.e. for research / non-commercial use under its terms. The text encoder weights come from Qwen3-VL-8B-Instruct (Apache-2.0). The Turbo LoRA is a derivative of the same base model, released by Viggle under the same Qwen Research License Agreement. TAEQI2.1 is released by madebyollin under the MIT license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for evankuo/Qwen-Image-2.1-MNN

Quantized
(113)
this model