Qwen-Image-2.1 · MNN (int4) for Android
Qwen/Qwen-Image-2.1 converted to MNN
for on-device text-to-image and image editing on Android, with the OpenCL GPU running the DiT. Any size with sides
a multiple of 32 works, from 256×256 up; the app offers 7 aspect ratios at three pixel budgets (~512², ~384², ~320²),
e.g. 512×512, 576×448, 672×384, 480×320, 384×288. An optional Turbo LoRA (dit_turbo.mnn) cuts a run from
20–40 steps to a fixed 6. The VAE is TAEQI2.1 by default — a
distilled VAE ~1/20th the size of the real one and ~200x faster to decode, on the GPU instead of a forced CPU
fallback — see Tiny VAE below.
Runtime, Android library and demo app: github.com/scsonic/libQwenImage21
Left two: base model, 20 steps. Right three: Turbo, 6 steps — a text-to-image portrait, then two edits of it (same face, new outfit and background). All generated on a phone; see Turbo (6-step) below for the rest of the set and per-run timings.
| Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 |
|---|---|
| Text to image, 448×576, 20 steps | 451 s total (DiT 19.1 s/step on OpenCL fp16) |
| Image edit, 352×448, 20 steps | 348 s total (DiT 12.6 s/step) |
| Text to image, 512×512, Turbo 6 steps | ~216–235 s total (DiT ~22 s/step) |
| Image edit, 352×448, Turbo 6 steps | ~196–200 s total |
Files
| Path | What | Size |
|---|---|---|
dit.mnn + .weight |
7B single-stream DiT (32 blocks + norm_out/proj_out). Block linears int4 (block 32), small layers int8 | 4.5 GB |
dit_turbo.mnn + .weight |
Same DiT with the Viggle-turbo LoRA applied unmerged (a small extra fp16 branch per targeted layer); fixed 6-step schedule. Optional — see below | 5.2 GB |
dit_2bit.mnn + .weight |
Same DiT from the GGUF repo's Q2_K build instead of Q4_K — smaller, not faster (see below). Optional | 3.6 GB |
dit_2bit_turbo.mnn + .weight |
dit_2bit.mnn with Turbo applied the same unmerged way. Optional |
4.3 GB |
txt_in.mnn, img_in.mnn |
text (int8) / latent (fp16) input projections | 36 MB |
vae_decoder_tiny.mnn, vae_encoder_tiny.mnn |
TAEQI2.1, a distilled VAE matching the real one's latent directly. Default — see below | 31 MB |
vae_decoder.mnn, vae_encoder.mnn |
The real Qwen-Image-2.1 VAE, 64-ch latent ↔ RGBA, fp16 weights, dynamic size. Optional — swap in for max fidelity | 0.66 GB |
text_encoder/ |
Qwen3-VL-8B-Instruct, MNN int4 (from taobao-mnn/Qwen3-VL-8B-Instruct-MNN). te_config.json runs it text-only; te_vl_config.json adds the vision tower (visual.mnn) for image editing. Both return the last decoder layer before the final norm |
5.4 GB |
Default download (tiny VAE, skips the real VAE's 0.66 GB) — add --exclude "dit_turbo.mnn*" to also skip Turbo, or
drop the --exclude on the VAE files for the real one instead/as well:
hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21 --exclude "vae_decoder.mnn" --exclude "vae_encoder.mnn"
Turbo (6-step)
Viggle-turbo distills Qwen-Image-2.1 to a fixed
6-step schedule with no CFG. dit_turbo.mnn applies it unmerged, the way diffusers and the LoRA's own
ComfyUI node do: the int4 base weights are untouched (identical to dit.mnn's), and the LoRA's rank-128
correction is added as a small extra fp16 branch per targeted layer, exported as its own MNN model file that
happens to carry a second copy of the base weights (MNN has no format for patching an already-compiled graph).
Merging the correction into the weights instead — especially into int4 — is what the LoRA's own README
specifically measures as lossy; unmerged is the accurate path. txt_in.mnn, img_in.mnn, both VAE models and
the text encoder are unaffected and shared with the base pipeline.
Text-to-image (studio portrait, kimono, Harajuku street fashion) and two edits of the studio portrait — same face, new outfit and setting. All on a Snapdragon 8 Gen 2, OpenCL, 6 steps.
| text encoder (+ prefix) | DiT (6 steps) | VAE | total | |
|---|---|---|---|---|
| Text to image, 448×576–512×512 | ~15 s | 6 × ~22 s ≈ 134 s | ~18 s | 216–235 s |
| Image edit → 352×448 | ~20–80 s (with vision; P≈680) | 6 × ~18 s ≈ 108 s | ~13 s | 196–200 s |
A Turbo DiT step (22 s at ~512²) is slower than a base-model step at the same size (19 s) — the extra fp16
branch costs roughly the 10–25% diffusers/ComfyUI themselves measure — but 6 steps instead of 20 still roughly
halves the total time for both modes. Base model at 6 steps without the LoRA is visibly worse (soft, muddy) —
the schedule alone isn't what's doing the work.
Load it like the base model, just with dit_turbo.mnn instead of dit.mnn; the demo app has a Turbo LoRA
checkbox that fixes the step count to 6. See the runtime repo
for the CLI/library API.
Tiny VAE (default)
TAEQI2.1 by madebyollin is a tiny distilled autoencoder — a handful
of convolution layers trained to match Qwen-Image-2.1's real VAE at its own "latent API" directly (16x spatial
compression, 64 latent channels, RGBA), so it's a drop-in replacement: same tensor names, shapes and value ranges
as vae_decoder.mnn / vae_encoder.mnn. The real VAE decoder needed ~4.3 GB of runtime memory at 512² — enough
that it exhausted GPU memory and rebooted the test phone, so it always ran on CPU, at ~19.5 s/decode. TAEQI2.1 is
~15 MB per direction, runs on the GPU with no memory concerns, and decodes in well under a tenth of a second
once the shader cache is warm.
Four prompts, Turbo 6 steps, decoded with TAEQI2.1. Measured against the real VAE on the same seeds/prompts
(Snapdragon 8 Gen 2, vae_on_cpu forced for the real VAE since GPU isn't safe for it):
| Prompt | Real VAE decode / RAM | Tiny VAE decode / RAM | Total time (real → tiny) |
|---|---|---|---|
| "A red apple on a wooden table..." | 19.63 s / 4326 MB | 0.57 s¹ / 187 MB | 220.9 s → 202.8 s |
| "A cat sitting on a windowsill at sunset" | 19.54 s / 4326 MB | 0.09 s / 187 MB | 226.0 s → 202.1 s |
| "A futuristic city skyline at night..." | 19.54 s / 4326 MB | 0.09 s / 187 MB | 225.9 s → 203.5 s |
| "A bowl of ramen noodles..., steam rising" | 19.54 s / 4326 MB | 0.09 s / 187 MB | 229.4 s → 207.5 s |
¹ First run after install pays a one-time OpenCL shader-compile cost; every run after that was 0.09 s.
export/export_vae_tiny.py in the runtime repo rebuilds TAEQI2.1's exact architecture (from
taesd.py, MIT) in plain PyTorch and exports it with the
same wrapper conventions as the real VAE's own export script, so the two are interchangeable at the runtime level.
Verified two ways before touching the phone: the published weights load strict=True into the reimplemented
architecture, and decoding the real VAE encoder's actual latents with the tiny decoder reproduces the source
photo essentially unchanged. Full write-up: docs/TINY_VAE.md.
It's the demo app's default (a Tiny VAE checkbox, checked) and the library's default (Options.tinyVae = true);
uncheck it / set it false for the original real VAE.
2-bit (optional)
dit_2bit.mnn is the same lossless-repack approach as the Q4_K dit.mnn, applied to the GGUF repo's Q2_K
quantization instead: every Q2_K sub-block (w = d·sc·q − dmin·m, q in [0,3], 16 weights per sub-block) maps
exactly onto MNN's asymmetric int2, copied without re-quantization. Combines with Turbo the same unmerged way
(dit_2bit_turbo.mnn).
The honest result: smaller, not faster. 3.6 GB instead of 4.5 GB (20%, not 50% — Q2_K's native block size is 16 instead of 32, so per-block overhead eats into the saving), but each DiT step measured slightly slower than int4's on this runtime, since MNN's OpenCL low-memory int2 kernel takes a less-optimized path than int4's. If you want speed, use Turbo; if you want a smaller download, 2-bit delivers that.
| DiT step | Total (512×512) | File size | |
|---|---|---|---|
| int4 | ~19.1 s | ~451 s (20 steps, 448×576) | 4.5 GB |
| int2 | ~20.5 s | 495.9 s (20 steps) | 3.6 GB |
| int4 + Turbo | ~22 s | ~216–235 s (6 steps) | 5.2 GB |
| int2 + Turbo | ~24.3 s | 242.1 s (6 steps) | 4.3 GB |
Getting this working also surfaced and fixed two bugs in MNN's own OpenCL low-memory int2/int3 path (a
use-after-free on load, and a missing dequant-offset correction that produced NaN) — not exercised before since
GGUF Q2_K's block size of 16 is new territory for that code path. Full write-up, images and the MNN fix details:
docs/TWOBIT.md.
How it was made
- DiT: taken from the GGUF Q4_K build (leejet/Qwen-Image-2.1-GGUF),
or optionally its Q2_K build (
dit_2bit.mnn— see 2-bit above). Every Q4_K/Q2_K sub-block (w = d·sc·q − dmin·m) maps exactly onto MNN's asymmetric int4/int2, so the weights are copied without re-quantization (scales stored as fp16). - Text encoder: Qwen-Image-2.1's
text_encoderis byte-identical to Qwen3-VL-8B-Instruct, so the existing MNN export is reused unchanged. - VAE: the residual stream is divided by 256 (exact, power of two) and RMSNorm pre-divides by max|x| so the decoder fits fp16 (it peaks at ~3.5e5 otherwise). The decoded image is unchanged.
- The pipeline caches the text K/V once per prompt (Qwen-Image-2.1's block-causal attention), so each denoising step
only runs the image tokens. The cache is one tensor per layer (
past_kv_0…past_kv_31): a single[32, 2, P, 32, 128]tensor is exactly 1 MiB per prefix token, and an image-edit prefix (P > 1000) would exceed OpenCL's 1 GiB maximum buffer size on an Adreno 740.
2026-09-23:
dit.mnn/dit.mnn.weightwere re-exported for that per-layer K/V cache. Older copies do not load with the current runtime — re-download both files. 2026-09-25: addeddit_turbo.mnn/.weight(optional, see Turbo above). 2026-10-01: addedvae_decoder_tiny.mnn/vae_encoder_tiny.mnn(default now — see Tiny VAE above); the realvae_decoder.mnn/vae_encoder.mnnare still here, now opt-in. 2026-10-01: addeddit_2bit.mnn/dit_2bit_turbo.mnn(optional — see 2-bit above).
Conversion scripts: export/ in the GitHub repo.
License
Derived from Qwen-Image-2.1 and released under the Qwen Research License Agreement (see LICENSE), i.e. for
research / non-commercial use under its terms. The text encoder weights come from Qwen3-VL-8B-Instruct (Apache-2.0).
The Turbo LoRA is a derivative of the same base model, released by Viggle under the same Qwen Research License
Agreement. TAEQI2.1 is released by madebyollin under the MIT license.
Model tree for evankuo/Qwen-Image-2.1-MNN
Base model
Qwen/Qwen-Image-2.1