GLM-5.3-Flash — W4A16 (INT4) + BF16 MTP

Model description

INT4 weight-only quantization of zai-org/GLM-5.3-Flash with the BF16 MTP draft head kept for speculative decoding. Only the 36,288 routed-expert GEMMs are INT4 (GPTQ, symmetric, group-size 128); attention, router, shared experts, embeddings, the vision tower and the MTP head stay in BF16. Not a new model — all capability comes from the base model.

Size 177.7 GiB (BF16 ≈ 599 GiB, −70%)
Runs on NVIDIA H100, H200, RTX PRO 6000, and DGX Spark (GB10) — validated GPU counts and context per config in Hardware & context limits
Context full 1,048,576 tokens on 2× DGX Spark and 2× H200; 262K–512K on the 4/8-GPU x86 configs (KV-memory-bound)
Quality AIME 2025 0.8833 (n=120) vs 0.9000 for the NVFP4 reference on H100, within noise; GSM8K 0.97; GPQA-Diamond within noise; vision: MMMU 0.747 vs 0.68 NVFP4 (same rig) · OCRBench 888 vs 882
Throughput 8× H100: 181.2 tok/s single-stream latency recipe (TP8 + DFlash2-G K=7, +46%), 2,418 tok/s @c256 aggregate (TP8 spec-off); 4× H200 TP=4: 195 → 1,954 tok/s from c1 to c256; beats NVIDIA's NVFP4 checkpoint on H100 (+47.8% c1, +5.1% c256); matches the NVFP4 reference on RTX PRO 6000

Full benchmark grids, comparison protocols and research notes: BENCHMARKS.md.

Uses & recommended recipes

Quick start

# 1. Download (~178 GiB)
huggingface-cli download canada-quant/GLM-5.3-Flash-W4A16-MTP --local-dir /models/glm53-flash-w4a16-mtp

# 2. Serve on 4× H100 / H200 (other hardware: see Serving recipes)
docker run --gpus '"device=0,1,2,3"' --ipc=host --network=host --rm \
  -v /models:/models vllm/vllm-openai:glm53-flash-x86_64-cu130 \
  vllm serve /models/glm53-flash-w4a16-mtp --served-model-name glm53-w4 \
    --tensor-parallel-size 4 --enable-expert-parallel \
    --max-model-len 262144 --max-num-seqs 512 \
    --gpu-memory-utilization 0.92 --no-enable-prefix-caching \
    --speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
    --reasoning-parser glm45 --tool-call-parser glm47 --enable-auto-tool-choice \
    --trust-remote-code --port 8000

# 3. Call it (OpenAI-compatible)
curl http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "glm53-w4",
  "messages": [{"role": "user", "content": "Prove there are infinitely many primes."}]
}'

Two things that bite. If your config.json predates 2026-09-08, re-download it — older copies fail in vLLM with KeyError: 'layers.0.mlp.gate_up_proj.weight' (weights are unchanged). And always pass --max-num-seqs ≤ 512 — the vLLM default of 1024 exceeds this hybrid linear-attention model's 512 Mamba/KDA-state cache blocks.

Serving recipes

Architecture Image
SM90 (H100 / H200) vllm/vllm-openai:glm53-flash-x86_64-cu130 (validated). Upstream vllm/vllm-openai:nightly-x86_64 ≥ 2026-09-08 also boots this checkpoint (vllm-project/vllm#53906); pass --attention-backend FLASH_ATTN_MLA_SPARSE there, its default backend faults on ≥131K prompts.
SM90, 8× H100 + DFlash2-G drafter ghcr.io/canada-quant/vllm-glm53-flash-h100:v2-w4a16-dflash2 — vLLM nightly pin + baked drafter patches (EAGLE3 aux taps, drafter-aware KV partitioning); image of record for the 2026-09-28 8×H100 sweep
SM120 (RTX PRO 6000) cstechdev/vllm:glm53-flash-nope-sm120-cu130-20260826-r1
SM121 (DGX Spark) ghcr.io/canada-quant/vllm-glm53-flash-sm121:v2-w4a16-dflash2e, built from canada-quant/vllm-glm53-flash-sm121 with the two SM121 serving patches baked in

All configurations use expert parallelism. fp8 KV is not available on Hopper for this NoPE model.

H100 / H200, TP=4 (pinned image) — the Quick start command is the benchmarked recipe. num_speculative_tokens: 2 is the sweet spot on Hopper (52–55% acceptance; N=5 collapses acceptance to ~30%). Keep prefix caching off — it measured −2…−5% on H200. On 8× H200, two independent TP=4 replicas behind a load balancer beat one TP=8 endpoint by +28–38% aggregate at c128–c512; TP=8 wins single-stream and holds one 7.8M-token pool. On 2× H200 (89.5 GiB weights per GPU) add --tensor-parallel-size 2 --max-num-seqs 128 --max-cudagraph-capture-size 128 for 262K, or --max-model-len 1048576 --max-num-seqs 16 --gpu-memory-utilization 0.95 --max-cudagraph-capture-size 64 --max-num-batched-tokens 4096 for 1M.

8× H100 + DFlash2-G (v2 image, 2026-09-28 closeout) — latency recipe = TP8 + DFlash2-G K=7: c1 181.2 tok/s, +46% vs spec-off (124.1). Aggregate recipe = TP8 spec-off or MTP-N2: c256 2,418 / 2,330 tok/s; TP8 MTP-N2 is the balanced middle with the best grid c1 (215.0). Do not run DFlash2 at TP4 past c128 — the drafter's 8 standalone FullAttention KV tensors cut the KV pool ~6× and aggregate collapses (structural, confirmed uncontended: 694 tok/s @c128 vs TP8's 2,220); MTP-N2's shared-embedding draft softens the cliff (solo-TP4 c1 193.2, the best single-stream shape measured on this rig, c256 1,292). With any drafter on this image, cap --max-model-len ≤ 131071 — 128K-token prefills die on the speculative path (spec-off survives them). Full grids: BENCHMARKS.md. On the same rig and harness, NVIDIA's own NVFP4 checkpoint — an emulated FP4 path on Hopper (no native SM90 FP4 kernels; the engine warns at boot) — delivers 122.56 tok/s c1 / 2,299.84 @c256 at its best arm against our 181.2 / 2,418.1: on Hopper, use our quant + drafter + recipe (full head-to-head).

8× H100, TP=8, DFlash2-G (bench-validated 2026-09-28) — single node, 8192/1024 prompts: the latency recipe is TP=8 + DFlash2-G K=7: 181.2 tok/s c1, +46% vs spec-off 124.1; the aggregate recipe is TP=8 spec-off: 2,418 tok/s @c256 (MTP N=2 at 2,330 is the balanced middle and the best c1 at 215.0). K=4 beats K=7 on aggregate at every concurrency. Avoid TP=4 with the DFlash2 drafter past c32: its 8 standalone full-attention KV tensors cut the KV pool ~6× vs TP8 (311,999 vs 1,811,949 tokens) and force a preemption/recompute TTFT cliff that is structural, not contention — solo-TP4 measured 160.5 K=7 / 165.7 K=4 / 123.9 spec-off / 193.18 MTP-N2 tok/s c1, MTP-N2 being the best single-stream shape measured on this rig at any TP (1,292.4 @c256, the only speculative arm still rising there). With any drafter, cap --max-model-len ≤ 131071 — the speculative path faults on 131K-token prefills (spec-off survives); image ghcr.io/canada-quant/vllm-glm53-flash-h100:v2-w4a16-dflash2, cold boot ≈ 12 min.

RTX PRO 6000, TP=4 — same command with the SM120 image and --max-num-seqs 64 --max-num-batched-tokens 8192 --kv-cache-dtype fp8 --enable-prefix-caching. fp8 KV is required at 262K on 96 GB cards. Keep MTP on at every concurrency here: it adds +70% at c1 and +48% at c32.

2× DGX Spark, TP=2, 1M context — prebuilt image and one-command launcher in canada-quant/vllm-glm53-flash-sm121; the launcher also ships in the drafter repo. Drafter: canada-quant/GLM-5.3-Flash-DFlash2-G (the authors' self-trained DFlash2 drafter, Apache-2.0; 3.676 mean acceptance at K=7 on the 500-prompt holdout vs 3.632 for the incoai reference measured on the same hardware; drop-in successor of -F and -E, E having carried the 2026-09-22 banked Spark row and G being the Spark drafter of record since 2026-09-25 (the 2026-09-27 H2H champion leg and the live production serve both ran G, sha256-verified) — DRAFTER_HOST_PATH=/models/GLM-5.3-Flash-DFlash2-G). Start the worker rank first, wait 25 s, then the head rank.

# on both nodes, rank1 (worker) first, then rank0 (head) 25 s later
MAX_MODEL_LEN=1048576 KV_CACHE_MEM=9663676416 GMU=0.90 EAGER=0 GRAPHS=1 bash launch_dflash2_tp2.sh <rank>

Hard constraints (Spark): num_speculative_tokens must be 7 (any other count wedges boot); confirm the boot log shows the mask-embedding load (mask_token_id 154856); keep single prompts ≤ ~310K tokens; stop with docker stop -t 30, never rm -f. Cold boot is 6–10 minutes.

Quality

Benchmark Hardware W4A16 (this) Reference
AIME 2025 — n=120, max thinking, 131,072-token budget H100 0.8833 NVFP4 0.9000 — within noise (0.42σ)
AIME 2025 — same protocol RTX PRO 6000 0.8083 raw · 0.8833 with a budget-commit fix NVFP4 0.9000
AIME 2026 — n=120, max thinking, 131,072-token budget 2× DGX Spark 85.0% (102/120) EXL3 80.0% (96/120), matched protocol
GSM8K H100, RTX PRO 6000 0.970–0.975 parity across quants
GPQA-Diamond — n=198 @131K H100 · RTX PRO 6000 0.8586 · 0.8586 NVFP4 0.8687 · 0.8737 — within noise
MMMU (validation) — n=900, lmms-eval 0.7.3 task spec 8× H100 TP=8, spec-off, greedy 0.74667 (672/900) NVIDIA NVFP4 ckpt, same rig (2026-09-28): 0.68; first formal for this artifact; MC 0.766 · open 0.434
OCRBench — n=1000, lmms-eval 0.7.3 task spec 8× H100 TP=8, spec-off, greedy 888 /1000 NVIDIA NVFP4 ckpt, same rig (2026-09-28): 882; weakest: handwritten math 56/100; strongest: doc-VQA 192/200

On RTX PRO 6000, ≈63% of the raw AIME 2025 deficit is a budget wall (empty-answer rate 11.7–14.2% vs 3.3%) and ≈37% is SM120 kernel numerics; a zero-cost commit hook closes it but is not part of the published recipes. Vision (formal, 2026-09-28): MMMU validation 0.74667 (672/900) and OCRBench 888/1000 on the 8×H100 closeout stack (lmms-eval 0.7.3 task specs, greedy temp 0, seed 42). The vision tower is BF16 passthrough, not covered by the text-only calibration — these are the artifact's own vision baseline; the only same-rig vision baseline measured to date is the NVIDIA NVFP4 checkpoint (MMMU 0.68 / OCRBench 882, 2026-09-28). Details in BENCHMARKS.md.

Benchmarks summary

Output tok/s, thinking ON, same hardware, flags and prompts within each row.

Hardware W4A16 (this) Reference Read
8× H100 (SM90), p5 grid, 8192/1024 latency recipe TP8 + DFlash2-G K=7: c1 181.2 (+46% vs spec-off 124.1); aggregate recipe TP8 spec-off: c256 2,418; TP8 MTP-N2: c1 215.0 · c256 2,330 — (cross-rig; same-rig row below) solo-TP4: MTP-N2 c1 193.2 = best single-stream shape measured · c256 1,292; spec-off 1,549.9 @c256 (no cliff without the drafter); DFlash TP4 falls past c128
8× H100 vs NVIDIA NVFP4 ckpt (emulated FP4 path) — same rig, 2026-09-28 latency recipe c1 181.2 · aggregate c256 2,418 · best measured c1 193.18 (solo MTP-N2) NVFP4 N2 (closeout mirror, TP8): 122.56 · 2,299.84; N1b card recipe (TP4 × 2EP, KV bf16): 81.87 · 861.95, TTFT p50 263 s @c256 ours leads both headline cells (+47.8% c1, +5.1% c256); c128 dead heat 2,066.8 vs 2,081.13 (NVFP4 +0.7%); ckpt ships no MTP layer — full grid
8× H100 (same rig, identical harness) MTP-N2 TP4: c1 175.1 · c8 219.4 · c32 726.8; TPOT p50 2.99 ms NVFP4 MTP-N2 TP4: 166.5 · 239.2 · 864.8; TPOT 3.28 ms +5.2% c1 and TPOT −9% vs NVFP4; NVFP4 edges c8/c32 in this TP4 pool-limited cell — the TP8 rows above are the answer shape
4× RTX PRO 6000, TP=4, MTP N=2 109.9 · 318.5 · 534.4 NVFP4: 109.3 · 319.5 · 530.8 parity (±0.7%)
4× H200, TP=4, MTP N=2 c1 195 · c8 698 · c32 1,258 · c64 1,529 · c128 1,789 · c256 1,954 — 8× H200 TP=8: 217 → 2,681; two TP=4 replicas: 391 → 3,911 aggregate
8× H100, TP=8, DFlash2-G K=7 (latency recipe) c1 181.2 TP8 spec-off 124.1 +46% single-stream; aggregate recipe = TP8 spec-off c256 2,418; solo-TP4 MTP-N2 c1 193.18 is the best single-stream shape measured on this rig
2× DGX Spark, TP=2, DFlash2, 8K/256, 1M serve c1 33.0 · c2 35.6 · c4 59.6 · c6 67.1 EXL3 (matched protocol): 29.9 · 59.6 · 112.7; LibertAI NVFP4 ckpt (2026-08-31): did not boot (9/9 OOM) +10.4% c1 and +31–39% long-prefill vs EXL3; EXL3 leads mid-concurrency (c2 +67%, c4 +89%)
2× DGX Spark, TP=2, DFlash2-G @800k — H2H final 2026-09-27 tg32 d0 31.35±3.52 (30.88 on 09-25) · c1@65k 15.34 (10.72 on 09-25) · c1@100k 9.97 NVIDIA NVFP4 quant + serving stack w/ incoai DFlash2 drafter (legB): 34.82±1.92 · 29.19 · 17.95 NVFP4 stack ahead on all 28 cells (+16.9%…+254.5%), widest at depth; our one measured win = stability at depth (their run2 HTTP 500; ours clean ×2). Full grid: BENCHMARKS.md

vs NVIDIA's NVFP4 checkpoint — 8× H100, same rig (2026-09-28)

nvidia/GLM-5.3-Flash-NVFP4 is a Blackwell-only release and Hopper has no native FP4 math: on SM90 every weight dequantizes through FP4-emulation kernels (the engine warns at boot) — an emulated path, not a native comparison. Measured on the same p5 rig and closeout harness (8192/1024, thinking ON). The card-verbatim fp8-KV recipe cannot boot on Hopper (scale item not float32 in the KV-quant kernel); the card-recipe arm runs with KV bf16, the one documented deviation. Output tok/s:

runtime c1 c8 c32 c64 c128 c256
W4A16 TP8 + DFlash2-G K=7 (latency recipe) 181.2 633.0 1009.1 1471.4 1645.3 1653.5
W4A16 TP8 spec-off (aggregate recipe) 135.4 651.0 1375.9 1622.3 2066.8 2418.1
W4A16 TP4 MTP-N2 (solo, best single-stream shape) 193.18 559.2 1078.4 1181.9 1258.5 1292.4
NVFP4 N2 (closeout mirror, TP8 spec-off) 122.56 629.32 1052.71 1564.07 2081.13 2299.84
NVFP4 N1b (card recipe TP4 × 2EP, KV bf16) 81.87 185.01 577.28 823.02 845.24 861.95

Our recipes lead NVFP4's best arm (N2) by 47.8% at c1 (181.2 vs 122.56; solo MTP-N2 193.18 = +57.6%) and 5.1% at c256 (2,418.1 vs 2,299.84); c128 is a dead heat (2,066.8 vs 2,081.13, NVFP4 +0.7%). The card-recipe arm (N1b) collapses under load — TP4 KV exhaustion drives TTFT p50 to 263 s @c256 — and the checkpoint ships no MTP layer, so it cannot express GLM's native speculative decode at all. On Hopper, use our quant + drafter + recipe. Boot-failure autopsy and sha256 custody: BENCHMARKS.md.

Single-stream decode is insensitive to KV length up to ≥486K (RTX PRO 6000); at batch, long-KV decode plateaus at ~2–4 tok/s per stream and long prefills serialize at a ~6–8.5K tok/s aggregate ceiling. All grids, protocols and the H200 extended table: BENCHMARKS.md.

Hardware & context limits

Each row is the largest context serving-validated on that configuration.

Configuration GPUs Validated context Stack
DGX Spark GB10 (SM121), TP=2 2× 128 GB UMA 1M — KV pool 1,360,420 tokens (1.30× a full 1M request) DFlash2 drafter, fp8 KV
RTX PRO 6000 (SM120), TP=4 4× 96 GB 512K (486K prompts measured) MTP N=2, fp8 KV
H100 (SM90), TP=4 4× 80 GB 262K (256K prompts measured) MTP N=2, bf16 KV
H100 (SM90), TP=8 + DFlash2-G 8× 80 GB 131,071 with any drafter — 128K-token prefills die on the speculative path (spec-off survives 131,072); with MTP-N2 the 262K row above applies v2 image, K=7 latency / spec-off aggregate
H200 (SM90), TP=4 4× 141 GB 262K — KV pool 6.19M tokens (≈23 concurrent 262K requests) MTP N=2, bf16 KV
H200 (SM90), TP=8 8× 141 GB 262K — KV pool 7.79M tokens MTP N=2, bf16 KV
H200 (SM90), TP=2 2× 141 GB 1M — KV pool 2.84M tokens (929K-token prompt measured) MTP N=2, bf16 KV

Known issues

  • config.json (2026-09-08): vLLM matches quantization_config.ignore against its own fused module names, so the ignore list now carries both the HF and vLLM spellings plus re:.*\.layers\.45\..* for the MTP head. Older 765-entry copies fail at load. Weights unchanged.
  • DFlash2 admission wedge (SM90 research stack only, MTP recipes unaffected): with the DFlash2 drafter at block size 2304, prompts above ~15.5K tokens are never admitted. A fix was validated to 256K prompts; block size 1536 avoids it. Filed as vllm-project/vllm#55800.
  • Marlin no-split-K path on SM121: deterministic illegal memory access at M=256 when forcing split_k=1; the stock heuristic used in serving is clean. Filed as vllm-project/vllm#56064.

Quantization details

Field Value
Architecture Glm5NextForConditionalGeneration (glm5_next) — 45 decoder layers + MTP layer 45, 288 routed experts (top-8) + 1 shared, KDA + DSA attention, 24-block vision tower
Quantized 36,288 tensors = 42 MoE layers × 288 experts × 3 GEMMs — W4A16, INT4, symmetric, group 128, GPTQ, compressed-tensors pack-quantized
Kept in BF16 attention (incl. DSA indexer), dense prefix layers 0–2, shared experts, router, mHC tensors, embeddings, lm_head, norms, vision tower (348 keys), MTP layer 45 (889 keys)
Kept in FP32 A_log, dt_bias, e_score_correction_bias, hc_* — verbatim from source
Calibration 256 samples × 4096 tokens, in-distribution chat/code mix, sequential per-layer GPTQ
Built on 8× NVIDIA B300, 2026-08-27

Build gates, all passing: exactly 36,288 packed tensors and nothing quantized outside routed experts; vision key set 348/348 identical to source; MTP layer present; zero dtype drift vs source; no collapsed expert scales. Loads with transformers ≥ 5.16; text generation and image captioning smoke tests pass.

Citation & license

@misc{canada_quant_glm53_flash_w4a16_mtp,
  title  = {GLM-5.3-Flash W4A16 (INT4) + BF16 MTP},
  author = {canada-quant},
  year   = {2026},
  url    = {https://huggingface.co/canada-quant/GLM-5.3-Flash-W4A16-MTP}
}

MIT, inherited from the base model. Follow the base model's usage terms.


Built, benchmarked and documented with the Digby.ai coding harness, developed by CQL.ca.

Downloads last month
11,366
Safetensors
Model size
50B params
Tensor type
F32
·
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for canada-quant/GLM-5.3-Flash-W4A16-MTP

Quantized
(132)
this model