Instructions to use JANGQ-AI/GLM-5.3-Flash-JANGHT2.4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use JANGQ-AI/GLM-5.3-Flash-JANGHT2.4 with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("JANGQ-AI/GLM-5.3-Flash-JANGHT2.4") config = load_config("JANGQ-AI/GLM-5.3-Flash-JANGHT2.4") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use JANGQ-AI/GLM-5.3-Flash-JANGHT2.4 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "JANGQ-AI/GLM-5.3-Flash-JANGHT2.4"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "JANGQ-AI/GLM-5.3-Flash-JANGHT2.4" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use JANGQ-AI/GLM-5.3-Flash-JANGHT2.4 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "JANGQ-AI/GLM-5.3-Flash-JANGHT2.4"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default JANGQ-AI/GLM-5.3-Flash-JANGHT2.4
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use JANGQ-AI/GLM-5.3-Flash-JANGHT2.4 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "JANGQ-AI/GLM-5.3-Flash-JANGHT2.4"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "JANGQ-AI/GLM-5.3-Flash-JANGHT2.4" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Run JANG models in MLX Studio / vMLX
⚠️ Runtime: JANGHT needs a vMLX (Python) build that is not released yet. The bundle was built and validated on a vMLX development build that adds the JANGT trellis experts (decode and prefill kernels) and the bundled DFlash2 drafter. vMLX 1.6.77 and older refuse to load it (they stop with an error; they never produce wrong output). Use it with the first vMLX release that lists JANGHT support. Need the previous JANGH2 weights, which released vMLX builds run? They are in this repo's history at commit
4017296(revision="4017296adf8c15c49c4743e6c646b89a2eb55673").
JANGQ-AI/GLM-5.3-Flash-JANGHT2.4
GLM-5.3-Flash for 128 GB Macs, now 4.3 GiB smaller and closer to the original model than our JANGH2 release, with a speculative-decoding drafter in the box.
A JANGHT bundle of zai-org/GLM-5.3-Flash: a 300B-class MoE (288 routed experts, top-8 + shared expert) with KDA linear attention, sparse attention, and vision + video towers.
- Routed experts: JANGHT, a per-layer mix of two formats, chosen by measurement:
- JANGT: a trellis code at 2, 2.5 or 3 bits per weight (Hadamard-128 rotation, 12-bit state). The 2.5-bit variant alternates 6- and 4-bit transitions, so a layer can sit between 2 and 3 bits.
- JANGH: our Hadamard-32 codebook format at 2, 3 or 4 bits, rounded with GPTQ.
- Everything else: 8-bit (MXFP8 where the source weights sit on an FP8-MX grid, affine 8-bit elsewhere) or full precision. Vision tower: bf16.
- DFlash2 drafter (
dflash2/): z-lab's block-diffusion draft model for GLM-5.3-Flash, unmodified. vMLX uses it for lossless speculative decoding; code and other low-entropy text get much faster. - Name: JANGHT + 2.4 bits per weight averaged over the whole model (91.57 GiB × 8 / 321.3B source parameters), drafter excluded.
This repository was JANGQ-AI/GLM-5.3-Flash-JANGH2 until 2026-10-10. Old links redirect here.
Quality vs the bf16 model
123 held-out sequences, 95,025 teacher-forced positions, scored against the full bf16 model run layer by layer from disk. Top-128 renormalized KL. Every row uses the same 8-bit non-expert weights, so the rows differ only in the routed experts. The set is hard on purpose: 44k of the positions are multi-turn tool conversations, plus long context and Chinese.
| experts | median KL ↓ | mean KL ↓ | p90 / p95 / p99 ↓ | top-1 ↑ | top-5 ↑ | top-10 ↑ | |
|---|---|---|---|---|---|---|---|
| control: bf16 experts | bf16 | 0.016 | 0.306 | — / — / 6.51 | 85.5% | 96.8% | — |
| GLM-5.3-Flash-JANGHT2.4 | JANGHT, 91.57 GiB bundle | 0.251 | 1.391 | 4.31 / 7.21 / 13.32 | 65.7% | 84.9% | 88.6% |
| GLM-5.3-Flash-JANGH2 (previous release, revision 2) | JANGH, 95.89 GiB bundle | 0.297 | 1.574 | 5.02 / 8.06 / 14.29 | 63.7% | 83.5% | 87.2% |
Earlier comparison: official FP8 reference, 20 short prompts (from the JANGH2 release)
15,830 teacher-forced positions, top-128 renormalized KL against the official FP8 release. This is a different, much easier set than the table above, and JANGHT2.4 was not run on it; the rows are kept for the comparison with other public quants.
| Bundle | Size | median KL ↓ | mean KL ↓ | p90 / p95 / p99 ↓ | top-1 ↑ | top-5 ↑ | top-10 ↑ |
|---|---|---|---|---|---|---|---|
| GLM-5.3-Flash-JANGH2 (revision 2) | 95.89 GiB | 0.0301 | 0.376 | 0.98 / 1.97 / 5.24 | 83.9% | 96.8% | 98.3% |
| GLM-5.3-Flash-JANG (affine, earlier release) | 95.35 GiB | 0.0882 | 0.528 | 1.50 / 2.55 / 5.67 | 78.6% | 94.6% | 96.8% |
orcarouter GLM-5.3-Flash-MLX 2bit-lite ¹ |
95.4 GiB | 0.2122 | 0.83 | — | 71.4% | 90.8% | 94.3% |
¹ Same 20 prompts (15,850 positions), loading its shipped quantized weights natively.
On the 95k-position set above, JANGHT2.4 beats JANGH2 on every aggregate (median KL −15%, top-1 +2.1 points).
Earlier comparison: agentic fidelity vs the bf16 model (from the JANGH2 release)
72 held-out tool-use conversations (25,867 positions; tools and phrasings disjoint from calibration), including 114 "call a tool or answer?" decision points, scored against the full bf16 model. JANGHT2.4 was not run on this set.
| JANGH2 (revision 2) | affine JANG (earlier release) | |
|---|---|---|
| median KL vs bf16 | 0.579 | 0.595 |
| top-1 agreement with bf16 | 57.3% | 54.9% |
| tool-call decisions: same choice as bf16 | 50 / 50 | 50 / 50 |
median P(<tool_call>) at call points (bf16: 0.998) |
0.999 | 0.981 |
lowest P(<tool_call>) at a call point |
0.914 | 0.784 |
| answer decisions: same next token as bf16 | 87.5% | 50.0% |
| decisions flipped call ↔ answer vs bf16 | 0 | 3 |
Earlier fix kept in JANGHT2.4: long reasoning ends on its own (JANGH2 revision 2)
JANGH2 revision 1 could not end long reasoning: error-minimizing row scales shrank every low-bit matrix (12% at 2
bits), which suppressed </think> after a few thousand reasoning tokens. Revision 2 corrected the row scales to unit
gain, and JANGHT2.4 applies the same rule to every expert unit.
| where the reference model ends its reasoning | bf16 | JANGH2 revision 2 | revision 1 | affine JANG |
|---|---|---|---|---|
P(</think>), held-out design task, after 8,025 reasoning tokens |
0.82 | 0.75 | 0.03 | 0.02 |
P(</think>), coding task, after 3,288 reasoning tokens |
0.99 | 0.94 | 0.19 | - |
By domain, median KL and top-1 agreement:
| domain | positions | JANGHT2.4 | JANGH2 |
|---|---|---|---|
| tool conversations | 44,214 | 0.498 · 59.2% | 0.661 · 55.9% |
| coding | 9,504 | 0.074 · 73.5% | 0.088 · 72.1% |
| agentic | 7,497 | 0.072 · 76.4% | 0.090 · 74.0% |
| cybersecurity | 7,324 | 0.081 · 73.3% | 0.098 · 72.3% |
| long context | 6,141 | 0.326 · 63.7% | 0.346 · 64.4% |
| general | 5,637 | 0.179 · 68.8% | 0.182 · 68.3% |
| science | 3,891 | 0.085 · 74.4% | 0.084 · 74.4% |
| academic multiple choice | 3,696 | 0.176 · 71.2% | 0.197 · 69.7% |
| Chinese | 3,668 | 0.264 · 64.8% | 0.288 · 63.9% |
| systems | 3,453 | 0.103 · 72.9% | 0.110 · 71.9% |
Tool decisions
Inside the 95k positions: 109 points where the bf16 model's transcript calls a tool, 96 where it answers.
| JANGHT2.4 | JANGH2 | |
|---|---|---|
| tool-call points: same next token as bf16 | 98.2% | 94.5% |
| tool-call points: decisions flipped call ↔ answer | 2 | 6 |
| tool-call points: model starts a tool call (bf16: 108 of 109) | 106 | 102 |
lowest P(<tool_call>) at a tool-call point |
0.467 | 0.182 |
| answer points: same next token as bf16 | 78.1% | 83.3% |
| answer points: false tool calls | 0 | 0 |
Reasoning ends on its own
JANGH revision 1 could not end long reasoning; revision 2 fixed the cause (shrunken low-bit row scales). JANGHT2.4 keeps that fix on every expert unit, JANGT and JANGH alike.
16 held-out transcripts written by the bf16 model in its native format (long design and coding reasoning, multi-step tool conversations), 60,243 assistant positions, scored against bf16:
| JANGHT2.4 | JANGH2 | |
|---|---|---|
| top-1 agreement with bf16, assistant text | 83.8% | 82.6% |
| median KL vs bf16, assistant text | 0.058 | 0.069 |
</think> points where the quant also picks </think> (of 34) |
18 | 14 |
median P(</think>) at those points |
0.94 | 0.85 |
lowest P(</think>) |
0.18 | 0.05 |
Free generation on a long design prompt (served, effort low, three seeds): the model closed its reasoning on its own
in 3 / 3 runs, after 8.5k, 9.7k and 12.4k reasoning tokens.
Live behavior (served, temperature 0, both serving lanes)
| suite | JANGHT2.4, DFlash2 lane (default) | JANGHT2.4, plain lane | JANGH2 |
|---|---|---|---|
| 48-case tool-use eval (required / auto × efforts) | 48 / 48 | 48 / 48 | 48 / 48 |
| 16 behavior probes: math, follow-up, 8 tool calls, tool round trip, efforts low/high/max, image, video | 16 / 16 | 16 / 16 | 16 / 16 |
| multi-turn chat: tool call → result → answer → second call → result → answer | 4 / 4 | 4 / 4 | — |
| reasoning-loop matrix (efforts × sampling) | clean | clean | — |
These suites check behavior; every bundle that works passes them.
Speed (M5 Max 128 GB)
Served, one request at a time; two interleaved runs per bundle (A, B, A, B), a fresh server for each run started after a 3-minute cool-down, each run the median of 3 probes on a never-seen prompt. Same runtime build for both bundles.
| JANGHT2.4 | JANGH2 | |
|---|---|---|
| decode, plain lane (tok/s) | 34.8 / 35.8 | 35.3 / 35.3 |
| prefill, ~5.1k-token prompt (tok/s) | 490 / 527 | 556 / 557 |
| time to first token, ~5.1k-token prompt | 10.4 s / 9.7 s | 9.2 s / 9.2 s |
Decode is on par with JANGH2 in a 4.3 GiB smaller bundle. Long-prompt prefill is ~8% slower (the trellis tiles cost more to decode than the JANGH codebook).
With the bundled DFlash2 drafter (the default lane) decode depends on how predictable the text is: 43.5 tok/s on prose, 65-68 tok/s on code, same conditions. Output is unchanged.
Prompt cache (served, 8.25k-token conversation)
| time to first token | |
|---|---|
| first request (cold) | 16.3 s |
| same request again | 0.30 s |
| conversation continued (new turn appended) | 0.71 s |
| same long prefix, different final question | 0.93 s |
| after a server restart, from the SSD cache (plain lane) | 0.40 s |
The plain lane keeps the cache on SSD across restarts. The DFlash2 lane keeps it in memory (0.32 s / 0.67 s / 1.03 s for the first three rows).
What's in the bundle
Experts (42 MoE layers):
projection JANGT 2 JANGT 2.5 JANGT 3 JANGH 2 JANGH 3 JANGH 4 gate / up 35 layers 4 2 1 — — down 1 19 — 10 8 4 Vision + video: the bf16 vision tower (byte-identical to JANGH2) + the consolidated image/video processor config.
dflash2/: the DFlash2 drafter (2.18 GiB), unmodified, under its own license (CC BY-NC-ND 4.0, seedflash2/README.md). vMLX picks it up automatically and adapts the draft length to how often drafts are accepted; output is identical to plain decoding at temperature 0 and has the same distribution when sampling. Start vMLX with--no-bundled-dflash2to serve without it.No MTP: layer 45 is omitted; its bytes went into expert precision.
Thinking + agentic:
- Thinking is ON by default (the template opens
<think>). - Reasoning efforts are
low/high/max(defaultmax). There is nomedium, and no thinking-off mode: the template renders any other value, and a missing value, as Max. - Budget for thinking: 8-12k tokens at
lowon a large design task, 17-23k atmax. Useloworhighfor interactive work and a generousmax_tokens. clear_thinking=falsepreserves thinking in history.
- Thinking is ON by default (the template opens
Tool calls: GLM's XML dialect (
<tool_call>name<arg_key>…</arg_key><arg_value>…</arg_value></tool_call>), declared astool_parser: glm_xml_args; tool results render as<|observation|>. Hermes-style JSON parsers will not work.Self-describing:
config.jsoncarries the per-modulequantizationmap: 102 JANGT expert projections ("mode": "jangt","kbits"), 24 JANGH ("mode": "jangtq2","bits"), 147 MXFP8, 214 affine 8-bit. Thejangtblock holds the trellis parameters and the per-layer outlier columns; thejangtqblock (version 2) holds the JANGH codebooks.jang_config.jsonrecords calibration, the allocation plan, the drafter and capabilities.
Memory: 91.57 GiB of weights + 2.18 GiB drafter. Fixed-size linear-attention state + compressed-latent KV (~6 KB/token), so long contexts do not balloon memory.
Every shard is alignment-safe (zero-copy memory mapping). 18 shards.
Serving contract
- Sampling:
temperature=1.0, top_p=0.95(vendor defaults), no repetition penalty - EOS:
[154820, 154827, 154829]· context: 1M native - Reasoning:
reasoning_effortchat-template kwarg (low/high/max), defaultmax - JANGH units keep their original on-disk identifiers (
jangtqblock,"mode": "jangtq2",*.tq2_packed/*.tq2_scales). They are not loadable by, and must not be routed to, JANGTQ v1 loaders.
Build details
- Source:
zai-org/GLM-5.3-Flash-BF16@a5b45eb - Calibration: 552k tokens rendered with the GLM chat template (coding, tool conversations, cybersecurity, agentic, general, Chinese, science, academic, long context, systems), referenced to the bf16 model run layer by layer; every tenth calibration sequence and all evaluation sets held out
- Allocation: per projection and layer, the format and width with the lowest measured cost on held-out rows, each layer's error weighted by its measured effect on the output distribution
- Row scales: unit gain along the source row on every expert unit (the JANGH2 revision-2 rule)
- Drafter: z-lab GLM-5.3-Flash DFlash2, included unmodified
Quantized and validated by Jinho Jang — eric@jangq.ai
- Downloads last month
- 491
8-bit
Model tree for JANGQ-AI/GLM-5.3-Flash-JANGHT2.4
Base model
zai-org/GLM-5.3-Flash
