infosave commited on
Commit
763a670
Β·
verified Β·
1 Parent(s): bade0a7

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +103 -0
README.md CHANGED
@@ -1,3 +1,106 @@
1
  ---
2
  license: apache-2.0
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: apache-2.0
3
+ base_model: Kwaipilot/KAT-Coder-V2.5-Dev
4
+ tags:
5
+ - cmf
6
+ - moe
7
+ - code
8
+ - expert-pruning
9
+ - rust
10
+ pipeline_tag: text-generation
11
  ---
12
+
13
+ # KAT-Coder-V2.5 β€” CMF coding specialist (expert-defragmented MoE)
14
+
15
+ **A 34.7B-A3B MoE coder in a single 12.7 GB file β€” 35% smaller than the
16
+ full quantized model, Γ—1.8 faster on a 24 GB MacBook, +2.8% code
17
+ perplexity.** This is [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev)
18
+ (Qwen3.6-35B-A3B architecture: 40 layers, 30 of them GatedDeltaNet
19
+ linear attention, 256 routed experts top-8 + a shared expert),
20
+ quantized to 4-bit tiles and then **physically stripped of the experts
21
+ that code generation never routes to**.
22
+
23
+ | file | size | held-out code ppl | decode, M4 24 GB | prefill |
24
+ |---|---:|---:|---:|---:|
25
+ | full q4t model | 19.6 GB | 5.058 | 7.6 tok/s | 6.1 tok/s |
26
+ | **this file** | **12.7 GB** | 5.198 (+2.8%) | **13.7 tok/s (Γ—1.8)** | **20.0 tok/s (Γ—3.3)** |
27
+
28
+ The speedup is not a kernel trick: the full model does not fit a 24 GB
29
+ machine and pages from disk on every token, while the specialist
30
+ resides in memory entirely. On machines with plenty of RAM the two
31
+ decode at similar speed and you simply save the 7 GB.
32
+
33
+ ## How to use
34
+
35
+ [CMF](https://github.com/infosave2007/cmf) is a single-file LLM format
36
+ with a small pure-Rust runtime β€” no torch, no CUDA install, no Python.
37
+ Install the CLI (one command, GPU backends included) and run:
38
+
39
+ ```sh
40
+ cargo install cortiq-cli # or a release binary: github.com/infosave2007/cmf/releases
41
+
42
+ cortiq run KAT-Coder-V2.5-CMF.cmf \
43
+ --prompt "Write a Python function that checks if a number is prime." --max-tokens 300
44
+
45
+ cortiq serve KAT-Coder-V2.5-CMF.cmf --port 8080 # OpenAI-compatible API + dashboard
46
+ cortiq bench KAT-Coder-V2.5-CMF.cmf # measure on your hardware
47
+ ```
48
+
49
+ The tokenizer and chat template are embedded in the file β€” `run` is a
50
+ real chat turn out of the box.
51
+
52
+ GPU: `CMF_GPU=1` enables the GPU path. On discrete Vulkan/DX12 cards
53
+ the entire decode β€” GatedDeltaNet recurrence, attention, the MoE
54
+ router, the on-device top-k expert selection and every selected
55
+ expert β€” executes as **one GPU submit per token** (the full-model
56
+ variant of this pipeline decodes at 32.8 tok/s on an RTX 5090 vs 14.4
57
+ on its 32-core host CPU). On Apple silicon a runtime probe arbitrates
58
+ Metal against CPU per operation and keeps whichever wins. Long
59
+ contexts: `--o1 all` converts the 10 softmax-attention layers into a
60
+ constant-memory streaming operator (KV+state at 4K context: 238 β†’ 83 MB).
61
+
62
+ ## The technology
63
+
64
+ MoE expert usage turns out to be **strongly task-conditional**.
65
+ Measured on the full KAT-Coder: the top-64 expert sets selected for
66
+ code vs for natural-language prose overlap with a Jaccard index of just
67
+ **0.25** β€” near-disjoint working sets β€” and a code-derived expert mask
68
+ captures only ~39% of prose routing mass. A model serving one task
69
+ therefore carries hundreds of experts it never routes to.
70
+
71
+ The pipeline that produced this file (two commands, reproducible with
72
+ [`cortiq` β‰₯ 0.5.27](https://github.com/infosave2007/cmf)):
73
+
74
+ 1. **Record the routing field.** A teacher-forced pass over a
75
+ representative code corpus with `CMF_MOE_STATS=stats.json` records
76
+ per-layer expert-selection frequencies β€” an empirical routing
77
+ field over the expert lattice (the "B-field" of the CMF patent
78
+ family's claim 12).
79
+ 2. **Drop what the task never uses.** `cortiq moe-defrag model.cmf
80
+ --stats stats.json --cover 0.95` keeps, per layer, the smallest
81
+ top set of experts reaching 95% of the recorded routing mass
82
+ (mean 158 of 256 per layer here; 3,906 experts dropped in total),
83
+ renumbers the kept experts into a contiguous prefix, slices the
84
+ router's rows to match, and writes a standard `.cmf`. At
85
+ inference the router's softmax renormalizes over the kept set β€”
86
+ no retraining, weights byte-identical to the quantized originals.
87
+
88
+ The same restriction can be previewed at runtime without rewriting the
89
+ file (`CMF_MOE_MASK=stats.json CMF_MOE_MASK_COVER=0.95`), which is how
90
+ the perplexity gate above was measured before the physical cut β€” the
91
+ two are mathematically identical.
92
+
93
+ **Scope, honestly.** This is a *code* specialist: calibration was
94
+ Rust-heavy source code, and quality outside the calibrated task
95
+ degrades by design (prose routes to experts that are no longer there).
96
+ For general-purpose use, take the full model through
97
+ [the step-by-step guide](https://github.com/infosave2007/cmf/blob/master/docs/KAT_CODER.md)
98
+ β€” it covers GGUF β†’ CMF conversion, Vulkan and Metal, and carving your
99
+ own specialist for any task with your own corpus.
100
+
101
+ ## Provenance
102
+
103
+ - Base model: [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) (Apache-2.0)
104
+ - Quantized source: [bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF](https://huggingface.co/bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF) (Q4_K_M), imported to CMF q4_tiled by `cortiq import-gguf` (pure Rust; every llama.cpp storage convention undone on import)
105
+ - Expert defrag: `cortiq moe-defrag`, cover 0.95, code-calibrated (3,072 tokens of Rust source)
106
+ - Runtime: [github.com/infosave2007/cmf](https://github.com/infosave2007/cmf) (Apache-2.0), `cortiq` β‰₯ 0.5.27