Instructions to use infosave/KAT-Coder-V2.5-CMF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/KAT-Coder-V2.5-CMF with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/KAT-Coder-V2.5-CMF --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq run FILE.cmf --prompt "What is the capital of France?"
cortiq serve FILE.cmf --port 8080 # OpenAI-compatible server
- Notebooks
- Google Colab
- Kaggle
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -1,3 +1,106 @@
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
+
base_model: Kwaipilot/KAT-Coder-V2.5-Dev
|
| 4 |
+
tags:
|
| 5 |
+
- cmf
|
| 6 |
+
- moe
|
| 7 |
+
- code
|
| 8 |
+
- expert-pruning
|
| 9 |
+
- rust
|
| 10 |
+
pipeline_tag: text-generation
|
| 11 |
---
|
| 12 |
+
|
| 13 |
+
# KAT-Coder-V2.5 β CMF coding specialist (expert-defragmented MoE)
|
| 14 |
+
|
| 15 |
+
**A 34.7B-A3B MoE coder in a single 12.7 GB file β 35% smaller than the
|
| 16 |
+
full quantized model, Γ1.8 faster on a 24 GB MacBook, +2.8% code
|
| 17 |
+
perplexity.** This is [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev)
|
| 18 |
+
(Qwen3.6-35B-A3B architecture: 40 layers, 30 of them GatedDeltaNet
|
| 19 |
+
linear attention, 256 routed experts top-8 + a shared expert),
|
| 20 |
+
quantized to 4-bit tiles and then **physically stripped of the experts
|
| 21 |
+
that code generation never routes to**.
|
| 22 |
+
|
| 23 |
+
| file | size | held-out code ppl | decode, M4 24 GB | prefill |
|
| 24 |
+
|---|---:|---:|---:|---:|
|
| 25 |
+
| full q4t model | 19.6 GB | 5.058 | 7.6 tok/s | 6.1 tok/s |
|
| 26 |
+
| **this file** | **12.7 GB** | 5.198 (+2.8%) | **13.7 tok/s (Γ1.8)** | **20.0 tok/s (Γ3.3)** |
|
| 27 |
+
|
| 28 |
+
The speedup is not a kernel trick: the full model does not fit a 24 GB
|
| 29 |
+
machine and pages from disk on every token, while the specialist
|
| 30 |
+
resides in memory entirely. On machines with plenty of RAM the two
|
| 31 |
+
decode at similar speed and you simply save the 7 GB.
|
| 32 |
+
|
| 33 |
+
## How to use
|
| 34 |
+
|
| 35 |
+
[CMF](https://github.com/infosave2007/cmf) is a single-file LLM format
|
| 36 |
+
with a small pure-Rust runtime β no torch, no CUDA install, no Python.
|
| 37 |
+
Install the CLI (one command, GPU backends included) and run:
|
| 38 |
+
|
| 39 |
+
```sh
|
| 40 |
+
cargo install cortiq-cli # or a release binary: github.com/infosave2007/cmf/releases
|
| 41 |
+
|
| 42 |
+
cortiq run KAT-Coder-V2.5-CMF.cmf \
|
| 43 |
+
--prompt "Write a Python function that checks if a number is prime." --max-tokens 300
|
| 44 |
+
|
| 45 |
+
cortiq serve KAT-Coder-V2.5-CMF.cmf --port 8080 # OpenAI-compatible API + dashboard
|
| 46 |
+
cortiq bench KAT-Coder-V2.5-CMF.cmf # measure on your hardware
|
| 47 |
+
```
|
| 48 |
+
|
| 49 |
+
The tokenizer and chat template are embedded in the file β `run` is a
|
| 50 |
+
real chat turn out of the box.
|
| 51 |
+
|
| 52 |
+
GPU: `CMF_GPU=1` enables the GPU path. On discrete Vulkan/DX12 cards
|
| 53 |
+
the entire decode β GatedDeltaNet recurrence, attention, the MoE
|
| 54 |
+
router, the on-device top-k expert selection and every selected
|
| 55 |
+
expert β executes as **one GPU submit per token** (the full-model
|
| 56 |
+
variant of this pipeline decodes at 32.8 tok/s on an RTX 5090 vs 14.4
|
| 57 |
+
on its 32-core host CPU). On Apple silicon a runtime probe arbitrates
|
| 58 |
+
Metal against CPU per operation and keeps whichever wins. Long
|
| 59 |
+
contexts: `--o1 all` converts the 10 softmax-attention layers into a
|
| 60 |
+
constant-memory streaming operator (KV+state at 4K context: 238 β 83 MB).
|
| 61 |
+
|
| 62 |
+
## The technology
|
| 63 |
+
|
| 64 |
+
MoE expert usage turns out to be **strongly task-conditional**.
|
| 65 |
+
Measured on the full KAT-Coder: the top-64 expert sets selected for
|
| 66 |
+
code vs for natural-language prose overlap with a Jaccard index of just
|
| 67 |
+
**0.25** β near-disjoint working sets β and a code-derived expert mask
|
| 68 |
+
captures only ~39% of prose routing mass. A model serving one task
|
| 69 |
+
therefore carries hundreds of experts it never routes to.
|
| 70 |
+
|
| 71 |
+
The pipeline that produced this file (two commands, reproducible with
|
| 72 |
+
[`cortiq` β₯ 0.5.27](https://github.com/infosave2007/cmf)):
|
| 73 |
+
|
| 74 |
+
1. **Record the routing field.** A teacher-forced pass over a
|
| 75 |
+
representative code corpus with `CMF_MOE_STATS=stats.json` records
|
| 76 |
+
per-layer expert-selection frequencies β an empirical routing
|
| 77 |
+
field over the expert lattice (the "B-field" of the CMF patent
|
| 78 |
+
family's claim 12).
|
| 79 |
+
2. **Drop what the task never uses.** `cortiq moe-defrag model.cmf
|
| 80 |
+
--stats stats.json --cover 0.95` keeps, per layer, the smallest
|
| 81 |
+
top set of experts reaching 95% of the recorded routing mass
|
| 82 |
+
(mean 158 of 256 per layer here; 3,906 experts dropped in total),
|
| 83 |
+
renumbers the kept experts into a contiguous prefix, slices the
|
| 84 |
+
router's rows to match, and writes a standard `.cmf`. At
|
| 85 |
+
inference the router's softmax renormalizes over the kept set β
|
| 86 |
+
no retraining, weights byte-identical to the quantized originals.
|
| 87 |
+
|
| 88 |
+
The same restriction can be previewed at runtime without rewriting the
|
| 89 |
+
file (`CMF_MOE_MASK=stats.json CMF_MOE_MASK_COVER=0.95`), which is how
|
| 90 |
+
the perplexity gate above was measured before the physical cut β the
|
| 91 |
+
two are mathematically identical.
|
| 92 |
+
|
| 93 |
+
**Scope, honestly.** This is a *code* specialist: calibration was
|
| 94 |
+
Rust-heavy source code, and quality outside the calibrated task
|
| 95 |
+
degrades by design (prose routes to experts that are no longer there).
|
| 96 |
+
For general-purpose use, take the full model through
|
| 97 |
+
[the step-by-step guide](https://github.com/infosave2007/cmf/blob/master/docs/KAT_CODER.md)
|
| 98 |
+
β it covers GGUF β CMF conversion, Vulkan and Metal, and carving your
|
| 99 |
+
own specialist for any task with your own corpus.
|
| 100 |
+
|
| 101 |
+
## Provenance
|
| 102 |
+
|
| 103 |
+
- Base model: [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) (Apache-2.0)
|
| 104 |
+
- Quantized source: [bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF](https://huggingface.co/bartowski/Kwaipilot_KAT-Coder-V2.5-Dev-GGUF) (Q4_K_M), imported to CMF q4_tiled by `cortiq import-gguf` (pure Rust; every llama.cpp storage convention undone on import)
|
| 105 |
+
- Expert defrag: `cortiq moe-defrag`, cover 0.95, code-calibrated (3,072 tokens of Rust source)
|
| 106 |
+
- Runtime: [github.com/infosave2007/cmf](https://github.com/infosave2007/cmf) (Apache-2.0), `cortiq` β₯ 0.5.27
|