HeliosLM β€” A Hackable DeepSeek-V3/K3-Style LLM Stack in Pure PyTorch

A from-scratch PyTorch reference implementation of a modern LLM stack: MLA attention with weight absorption, sigmoid-gated MoE with auxiliary-loss-free load balancing, hybrid linear attention, speculative decoding, FP8 training, a DualPipe schedule simulation, a vLLM-style serving engine, a verifiable agent layer (strict tool schema, bitwise-replay oracle), DSA sparse attention, Mooncake-style prefill/decode disaggregation, and tool-tuned checkpoints. Built to be read, modified, and verified β€” every core path is unit-tested and many are checked with bitwise-equivalence tests. Everything runs on CPU.

One-liner: If you want to understand (or hack on) how DeepSeek-V3/K3-class models actually work β€” without needing a GPU cluster first β€” this repo is for you.

Also: the open-source reference for the RLCD "decision layer" paradigm β€” calibrated confidence on acting models, verified across three scales and on the 7,193-instance RLCDAlignBench. Working paper: docs/paper_draft_2026-10-03.tex (repo | HF).

CI Tests License PyTorch

Demo

Who is this for?

You are... What HeliosLM gives you
A learner who wants to understand MLA, MoE routing, DualPipe, speculative decoding Annotated, review-hardened PyTorch with 115+ tests (T11-T28) that act as executable documentation
A researcher who wants a stack to modify, ablate, and extend quickly Single-process, CPU-iterable training + serving code β€” change one file, run one test
A practitioner evaluating serving/quantization techniques vLLM-style paged engine, GPTQ/AWQ/FP8/MXFP4 quantization, MTP speculative decoding β€” all inspectable

Honest positioning: this is a correctness-focused reference implementation, not a throughput-optimized production engine (see Known Limitations).

Paper

Calibrated Agency: An Open-Source RLCD Stack, from Toy Scale to 360M, with a Canonical-Benchmark Comparison β€” Chien-Hsin Lin, working draft 2026-10-03.

πŸ“„ Read the PDF (7pp, figures included) Β· LaTeX source (arXiv-ready) Β· markdown Β· HF copy. Headline numbers: overconfidence on wrong answers grows with scale (0.94 -> 0.9999 -> ~1.0 at 8.5M/360M/Qwen3-0.6B); deterministic grounding cures a 360M model's copy-shaped agentic failure (0/12 -> 9/12 + 3 abstained); trust is f(state, intervention) with an 8.4x measured contrast; on RLCDAlignBench our open readout reaches 0.726 median AUROC and beats the commercial Jev detector's zero-shot numbers on 11/41 benchmarks. Every claim traces to a versioned artifact in benchmarks/.

Citation

@article{lin2026calibrated,
  title={Calibrated Agency: An Open-Source {RLCD} Stack,
         from Toy Scale to 360M, with a Canonical-Benchmark Comparison},
  author={Lin, Chien-Hsin},
  year={2026},
  note={Working draft. arXiv ID: [pending endorsement]},
  url={https://github.com/tonythetiger168/helioslm}
}

Every numeric claim traces to a versioned artifact: code and tests in this repository (CHANGELOG v5.30.2--v5.38j), weights/tokenizer/ fingerprints and benchmark data in huggingface.co/chienhsinlin/helioslm, and the paper source + PDF in docs/.

Deployment tiers

Three rungs, each with one recorded reason to exist (no spectrum theater):

Tier Params Vocab Weights (bf16) Deployment RAM Why it exists
lite 8.5M 1,024 char-level ~17 MB <500 MB, CPU, millisecond latency Protocol/audit research at $0: agent, chat, decision layer, replay verification
mid (v5.33) 360M 32,768 BPE (HeliosBPE, T31) ~720 MB ~1.5 GB, CPU-runnable inference Scale validation: separate toy artifacts from scale-invariant findings (acceptance oracles T32; GPU training run pending)
full DeepSeek-V3/Kimi-K3-class spec 160,000 fp8 still needs hundreds of GB HBM H200/B200-class servers Architecture decision reference (MLA + MoE + FP8 + MTP), not a local target

Any intermediate size is constructible by explicit field overrides -- caller-provided non-None values always win over presets and are validated in __post_init__ (no silent truncation).

Quick Start

pip install torch
python -m helioslm_v5.tests.test_v5     # 48 unit tests
python integration_test_v51.py          # 9 end-to-end integration tests
python -m helioslm_v5.tests.test_agent  # 9 agent-layer oracles (v5.23)
python -m helioslm_v5.tests.test_v5_stage_a   # T15 real-model oracles (v5.26)
python -m helioslm_v5.tests.test_v5_stage_b   # T17 tool-tuned end-to-end (v5.27)
python -m helioslm_v5.tests.test_disagg_pareto  # T18 Pareto sweep oracles (v5.28)
python helioslm_v5/tests/test_chat.py            # T20 chat capability oracles (v5.30)
python helioslm_v5/tests/test_mhc.py             # T22 mHC oracles (v5.31)
python helioslm_v5/tests/test_kv_compress.py     # T23 compressed-attention causality (v5.31)
python helioslm_v5/tests/test_file_env.py        # T24 long-horizon env (v5.31)
python helioslm_v5/tests/test_think.py           # T25 think + experience reuse (v5.31)
python helioslm_v5/tests/test_async_grpo.py      # T26 async GRPO sync-parity (v5.31)
python helioslm_v5/tests/test_decision.py        # T27 typed decision primitives (v5.32)
python helioslm_v5/tests/test_decision_head.py   # T28 DecisionHead calibration (v5.32)
from helioslm_v5.configs.config_v5 import HeliosLMv5Config
from helioslm_v5.src.model_v5 import HeliosLMv5

model = HeliosLMv5(HeliosLMv5Config(size="lite"))   # CPU-friendly
out = model.generate([[1, 2, 3]], max_new_tokens=20, temperature=0)
print(out)

Capabilities at a glance

Area Implementation
Attention MLA with weight absorption β€” latent-only KV cache, βˆ’97.7% memory vs MHA (full config), verified equivalent to the expanded path (<1e-4). Hybrid linear attention: Gated Delta Rule layers interleaved with MLA, fixed-size recurrent state cache, decode ≑ one-shot (<2e-7). RoPE scaling: linear / NTK / YaRN. DSA-style sparse top-k decode over the latent cache (kβ‰₯L exactly dense). Sliding-window attention with StreamingLLM sinks (O(W) decode), per-head QK-norm, Gemma-style logit soft-capping.
MoE Sigmoid-gated fine-grained experts with auxiliary-loss-free load balancing (selection-only bias, heuristic or quantile updates). LatentMoE: routed experts in a shared latent space. SiTU-GLU tanh soft-capped activation.
Cross-layer Attention Residuals β€” per-layer gated injection of accumulated lower-layer attention outputs, threaded through DualPipe (gradient-exact, bitwise-verified).
Agent v5.23 agent layer: strict tool-call schema/parser (16 error classes), deterministic sandboxed tools, trajectory bitwise-replay oracle, ground-truth-by-construction toy envs, routing gates (v5.22 decision-audit discipline); v5.26 real-model oracles (T15); v5.27 tool-tuned checkpoint trained on agent-loop replays (data format == inference by construction) β€” T17 end-to-end baseline parse 0.15 / finish 1/9, sparse_top_k=4 ~= dense; hardened by real-model findings (TOOL_ERROR recovery, ASCII-safe docs) ; v5.29 three-modes benchmark with strict-monotone tau routing curves and a recorded overconfidence finding (max conf 0.925-0.944 on wrong answers), artifact oracles T19
Chat v5.30: dual-mode protocol (plain text OR @@tool@@ block; text bypasses the gate, gate governs tools only), multi-turn ChatSession with transcript replay, chat SFT data with inference-identical prompts; v5.30.2 fixed a dataset filter bias that had hidden the text channel (mode-choice: text 11/20) β€” T20/T21
Frontier references v5.31: five readable toy-scale references distilled from the 2026 frontier β€” manifold-constrained hyper-connections (DS-V4), HCA/CSA-style compressed attention with exact causality, long-horizon file env (8-14 steps), think-mode + experience reuse (Qwen3-Max direction), async GRPO with bitwise sync-parity (GLM-5 direction) β€” T22-T26
Decision layer v5.32: typed decision primitives Choice/Score/Noul (System One / Jev direction) with schema enforcement, gate ask/ask_batch extension point, non-autoregressive DecisionHead answering K questions in one pass, outcome-targeted Brier calibration (RLCD direction; the label-targeted variant is provably redundant with CE β€” recorded) β€” T27/T28 ; v5.36 continuous Noul (P(yes), floor retired); v5.34 deterministic grounding (0/12 -> 9/12 + 3 abstained on a real 360M checkpoint, zero hallucination leakage); v5.37 TrustGate calibrated abstention + TherapyPair composition (trust is f(state, intervention) β€” 8.4x measured contrast); v5.38 RLCDAlignBench: our supervised TF-IDF readout median 0.726 AUROC, beats the commercial Jev detector's zero-shot numbers on 11/41 benchmarks (charts in benchmarks/charts/) β€” T27-T39
Benchmarks Top-10 LLM position paper with verified leaderboard data and the calibration/replay axes no vendor publishes (helioslm_v5/docs/benchmark_top5_2026-09-27.md); probe suite (toy + file-env + chat) with scripted oracles, runnable against any API
Disaggregation v5.25 Mooncake-style prefill/decode module behind a monotonicity gate; v5.28 three-axis Pareto sweep (makespan / workers / worker-seconds) with latency-cost curves per workload β€” cache-aware anti-monotonicity recorded as a structural finding, not hidden
Speculative decoding DeepSeek-style MTP with strict verification (residual (pβˆ’q)β‚Š resampling), batch support, O(1) cache-truncation rollback; hybrid recurrent-state rollback via restore+replay.
Serving vLLM-style engine: paged KV accounting, copy-on-write forks, watermark-aligned continuous batching.
Training FP8 trainer (native float8 + STE, E5M2 gradient hooks, AdamW master weights), DualPipe schedule simulation (recompute-based, gradient-exact), GRPO (real sampling, k3 KL, answer-extraction rewards), Muon optimizer (Newton–Schulz orthogonalized momentum, optional per-head blocks), QAT straight-through fake-quant training.
Adaptation AdaptationLoop: transferred harnesses must be re-accepted under the target workload's gate or dropped; cold/prefix-free targets shed draft+pool, matching the v5.16 break-even data
Harness evolution inference/harness_evolver.py: ModularRSI-style module-wise search over draft/pool/tier configs; the temp-0 bitwise gate makes 'latency evolves, answers never change' an enforced invariant (1.41x modeled speedup, 0 gate rejects)
Prefix pool inference/prefix_pool.py: blake2b(token-block + config fingerprint) keyed KV snapshots, LRU, exact-length past; pooled greedy == from-scratch bitwise
Stream scoring eval/score_stream.py: score any engine's (prompt, output) JSONL under a reference model; A/B compare with bootstrap CI β€” the audit-side complement to serving engines
Spec telemetry bench_spec_breakeven.py: draft x cache-state sweep in the colibri-P3 schema; acceptance + expert hit-rate per decode context, best_draft_per_cache_state() picker
Expert streaming expert_store.py: routed experts tiered to a memory-mapped file, LRU residency with hit/miss/eviction telemetry; streaming forward is bitwise-identical to dense (oracle-verified, roadmap #6)
Toy checkpoints checkpoints/toy_v5.13.pt β€” 8.5M char-level model trained on the repo's own source in ~10 CPU-minutes (examples/train_toy_checkpoint.py); checkpoints/tool_tuned_v5.27.pt β€” tool-tuned on agent-loop replays (examples/train_tool_tuned.py, periodic save + resume); generate() / harness / MTP / agent loop run against trained weights
Quantization True GPTQ (Hessian OBS with error compensation, optional act-order), AWQ with activation-aware grid search, native FP8, MXFP4 β€” all with from_linear real-weight packing.
Eval Log-likelihood harness (helioslm_v5/eval/harness.py): loglikelihood / multiple_choice / run_harness + built-in synthetic tasks (v5.11), token-id based, lm-eval-harness spirit
Multimodal NaViT vision encoder (row/col position decomposition, mixed-resolution packing), streaming audio encoder (causal, sliding-window memory, bit-equivalent to one-shot).

Why HeliosLM vs. alternatives?

HeliosLM transformers vLLM nanoGPT-style
Purpose Understand + hack the full stack Run pre-trained models Max serving throughput Learn basics
Runs on CPU end-to-end βœ… partial ❌ βœ…
Training + serving + quantization in one repo βœ… ❌ ❌ ❌
Bitwise/strict correctness checks on core paths βœ… β€” β€” β€”
Production throughput ❌ (by design) β€” βœ… ❌

Benchmarks

KV cache

The full config holds 99.2% less KV-cache memory than MHA at 128k context (2.0 GB vs 257.7 GB): 36 of 48 layers are Gated-Delta linear attention with a fixed ~1 MB recurrent state, and the 12 MLA layers store only the compressed latent (512 + 64 values/token). Full numbers and methodology: docs/BENCHMARKS.md β€” reproducible on CPU via python benchmarks/bench_cpu.py.

v5.28 serving Pareto (benchmarks/disagg_pareto_2026-09-25.json, regenerate via python examples/disagg_pareto.py): latency-cost curves for cache-heavy / cold / mixed workloads β€” the cost-axis alignment artifact, see docs/benchmark_alignment.md.

Repository layout

helioslm_v5/          # source (configs, src/{attention,moe,inference,training,vision,audio,quantization}, agent/, tests)
examples/             # train_toy_checkpoint.py, train_tool_tuned.py (v5.27), disagg_pareto.py (v5.28)
benchmarks/           # CPU bench results + disagg_pareto_2026-09-25.json (v5.28 artifact)
docs/                 # code review reports, benchmark_alignment.md, k3_alignment_targets(.md/.csv),
                      # competitive_intel_2026-09-25.md (+ raw claims CSV), helioslm_handoff.md
integration_test_v51.py
CHANGELOG.md          # full version history (v5.0 β†’ v5.28)

Roadmap

See the GitHub Project board for the live plan. Highlights:

  • Pre-trained toy checkpoints (v5.13 text, v5.27 tool-tuned) β€” load and generate() / agent-loop immediately
  • Epoch-2 + scaled tool-tuning (T17 parse 0.15 β†’ target 0.4+; trainer resume-ready)
  • Fused quantization kernels
  • CUDA end-to-end verification (paths are currently static-checked; CPU-verified)
  • Hugging Face Hub: chienhsinlin/helioslm-agent hosts the agent layer + tool-tuned artifacts
  • Example notebooks: "Train a tiny HeliosLM on your laptop" / "Add a new attention variant in 30 lines"

Contributing

Contributions are very welcome β€” see CONTRIBUTING.md. Issues labeled good first issue are the best entry points.

Known Limitations

All five previously known limitations are resolved as of v5.14. Remaining hardware-dependent item: CUDA end-to-end verification (CPU-verified paths are static-checked) β€” tracked as a community issue.

Version history

Headlines (full details in CHANGELOG.md):

  • v5.29 β€” Three-mode benchmark on the real checkpoint (direct/routed/oracle): tau-routing curve strictly monotone (v5.22 gate PASS on real confidence); headline finding = systematic overconfidence on wrong answers (conf 0.944) β€” the exact failure class the v5.22 audit toolkit measures
  • v5.28 β€” Cost-axis alignment: disagg three-axis Pareto sweep (latency-cost curves per workload; cache-aware anti-monotonicity recorded as structural finding)
  • v5.27 β€” Tool-tuned checkpoint: agent-loop-replay training, T17 end-to-end (parse 0.15/finish 1/9 baseline, regression-guard floors), sparse_top_k=4 ~= dense in agent inference; agent hardened (TOOL_ERROR recovery, ASCII-safe docs)
  • v5.26 β€” Stage A real-model oracles (T15): zero-gate attention residuals bitwise-verified on HeliosLMv5; sparse top-k decode oracle (kβ‰₯L bit-identical, selection validity + determinism); agent-loop smoke on the real toy checkpoint
  • v5.25 β€” Disagg evolver module: Mooncake-style prefill/decode separation as a HarnessEvolver search module, monotonicity gate, three-axis Pareto (makespan / workers / worker-seconds)
  • v5.24 β€” Attention variants with two-layer oracles: DSA sparse decode (fp32 certificate β‡’ fp64 gate), AttnRes mixing (zero-init β‡’ bitwise migration gate)
  • v5.23 β€” Agent layer: strict tool-call schema + parser, deterministic sandboxed tools, trajectory bitwise-replay oracle, ground-truth-by-construction envs, routing gates in the agent loop
  • v5.22 β€” Decision-layer audit toolkit: calibration metrics (ECE / Brier) against constructed ground truth β€” no reference LLM required
  • v5.20 β€” Pareto-aware integration (memory axis) + cross-workload adaptation loop β€” ModularRSI gaps 2/3 closed at the inference layer
  • v5.19 β€” Evolvable serving harness: ModularRSI-style module-wise search (draft/pool/tier) behind a deterministic bitwise oracle gate
  • v5.18 β€” Content-addressed KV prefix pool: cross-session prefix reuse, fingerprint-guarded, bit-exact oracle (roadmap #6 complete)
  • v5.17 β€” Standalone token-stream scorer: quality-gate any engine's output (A/B compare + bootstrap CI, colibri-container-style)
  • v5.16 β€” Speculation break-even sweep: colibri-P3-compatible JSONL schema, first honest data point (draft pays only when warm)
  • v5.15 β€” Disk-tier expert store: bit-exact streaming oracle, mmap + LRU, router/shared stay resident (roadmap #6)
  • v5.14 β€” Multi-process DualPipe (one stage per process, phased queue protocol, gradient-exact vs single-process)
  • v5.13 β€” CPU-trained toy char-level checkpoint (MTP aux loss, acceptance 1.00 on greedy) + train_toy_checkpoint.py
  • v5.12 β€” Batched equal-length prefill in the engine (hybrid recurrent-state models included)
  • v5.11 β€” Fused dequantΓ—matmul kernels (AWQ/GPTQ/MXFP4), MXFP4 decode fix, eval loglikelihood harness
  • v5.10 β€” CPU benchmark suite: analytic KV-cache accounting + wall-clock generation, BENCHMARKS.md
  • v5.9 β€” Gemma-style attention + final logit soft-capping, per-head QK-norm, sliding-window attention with StreamingLLM sinks
  • v5.8 β€” YaRN RoPE scaling, DSA-style sparse top-k decode, per-head Muon, GPTQ act-order
  • v5.7 β€” RoPE scaling (linear/NTK), FP8 latent KV cache, Hyper-Connections, QAT training
  • v5.6 β€” Hybrid packed-sequence training, MXFP4, Muon optimizer
  • v5.5 β€” K3-aligned feature wave: hybrid Gated-Delta attention, LatentMoE, quantile balancing, attention residuals, SiTU-GLU
  • v5.0–v5.4 β€” Review-hardened core: MLA weight absorption, true GPTQ, strict speculative sampling, NaViT + audio encoders

License

Apache License 2.0 β€” see LICENSE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support