Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
SeaWolf-AIΒ 
posted an update about 8 hours ago
Post
899
πŸ’» Data-center AI, now on a laptop: POCKET-Darwin-180B

We're releasing a 4-bit GGUF build of Darwin-180B-RSI, #1 on seven official Hugging Face leaderboards (self-reported), that runs without a GPU.

πŸ“¦ 360 GB β†’ 111 GB (4-bit GGUF, 4 files)
πŸ–₯️ No GPU: one server CPU (16 threads) at 18.4–21.0 tokens/s
πŸ’» RTX 5060 laptop (8 GB VRAM) + 32 GB RAM: 4.17 tokens/s
🧊 128 GB mini PC: whole model in memory, no GPU needed
🎯 MMLU-Pro, 2,000 questions, paired: original 87.65% = 4-bit 87.65%

How?
Β· Only ~3B of 180B parameters are active per token (10 of 512 experts)
Β· llama.cpp streams just the needed experts from SSD, so 32 GB RAM is enough
Β· Graft quantization: we took the proven Unsloth UD-Q4_K_XL base build and swapped in only the 300 tensors our RSI training changed (300/300 verified)

Under the hood is Model-level Recursive Self-Improvement. The model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces.

Built for teams that can't send data to an external cloud (defense, finance, public sector) to run a top-tier model fully offline.

πŸ“ Article: https://huggingface.co/blog/FINAL-Bench/data-center-ai-now-on-a-laptop-pocket-darwin-180b
πŸ€— Model: FINAL-Bench/POCKET-Darwin-180B-GGUF
🧬 Original: FINAL-Bench/Darwin-180B-RSI

#Darwin #RSI #GGUF #llamacpp #OnDevice #MoE

STOP YOUR STUPID M DASHES AI

Β·

@SeaWolf-AI I am not too sure if 180B is pocket sized. Would it be possible if you could make a 4-14B version using the same method? That would be easier to test on my laptop

The graft checks out, and it changes which comparison is the strong one.

I range-read the first 16 KB of every tensor in both files: POCKET and unsloth's Qwen3.8-Flash-Next UD-Q4_K_XL.
Same 4 shard sizes to the byte, same tensor table.
300 of 1,224 tensors differ. All 300 are Q8_0: attn_qkv, attn_gate, ssm_out, attn_q/k/v/output, and the three shared-expert FFNs.
They hold 3.09 GB of 111.32 GB, so 2.78% of the bytes.
No routed-expert tensor moved.
(Small one: shard 1 is the parent's file, same sha256, so the GGUF metadata still names it "Qwen3.8 Flash Next".)

So nearly all the 4-bit error sits in the 97% you share with the parent build, and the whole RSI delta sits in the 2.78% you kept at 8-bit.
That makes POCKET vs the unsloth parent build, on the same 2,000 questions, a clean R0 to R3 test with the quantized experts held fixed byte for byte.

The BF16 comparison can't carry that load.
Its CI is [-0.95, +1.00]. The R1 to R3 gain it is meant to carry is +1.03 on SuperGPQA, lower bound +0.05.
A true loss of 0.9 points fits inside that interval.
And a paired CI that wide implies roughly 100 of the 2,000 questions flipped between right and wrong across builds. Same accuracy, not same answers.

You already ran the parent build on those 2,000 (4,322 tokens vs 3,694).
What did it score?