minLillemus: "my little mouse"

An experimental weight-surgery variant of SmolLM2-135M-Instruct, made with the HyperTensor CECI "graft" procedure. It is a research artifact: it does not outperform its base model.

What was changed

Only the layer-12 MLP was changed: W12 <- W12 + 0.5 * (W2 - W12) @ P (rank-230 projector from layer 12's q_proj). Created by scripts/graft_proof.py ("Splejsning").

The other 269 of 272 tensors are bit-identical to the base model. Weights are stored in float32. The tokenizer and chat template are the base model's.

Measured quality (audit, 2026-10-09)

Teacher-forced next-token loss on 30 web and 30 code documents (512 tokens each, data/corpus/{web,code}.dev.jsonl), compared with the base model:

text base NLL this model NLL perplexity change top-1 agreement with base KL(base‖model)
web 2.865 2.897 +3.2% 89.7% 0.037
code 1.783 1.846 +6.4% 89.8% 0.068
all 2.324 2.371 +4.8% 89.8% 0.053
original HyperTensor test base this model
graft_proof 15-sentence PPL (mean of per-sentence PPL) 60.96 69.87
graft_benchmark MMLU-lite (50 items) 31/50 33/50 (+2 / -0 items, sign-test p=0.50)
graft_benchmark BoolQ (15 items) 6/15 5/15 (+0 / -1, p=1.00)

The earlier card claimed "29% PPL recovery". The "% recovery" figures compared this model with a copy of the base whose MLP had been zeroed. This model was never ablated, so they do not measure recovery. On real text the model is slightly worse than its base, and the small MMLU-lite/BoolQ differences are a few items out of 50/15 and are not statistically significant.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("NagusameCS/minLillemus")
model = AutoModelForCausalLM.from_pretrained("NagusameCS/minLillemus")
msgs = [{"role": "user", "content": "What is the capital of France?"}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt", return_dict=True)
out = model.generate(**inputs, max_new_tokens=64, do_sample=False)
print(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Ollama

The repo includes minLillemus-Q8_0.gguf and a Modelfile:

hf download NagusameCS/minLillemus --local-dir minLillemus && cd minLillemus
ollama create minlillemus -f Modelfile && ollama run minlillemus
# or directly: ollama run hf.co/NagusameCS/minLillemus

Source: HyperTensor.

Downloads last month
38
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for NagusameCS/minLillemus

Quantized
(143)
this model