Access to this repository is granted by the TrustLLM project on request

Research-only artifact of the TrustLLM project (EU). Access requests are reviewed manually by project administrators; a request does not imply approval.

Log in or Sign Up to review the conditions and access this model content.

TrustLLM v2 8B SFT, 8k context

An 8,192-token-context version of the TrustLLM v2 8B nine-language SFT model, prepared by the TrustLLM project.

Revision note (7 October 2026): exact FP32 router biases. Earlier revisions stored the router's expert_bias in BF16. These 23 small tensors (one per MoE layer, 64 values each) are added to the router scores and alone decide which 8 of the 64 experts process each token. Their values are about 70 to 90, where BF16 can only represent steps of 0.5, so the rounding (by up to 0.243 in this model) changed the expert choice for some tokens. This revision stores the exact FP32 values from the training checkpoint. Every other tensor is byte-identical to the previous revision. The included model code now keeps expert_bias in FP32 even after model.to(torch.bfloat16), and it also runs on Transformers 5.13 and later, where the previous code failed in forward and generate because a masking function renamed one of its arguments.

Measured effect, this revision minus the previous one (same items, paired bootstrap 95% intervals; loss in nats per token on 128 held-out SFT conversations per language, lower is better): loss English −0.012 [−0.024, −0.001], Icelandic −0.009 [−0.018, +0.002], Faroese −0.005 [−0.018, +0.009]; multiple-choice suite (5,532 items) +0.8 points [+0.0, +1.6]; GSM8K test (1,319 items) English +4.3 points [+2.0, +6.7], Icelandic +3.6 [+1.2, +6.1]. Faroese grammar pairs did not change. Correction (8 October 2026): an earlier version of this note also said that 8k passkey retrieval did not change; that was checked at 10% depth only. With the same test, retrieval at 8k context is unchanged at 10% and 90% depth (10/10 each), but at 50% depth it fell from 10/10 (previous revision) to 2/10, 4/10 and 3/10 in three runs of this revision (n = 10 each; the previous revision scored 10/10 again in a rerun the same day). The test is synthetic and small, we have not found the cause, and we have not evaluated long-document tasks. The evaluation tables below were measured on the previous revision and are not updated.

Which revision to use. For long-context retrieval, use the previous revision, which has the BF16-rounded router biases: revision="bf16-bias-2026-10-07" (10/10 at 10%, 50% and 90% depth at 8k in the same test). main (exact FP32 biases) is better on math (GSM8K +4.3 points English, +3.6 Icelandic) and on text prediction loss. The model code stored in that revision fails in generate() on Transformers 5.13 and later. Either use Transformers below 5.13, or load the code from main with code_revision="main":

model = AutoModelForCausalLM.from_pretrained("TrustLLMeu/trustllm-v2-8b-sft-8k", revision="bf16-bias-2026-10-07",
    code_revision="main", trust_remote_code=True, torch_dtype=torch.bfloat16)

Known training issue. After this model was trained, we found a bug in how packed training sequences were masked. Several conversations are packed into each 8,192-token training sequence, and each should attend only to its own tokens; because of the bug, later conversations in a sequence could also attend to the earlier ones. This 8k continuation was trained with the bug, and so was the M3 model it continues from. The bug is fixed for future TrustLLM training runs. We have not measured its effect on this model, and this revision does not change it.

The previous revision is kept under the tag bf16-bias-2026-10-07.

It is TrustLLMeu/trustllm-v2-8b-midtrain-sft (the "M3" model, also main of trustllm-v2-8b-sft since 27 September 2026) after a further 256 training steps at 8,192 tokens. M3 was fine-tuned at 4,096 tokens. It still predicts text well up to 8k, but it lost long-range recall: it fails to retrieve a passkey placed about 7,000 tokens back. This continuation restores that recall.

A separate release, confirmed over three seeds (updated 3 October 2026). Before training we set pass limits for this continuation. This checkpoint (seed 420000), taken alone, missed two of them by small margins (English GSM8K −2.05 points against a limit of −2; English text-prediction loss, upper confidence bound +0.021 nats against +0.02). We then trained the same recipe with two more seeds. Pooled over the three seeds, no check shows a regression: English GSM8K is −0.08 points [−2.3, +2.1], and Icelandic GSM8K improves by +2.4 [+0.4, +4.6]. So the two misses were seed noise. The model is still released under its own name, and the M3 repositories are unchanged. Use M3 if you only need 4k context; use this model if you need recall across 8k.

Lineage:

  1. Base: TrustLLMeu/trustllm-v2-8b, final pretraining step 120,000.
  2. General midtraining by Jiangtao Wang and Andreas Bueff (TrustLLM project): 2,500 steps, about 84B tokens, 8,192-token context. Published as TrustLLMeu/trustllm-v2-8b-midtrain.
  3. Nine-language SFT (M3): 2,048 steps at 4,096 tokens. Published as TrustLLMeu/trustllm-v2-8b-midtrain-sft.
  4. 8k continuation (this model): 256 steps at 8,192 tokens, described below.

Research-only; not for redistribution or commercial use. Access requests are reviewed manually by the TrustLLM project, under the same access conditions as the base model.

Model and training

Same architecture as M3: approximately 7.19 billion total parameters, 24 layers, 64 routed experts with top-8 routing, and one shared expert. The repository contains BF16 weights (except the 23 router expert_bias tensors, which are FP32), the tokenizer, chat template, generation configuration, and the custom OptMoE model code. The configuration sets max_position_embeddings to 8,192. The global attention layers have no position encoding; RoPE is used only in the 128-token sliding-window layers, as in midtraining.

Setting Value
Starting point M3 weights (SFT step 2,048), fresh optimizer state
Context length 8,192 tokens
Global batch size 64 packed sequences (the same tokens per step as M3)
Optimizer DiSCO, learning rate 0.02 (one fifth of M3's peak): 16 warmup steps, constant, then linear decay to 0 over the last 50% of steps
Training steps 256
Hardware 16 H100 GPUs at MareNostrum 5, 29 minutes
Objective Assistant-only loss with conversation packing

Data.

  • 75%: M3's nine-language SFT mix, with the same relative shares (an export of allenai/Dolci-Instruct-SFT and project translations into Icelandic, Faroese, Norwegian Bokmål and Nynorsk, Swedish, Danish, Dutch and German).
  • 25%: long-context retrieval. Each example is one Wikipedia article of 4,500–7,800 tokens (wikimedia/wikipedia, 20231101 dump) in one of the nine languages, followed by a question. The question asks which sentence comes right after (70%) or right before (30%) a quoted sentence. Quoted sentences are spread evenly over the article. Answers are copied from the article. Only the answer is trained.

Training note. The run had loss spikes early in training (for example at steps 4, 20, 48 and 88). Validation loss returned to its starting value by step 128 and stayed flat through the decay. The two replicate seeds read different batches but spiked at the same steps, so the spikes follow the training step, not particular data; their cause is not yet known. We checked the Icelandic output for the damage a late spike caused in an earlier checkpoint (Scandinavian word endings grafted onto Icelandic) and found none beyond noise (table below).

Evaluation

All comparisons are against M3 (same seed), scored the same way. Multiple-choice tasks use option-text log-likelihood under the chat template. Math is greedy step-by-step generation, scored on the last number in the answer.

this model M3
Passkey at 8k context, 10% depth (n = 10) 10/10 0/10
Passkey at 8k context, 50% and 90% depth 10/10, 10/10 10/10, 10/10
Passkey at 4k context, 10/50/90% depth 10/10 each 10/10 each
ARC-Challenge, English (acc_norm, n = 1,172) 0.366 0.359
ARC-Challenge, Icelandic translation (acc_norm) 0.325 0.324
Belebele (acc_norm, 300-item subsets): is / en / da / sv / nb / nl / de 0.340 / 0.343 / 0.317 / 0.370 / 0.350 / 0.317 / 0.353 0.343 / 0.327 / 0.313 / 0.380 / 0.337 / 0.297 / 0.323
WinoGrande, Icelandic (partial scoring, n = 1,088) 0.523 0.536
GSM8K test, English (n = 1,319) 0.355 0.375
GSM8K test, Icelandic translation (n = 1,319) 0.260 0.243
Faroese grammar pairs (accuracy) 0.757 0.749
Icelandic answers: non-words with Scandinavian endings per 1,000 words 0.16 0.08

Text prediction. Change in mean loss per token on held-out SFT conversations, against M3 (paired bootstrap 95% interval, 128 items): English +0.008 [−0.006, +0.021], Faroese −0.005 [−0.028, +0.015], Icelandic −0.009 [−0.039, +0.015] nats. On long English and Icelandic documents, loss on tokens 6k–8k is lower than on tokens 2k–4k, so prediction does not degrade past 4k.

What is and is not established:

  • The continuation restores 8k recall in our passkey test (0/10 → 10/10 at 10% depth). The test is synthetic. We have not evaluated long-document question answering or summarisation.
  • The table shows this checkpoint only. Over three seeds of the same recipe (pooled, paired item-bootstrap 95% intervals against M3): passkey 30/30 at every depth at 8k and 4k; Icelandic GSM8K +2.4 [+0.4, +4.6] (the one change we claim); English GSM8K −0.08 [−2.3, +2.1]; the largest multiple-choice drop, Swedish Belebele, −1.4 [−4.4, +1.6]; English text-prediction loss +0.007 [−0.004, +0.019] nats. None of the other differences is distinguishable from zero.
  • The non-word count rose from 0.08 to 0.16 per 1,000 words for this checkpoint (0.14 and 0.36 for the replicate seeds; preset limit: M3 + 1). Native Icelandic text scores 0.46 on the same heuristic.
  • Four other runs (lower learning rate, longer warmup, a smaller or capped retrieval share, skipping spiking batches) either did not restore recall or regressed more.
  • Absolute scores are modest: this is a short multilingual SFT, not a tuned assistant. No safety evaluation has been performed. Translated training and evaluation data may contain translation artifacts.

Loading notes

Access approval and Hugging Face authentication are required to download the weights. The architecture uses the included custom Transformers code, so loading through Transformers requires trust_remote_code=True. Use the supplied chat template with enable_thinking=False. The generation configuration stops on both the end-of-text and end-of-message token IDs.

Keep expert_bias in FP32. The included code does this when the model is loaded with torch_dtype=torch.bfloat16 and after model.to(torch.bfloat16). If you convert the weights to another format or write your own loader, do not cast expert_bias to BF16 or FP16: it changes which experts are used. In our test with vLLM 0.19.1 through its Transformers backend (model_impl="transformers", trust_remote_code=True, dtype="bfloat16"), the loaded expert_bias was also the exact FP32 value. We tested loading and generation of this revision with Transformers 4.51, 4.53, 4.57, 5.13 and 5.19 (7 October 2026).

Evaluation used Transformers with the included remote code (BF16, eager attention, batch size 1). Batched left-padded generation did not match single-sequence generation in our tests, so use batch size 1 or verify batching. The repository does not claim compatibility with a particular vLLM release.

Downloads last month
6
Safetensors
Model size
7B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TrustLLMeu/trustllm-v2-8b-sft-8k

Datasets used to train TrustLLMeu/trustllm-v2-8b-sft-8k