Laguna-S-2.1 β€” ROCmFP4 COHERENT (Strix Halo optimized)

AMD-optimized 4-bit quant of poolside/Laguna-S-2.1 (118B MoE, 8B active, 1M ctx) using the Q4_0_ROCMFP4_COHERENT tensor-protected format from charlie12345/ROCmFPX.

File: Laguna-S-2.1-Q4_0_ROCMFP4_COHERENT.gguf β€” 58.3 GiB, ~4.3 bpw effective.

Why this quant

  • COHERENT recipe protects agent-critical tensors (attention, embeddings, structured-output paths) at higher precision β€” in our testing it had the best long-context integrity of any Laguna quant we measured (correct refusals instead of confabulation on absent-information probes).
  • Corrected metadata: laguna.rope.scaling.yarn_attn_factor is baked to 1.0 per poolside's upstream fix ("llama.cpp derives mscale") β€” upstream-converted GGUFs from before 2026-07-24 carry 1.4852, which double-applies YaRN attention scaling. No --override-kv needed with this file.
  • Quantized directly from poolside's BF16 with poolside's official imatrix (no intermediate quant).

Requirements

⚠️ Does not load in stock llama.cpp. Requires the ROCmFPX runtime (built/tested at commit c190e435, 2026-07-23):

git clone https://github.com/charlie12345/ROCmFPX
env JOBS=16 scripts/build-strix-rocmfp4-mtp.sh   # gfx1151 / Strix Halo

Measured on AMD Ryzen AI MAX+ 395 (Radeon 8060S, Strix Halo, 128GB)

Metric Value
Decode (tg128) ~32–35 t/s (HIP), ~37 t/s (Vulkan)
Prefill (pp512) ~405 t/s (HIP)
256k context serving validated (HIP device; ~194k-token cold prompt in ~12 min)
Warm-turn TTFT (49k-token cached prefix) ~0.3 s prefill (--cache-reuse 256)
KV q8_0 decode cost negligible (<1%)

Recommended serve command:

HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
llama-server -m Laguna-S-2.1-Q4_0_ROCMFP4_COHERENT.gguf \
  -ngl 999 -fa on -dev ROCm0 -c 262144 --jinja \
  -b 2048 -ub 1024 --cache-reuse 256 \
  --temp 0.7 --top-p 0.95 --top-k 20 --min-p 0

Notes: thinking is on by default (disable per request via chat_template_kwargs: {"enable_thinking": false}); give thinking β‰₯32k max_tokens or it can exhaust the budget mid-reasoning. Vulkan devices failed very large (>190k-token) single fills in our testing across all runtimes β€” use the HIP device for extreme contexts.

Provenance & credits

  • Base model & imatrix & DFlash draft: poolside (Laguna S 2.1, OpenMDW 1.1 + model terms β€” review before use)
  • ROCmFP4 codebook & ROCmFPX runtime: charlie12345/caf
  • Quantization & benchmarking: raulvidis, 2026-07-24

License follows the base model (OpenMDW 1.1 + poolside model terms).

Downloads last month
32
GGUF
Model size
118B params
Architecture
laguna
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF

Quantized
(96)
this model