Instructions to use raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: llama cli -hf raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: llama cli -hf raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: ./llama-cli -hf raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP # Run inference directly in the terminal: ./build/bin/llama-cli -hf raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP
Use Docker
docker model run hf.co/raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP
- LM Studio
- Jan
- Ollama
How to use raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF with Ollama:
ollama run hf.co/raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP
- Unsloth Desktop
- Pi
How to use raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF with Docker Model Runner:
docker model run hf.co/raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP
- Lemonade
How to use raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP
Run and chat with the model
lemonade run user.Laguna-S-2.1-ROCmFP4-COHERENT-GGUF-Q4_0_ROCMFP
List all available models
lemonade list
- Hermes Agent
How to use raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF:Q4_0_ROCMFP" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Laguna-S-2.1 β ROCmFP4 COHERENT (Strix Halo optimized)
AMD-optimized 4-bit quant of poolside/Laguna-S-2.1 (118B MoE, 8B active, 1M ctx) using the Q4_0_ROCMFP4_COHERENT tensor-protected format from charlie12345/ROCmFPX.
File: Laguna-S-2.1-Q4_0_ROCMFP4_COHERENT.gguf β 58.3 GiB, ~4.3 bpw effective.
Why this quant
- COHERENT recipe protects agent-critical tensors (attention, embeddings, structured-output paths) at higher precision β in our testing it had the best long-context integrity of any Laguna quant we measured (correct refusals instead of confabulation on absent-information probes).
- Corrected metadata:
laguna.rope.scaling.yarn_attn_factoris baked to1.0per poolside's upstream fix ("llama.cpp derives mscale") β upstream-converted GGUFs from before 2026-07-24 carry1.4852, which double-applies YaRN attention scaling. No--override-kvneeded with this file. - Quantized directly from poolside's BF16 with poolside's official imatrix (no intermediate quant).
Requirements
β οΈ Does not load in stock llama.cpp. Requires the ROCmFPX runtime (built/tested at commit c190e435, 2026-07-23):
git clone https://github.com/charlie12345/ROCmFPX
env JOBS=16 scripts/build-strix-rocmfp4-mtp.sh # gfx1151 / Strix Halo
Measured on AMD Ryzen AI MAX+ 395 (Radeon 8060S, Strix Halo, 128GB)
| Metric | Value |
|---|---|
| Decode (tg128) | ~32β35 t/s (HIP), ~37 t/s (Vulkan) |
| Prefill (pp512) | ~405 t/s (HIP) |
| 256k context serving | validated (HIP device; ~194k-token cold prompt in ~12 min) |
| Warm-turn TTFT (49k-token cached prefix) | ~0.3 s prefill (--cache-reuse 256) |
| KV q8_0 decode cost | negligible (<1%) |
Recommended serve command:
HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
llama-server -m Laguna-S-2.1-Q4_0_ROCMFP4_COHERENT.gguf \
-ngl 999 -fa on -dev ROCm0 -c 262144 --jinja \
-b 2048 -ub 1024 --cache-reuse 256 \
--temp 0.7 --top-p 0.95 --top-k 20 --min-p 0
Notes: thinking is on by default (disable per request via chat_template_kwargs: {"enable_thinking": false}); give thinking β₯32k max_tokens or it can exhaust the budget mid-reasoning. Vulkan devices failed very large (>190k-token) single fills in our testing across all runtimes β use the HIP device for extreme contexts.
Provenance & credits
- Base model & imatrix & DFlash draft: poolside (Laguna S 2.1, OpenMDW 1.1 + model terms β review before use)
- ROCmFP4 codebook & ROCmFPX runtime: charlie12345/caf
- Quantization & benchmarking: raulvidis, 2026-07-24
License follows the base model (OpenMDW 1.1 + poolside model terms).
- Downloads last month
- 32
4-bit
Model tree for raulvidis/Laguna-S-2.1-ROCmFP4-COHERENT-GGUF
Base model
poolside/Laguna-S-2.1