ArseneLupin v1.1

ArseneLupin is an open decision model: you give it a state (a document, a ticket, a trace, a JSON record) and typed questions, and it returns a probability distribution for each question, in a single forward pass per option, with no generated text. Three question types:

type answer example
noul yes / no "Does this alert describe unauthorized activity?"
choice one option among several "Which department should handle this ticket?"
score an ordered level "How severe is this incident, from Low to Critical?"

It exposes the same request shape as a System One API: POST /v1/systemone with {"state": ..., "questions": {...}}.

Two versions

version where size runs with
complete model, bfloat16 repository root 9.3 GB transformers 5.17.0
8-bit, GGUF Q8_0 gguf/ 4.5 GB llama.cpp

Both are served by the code in code/ (the arsenelupin package), with the same API.

Model

  • Base: Qwen/Qwen3.5-4B (Apache 2.0), revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a; same architecture, same tokenizer. The checkpoint keeps the structure of its base: 4.66B parameters, of which 4.21B are used for decisions.
  • Fine-tuning: LoRA rank 32, α 64, on all linear layers, merged into the weights.
  • Read-out: a small trained head (2,561 parameters, readout/head.safetensors) turns the model's last hidden state into the probability of yes for each option.
  • Calibration: optional per-workflow temperature files (calibration/).

Results

All figures below were measured by us, with the same inputs for every system; each row gives the question type and the number of options. Accuracy counts the option given the highest probability; when several options tie at the top, they share the point. Higher is better. B/2 (half Brier score against a reference distribution): lower is better.

benchmark ArseneLupin v1.1 Jev 1.13 openjev 26B laya 421M
Multi-family decisions, yes/no, choice and score (173) — proprietary task 93.1 % 93.6 % 92.5 % 74.0 %
Transfer to an untrained decision, choice and yes/no (200) — proprietary task 92.0 % 85.0 % 91.5 % 47.0 %
Median without list, list mode, choice (50) — proprietary task 98.0 % 98.0 % 86.0 % 14.0 %
jabr v2 — public external benchmark, 49 tasks (869), macro accuracy per task 88.4 % 97.1 % 95.3 % 58.5 %
jabr v2 — choice subset, 20 tasks, macro accuracy per task 92.3 % 97.7 % 98.6 % 70.8 %
Banking77 intents, choice among 12 (800) 88.5 % 92.8 % 88.9 % 86.4 %
CLINC150 intents, choice among 12 (800) 97.0 % 98.4 % 95.4 % 98.8 %
Amazon ESCI product relevance, ja, choice among 4 (500) 69.6 % 64.8 % 62.8 % 32.0 %
MASSIVE intents, 4 languages, choice among 2 to 9 (482) 82.0 % 81.7 % 80.5 % 65.8 %
French search intent, human gold, choice among 10 (1,016) 50.7 % 43.9 % 43.8 % 15.5 %
Hate speech, 3-level score, B/2 (300) 0.0658 0.2016 0.1531 0.2511
Security incidents workflow, yes/no and score, B/2 (26), calibrated 0.0844 0.0385 0.1076 0.1449
Agent traces workflow, yes/no, choice and score, B/2 (52), calibrated 0.0780 0.0701 0.1176 0.2164
Customer service workflow, yes/no, choice and score, B/2 (84), calibrated 0.0564 0.0346 0.0712 0.2382

Highlights: ahead of Jev 1.13 on transfer to an untrained decision (92.0 vs 85.0) and on hate speech (B/2 0.0658 vs 0.2016); level with it on the median test.

Decision Index 0.1

The Decision Index is a benchmark and public leaderboard for typed decision engines by Apolinário Passos (multimodalart), with an open reproduction kit (MIT licence). Measured by us with that kit, on the benchmark suite rebuilt from its public sources: 49.05 on the Decision Index 0.1 headline panel (19 benchmarks, five areas weighted equally). This is not an official submission. On the Decision Index 0.1 leaderboard of 22 Sep 2026, Jev is listed at 59.51.

area ArseneLupin v1.1 Jev
Knowledge & Reasoning 42.7 68.9
Language Understanding 53.6 62.3
Retrieval & Classification 33.9 37.0
Tools & Automation 67.0 73.4
Arts & Human Judgment 48.1 56.0

Speed

By default, the server of the complete model reads the text shared by the options of a request once, then scores all the options together (engine.mode: shared_prefix in serve.yaml). Measured on one NVIDIA RTX PRO 6000 (Blackwell, 96 GB), bfloat16, transformers 5.17.0 with flash-linear-attention 0.5.2, against scoring one option at a time:

one option at a time default
time per option, over 390 real Decision Index requests 79 ms 19 ms
time per request of 32 yes/no questions on a short text 1.69 s 0.12 s
time per request of 32 yes/no questions on a text of about 2,500 tokens 3.26 s 0.47 s

On those 390 requests, 99.6 % of the decisions are the same in both modes; the others were near ties, whose probabilities moved by at most 0.03. The results tables above were computed one option at a time: to compute them the same way, set engine.mode: independent and engine.branch_microbatch: 1 in serve.yaml. If the card runs out of memory, lower engine.max_batch_tokens.

Calibration

Each file in calibration/ holds one correction factor (a temperature), fitted on one kind of workflow so that the probabilities better match how often the answer is right on that workflow. It changes the probabilities, never the chosen answer. Pass the file for your workflow with --calibration, or calibration/default.json for any other workflow.

Usage

Install the serving code (it needs transformers 5.17.0, the version the model was trained and measured with), then run the commands from the repository folder: in serve.yaml, the model is ., this folder.

hf download IBOU-SEARCH/ArseneLupin-v1.1 --local-dir arsenelupin-v1.1 && cd arsenelupin-v1.1
pip install ./code "transformers==5.17.0"
# serve; set model.device in serve.yaml to cuda:0, mps (Mac) or cpu
python -m arsenelupin serve --config serve.yaml --calibration calibration/default.json --port 8000
curl -s http://127.0.0.1:8000/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": "Customer: my parcel shows as delivered but it never arrived.",
  "questions": {
    "team": {"type": "choice", "instructions": "Which team should handle this ticket?",
             "criteria": {"billing": "Billing", "shipping": "Shipping and delivery",
                          "tech": "Technical support"}},
    "urgent": {"type": "noul", "instructions": "Is this ticket urgent?"}
  }
}'

Each question comes back with its probabilities and a confidence. GET /health reports the model and calibration in use. Omit --calibration for raw probabilities; use one of the workflow files when you run that workflow. Each option's prompt can hold up to 32,768 tokens (limits.max_branch_tokens in serve.yaml). The other limits sit in the same limits section: by default up to 256 questions and 4,096 options per request, 255 options per choice question and 10 levels per score question. GET /health reports the limits in force.

8-bit version (GGUF)

llama.cpp runs the model; arsenelupin keeps the template, the tokenization, the read-out head, the calibration and the API.

cd gguf
llama-server -m arsenelupin-v1.1-Q8_0.gguf --embedding --pooling last -c 8192 -b 8192 -ub 8192 --port 8081 &
python -m arsenelupin serve --config serve.yaml --calibration calibration/default.json --port 8000

This setting takes prompts of up to 8,192 tokens per option; for longer ones, raise -c, -b and -ub together, and limits.max_branch_tokens in gguf/serve.yaml.

Integrity: SHA256SUMS lists the files of the complete model and gguf/SHA256SUMS those of the 8-bit version; at start-up the server also checks the fingerprints of readout/token_ids.json and readout/head.safetensors against readout/environment.json, which records how the model was built.

Licence

Apache 2.0, like the base model (its licence text is in LICENSE).

Downloads last month
172
Safetensors
Model size
5B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IBOU-SEARCH/ArseneLupin-v1.1

Finetuned
Qwen/Qwen3.5-4B
Quantized
(465)
this model

Collection including IBOU-SEARCH/ArseneLupin-v1.1