Instructions to use autotrust/JEV-35B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use autotrust/JEV-35B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="autotrust/JEV-35B")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("autotrust/JEV-35B") model = AutoModelForMultimodalLM.from_pretrained("autotrust/JEV-35B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
autotrust/JEV-35B
🔴 Live: AutoTrust/GuruSearch recommendation & search demo → news.guru.so
Try AutoTrust/GuruSearch live (new, 11 October 2026). A recommendation and search demo in which every ranking is a calibrated System 1 decision, made in about half a second:
- Recommendation: the latest headlines from 6 news feeds, ranked by importance.
- Search: results from Google News, Bing News and Yahoo News re-ranked by relevance, side by side with the search engines' own order, plus a short answer with citations.
A System 1 decision model on Qwen3.5-35B-A3B (35 B parameters, about 3 B active per token), up to 256 options in one pass
Decision Index 0.3, public suite: 59.73 (our scoring with the kit), against 53.64 for autotrust/JEV-27B-VL on the board. Approximate public vision score 71.94 (our rebuild of the vision benchmarks; JEV-27B-VL 71.67 on the same rebuild, a tie within noise). Median single-request latency 204 ms on one B200 (JEV-27B-VL 271 ms).
NVFP4 version: autotrust/JEV-35B-NVFP4, the same model with its routed experts in NVFP4: 25.6 GB download, 23 GiB of GPU memory (bf16: 72 GB, 66 GiB), so it runs on one 32–48 GB GPU. Decision Index 0.3 public 59.18 (bf16 59.73), vision rebuild 71.23 (bf16 71.94), 96.6 % of the Decision Index answers identical to bf16, same latency, same
serve.shand API.
| version | download | GPU memory (weights) | Decision Index 0.3 public | vision (rebuild) |
|---|---|---|---|---|
| autotrust/JEV-35B (bf16, this repo) | 72 GB | 66.5 GiB | 59.73 | 71.94 |
| autotrust/JEV-35B-NVFP4 | 25.6 GB | 23.3 GiB | 59.18 | 71.23 |
Two models, two organisations. TypeSafe Jev 1.13 is the hosted, closed model made by TypeSafe AI. autotrust/JEV-35B is an independent open-weights model built by AutoTrust AI; it is not affiliated with, endorsed by, or a product of TypeSafe AI.
Decision Index 0.3 (public suite)
| public index | |
|---|---|
| autotrust/JEV-35B, System 1 (our run with the kit) | 59.73 |
| autotrust/JEV-27B-VL, System 1 (board, public part) | 53.64 |
| area (skill) | Knowledge & Reasoning | Language | Retrieval & Classification | Tools & Automation | Arts & Taste |
|---|---|---|---|---|---|
| JEV-35B | 0.426 | 0.632 | 0.710 | 0.764 | 0.422 |
| JEV-27B-VL | 0.414 | 0.564 | 0.537 | 0.735 | 0.416 |
JEV-35B: all 140,178 scoreable requests of the 0.3 public suite answered (0 errors), System 1 only (thinking off),
every choice read in one pass (up to 256 options), scored with the kit's score --edition 0.3. Our scoring, not a board
entry: the board's Full score also counts private tests (80 %), which only its maintainers run. JEV-27B-VL: the board's
public numbers for JEV-27B, whose text decisions JEV-27B-VL reproduces (see the JEV-27B-VL card).
JEV-35B is ahead in all five areas and on 24 of 37 benchmarks. Largest gains (skill points): PhishNChips +42.2, HoVer +28.4, Habermas Machine +28.1, VAST +22.7, WinoGrande +18.9, BANKING77 +18.3, iSarcasmEval +17.2, When2Call +14.9, GPQA +9.5. JEV-27B-VL is ahead on POP909-CL (+25.8), GSM8K (+18.2), ANLI (+11.4), BPoMP (+10.9), NLI4CT (+6.8) and BBH (+4.7).
Every benchmark
Skill rescales the benchmark's own metric so that chance is 0 (below chance counts as 0); ★ = gold benchmark (weight 1.2); the higher score is in bold.
Knowledge & Reasoning (area skill 0.426; JEV-27B-VL 0.414)
| benchmark | JEV-35B | JEV-27B-VL |
|---|---|---|
| GPQA Diamond ★ | 0.360 | 0.265 |
| GSM8K (0.3 rebuild) | 0.360 | 0.542 |
| ChessBench | 0.127 | 0.091 |
| MuSR | 0.361 | 0.402 |
| SATA-Bench | 0.284 | 0.328 |
| CRUXEval | 0.546 | 0.583 |
| CLadder | 0.441 | 0.410 |
| HLE ★ | 0.000 | 0.000 |
| MMLU-Pro ★ | 0.642 | 0.567 |
| BBH ★ | 0.647 | 0.695 |
| WinoGrande ★ | 0.841 | 0.651 |
Language Understanding (area skill 0.632; JEV-27B-VL 0.564)
| benchmark | JEV-35B | JEV-27B-VL |
|---|---|---|
| ContractNLI | 0.732 | 0.639 |
| ANLI ★ | 0.545 | 0.658 |
| HellaSwag ★ | 0.962 | 0.900 |
| ACOS | 0.276 | 0.169 |
| FinEntity | 0.855 | 0.755 |
| iSarcasmEval | 0.426 | 0.254 |
| VAST | 0.549 | 0.322 |
| NLI4CT | 0.626 | 0.694 |
| RAGTruth | 0.665 | 0.604 |
Retrieval & Classification (area skill 0.710; JEV-27B-VL 0.537)
| benchmark | JEV-35B | JEV-27B-VL |
|---|---|---|
| BANKING77 ★ | 0.921 | 0.739 |
| CLINC150+OOS ★ | 0.943 | 0.845 |
| BRIGHT ★ | 0.421 | 0.417 |
| Amazon ESCI | 0.520 | 0.430 |
| PhishNChips | 0.660 | 0.238 |
| HoVer | 0.762 | 0.478 |
Tools & Automation (area skill 0.764; JEV-27B-VL 0.735)
| benchmark | JEV-35B | JEV-27B-VL |
|---|---|---|
| BFCL ★ | 0.950 | 0.946 |
| ToolRet | 0.600 | 0.604 |
| API-Bank ★ | 0.856 | 0.825 |
| Home appliance simulator | 0.500 | 0.523 |
| When2Call | 0.862 | 0.714 |
Arts & Human Taste (area skill 0.422; JEV-27B-VL 0.416)
| benchmark | JEV-35B | JEV-27B-VL |
|---|---|---|
| BPoMP | 0.750 | 0.859 |
| Humicroedit | 0.233 | 0.223 |
| POP909-CL | 0.132 | 0.390 |
| cfcolor | 0.256 | 0.283 |
| Habermas Machine | 0.416 | 0.135 |
| New Yorker | 0.744 | 0.607 |
Images (approximate public vision score)
The vision benchmarks of the Decision Index are not published; this is our rebuild of their public datasets (CV-Bench, BLINK, RealWorldQA, CharXiv, InfographicVQA, Mind2Web, CORD + FUNSD, Hateful Memes, R-Bench-M, MMMU-Pro vision; Winoground not included), scored with the board's chance correction and weights. The same rebuild gives JEV-27B-VL 71.67 against its board score of 71.53 on the same benchmarks.
| approximate public vision score | accuracy | ECE | |
|---|---|---|---|
| JEV-35B | 71.94 | 78.5 % | 0.050 |
| JEV-27B-VL (same rebuild) | 71.67 | 78.0 % | 0.044 |
| benchmark | JEV-35B skill | JEV-27B-VL skill |
|---|---|---|
| CV-Bench | 76.0 | 76.9 |
| BLINK | 55.8 | 56.7 |
| RealWorldQA | 68.3 | 72.2 |
| CharXiv | 79.7 | 80.7 |
| InfographicVQA | 95.0 | 95.9 |
| Mind2Web | 83.0 | 83.2 |
| KIE (CORD+FUNSD) | 98.6 | 98.6 |
| Moderation (Hateful Memes) | 48.4 | 38.0 |
| R-Bench-M | 28.1 | 30.3 |
| MMMU-Pro vision | 42.3 | 39.2 |
Overall the two models are level: the difference (+0.3) is inside the noise (paired bootstrap 95% interval −0.9 to +1.4). The only statistically significant per-benchmark difference (paired McNemar test, p < 0.01) is moderation, in favour of JEV-35B (+10.4); MMMU-Pro (+3.1, p = 0.03) also leans to JEV-35B. The small deficits on RealWorldQA, R-Bench-M, CharXiv, InfographicVQA, CV-Bench and BLINK (−1 to −4) are not significant. The vision tower is Qwen3.5-35B-A3B's; image decisions are zero-shot.
Speed
One sequential client, the same 587 rows (387 text rows sampled across the Decision Index, 200 image rows), POST
/v1/decide on serve.sh (vLLM), one B200:
| median (p95) | text | image | all |
|---|---|---|---|
| JEV-35B | 184 ms (345) | 346 ms (744) | 204 ms (635) |
| JEV-27B-VL (same benchmark) | 207 ms (617) | 457 ms (910) | 271 ms (746) |
Computer use, robot arm and games
Same demo code, seeds, scenes and opponents as the JEV-27B-VL card; every
step is one System 1 decision (POST /v1/decide, thinking off), one model on one B200.
Computer use: screenshot → which element to click. A real browser (headless Chromium). Every clickable element gets a numbered box; System 1 picks the next click (or "the task is complete"), the browser clicks it, and the loop repeats.
Robot arm: pick and place from a camera image (MuJoCo). At every step System 1 answers two questions from the top camera: is the target left or right of the gripper, and above or below it? The arm halves its step whenever an answer flips.
| JEV-35B | JEV-27B-VL | |
|---|---|---|
| Computer use, numbered boxes + element text (60 tasks: shop, settings, mail) | 95% | 95% |
| Computer use, numbered boxes only (60 tasks) | 38% | 10% |
| ms per click decision, median (6 browsers in parallel) | 397 | 720 |
| Robot arm, binary-decision servo (20 scenes): pick-and-place success | 75% | 75% |
| Robot arm: median distance from the cube centre when grasping | 2.5 cm | 2.7 cm |
| Robot arm: placed in the tray, once grasped | 15/15 | 15/15 |
| Robot arm: ms per decision, median | 239 | 239 |
| Robot arm, direct choice among 8 motor actions (10 scenes) | 0% | 0% |
- Computer use with element text: 95%, the same three failures as JEV-27B-VL (shop seeds 12, 15, 20: the colour swatch carries no text and is skipped).
- Numbered boxes only (every element read from pixels): 38% against 10%. JEV-35B completes 90% of the settings tasks but only 15% of mail and 10% of shop; most failures still declare the task complete too early (24 of 37).
- Robot arm: same success rate, different scenes. Each model misses 5 of 20 grasps (both miss scenes 4 and 19), every miss 3 cm or more off the cube centre; once grasped, every cube reaches the tray.
- Choosing directly among 8 motor commands fails for both models; decompose control into simple visual questions.
Games (seed 0 for both models; board as image + text):
| game | JEV-35B | JEV-27B-VL |
|---|---|---|
| 2048: score / largest tile | 336 / 32 | 2,080 / 128 |
| Connect Four against a rule-based opponent (6 games) | 0 wins, 6 losses | 0 wins, 6 losses |
| Flappy Bird: pipes passed | 0 | 3 |
| Snake: food eaten | 11 | 21 |
| Quick, Draw! top-1 among 16 (320 sketches; 30% / 60% / 100% of strokes) | 35.0 / 57.5 / 86.3% | 41.6 / 62.2 / 87.8% |
| Chess mate in one, choice among 16 moves (200 Lichess puzzles) | 47.0% | 52.0% |
For game play, use JEV-27B-VL. Over five seeds JEV-35B's largest 2048 tile averages 64 (random play: 102), it passes
no Flappy Bird pipe and eats 3–18 pieces of food in Snake (mean 10). Per-episode results: reports/demos/.
Quick start (vLLM)
hf download autotrust/JEV-35B --local-dir JEV-35B
bash JEV-35B/serve.sh # vLLM on :8000; one GPU with 80 GB or more
On a smaller GPU, use autotrust/JEV-35B-NVFP4 (one GPU with 32 GB or
more; same serve.sh and API):
hf download autotrust/JEV-35B-NVFP4 --local-dir JEV-35B-NVFP4
bash JEV-35B-NVFP4/serve.sh
serve.sh runs serve_decide.py: the standard vLLM OpenAI server (same flags as vllm serve) with a POST /v1/decide
route. The decision model is merged into the weights, so vLLM loads one plain checkpoint, with no LoRA. Tested with a
vLLM development build from September 2026.
System 1: POST /v1/decide
curl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
"kind": "choice",
"state": "Customer message: my card was charged twice for the same order.",
"question": "Which team should handle this ticket?",
"options": ["billing", "shipping", "technical support", "account security"]}'
| field | value |
|---|---|
kind |
noul: yes/no, probabilities for ["false", "true"] · score: 0–5 · choice: your options |
state |
what the decision is about: a string, a JSON object, or a list mixing text and images |
question |
one question about the state |
options |
choice only: 2–256 strings |
The response has options, probabilities, choice, choice_index and usage. GET /v1/decide/info lists the
defaults.
If you call vLLM's /v1/completions yourself instead of /v1/decide, pass top_k: 0 and top_p: 1.0 so that the
returned probabilities are not truncated, then add the head bias and apply the temperature as below. JEV-35B is a decision
model only: it does not chat or generate text.
Engine and read-out
System 1 reads the hidden state at the last token of a bare-text prompt:
[kind] choice
[state] ...
[question] ...
[options]
A) ...
B) ...
[decision]:
A 264-slot decision head gives one logit per slot; the active slots of the question's kind are soft-maxed with a
per-kind temperature (calibration.json). The head is merged into the lm_head rows of the slot tokens (A–JU,
false/true, 0–5); serve_decide.py adds the head bias (adapter_vllm/decision_head.json) and the temperature.
Limitations
- A System 1 decision model only: no chat, no text generation, no thinking mode.
- Knowledge & Reasoning is the weak area (GPQA 0.360, MMLU-Pro 0.642, HLE below chance); JEV-27B-VL is better on GSM8K, ANLI and BBH.
- Weaker than JEV-27B-VL at sequential game play (2048, Flappy Bird, Snake).
- The vision score is our rebuild, not the board's; image decisions are zero-shot.
- English-centric; not for high-stakes decisions without confidence gating.
Files
model.safetensors-* · config.json · tokenizer* · chat_template.jinja · generation_config.json · *_config.json
the decision model: Qwen3.5-35B-A3B architecture, decision head in the lm_head rows of the slot tokens
judge_config.json slot layout, verbalizer ids, read-out
calibration.json per-kind temperatures
adapter_vllm/decision_head.json
head bias and slot layout, applied by serve_decide.py
serve_decide.py · serve.sh
vLLM server with POST /v1/decide
videos/ computer-use and robot-arm episodes
reports/demos/ per-episode computer-use, robot-arm and game results
NVFP4 weights (routed experts in NVFP4): autotrust/JEV-35B-NVFP4.
License
Apache-2.0. A fine-tune of Qwen/Qwen3.5-35B-A3B (Apache-2.0).
- Downloads last month
- 360
Model tree for autotrust/JEV-35B
Evaluation results
- Decision Index 0.3 public index on Decision Index 0.3 public suiteself-reported59.730
