AutoJev-27B

Built with autonomous agents, from research and data generation to training, evaluation, and deployment. The user set the goals and refined the scope; agents executed the work.

AutoJev-27B is a multimodal decision model based on Qwen3.8-27B. It returns probabilities over supplied choices in one forward pass per question, with a TypeSafe-compatible API and browser playground.

Code ยท Model weights

Results

Benchmark accuracy: Qwen3.8-27B, AutoJev-27B and Jev

Model Overall accuracy โ†‘ ECE โ†“ Brier โ†“
Qwen3.8-27B 69.83% 0.06483 0.40834
AutoJev-27B 84.60% 0.04282 0.22027
Jev 82.79% 0.05274 0.25400

Training

One H200 ยท full-weight SFT ยท 73,000 unique training examples ยท 286 updates. The released model is checkpoint 200. Training uses cross-entropy; calibration fits a scalar temperature separately.

Training loss

Run bash configs/train.sh --help for training arguments. The exact curated training corpus is not bundled.

Run

Python 3.12+, uv and a GPU with space for approximately 49 GiB of BF16 weights plus runtime overhead.

git clone https://github.com/denis-pplx/autojev.git
cd autojev
uv sync --frozen --python 3.12
uv run hf download denis-pplx/autojev-27b --local-dir checkpoints/selected
AUTOJEV_CHECKPOINT=checkpoints/selected uv run autojev-serve

Open http://localhost:8000 for the playground or /docs for the API. POST /v1/systemone supports choice, noul, score, and optional base64 images. Set AUTOJEV_API_KEY to enable authentication. While the weights are private, authenticate with uv run hf auth login before downloading.

Use the included DecisionModel loader or server. Published benchmarks measure text decisions; image support is not a natural-image accuracy claim.

Code: MIT ยท Weights: Apache 2.0 ยท Independent implementation inspired by Jev.

Downloads last month
905
Safetensors
Model size
26B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for denis-pplx/autojev-27b

Base model

Qwen/Qwen3.8-27B
Finetuned
(427)
this model
Finetunes
1 model
Quantizations
1 model

Spaces using denis-pplx/autojev-27b 2