This GGUF does not run with mainline llama.cpp or LM Studio. Build the AutoJev-enabled llama.cpp fork to serve /v1/systemone.

AutoJev-27B GGUF

Model creator: denis-pplx
Original model: autojev-27b
GGUF conversion: AutoJev-enabled llama.cpp fork (requires support for the AutoJev classifier head and /v1/systemone).

This repository contains Q4_K_M quants autojev-Q4_K_M.gguf, and the vision projector mmproj-autojev-bf16.gguf. The projector is needed for image requests; text requests do not require it.

Clone and build the fork (CPU; other backends):

git clone https://github.com/thomasgauthier/jev.cpp.git
cd jev.cpp
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release --target llama-server -j 8

If using CUDA, configure and build with GPU support instead (from jev.cpp):

cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON
cmake --build build --config Release --target llama-server -j 8

Download the GGUF files from this repository, then run (replace /path/to/ with their location):

./build/bin/llama-server -m /path/to/autojev-Q4_K_M.gguf --system-one
# For image requests:
./build/bin/llama-server -m /path/to/autojev-Q4_K_M.gguf --mmproj /path/to/mmproj-autojev-bf16.gguf --system-one

For a CUDA build, add --n-gpu-layers 99 to either launch command to offload model layers to the GPU.

--system-one enables POST /v1/systemone. For the request format, see the server documentation.

The original model is licensed under Apache 2.0. See its model card for training and evaluation details; the reported benchmarks are for the original model, not a GGUF quantization.

Downloads last month
3,105
GGUF
Model size
26B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thomasgauthier/autojev-27b-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(1)
this model