KVAE: Family of Tokenizers for Multimodal Generative Models Paper • 2608.05798 • Published Aug 6 • 32
Kandinsky WM 1.0 Collection Image-to-Video models for Physical AI: autonomous driving, robotics, general physics. • 3 items • Updated 6 days ago • 7
Decision models Collection GGUF decision models for the /v1/systemone API in llama.cpp • 10 items • Updated 2 days ago • 27
PII-TRACE: A Benchmark for Context-Aware PII Detection in Multi-Turn LLM Conversations Paper • 2609.22200 • Published Aug 31 • 1
pplx-embed-v2 Collection Collection of our second set of text embedding models • 3 items • Updated 4 days ago • 19
Clef Collection Clef is a family of open-weight decision models trained by Cloudflare, optimized for speed, quality and a new vision encoder. • 2 items • Updated 6 days ago • 5
LightOn-rerank Collection LightOn-rerank models are unified cross-encoder rerankers: Qwen3.5 LoRAs scoring both text passages and document page images against a query. • 6 items • Updated about 3 hours ago • 8
mDenseOn & mLateOn Collection A collection of open state-of-the-art single and multi-vector models and datasets for multilingual, long context and code retrieval • 19 items • Updated about 3 hours ago • 21
LightOnOCR-3 🦉 Collection High-Performance OCR and Layout Extraction in One Model • 4 items • Updated about 3 hours ago • 13
NeuCodec Collection We introduce NeuCodec, a 0.8kbps audio codec that outputs audio at 24kHz. • 5 items • Updated Jun 2 • 9
NeuTTS Air Collection NeuTTS Air is a speech foundation model that runs on CPU in real-time, with instant voice cloning. • 3 items • Updated Jun 2 • 22
NeuTTS Nano Multilingual Collection Collection NeuTTS Nano is a TTS model, 3x smaller than NeuTTS Air, that runs on CPU in real-time - now in English, Spanish, French, and German versions! • 13 items • Updated Jul 21 • 22
NeuTTS-2E Collection NeuTTS-2E is a super-fast, highly realistic TTS model with rich emotional control - happy, sad, angry, disgusted, surprised, fearful, and neutral! • 4 items • Updated Jul 22 • 8
EgoSuite-Open100K Collection The largest fully-annotated open egocentric human dataset. 100,000 hours across 15,000+ tasks and scenes. • 3 items • Updated Aug 19 • 122
Qwen3.8-3.6-27B Blend Collection Junie Local's unofficial JetBrains 50/50 blend of Qwen3.6-27B and Qwen3.8-27B: BF16, GGUF, and MLX weights. • 5 items • Updated 13 days ago • 9
Tfree-HAT-7b-pretrained Collection Tokenizer free models based on Hierarchical Autoregressive Transformer (https://arxiv.org/abs/2501.10322) trained from scratch. • 2 items • Updated Aug 1, 2025 • 12