AI & ML interests
Open RL Environments at Scale
Recent Activity
๐ค FineEnvs: Open RL Environments
FineEnvs is a home for end-to-end RL environment recipes, built to make it easier to explore, reproduce, train, and evaluate agent systems.
Explore complete and reproducible environment projects from us and the community, including:
- ๐ Open RL environments
- ๐งฉ End-to-end environment recipes
- ๐ป Complete implementations
- ๐ฆ Models, datasets, and artifacts
- ๐งช Training and evaluation setups
- ๐ Demos and Spaces
- ๐ Tutorials and guides
All the reproducible code โ environments, rollouts, training configs, notebooks, article and slide sources โ lives in one repo: github.com/adithya-s-k/FineEnvs. The artifacts those produce live here on the Hub.
FineEnvs Projects
A growing collection of open projects, environments, resources, and artifacts.
| Project | What it is | Explore |
|---|---|---|
| FineEnvs Academy | Articles, guides, tutorials, slides, and hands-on resources for learning how to build RL environments and agent systems. | Explore โ |
| Data Agent | Training SLMs for data science with multi-harness RL environments. | Explore โ |
| MiMo-V2.6-RL in Harbor | All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, plus an explorer to browse them and run graded rollouts. | Explore โ ยท Explorer โ |
| Repo2RLEnv | Verifiable coding and terminal RL environments in Harbor format, with per-task quality labels and provenance. | Explore โ |
Each numbered project below is a self-contained recipe: an environment, a training run, and every artifact it produced.
| # | Project | What it is | Explore |
|---|---|---|---|
| 00 | RL Environments 101 | Three environments implemented six times over, one per framework. Same logic, six dialects. | Source โ ยท Collection โ |
| 01 | LaTeX OCR | Qwen3-VL-2B trained to read rendered math into LaTeX, scored by a reward served from a live Space. | Collection โ |
| 02 | Watercolour | Qwen3.5-35B-A3B trained to paint watercolours by writing p5.brush sketches, rewarded by taste rather than correctness. | Collection โ |
| 03 | GeoGuesser | A multi-turn visual geolocation environment, and the 4B trained on it until it outscored gpt-5.4-mini and claude-haiku-4.5. |
Collection โ |
| 04 | SmolDataEnvs | 5.5K+ data-analysis tasks for hill-climbing small models, graded deterministically with no LLM judge: plain prompts, verified SFT traces, and Harbor task suites. | Collection โ |
| 05 | SmolDataEnvs: Multi-harness RL | One small model trained with GRPO inside four unmodified coding agents (OpenCode, Claude Code, Codex, Mini-SWE-Agent), with LFM2.5-2.6B and Qwen3.5-2B checkpoints. | Article โ ยท Collection โ |
| 06 | Multilingual | Two OpenEnv servers: a million document pages in 22 languages with Sarvam Indic OCR Bench, and read speech in all 102 FLEURS languages. Gemma 4 trained on each to read and to hear Kannada. | Collection โ |
Articles & Talks
| What it covers | Read / Watch | |
|---|---|---|
| ๐ The Ultimate Guide to RL Environments | Building and scaling RL environments in the LLM era โ how frameworks are built, how rewards are wired, how they scale to thousands of concurrent sessions. | Read โ |
| ๐๏ธ RL Environments 101 | From "what is an env?" to training your own: RL fundamentals โ environment anatomy โ OpenEnv โ training with TRL. | Watch โ |
| ๐ Scaling RL for LLMs | RL environments and RL training โ what an environment is, how reward hacking happens, how to train against your own. AMD AI Dev Day. | Watch โ |
| ๐ Multi-Harness Training | OpenEnv ร Harbor โ why an environment's failure model decides whether it can be trained against. | Watch โ |
| ๐งญ The Ultimate Guide to Multi-Harness RL | Training small models on SmolDataEnvs with the same tasks and reward but a different tool loop each time (TRL, native OpenCode, Harbor), and what changes. | Read โ |
| ๐ค Training a Coding Agent Through a Harness You Did Not Write | Multi-harness RL talk by Sergio Paniego Blanco: one model, four unmodified coding agents, GRPO. | Watch โ |
| ๐ How to turn a game into an RL environment | The technical intuition, end to end: curating the data, designing the environment, shipping it with OpenEnv, and training a 4B against it with TRL. | Read โ |
Environments
Three reference environments, each implemented across six frameworks โ openenv, ors, nemo_gym, verifiers, skyrl_gym, gem. Same logic, six dialects. Source โ ยท RL Envs 101 collection โ
| Environment | Tools | OpenEnv | ORS | NeMo Gym |
|---|---|---|---|---|
| Jupyter agent โ real code execution in an E2B sandbox | 4 | Space | Space | Space |
| Wordle โ multi-turn, pure Python, no backend | 1 | Space | Space | Space |
| Desktop โ computer-use, vision-driven Linux desktop | 19 | Space | Space | โ |
Build your own
Five agent skills turn a plain-English description into a runnable RL environment across four frameworks โ works with Claude Code, Cursor, Codex, OpenCode, Gemini CLI and others.
npx skills add adithya-s-k/FineEnvs
We're looking for new end-to-end recipes โ a task, an environment, a training run, and honest results. Contributing guide โ
Citation
@misc{fineenvs,
author = {Kolavi, Adithya S},
title = {FineEnvs: Open Source RL Environments for LLM Agents},
year = {2026},
url = {https://github.com/adithya-s-k/FineEnvs}
}
-
Nayana Multilingual OCR
๐OpenEnv document OCR, layout and VQA over 22 languages
-
FLEURS Multilingual ASR
๐1OpenEnv speech recognition over 102 FLEURS languages
-
FineEnvs/gemma-4-E4B-it-kannada-ocr-grpo
Image-Text-to-Text โข Updated โข 20 -
FineEnvs/gemma-4-E4B-it-kannada-asr-grpo
Automatic Speech Recognition โข Updated โข 17
-
Nayana Multilingual OCR
๐OpenEnv document OCR, layout and VQA over 22 languages
-
FLEURS Multilingual ASR
๐1OpenEnv speech recognition over 102 FLEURS languages
-
FineEnvs/gemma-4-E4B-it-kannada-ocr-grpo
Image-Text-to-Text โข Updated โข 20 -
FineEnvs/gemma-4-E4B-it-kannada-asr-grpo
Automatic Speech Recognition โข Updated โข 17
spaces 34
Simulation RL Environments
Turning real-world work into RL environments: PortSimEnv v1
PortSimEnv-Eval
PortSimEnv v1 eval, every model rollout in 3D
PortSimEnv v1
Dock planning on real Port of Barcelona calls
DesignGym
Recreate real graphic designs via MCP, HTML or a real editor
Nayana Multilingual OCR
OpenEnv document OCR, layout and VQA over 22 languages
