Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
Abstract
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that strengthening the understanding capability of the system, through a stronger multimodal encoder, agentic prompt rewriting, and related techniques, together with improvements in data quality, training pipelines, and agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.
Community
Boogu-Image-0.1 is a strongly competitive Apache-2.0 open-source unified image generation and editing model family, including Base, Turbo, Edit, and other variants that provide stable, practical capabilities for high-quality text-to-image generation, fast generation, image editing, and Chinese-English text rendering, with performance that matches top closed-source models in many scenarios.
impressive!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models (2026)
- Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning (2026)
- TextSculptor: Training and Benchmarking Scene Text Editing (2026)
- UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation (2026)
- UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation (2026)
- Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers (2026)
- EquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
I'm really enjoying the quality of the outputs from these, even the turbo editing model performs well with minimal fiddling, and its comprehension is far better than other models in its size range. I've been able to feed it complex prompts step by step to perform multiple edits in a single pass, and have had surprising success. This family of models is Pheeby-Approved~!
Thank you for choosing the Apache2.0 license, and thank you for what you're doing for the open models community.
Models citing this paper 10
Boogu/Boogu-Image-0.1-Base
Datasets citing this paper 0
No dataset linking this paper