Configuration Parsing Warning:Invalid JSON for config file config.json

UMM-Reflection BAGEL-RL

The final model of UMM-Reflection (Learning Native Reflection in Unified Models). Starting from UMM-Reflection-BAGEL-SFT, it is trained with multi-round Flow-GRPO. For each prompt, it samples 16 complete reflection trajectories from the same first image and optimizes the text (controller) and flow (image) heads jointly against a graded six-family GenEval reward.

Results

Local evaluation with the same prompts, resolution, and scorer for every row. GenEval uses the 553 official prompts with up to three repair rounds.

Model GenEval WISE OneIG T2I-CompBench++
BAGEL-Base 0.71 0.55 0.80 0.49
BAGEL-SFT 0.72 0.63 0.79 0.50
UMM-Reflection (this model) 0.84 0.74 0.83 0.55

On GenEval-553, 64.9% of images that are wrong after the first render end up correct after reflection, against 20.6% for the SFT model.

Contents

The layout is the same as the base BAGEL repository. ema.safetensors merges the 784 trained language-model tensors and the nine trained auxiliary tensors (time_embedder, vae2llm, llm2vae, latent_pos_embed) of RL checkpoint-1000 into the SFT weights. Every other file is byte-identical to BAGEL-7B-MoT revision 265d1d48ec8e850a29d3a1f208c2a2ec3cd7577b.

File sha256
ema.safetensors 34b1b232d971c24a279ab4863162ddb9155adc757a2f39b5f934a4396d824345

Training

  • Initialization. UMM-Reflection-BAGEL-SFT.
  • Prompts. A 3,000-prompt pool over the six GenEval families, built from Flow-GRPO's GenEval training metadata and disjoint from the GenEval-553 test prompts.
  • Schedule. 1,000 steps on 16 GPUs. Each rank takes 2 prompts with 16 sibling trajectories each, 20 denoising steps, and up to three repair rounds. Text and flow learning rates are both a constant 5e-6.

Usage

git clone https://github.com/waltstephen/UMM-Reflection && cd UMM-Reflection
huggingface-cli download YijiaFan/UMM-Reflection-BAGEL-RL \
    --local-dir pretrained/UMM-Reflection-BAGEL-RL

ARM=rl MODEL_DIR=pretrained/UMM-Reflection-BAGEL-RL bash scripts/eval/geneval553.sh

Evaluation uses 512 x 512 images, 50 denoising steps, CFG 3.0, controller temperature 0.5, and at most three repair rounds. REPAIR_RNG_STEP defaults to 500 for ARM=rl, the convention of the reported numbers. See the repository README for the GenEval verifier assets.

License

Apache 2.0, the same as BAGEL. ae.safetensors is the FLUX.1 autoencoder shipped with BAGEL, under its own Apache 2.0 license.

Downloads last month
188
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YijiaFan/UMM-Reflection-BAGEL-RL

Base model

Qwen/Qwen2.5-7B
Finetuned
(1)
this model

Collection including YijiaFan/UMM-Reflection-BAGEL-RL