YuhengSSS/SDRPN-Qwen3.5-4B

Papers: https://arxiv.org/abs/2609.19745 (Vision-RL²) · https://arxiv.org/abs/2509.16944 (SD-RPN)

SD-RPN stage-1 checkpoint: a self-distilled RoI predictor twig (K = 21, T = 3) trained on a frozen Qwen/Qwen3.5-4B. This is the initialisation of the Vision-RL² RL run YuhengSSS/VisionRL2-Qwen3.5-4B.

The backbone weights are unchanged from the base model; only the three attached twig blocks are trained, from self-distilled attention pseudo-labels (no human RoI annotation).

Configuration

item value
backbone Qwen/Qwen3.5-4B (frozen)
twig K = 21, T = 3
stage 1 (SD-RPN self-distilled pseudo-label training)
weights full model directory (backbone + twig), bfloat16 safetensors
next stage region-level RL → YuhengSSS/VisionRL2-Qwen3.5-4B

How to use

These weights need the modeling code in YuHengsss/VisionRL2. They are not loadable for RoI inference through a plain AutoModel / AutoModelForCausalLM call: the RoI gating path (heatmap head, peak-relative gate, connected-component crop, sub-image splice) lives in the repository's model classes and evaluation harness.

git clone https://github.com/YuHengsss/VisionRL2
cd VisionRL2
hf download YuhengSSS/SDRPN-Qwen3.5-4B --local-dir output/sdrpn/qwen3_5-4b-sdrpn-K21T3

# stage 2: region-level RL starting from this checkpoint
PHASE_A_CKPT=output/sdrpn/qwen3_5-4b-sdrpn-K21T3 DATASET_ROOT=datasets \
  bash scripts/train_rl_qwen3_5_4b.sh

# evaluate the stage-1 checkpoint itself (training-aligned protocol)
MODEL=qwen3_5 CHECKPOINT=output/sdrpn/qwen3_5-4b-sdrpn-K21T3 CAP=576 bash scripts/aligned_eval.sh

Expected numbers (sanity check after download)

Training-aligned protocol (rule metrics, no judge), source limit 576 tokens - the SD-RPN (stage 1) row of the Vision-RL² README:

Model V* ZoomBench HR-4K HR-8K MME-RW Lite InfoVQA Avg.
Qwen3.5-4B (base) 66.0 40.5 63.5 56.4 41.0 69.8 56.2
SD-RPN (stage 1), this checkpoint 82.7 55.6 71.1 63.3 48.9 78.1 66.6
Vision-RL² (stage 2) 85.3 61.8 77.4 70.9 51.0 80.5 71.1

Links

Citation

@article{shi2026visionrl2,
  title   = {Region-Level Policy Optimization for Fine-grained MLLM Perception},
  author  = {Shi, Yuheng and Pei, Xiaohuan and Dong, Minjing and Xu, Chang},
  journal = {arXiv preprint arXiv:2609.19745},
  year    = {2026}
}

@inproceedings{shi2026sdrpn,
  title     = {Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception},
  author    = {Shi, Yuheng and Pei, Xiaohuan and Dong, Minjing and Xu, Chang},
  booktitle = {ICLR},
  year      = {2026}
}
Downloads last month
9
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YuhengSSS/SDRPN-Qwen3.5-4B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(788)
this model

Collection including YuhengSSS/SDRPN-Qwen3.5-4B

Papers for YuhengSSS/SDRPN-Qwen3.5-4B