Qwen3-8B teacher-regularized RL method collection for the ToolUse dataset. Includes GRPO-TR, RLSD-TR, SDPO-TR, and SRPO-TR models.
-
SeongryongJung/Qwen3-8B-Tooluse-GRPO-TR
Text Generation • 8B • Updated • 50 -
SeongryongJung/Qwen3-8B-Tooluse-RLSD-TR
Text Generation • 8B • Updated • 222 -
SeongryongJung/Qwen3-8B-ToolUse-SDPO-TR
Reinforcement Learning • 8B • Updated • 18 -
SeongryongJung/Qwen3-8B-ToolUse-SRPO-TR
Reinforcement Learning • 8B • Updated • 24 • 1