An Empirical Study of Path Feasibility Queries
Paper โข 1302.4798 โข Published
None defined yet.
curl -sLO https://huggingface.co/spaces/bench-labs/BenchLabs-Leaderboard/resolve/main/script.py
pip install torch transformers
python script.py --model your/modelacc, acc_norm, soft_score_norm) โ and reports every category and subcategory, not just one number.leaderboard.json: a ready-to-paste models.json entry. Run the script, paste it, open a PR on the [leaderboard]( bench-labs/BenchLabs-Leaderboard). Done. leaderboard = """
| Model | Params | Score |
|----------------------------------|--------------|-----------|
| Qwen/Qwen2.5-Math-1.5B | 1.54B | 82.08% |
| WhirlwindAI/Arithmetic-SLM | 31.70M | 78.60% | <=
| Qwen/Qwen2.5-3B | 3.09B | 78.44% |
| Qwen/Qwen2.5-1.5B | 1.54B | 77.72% |
| Qwen/Qwen2.5-Coder-1.5B | 1.54B | 74.88% |
| HuggingFaceTB/SmolLM2-1.7B | 1.71B | 66.12% |
| Qwen/Qwen2.5-0.5B | 494M | 63.04% |
| facebook/MobileLLM-R1-140M-base | 140M | 53.88% |
| SupraLabs/Supra-50M-Base | 52M | 27.12% |
"""