leejunhyeok commited on
Commit
4829c8d
·
verified ·
1 Parent(s): 4545720

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +17 -17
README.md CHANGED
@@ -78,28 +78,28 @@ For contextual comparison, Motif 3 is compared with strong open-weight models us
78
  (*: public dataset only)
79
  <div align="center">
80
 
81
- | Benchmark | **Motif 3**<br><sup>314B-A13B</sup> | MiniMax-3<br><sup>428B-A23B</sup> | GLM-5.1<br><sup>744B-A40B</sup> | Kimi-K2.6<br><sup>1T-A32B</sup> | Qwen-3.5<br><sup>397B-A17B</sup> | DS-v4-Pro<br><sup>1.6T-A49B</sup> |
82
  |:---|:---:|:---:|:---:|:---:|:---:|:---:|
83
  | **Agentic** | | | | | | |
84
- | GDPVal v2 | 38.7 | 44.4 | 37.8 | 34.4 | 23.0 | 40.2 |
85
- | τ²-Bench Telecom | 94.7 | 88.9 | 97.7 | 95.9 | 95.6 | 96.2 |
86
- | τ³-Banking | 35.3 | 15.3 | 13.6 | 23.3 | 13.4 | 30.1 |
87
- | ITBench* | 51.5 | — | 40.3 | 31.2 | 34.1 | 38.3 |
88
  | **Coding** | | | | | | |
89
- | SWE-Bench Verified | 76.2 | 75.0 | 76.4 | 76.2 | 80.0 | 77.4 |
90
- | Terminal-Bench 2.1 | 74.9 | 65.2 | 61.8 | 65.9 | 51.3 | 64.0 |
91
- | SciCode | 40.6 | 45.4 | 43.8 | 53.5 | 42.0 | 50.0 |
92
  | **Reasoning & Knowledge** | | | | | | |
93
- | IMOAnswerBench | 83.2 | — | 83.8 | 81.8 | 80.9 | 89.8 |
94
- | Apex-Shortlist | 75.5 | — | 71.1 | 77.4 | 61.4 | 85.8 |
95
- | GPQA Diamond | 83.4 | 92.9 | 86.8 | 91.1 | 89.3 | 88.8 |
96
- | HLE | 37.0 | 39.0 | 30.1 | 37.5 | 29.0 | 37.5 |
97
- | CritPt | 6.6 | 3.7 | 4.6 | 8.0 | 1.7 | 12.9 |
98
- | OmniScience — Accuracy | 30.1 | 16.7 | 23.7 | 32.6 | 30.8 | 42.9 |
99
- | OmniScience — Non-Hallucination | 71.6 | 81.6 | 70.1 | 59.5 | 11.1 | 5.9 |
100
  | **Long Context & Instruction Following** | | | | | | |
101
- | AA-LCR | 72.3 | 80.3 | 68.0 | 76.7 | 72.7 | 70.0 |
102
- | IFBench | 78.2 | 82.9 | 76.3 | 76.0 | 78.8 | 76.5 |
103
 
104
  </div>
105
 
 
78
  (*: public dataset only)
79
  <div align="center">
80
 
81
+ | Benchmark | **Motif 3**<br><sup>314B-A13B</sup> | MiniMax-3<br><sup>428B-A23B</sup> | GLM-5.1<br><sup>744B-A40B</sup> | Kimi-K2.6<br><sup>1T-A32B</sup> | Qwen-3.7<br><sup>max<sup> | DS-v4-Pro<br><sup>1.6T-A49B</sup> |
82
  |:---|:---:|:---:|:---:|:---:|:---:|:---:|
83
  | **Agentic** | | | | | | |
84
+ | GDPVal v2 | 38.7 | 44.4 | 37.8 | 34.4 | 39.0 | 40.2 |
85
+ | τ²-Bench Telecom | 94.7 | 88.9 | 97.7 | 95.9 | 94.7 | 96.2 |
86
+ | τ³-Banking | 35.3 | 15.3 | 13.6 | 23.3 | 12.0 | 30.1 |
87
+ | ITBench* | 51.5 | — | 40.3 | 31.2 | 42.5 | 38.3 |
88
  | **Coding** | | | | | | |
89
+ | SWE-Bench Verified | 76.2 | 75.0 | 76.4 | 76.2 | 80.4 | 77.4 |
90
+ | Terminal-Bench 2.1 | 74.9 | 65.2 | 61.8 | 65.9 | 75.0 | 64.0 |
91
+ | SciCode | 40.6 | 45.4 | 43.8 | 53.5 | 53.5 | 50.0 |
92
  | **Reasoning & Knowledge** | | | | | | |
93
+ | IMOAnswerBench | 83.2 | — | 83.8 | 81.8 | 90.0 | 89.8 |
94
+ | Apex-Shortlist | 75.5 | — | 71.1 | 77.4 | 44.5 | 85.8 |
95
+ | GPQA Diamond | 83.4 | 92.9 | 86.8 | 91.1 | 92.4 | 88.8 |
96
+ | HLE | 37.0 | 39.0 | 30.1 | 37.5 | 41.4 | 37.5 |
97
+ | CritPt | 6.6 | 3.7 | 4.6 | 8.0 | 11.4 | 12.9 |
98
+ | OmniScience — Accuracy | 30.1 | 16.7 | 23.7 | 32.6 | 31.0 | 42.9 |
99
+ | OmniScience — Non-Hallucination | 71.6 | 81.6 | 70.1 | 59.5 | 74 | 5.9 |
100
  | **Long Context & Instruction Following** | | | | | | |
101
+ | AA-LCR | 72.3 | 80.3 | 68.0 | 76.7 | 75.0 | 70.0 |
102
+ | IFBench | 78.2 | 82.9 | 76.3 | 76.0 | 79.1 | 76.5 |
103
 
104
  </div>
105