StepFun · Step 3.7
Open-weight (Apache-2.0) 198B/11B-active multimodal MoE reasoning model from StepFun built for high-throughput agent and coding workloads.
Leader: Claude Opus 5.5 at 58
Leader: Claude Fable 5.1 at 91.4%
Leader: Claude Opus 5.5 at 66.9%
Leader: Claude Opus 5.5 at 89.9%
Leader: Claude Opus 5.5 at 61.4%
Leader: GPT-6 Astra at 96.1%
Leader: GPT-5.6 Sol at 32.3%
Leader: Kimi K3 at 88.7%
Leader: Claude Opus 5.5 at 87.7%
| Result | Score | Reported by | Source |
|---|---|---|---|
| Terminal-Bench Hard (AA) · AA | 35.6% | independent | Artificial Analysis: Step 3.7 Flash ↗ |
| IFBench (AA) · AA | 67.3% | independent | Artificial Analysis: Step 3.7 Flash ↗ |
| AA-Omniscience Index · AA | -37.3 index (-100..100) | independent | Artificial Analysis: Step 3.7 Flash ↗ |
| Toolathlon | 49.5% | vendor | Hugging Face: stepfun-ai/Step-3.7-Flash ↗ |
| ClawEval-1.1 | 67.1% | vendor | Hugging Face: stepfun-ai/Step-3.7-Flash ↗ |
| SimpleVQA (search) | 79.2% | vendor | Hugging Face: stepfun-ai/Step-3.7-Flash ↗ |
| V* (python) | 95.3% | vendor | Hugging Face: stepfun-ai/Step-3.7-Flash ↗ |
| GDPval-AA (vendor-run) | 45.8 as stated on model card (unit unspecified) | vendor | Hugging Face: stepfun-ai/Step-3.7-Flash ↗ |