Coverage
Recorded results across recent models and the most covered benchmark families. Empty cells indicate gaps in this registry.
Updated
Recent models are selected in rounds across providers, up to 25 models total. Cells link to the newest reported observation in each family.
Coverage matrix
By provider
| Provider | Models | Benchmarks covered | Records |
|---|---|---|---|
| Alibaba (Tongyi) | 8 | 27 | 80 |
| Anthropic | 16 | 21 | 199 |
| Cursor | 4 | 3 | 15 |
| DeepSeek AI | 2 | 10 | 54 |
| Google (DeepMind) | 16 | 37 | 120 |
| Meta AI (originally Facebook AI Research) | 5 | 26 | 66 |
| Microsoft AI | 3 | 14 | 19 |
| MiniMax | 3 | 8 | 15 |
| Mistral AI | 5 | 10 | 23 |
| Moonshot AI | 3 | 17 | 45 |
| NVIDIA | 3 | 7 | 20 |
| OpenAI | 16 | 24 | 214 |
| SpaceXAI | 3 | 6 | 12 |
| Thinking Machines Lab | 2 | 11 | 23 |
| Z.ai | 6 | 13 | 40 |
Stale results
Benchmarks whose newest reported result is more than 90 days before the latest data timestamp.
- Codeforces · newest reported
- Massive Multitask Language Understanding · newest reported
- Multi-SWE-bench · newest reported
- ChartQA · newest reported
- MathVerse · newest reported
- MMStar · newest reported
- OCRBench · newest reported
- Embodied Reasoning Question Answer · newest reported
- FrontierScience · newest reported
- International Physics Olympiad · newest reported
- ScreenSpot-Pro · newest reported
- AI2D · newest reported
- HallusionBench · newest reported
- MathVista · newest reported
- Massive Multi-discipline Multimodal Understanding · newest reported
- GDPval · newest reported
- BeyondAIME · newest reported
- τ³-bench · newest reported
- APEX-Agents · newest reported
- CL-bench · newest reported
- SWE-fficiency · newest reported