Multi-agent model for routing expert agents across complex analytical tasks
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| Sakana AIOfficial | First-party | — | — | — | — | 1,000,000 | 1,000,000 | hostTable.undisclosed |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
This model has no public prices; can't estimate.
| Name | Conditions | Score | Metric | Source |
|---|---|---|---|---|
| SWE Bench Pro | — | 59 | — | Source ↗ |
| Terminal Bench 2.1 | — | 80.2 | — | Source ↗ |
| LiveCodeBench | — | 92.9 | — | Source ↗ |
| LiveCodeBench Pro | — | 87.8 | — | Source ↗ |
| Humanity’s Last Exam | — | 47.2 | — | Source ↗ |
| CharXiv Reasoning | — | 85.1 | — | Source ↗ |
| GPQA Diamond | — | 95.5 | — | Source ↗ |
| SciCode | — | 60.1 | — | Source ↗ |
| τ3 Banking | — | 21.7 | — | Source ↗ |
| Long Context Reasoning | — | 74.7 | — | Source ↗ |
| MRCRv2 | — | 86.6 | — | Source ↗ |
| CTI-REALM | — | 67.5 | — | Source ↗ |
The same benchmark scores very differently across harness / dataset, so the qualifying conditions must be shown together.