Speech generation model for controllable voice, narration, and audio delivery
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| StepFun (Global)Official | First-party | — | — | — | — | — | — | hostTable.undisclosed |
| StepFun (China)Official | First-party | — | — | — | — | — | — | hostTable.undisclosed |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
This model has no public prices; can't estimate.
Elo from human preference matchups run by Artificial Analysis. Each arena is anchored separately, so Elo is only comparable within one arena, against the other models on that board, and the 95% confidence interval (shown as ±) indicates how much of a gap is noise. Rank is over the whole arena, including models not listed here.