Compact GPT model for low-latency assistance and high-volume workloads
GPT-3.5 Turbo 0125 by openai is offered by 2 providers on this page. Its reference price is $0.50 per 1M input tokens and $1.50 per 1M output tokens.
The context window is 16,384 tokens, with an output limit of 16,384 tokens. Its training knowledge cuts off at 2021-08.
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| Azure Cognitive Services | First-party | $0.50 | $1.50 | — | — | 16,384 | 16,384 | deprecated |
| Azure | First-party | $0.50 | $1.50 | — | — | 16,384 | 16,384 | deprecated |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.