GLM-4.6V scales its context window to 128k tokens in training, and achieves SoTA performance in visual understanding among models of similar parameter scales. Integrates native Function Calling capabilities, bridging 'visual perception' and 'executable action' for multimodal agents. Direct via Z-AI (Zhipu).
GLM 4.6V Original by zhipuai is currently listed from a single provider. Its reference price is $0.60 per 1M input tokens and $0.90 per 1M output tokens.
The context window is 128,000 tokens, with an output limit of 24,000 tokens. Accepted input modalities are Text and Image. The weights are open, so it can also be self-hosted or served through a gateway of your choice.
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| NanoGPT | Gateway | $0.60 | $0.90 | $0.30 | — | 128,000 | 24,000 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.