Inference is a third-party gateway that routes to many upstream models through one API. Gateway quotes are often among the lowest on the site, but availability, rate limits and routing can vary between models, so a low headline price is worth verifying against your own workload.
It lists 9 models on this page. Coverage spans text and chat (8) and embeddings (1). The lowest-priced model hosted here is Qwen 3 Embedding 4B.
Inference is the lowest paid channel for 4 models that have at least two competing paid hosts, which is where its pricing genuinely leads rather than being the only option.
| Cheapest here? | |||||
|---|---|---|---|---|---|
| Qwen 3 Embedding 4BEmbeddingqwen/qwen3-embedding-4b | $0.01 | $0 | — | 32,000 | Lowest anywhere |
| Llama 3.2 1B Instructmeta/llama-3.2-1b-instruct | $0.01 | $0.01 | — | 16,000 | Lowest anywhere |
| Llama 3.2 3B Instructmeta/llama-3.2-3b-instruct | $0.02 | $0.02 | — | 16,000 | Lowest anywhere |
| Llama-3.1-8B-Instructmeta/llama-3.1-8b-instruct | $0.025 | $0.025 | — | 16,000 | 1.0× pricier |
| Mistral Nemo 12B Instructmistral/mistral-nemo-12b-instruct | $0.038 | $0.10 | — | 16,000 | Only channel |
| Llama-3.2-11B-Vision-Instructmeta/llama-3.2-11b-vision-instruct | $0.055 | $0.055 | — | 16,000 | Lowest anywhere |
| Google Gemma 3google/gemma-3 | $0.15 | $0.30 | — | 125,000 | Only channel |
| Osmosis Structure 0.6Bosmosis/osmosis-structure-0.6b | $0.10 | $0.50 | — | 4,000 | Only channel |
| Qwen 2.5 7B Vision Instructqwen/qwen-2.5-7b-vision-instruct | $0.20 | $0.20 | — | 125,000 | Only channel |