LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Providers

Inference

inference·Gateway·9 models·Official docs ↗

Inference is a third-party gateway that routes to many upstream models through one API. Gateway quotes are often among the lowest on the site, but availability, rate limits and routing can vary between models, so a low headline price is worth verifying against your own workload.

It lists 9 models on this page. Coverage spans text and chat (8) and embeddings (1). The lowest-priced model hosted here is Qwen 3 Embedding 4B.

Inference is the lowest paid channel for 4 models that have at least two competing paid hosts, which is where its pricing genuinely leads rather than being the only option.

Coverage & pricing
Models hosted9With public price 9
Cheapest anywhere4
Model typesText / Chat 8Embedding 1
Cheapest model hereQwen 3 Embedding 4B
SDK package@ai-sdk/openai-compatible
APIhttps://inference.net/v1

Hosted models by blended price asc

⌕
9 of 9
Cheapest here?
Qwen 3 Embedding 4BEmbeddingqwen/qwen3-embedding-4b$0.01$0—32,000Lowest anywhere
Llama 3.2 1B Instructmeta/llama-3.2-1b-instruct$0.01$0.01—16,000Lowest anywhere
Llama 3.2 3B Instructmeta/llama-3.2-3b-instruct$0.02$0.02—16,000Lowest anywhere
Llama-3.1-8B-Instructmeta/llama-3.1-8b-instruct$0.025$0.025—16,0001.0× pricier
Mistral Nemo 12B Instructmistral/mistral-nemo-12b-instruct$0.038$0.10—16,000Only channel
Llama-3.2-11B-Vision-Instructmeta/llama-3.2-11b-vision-instruct$0.055$0.055—16,000Lowest anywhere
Google Gemma 3google/gemma-3$0.15$0.30—125,000Only channel
osOsmosis Structure 0.6Bosmosis/osmosis-structure-0.6b$0.10$0.50—4,000Only channel
Qwen 2.5 7B Vision Instructqwen/qwen-2.5-7b-vision-instruct$0.20$0.20—125,000Only channel