LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS

LLM pricing at a glance

Compare price, capability and every provider's quote for mainstream LLMs in one place

The data here is each provider's published API pricing and specs, with prices normalized to US dollars per million tokens. Filter by capability, context window or budget, or open any model to compare its quote across every provider, with quality and speed benchmarks where available.

1,922 models · 213 providers · 7,692 quotes · synced 2026-09-11

⌕

Filters

Model type

Type is derived from family + output modality. Chat models only by default.

Capability

Input modality

Weights

Context window

Any

Min input price (USD/1M)

Any

Lifecycle

Labs (0)

Providers (0)

Showing 0 / 0 models

Explore models

Browse 1,922 large language models by intelligence, price, context window and availability. Open any model for its full pricing across every provider.

Highest intelligence

  • Claude Fable 5.1AA 65.7
  • Claude Opus 5AA 63.1
  • Claude Fable 5AA 62.1
  • GPT-5.6 SolAA 60.9
  • Grok 4.6AA 60.9
  • Kimi K3AA 59.7
  • GLM-5.3AA 59.5
  • Qwen3.8 MaxAA 58.1
  • Qwen3.8 2.4T A95BAA 57.7
  • GLM-5.3-FlashAA 57.5
  • Claude Opus 4.8AA 57.3
  • Muse Spark 1.2AA 56.8

Lowest blended price

  • Llama 3.2 1B Instruct$0.010
  • Google Gemma 2$0.015
  • Llama 3.2 3B Instruct$0.020
  • E5 Mistral 7B$0.020
  • PaddleOCR-VL$0.020
  • Mistral Nemo$0.022
  • Llama 3.1 8B (decentralized)$0.022
  • Meta Llama 3.1 8B Instruct Turbo$0.022
  • GLM-5.3-Flash$0.024
  • Llama-3.1-8B-Instruct$0.025
  • DeepSeek OCR 2$0.030
  • DeepSeek OCR$0.030

Largest context window

  • Pokee-Isaac 28B1M
  • Qwen Long1M
  • Llama 4 Scout 17B Instruct (US)3.5M
  • Gemini 2.0 Pro 02052.1M
  • Gemini 2.0 Pro 12062.1M
  • Gemini 2.0 Flash-Lite2M
  • Grok 4.202M
  • grok-4-fast-non-reasoning2M
  • grok-4-fast-reasoning2M
  • Auto Router2M
  • X-Ai/Grok 4.1 Fast Non Reasoning2M
  • Grok 4.1 Fast Reasoning (Azure AI Foundry)2M

Most providers

  • GLM-5.2102 hosts
  • GPT OSS 120B82 hosts
  • DeepSeek V4 Pro79 hosts
  • DeepSeek V4 Flash77 hosts
  • Kimi K372 hosts
  • Kimi K2.7 Code64 hosts
  • GLM-5.3-Flash57 hosts
  • GLM-5.352 hosts
  • DeepSeek V4 Flash 073152 hosts
  • GPT-5.549 hosts
  • Claude Sonnet 4.649 hosts
  • Claude Opus 4.847 hosts

How this comparison works

Every figure here is the per-token API rate published by each provider, normalized to US dollars per one million tokens. When a model is hosted by several providers, the table keeps each provider's own quote instead of averaging them, because the same model often differs in price, context window and capability flags from one host to the next.

The blended price ranks models on one number: input price times 0.75 plus output price times 0.25, matching a 3-to-1 input-to-output usage assumption. The lowest paid channel excludes $0 tiers, since free access usually carries rate limits and no service-level guarantee. Quality scores, where shown, come from Artificial Analysis; most models have no independent quality data, so those rows show pricing only.

Pricing and capability data is aggregated from models.dev, and quality and speed data from Artificial Analysis. Figures are indicative and can change, so the official provider page is the final authority before you commit to a channel.

Frequently asked questions

What is the cheapest LLM API in 2026?
Based on our comparison of 1922+ models, Llama 3.2 1B Instruct from Inference offers the lowest blended price at about $0.010 per 1M tokens.
How do I estimate LLM API costs?
Multiply your input tokens by the model's input price per 1M tokens, do the same for output tokens, then add them. Use the built-in cost calculator to estimate totals for any model and provider.
Which LLM has the largest context window?
Pokee-Isaac 28B from misc supports the largest context window, up to 10,000,000 tokens.
Which providers and models are compared here?
This site compares 1922+ models across providers like Alibaba, OpenAI, Anthropic, Google, Zhipu AI, Mistral, with input, output and cache pricing per 1M tokens plus quality leaderboards.