LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS

LLM pricing at a glance

Compare price, capability and every provider's quote for mainstream LLMs in one place

The data here is each provider's published API pricing and specs, with prices normalized to one million tokens and quoted in US dollars, switchable to a local currency at the latest exchange rate. Filter by capability, context window or budget, or open any model to compare its quote across every provider, with quality and speed benchmarks where available.

2,025 models · 226 providers · 8,389 quotes · synced 2026-10-05

⌕

Filters

Model type

Type is derived from family + output modality. Chat models only by default.

Capability

Input modality

Weights

Context window

Any

Min input price (USD/1M)

Any

Lifecycle

Labs (0)

Providers (0)

Showing 0 / 0 models

Explore models

Browse 2,025 large language models by intelligence, price, context window and availability. Open any model for its full pricing across every provider.

Highest intelligence

  • Claude Opus 5.5AA 57.6
  • Claude Sonnet 5.5AA 56
  • Claude Fable 5.1AA 53.4
  • GPT-6 AstraAA 52.7
  • GPT-6.1 SolAA 51.8
  • Claude Opus 5AA 50.8
  • Claude Fable 5AA 49.6
  • Muse Spark 1.3AA 48.1
  • GPT-6 SolAA 47.6
  • GPT-5.6 SolAA 47
  • Grok 4.7AA 46.4
  • MiMo-V2.6-ProAA 46.3

Lowest blended price

  • Mistral Nemo$0.022
  • DeepSeek-OCR$0.03
  • Ling 3.0 Flash VL$0.031
  • Ling 3.0 Flash$0.032
  • Ministral 3B$0.04
  • qwen3.5-2b$0.04
  • Granite 4.0 Micro$0.041
  • DeepSeek V4 Flash 0731$0.044
  • Nex AGI: Nex-N2.5-Mini$0.044
  • GLM-5.3-Flash$0.048
  • Qwen3.5 4B$0.048
  • Qwen3.7 Flash$0.049

Largest context window

  • Grok 4.1 Fast (Non-Reasoning)2M
  • Grok 4.202M
  • grok-4-fast-non-reasoning2M
  • grok-4-fast-reasoning2M
  • GPT-5.51.05M
  • GPT-5.41.05M
  • GPT-5.6 Luna1.05M
  • GPT-5.6 Sol1.05M
  • GPT-5.6 Terra1.05M
  • GPT-6 Astra1.05M
  • GPT-6 Luna1.05M
  • GPT-6 Sol1.05M

Most providers

  • GLM-5.2102 hosts
  • Kimi K388 hosts
  • GPT OSS 120B84 hosts
  • GLM-5.381 hosts
  • GLM-5.3-Flash80 hosts
  • DeepSeek V4.1 Flash73 hosts
  • DeepSeek V4 Flash 073154 hosts
  • Claude Sonnet 4.650 hosts
  • GPT-5.549 hosts
  • Claude Opus 4.848 hosts
  • Claude Opus 4.747 hosts
  • Qwen3.8 27B46 hosts

How this comparison works

Every figure here is the per-token API rate published by each provider, normalized to one million tokens and quoted in US dollars. Other currencies are converted from that dollar figure at a daily exchange rate, so they are reference values rather than what the provider publishes. When a model is hosted by several providers, the table keeps each provider's own quote instead of averaging them, because the same model often differs in price, context window and capability flags from one host to the next.

The blended price ranks models on one number: input price times 0.75 plus output price times 0.25, matching a 3-to-1 input-to-output usage assumption. The lowest paid channel excludes $0 tiers, since free access usually carries rate limits and no service-level guarantee. Quality scores, where shown, come from Artificial Analysis; most models have no independent quality data, so those rows show pricing only.

Pricing and capability data is aggregated from models.dev, and quality and speed data from Artificial Analysis. Figures are indicative and can change, so the official provider page is the final authority before you commit to a channel.

Frequently asked questions

What is the cheapest LLM API in 2026?
Based on our comparison of 2025+ models, Mistral Nemo from OpenRouter offers the lowest blended price at about $0.022 per 1M tokens.
How do I estimate LLM API costs?
Multiply your input tokens by the model's input price per 1M tokens, do the same for output tokens, then add them. Use the built-in cost calculator to estimate totals for any model and provider.
Which LLM has the largest context window?
Grok 4.1 Fast (Non-Reasoning) from xai claims the largest context window, up to 2,000,000 tokens. Its hosts report different limits, from 128,000 to 2,000,000 tokens, so the usable window depends on the provider you pick.
Which providers and models are compared here?
This site compares 2025+ models across providers like OpenAI, Alibaba, Anthropic, Google, Zhipu AI, Mistral, with input, output and cache pricing per 1M tokens plus quality leaderboards.