LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Model list

Gemma 4 12B IT

google·google/gemma-4-12b-it·GA·Open weights·gemma series
Compact Gemma 4 instruction model for open, self-hosted chat and reasoning

Gemma 4 12B IT by google is offered by 3 providers on this page. Its reference price is $0.25 per 1M tokens. The lowest paid channel is NanoGPT at $0.05 / $0.25 per 1M, about 5.0× below the reference price. That channel is a third-party gateway, so confirm its availability and rate limits before depending on it.

The context window is 32,768 tokens at the reference host, but hosts report different limits, from 32,768 to 262,144, so the usable window depends on the provider you pick. It supports reasoning, tool use, and structured output. Accepted input modalities are Text, Image, Video, and Audio. The weights are open, so it can also be self-hosted or served through a gateway of your choice. Providers report the capability flags inconsistently, so verify a specific feature against the host you plan to use.

Specs & pricing

Input / output per 1M tokens
Reference price·Pioneer
$0.25 / $0.25
Blended $0.25 · Cache read $0.25
Lowest paid·NanoGPTGateway
$0.05 / $0.25
Blended $0.10 · 5× spread
Context
32,768
Output limit
32,768
Knowledge cutoff
—
Released / updated
2026-06-09 / 2026-06-09
Capabilities
✓ Reasoning✓ Tool use✓ Structured output✓ Temperature✓ Attachments
⚠ Providers report capability flags inconsistently
Modalities
TextImageVideoAudio
Weights
Hugging Face

Available at 3 providers3 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
NanoGPTGateway$0.05$0.25$0.025—131,072 ⚠32,768
SiliconFlow
google/gemma-4-12B-it
Cloud$0.10$0.30——262,144 ⚠262,144
Pioneer
google/gemma-4-12B-it
Gateway$0.25$0.25$0.25$0.2532,76832,768

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Reasoning control

effort = loweffort = mediumeffort = high

Interleaved thinking (reasoning between tool calls) is declared by 1 of 3 providers.

Your usage cost

1NanoGPT$19.50
2SiliconFlow$35.00
3Pioneer$62.50
The cheapest paid channel is the only channel.

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$0.30$02026-08-052026-09-032026-08-05 · Input list $0.062026-08-06 · Input list $0.062026-08-07 · Input list $0.062026-08-08 · Input list $0.062026-08-09 · Input list $0.062026-08-10 · Input list $0.062026-08-11 · Input list $0.062026-08-12 · Input list $0.062026-08-13 · Input list $0.062026-09-01 · Input list $0.052026-09-03 · Input list $0.252026-08-05 · Output list $0.302026-08-06 · Output list $0.302026-08-07 · Output list $0.302026-08-08 · Output list $0.302026-08-09 · Output list $0.302026-08-10 · Output list $0.302026-08-11 · Output list $0.302026-08-12 · Output list $0.302026-08-13 · Output list $0.302026-09-01 · Output list $0.252026-09-03 · Output list $0.252026-08-05 · Min blended $0.122026-08-06 · Min blended $0.122026-08-07 · Min blended $0.122026-08-08 · Min blended $0.122026-08-09 · Min blended $0.122026-08-10 · Min blended $0.122026-08-11 · Min blended $0.122026-08-12 · Min blended $0.122026-08-13 · Min blended $0.122026-09-01 · Min blended $0.102026-09-03 · Min blended $0.10

Related models

Gemma 4 26B A4B Cybersecuritysame series$0.11 / $0.33Gemma 4 31B Split-Untiedsame series$0.10 / $0.30Gemma 4 12B Semancersame series$0.05 / $0.25GPT-5 Nanocheaper alternative$0.05 / $0.40Hy3cheaper alternative$0 / $0Qwen3.7 Flashcheaper alternative$0.03 / $0.13Nemotron 3.5 Lightning 30B A3Bcheaper alternative$0 / $0
Data partly from models.dev (MIT) · Pioneer official docs ↗