LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Model list
ib

Granite-4.0-H-Small

ibm·ibm/granite-4-h-small·GA·Open weights·granite series
Open-weight hybrid model for enterprise chat, coding, retrieval-augmented generation, and tool-calling workloads

Granite-4.0-H-Small by ibm is currently listed from a single provider. Its reference price is $0.064 per 1M input tokens and $0.265 per 1M output tokens.

Artificial Analysis rates it 6 on the Intelligence Index. Against its lowest blended price of $0.114 per 1M, that is roughly 53 index points per dollar, which is the value ratio the leaderboards rank on. Median output speed is 14 tokens per second, with 35.79s to the first token.

The context window is 131,072 tokens, with an output limit of 131,072 tokens. It supports tool use and structured output. The weights are open, so it can also be self-hosted or served through a gateway of your choice.

Specs & pricing

Input / output per 1M tokens
Reference price·watsonx.ai
$0.064 / $0.27
Blended $0.11 · Cache read —
Lowest paid·watsonx.aiCloud
$0.064 / $0.27
Blended $0.11
Context
131,072
Output limit
131,072
Knowledge cutoff
—
Released / updated
2025-10-02 / 2025-10-02
Capabilities
Reasoning✓ Tool use✓ Structured output✓ TemperatureAttachments
Modalities
Text
Weights
Hugging Face

Quality & performanceArtificial Analysis · Intelligence Index v4.3

Intelligence6
Value53
Coding—
Agentic—
Output speed14 tok/s
TTFT35.79 s
Value formulaIQ 6 ÷ min blended $0.114 = 53

Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.

Available at 1 providers1 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
watsonx.aiCloud$0.064$0.27——131,072131,072

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Your usage cost

1watsonx.ai$25.97
The cheapest paid channel is the only channel.

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$0.27$02026-08-092026-08-132026-08-09 · Input list $0.0642026-08-10 · Input list $0.0642026-08-11 · Input list $0.0642026-08-12 · Input list $0.0642026-08-13 · Input list $0.0642026-08-09 · Output list $0.272026-08-10 · Output list $0.272026-08-11 · Output list $0.272026-08-12 · Output list $0.272026-08-13 · Output list $0.272026-08-09 · Min blended $0.112026-08-10 · Min blended $0.112026-08-11 · Min blended $0.112026-08-12 · Min blended $0.112026-08-13 · Min blended $0.11

Artificial Analysis evaluations9 items

GPQA Diamond41.6%
Humanity's Last Exam3.8%
MMLU-Pro62.4%
LiveCodeBench25.1%
Terminal-Bench Hard2.3%
AIME 202513.7%
τ²-Bench Telecom17.3%
AA-LCR11.3%
IFBench31.5%

Individual evaluations run by Artificial Analysis, on the same reasoning tier as the intelligence score above. Each benchmark has its own task set and harness, so rows are not comparable with one another. The Intelligence Index above draws on a different, newer set of evaluations.

Related models

ibGranite 4.2 8Bsame series$0.10 / $0.15Granite 4.2 8B (DeepInfra)same series$0.06 / $0.25ibGranite 4.1 8Bsame series$0.05 / $0.10Ministral 3Bcheaper alternative$0.04 / $0.04Command R7Bcheaper alternative$0.038 / $0.15inSchematron V2 Turbocheaper alternative$0.03 / $0.15Lyria 3 Clip Previewcheaper alternative$0 / $0
Data partly from models.dev (MIT) · watsonx.ai official docs ↗