← Model list
Granite-4.0-H-Small
ibm·ibm/granite-4-h-small·GA·Open weights·granite series
Open-weight hybrid model for enterprise chat, coding, retrieval-augmented generation, and tool-calling workloads
At a glance
Reference price$0.064 / $0.265 per 1M
Cache read — · watsonx.ai
Lowest paid$0.064 / $0.265
watsonx.ai Gateway
Context131,072
Output limit131,072
Capabilities
Reasoning✓ Tool use✓ Structured output✓ TemperatureAttachments
Modalities
Text
Knowledge cutoff—
Released / updated2025-10-02 / 2025-10-02
Quality & performanceArtificial Analysis · Intelligence Index v4.1
Intelligence4.9
Value43
Coding—
Agentic—
Output speed25 tok/s
TTFT18.74 s
Value formulaIQ 4.9 ÷ min blended $0.114 = 43
Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.
Available at 1 providers1 with public prices
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| watsonx.ai | Gateway | $0.064 | $0.265 | — | — | 131,072 | 131,072 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Your usage cost
1watsonx.ai$25.97
The cheapest paid channel is the only channel.
Related models
Granite 4.1 8Bsame series$0.05 / $0.10IBM: Granite 4.1 8Bsame series$0.05 / $0.10Granite 4.0 Microsame series$0.017 / $0.112Llama 4 Maverick 17B 128E Instruct FP8cheaper alternative$0 / $0Lyria 3 Clip Previewcheaper alternative$0 / $0Command R7Bcheaper alternative$0.037 / $0.15
Price historyone sample accumulated per data sync
Input listOutput listMin blended