← Model list

Granite-4.0-H-Small

ibm·ibm/granite-4-h-small·GA·Open weights·granite series

Open-weight hybrid model for enterprise chat, coding, retrieval-augmented generation, and tool-calling workloads

At a glance
Reference price$0.064 / $0.265 per 1M
Cache read · watsonx.ai
Lowest paid$0.064 / $0.265
watsonx.ai Gateway
Context131,072
Output limit131,072
Capabilities
ReasoningTool useStructured outputTemperatureAttachments
Modalities
Text
Knowledge cutoff
Released / updated2025-10-02 / 2025-10-02

Quality & performanceArtificial Analysis · Intelligence Index v4.1

Intelligence4.9
Value43
Coding
Agentic
Output speed25 tok/s
TTFT18.74 s
Value formulaIQ 4.9 ÷ min blended $0.114 = 43

Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.

Available at 1 providers1 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
watsonx.aiGateway$0.064$0.265131,072131,072

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Your usage cost

1watsonx.ai$25.97
The cheapest paid channel is the only channel.

Related models

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$0.265$02026-08-092026-08-132026-08-09 · Input list $0.0642026-08-10 · Input list $0.0642026-08-11 · Input list $0.0642026-08-12 · Input list $0.0642026-08-13 · Input list $0.0642026-08-09 · Output list $0.2652026-08-10 · Output list $0.2652026-08-11 · Output list $0.2652026-08-12 · Output list $0.2652026-08-13 · Output list $0.2652026-08-09 · Min blended $0.1142026-08-10 · Min blended $0.1142026-08-11 · Min blended $0.1142026-08-12 · Min blended $0.1142026-08-13 · Min blended $0.114