← Model list
Hermes 3 70B
nousresearch·nousresearch/hermes-3-llama-3.1-70b·GA·Open weights·nousresearch series
Open Llama instruction model for multilingual chat, reasoning, and coding
At a glance
Reference price$0.408 / $0.408 per 1M
Cache read $0.204 · NanoGPT
Lowest paid$0.408 / $0.408
NanoGPT Gateway
Context65,536
Output limit8,192
Capabilities
Reasoning✓ Tool useStructured output? TemperatureAttachments
Modalities
Text
Knowledge cutoff—
Released / updated2026-01-07 / 2026-01-07
Quality & performanceArtificial Analysis · Intelligence Index v4.1
Intelligence4.8
Value12
Coding—
Agentic—
Output speed34 tok/s
TTFT1.97 s
Value formulaIQ 4.8 ÷ min blended $0.408 = 12
Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.
Available at 1 providers1 with public prices
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| NanoGPT | Gateway | $0.408 | $0.408 | $0.204 | — | 65,536 | 8,192 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Your usage cost
1NanoGPT$77.52
The cheapest paid channel is the only channel.
Related models
Hermes 4 (Thinking)same series$0.201 / $0.40Hermes 4 Largesame series$0.30 / $1.20Nous: Hermes 4 70Bsame series$0.13 / $0.40Llama-3.1-8B-Instructcheaper alternative$0.15 / $0.45Llama 3.2 3B Instructcheaper alternative$0.10 / $0.335Mistral Nemocheaper alternative$0.15 / $0.15DeepSeek Chatcheaper alternative$0.14 / $0.28
Price historyone sample accumulated per data sync
Input listOutput listMin blended