← Model list

Llama 3.3 70B

meta·meta/llama3-3-70b·GA·Open weights·llama series·NEW

Open Llama instruction model for multilingual chat, reasoning, and coding

At a glance
Reference price$1.75 / $2.75 per 1M
Cache read $1.75 · NanoGPT
Lowest paid$0.53 / $0.76
STACKIT Cloud · 3.3× spread
Context128,000
⚠ Providers report 128,000–131,072; the table below is authoritative
Output limit16,384
Capabilities
ReasoningTool useStructured outputTemperatureAttachments
⚠ Providers report capability flags inconsistently
Modalities
Text
Knowledge cutoff2023-12
Released / updated2024-12-06 / 2026-07-02

Quality & performance

Artificial Analysis doesn't cover this model (267 of 2059 have data). Quality data comes from independent evals covering widely used models.

Available at 6 providers6 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
STACKIT
cortecs/Llama-3.3-70B-Instruct-FP8-Dynamic
Cloud$0.53$0.76128,0004,096
Groq
llama-3.3-70b-versatile
Cloud$0.59$0.79131,07232,768
Weights & Biases
meta-llama/Llama-3.3-70B-Instruct
Cloud$0.71$0.71$0.71128,000128,000
Together AI
meta-llama/Llama-3.3-70B-Instruct-Turbo
Cloud$1.04$1.04131,072131,072
Venice AI
llama-3.3-70b
Gateway$0.70$2.80128,0004,096
NanoGPTGateway$1.75$2.75$1.75128,00016,384

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Your usage cost

1STACKIT$144.00
2Groq$157.50
3Weights & Biases$177.50
4Together AI$260.00
5Venice AI$280.00
6NanoGPT$487.50
The cheapest paid channel is the only channel.

Related models

Price historyone sample accumulated per data sync

Input list $1.75Output list $2.75Min blended $0.588

Price history accumulates from each data sync; currently only 1 sample(s) (2026-08-13). Each future sync adds a point, and once accumulated a line is drawn here.