← Model list
llama-3_2-nemoretriever-300m-embed-v1
nvidia·nvidia/llama-3-2-nemoretriever-300m-embed-v1·GA·Open weights·Embedding
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
At a glance
Official price$0 / $0 per 1M
Cache read — · Nvidia
Lowest paidNo paid channels
1 more $0 channels
Context32,768
Output limit2,048
Capabilities
ReasoningTool use? Structured outputTemperatureAttachments
Modalities
Text
Knowledge cutoff—
Released / updated2025-07-24 / 2025-07-24
Quality & performance
Artificial Analysis doesn't cover this model (267 of 2059 have data). Quality data comes from independent evals covering widely used models.
Available at 1 providers0 with public prices · 1 free
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| NvidiaOfficial nvidia/llama-3_2-nemoretriever-300m-embed-v1 | First-party | Free | — | — | 32,768 | 2,048 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Your usage cost
This model is $0 across all listed channels (Nvidia).
Price historyone sample accumulated per data sync
Input listOutput listMin blended