Embedding model for semantic search, retrieval, clustering, and ranking pipelines
llama-3_2-nemoretriever-300m-embed-v1 by nvidia is currently listed from a single provider. Its official list price is $0 per 1M tokens. It also has 1 free ($0) channel; free tiers usually carry rate limits, and subscription-covered access bills $0 per token only after the subscription fee.
The context window is 32,768 tokens, with an output limit of 2,048 tokens. The weights are open, so it can also be self-hosted or served through a gateway of your choice.
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| NvidiaOfficial nvidia/llama-3_2-nemoretriever-300m-embed-v1 | First-party | Free | — | — | 32,768 | 2,048 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
This model is $0 across all listed channels (Nvidia).