LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Model list

llama-3_2-nemoretriever-300m-embed-v1

nvidia·nvidia/llama-3-2-nemoretriever-300m-embed-v1·GA·Open weights·Embedding
Embedding model for semantic search, retrieval, clustering, and ranking pipelines

llama-3_2-nemoretriever-300m-embed-v1 by nvidia is currently listed from a single provider. Its official list price is $0 per 1M tokens. It also has 1 free ($0) channel; free tiers usually carry rate limits, and subscription-covered access bills $0 per token only after the subscription fee.

The context window is 32,768 tokens, with an output limit of 2,048 tokens. The weights are open, so it can also be self-hosted or served through a gateway of your choice.

Specs & pricing

Input / output per 1M tokens
Official price·Nvidia
$0 / $0
Blended $0 · Cache read —
Lowest paid
No paid channels
1 more $0 channels
Context
32,768
Output limit
2,048
Knowledge cutoff
—
Released / updated
2025-07-24 / 2025-07-24
Capabilities
ReasoningTool use? Structured outputTemperatureAttachments
Modalities
Text

Available at 1 providers0 with public prices · 1 free

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
NvidiaOfficial
nvidia/llama-3_2-nemoretriever-300m-embed-v1
First-partyFree——32,7682,048

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Your usage cost

This model is $0 across all listed channels (Nvidia).

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$0.0001$02026-08-052026-08-132026-08-05 · Input list $02026-08-06 · Input list $02026-08-07 · Input list $02026-08-08 · Input list $02026-08-09 · Input list $02026-08-10 · Input list $02026-08-11 · Input list $02026-08-12 · Input list $02026-08-13 · Input list $02026-08-05 · Output list $02026-08-06 · Output list $02026-08-07 · Output list $02026-08-08 · Output list $02026-08-09 · Output list $02026-08-10 · Output list $02026-08-11 · Output list $02026-08-12 · Output list $02026-08-13 · Output list $0
Data partly from models.dev (MIT) · Nvidia official docs ↗