Skip to content
01

Providers + pricing

CapabilitiesProvider details
● Credits · billed by AnyRouter
NVIDIA
nvidia
$0
$0
● Your own key · billed by your provider · AnyRouter fee $0
NVIDIA
nvidia-byok
$0
$0
Unavailable
● Free pool · donated keys · not available for this model — donate a key
02

Try it

POST /api/v1/embeddings
import OpenAI from "openai" const client = new OpenAI({  apiKey: process.env.ANYROUTER_API_KEY,  baseURL: "https://anyrouter.dev/api/v1",}) const resp = await client.embeddings.create({  model: "nvidia/llama-nemotron-embed-vl-1b-v2",  input: "The quick brown fox jumps over the lazy dog",})console.log(resp.data[0].embedding.slice(0, 8))
Set ANYROUTER_API_KEY · edits on the demo update this code
03

Uptime + latency

—recent checks · all routes

No health checks recorded for this model yet.

24h
No traffic in the last 24h.
More performance detail
Usage analytics

Loading usage…

Uptime & Health
No uptime data yet

These providers haven't been health-probed for this model yet. The router still routes around upstreams that fail live requests — uptime fills in once probe history accrues.

Llama Nemotron Embed VL 1B v2Removed

This model has been removed. Requests that still use this id are routed to nvidia/nemotron-3-embed-1b.

  • NVIDIA
  • NVIDIA
Embedtext → [0.12, -0.4, …]8Ktoken context
Context
8K
Input
$0
Output
$0
TTFT
—
Uptime
—
Routes
2
04

Your access

Credits
Your own key · AnyRouter fee $0

Run Llama Nemotron Embed VL 1B v2 on your own key — your requests are billed by the provider. Pool callers pay AnyRouter credits.

No BYOK keys configured for this model yet.

Share a key with the pool to earn credits for every request it serves, covering your plan cost.

Free poolnot in pool yetDonate key →
Create API key for this model
Embedding vectors
Vector dimensionsNot published
Max input8,192 tokens
PriceNot published
Request parameters
inputmodelencoding_formatinput_type
ArchitectureTransformer
Categoryembedding
ReleasedJun 1, 2026
Modalities
→
Capabilities
Embeddings are fixed-length vectors — compare them with cosine similarity for semantic search, RAG retrieval, clustering, and deduplication. Embed queries and documents with the same model, or the distances are meaningless.
05

About

NVIDIA Llama-Nemotron-Embed-VL-1B-v2 is a vision-language embedding model for multimodal question-answering and retrieval over text, images, or combined image-text documents. A transformer encoder fine-tuned from Llama 3.2 1B with SigLip2 400M, it uses a tiling-based VLM architecture (Eagle 2 + nemoretriever-parse) for high-resolution image and complex visual-document understanding. Served via NVIDIA NIM.

Released 2026-06-01 · params: input · model · encoding_format · input_type

Share cards
Llama Nemotron Embed VL 1B v2 share card
Llama Nemotron Embed VL 1B v2
NVIDIA upstream share card
NVIDIA upstream
NVIDIA upstream share card
NVIDIA upstream
Back to models