# NVIDIA Nemotron 3 Ultra

> NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context. Designed for advanced function calling, structured output, and complex reasoning tasks.

**ID**: `venice/nvidia-nemotron-3-ultra-550b-a55b`  
**Creator**: nvidia  
**Category**: text  
**Context**: 1M tokens  
**Pricing**: $0.63 / $3.13 per 1M · billed by your provider (BYOK)  
**Released**: 2025-12-10  
**Web page**: https://anyrouter.dev/model/venice/nvidia-nemotron-3-ultra-550b-a55b

**Input modalities**: text  
**Output modalities**: text  

**Capabilities**: chat, reasoning, function-calling, structured-output, streaming

## Usage

**Endpoint**: `POST https://anyrouter.dev/api/v1/chat/completions`

```bash
curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer $ANYROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "venice/nvidia-nemotron-3-ultra-550b-a55b",
  "messages": [
    {
      "role": "user",
      "content": "Say hi in 3 words."
    }
  ]
}'
```

## Providers

| Provider | Input | Output |
| --- | --- | --- |
| venice-byok | $0.63/M (list) | $3.13/M (list) |

BYOK routes are billed by your own provider key — AnyRouter charges $0. Rates marked (list) are the provider's published list price.

## Benchmarks

- **GPQA Diamond**: 0.9
- **HLE**: 0.3
- **τ²-Bench Telecom**: 0.8
- **SciCode**: 0.4

## Source

- Upstream docs: https://docs.venice.ai
