Skip to content
01

Providers + pricing

CapabilitiesProvider details
● Credits · billed by AnyRouter
NVIDIA
nvidia
$0
$0
● Your own key · billed by your provider · AnyRouter fee $0
NVIDIA
nvidia-byok
$0
$0
Unavailable
● Free pool · donated keys · not available for this model — donate a key
02

Try it

POST /api/v1/chat/completions
import { createAnyRouter } from "@anyr/ai-sdk-provider"import { generateText } from "ai" const anyrouter = createAnyRouter() const { text } = await generateText({  model: anyrouter("google/diffusiongemma-26b-a4b-it"),  prompt: "Say hi in 3 words.",})console.log(text)
Set ANYROUTER_API_KEY · edits on the demo update this code
03

Uptime + latency

—recent checks · all routes

No health checks recorded for this model yet.

24h
No traffic in the last 24h.
More performance detail
Usage analytics

Loading usage…

Uptime & Health
No uptime data yet

These providers haven't been health-probed for this model yet. The router still routes around upstreams that fail live requests — uptime fills in once probe history accrues.

DiffusionGemma 26B A4BDisabled since Aug 19, 2026

Chat completions return empty content for this discrete-diffusion model under normal chat token budgets. It is unlisted rather than advertised as a live chat model.

  • NVIDIA
  • NVIDIA
Visiontext + image + video → text262Ktoken context
Context
262K
Input
$0
Output
$0
TTFT
200ms
Uptime
—
Routes
2
04

Your access

Credits
Your own key · AnyRouter fee $0

Run DiffusionGemma 26B A4B on your own key — your requests are billed by the provider. Pool callers pay AnyRouter credits.

No BYOK keys configured for this model yet.

Share a key with the pool to earn credits for every request it serves, covering your plan cost.

Free poolnot in pool yetDonate key →
Create API key for this model
Text generation
Context length262,144 tokens
Max output209,715 tokens
ArchitectureTransformer
Categorytext
ReleasedJun 10, 2026
Modalities
→
Capabilities
05

About

DiffusionGemma 26B A4B IT is an open-weights multimodal model from Google DeepMind that generates text via discrete diffusion. Built on the Gemma 4 26B A4B MoE architecture (25.2B total / 3.8B active), it emits tokens in parallel 256-token blocks for high-throughput generation, with a 256K context window, configurable thinking mode, native function calling, and 35+ language support.

Released 2026-06-10 · params: max_tokens · temperature · top_p · stop

API & code
Share cards
DiffusionGemma 26B A4B share card
DiffusionGemma 26B A4B
NVIDIA upstream share card
NVIDIA upstream
NVIDIA upstream share card
NVIDIA upstream
Back to models