01
Providers + pricing
| Capabilities | Provider details | ||||
|---|---|---|---|---|---|
| ● Your own key · billed by your provider · AnyRouter fee $0 | |||||
AIHubMix aihubmix-byok | $0.15list | $0.50list | Unavailable | ||
OpenCode Zen opencode-zen-byok | $0.15list | $0.50list | Unavailable | ||
CommandCode commandcode-byok | $0.15list | $0.50list | Unavailable | ||
OpenRouter openrouter-byok | $0.15list | $0.50list | Unavailable | ||
Venice AI venice-byok | $0.15list | $0.50list | Unavailable | ||
Cline cline-byok | $0.15list | $0.50list | Unavailable | ||
Ramp Router router-byok | $0.15list | $0.50list | Unavailable | ||
| ● Free pool · donated keys · not available for this model — donate a key | |||||
02
Try it
POST /api/v1/chat/completions
import { createAnyRouter } from "@anyr/ai-sdk-provider"import { generateText } from "ai" const anyrouter = createAnyRouter() const { text } = await generateText({ model: anyrouter("z-ai/glm-flash-latest"), prompt: "Say hi in 3 words.",})console.log(text)
Set
ANYROUTER_API_KEY · edits on the demo update this codeLive demo
03
Uptime + latency
—recent checks · all routes
No health checks recorded for this model yet.
24h
No traffic in the last 24h.
More performance detail
Glm Flash (latest)
Also accepted:
zai-org/glm-5.3-flash- AIHubMix
- OpenCode Zen
- CommandCode
- OpenRouter
- Venice AI
- Cline
- Ramp Router
Visiontext + image + video → text1Mtoken context
Context
1MInput
$0.15Output
$0.50TTFT
—Uptime
—Routes
704
Your access
CreditsTín dụngCredits
Your own key · AnyRouter fee $0
Run Glm Flash (latest) on your own key — your requests are billed by the provider. Pool callers pay AnyRouter credits.
No BYOK keys configured for this model yet.
Share a key with the pool to earn credits for every request it serves, covering your plan cost.
Free poolnot in pool yetDonate key →
Create API key for this model05
About
Always resolves to the newest live Glm Flash model — currently GLM-5.3-Flash (z-ai/glm-5.3-flash). GLM-5.3-Flash is a native multimodal model from Z.ai (320B-A18B, MIT License, 1M-token context). It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
Released 2026-08-26
Share cards11 images










