Skip to content

Google

17 modelsModel creator

Models published by Google, available through the AnyRouter API. Each can route across multiple upstream providers for availability and price.

$0.10 in
$0.40 out
provider list · BYOK
1M

Google's lightest Gemini 2.5 Flash variant, designed for high-throughput, cost-efficient inference with a 1M-token context window and fast response times.

$0.30 in
$2.50 out
provider list · BYOK
1M

Google's fast and cost-efficient Gemini model with strong reasoning capabilities, designed for high-throughput tasks.

$1.25 in
$10.00 out
provider list · BYOK
1M

Google's advanced Gemini 2.5 Pro model offering deep reasoning, long context, and multimodal capabilities for complex research, coding, and analytical workflows.

1MFree

Google's Gemini 3 Flash model combining fast inference with strong reasoning, coding, and multimodal understanding across a 1M-token context window.

$0.25 in
$1.50 out
provider list · BYOK
1M

Google's lightest and most cost-efficient Gemini model for high-throughput tasks.

Terms5
TTFT 180ms250 tok/s
$2.00 in
$12.00 out
provider list · BYOK
1M

Google's most intelligent Gemini model with improved reasoning, a medium thinking level, and a 1M token context window.

Terms3
TTFT 350ms120 tok/s
$0.30 in
$2.50 out
provider list · BYOK
1M

Google's Gemini 3.5 Flash Lite is a cost-efficient, high-throughput multimodal model available as a partner-hosted SKU on Cloudflare Workers AI. Optimized for agentic workflows, complex coding, and long-horizon tasks across a 1M-token context window.

$1.50 in
$9.00 out
provider list · BYOK
1M

Google's high-performance multimodal AI model designed for agentic workflows, complex coding, and long-horizon tasks.

Terms7
TTFT 120ms218 tok/s
$0.75 in
$3.75 out
provider list · BYOK
1M

Google's Gemini 3.6 Flash is a frontier-level multimodal model available as a partner-hosted SKU on Cloudflare Workers AI. Combines fast inference with strong reasoning, coding, and multimodal understanding across a 1M-token context window.

$0.38 in
$1.88 out
provider list · BYOK
1M

Google's Gemini 3.7 Flash is a multimodal model for fast agentic workflows, coding, and complex multi-step reasoning, with a 1,048,576-token context window.

$0.38 in
$1.88 out
provider list · BYOK
1M

Google's Gemini 3.8 Flash is a multimodal model for fast agentic workflows, coding, and complex multi-step reasoning, with a 1,048,576-token context window.

$0.20 in
$0 out
provider list · BYOK
8K

Gemini Embedding 2 is Google's first multimodal embedding model, mapping text, images, video, audio, and PDFs into one 3,072-dimension vector space for cross-modal semantic search, document retrieval, and recommendations over 100+ languages. Upstream accepts 8,192 input tokens and MRL-truncates to 128–3,072 dimensions. AnyRouter exposes the text path only — the OpenAI-compatible /embeddings body is a text `input`, so the extra upstream modalities are deliberately not declared here.

$0.05 in
$0.10 out
provider list · BYOK
131K

Gemma 3 4B IT is Google's instruction-tuned 4B open model — multimodal (text and image input), with a 128K context window and multilingual support in over 140 languages.

33KZDRFree

Google's Gemma 3n E2B, run on a user's own paired phone and relayed to AnyRouter over an outbound WebSocket (no Google API key, no server-side compute). Ships $0 in/out — the user supplies their own hardware. Text in, text out.

$0.10 in
$0.30 out
provider list · BYOK
262K

Gemma 4 is Google's most intelligent family of open models, built from Gemini 3 research to maximize intelligence-per-parameter.

Terms6
TTFT 250ms100 tok/s
$0.10 in
$0.34 out
provider list · BYOK
262K

Gemma 4 31B is Google's dense frontier-level open reasoning model, optimized for complex reasoning, agentic workflows, coding, and multimodal understanding.

256K

Gemma 4 26B A4B is a Mixture-of-Experts model from Google DeepMind with 26B total parameters and only 4B active per token, offering fast inference at high quality. It handles text, image, and video input, supports 256K context, function calling, and reasoning with configurable thinking modes.