Models

Every model routable through LokaRouter, with pricing, throughput, and latency at a glance.

ModelModalityContextInput / MTokOutput / MTokThroughputLatencyProviders
openai/gpt-5.2Current flagship for reasoning, coding, and agentic work.
TextImage
400K$1.75$14.00118 tok/s620 ms6
openai/gpt-5.2-proExtended-compute GPT-5.2 for the hardest problems.
TextImage
400K$15.00$120.0076 tok/s940 ms4
anthropic/claude-opus-4.6Deepest reasoning Claude with adaptive thinking budgets.
TextImage
200K$5.00$25.0092 tok/s760 ms4
anthropic/claude-sonnet-4.6Balanced production flagship with stronger code gains.
TextImage
200K$3.00$15.00104 tok/s620 ms5
google/gemini-3-proCurrent Gemini flagship with deep-think mode.
TextImageAudioVideo
1M$2.00$12.00112 tok/s590 ms5
openai/gpt-5.1Previous flagship with adaptive reasoning and warm tuning.
TextImage
400K$1.25$10.00116 tok/s600 ms6
google/gemini-3-flashSpeed flagship: Pro-class quality at Flash latency.
TextImageAudioVideo
1M$0.50$3.00178 tok/s330 ms6
openai/gpt-5.1-codex-maxLong-horizon agentic coding flagship for real repos.
TextImage
400K$1.25$10.00112 tok/s640 ms5
anthropic/claude-sonnet-4.5Proven all-round Sonnet for production workloads.
TextImage
200K$3.00$15.00102 tok/s640 ms5
openai/gpt-5-miniCost-efficient GPT-5 tier for high-volume tasks.
TextImage
400K$0.25$2.00165 tok/s480 ms5
google/gemini-2.5-flashAdaptive-thinking multimodal flash tier.
TextImageAudio
1M$0.30$2.50178 tok/s330 ms6
anthropic/claude-opus-4.5Previous Opus flagship, still favored for long agent runs.
TextImage
200K$5.00$25.0088 tok/s780 ms4
meta/llama-4.1-maverick-405bOpen MoE flagship of the Llama 4.1 family.
TextImage
1M$0.28$0.95148 tok/s380 ms7
anthropic/claude-haiku-4.5Fast, inexpensive near-flagship quality.
TextImage
200K$1.00$5.00158 tok/s410 ms5
deepseek/deepseek-v4Sparse-attention flagship for coding and agents.
Text
1M$0.35$1.40132 tok/s420 ms6
google/gemini-2.5-proMillion-token thinking model in wide production use.
TextImageAudio
1M$1.25$10.00112 tok/s590 ms5
deepseek/deepseek-r2Second-generation open reasoning model trained with RL.
Text
1M$0.55$2.20104 tok/s780 ms5
meta/llama-4.1-scout-109bEfficient multimodal MoE for broad deployment.
TextImage
1M$0.18$0.65163 tok/s350 ms8
qwen/qwen3.5-plusFlagship 1M-context multimodal tier for agentic work.
TextImage
1M$0.90$4.50124 tok/s440 ms4
openai/gpt-5.2-chatNon-reasoning tuning of GPT-5.2 for conversation and drafting.
TextImage
400K$1.75$14.00128 tok/s540 ms4
meta/llama-4.1-behemoth-2tTrillion-parameter MoE teacher model now served broadly.
TextImage
1M$0.75$1.0098 tok/s520 ms5
deepseek/deepseek-v3.2Cost flagship with sparse attention and wide adoption.
Text
164K$0.27$1.10132 tok/s420 ms6
qwen/qwen3-coder-nextOpen MoE code model with 256K context for agentic coding.
Text
262K$0.14$0.50150 tok/s360 ms5
qwen/qwen3.5-flashFast tier with strong instruction following.
TextImage
1M$0.15$0.90176 tok/s320 ms5
openai/gpt-5.1-codex-miniCompact codex tier for fast, cheap code edits.
Text
400K$0.25$2.00172 tok/s440 ms5
google/gemini-2.5-flash-liteCheapest Gemini tier for high-volume routing.
TextImage
1M$0.10$0.40205 tok/s290 ms5
openai/gpt-4.1Million-token workhorse still in wide production use.
TextImage
1M$2.00$8.00105 tok/s560 ms5
qwen/qwen3.5-397b-a17bOpen MoE flagship of the Qwen3.5 lineup.
TextImage
262K$0.40$1.40138 tok/s380 ms6
meta/llama-4-scoutWidely deployed multimodal Llama 4.
TextImage
328K$0.15$0.60163 tok/s350 ms8
x-ai/grok-4.2Latest Grok flagship with real-time knowledge hooks.
TextImage
256K$3.00$15.0094 tok/s760 ms3
qwen/qwen3-maxLargest Qwen3 tier for complex agentic work.
Text
262K$1.20$6.00118 tok/s470 ms4
openai/gpt-5-nanoSmallest GPT-5 tier for latency-sensitive workloads.
Text
400K$0.050$0.40210 tok/s350 ms4
mistralai/mistral-large-3European open-weights flagship for reasoning and code.
Text
256K$2.00$6.00110 tok/s520 ms5
google/gemini-2.5-flash-imageConversational image generation and editing.
TextImage
33K$0.30$2.5096 tok/s640 ms4
openai/gpt-4.1-miniAffordable long-context mini in current deployments.
TextImage
1M$0.40$1.60152 tok/s430 ms5
qwen/qwen3-coder-480bOpen code flagship for agentic coding.
Text
262K$0.29$1.20129 tok/s400 ms6
mistralai/codestral-25.08Latest low-latency code completion specialist.
Text
256K$0.30$0.90187 tok/s240 ms5
openai/gpt-5.1-chatConversational 5.1 tuning for support and companions.
TextImage
400K$1.25$10.00132 tok/s520 ms4
x-ai/grok-4.1Frontier Grok generation with strong agentic tool use.
TextImage
256K$3.00$15.0096 tok/s740 ms3
mistralai/mistral-medium-3.1Frontier-class quality at mid-tier pricing.
Text
128K$0.40$2.00139 tok/s420 ms5
deepseek/deepseek-v3.1-terminusStabilized V3.1 with better agent reliability.
Text
164K$0.55$1.68118 tok/s470 ms5
google/gemini-3-deep-researchAutonomous research agent producing cited reports.
TextImage
1M$2.00$12.0046 tok/s1900 ms3
qwen/qwen3-coder-plusProprietary code model with repo-scale context.
Text
1M$0.45$1.80124 tok/s420 ms4
x-ai/grok-4.1-fastSpeed tier with 2M context for agents.
TextImage
2M$0.20$0.50172 tok/s290 ms3
qwen/qwen3-vl-235bOpen multimodal flagship for document agents.
TextImageVideo
262K$0.30$1.20122 tok/s430 ms4
x-ai/grok-code-fast-1Cheap, fast coding model for agentic loops.
Text
256K$0.15$0.60190 tok/s280 ms3
anthropic/claude-3.7-sonnetHybrid-reasoning Sonnet still common in enterprises.
TextImage
200K$3.00$15.0096 tok/s660 ms4
openai/gpt-image-1.5Latest instruction-following image generation.
Image
32K$5.00$40.003 tok/s3800 ms3
mistralai/magistral-2Transparent reasoning model with visible planning.
Text
128K$2.00$5.00104 tok/s690 ms3
google/gemini-2.0-flashStable previous-gen flash still common in pipelines.
TextImageAudio
1M$0.10$0.40196 tok/s310 ms4

Sample data for demonstration — connect your workspace for live stats.