⚡ New — Kimi K3 is live: bring your own Moonshot key →

Which model should you use?

The cheapest model, the snappiest model and the fastest-streaming model are usuallythree different models — and the winner flips with your workload. Pick your use case; we rank the catalog by what that workload actually feels:₹ per task, time to first token andsustained throughput, measured through the production gateway (sweep of 2026-07-02), not read off datasheets.

Start from your application

Agent loops make dozens of short, tool-calling turns per task — time to first token dominates how fast the agent feels, throughput matters for big diffs, and cost adds up across the loop. Reasoning support is required.

#ModelBest route₹/Mtok (blended)First tokentok/s
1gpt-oss-120bopen weights🇮🇳 India routefireworks₹391.2s469
2qwen3-32bopen weightsgroq₹222.1s382
3gpt-oss-20bopen weights🇮🇳 India routekrutrim₹244.7s606
4glm-4.7open weightsopenrouterBYOK₹1361.9s68
5gemini-2.5-flashopenrouter₹1872.2s132
6glm-4.7-flashopen weightsprice-rankedzhipuBYOKfree
7glm-4.5-flashopen weightsprice-rankedzhipuBYOKfree
8qwen3.5-9bopen weights🇮🇳 India routeprice-rankedkrutrim₹6.8

Blended ₹/Mtok = cheapest route at a 1:3 input:output token mix (generation-dominant). First token and tok/s are the best measured route per model. Rankings are a weighted percentile score per lens — details in the methodology below.

The full picture — every chat model, three lenses

Click a metric column to sort by that lens. The same model often wins one and loses another.

ModelRoutes₹/Mtok ↕First token ↕tok/s ↕
glm-4.7-flashreasoningzhipuopenrouterfree
glm-4.5-flashreasoningzhipufree
llama-3.1-8b-instructbharatrouter 🇮🇳openroutergroqfireworks₹2.8428ms groq548
qwen2.5-coder-7bbharatrouter 🇮🇳₹3.5
qwen2.5-7b-instructbharatrouter 🇮🇳₹3.5
qwen2.5-vl-7b-instructbharatrouter 🇮🇳₹5.3
gemma-4-e4b-itkrutrim 🇮🇳₹6.8
qwen3.5-9breasoningkrutrim 🇮🇳₹6.8
glm-4-32b-0414-128kzhipu₹10
command-r7bcohere₹12
mistral-smallmistral₹16
qwen3-32breasoninggroqopenrouterfireworks₹222.1s groq382
gemma-4-26b-a4b-itkrutrim 🇮🇳₹23
qwen3.6-35b-a3breasoningkrutrim 🇮🇳₹23
gpt-oss-20breasoningkrutrim 🇮🇳groqfireworks₹244.7s krutrim606
deepseek-v4-flashreasoningdeepseek₹24
devstral-smallmistral₹24
llama-3.3-70bgroqopenrouterfireworks₹262.0s openrouter218
gemma-4-31b-itkrutrim 🇮🇳₹27
llama-4-scoutgroqfireworks₹28
glm-4.6v-flashxreasoningzhipu₹30
glm-4.7-flashxreasoningzhipu₹30
gemini-2.5-flash-litereasoninggemini₹31
gpt-oss-120breasoningkrutrim 🇮🇳groqbasetenfireworks₹391.2s fireworks469
command-rcohere₹47
gpt-4o-miniopenai₹47
gpt-4o-mini-search-previewopenai₹47
nemotron-superreasoningbaseten₹61
deepseek-v3openrouter₹6316.3s openrouter38
glm-4.5-airreasoningzhipuopenrouterfireworks₹65
glm-4.6vreasoningzhipuopenrouter₹72
codestralmistral₹72
deepseek-v4-proreasoningdeepseekbasetenfireworks₹74
sonarperplexity₹96
grok-code-fast-1reasoningxai₹113
gemini-3.1-flash-litereasoninggemini₹114
mistral-largereasoningmistral₹120
glm-4.6reasoningzhipuopenrouter₹136
glm-4.7reasoningbasetenzhipuopenrouter₹1361.9s openrouter68
qwen3.6-27breasoningkrutrim 🇮🇳groq₹147
gpt-5-minireasoningopenai₹15010.2s openai63
glm-5reasoningbasetenzhipuopenrouter₹15312.2s zhipu53
kimi-k2.5reasoningbasetenopenrouter₹155
grok-build-0.1reasoningxaiopenrouter₹168
glm-4.5reasoningzhipuopenrouter₹173
kimi-k2groqopenrouter₹180688ms groq178
nemotron-ultrareasoningbaseten₹187
gemini-2.5-flashreasoninggeminiopenrouter₹1872.2s openrouter132
gemini-3.5-flash-litereasoninggemini₹187
grok-4.3reasoningxaiopenrouter₹210
grok-4.20reasoningxaiopenrouter₹210
glm-5.2reasoningbasetenzhipuopenrouterfireworks₹2385.7s zhipu68
kimi-k2.6reasoningmoonshotbasetenopenrouter₹244
kimi-k2.7-codereasoningmoonshotbasetenopenrouter₹27019.6s openrouter39
glm-5-turboreasoningzhipu₹317
glm-5v-turboreasoningzhipuopenrouter₹317
glm-5.1reasoningbasetenzhipuopenrouter₹333
glm-4.5-airxreasoningzhipu₹351
claude-haiku-4.5reasoninganthropicopenrouter₹38416.8s openrouter102
gpt-5.6-lunareasoningopenai₹456
pixtral-largemistral₹480
grok-4.5reasoningxai₹480
gemini-3.6-flashreasoninggemini₹576
mistral-medium-3.5mistral₹576
kimi-k2.7-code-highspeedreasoningmoonshot₹622
sonar-reasoning-proreasoningperplexity₹624
sonar-deep-researchreasoningperplexity₹624
gemini-3.5-flashreasoninggemini₹684
glm-4.5-xreasoningzhipu₹693
gemini-2.5-proreasoninggeminiopenrouter₹750
gpt-5reasoningopenai₹750
command-acohere₹780
gpt-4o-search-previewopenai₹780
gpt-5.6-terrareasoningopenai₹1140
kimi-k3reasoningmoonshotopenrouterbaseten₹1152
sonar-properplexity₹1152
claude-sonnet-5reasoninganthropicopenrouter₹1152
claude-opus-4.8reasoninganthropicopenrouter₹1920
claude-opus-5reasoninganthropic₹1920
gpt-5.6-solreasoningopenai₹2280
claude-fable-5reasoninganthropicopenrouter₹3840

In the open — how these numbers are made

Perf numbers are medians from a multi-run streamed sweep through the production gateway on 2026-07-02: every (model × provider) route gets the same ~300-token prompt, rounds interleaved across hosts so no provider owns a time-of-day advantage. First token counts reasoning tokens (it's what you see). tok/s is the post-first-token decode rate. Models not yet swept show “—” and rank on price with a neutral perf score. Routes we could not measure are listed openly in theAPI response(9 skipped this sweep), never silently dropped. Numbers refresh with each sweep; live per-route health is on /models.

Agents get this same chooser as JSON:GET /v1/compare/models — rankings, per-route pricing, measured perf and live failure rates, no auth required.

Get a keyBrowse the catalog