Esc to close · ⌘K / Ctrl-K opens search anywhere
The cheapest model, the snappiest model and the fastest-streaming model are usuallythree different models — and the winner flips with your workload. Pick your use case; we rank the catalog by what that workload actually feels:₹ per task, time to first token andsustained throughput, measured through the production gateway (sweep of 2026-07-02), not read off datasheets.
Agent loops make dozens of short, tool-calling turns per task — time to first token dominates how fast the agent feels, throughput matters for big diffs, and cost adds up across the loop. Reasoning support is required.
| # | Model | Best route | ₹/Mtok (blended) | First token | tok/s |
|---|---|---|---|---|---|
| 1 | gpt-oss-120bopen weights🇮🇳 India route | fireworks | ₹39 | 1.2s | 469 |
| 2 | qwen3-32bopen weights | groq | ₹22 | 2.1s | 382 |
| 3 | gpt-oss-20bopen weights🇮🇳 India route | krutrim | ₹24 | 4.7s | 606 |
| 4 | glm-4.7open weights | openrouterBYOK | ₹136 | 1.9s | 68 |
| 5 | gemini-2.5-flash | openrouter | ₹187 | 2.2s | 132 |
| 6 | glm-4.7-flashopen weightsprice-ranked | zhipuBYOK | free | — | — |
| 7 | glm-4.5-flashopen weightsprice-ranked | zhipuBYOK | free | — | — |
| 8 | qwen3.5-9bopen weights🇮🇳 India routeprice-ranked | krutrim | ₹6.8 | — | — |
Blended ₹/Mtok = cheapest route at a 1:3 input:output token mix (generation-dominant). First token and tok/s are the best measured route per model. Rankings are a weighted percentile score per lens — details in the methodology below.
Click a metric column to sort by that lens. The same model often wins one and loses another.
| Model | Routes | ₹/Mtok ↕ | First token ↕ | tok/s ↕ |
|---|---|---|---|---|
| glm-4.7-flashreasoning | zhipuopenrouter | free | — | — |
| glm-4.5-flashreasoning | zhipu | free | — | — |
| llama-3.1-8b-instruct | bharatrouter 🇮🇳openroutergroqfireworks | ₹2.8 | 428ms groq | 548 |
| qwen2.5-coder-7b | bharatrouter 🇮🇳 | ₹3.5 | — | — |
| qwen2.5-7b-instruct | bharatrouter 🇮🇳 | ₹3.5 | — | — |
| qwen2.5-vl-7b-instruct | bharatrouter 🇮🇳 | ₹5.3 | — | — |
| gemma-4-e4b-it | krutrim 🇮🇳 | ₹6.8 | — | — |
| qwen3.5-9breasoning | krutrim 🇮🇳 | ₹6.8 | — | — |
| glm-4-32b-0414-128k | zhipu | ₹10 | — | — |
| command-r7b | cohere | ₹12 | — | — |
| mistral-small | mistral | ₹16 | — | — |
| qwen3-32breasoning | groqopenrouterfireworks | ₹22 | 2.1s groq | 382 |
| gemma-4-26b-a4b-it | krutrim 🇮🇳 | ₹23 | — | — |
| qwen3.6-35b-a3breasoning | krutrim 🇮🇳 | ₹23 | — | — |
| gpt-oss-20breasoning | krutrim 🇮🇳groqfireworks | ₹24 | 4.7s krutrim | 606 |
| deepseek-v4-flashreasoning | deepseek | ₹24 | — | — |
| devstral-small | mistral | ₹24 | — | — |
| llama-3.3-70b | groqopenrouterfireworks | ₹26 | 2.0s openrouter | 218 |
| gemma-4-31b-it | krutrim 🇮🇳 | ₹27 | — | — |
| llama-4-scout | groqfireworks | ₹28 | — | — |
| glm-4.6v-flashxreasoning | zhipu | ₹30 | — | — |
| glm-4.7-flashxreasoning | zhipu | ₹30 | — | — |
| gemini-2.5-flash-litereasoning | gemini | ₹31 | — | — |
| gpt-oss-120breasoning | krutrim 🇮🇳groqbasetenfireworks | ₹39 | 1.2s fireworks | 469 |
| command-r | cohere | ₹47 | — | — |
| gpt-4o-mini | openai | ₹47 | — | — |
| gpt-4o-mini-search-preview | openai | ₹47 | — | — |
| nemotron-superreasoning | baseten | ₹61 | — | — |
| deepseek-v3 | openrouter | ₹63 | 16.3s openrouter | 38 |
| glm-4.5-airreasoning | zhipuopenrouterfireworks | ₹65 | — | — |
| glm-4.6vreasoning | zhipuopenrouter | ₹72 | — | — |
| codestral | mistral | ₹72 | — | — |
| deepseek-v4-proreasoning | deepseekbasetenfireworks | ₹74 | — | — |
| sonar | perplexity | ₹96 | — | — |
| grok-code-fast-1reasoning | xai | ₹113 | — | — |
| gemini-3.1-flash-litereasoning | gemini | ₹114 | — | — |
| mistral-largereasoning | mistral | ₹120 | — | — |
| glm-4.6reasoning | zhipuopenrouter | ₹136 | — | — |
| glm-4.7reasoning | basetenzhipuopenrouter | ₹136 | 1.9s openrouter | 68 |
| qwen3.6-27breasoning | krutrim 🇮🇳groq | ₹147 | — | — |
| gpt-5-minireasoning | openai | ₹150 | 10.2s openai | 63 |
| glm-5reasoning | basetenzhipuopenrouter | ₹153 | 12.2s zhipu | 53 |
| kimi-k2.5reasoning | basetenopenrouter | ₹155 | — | — |
| grok-build-0.1reasoning | xaiopenrouter | ₹168 | — | — |
| glm-4.5reasoning | zhipuopenrouter | ₹173 | — | — |
| kimi-k2 | groqopenrouter | ₹180 | 688ms groq | 178 |
| nemotron-ultrareasoning | baseten | ₹187 | — | — |
| gemini-2.5-flashreasoning | geminiopenrouter | ₹187 | 2.2s openrouter | 132 |
| gemini-3.5-flash-litereasoning | gemini | ₹187 | — | — |
| grok-4.3reasoning | xaiopenrouter | ₹210 | — | — |
| grok-4.20reasoning | xaiopenrouter | ₹210 | — | — |
| glm-5.2reasoning | basetenzhipuopenrouterfireworks | ₹238 | 5.7s zhipu | 68 |
| kimi-k2.6reasoning | moonshotbasetenopenrouter | ₹244 | — | — |
| kimi-k2.7-codereasoning | moonshotbasetenopenrouter | ₹270 | 19.6s openrouter | 39 |
| glm-5-turboreasoning | zhipu | ₹317 | — | — |
| glm-5v-turboreasoning | zhipuopenrouter | ₹317 | — | — |
| glm-5.1reasoning | basetenzhipuopenrouter | ₹333 | — | — |
| glm-4.5-airxreasoning | zhipu | ₹351 | — | — |
| claude-haiku-4.5reasoning | anthropicopenrouter | ₹384 | 16.8s openrouter | 102 |
| gpt-5.6-lunareasoning | openai | ₹456 | — | — |
| pixtral-large | mistral | ₹480 | — | — |
| grok-4.5reasoning | xai | ₹480 | — | — |
| gemini-3.6-flashreasoning | gemini | ₹576 | — | — |
| mistral-medium-3.5 | mistral | ₹576 | — | — |
| kimi-k2.7-code-highspeedreasoning | moonshot | ₹622 | — | — |
| sonar-reasoning-proreasoning | perplexity | ₹624 | — | — |
| sonar-deep-researchreasoning | perplexity | ₹624 | — | — |
| gemini-3.5-flashreasoning | gemini | ₹684 | — | — |
| glm-4.5-xreasoning | zhipu | ₹693 | — | — |
| gemini-2.5-proreasoning | geminiopenrouter | ₹750 | — | — |
| gpt-5reasoning | openai | ₹750 | — | — |
| command-a | cohere | ₹780 | — | — |
| gpt-4o-search-preview | openai | ₹780 | — | — |
| gpt-5.6-terrareasoning | openai | ₹1140 | — | — |
| kimi-k3reasoning | moonshotopenrouterbaseten | ₹1152 | — | — |
| sonar-pro | perplexity | ₹1152 | — | — |
| claude-sonnet-5reasoning | anthropicopenrouter | ₹1152 | — | — |
| claude-opus-4.8reasoning | anthropicopenrouter | ₹1920 | — | — |
| claude-opus-5reasoning | anthropic | ₹1920 | — | — |
| gpt-5.6-solreasoning | openai | ₹2280 | — | — |
| claude-fable-5reasoning | anthropicopenrouter | ₹3840 | — | — |
Perf numbers are medians from a multi-run streamed sweep through the production gateway on 2026-07-02: every (model × provider) route gets the same ~300-token prompt, rounds interleaved across hosts so no provider owns a time-of-day advantage. First token counts reasoning tokens (it's what you see). tok/s is the post-first-token decode rate. Models not yet swept show “—” and rank on price with a neutral perf score. Routes we could not measure are listed openly in theAPI response(9 skipped this sweep), never silently dropped. Numbers refresh with each sweep; live per-route health is on /models.
Agents get this same chooser as JSON:GET /v1/compare/models — rankings, per-route pricing, measured perf and live failure rates, no auth required.