⚡ New — Kimi K3 is live: bring your own Moonshot key →

Benchmarking Sangam — measured, not marketed.

Consensus routers are easy to hype and hard to compare fairly. Here's how we put Sangam up against OpenRouter Fusion — with a real, money-spending benchmark harness, ablind LLM judge, and a cost model whose assumptions are all on the table. We don't quote a number we haven't measured. Where we only have an estimate, we say so.

Try it yourselfSee the methodology

Honest up front: open-sangam is an open-weight, BYOK panel — transparent and reproducible, running on your own provider keys — and it sits below Fusion on raw frontier quality by design; sangam uses a same-class frontier panel and should land roughly level with Fusion. We don't fight Fusion on the ceiling — we win on the axes it can't follow: openness, reproducibility, cost, and BYOK.

The three engines

Every run puts the same prompts through three consensus engines. Two are ours; one is the router we're measuring against. The cost column is a rate-card estimate(assumptions below) — the harness replaces it with exact billed cost.

EnginePanelResidencyOpennessEst. ₹/req
bharatrouter/open-sangamheroopen-weight · BYOK-ready
Qwen2.5 · Llama 3.1 · Qwen2.5-VL — open-weight, on your own keys (BYOK)
🇮🇳 India available✓ open-weightyour open-weight rate
bharatrouter/sangamfull frontier
gpt-5 · gemini-2.5-flash · claude-haiku-4.5 → gpt-5-mini
🌐 globalclosed frontierbetween — frontier panel, INR catalog rates
openrouter/fusionfull frontier
Opus · GPT · Gemini + a fuser
🌐 globalclosed frontier≈ ₹1.5–2.0 / req

bharatrouter/sangam is the apples-to-apples Fusion competitor: a same-class frontier panel, but billed at our INR catalog rates — on your own keys (BYOK).bharatrouter/open-sangam is the open-weight, BYOK route — transparent and reproducible, with India residency available via data_policy — the one to try first.

Where each engine wins

Raw benchmark quality is one axis, not the only one. On a hard reasoning leaderboard, Fusion's frontier panel is hard to top — so we say so plainly. But openness, reproducibility, cost, BYOK, and VPC-deployability are axes a closed frontier router can't follow. Pick the engine your constraints demand.

 open-sangamsangamFusion
India data-residency available (via data_policy)✓ YesNo — frontier panelNo — foreign-deployed providers
Open-weight panel (self-hostable)✓ YesNoNo
BYOK — bring your own keys✓ Yes✓ YesNo
Foreign-deployed frontier model in the loop✓ NoneYes — by designYes — that's its panel
Reproducible (panel composition published)✓ Published✓ PublishedNot disclosed
Runs inside your own VPC (engine-mode)On the roadmap — open weights make it possibleNo — closed panelNo — closed weights
Est. cost / requestyour open-weight rate (BYOK)between — INR catalog rates≈ ₹1.5–2.0
Raw frontier-quality ceilingLower — open-weight panel, by design≈ Fusion — same-class panelHighest — frontier panel

Read the last two rows together with the rest. If the question is "best possible answer, cost and residency no object," Fusion is a fair pick — and so is our own sangam, at INR rates. If it's "best answer on open weights I can audit, reproduce, and eventually run myself — on my own provider keys (BYOK), with India residency when I need it," that'sopen-sangam.

Cost model

Rate-card estimateNot a measured result. Calculated from published per-token rates under the stated assumptions — the harness replaces every figure here with OpenRouter's exact billed cost.

Under the same workload, open-sangam landsroughly an order of magnitude cheaper per request than Fusion.

  • open-sangam — your own open-weight rate (BYOK) — the panel is all open-weight, so you run it on your own provider keys (e.g. Groq) and pay your provider's open-weight per-token rate, far below frontier rates.
  • Fusion ≈ ₹1.5–2.0 / request — sum of frontier completions at frontier-provider rates, converted at ~₹86/$.
  • bharatrouter/sangam sits between — a frontier panel like Fusion's, but billed at our INR catalog rates rather than raw frontier-provider rates.

Assumptions: ≈400 input + ≈400 output tokens per panel member; USD→INR at ~₹86/$; published per-token rates at time of writing. Token counts, model mix, and FX move the figure — these are order-of-magnitude, not a quote.

When you run the harness, the cost columns stop being estimates: it records the real billed cost OpenRouter reports for each call, so you see what these engines actually charged onyour prompts.

Methodology — the harness

The benchmark isn't a slide; it's runnable code in the repo:bench/sangam-vs-fusion.mjs. It runs the same prompt set through all three engines, has a blind LLM judge pick the best answer, and records latency, tokens, and measured cost. It refuses to run without keys — because it spends real money on real provider calls, there's no fake-data mode.

📋
Shared prompt set

Same prompts, every engine

A fixed set spanning reasoning, India-tax, Hindi, code, and summarization — so each engine answers identical work, not a curated home-field selection.

🙈
Blind best-of-3 judge

The judge can't see who's who

An LLM judge reads the three answers in randomized, anonymized order and picks the best. It never learns which engine produced which answer — no house bias.

🔁
ROUNDS for stability

Repeat, don't trust one shot

Each prompt runs over multiple ROUNDS so a single lucky or unlucky generation doesn't decide the outcome. One judge call is never the verdict — results are aggregated.

📊
Measured, recorded

Latency · tokens · real cost

For every call it logs wall-clock latency, token usage, and the cost the provider actually billed — the numbers that replace this page's estimates.

Run it yourself — supply your own keys (it will spend real money):

BR_KEY=… OPENROUTER_API_KEY=… node bench/sangam-vs-fusion.mjs

One judge model is not the last word — a single judge has its own preferences. Treat the harness as a method for getting comparable numbers under controlled conditions, and read its output as a distribution across prompts and rounds, not a single score.

Measured results — first run

A first pass through the harness — 6 prompts (reasoning, India-tax, Hindi, code, summarization), 1 round, blind best-of-3 judge (gpt-5). A small, directional sample, not a leaderboard — but real numbers, nothing fabricated.

EngineMedian latencyCost / requestBlind quality
bharatrouter/open-sangam~17 syour open-weight rate (BYOK — your provider's open-weight per-token rate)tie 6/6
bharatrouter/sangam~22 s≈ ₹8–13 (est. — frontier panel, not free)tie 6/6
OpenRouter Fusion~55 s≈ ₹13 (measured — $0.89 / 6 calls)tie 6/6

Three honest takeaways: open-sangam was the fastest (~3× quicker than Fusion), ran at your own open-weight rate — its panel is all open-weight, so you run it on your own provider keys (BYOK) and pay your provider's open-weight rate, far below Fusion's ~₹13/request — and the judge scored every prompt a tie — open-weight consensus held its own against a frontier panel on these general questions.sangam is the frontier-panel route, so it is not free: its cost is comparable to Fusion (~₹13), just billed at our INR catalog rates.

Caveats, kept honest: one round, six prompts, a single judge — ties may partly reflect judge conservatism, and harder reasoning would likely separate the field. Latency is wall-clock with max_tokens 700. On cost: only Fusion (OpenRouter-billed) was directly metered this run; open-sangam runs on your own BYOK keys, so its cost is your provider's open-weight rate, and sangam's figure is a rate-card estimate (BR responses don't return a cost field), frontier-class and explicitly not the open-weight rate of open-sangam. The bigger axes — residency, openness, reproducibility — are categorical wins above, not in this table. Re-run it yourself:node bench/sangam-vs-fusion.mjs.

Consensus uplift — same base model

The result above compares whole engines. This one isolates a narrower question: does theconsensus + verify machinery itself add accuracy — separate from just using a bigger model? To find out we pit open-sangam (3-model panel + verifier + synthesizer) against its own lead model, llama-3.1-8b-instruct, runningalone. Same base model on both sides — so any gap is the ensemble machinery, not extra parameters. Grading is objective (every task has a known correct answer, checked by answer-match) — no LLM judge, so no judge bias.8 classic reasoning & arithmetic traps × 5 rounds =40 graded answers. We ran it twice, independently, and publish both — so what you see is a range across real runs, not one lucky shot.

Runopen-sangam (panel + verify + synth)llama-3.1-8b (lead model, alone)UpliftLatency (ens / single)
Run 197.5% (39/40)87.5% (35/40)+10.0 pts~11.8s / ~6.2s
Run 295% (38/40)87.5% (35/40)+7.5 pts~11.8s / ~5.9s

Across two independent runs the ensemble scored 95.0–97.5% against the single model's stable 87.5% — a +7.5 to +10 point uplift on the same base model. That gap is consensus and verification doing their job, not a heavier model. The per-task breakdown (most recent run) shows where: the ensemble rescues traps a single small model slips on — e.g. clock-angle and speed-convert(5/5 ensemble vs 4/5 alone), and the harddigit-7 count (3/5 vs 2/5) where both still struggle — honest about the ceiling.

TaskCategoryCorrect answeropen-sangamsingle
avg-speedreasoning40 km/h (not 45)5/55/5
eggsreasoning245/55/5
widgetsreasoning5 minutes5/55/5
speed-convertquant27.78 m/s5/54/5
days-100logicWednesday5/55/5
digit-7counting203/52/5
clock-anglequant7.5°5/54/5
montyreasoningswitch — 2/35/55/5

How this compares to what the field has published

Consensus isn't ours alone — Together, OpenRouter and Sakana have all published results showing a panel beats its own members. These are different benchmarks on different baselines, so this is not a ranking — read the direction, and note who is independently reproducible.

SystemBenchmarkPanelvs best singleUpliftReproducible?
BharatRouter open-sangam8 reasoning/arithmetic traps · objective grading, no judge95.0–97.5%its own lead model, 87.5%+7.5 to +10 pts✓ harness in repo · 2 runs published
Together MoAAlpacaEval 2.0 (LC win-rate)65.1%GPT-4o, 57.5%+7.6 pts✓ open code · ICLR 2025
OpenRouter FusionDRACO (Perplexity)69.0%Fable 5 solo, 65.3%+3.7 ptsvendor-published
Sakana Fugu UltraSWE-Bench Pro · GPQA-D · LiveCodeBench73.7% · 95.5 · 93.2frontier panel membersleads pool on 10/11✗ vendor-reported, not independently reproduced

Read honestly: Fusion and Fugu lift frontier panels a few points higher (a higher ceiling than our open-weight panel, by design); Together MoA shows open models beating GPT-4o. Ours is the one that isolates the machinery — same base model both sides — and ships a runnable, objective-graded harness with two published runs. Fugu's numbers are vendor-reported and not yet independently reproduced; ours you can reproduce from the repo today. Sources:Together MoA ·OpenRouter Fusion ·Sakana Fugu.

On latency — it's provider-bound, not inherent. The ~11.8s above is the open-weight panel running on our own infra. The consensus pattern is mostly waiting on token generation, so it tracks your provider's speed: in a separate latency probe on Groq's fast inference (open-weight models, 3-model panel + a synthesis pass), the same ensemble shape completed in ~1.25s end-to-end (single ~0.42s). So on production-grade inference, consensus latency stops being the tradeoff — you keep the +10 points without the wait.

Kept honest: the ~1.25s Groq figure is a latency probe on a different panel and prompt than the 40-answer accuracy run above — it measures speed, not quality, and is not conflated with the accuracy numbers. On cost — consensus is not free:a request fans out to roughly 5 calls (a multi-model panel plus verify and synthesis passes), so it spends more tokens than a single call — just on open-weight models. On BYOK you pay your provider's open-weight rate across those calls: a real cost, but typically an order of magnitude below a frontier consensus panel like Fusion — not ₹0. Re-run it yourself: BR_KEY=… node bench/sangam-vs-single.mjs.

Try it yourselfHow Sangam works