Skip to main content

OpenRouter’s router benchmarks: how to choose a route for your workload

2 min read

A model router chooses which model handles a request. OpenRouter’s new benchmark page offers a comparison framework, but a leaderboard cannot replace testing your own prompts.

OpenRouter’s router benchmarks: how to choose a route for your workload

A model router chooses which model handles a request. OpenRouter’s new benchmark page offers a comparison framework, but a leaderboard cannot replace testing your own prompts.

What the benchmark measures

OpenRouter published its model-router benchmark announcement on October 2. It describes a combined index with default weights of 60% quality, 20% time, and 20% cost. Readers can adjust those priorities. Those weights express a preference; they are not a universal definition of the best router—openRouter announcement.

Model routing also differs from selecting an infrastructure provider for a fixed model. Keep those two decisions separate when discussing costs or reliability. The announcement notes that switching models can affect cache reuse and latency.

We are not naming a winner here. A defensible ranking needs the current results, the tested model versions, and the evaluation conditions. The router benchmark page is where you can inspect those details.

Build a small evaluation set.

Collect examples from the work the application actually does. Include short requests, long context, ambiguous instructions, and cases where a wrong answer has a meaningful cost. Remove private data before sharing the set with an external service.

Write the acceptance rules before running the comparison. A support answer might need the correct policy and no invented refund promise. A structured extraction might need valid fields and exact values. These are different quality tests.

Compare the complete request path.

  • Use the same prompts and acceptance rules for each candidate.
  • Record the selected model when the service exposes it.
  • Measure total latency, including retries and routing overhead.
  • Include failed requests in the cost calculation.
  • Compare warm-cache and cold-cache behavior separately.
  • Track escalation to a stronger model and the reason for it.
  • Recheck the result after a model or router changes.

For a practical example of typed decisions, read our TypeSafe Jev guide. Our GPT-6.1 Sol migration guide covers another part of the model-selection decision.

Editorial takeaway: choose the route that meets your acceptance rules at an acceptable total cost, rather than copying a headline rank.

Leave a comment

Your email address will not be published. Required fields are marked *