Start with a decision, not a leaderboard
A new large language model (LLM) appears. Its announcement looks impressive. Your existing model sometimes fails, so replacing it feels like progress. But a better aggregate benchmark does not tell you whether the new model can extract your invoice fields, choose your support policy, or safely use your tools.
The practical question is narrower: which configuration completes our task within our quality, safety, latency and cost limits? A configuration includes the model version, prompt, tool definitions, retrieval settings and generation budget. Changing all of them at once makes it difficult to explain a gain or diagnose a regression.
This guide builds a decision process you can reproduce. The numerical lab uses explicitly fictional candidates, not measured provider performance. Current names below are discovery leads with source links, not recommendations or a ranking.
What is current, and what is actually verified?
Source check: October 5, 2026. Our model directory and comparison page use OpenRouter catalog metadata. The refreshed fallback contains 441 validated text-output entries; the count includes variants and is not a count of unique base models. The live count can change.
| Discovery lead | Evidence available | What it does not establish |
|---|---|---|
| OpenAI GPT-6.1 Sol | OpenRouter lists the model, context and token prices | Performance or provider-direct availability in your region |
| Anthropic Claude Sonnet 5.5 | OpenRouter lists standard and batch variants separately | Batch pricing for interactive requests |
| Google Gemini 4 Argon | Google's September 30 announcement describes coding, multimodal and enterprise evaluations | Independent replication or success on your task |
| inclusionAI Ling 3.1 Flash | An October 2 catalog-added timestamp | The model's original release date or unlimited free use |
Google reports results on several evaluations in its Argon announcement. Treat those as provider-reported evidence. Before adopting any candidate, verify the exact endpoint, current pricing, terms, supported features and your region's availability. Keep the evidence date in your decision record.
The whole system in one picture
Read this from top to bottom. The model is only one box. Your verifier, test cases and rollback mechanism determine whether a promising demonstration becomes a dependable service. The canary is a limited rollout to eligible traffic, not permission to experiment with sensitive actions.
For a support assistant, create separate cases for factual answers, unclear requests, policy exceptions, missing documents and prohibited actions. Include the correct answer and the evidence required to support it. For a coding assistant, require executable tests and inspect unintended file changes. For retrieval-augmented generation (RAG), distinguish retrieval failure from an unsupported answer despite adequate evidence.
Separate hard gates from preferences
A weighted average can conceal a serious failure. Imagine a configuration with excellent speed and low cost that exposes another customer's record in one test. Giving privacy a small penalty inside a composite score is the wrong decision rule.
These gates are an example policy, not universal thresholds. Define a critical violation before running the experiment. Evaluate severe risks independently, including unauthorized tool calls, unsupported financial changes and disclosure of private records. Zero observed violations in a small dataset does not prove zero real-world risk.
After hard gates, compare costs and operational preferences. The 95th percentile (p95) latency is the duration below which roughly 95% of measured requests fall. It needs enough representative observations; a p95 computed from a handful of requests is unstable.
Calculate the cost of a completed task
For ordinary uncached text token billing:
Request cost = input tokens × input price / 1,000,000 + output tokens × output price / 1,000,000.
For a workflow, sum every model call, retry and tool charge. If a provider bills additional reasoning, cache or media categories, account for those separately using that provider's usage fields. Do not assume displayed text length equals billed output.
Observed cost per successful task = total measured spend / number of successful tasks.
This is a cohort accounting metric. It is not a guarantee about the cost of eventually solving an individual task. Dividing an average request cost by a success rate assumes that the measured average cost represents both successes and failures. Actual event-level totals are better.
The dotted edges mean measurement, not additional model calls. Track end-to-end time separately from first-token time. A fast first sentence can hide a long tool sequence; a low token rate can hide expensive repairs.
Runnable Python: choose an eligible candidate
Run this with Python 3.10 or later. It needs no dependencies, account, API key or network connection. All prices, success rates and latencies are synthetic teaching inputs.
from dataclasses import dataclass
@dataclass(frozen=True)
class Candidate:
name: str
input_price: float
output_price: float
success_rate: float
p95_seconds: float
critical_violations: int
def request_cost(candidate, input_tokens=4000, output_tokens=800):
return (input_tokens * candidate.input_price
+ output_tokens * candidate.output_price) / 1_000_000
candidates = [
Candidate("fast", 0.5, 2.0, 0.65, 2.0, 0),
Candidate("balanced", 1.0, 4.0, 0.90, 5.0, 0),
Candidate("deep", 2.0, 10.0, 0.96, 12.0, 0),
Candidate("unsafe", 0.2, 0.5, 0.98, 1.0, 1),
]
eligible = [
c for c in candidates
if c.critical_violations == 0
and c.success_rate >= 0.90
and c.p95_seconds <= 8.0
]
if not eligible:
raise RuntimeError("No candidate meets the acceptance policy")
winner = min(eligible, key=lambda c: request_cost(c) / c.success_rate)
for c in candidates:
print(c.name, f"request=${request_cost(c):.4f}",
f"per_success=${request_cost(c) / c.success_rate:.4f}",
"eligible" if c in eligible else "rejected")
print("Selected:", winner.name)
assert winner.name == "balanced"
assert abs(request_cost(winner) - 0.0072) < 1e-12
Expected selection: balanced, with $0.0072 per request and $0.0080 per successful task under the simplified assumptions. The fast candidate misses the quality gate. The deep candidate misses the latency gate. The unsafe candidate is rejected despite its attractive average score.
Now change the latency budget from 8 to 15 seconds. Both balanced and deep become eligible, but balanced still has lower cost per successful task. That is a useful exercise: eligibility and optimization are different stages. If your workload values the additional six percentage points of success enough, define that preference explicitly rather than silently changing the metric.
Build a real comparison without leaking the answers
Use a development set to revise prompts and a separate held-out set for the decision. Freeze case identifiers and expected outputs before the final comparison. Reusing the same cases after every prompt adjustment gradually turns evaluation into training.
Run candidates on the same case distribution. Keep tool permissions, retrieval documents, timeouts and retry budgets fixed. Randomize execution order to reduce the effect of provider-load changes. Repeat stochastic tasks from clean state and report results by difficulty and task family, not only overall.
For a binary success measure, include numerator, denominator and uncertainty. A result of 18/20 and a result of 900/1,000 can have the same observed rate but very different precision. Compare paired outcomes: which exact cases improved and which regressed? For workflow tasks, inspect policy compliance and final state even when the answer reads well.
Long context is capacity, not a comprehension guarantee
A million-token advertised context window is an input limit, not proof that a model will reliably use every detail. More text also increases cost, latency and the chance of contradictory evidence. Start with the smallest evidence set that supports the answer; expand it when evaluation shows missing context.
Measure retrieval coverage, citation correctness and abstention separately. Add cases where the relevant fact is near the beginning, middle and end, plus distracting documents that are similar but wrong. If a configuration fails, do not immediately buy more context: verify document quality, chunk boundaries and access filters first.
Your rollout checklist
Record the exact model endpoint and version, prompt hash, case-set version, evaluation date, region, price source and billed token categories. Keep failure traces with sensitive content minimized. Distinguish expected abstention, infrastructure failure and an incorrect completed answer.
Choose a rollback trigger before rollout: for example, a verified critical violation, a tail-latency breach or a sustained increase in task failure. Keep the previous configuration deployable. A provider changing a model alias, rate limit or tool behavior should trigger revalidation, not an automatic claim that the platform has improved.
Use new releases to refresh your shortlist. Use your own controlled task evidence to choose the configuration.
Sources and next steps
- •OpenRouter models API: aggregator catalog metadata, checked October 5, 2026.
- •GPT-6.1 Sol listing and Claude Sonnet 5.5 listing: verify current endpoint details before use.
- •Google's Gemini 4 Argon announcement: provider-reported evidence, September 30, 2026.
- •Microsoft and Hugging Face on ThinkingBox: why final state and repeatability matter, October 3, 2026.
Continue with our state-based agent evaluation lab, browse recent research, or compare current catalog entries. The lab below those ideas is intentionally small enough to understand before you connect a paid model.
