Inference Routing Is Architecture: Matching Each Request to the Cheapest Model That Works

·11 min read·Evergreen Tools Team

If 2026 needed one sentence about enterprise AI cost, HPCwire's August analysis supplied it: the story is not about training budgets or which frontier model tops the latest benchmark. It is about inference, and specifically about routing, meaning matching each request to the cheapest model that can handle it and tuning the threshold at which you spend more. The same piece has a line worth pinning to the wall: the cost of an AI system is set by its architecture, not by the model on the invoice. That is especially visible in 2026 pricing, where input prices span roughly two orders of magnitude across tiers. Treat routing as a first-class part of your infrastructure and measure it like everything else, and the bill stops being a mystery and becomes a design decision.

1. Your Bill Is a Chain of Routing Decisions

The same question might cost cents on a cheap model and tens of times more on a frontier one. If every request goes to the most expensive model available, no saving is happening, because nobody ever asked whether the request needed that model. HPCwire's frame puts routing at the center: match requests to model capability, then tune the threshold that decides when to escalate. Analyses from CloudZero and Wavect point to the same drivers, only finer-grained: context length, cache hits, retries, output length, all of which the routing layer is uniquely positioned to influence. The conclusion is plain. To control cost, make every request go to the model it deserves.

// Route to the cheapest model that can actually do the job, then escalate
// only when a cheap attempt fails a check you trust.
type Tier = { model: string; usdPerMTokIn: number; usdPerMTokOut: number };

const TIERS: Tier[] = [
  { model: "budget-flash-lite", usdPerMTokIn: 0.10, usdPerMTokOut: 0.40 },
  { model: "mid-sonnet",       usdPerMTokIn: 2.00, usdPerMTokOut: 10.00 },
  { model: "frontier-fable",   usdPerMTokIn: 10.00, usdPerMTokOut: 50.00 },
];

export async function answer(prompt: string) {
  for (const tier of TIERS) {
    const out = await call(tier.model, prompt);
    const verdict = await check(out);
    metrics.record(tier.model, verdict, out);
    if (verdict.passed) return { text: out, tier: tier.model };
  }
  throw new Error("no tier passed checks");
}
Routing requests across model tiers

Cost is set by architecture, not by the model

2. Two Tables Explain Everything: Price and Quality

The first table is price. In published 2026 pricing, cheap tiers start at cents per million input tokens while frontier tiers sit in the ten-dollars-per-million range, so input prices span roughly two orders of magnitude and output spreads are just as wide. The second table is quality. Cheap models are genuinely sufficient for some tasks and genuinely fail on others. Routing aligns the two tables: send the majority of requests to the cheap tier and escalate the minority that need it. The classic mistake is reading price without quality, because a cheap request that keeps failing gets more expensive through retries. The goal is not the lowest unit price; it is the lowest cost per success.

// Cost per request, not cost per month. If you cannot attribute spend to a
// request, you cannot route it - you can only hope.
function charge(
  req: { id: string },
  usage: { input: number; output: number },
  tier: Tier,
) {
  const usd = (usage.input / 1e6) * tier.usdPerMTokIn
            + (usage.output / 1e6) * tier.usdPerMTokOut;
  db.requests.update(req.id, { $set: { model: tier.model, usd } });
  return usd;
}

3. Cascades and Thresholds: Making 'Good Enough' Decidable

The cascade is the most common form of routing: try the cheapest model first, accept its output if it passes checks you trust, and escalate otherwise. Everything hinges on those checks. Too loose and errors slip through; too strict and nearly every request escalates, making the cheap tier decorative. Good checks are hard, task-specific constraints: does the output match the schema, does it leak PII, is it grounded in the provided context, is it within length bounds. One prerequisite is often forgotten: re-run your eval set whenever a model version changes, because a threshold tuned to one model's failure modes does not transfer cleanly to another's.

// The cheapest token is the one you never send. Cache exact and fuzzy hits.
const key = hash(normalize(prompt));
const hit = await cache.get(key);
if (hit) return { text: hit, tier: "cache" };

// Normalize before hashing: whitespace and casing are formatting, not intent.
function normalize(s: string) {
  return s.trim().replace(/\s+/g, " ").toLowerCase();
}
Spend and quality by tier

One row per tier: volume, spend, pass rate

4. Measure Before You Route: Three Numbers You Need

You cannot route what you do not measure. At minimum, answer three questions. First, what did each request cost, attributed per request rather than summed per month. Second, what is the pass rate per tier, because a cheap tier with a low pass rate is not saving money, it is spending it twice. Third, what is your cache and dedupe hit rate, since the cheapest token is always the one you never send. Wire those up and routing stops being a feeling and becomes data. Watch context length too, since long contexts are expensive and much of that length is accumulated conversation rather than intent; compressing it often beats switching models.

// A router is only as good as its checks. Score every cheap attempt, and
// re-run the eval set whenever a model version changes underneath you.
const CHECKS = [schemaValid, noPii, groundedInContext, withinLength];

function passesAll(out: string) {
  return CHECKS.every((c) => c(out));
}
// A threshold tuned for one model's failure modes will not transfer to
// another model's. Treat routing like a model swap with a regression suite.

5. In Practice: From Cascade to Report

The first snippet is a cascade router: try tiers from cheapest to most capable, return as soon as one passes your checks, and fail loudly only when none does. The second attributes cost to each request by converting tokens and unit prices into dollars. The third is caching with normalization, because whitespace and casing are formatting rather than intent, and the cache key should reflect that. The fourth fixes the quality gate into a list of functions and insists on re-running evals when a model version moves. The fifth exports a per-tier report of volume, spend, and pass rate, which is the input to threshold tuning.

// Export spend and quality by tier on a schedule. Routing is an experiment,
// so log the outcome instead of trusting the intuition that started it.
const rows = await db.requests.aggregate([
  { $group: {
      _id: "$model",
      usd: { $sum: "$usd" },
      n: { $sum: 1 },
      passRate: { $avg: "$passed" },
  } },
]);

await exportCsv("routing-report.csv", rows);
Cascading and caching

The cheapest token is the one you never send

6. A Checklist: Run Routing Like a Product

Five checks. First, can you compute the cost of a single request? Second, do you know the pass rate of each tier? Third, what is your cache and dedupe hit rate? Fourth, do model version changes trigger a re-run of the eval set? Fifth, do you export a per-tier spend and quality report on a schedule? Of these, the second is quietly the most important. A cheap tier that passes half the time does not save money; it spends it twice, once on the failed attempt and once on the escalation. The real goal of routing was never the cheapest model. It is the fewest dollars for a result you can trust.

📌 Frequently Asked Questions

Why is 2026 AI cost an inference problem rather than a training one?

Per HPCwire's analysis, enterprise AI cost is driven mainly by inference and routing rather than training budgets or which frontier model leads a benchmark. It states that the cost of an AI system is set by its architecture, not by the model on the invoice, and recommends treating routing as first-class infrastructure to measure and manage.

How wide is the price gap across model tiers?

In published 2026 pricing, cheap tiers can start at fractions of a dollar per million input tokens while frontier tiers reach roughly ten dollars per million, so input prices span about two orders of magnitude, with comparable spreads on output, which is why routing moves the bill so much.

What is a cascade router?

A cascade calls the cheapest model first, accepts its output if it passes checks you trust, and escalates to a more capable tier when it does not. Its effectiveness depends entirely on check quality: too loose lets errors through, too strict escalates almost everything and makes the cheap tier pointless.

Which numbers must routing measure?

At least three: cost per request attributed per request rather than summed per month, pass rate per tier because a low-pass cheap tier spends money twice, and cache and dedupe hit rate because the cheapest token is the one never sent.

What should I watch when a model version changes?

Re-run the eval set. A threshold tuned to one model's failure modes does not transfer cleanly to another, so a model swap is effectively a regression test, and the routing layer should treat it that way.