3.7x on the Same Rack: Reading MLPerf Inference v6.1

·10 min read·Evergreen Tools Team

On September 16, 2026 the MLPerf Inference v6.1 results landed, and NVIDIA's Vera Rubin NVL72 made its first appearance as a preview submission: up to 3.7x the throughput of the previous-generation GB300 NVL72 on Qwen3-VL, up to 2.5x on DeepSeek-R1. Buried in the same round is the number engineering teams should read twice: GB300 NVL72 itself gained up to 1.6x over its v6.0 results on Qwen3-VL from software optimisations alone. Put those two numbers side by side and the story of this round changes shape.

The right unit for inference is tokens per second per rack

The right unit for inference is tokens per second per rack

1. Read an Inference Benchmark in Tokens per Rack

NVIDIA's framing of why this metric matters is a revenue statement: higher system performance means more tokens generated, which means higher revenue. That framing tells you how to read the submission. Do not look at a per-chip figure; look at tokens per second per rack, per watt, and per hour. A 3.7x claim is a statement about that ratio, not about a part number. For a buyer it is also the only view that maps onto an actual business: your product is billed per token, your hardware is billed per rack and per kilowatt, and the quotient between them is your cost structure. Any performance number that cannot be traced back to that quotient is a spec-sheet decoration.

# 1. Read an inference benchmark as tokens per rack, not per chip
def token_economics(result):
    # NVIDIA's framing of why this metric matters is a revenue statement:
    # "Higher system performance means more tokens generated, resulting in
    # higher revenue." Everything else in the submission is an input to this.
    rack_throughput = result["tokens_per_second"] * result["accelerators_per_rack"]
    return {
        "tokens_per_second_per_rack": rack_throughput,
        "power_per_rack_kw": result["rack_power_kw"],
        "tokens_per_joule": rack_throughput / max(result["rack_power_kw"] * 1000, 1),
        "tokens_per_dollar_hour": rack_throughput / max(result["cost_per_hour"], 0.01),
    }
# A 3.7x on one workload is a claim about this ratio, not about a chip.

2. Keep Every Multiplier Attached to a Workload

This round produced three headline multipliers: up to 3.7x on Qwen3-VL, up to 2.5x on DeepSeek-R1, and up to 30x on the SemiAnalysis AgentX benchmark in preview testing per NVIDIA. They span the offline, server and interactive scenarios, and they run on vLLM with the open-source NVIDIA Dynamo inference framework. A practical rule: a multiplier without a workload name is a marketing number; three multipliers with workload names are an engineering claim. The other distinction to hold onto is preview versus submitted. A preview submission describes a platform that is not shipping yet - you are buying performance from a roadmap, not from a rack.

// 2. The v6.1 headline, with its workload attached
const mlperfInferenceV61 = {
  date: "2026-09-16",
  submitter: "NVIDIA",
  system: "Vera Rubin NVL72",
  submissionType: "preview",          // first appearance of the platform
  versus: "GB300 NVL72",
  gains: {
    "Qwen3-VL": "up to 3.7x higher throughput",
    "DeepSeek-R1": "up to 2.5x",
    "SemiAnalysis AgentX": "up to 30x in preview testing (per NVIDIA)",
  },
  scenarios: ["offline", "server", "interactive"],
  stack: ["vLLM", "NVIDIA Dynamo"],
};
// Rule of thumb: a single multiplier without a workload name is a marketing
// number. Three multipliers with workload names are an engineering claim.
A multiplier must stay attached to its workload

A multiplier must stay attached to its workload

3. Software Is the Part You Can Buy Today

The most underrated result of the round is what GB300 NVL72 did to itself: up to 1.6x higher performance on Qwen3-VL than in v6.0, from software alone. NVIDIA lists the techniques - lower KV cache precision, additional kernel fusion, better kernels, and disaggregated serving with vLLM and NVIDIA Dynamo. Same silicon, same rack, 1.6x more throughput inside one release cycle. That yields a concrete budget ordering: before you approve a hardware refresh, price the software upgrade. It is usually cheaper, and it does not sit in a shipping queue.

# 3. Software is the part you can buy today
SOFTWARE_GAINS_V61 = {
    "platform": "GB300 NVL72",
    "workload": "Qwen3-VL",
    "gain_vs_v60": 1.6,   # up to 1.6x, from software alone
    "techniques": [
        "lower KV cache precision",
        "additional kernel fusion",
        "better kernels",
        "disaggregated serving with vLLM and NVIDIA Dynamo",
    ],
}

# The same silicon, the same rack, 1.6x more throughput in one release cycle.
# Before budgeting for a hardware refresh, price the software upgrade.

4. Keep Verified Results and Post-Deadline Previews in Separate Columns

MLPerf's comparability comes from MLCommons verification, so classify every claim before you use it. A verified submission can be compared across submitters. A preview submission indicates direction and is not commercially available. And NVIDIA notes that optimisations continued after the v6.1 submission deadline, producing further gains on GPT-OSS-120B and DLRMv3 that are not yet verified by MLCommons - directional only, never a procurement input. Mixing those three classes is the most common and most expensive mistake in performance evaluation: you end up buying a system that does not exist, justified by a number nobody verified.

# 4. Separate verified results from post-deadline previews
def classify(claim):
    if claim["verified_by"] == "MLCommons":
        return "comparable across submitters"
    if claim["stage"] == "post_submission_not_yet_verified":
        # NVIDIA reports further gains on GPT-OSS-120B and DLRMv3 after the
        # v6.1 submission deadline; those numbers have not been verified.
        return "directional only - do not put it in a procurement decision"
    if claim["stage"] == "preview":
        return "promising, but the platform is not shipping today"
    return "unclassified"

def procurement_grade(claims):
    return [c for c in claims if classify(c) == "comparable across submitters"]
Price the software upgrade before the hardware refresh

Price the software upgrade before the hardware refresh

5. A Capacity Decision Record That Survives Questions

If your team is sizing inference capacity for an agent product, write the conclusion down in this shape: the workload (long-context, Qwen3-VL class); the metric (tokens per second per rack, measured at the interaction pattern you actually serve); the measurement method (a replay of production traffic on borrowed capacity, not a vendor submission); the vendor claims used, itemised (up to 3.7x versus GB300 NVL72 on Qwen3-VL in the v6.1 preview; up to 1.6x for GB300 NVL72 from software); an explicit excluded column for unverified numbers; and a review date tied to the next submission round. The point of the record is that six months later, when someone asks why you bought what you bought, there is an answer that does not depend on memory.

{
  "capacity_decision_record": {
    "workload": "long-context Qwen3-VL style inference for an agent product",
    "metric": "tokens per second per rack, at the interaction pattern we actually serve",
    "measured_by": "our own replay of production traffic on borrowed capacity",
    "vendor_claims_used": [
      "MLPerf Inference v6.1 Vera Rubin NVL72 preview: up to 3.7x vs GB300 NVL72 on Qwen3-VL",
      "GB300 NVL72 software-only gain of up to 1.6x over v6.0"
    ],
    "claims_excluded": ["post-submission, not-yet-verified results"],
    "decision": "keep GB300 fleet for 2 quarters; adopt the v6.1 software stack first",
    "review_due": "next MLPerf submission round"
  }
}

📌 Frequently Asked Questions

What exactly does 3.7x refer to?

Per NVIDIA's announcement: in the MLPerf Inference v6.1 preview submission, Vera Rubin NVL72 delivered up to 3.7x better throughput than GB300 NVL72 on Qwen3-VL across the offline, server and interactive scenarios, using vLLM with the NVIDIA Dynamo inference framework.

Why does the 1.6x software gain matter more?

Because it applies to hardware you may already own. GB300 NVL72 improved up to 1.6x over its v6.0 results on Qwen3-VL through lower KV cache precision, additional kernel fusion, better kernels and disaggregated serving with vLLM and Dynamo. Same silicon, more throughput - the best return available in the round.

What is the difference between a preview and a submitted result?

Submitted results are verified by MLCommons and comparable across submitters. A preview submission signals direction for a platform still in preparation. NVIDIA also reports further post-submission optimisations on GPT-OSS-120B and DLRMv3 that are not yet verified by MLCommons, which should be treated as directional only.

What should procurement actually do?

Price the software upgrade before the hardware refresh, use tokens per second per rack (and per watt, per hour) as the single unit of comparison, and measure it with your own production replay rather than adopting a vendor multiplier as a planning figure.

Why can benchmark numbers not go straight into a roadmap?

Because a benchmark measures system throughput under a specific workload and scenario mix. Your product has its own context lengths, concurrency patterns and latency constraints. The benchmark tells you the direction of the ceiling; only your own replay tells you the floor of your cost.