同一机架 3.7 倍:读懂 MLPerf Inference v6.1

·阅读约10分钟·Evergreen Tools Team

2026 年 9 月 16 日,MLPerf Inference v6.1 结果公布,NVIDIA 的 Vera Rubin NVL72 完成了它的第一次(预览性质)提交:在 Qwen3-VL 上吞吐最高达到上一代 GB300 NVL72 的 3.7 倍,在 DeepSeek-R1 上最高 2.5 倍。但同一轮结果里还有一个更值得工程团队注意的数字:GB300 NVL72 自己,仅靠软件优化,在 Qwen3-VL 上就比 v6.0 提升最高 1.6 倍。把这两个数字并排看,你才会明白 MLPerf 这一轮真正在讲什么故事。

推理基准的正确单位是每机架每秒 token

推理基准的正确单位是每机架每秒 token

一、把推理基准读成「每机架每秒多少 token」

NVIDIA 对 MLPerf 意义的表述非常商业化:更高的系统性能意味着生成更多 token,进而带来更高收入。这句话决定了正确的读法——不要看单卡指标,要看每机架、每瓦、每小时的 token 产出。一次 3.7 倍的提升,说的是这个比值的变化,而不是某颗芯片的规格。对采购者来说,这也是唯一能和自己业务对齐的口径:你的产品按 token 计费,机器按机架与千瓦计费,中间那个商就是你真正的成本结构。任何不能落回这个商数的「性能」数字,都只是参数表上的装饰。

# 1. Read an inference benchmark as tokens per rack, not per chip
def token_economics(result):
    # NVIDIA's framing of why this metric matters is a revenue statement:
    # "Higher system performance means more tokens generated, resulting in
    # higher revenue." Everything else in the submission is an input to this.
    rack_throughput = result["tokens_per_second"] * result["accelerators_per_rack"]
    return {
        "tokens_per_second_per_rack": rack_throughput,
        "power_per_rack_kw": result["rack_power_kw"],
        "tokens_per_joule": rack_throughput / max(result["rack_power_kw"] * 1000, 1),
        "tokens_per_dollar_hour": rack_throughput / max(result["cost_per_hour"], 0.01),
    }
# A 3.7x on one workload is a claim about this ratio, not about a chip.

二、把倍数和工作负载绑在一起读

这一轮的头条数字有三个:Qwen3-VL 上最高 3.7 倍、DeepSeek-R1 上最高 2.5 倍,以及 NVIDIA 在预览测试中提到的 SemiAnalysis AgentX 基准上最高 30 倍。它们覆盖 offline、server、interactive 三种场景,技术栈是 vLLM 加上开源的 NVIDIA Dynamo 推理框架。这里有一条简单的判别规则:脱离工作负载名字的倍数,是营销数字;带着三个工作负载名字的倍数,才是工程声明。同时要分清「预览提交」与「正式提交」——预览意味着平台尚未上市,你买到的是路线图上的性能,不是货架上的性能。

// 2. The v6.1 headline, with its workload attached
const mlperfInferenceV61 = {
  date: "2026-09-16",
  submitter: "NVIDIA",
  system: "Vera Rubin NVL72",
  submissionType: "preview",          // first appearance of the platform
  versus: "GB300 NVL72",
  gains: {
    "Qwen3-VL": "up to 3.7x higher throughput",
    "DeepSeek-R1": "up to 2.5x",
    "SemiAnalysis AgentX": "up to 30x in preview testing (per NVIDIA)",
  },
  scenarios: ["offline", "server", "interactive"],
  stack: ["vLLM", "NVIDIA Dynamo"],
};
// Rule of thumb: a single multiplier without a workload name is a marketing
// number. Three multipliers with workload names are an engineering claim.
倍数必须和工作负载绑在一起

倍数必须和工作负载绑在一起

三、软件才是你今天就买得到的部分

同一轮里最被低估的结果是 GB300 NVL72 自身的进步:在 Qwen3-VL 上,仅靠软件优化就比 v6.0 提升最高 1.6 倍。NVIDIA 列出的手段包括更低的 KV cache 精度、更多的算子融合、更好的 kernel,以及基于 vLLM 与 Dynamo 的分离式服务(disaggregated serving)。同样的硅、同样的机架,在一个发布周期里多出 1.6 倍吞吐。这给出一条很实际的预算顺序:在批准一次硬件换代之前,先给软件升级定价——它通常更便宜,而且不需要等货期。

# 3. Software is the part you can buy today
SOFTWARE_GAINS_V61 = {
    "platform": "GB300 NVL72",
    "workload": "Qwen3-VL",
    "gain_vs_v60": 1.6,   # up to 1.6x, from software alone
    "techniques": [
        "lower KV cache precision",
        "additional kernel fusion",
        "better kernels",
        "disaggregated serving with vLLM and NVIDIA Dynamo",
    ],
}

# The same silicon, the same rack, 1.6x more throughput in one release cycle.
# Before budgeting for a hardware refresh, price the software upgrade.

四、把「已核验」与「截止后预览」分开记账

MLPerf 的可比性来自 MLCommons 的核验规则,所以读结果时必须做一次分类:经 MLCommons 核验的提交,可以跨厂商横向比较;预览阶段的提交,说明方向但平台未出货;而 NVIDIA 表示在 v6.1 提交截止后又有进一步优化(涉及 GPT-OSS-120B 与 DLRMv3,尚未经 MLCommons 核验)——这类数字只能当作方向信息,不能写进采购决策。把这三类混在一起,是性能评估里最常见也最贵的错误:你会用「未经核验的最好数字」去买「尚不存在的系统」。

# 4. Separate verified results from post-deadline previews
def classify(claim):
    if claim["verified_by"] == "MLCommons":
        return "comparable across submitters"
    if claim["stage"] == "post_submission_not_yet_verified":
        # NVIDIA reports further gains on GPT-OSS-120B and DLRMv3 after the
        # v6.1 submission deadline; those numbers have not been verified.
        return "directional only - do not put it in a procurement decision"
    if claim["stage"] == "preview":
        return "promising, but the platform is not shipping today"
    return "unclassified"

def procurement_grade(claims):
    return [c for c in claims if classify(c) == "comparable across submitters"]
先给软件升级定价,再考虑硬件换代

先给软件升级定价,再考虑硬件换代

五、一份经得起追问的容量决策记录

如果你的团队正在为智能体产品做推理容量规划,可以把结论写成这样一份记录:明确工作负载(长上下文、Qwen3-VL 类)、明确指标(每机架每秒 token,且按你真实服务的交互模式测)、明确测量方式(在生产流量回放上自测,而不是只看厂商提交),把用到的厂商声明逐条列出(例如 v6.1 预览提交的 3.7 倍、GB300 软件优化的 1.6 倍),把未核验的数字单列在「不采纳」栏,然后给出决策与复核时间点(比如下一轮 MLPerf 提交结果公布时)。这份记录的价值在于:半年后再被问「你当初凭什么这么买」时,你有答案。

{
  "capacity_decision_record": {
    "workload": "long-context Qwen3-VL style inference for an agent product",
    "metric": "tokens per second per rack, at the interaction pattern we actually serve",
    "measured_by": "our own replay of production traffic on borrowed capacity",
    "vendor_claims_used": [
      "MLPerf Inference v6.1 Vera Rubin NVL72 preview: up to 3.7x vs GB300 NVL72 on Qwen3-VL",
      "GB300 NVL72 software-only gain of up to 1.6x over v6.0"
    ],
    "claims_excluded": ["post-submission, not-yet-verified results"],
    "decision": "keep GB300 fleet for 2 quarters; adopt the v6.1 software stack first",
    "review_due": "next MLPerf submission round"
  }
}

📌 常见问题 FAQ

3.7 倍具体指什么?

据 NVIDIA 的公告:在 MLPerf Inference v6.1 的首次预览提交中,Vera Rubin NVL72 在 Qwen3-VL 工作负载上的吞吐最高达到 GB300 NVL72 的 3.7 倍,覆盖 offline、server 与 interactive 三种场景,使用的是 vLLM 与 NVIDIA Dynamo。

为什么「软件提升 1.6 倍」更值得关注?

因为它是同一代硬件(GB300 NVL72)在一轮发布周期内由软件取得的提升:更低的 KV cache 精度、更多算子融合、更好的 kernel、以及基于 vLLM 与 Dynamo 的分离式服务。硬件不变而吞吐提升,是投入产出比最高的那一类优化。

预览提交和正式提交有什么区别?

提交经 MLCommons 核验的结果可以跨厂商比较;预览提交用于展示尚在准备中的平台。此外 NVIDIA 提到提交截止后的进一步优化结果尚未经 MLCommons 核验,这类数字只能作为方向参考。

对采购有什么直接建议?

先给软件升级定价,再考虑硬件换代:把「每机架每秒 token、每瓦 token、每小时 token」作为唯一口径,并用自己的生产流量回放来测量,而不是依赖任何单一厂商提交的倍数。

为什么基准数字不能直接搬进产品路线图?

因为基准测的是特定工作负载与场景组合下的系统吞吐,而你的产品有其自身的上下文长度、并发模式与延迟约束。基准能告诉你上限方向,只有你自己的回放实验能告诉你成本底线。