同一机架 3.7 倍:读懂 MLPerf Inference v6.1
💡 工具推荐:API 响应时间计算, AI Token 计数器, 百分比计算
2026 年 9 月 16 日,MLPerf Inference v6.1 结果公布,NVIDIA 的 Vera Rubin NVL72 完成了它的第一次(预览性质)提交:在 Qwen3-VL 上吞吐最高达到上一代 GB300 NVL72 的 3.7 倍,在 DeepSeek-R1 上最高 2.5 倍。但同一轮结果里还有一个更值得工程团队注意的数字:GB300 NVL72 自己,仅靠软件优化,在 Qwen3-VL 上就比 v6.0 提升最高 1.6 倍。把这两个数字并排看,你才会明白 MLPerf 这一轮真正在讲什么故事。
推理基准的正确单位是每机架每秒 token
一、把推理基准读成「每机架每秒多少 token」
NVIDIA 对 MLPerf 意义的表述非常商业化:更高的系统性能意味着生成更多 token,进而带来更高收入。这句话决定了正确的读法——不要看单卡指标,要看每机架、每瓦、每小时的 token 产出。一次 3.7 倍的提升,说的是这个比值的变化,而不是某颗芯片的规格。对采购者来说,这也是唯一能和自己业务对齐的口径:你的产品按 token 计费,机器按机架与千瓦计费,中间那个商就是你真正的成本结构。任何不能落回这个商数的「性能」数字,都只是参数表上的装饰。
# 1. Read an inference benchmark as tokens per rack, not per chip
def token_economics(result):
# NVIDIA's framing of why this metric matters is a revenue statement:
# "Higher system performance means more tokens generated, resulting in
# higher revenue." Everything else in the submission is an input to this.
rack_throughput = result["tokens_per_second"] * result["accelerators_per_rack"]
return {
"tokens_per_second_per_rack": rack_throughput,
"power_per_rack_kw": result["rack_power_kw"],
"tokens_per_joule": rack_throughput / max(result["rack_power_kw"] * 1000, 1),
"tokens_per_dollar_hour": rack_throughput / max(result["cost_per_hour"], 0.01),
}
# A 3.7x on one workload is a claim about this ratio, not about a chip.二、把倍数和工作负载绑在一起读
这一轮的头条数字有三个:Qwen3-VL 上最高 3.7 倍、DeepSeek-R1 上最高 2.5 倍,以及 NVIDIA 在预览测试中提到的 SemiAnalysis AgentX 基准上最高 30 倍。它们覆盖 offline、server、interactive 三种场景,技术栈是 vLLM 加上开源的 NVIDIA Dynamo 推理框架。这里有一条简单的判别规则:脱离工作负载名字的倍数,是营销数字;带着三个工作负载名字的倍数,才是工程声明。同时要分清「预览提交」与「正式提交」——预览意味着平台尚未上市,你买到的是路线图上的性能,不是货架上的性能。
// 2. The v6.1 headline, with its workload attached
const mlperfInferenceV61 = {
date: "2026-09-16",
submitter: "NVIDIA",
system: "Vera Rubin NVL72",
submissionType: "preview", // first appearance of the platform
versus: "GB300 NVL72",
gains: {
"Qwen3-VL": "up to 3.7x higher throughput",
"DeepSeek-R1": "up to 2.5x",
"SemiAnalysis AgentX": "up to 30x in preview testing (per NVIDIA)",
},
scenarios: ["offline", "server", "interactive"],
stack: ["vLLM", "NVIDIA Dynamo"],
};
// Rule of thumb: a single multiplier without a workload name is a marketing
// number. Three multipliers with workload names are an engineering claim.倍数必须和工作负载绑在一起
三、软件才是你今天就买得到的部分
同一轮里最被低估的结果是 GB300 NVL72 自身的进步:在 Qwen3-VL 上,仅靠软件优化就比 v6.0 提升最高 1.6 倍。NVIDIA 列出的手段包括更低的 KV cache 精度、更多的算子融合、更好的 kernel,以及基于 vLLM 与 Dynamo 的分离式服务(disaggregated serving)。同样的硅、同样的机架,在一个发布周期里多出 1.6 倍吞吐。这给出一条很实际的预算顺序:在批准一次硬件换代之前,先给软件升级定价——它通常更便宜,而且不需要等货期。
# 3. Software is the part you can buy today
SOFTWARE_GAINS_V61 = {
"platform": "GB300 NVL72",
"workload": "Qwen3-VL",
"gain_vs_v60": 1.6, # up to 1.6x, from software alone
"techniques": [
"lower KV cache precision",
"additional kernel fusion",
"better kernels",
"disaggregated serving with vLLM and NVIDIA Dynamo",
],
}
# The same silicon, the same rack, 1.6x more throughput in one release cycle.
# Before budgeting for a hardware refresh, price the software upgrade.四、把「已核验」与「截止后预览」分开记账
MLPerf 的可比性来自 MLCommons 的核验规则,所以读结果时必须做一次分类:经 MLCommons 核验的提交,可以跨厂商横向比较;预览阶段的提交,说明方向但平台未出货;而 NVIDIA 表示在 v6.1 提交截止后又有进一步优化(涉及 GPT-OSS-120B 与 DLRMv3,尚未经 MLCommons 核验)——这类数字只能当作方向信息,不能写进采购决策。把这三类混在一起,是性能评估里最常见也最贵的错误:你会用「未经核验的最好数字」去买「尚不存在的系统」。
# 4. Separate verified results from post-deadline previews
def classify(claim):
if claim["verified_by"] == "MLCommons":
return "comparable across submitters"
if claim["stage"] == "post_submission_not_yet_verified":
# NVIDIA reports further gains on GPT-OSS-120B and DLRMv3 after the
# v6.1 submission deadline; those numbers have not been verified.
return "directional only - do not put it in a procurement decision"
if claim["stage"] == "preview":
return "promising, but the platform is not shipping today"
return "unclassified"
def procurement_grade(claims):
return [c for c in claims if classify(c) == "comparable across submitters"]先给软件升级定价,再考虑硬件换代
五、一份经得起追问的容量决策记录
如果你的团队正在为智能体产品做推理容量规划,可以把结论写成这样一份记录:明确工作负载(长上下文、Qwen3-VL 类)、明确指标(每机架每秒 token,且按你真实服务的交互模式测)、明确测量方式(在生产流量回放上自测,而不是只看厂商提交),把用到的厂商声明逐条列出(例如 v6.1 预览提交的 3.7 倍、GB300 软件优化的 1.6 倍),把未核验的数字单列在「不采纳」栏,然后给出决策与复核时间点(比如下一轮 MLPerf 提交结果公布时)。这份记录的价值在于:半年后再被问「你当初凭什么这么买」时,你有答案。
{
"capacity_decision_record": {
"workload": "long-context Qwen3-VL style inference for an agent product",
"metric": "tokens per second per rack, at the interaction pattern we actually serve",
"measured_by": "our own replay of production traffic on borrowed capacity",
"vendor_claims_used": [
"MLPerf Inference v6.1 Vera Rubin NVL72 preview: up to 3.7x vs GB300 NVL72 on Qwen3-VL",
"GB300 NVL72 software-only gain of up to 1.6x over v6.0"
],
"claims_excluded": ["post-submission, not-yet-verified results"],
"decision": "keep GB300 fleet for 2 quarters; adopt the v6.1 software stack first",
"review_due": "next MLPerf submission round"
}
}📌 常见问题 FAQ
3.7 倍具体指什么?
据 NVIDIA 的公告:在 MLPerf Inference v6.1 的首次预览提交中,Vera Rubin NVL72 在 Qwen3-VL 工作负载上的吞吐最高达到 GB300 NVL72 的 3.7 倍,覆盖 offline、server 与 interactive 三种场景,使用的是 vLLM 与 NVIDIA Dynamo。
为什么「软件提升 1.6 倍」更值得关注?
因为它是同一代硬件(GB300 NVL72)在一轮发布周期内由软件取得的提升:更低的 KV cache 精度、更多算子融合、更好的 kernel、以及基于 vLLM 与 Dynamo 的分离式服务。硬件不变而吞吐提升,是投入产出比最高的那一类优化。
预览提交和正式提交有什么区别?
提交经 MLCommons 核验的结果可以跨厂商比较;预览提交用于展示尚在准备中的平台。此外 NVIDIA 提到提交截止后的进一步优化结果尚未经 MLCommons 核验,这类数字只能作为方向参考。
对采购有什么直接建议?
先给软件升级定价,再考虑硬件换代:把「每机架每秒 token、每瓦 token、每小时 token」作为唯一口径,并用自己的生产流量回放来测量,而不是依赖任何单一厂商提交的倍数。
为什么基准数字不能直接搬进产品路线图?
因为基准测的是特定工作负载与场景组合下的系统吞吐,而你的产品有其自身的上下文长度、并发模式与延迟约束。基准能告诉你上限方向,只有你自己的回放实验能告诉你成本底线。
🔧 推荐工具
📚 参考资料
- NVIDIA blog (September 16, 2026) - NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut: up to 3.7x better throughput than GB300 NVL72, GB300 software gains of up to 1.6x over v6.0, and post-submission results not yet verified by MLCommons
- MLCommons - MLPerf Inference: datacenter benchmark rules, scenarios and the published results submissions the numbers above come from
- diyai.io - analysis of the MLCommons submission entry for VR200 NVL72 (72 accelerators across 18 nodes, 288GB HBM4 per accelerator, liquid cooling); secondary source, used only for the system configuration