开源权重已逼近前沿:如何正确解读 2026 年 9 月的差距
💡 工具推荐:AI Token 计数器、API 限流计算器、AI 代码审查
2026 年 9 月的开源权重格局,用一句话概括:头部开源模型与闭源前沿之间,只差个位数。在 Artificial Analysis 智能指数(AA Intelligence Index)上,GLM-5.3 以 45 分领跑开源阵营,Kimi K3 以 44 分紧随,Qwen3.8-Max 约 40 分;而美国最强的开源模型 Inkling(Thinking Machines)只有 26 分。闭源第一梯队——Claude Fable 5.1 与 GPT-6 Astra——同为 53 分。也就是说,顶级开源模型与闭源前沿之间约 9 分,而「开源 vs 开源」之间的差距反而更大。
先厘清数字,别急着下结论
这里必须先履行一条纪律:区分事实与来源分歧。Interconnects AI 在 9 月 14 日的分析中给出:GLM-5.3 为 45、Kimi K3 为 44、Kimi K3 与 GLM-5.3-Flash 分别为 44 与 42,美国最好的开源模型 Inkling 为 26、Nvidia Nemotron 3 Ultra 为 23。Maxime Labonne 的分析同样给出 GLM-5.3 领先开源阵营 45 分、K3 落后 1 分,并指出 Fable 5.1 已把差距拉大到 8 分。而 24/7 Wall St. 的表述是 GLM-5.3 与 Kimi K3 同为 44 分、闭源领先 53 分,差距 9 分。数字差 1 到 2 分,结论方向一致:差距是个位数。
// Rule one when reading a leaderboard: record the index version and the date.
// A score without its snapshot is a rumour, and the snapshots disagree.
type Score = { model: string; index: string; value: number; asOf: string };
const snapshot: Score[] = [
{ model: "GLM-5.3", index: "AA Intelligence Index", value: 45, asOf: "2026-09" },
{ model: "Kimi K3", index: "AA Intelligence Index", value: 44, asOf: "2026-09" },
{ model: "Qwen3.8-Max", index: "AA Intelligence Index", value: 40, asOf: "2026-09" },
{ model: "Inkling (US open)", index: "AA Intelligence Index", value: 26, asOf: "2026-09" },
];
// Report the spread, not one number. If two sources differ by two points, say so.
function spread(a: number, b: number) { return Math.abs(a - b); }第一梯队开源模型与闭源前沿只差个位数
「同分」之下,价格才是真正的变量
当能力接近时,决策权就交给了成本。GLM-5.3 于 2026 年 8 月 14 日发布,与 GLM-5.2 共用同一基座,增益主要来自扩大后训练;据 Morph 的整理,它在 Terminal-Bench 3.0 上从 4.6 提升到 28.3,在 DeepSWE v1.1 上从 46.2 提升到 66.9,并以 45 分略胜 Kimi K3 的 44 分,而 token 价格约为 K3 的五分之一(GLM 定价维持在 1.40/4.40 美元每百万 token)。在 DeepSWE 的单任务成本上,GLM-5.3 为 3.99 美元,Kimi K3 为 4.65 美元。能力打平的地方,价格就是答案。
# Capability is half a decision. The other half is what a task costs at that
# capability. Two models tied on score can be far apart on the bill.
models = [
# name, open_weights, price_in_per_M, price_out_per_M, cost_per_task_usd
("GLM-5.3", True, 1.40, 4.40, 3.99),
("Kimi K3", True, None, None, 4.65),
("Claude Fable 5.1", False, None, None, None),
]
for name, weights, pin, pout, cpt in models:
tag = "open" if weights else "closed"
print(f"{name:20s} {tag:6s} cost/task={cpt}")
# At equal score, the cheaper open model is the one to benchmark first for text
# coding. That is a procurement fact, not a religious one.许可不是脚注,是运行时约束
「开源」这个词需要拆开看。GLM-5.3 的权重自 2026 年 8 月 25 日起公开,但采用的是自定义的 GLM-5.3 许可;Kimi K3 的权重自 7 月 27 日起在 Hugging Face 上按 Kimi K3 许可提供。两者都不是「随便用」的 MIT 级别。计划自托管前,先把许可读一遍,并把「选这个模型」和「依据什么许可」一起记录在案。把任何自定义命名的许可当成需要法务复核的事项,而不是一个勾选框。
// A licence is a runtime constraint, not fine print. GLM-5.3's weights are public
// under a custom GLM-5.3 licence, and Kimi K3 ships under its own Kimi K3 licence.
// Read both before you plan a deployment, and record the decision next to the model.
class Deployment {
constructor(model, license) { this.model = model; this.license = license; }
requiresReview() {
// Any non-standard, custom-name licence is a legal review, not a checkbox.
return /custom|non-commercial|research/i.test(this.license);
}
}
const plan = new Deployment("GLM-5.3", "GLM-5.3 License (custom)");
if (plan.requiresReview()) console.log("route to legal before self-hosting");同分之下,价格差才是真变量
把自托管算成一笔账,而不是一种信仰
自托管的经济性是个盈亏平衡问题。你需要用你自己的数字去算:硬件或 GPU 租用、利用率、电力,以及运行它的工程时间。许可只是第一道门槛,盈亏平衡是第二道。一个简单原则:只有当自托管的总成本低于它所替代的 API 账单时,才值得买这台机器。而这个式子两边都会移动,所以每个季度都该重算一次,别把它当成一次性决定。
// Instrument the router, then let the data settle the argument. Route by task
// class, cap cost, and log which tier actually satisfied the request.
type Tier = "open-cheap" | "open-strong" | "closed-frontier";
type Decision = { task: string; tier: Tier; usd: number; accepted: boolean };
function route(task: string, difficulty: number, budgetUsd: number): Tier {
if (difficulty < 0.4 && budgetUsd < 0.01) return "open-cheap";
if (difficulty < 0.8) return "open-strong";
return "closed-frontier";
}
function logRoute(d: Decision) {
db.routes.insert({ ...d, ts: Date.now() });
}
// If 90% of accepted requests never leave the open tiers, the frontier is an
// escalation path, not your default. Measure it before you assume otherwise.用路由把争论变成可辩护的决策
把模型差异交给路由去裁决。按任务类别分流:低难度、低预算的请求交给便宜的开源档;中等难度交给强开源档;只有高难度才升级到闭源前沿。然后记录每一次决策——任务类别、档位、花费、是否被接受。如果 90% 被接受的请求从未离开开源档,那么闭源前沿对你而言就是一条升级通道,而不是默认选项。先测量,再假设。
# Self-hosting economics are a break-even question, not a vibe. Compute it with
# your own numbers: hardware, utilisation, power, and the engineering time to run
# it. The licence is the first gate; the break-even is the second.
def break_even_monthly(api_spend_usd, gpu_rent_usd, eng_hours_usd):
# You buy the box when renting it costs more than the API bill it replaces.
total_self_host = gpu_rent_usd + eng_hours_usd
return {
"api": api_spend_usd,
"self_host": total_self_host,
"verdict": "self-host" if total_self_host < api_spend_usd else "stay on API",
}
print(break_even_monthly(4200, 3100, 900))
# Then re-run it every quarter: both terms move, and the answer moves with them.许可、权重与硬件共同定义「可行」
差距会动,方法不动
把这条新闻读成一种方法。数字会随快照变化——有的来源给你 44,有的给你 45,有的说差距 8 分,有的说 9 分——所以任何单个分数都不该进你的技术选型文档,除非附带「指数版本 + 日期」。真正稳定的是你的评估方式:记录快照与来源、按能力与成本双轴评估、把许可当作门槛、把自托管算成账、用路由把差异变成决策。方法比分数活得久,而这正是 2026 年 9 月这条新闻真正值得抄下来的部分。
📌 常见问题 FAQ
目前开源权重与闭源前沿差多少?
按 Artificial Analysis 智能指数,截至 2026 年 9 月,顶级开源模型 GLM-5.3 为 45、Kimi K3 为 44,闭源第一梯队 Claude Fable 5.1 与 GPT-6 Astra 同为 53,差距约个位数。不同来源给出的数字相差 1 到 2 分,方向一致。
为什么不同来源的分数不一样?
因为指数有版本、快照有日期。Interconnects AI(9 月 14 日)给出 GLM-5.3 为 45;24/7 Wall St. 的表述是 GLM-5.3 与 Kimi K3 同为 44。记录任何分数时都应附带指数版本与日期,否则它只是一个传闻。
GLM-5.3 的许可可以自由商用吗?
不能想当然。GLM-5.3 的权重自 2026 年 8 月 25 日起公开,但采用自定义的 GLM-5.3 许可;Kimi K3 使用自己的 Kimi K3 许可。两者都非 MIT 级别的宽松许可,自托管前应读许可并做法务复核。
什么时候该自托管而不是用 API?
把它当作盈亏平衡问题:只有当自托管的总成本(GPU、利用率、电力、工程时间)低于它所替代的 API 账单时,才值得买机器。许可只是第一道门槛,且这个式子每个季度都会变,需要重算。
怎样把模型差异变成可辩护的决策?
用路由按任务类别分流——低难度走便宜开源档、中难度走强开源档、高难度才升级闭源前沿——并记录每次决策的任务类别、档位、花费与是否被接受。若多数被接受的请求从未离开开源档,前沿就只是升级通道而非默认。