DeepSeek 联手华为开源 TileLang:给昇腾芯片补上「高级语言」这一层
2026 年 9 月 30 日,DeepSeek 在官方微信公众号宣布与华为合作,为华为昇腾(Ascend)芯片开发编程工具,并开源了昇腾平台上的编程基础设施。路透社的报道说得很具体:开源的清单里包括一门名为 TileLang 的高级编程语言,以及配套的计算库与通信库。DeepSeek 给出的理由是——要建新一代自主可控的 GPU 软件生态,第一优先级是拥有人人都能用、好写、还能榨干硬件性能的高级语言。这句话里藏着三个互相拉扯的约束,也是本文要讲清楚的地方。
一、这次到底发布了什么
先把事实钉死。路透社 9 月 30 日发自香港的报道给出了三条明确信息:DeepSeek 表示已与华为合作开发昇腾芯片的编程工具;在官方微信公众号的帖子里,DeepSeek 称正在开源昇腾平台的编程基础设施;开源内容包括一门名为 TileLang 的高级编程语言,以及相关的计算与通信库。报道同时点明了这件事的背景——中国科技公司正在加深彼此之间的合作,寻找英伟达生态之外的替代方案。没有发布会,没有产品页,一条公众号帖子就把最敏感的一层软件开源了出去。
# What DeepSeek says it released, restated as a mental model.
# This is a shape for reasoning, not a vendor API.
ascend_software_stack = {
"high_level_language": "TileLang", # the headline item
"compute_libraries": True, # open-sourced
"communication_libs": True, # chip-to-chip / collective
"target_hardware": "Huawei Ascend",
"comparison_claimed": "simpler programming model than NVIDIA CUDA",
}
# The claim worth testing is the language, not the libraries.
# Libraries shave weeks off one port. A language changes who is
# allowed to write kernels in the first place -- and that is the moat.CUDA 的护城河从来不是硅,而是语言
二、为什么「高级语言」是最难的一层
DeepSeek 在帖子里把优先级讲得很直白:要建新一代自主可控的 GPU 软件生态,第一优先级是建立一门通用、易于编程、同时仍能触及硬件完整性能潜力的高级语言,并称 TileLang 正是为此而生。请注意这三个约束是互相拉扯的:「通用」要求它能跨越不同代的加速器;「易于编程」要求应用团队能直接上手,而不是只有内核专家能碰;「完整性能潜力」则意味着不能因为抽象而付出几倍的性能税。任何一层软件都可以替换,只有语言的替换成本最高——因为语言决定的是谁能写内核、生态里有多少现成的代码可以复用。
# DeepSeek's own framing, quoted by Reuters on 2026-09-30.
# Read it as a requirements document, not marketing:
DEEPSEEK_REASONING = """
To build a new generation of independent, self-controlled GPU
software ecosystem, the first priority is establishing a high-level
language that is universal, easy to program, and still capable of
reaching the hardware's full performance potential.
"""
# Three constraints, and the third one is the hard one.
CONSTRAINTS = {
"universal": "one language, many accelerator generations",
"easy_to_program": "reachable by app teams, not just kernel specialists",
"full_performance": "no 2-5x tax versus hand-tuned kernels",
}三、超级节点:把 128 颗芯片当成一台机器
路透社的报道还提到一件容易被忽略的事:两家公司联合推进了一套基于 128 颗昇腾 950 芯片的超级节点配置,并对集群的算力与芯片间通信做了联合优化。这条信息的重要性在于它指向了真正难的部分。单卡峰值算力是纸面数字,能买到的是可持续算力;在大规模并行下,性能损耗几乎全部来自通信与调度。也就是说,这个超级节点本质是一个软件交付物——它检验的正是「高级语言 + 通信库」这套组合能不能把规模化的损耗压下来。
# A portability check you can actually run. The question is never
# "does the hardware work?" -- it is "how much source has to change?"
PORT_SURFACE = [
"kernel language", # TileLang vs CUDA C++ / Triton
"collective comms", # chip-to-chip libraries
"graph capture", # does the runtime expose CUDA graphs equivalent?
"quantisation kernels", # FP8 / INT4 paths
"profiler + debugger", # what do you read when it is slow?
"CI on real silicon", # can you test without renting a cluster?
]
def portability_score(available):
missing = [k for k in PORT_SURFACE if k not in available]
return {"coverage": 1 - len(missing) / len(PORT_SURFACE), "missing": missing}高级语言决定谁来写内核
四、对开发团队意味着什么
如果你只在一种加速器上跑模型,这次发布短期内不会改变你的预算。真正受影响的是三类人:一是在做硬件多元化、需要给供应链留后路的基础设施团队;二是被 CUDA 特有的内核写法绑死、难以迁移的推理团队;三是评估自研或国产推理集群、需要判断三年后软件成熟度的技术决策者。对第二类人来说,判断标准很朴素:把「可移植性面积」列成清单,逐项对——内核语言、集合通信、图捕获、量化内核、性能分析工具、真机 CI。缺哪一项,就意味着哪一项要你自己补。
# The supernode work Reuters reported alongside the release: a system
# built on 128 Ascend 950 chips, with compute throughput and
# chip-to-chip communication tuned together. Scale-out is a software
# problem, so treat it as a software deliverable, not a spec sheet.
SUPERNODE = {"chips": 128, "model": "Ascend 950"}
def effective_throughput(chips, per_chip, comm_efficiency):
"""Peak flops are easy. Sustained flops are what you buy."""
ideal = chips * per_chip
return {
"ideal": ideal,
"sustained": ideal * comm_efficiency,
"loss_pct": round((1 - comm_efficiency) * 100, 1),
}
# Ask every accelerator vendor for `comm_efficiency` at your tensor
# parallel width, not at their demo width.五、像工程师一样验证,而不是像分析师一样转述
「比 CUDA 更简单」是一个可以设计实验去证伪的命题。不要拿厂商的演示宽度去问通信效率,要拿你自己的张量并行宽度去问;不要只比峰值吞吐,要同时对吞吐比、P99 延迟比、以及最关键的「是否需要重新训练权重」。一个诚实的评估框架应该包含三个数字:相同模型与精度下的每秒 token 比值、p99 延迟比值、以及完成一次移植所改动的人日。任何没有这三个数字的「已适配」声明,都还停留在演示阶段。
# Vendor claims deserve vendor-grade verification. Put the numbers
# behind a gate so "it works on paper" never becomes a roadmap item.
def evaluate_accelerator_claim(native_kernels, ascend_kernels, prompt_suite):
"""Compare like for like: same model, same batch, same dtype."""
results = {}
for name, kernels in (("native", native_kernels), ("ascend", ascend_kernels)):
results[name] = benchmark(kernels, suite=prompt_suite)
native, ascend = results["native"], results["ascend"]
return {
"tok_per_s_ratio": ascend.tokens_per_second / native.tokens_per_second,
"p99_latency_ratio": ascend.p99_ms / native.p99_ms,
"retrain_required": not ascend.weights_compatible,
"verdict": "port" if ascend.tokens_per_second / native.tokens_per_second > 0.9
else "re-benchmark next generation",
}把 128 颗芯片当成一台机器来调优
六、对 CUDA 护城河的实际意义
英伟达的护城河从来不只是硅,而是那一层让全世界的内核、编译器、调优经验都长在上面的软件惯性。DeepSeek 这次的动作,打的正是这层惯性里最上游的位置:语言。它不会在明天改变谁买什么卡,但它改变了「替代方案的成本曲线」——只要有一个足够好用、且被真实工作负载验证过的高级语言,上层的库、框架、调优经验就有机会跟着一起迁移。对任何一家不想把三年路线图押在单一供应商上的团队来说,这值得放进观察名单,而不是放进新闻摘要。
📌 常见问题 FAQ
DeepSeek 在 2026 年 9 月 30 日宣布了什么?
DeepSeek 在官方微信公众号发帖称,已与华为合作开发面向华为昇腾芯片的编程工具,并开源了昇腾平台的编程基础设施,其中包括一门名为 TileLang 的高级编程语言,以及相关的计算与通信库。
为什么说「高级语言」是最关键的一层?
DeepSeek 自己的表述是:要建新一代自主可控的 GPU 软件生态,第一优先级是建立一门通用、易于编程、同时仍能触及硬件完整性能潜力的高级语言。库能省下移植的时间,语言决定的是谁能写内核——这正是 CUDA 生态长期占据护城河的原因。
TileLang 和 CUDA 相比,DeepSeek 的说法是什么?
据路透社报道,DeepSeek 称 TileLang 提供了一种比英伟达 CUDA 更简单的编程模型,可以提升开发效率、简化代码逻辑。这是 DeepSeek 的说法,独立第三方基准尚未在报道中给出。
报道里还提到了什么硬件层面的进展?
路透社报道还提到,两家公司联合推进了一套基于 128 颗昇腾 950 芯片的超级节点(supernode)配置,并对该集群的整体算力与芯片间通信做了联合优化。
工程团队应该怎么评估这类声明?
不要按厂商的演示宽度去问通信效率,要按你自己的张量并行宽度问;把「可移植性面积」列成清单(内核语言、集合通信、图捕获、量化内核、性能分析工具、真机 CI),逐项打分;最后用同一模型、同一批次、同一精度做同场对比,看吞吐比、P99 延迟比,以及是否需要重新训练权重。
🔧 推荐工具
📚 参考资料
- Reuters — DeepSeek partners with Huawei to develop chip programming tools, reducing reliance on Nvidia (2026-09-30)
- Reuters via The Bellingham Herald — full wire text (2026-09-29/30)
- Reuters via ANI — DeepSeek joins hands with Huawei to build AI chip software (128 x Ascend 950 supernode)
- Reuters via TradingView / GuruFocus — DeepSeek Just Gave Huawei a New Weapon in the Race Against Nvidia