DeepSeek Open-Sources TileLang for Huawei's Ascend Chips: Attacking the CUDA Moat

·10 min read·Evergreen Tools Team

On September 30, 2026, DeepSeek said it had partnered with Huawei to develop programming tools for Huawei's Ascend chips, and that it was open-sourcing programming infrastructure for the Ascend platform. Reuters is specific about what is in the release: a high-level programming language called TileLang, alongside related compute and communication libraries. DeepSeek's stated reason is blunt: to build a new generation of independent, self-controlled GPU software ecosystem, the first priority is a high-level language that is universal, easy to program, and still capable of reaching the hardware's full performance potential. Those are three constraints that fight each other. That tension is the story.

1. What was actually announced

Pin the facts first. Reuters, reporting from Hong Kong on September 30, gives three concrete items: DeepSeek said it has partnered with Huawei to develop programming tools for Huawei's Ascend chips; in a post on its official WeChat account it said it was open-sourcing infrastructure for the Ascend platform; and the release includes a high-level programming language called TileLang along with related compute and communication libraries. The story also names the backdrop — Chinese technology firms deepening ties as they look for alternatives to the NVIDIA ecosystem. No launch event, no product page. A WeChat post moved the most sensitive layer of the stack into open source.

# What DeepSeek says it released, restated as a mental model.
# This is a shape for reasoning, not a vendor API.

ascend_software_stack = {
    "high_level_language": "TileLang",          # the headline item
    "compute_libraries":   True,                # open-sourced
    "communication_libs":  True,                # chip-to-chip / collective
    "target_hardware":     "Huawei Ascend",
    "comparison_claimed":  "simpler programming model than NVIDIA CUDA",
}

# The claim worth testing is the language, not the libraries.
# Libraries shave weeks off one port. A language changes who is
# allowed to write kernels in the first place -- and that is the moat.
An AI accelerator chip and its software stack

CUDA's moat was never the silicon. It is the language

2. Why the language layer is the hard part

DeepSeek states its priority plainly: to build a new generation of independent, self-controlled GPU software ecosystem, the first priority is establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware's full performance potential. TileLang, it says, was created precisely to meet that need. Notice how those three constraints pull against each other. Universal means it has to span accelerator generations. Easy to program means application teams, not just kernel specialists, can use it. Full performance potential means the abstraction cannot cost you a 2x to 5x tax. Every layer of a stack can be swapped. A language is the most expensive thing to replace, because it decides who writes kernels and how much existing code is reusable.

# DeepSeek's own framing, quoted by Reuters on 2026-09-30.
# Read it as a requirements document, not marketing:

DEEPSEEK_REASONING = """
To build a new generation of independent, self-controlled GPU
software ecosystem, the first priority is establishing a high-level
language that is universal, easy to program, and still capable of
reaching the hardware's full performance potential.
"""

# Three constraints, and the third one is the hard one.
CONSTRAINTS = {
    "universal":        "one language, many accelerator generations",
    "easy_to_program":  "reachable by app teams, not just kernel specialists",
    "full_performance": "no 2-5x tax versus hand-tuned kernels",
}

3. The supernode is the real deliverable

Reuters also reports something easy to skim past: the two companies progressed a supernode configuration built on 128 Ascend 950 chips, optimising compute performance and chip-to-chip communication across the clustered system. That detail points at the genuinely hard part. Per-chip peak floating point is a paper number; sustained throughput is what you actually buy, and at scale almost all of the loss comes from communication and scheduling. A 128-chip supernode is therefore a software deliverable. It is the test of whether the language-plus-communication-library combination can keep scale-out loss from eating the theoretical gain.

# A portability check you can actually run. The question is never
# "does the hardware work?" -- it is "how much source has to change?"

PORT_SURFACE = [
    "kernel language",        # TileLang vs CUDA C++ / Triton
    "collective comms",       # chip-to-chip libraries
    "graph capture",          # does the runtime expose CUDA graphs equivalent?
    "quantisation kernels",   # FP8 / INT4 paths
    "profiler + debugger",    # what do you read when it is slow?
    "CI on real silicon",     # can you test without renting a cluster?
]

def portability_score(available):
    missing = [k for k in PORT_SURFACE if k not in available]
    return {"coverage": 1 - len(missing) / len(PORT_SURFACE), "missing": missing}
Compiling accelerator kernels in a terminal

The high-level language decides who gets to write kernels

4. What it changes for development teams

If you run models on one accelerator family, this announcement will not move your budget this quarter. The teams affected are narrower and more specific: infrastructure groups hedging their supply chain across hardware vendors; inference teams whose custom kernels are welded to CUDA-specific idioms; and the people evaluating a domestic or in-house inference cluster who need to judge how mature the software will be in three years. For that last group the test is unglamorous. Write down the portability surface — kernel language, collective communications, graph capture, quantisation kernels, profiler and debugger, and CI on real silicon — and score each item. Whatever is missing is work you will own. A useful discipline is to write that list down before any vendor conversation, then ask each candidate to fill it in with public documentation links rather than slides. The gaps become your migration plan, and the size of the gap is the real answer to whether a second accelerator family is a hedge or a second job.

# The supernode work Reuters reported alongside the release: a system
# built on 128 Ascend 950 chips, with compute throughput and
# chip-to-chip communication tuned together. Scale-out is a software
# problem, so treat it as a software deliverable, not a spec sheet.

SUPERNODE = {"chips": 128, "model": "Ascend 950"}

def effective_throughput(chips, per_chip, comm_efficiency):
    """Peak flops are easy. Sustained flops are what you buy."""
    ideal = chips * per_chip
    return {
        "ideal": ideal,
        "sustained": ideal * comm_efficiency,
        "loss_pct": round((1 - comm_efficiency) * 100, 1),
    }

# Ask every accelerator vendor for `comm_efficiency` at your tensor
# parallel width, not at their demo width.

5. Evaluate it like an engineer, not a news aggregator

"A simpler programming model than CUDA" is a falsifiable proposition, so design the experiment. Do not ask for communication efficiency at the vendor's demo width; ask for yours. Do not compare only peak throughput; compare throughput ratio, p99 latency ratio, and the hardest number of all — whether weights must be retrained. A defensible evaluation framework needs three figures: tokens per second at the same model, batch and dtype; the p99 latency ratio; and the person-days a real port consumes. Any "we support it" claim that arrives without those three numbers is still a demo.

# Vendor claims deserve vendor-grade verification. Put the numbers
# behind a gate so "it works on paper" never becomes a roadmap item.

def evaluate_accelerator_claim(native_kernels, ascend_kernels, prompt_suite):
    """Compare like for like: same model, same batch, same dtype."""
    results = {}
    for name, kernels in (("native", native_kernels), ("ascend", ascend_kernels)):
        results[name] = benchmark(kernels, suite=prompt_suite)
    native, ascend = results["native"], results["ascend"]
    return {
        "tok_per_s_ratio": ascend.tokens_per_second / native.tokens_per_second,
        "p99_latency_ratio": ascend.p99_ms / native.p99_ms,
        "retrain_required": not ascend.weights_compatible,
        "verdict": "port" if ascend.tokens_per_second / native.tokens_per_second > 0.9
                   else "re-benchmark next generation",
    }
Accelerator cards in a data centre

Tuning 128 chips as if they were one machine

6. What it means for the CUDA moat

NVIDIA's moat was never only the silicon. It is the software inertia that let a decade of kernels, compilers and tuning knowledge accrete in one place. This announcement aims at the most upstream point of that inertia: the language. It will not change who buys which card tomorrow, but it moves the cost curve of the alternative. Once a high-level language is good enough and has been validated on real workloads, the libraries, frameworks and tuning lore above it have a reason to follow. For any team that does not want its three-year roadmap pinned to a single vendor, that belongs on a watch list rather than in a news roundup.

📌 Frequently Asked Questions

What did DeepSeek announce on September 30, 2026?

In a post on its official WeChat account, DeepSeek said it has partnered with Huawei Technologies to develop programming tools for Huawei's Ascend chips and is open-sourcing programming infrastructure for the Ascend platform, including a high-level programming language called TileLang plus related compute and communication libraries.

Why is the high-level language the hardest layer to replace?

DeepSeek's own framing is that the first priority for a self-controlled GPU software ecosystem is a language that is universal, easy to program, and still able to reach the hardware's full performance potential. Libraries save weeks on a single port. A language decides who is allowed to write kernels, and that is exactly what has kept the CUDA ecosystem sticky.

What does DeepSeek claim about TileLang versus CUDA?

According to Reuters, DeepSeek said TileLang offers a simpler programming model than NVIDIA's CUDA, improving development efficiency and simplifying code logic. That is DeepSeek's claim; the reporting does not include independent third-party benchmarks.

Was there any hardware-level detail in the reporting?

Yes. Reuters reported that the two companies also progressed a supernode configuration built on 128 Ascend 950 chips, jointly optimising compute performance and chip-to-chip communication across the clustered system.

How should an engineering team evaluate a claim like this?

Ask for communication efficiency at your tensor-parallel width, not the vendor's demo width. Score the portability surface item by item: kernel language, collective communications, graph capture, quantisation kernels, profiler and debugger, and real-silicon CI. Then benchmark the same model, batch and dtype on both stacks and compare throughput ratio, p99 latency ratio, and whether weights must be retrained.