Crusoe 融资 39 亿美元、估值 309 亿:从电子到 token 的垂直整合经济学

·阅读约11分钟·Evergreen Tools Team

2026 年 9 月 17 日,Crusoe 宣布其 39 亿美元 F 轮融资完成首次交割,投后估值 309 亿美元。这是一轮超额认购,由 Atreides Management、Mubadala Capital 与 Valor Equity Partners 联合领投,Founders Fund、GIC、NVIDIA、卡塔尔投资局(QIA)、Radical Ventures 与 TPG 等新老投资方参与。公司称,其垂直整合平台的总合同价值(TCV)已超过 1400 亿美元,已签容量超过 6GW,其中 1GW 已交付并投入运营。对开发者而言,这条新闻的信号不是估值本身,而是推理的单位成本正在被「谁拥有发电、数据中心和推理栈」这件事决定。

一、先看数字,再看叙事

发布稿里几个数字值得记住。总合同价值超过 1400 亿美元;数据中心与云业务合计已签约容量超过 6GW,其中 1GW 已交付并运营;Crusoe Cloud 今年迄今的订单同比增长超过 20 倍;去年上线的 CrusoeManaged Inference 已经签下超过 1 亿美元的年度经常性收入(ARR),公司称其推理引擎的首次 token 时延最高快 9.9 倍、吞吐比 vLLM 高 5 倍,依托其 MemoryAlloy 技术。公司同时被列为 Gartner 魔力象限「远见者」、Artificial Analysis 推理速度第一、NVIDIA Exemplar Cloud,员工超过 1800 人、分布在五个国家。

# Time-to-first-token is the metric users feel, and it is the one an
# inference vendor's engine choice actually moves. Measure it yourself
# rather than reading a chart.

import statistics, time, httpx

def ttft(url, payload, token, n=9):
    samples = []
    for _ in range(n):
        t0 = time.perf_counter()
        with httpx.stream("POST", url,
                          headers={"Authorization": "Bearer " + token},
                          json={**payload, "stream": True}, timeout=60) as r:
            for line in r.iter_lines():
                if line and line.startswith("data:"):
                    samples.append(time.perf_counter() - t0)
                    break
    return {"p50_ms": round(statistics.median(samples) * 1000),
            "p90_ms": round(sorted(samples)[int(len(samples) * 0.9)] * 1000),
            "n": n}

print(ttft("https://api.example-inference.com/v1/chat/completions",
           {"model": "example-70b", "messages": [{"role": "user", "content": "hi"}]},
           "TOKEN"))
从电子到 token

垂直整合的意思是:发电、数据中心与推理栈都由一家所有

二、所谓「垂直整合」到底整合了什么

常规数据中心开发商把电力当作选址之后要去规避的约束;Crusoe 的路线是反过来的——从源头直接发起并管理电力,再往上盖 AI 优化的数据中心,再往上提供 AI 云与推理服务。CEO Chase Lochmiller 的说法是:要到达那个丰裕的时代,就必须「从电子到 token」控制整条基础设施。Atreides 的 Gavin Baker 把它讲成了经济学:随着 AI 增长,经济性会流向「最低成本的智能生产者」,而垂直整合恰好让一家公司拥有整条价值链。你可以把这套说法理解为投资叙事;但对你的账单而言,它确实提出了一个具体问题——你的推理供应商,是自己发电,还是转售别人的算力?

# Throughput is what your bill is actually made of. Request throughput and
# token throughput pull in opposite directions as concurrency rises, so
# measure both and find the knee, not the peak.

import asyncio, time, httpx

async def one(client, url, payload, token, out):
    t0 = time.perf_counter()
    r = await client.post(url, headers={"Authorization": "Bearer " + token},
                          json=payload, timeout=120)
    dt = time.perf_counter() - t0
    body = r.json()
    toks = body.get("usage", {}).get("output_tokens", 0)
    out.append((dt, toks))

async def sweep(url, payload, token, concurrency):
    out = []
    async with httpx.AsyncClient() as client:
        # warm up, then measure
        await one(client, url, payload, token, out)
        out.clear()
        await asyncio.gather(*[
            one(client, url, payload, token, out) for _ in range(concurrency)])
    elapsed = max(d for d, _ in out)
    tokens = sum(t for _, t in out)
    return {"concurrency": concurrency,
            "tokens_per_sec": round(tokens / elapsed, 1),
            "req_per_sec": round(len(out) / elapsed, 2)}

print(asyncio.run(sweep("https://api.example-inference.com/v1/chat/completions",
                        {"model": "example-70b", "max_tokens": 512,
                         "messages": [{"role": "user", "content": "summarise"}]},
                        "TOKEN", 16)))

三、为什么这跟写代码的人有关

开发者感受到的不是估值,而是三件事:首 token 时延、吞吐,以及每百万 token 的价格。这三件事在物理上高度相关:如果推理引擎的首 token 更快、单位 GPU 能服务的请求更多,同一份工作负载需要的卡就更少,单位成本就能更低。这也是发布稿里那个「最高快 9.9 倍首 token、5 倍吞吐」值得关注的原因——它不是为了排行榜,而是为了「更少卡跑同样多的活」。Crusoe 的模块化 Crusoe Spark 数据中心单元,也是同一个逻辑:把容量做成可复制的模块,而不是一次性的大工程。

# A $/million-token headline is only comparable once you hold the workload
# fixed. Blend the numbers that actually move: cache hit rate, output length,
# and whether the provider charges for reasoning tokens.

RATES = {  # USD per 1M tokens, illustrative -- substitute your vendors
    "vendor-a": {"in_miss": 0.435, "in_hit": 0.0036, "out": 0.87},
    "vendor-b": {"in_miss": 0.55,  "in_hit": 0.05,   "out": 1.10},
}

def blended(rate, steps, prefix_tokens, new_tokens, out_per_step):
    cost = 0.0
    for step in range(steps):
        hits = prefix_tokens if step else 0
        misses = new_tokens if step else prefix_tokens + new_tokens
        cost += (misses * rate["in_miss"] + hits * rate["in_hit"]
                 + out_per_step * rate["out"]) / 1e6
    return round(cost, 5)

for name, rate in RATES.items():
    print(name, blended(rate, steps=8, prefix_tokens=12_000,
                        new_tokens=800, out_per_step=400))
已签容量与交付容量

已签容量超 6GW,其中 1GW 已交付并运营

四、一个有意思的注脚

发布稿提到,OpenAI 在 Crusoe 设计并建造的 Abilene 园区训练了其首个通用人工智能(AGI)Astra。这句话的分量在于它把「AI 基础设施」从一个抽象品类,变成了一个有具体客户的产业环节:训练前沿模型这件事,需要有人能同时搞定电、地、机房和散热。同一个 9 月,Crusoe 还宣布三位新董事加入,包括 Cloudflare 的 CFO Thomas Seifert、Digital Realty 前 CEO Bill Stein 以及 Redwood Materials 创始人兼 CEO JB Straubel——一位金融、一位地产、一位材料,这个组合本身就是对「AI 基础设施是能源+地产+材料」这句话的注解。

# Second-order effects show up in the queue, not the price list. A provider
# with faster time-to-first-token can serve the same workload on fewer GPUs,
# which is exactly how an energy-first operator turns into a cheaper token.

def queueing(WORKLOAD_RPS, TTFT_s, TOKENS_PER_REQ, TPS_PER_REPLICA):
    service_s = TOKENS_PER_REQ / TPS_PER_REPLICA
    util = WORKLOAD_RPS * service_s
    # Little's Law gives the average in-system concurrency you must provision.
    in_system = WORKLOAD_RPS * (service_s + TTFT_s)
    return {"utilisation": round(util, 2),
            "replicas_needed": max(1, round(in_system)),
            "verdict": "headroom" if util < 0.7 else "add capacity"}

print(queueing(WORKLOAD_RPS=12, TTFT_s=0.18,
               TOKENS_PER_REQ=700, TPS_PER_REPLICA=90))

五、给你的采购清单:怎么比推理供应商

把营销数字换成你自己的测量。第一,测首 token 时延,而且要测 p50 和 p90,不要只看平均值——用户感受到的是尾部。第二,测吞吐随并发的曲线,找到拐点,而不是峰值。第三,把价格放到固定工作负载上比较:同样的前缀长度、同样的输出长度、同样的步数,用缓存命中率折算;示例 3 就是一个混合成本计算器。第四,把排队效应算进去:首 token 更快的供应商,在同等负载下需要的副本更少,这是价格表里看不到的成本。第五,也是最重要的一条——保持可切换:示例 5 展示了一个按「先满足延迟预算、再比成本」排序的小函数,你的架构应该允许它每月重新跑一次。

# The durable lesson for teams: keep your inference layer swappable. Put the
# vendor behind one interface, keep prompts and tool schemas portable, and
# re-measure on a schedule, because the cost floor moves faster than your
# architecture does.

from dataclasses import dataclass

@dataclass
class Provider:
    name: str
    base_url: str
    ttft_ms: int
    cost_per_mtok: float

def rank(providers, ttft_budget_ms=250):
    viable = [p for p in providers if p.ttft_ms <= ttft_budget_ms]
    # Among providers that meet the latency budget, cost decides.
    return sorted(viable, key=lambda p: p.cost_per_mtok)

print(rank([
    Provider("a", "https://a.example/v1", 180, 0.87),
    Provider("b", "https://b.example/v1", 340, 0.62),
    Provider("c", "https://c.example/v1", 210, 0.74),
]))
单位成本决定权

推理的单位成本,最终由谁拥有最便宜的电决定

六、这件事对团队的长期含义

有几件事在 2026 年变得越来越清楚:智能的价格在快速下降,而下降的主要来源不是模型变小,而是**供给侧的垂直整合**——谁拥有更便宜的电、更高效的推理引擎、更模块化的机房。这意味着两件事。第一,不要把推理能力当成一次性的架构决策:今天最便宜的供应商,两年后不一定还在榜上,而你的提示词、工具 schema 与评估应该能扛住一次换供应商。第二,警惕「隐性锁定」:如果一个供应商的便宜来自它专有的推理引擎与缓存语义,那么你的成本优势同时也是一张锁定合同。把接口、数据与评估留在你这一侧,你就能一直在最低成本的智能生产者之间自由选择。

📌 常见问题 FAQ

Crusoe 宣布了什么?

2026 年 9 月 17 日宣布其 39 亿美元 F 轮融资完成首次交割,投后估值 309 亿美元;本轮由 Atreides Management、Mubadala Capital 与 Valor Equity Partners 联合领投。

公司的规模数据有哪些?

总合同价值超过 1400 亿美元;已签约容量超过 6GW,其中 1GW 已交付并运营;Crusoe Cloud 今年迄今订单同比增长超过 20 倍;CrusoeManaged Inference 已签约超 1 亿美元 ARR。

什么是「垂直整合」的 AI 基础设施?

指一家公司同时拥有并运营发电、AI 优化数据中心与 AI 云/推理服务,也就是 CEO 所说「从电子到 token」控制整条价值链。

这和开发者有什么关系?

推理的首 token 时延、吞吐与每百万 token 价格密切相关;更快的引擎与更高的单位 GPU 吞吐意味着更低的单位成本,而这会体现在你选的推理服务与账单上。

我该怎么比较推理供应商?

用你自己的工作负载测量:首 token 时延的 p50/p90、吞吐随并发的拐点、固定负载下的混合成本(含缓存命中),以及同负载下需要的副本数;并保持可切换。