Crusoe Raises $3.9B at a $30.9B Valuation: The Economics of Owning Electrons to Tokens
On September 17, 2026, Crusoe announced the initial closing of a $3.9 billion Series F at a $30.9 billion post-money valuation. The oversubscribed round was co-led by Atreides Management, Mubadala Capital and Valor Equity Partners, with participation from Founders Fund, GIC, NVIDIA, the Qatar Investment Authority, Radical Ventures, TPG and others. Crusoe says its vertically integrated platform carries more than $140 billion in total contracted value, with over 6GW of gross contracted capacity and 1GW already delivered and operational. For developers, the signal is not the valuation. It is that the cost of a million tokens is increasingly set by who owns the power, the data centre and the inference stack.
1. The numbers first, the narrative second
A few figures are worth memorising. More than $140 billion in total contracted value. Over 6GW of gross contracted capacity across data centres and cloud, with 1GW already delivered and operational. More than 20x year-over-year growth in Crusoe Cloud bookings year to date. And CrusoeManaged Inference, launched last year, has already contracted over $100 million in annual recurring revenue, with the company claiming up to 9.9x faster time-to-first-token and 5x higher throughput than vLLM, powered by its MemoryAlloy technology. Crusoe also notes it is a Gartner Magic Quadrant Visionary, ranked first for inference speed on Artificial Analysis, and an NVIDIA Exemplar Cloud, with more than 1,800 employees across five countries.
# Time-to-first-token is the metric users feel, and it is the one an
# inference vendor's engine choice actually moves. Measure it yourself
# rather than reading a chart.
import statistics, time, httpx
def ttft(url, payload, token, n=9):
samples = []
for _ in range(n):
t0 = time.perf_counter()
with httpx.stream("POST", url,
headers={"Authorization": "Bearer " + token},
json={**payload, "stream": True}, timeout=60) as r:
for line in r.iter_lines():
if line and line.startswith("data:"):
samples.append(time.perf_counter() - t0)
break
return {"p50_ms": round(statistics.median(samples) * 1000),
"p90_ms": round(sorted(samples)[int(len(samples) * 0.9)] * 1000),
"n": n}
print(ttft("https://api.example-inference.com/v1/chat/completions",
{"model": "example-70b", "messages": [{"role": "user", "content": "hi"}]},
"TOKEN"))Vertical integration means owning power, data centres and the inference stack
2. What vertically integrated actually means
Conventional data centre developers treat power as a constraint to navigate after site selection. Crusoe starts from power, originating and managing it at the source, then builds AI-optimised data centres on top, then sells the AI cloud and inference on top of that. CEO Chase Lochmiller puts it as controlling the infrastructure from electrons to tokens. Atreides' Gavin Baker translates it into economics: as AI grows, the economics flow to the lowest-cost producer of intelligence, and owning the whole value chain is a structural advantage. You can read that as investor narrative if you like. Either way it raises a concrete question about your own bill: does your inference provider generate its own power, or resell someone else's compute?
# Throughput is what your bill is actually made of. Request throughput and
# token throughput pull in opposite directions as concurrency rises, so
# measure both and find the knee, not the peak.
import asyncio, time, httpx
async def one(client, url, payload, token, out):
t0 = time.perf_counter()
r = await client.post(url, headers={"Authorization": "Bearer " + token},
json=payload, timeout=120)
dt = time.perf_counter() - t0
body = r.json()
toks = body.get("usage", {}).get("output_tokens", 0)
out.append((dt, toks))
async def sweep(url, payload, token, concurrency):
out = []
async with httpx.AsyncClient() as client:
# warm up, then measure
await one(client, url, payload, token, out)
out.clear()
await asyncio.gather(*[
one(client, url, payload, token, out) for _ in range(concurrency)])
elapsed = max(d for d, _ in out)
tokens = sum(t for _, t in out)
return {"concurrency": concurrency,
"tokens_per_sec": round(tokens / elapsed, 1),
"req_per_sec": round(len(out) / elapsed, 2)}
print(asyncio.run(sweep("https://api.example-inference.com/v1/chat/completions",
{"model": "example-70b", "max_tokens": 512,
"messages": [{"role": "user", "content": "summarise"}]},
"TOKEN", 16)))3. Why developers should care
What developers feel is not a valuation. It is time-to-first-token, throughput, and price per million tokens. Those three are physically linked. If an inference engine gets the first token out faster and serves more requests per GPU, the same workload needs fewer cards, and the unit cost drops. That is why the up-to-9.9x time-to-first-token and 5x throughput claims matter: not for a leaderboard, but for running the same work on fewer GPUs. Crusoe Spark, the company's modular data centre unit, follows the same logic, turning capacity into a repeatable module instead of a one-off megaproject.
# A $/million-token headline is only comparable once you hold the workload
# fixed. Blend the numbers that actually move: cache hit rate, output length,
# and whether the provider charges for reasoning tokens.
RATES = { # USD per 1M tokens, illustrative -- substitute your vendors
"vendor-a": {"in_miss": 0.435, "in_hit": 0.0036, "out": 0.87},
"vendor-b": {"in_miss": 0.55, "in_hit": 0.05, "out": 1.10},
}
def blended(rate, steps, prefix_tokens, new_tokens, out_per_step):
cost = 0.0
for step in range(steps):
hits = prefix_tokens if step else 0
misses = new_tokens if step else prefix_tokens + new_tokens
cost += (misses * rate["in_miss"] + hits * rate["in_hit"]
+ out_per_step * rate["out"]) / 1e6
return round(cost, 5)
for name, rate in RATES.items():
print(name, blended(rate, steps=8, prefix_tokens=12_000,
new_tokens=800, out_per_step=400))Over 6GW contracted, with 1GW delivered and operational
4. An instructive footnote
The release notes that OpenAI trained Astra, its first artificial general intelligence, at the Abilene campus Crusoe designed and built. That sentence turns AI infrastructure from an abstract category into an industrial link with a named customer: training a frontier model needs someone who can line up power, land, buildings and cooling at once. In the same month, Crusoe added three board members, including Cloudflare CFO Thomas Seifert, former Digital Realty CEO Bill Stein, and Redwood Materials founder and CEO JB Straubel. Finance, real estate, materials. The composition is itself a comment on what AI infrastructure is made of.
# Second-order effects show up in the queue, not the price list. A provider
# with faster time-to-first-token can serve the same workload on fewer GPUs,
# which is exactly how an energy-first operator turns into a cheaper token.
def queueing(WORKLOAD_RPS, TTFT_s, TOKENS_PER_REQ, TPS_PER_REPLICA):
service_s = TOKENS_PER_REQ / TPS_PER_REPLICA
util = WORKLOAD_RPS * service_s
# Little's Law gives the average in-system concurrency you must provision.
in_system = WORKLOAD_RPS * (service_s + TTFT_s)
return {"utilisation": round(util, 2),
"replicas_needed": max(1, round(in_system)),
"verdict": "headroom" if util < 0.7 else "add capacity"}
print(queueing(WORKLOAD_RPS=12, TTFT_s=0.18,
TOKENS_PER_REQ=700, TPS_PER_REPLICA=90))5. Your procurement checklist for inference providers
Replace marketing numbers with your own measurements. First, measure time-to-first-token at p50 and p90, not just the mean, because users feel the tail. Second, plot throughput against concurrency and find the knee, not the peak. Third, compare price on a fixed workload, holding prefix length, output length and step count constant, with cache hit rate folded in; code sample 3 is a blended-cost calculator for exactly that. Fourth, price in queueing: a provider with faster time-to-first-token needs fewer replicas for the same load, which never appears on the price list. Fifth, and most important, stay swappable. Code sample 5 is a small function that ranks providers by meeting a latency budget and then by cost, and your architecture should let it run again every month.
# The durable lesson for teams: keep your inference layer swappable. Put the
# vendor behind one interface, keep prompts and tool schemas portable, and
# re-measure on a schedule, because the cost floor moves faster than your
# architecture does.
from dataclasses import dataclass
@dataclass
class Provider:
name: str
base_url: str
ttft_ms: int
cost_per_mtok: float
def rank(providers, ttft_budget_ms=250):
viable = [p for p in providers if p.ttft_ms <= ttft_budget_ms]
# Among providers that meet the latency budget, cost decides.
return sorted(viable, key=lambda p: p.cost_per_mtok)
print(rank([
Provider("a", "https://a.example/v1", 180, 0.87),
Provider("b", "https://b.example/v1", 340, 0.62),
Provider("c", "https://c.example/v1", 210, 0.74),
]))The cost of a token is ultimately set by whoever owns the cheapest power
6. What this means over the long run
Two things are becoming clear in 2026. Intelligence is getting cheaper fast, and the main source of that decline is not smaller models but supply-side vertical integration: whoever owns cheaper power, a more efficient inference engine and more modular data centres. That leads to two conclusions. First, do not treat inference capacity as a one-time architecture decision. The cheapest provider today may not be on the board in two years, so your prompts, tool schemas and evaluations should survive a provider swap. Second, watch for hidden lock-in. If a provider's low price comes from a proprietary inference engine and its own caching semantics, then your cost advantage is also a lock-in contract. Keep the interface, the data and the evaluations on your side, and you can keep choosing the lowest-cost producer of intelligence.
📌 Frequently Asked Questions
What did Crusoe announce?
On September 17, 2026, the initial closing of a $3.9 billion Series F at a $30.9 billion post-money valuation, co-led by Atreides Management, Mubadala Capital and Valor Equity Partners.
What are the headline scale numbers?
More than $140 billion in total contracted value, over 6GW of gross contracted capacity with 1GW delivered and operational, 20x-plus year-over-year growth in Crusoe Cloud bookings year to date, and over $100 million in ARR contracted for CrusoeManaged Inference.
What does vertically integrated AI infrastructure mean?
It means one company owns and operates power generation, AI-optimised data centres and an AI cloud and inference platform, which Crusoe's CEO describes as controlling the infrastructure from electrons to tokens.
Why does this matter to developers?
Time-to-first-token, throughput and price per million tokens are physically linked. Faster engines and higher per-GPU throughput mean lower unit cost, which shows up in the inference providers and prices you choose.
How should I compare inference providers?
Measure on your own workload: p50 and p90 time-to-first-token, the throughput knee as concurrency rises, blended cost on a fixed workload including cache hits, and the replicas a given load requires. Then keep the layer swappable.
🔧 Recommended Tools
API Response Time Calculator
Measure time-to-first-token and tail latency
AI Token Counter
Turn price lists into your blended cost
AI Data Analyzer
Analyse throughput against concurrency
API Rate Limit Calculator
Estimate replicas for a given load
API Documentation Generator
Document a swappable inference interface
📚 Sources
- Crusoe — Crusoe Raises $3.9 Billion Series F for its Vertically-Integrated AI Infrastructure Platform (Sep 17, 2026)
- Crusoe — Crusoe Adds Cloudflare CFO Thomas Seifert, Bill Stein and JB Straubel to its Board (Sep 17, 2026)
- TechCrunch — Crusoe raises $3.9B to build massive data centers and small modular AI factories (Sep 17, 2026)
- Crusoe — Crusoe Cloud and CrusoeManaged Inference