Xiaomi Open-Sources MiMo-V2.6: A Trillion-Parameter Omnimodal Model Tops Open Weights

·11 min read·Evergreen Tools Team

On September 21 and 22, 2026, Xiaomi released and open-sourced the MiMo-V2.6 series: two natively omnimodal models, MiMo-V2.6-Pro and MiMo-V2.6-Flash, plus a Pro-UltraSpeed variant for latency-sensitive work. Pro is Xiaomi's most powerful flagship reasoning model and, in the company's words, trillion-parameter. Both checkpoints take text, images, video, and audio as input and carry a one-million-token context window with up to 128K output tokens. Weights are open, the API is OpenAI- and Anthropic-compatible, and Xiaomi's own claim is blunt: MiMo-V2.6-Pro scores 46.32 on Artificial Analysis' composite intelligence index, which Xiaomi calls the strongest open-source result today. Claims are cheap, so here is what is verifiable and what it changes for a working developer.

1. The specs, from the source

The official model pages list a 1M-token context window, 128K max output, 100 requests per minute, and 10 million tokens per minute. Input is text and, through the omnimodal capability, images, video, and audio; output is text. Capabilities include tool calling, streaming, web search, structured output, context caching, and a deep-thinking mode you can toggle. Xiaomi positions Pro for long-horizon and high-stakes work such as coding, research, cybersecurity and computer operation, and Flash as the balance of intelligence, cost and token efficiency for high-frequency calls. The third model, Pro-UltraSpeed, targets latency-sensitive workloads.

# The migration path is deliberately boring: the API is OpenAI-compatible,
# so pointing an existing tool at MiMo is a base URL and a model name.

from openai import OpenAI

client = OpenAI(
    api_key="MIMO_API_KEY",
    base_url="https://api.xiaomimimo.com/v1",   # official endpoint
)

resp = client.chat.completions.create(
    model="mimo-v2.6-flash",                    # or mimo-v2.6-pro
    messages=[
        {"role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi."},
        {"role": "user", "content": "Review this function for edge cases."},
    ],
    max_tokens=1024,
)
print(resp.choices[0].message.content)
A trillion-parameter omnimodal model

Xiaomi calls MiMo-V2.6-Pro its trillion-parameter flagship

2. Pricing is the actual headline

Open weights get the attention, but the pricing table is what changes procurement. On Xiaomi's USD pay-as-you-go rates, MiMo-V2.6-Pro costs $0.435 per million input tokens on a cache miss, $0.0036 on a cache hit, and $0.87 per million output tokens. Flash is $0.14 per million input on a cache miss, $0.0028 on a cache hit, and $0.28 per million output. The cache-hit numbers are not a rounding detail: agent workflows repeatedly resend the same system prompt, tool schemas, and repository context, so the hit rate is where the cost is decided. A cache hit at roughly a hundred-and-twentieth of the miss price rewards exactly the pattern agents produce.

# The same endpoint also speaks the Anthropic protocol, which is what makes
# a CLI agent swap a two-line change. Xiaomi documents Claude Code and Cline.

# Claude Code / Cline environment variables
export ANTHROPIC_BASE_URL="https://api.xiaomimimo.com"
export ANTHROPIC_AUTH_TOKEN="$MIMO_API_KEY"
export ANTHROPIC_MODEL="mimo-v2.6-pro"

# Keep the model name in a variable so you can A/B it:
#   ANTHROPIC_MODEL="mimo-v2.6-flash"   # cheaper, higher throughput
claude "summarise the failing test and propose a fix"

3. Why native omnimodal is more than a label

Natively omnimodal means one set of weights handles text, image, video, and audio jointly, rather than routing each modality to a separate specialist model stitched together behind a router. For developers, the practical difference is fewer moving parts and one context to manage. Xiaomi's own demos illustrate the intended range: coordinating multiple agents to build a 3D open-world game from images, video, or text; a materials-research copilot that reviews literature, forms hypotheses, simulates binding strength, and shortlists candidates; and code-driven content creation that turns concepts like Fourier decomposition into animations. Not every team needs any of that, but the pattern worth noticing is the model being asked to plan and coordinate, not just answer.

# Native omnimodality means one set of weights handles images, video and
# audio jointly -- no router stitching separate specialist models together.

import base64, httpx

def describe(image_path: str, question: str) -> str:
    with open(image_path, "rb") as fh:
        b64 = base64.b64encode(fh.read()).decode()
    r = httpx.post(
        "https://api.xiaomimimo.com/v1/chat/completions",
        headers={"Authorization": "Bearer MIMO_API_KEY"},
        json={
            "model": "mimo-v2.6-pro",
            "messages": [{
                "role": "user",
                "content": [
                    {"type": "text", "text": question},
                    {"type": "image_url",
                     "image_url": {"url": "data:image/png;base64," + b64}},
                ],
            }],
        }, timeout=120)
    return r.json()["choices"][0]["message"]["content"]
Agent coordination and multimodality

In Xiaomi's demos it splits tasks and coordinates multiple agents

4. Open weights under MIT, and a real technical report

The weights are released under the MIT licence, which permits commercial use and further training. Xiaomi also published a technical report alongside the models, and independent analysis has begun. Sebastian Raschka's read of the architecture is refreshingly unromantic: Pro is largely a conventional design, grouped-query attention with a small sliding-window attention layer, and the gains come mostly from data and post-training recipe improvements rather than exotic attention variants. The report also describes training across multiple agent harnesses and an agentic grader that inspects execution traces rather than only checking final answers. If you care about agent quality, that detail, rewarding the trace rather than just the answer, is the transferable idea.

# Cache hits are where the bill is decided in agent loops, because the same
# system prompt and tool schemas get resent on every step. Check your rate
# before you move traffic -- a hit costs roughly a hundredth of a miss.

PRICES = {                      # USD per 1M tokens, Xiaomi pay-as-you-go
    "mimo-v2.6-pro":   {"in_miss": 0.435, "in_hit": 0.0036, "out": 0.87},
    "mimo-v2.6-flash": {"in_miss": 0.14,  "in_hit": 0.0028, "out": 0.28},
}

def run_cost(model, in_miss, in_hit, out, steps=8):
    total = 0.0
    for _ in range(steps):      # each step resends the cached prefix
        total += (in_miss * PRICES[model]["in_miss"]
                  + in_hit * PRICES[model]["in_hit"]
                  + out * PRICES[model]["out"]) / 1e6
    return round(total, 4)

print(run_cost("mimo-v2.6-pro", 12_000, 12_000, 800))
print(run_cost("mimo-v2.6-flash", 12_000, 12_000, 800))

5. How to try it without committing

The migration path is deliberately boring, which is a compliment. Xiaomi's API is compatible with both the OpenAI and Anthropic protocols, so pointing an existing tool at it is usually a base URL and a model name change; Xiaomi documents Claude Code and Cline as supported tools. Code sample 1 shows the OpenAI-compatible swap and code sample 2 the Anthropic-style environment variables for a CLI agent. Before you move real traffic, do two things: benchmark on your own tasks rather than trusting a composite index, and run the same prompt twice to measure the cache-hit rate you actually get, because that is the number your bill will reflect. Code sample 4 is a small blended-cost calculator for exactly that.

# Route by task, not by habit: keep the cheap model for the bulk of routine
# calls and spend the flagship only where the task actually needs it.

ROUTES = {
    "classify":       "mimo-v2.6-flash",
    "summarise":      "mimo-v2.6-flash",
    "extract":        "mimo-v2.6-flash",
    "code_review":    "mimo-v2.6-pro",
    "long_plan":      "mimo-v2.6-pro",
    "multimodal_doc": "mimo-v2.6-pro",
}

def pick(task: str, complexity: float) -> str:
    base = ROUTES.get(task, "mimo-v2.6-flash")
    if complexity > 0.8 and base.endswith("flash"):
        return "mimo-v2.6-pro"   # escalate only when the task earns it
    return base

for job in ["classify", "code_review", "long_plan"]:
    print(job, "->", pick(job, 0.9))
Open weights and pricing

MIT-licensed weights plus cache-aware API pricing

6. What it means for the open-weights race

A phone maker shipping a trillion-parameter omnimodal model under MIT, with a technical report and cache-aware pricing, is a statement about how the open-weights market now works. Capability is no longer the only axis; licence, context economics, and tool compatibility are. For teams, that is good news: the decision is no longer frontier-or-nothing. You can run a capable open model for the bulk of routine work, keep a hosted model for the hard tail, and switch between them behind one protocol. The teams that benefit most are the ones that keep their prompts, tools and evaluation in shapes that survive a model swap, which is a design habit rather than a purchase.

📌 Frequently Asked Questions

What did Xiaomi release?

The MiMo-V2.6 series, released on September 21 and 22, 2026: MiMo-V2.6-Pro, MiMo-V2.6-Flash, and a Pro-UltraSpeed variant, with open weights and an API.

What is the context window?

The official pages list a 1M-token context window with up to 128K output tokens.

Is it genuinely open source?

Yes. The weights are released under the MIT licence, permitting commercial use and further training, and Xiaomi published a technical report alongside the models.

What does the API cost?

On Xiaomi's USD pay-as-you-go rates, Pro is $0.435 per million input tokens on a cache miss, $0.0036 on a cache hit, and $0.87 per million output; Flash is $0.14, $0.0028 and $0.28 respectively.

Can I use it with existing tools?

Yes. The API is compatible with both the OpenAI and Anthropic protocols, and Xiaomi documents tools including Claude Code and Cline.