MiniCPM5-2B:2.5B 开源模型,在笔记本上跑赢 4B 对手

·阅读约11分钟·Evergreen Tools Team

2026 年 9 月 7 日,OpenBMB 发布了 MiniCPM5-2B,这是 MiniCPM5 系列的第二款模型,接在 MiniCPM5-1B 之后。它是一个稠密的 20 亿级 Transformer,沿用并放大了同一套训练配方,明确为端侧、本地部署与资源受限场景而设计。权重以 Apache 2.0 许可开源,原生上下文窗口 131,072 token。OpenBMB 自己的评测显示,它在 34 项基准上平均 53.9 分,高过对比集里所有更大的模型(最高 51.1);Artificial Analysis 则独立给出「4B 以下开源权重模型中最高分」的结论。小模型值得看的原因很简单:不是每个 Agent 任务都值得把数据送到云端。

一、规格,来自模型卡

MiniCPM5-2B 是一个稠密因果语言模型,参数量 2,516,756,480,其中 1,981,982,720 位于嵌入层之外。42 层,分组查询注意力(16 个 query 头、2 个 key/value 头),原生上下文窗口 131,072 token,结构是标准的 LlamaForCausalLM。许可为 Apache 2.0——可以商用、可以继续训练、不必谈判。部署路径覆盖 Transformers、vLLM、SGLang 与 Docker Model Runner,官方模型卡还提到配套的「部署/微调 Agent Skills」,以及针对 NVIDIA 平台的 FlagOS 适配版本。换句话说,它不是一篇论文,而是一个可以今天就拉下来跑的权重。

# Serving it locally is one command if you already run vLLM. The weights are
# Apache 2.0, so commercial use needs no negotiation.

# vLLM (OpenAI-compatible server)
vllm serve openbmb/MiniCPM5-2B \
  --served-model-name minicpm5-2b \
  --max-model-len 131072 \
  --port 8000

# Then talk to it like any OpenAI endpoint.
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"minicpm5-2b","messages":[{"role":"user","content":"Explain GQA in two sentences."}],"max_tokens":200}'
本地部署的小模型

25 亿参数、Apache 2.0,为设备端与资源受限场景而做

二、53.9 这个数字该怎么读

OpenBMB 把 MiniCPM5-2B 与同尺寸的 LFM2.5-2.6B、Qwen3.5-2B、Gemma-4-E2B-it 对比,同时列出更大的 Qwen3.5-4B、granite-4.2-3B、Nemotron-3-Nano-4B 等作参考。在这个对比集里,它的平均分 53.9 高过所有更大的模型,最高的 Qwen3.5-4B 是 51.1。优势最明显的地方是代码推理(LiveCodeBench v6 得 69.1,对比 56.4)、数学推理(AIME 2025 得 86.5)、长上下文(NoLiMa 68.1)、工具调用(BFCL v4 66.6)以及若干 Agent 任务。要提醒的是:这是厂商自己设定的对比集,62 分的差距不大,真正的结论应该是「2B 级别已经可以承担相当一部分实际工作」,而不是「它全面超过 4B」。

# If you would rather not run a server, the Transformers path is short.
# Everything stays on the machine, which is the point of a 2B-class model.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "openbmb/MiniCPM5-2B"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, trust_remote_code=True, dtype="bfloat16", device_map="auto")

messages = [{"role": "user", "content": "Summarise this diff in one paragraph."}]
prompt = tok.apply_chat_template(messages, tokenize=False,
                                 add_generation_prompt=True)
ids = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**ids, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][ids["input_ids"].shape[-1]:], skip_special_tokens=True))

三、独立评测的验证与保留

Artificial Analysis 在 9 月 7 日的评述里给了独立参照:MiniCPM5-2B 在 Intelligence Index v4.2 上得 15 分,是 4B 以下开源权重模型中最高,比 Granite 4.2 3B(11 分)高 4 分,也比参数多约 44% 的 Qwen3.5 4B(14 分,估算)高 1 分。它的 Agent 能力在同等尺寸里确实突出:GDPval-AA v2 的 Elo 为 831,在 4B 以下模型里领先;τ³-Banking 上以 21% 并列第一。它同时明确指出短板:知识、代码与长上下文的硬仗仍然吃亏——Humanity's Last Exam 9%、Terminal-Bench v2.1 9%、CritPt 0%。还有一点很有意思:它在一个知识类评测上的低分,是「拒答」换来的——只尝试了 29% 的问题,非幻觉率因此达到 78%。对小模型来说,会拒答往往比胡乱自信更有用。

# Tool use is where a small model earns its place in an agent. Keep the tool
# schemas small and the decisions binary -- that is exactly the regime the
# model is strongest in.

TOOLS = [{
    "type": "function",
    "function": {
        "name": "read_file",
        "description": "Read a UTF-8 text file from the workspace.",
        "parameters": {"type": "object",
                       "properties": {"path": {"type": "string"}},
                       "required": ["path"]},
    },
}]

def plan(client, user_request):
    r = client.chat.completions.create(
        model="minicpm5-2b",
        messages=[{"role": "user", "content": user_request}],
        tools=TOOLS, tool_choice="auto", max_tokens=256)
    msg = r.choices[0].message
    if msg.tool_calls:
        call = msg.tool_calls[0]
        # Parse and validate before executing anything.
        return {"tool": call.function.name, "args": json.loads(call.function.arguments)}
    return {"tool": None, "answer": msg.content}
工具调用与本地 Agent

代码推理、工具调用与 Agent 任务,是它在同尺寸里的强项

四、为什么「端侧」这次不是噱头

对小模型最实际的批评是「跑得动但不好用」。MiniCPM5-2B 的定位绕开了这一点,因为它瞄准的是明确受限的场景:本地助手、编码 Agent、工具调用工作流,以及必须在设备上完成的推理。把这些场景拆开看,需求是清楚的——数据不能出机器、调用频率很高但每次都简单、延迟要求苛刻。这类工作恰恰是 2B 模型的主场,也是云端大模型最不划算的地方。Artificial Analysis 还注意到它作为推理模型却很省 token(每个任务约 1.9 万输出 token,与 Granite 4.2 3B 并列最低),而输出 token 数量正是端侧部署最敏感的成本项。

# Size the machine before you promise anything. Weights dominate memory at
# batch size 1, and the KV cache grows with context -- the two numbers are
# separate budgets, and the second one bites on 131K contexts.

def memory_budget(params_b=2.52, bits=16, ctx=32_768,
                  layers=42, kv_heads=2, head_dim=128, batch=1):
    bytes_per_param = bits / 8
    weights_gb = params_b * 1e9 * bytes_per_param / 1e9
    kv_gb = (2 * layers * kv_heads * head_dim * ctx * batch * 2) / 1e9
    return {"weights_gb": round(weights_gb, 2),
            "kv_cache_gb": round(kv_gb, 2),
            "total_gb": round(weights_gb + kv_gb, 2),
            "note": "KV scales with context; trim ctx before shrinking the model"}

print(memory_budget(bits=16, ctx=32_768))
print(memory_budget(bits=4,  ctx=131_072))

五、怎么把它接进你的本地 Agent

落地路径很短。第一步,用 vLLM 起一个 OpenAI 兼容的服务(示例 1),或者直接用 Transformers 加载(示例 2),两者都在本机完成。第二步,给工具调用留出空间:把工具 schema 写小、把决策做成分支明确的二选一,这正是小模型最强的区间(示例 3)。第三步,先把显存算清楚:权重和 KV 缓存是两笔独立预算,131K 上下文下 KV 才是那个会咬人的数字,示例 4 给了一个可调的计算函数。第四步,做好路由:把「必须留在本地」和「高频但简单」的任务交给本地模型,把困难的长尾留给云端(示例 5)。第五步,别只看综合分——用你自己的 20 条任务测一遍,小模型的短板往往只在你特定的工作负载上才暴露。

# Route by privacy and cost, not by habit. Local wins on data that must not
# leave the machine and on high-frequency trivial calls; the cloud keeps the
# hard tail.

LOCAL_FIRST = {"classify", "extract", "format", "rename", "summarise_small"}

def route(task, doc_tokens, sensitivity):
    if sensitivity == "restricted":
        return "local"            # the file never leaves the machine
    if task in LOCAL_FIRST and doc_tokens <= 100_000:
        return "local"            # cheap, fast, good enough
    return "cloud"                # hard reasoning, huge context, long horizon

print(route("extract", 4_000, "restricted"))
print(route("long_plan", 220_000, "internal"))
尺寸与能力的前沿

Artificial Analysis:4B 以下开源权重模型中最高分

六、它对开源生态意味着什么

MiniCPM5-2B 与它一同发布的还有训练数据:UltraX 网页预训练集、分层的 UltraData-Code、50 万条 Agent 训练样本的 UltraData-SFT-Agent-2609,以及 8 万余条覆盖数学、代码、常识与长上下文推理的 UltraData-RL-2609。把数据一起开放,比只放权重更有意义——它让「这个尺寸能做到什么」变成一个可以复现、也可以继续推进的工程问题。如果再叠加上 9 月里那些万亿参数的开源旗舰,2026 年的开源权重格局已经很清楚了:不是「能不能用」的问题,而是「在哪个尺寸、哪个许可、哪种成本下用」,而这个选择权正在回到团队自己手里。

📌 常见问题 FAQ

MiniCPM5-2B 是什么?

OpenBMB 于 2026 年 9 月 7 日发布的 25 亿参数稠密语言模型,MiniCPM5 系列第二款(接 MiniCPM5-1B),Apache 2.0 开源,面向端侧与资源受限场景。

它的规格如何?

2,516,756,480 参数(其中 1,981,982,720 在嵌入层之外)、42 层、分组查询注意力(16 个 query 头 / 2 个 KV 头)、原生 131,072 token 上下文,结构为 LlamaForCausalLM。

评测成绩怎么样?

OpenBMB 自有对比集中 34 项基准平均 53.9 分,高于对比集内所有更大模型(最高 51.1);Artificial Analysis 给出 15 分的 Intelligence Index v4.2,是 4B 以下开源权重模型最高。

它有哪些短板?

按 Artificial Analysis 的独立评测,知识、代码与长上下文的硬指标偏弱,例如 Humanity's Last Exam 9%、Terminal-Bench v2.1 9%、CritPt 0%;优势集中在工具调用与若干 Agent 任务。

怎么部署?

支持 Transformers、vLLM、SGLang 与 Docker Model Runner 等路径;权重以 Apache 2.0 许可发布,可商用与继续训练。