MiniCPM5-2B: A 2.5B Open Model That Beats 4B Rivals on Your Laptop
On September 7, 2026, OpenBMB released MiniCPM5-2B, the second model in the MiniCPM5 series after MiniCPM5-1B. It is a dense 2B-class Transformer that scales up the same training recipe, built explicitly for on-device, local deployment and resource-constrained scenarios. The weights are open under Apache 2.0 with a native 131,072-token context window. OpenBMB's own evaluation reports a 53.9 average across its 34-benchmark set, above every larger model it includes, the highest of which is 51.1. Artificial Analysis independently scored it highest of any open-weights model under 4B total parameters. Small models matter for a simple reason: not every agent task deserves to send your data to the cloud.
1. The specs, from the model card
MiniCPM5-2B is a dense causal language model with 2,516,756,480 parameters, of which 1,981,982,720 sit outside the embeddings. It uses 42 layers and grouped-query attention with 16 query heads and 2 key/value heads, has a native context window of 131,072 tokens, and is built on the standard LlamaForCausalLM architecture. The licence is Apache 2.0, which means commercial use and further training without negotiation. Deployment paths cover Transformers, vLLM, SGLang and Docker Model Runner, and the model card mentions companion deployment and fine-tuning Agent Skills plus NVIDIA-focused FlagOS builds. In other words, this is not a paper. It is weights you can pull today.
# Serving it locally is one command if you already run vLLM. The weights are
# Apache 2.0, so commercial use needs no negotiation.
# vLLM (OpenAI-compatible server)
vllm serve openbmb/MiniCPM5-2B \
--served-model-name minicpm5-2b \
--max-model-len 131072 \
--port 8000
# Then talk to it like any OpenAI endpoint.
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"minicpm5-2b","messages":[{"role":"user","content":"Explain GQA in two sentences."}],"max_tokens":200}'2.5B parameters, Apache 2.0, built for on-device and constrained use
2. How to read the 53.9
OpenBMB compares MiniCPM5-2B against same-size peers including LFM2.5-2.6B, Qwen3.5-2B and Gemma-4-E2B-it, while listing larger models such as Qwen3.5-4B, granite-4.2-3B and Nemotron-3-Nano-4B for reference. Within that comparison set its 53.9 average beats every larger model, the highest of which is Qwen3.5-4B at 51.1. The advantages are most visible in code reasoning, where it scores 69.1 on LiveCodeBench v6 against 56.4, in mathematics with 86.5 on AIME 2025, in long-context understanding with 68.1 on NoLiMa, in tool use with 66.6 on BFCL v4, and across several agentic tasks. The caveat is that this is a vendor-defined comparison set and a two-point lead is small. The honest conclusion is that the 2B class can now carry a real share of production work, not that it beats every 4B model everywhere.
# If you would rather not run a server, the Transformers path is short.
# Everything stays on the machine, which is the point of a 2B-class model.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "openbmb/MiniCPM5-2B"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id, trust_remote_code=True, dtype="bfloat16", device_map="auto")
messages = [{"role": "user", "content": "Summarise this diff in one paragraph."}]
prompt = tok.apply_chat_template(messages, tokenize=False,
add_generation_prompt=True)
ids = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**ids, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][ids["input_ids"].shape[-1]:], skip_special_tokens=True))3. Independent verification, and the caveats
Artificial Analysis published an independent read on September 7. MiniCPM5-2B scores 15 on its Intelligence Index v4.2, the highest of any open-weights model under 4B total parameters, four points clear of Granite 4.2 3B at 11, and a point ahead of Qwen3.5 4B at 14 despite about 44 percent fewer parameters. Its agentic performance at this size is genuinely strong: an Elo of 831 on GDPval-AA v2, leading models under 4B, and joint-first at 21 percent on tau-squared-Banking. It is equally clear about the weaknesses. Knowledge, coding and long-context heavy lifting still suffer: 9 percent on Humanity's Last Exam, 9 percent on Terminal-Bench v2.1, and 0 percent on CritPt. One detail is worth noting: its low score on a knowledge evaluation comes from abstaining rather than guessing, attempting only 29 percent of questions for a 78 percent non-hallucination rate. For a small model, knowing when not to answer is often worth more than confidence.
# Tool use is where a small model earns its place in an agent. Keep the tool
# schemas small and the decisions binary -- that is exactly the regime the
# model is strongest in.
TOOLS = [{
"type": "function",
"function": {
"name": "read_file",
"description": "Read a UTF-8 text file from the workspace.",
"parameters": {"type": "object",
"properties": {"path": {"type": "string"}},
"required": ["path"]},
},
}]
def plan(client, user_request):
r = client.chat.completions.create(
model="minicpm5-2b",
messages=[{"role": "user", "content": user_request}],
tools=TOOLS, tool_choice="auto", max_tokens=256)
msg = r.choices[0].message
if msg.tool_calls:
call = msg.tool_calls[0]
# Parse and validate before executing anything.
return {"tool": call.function.name, "args": json.loads(call.function.arguments)}
return {"tool": None, "answer": msg.content}Code reasoning, tool use and agentic tasks are its strengths at this size
4. Why on-device is a real argument this time
The fair criticism of small models is that they run but fail to be useful. MiniCPM5-2B sidesteps that by aiming at genuinely constrained work: local assistants, coding agents, tool-use workflows, and reasoning that has to happen on the device. Break those down and the requirements are clear. The data cannot leave the machine. The calls are frequent but individually simple. The latency has to be tight. That is exactly where a 2B model belongs and where a cloud flagship is least economical. Artificial Analysis also notes it is token-efficient for a reasoning model, using about 19,000 output tokens per Intelligence Index task, joint-lowest in its comparison group, and output token volume is precisely the cost line that hurts on-device.
# Size the machine before you promise anything. Weights dominate memory at
# batch size 1, and the KV cache grows with context -- the two numbers are
# separate budgets, and the second one bites on 131K contexts.
def memory_budget(params_b=2.52, bits=16, ctx=32_768,
layers=42, kv_heads=2, head_dim=128, batch=1):
bytes_per_param = bits / 8
weights_gb = params_b * 1e9 * bytes_per_param / 1e9
kv_gb = (2 * layers * kv_heads * head_dim * ctx * batch * 2) / 1e9
return {"weights_gb": round(weights_gb, 2),
"kv_cache_gb": round(kv_gb, 2),
"total_gb": round(weights_gb + kv_gb, 2),
"note": "KV scales with context; trim ctx before shrinking the model"}
print(memory_budget(bits=16, ctx=32_768))
print(memory_budget(bits=4, ctx=131_072))5. How to wire it into a local agent
The path is short. First, serve it with vLLM as an OpenAI-compatible endpoint, as in code sample 1, or load it directly with Transformers as in code sample 2; both stay on the machine. Second, make room for tool calling by keeping schemas small and decisions binary, which is the regime this class of model is strongest in, shown in code sample 3. Third, do the memory maths before you promise anything: weights and KV cache are separate budgets, and at a 131K context the KV cache is the number that bites, which code sample 4 estimates. Fourth, route deliberately, sending restricted data and high-frequency trivial calls local while the hard tail goes to the cloud, as in code sample 5. Fifth, do not trust a composite score. Run your own twenty tasks, because a small model's weaknesses only show up on your particular workload.
# Route by privacy and cost, not by habit. Local wins on data that must not
# leave the machine and on high-frequency trivial calls; the cloud keeps the
# hard tail.
LOCAL_FIRST = {"classify", "extract", "format", "rename", "summarise_small"}
def route(task, doc_tokens, sensitivity):
if sensitivity == "restricted":
return "local" # the file never leaves the machine
if task in LOCAL_FIRST and doc_tokens <= 100_000:
return "local" # cheap, fast, good enough
return "cloud" # hard reasoning, huge context, long horizon
print(route("extract", 4_000, "restricted"))
print(route("long_plan", 220_000, "internal"))Artificial Analysis: highest score of any open-weights model under 4B
6. What it means for the open ecosystem
MiniCPM5-2B shipped alongside its training data: the UltraX web pre-training set, tiered UltraData-Code, UltraData-SFT-Agent-2609 with 500,000 agent training samples, and UltraData-RL-2609 with more than 80,000 RL samples across mathematics, code, general knowledge and long-context reasoning. Releasing the data matters more than releasing the weights, because it turns what this size can do into a reproducible engineering question. Combine that with the trillion-parameter open flagships that also shipped this September and the 2026 picture is clear. The question in open weights is no longer whether the models are usable. It is which size, which licence and which cost profile fits your workload, and that choice is moving back to the teams doing the work.
📌 Frequently Asked Questions
What is MiniCPM5-2B?
A 2.5-billion-parameter dense language model released by OpenBMB on September 7, 2026, the second in the MiniCPM5 series after MiniCPM5-1B. It is Apache 2.0 and built for on-device and resource-constrained use.
What are its specs?
2,516,756,480 parameters with 1,981,982,720 outside the embeddings, 42 layers, grouped-query attention with 16 query and 2 key/value heads, a native 131,072-token context window, and a LlamaForCausalLM architecture.
How does it score?
OpenBMB reports a 53.9 average across its own 34-benchmark comparison set, above every larger model included, the highest of which is 51.1. Artificial Analysis independently scores it 15 on the Intelligence Index v4.2, the highest of any open-weights model under 4B total parameters.
What are its weaknesses?
Per Artificial Analysis, it gives ground on knowledge, coding and long-context benchmarks, including 9 percent on Humanity's Last Exam, 9 percent on Terminal-Bench v2.1 and 0 percent on CritPt, while its strengths concentrate in tool use and several agentic tasks.
How do I deploy it?
It supports Transformers, vLLM, SGLang and Docker Model Runner, and the weights are Apache 2.0, permitting commercial use and further training.