Abacus.AI 的 Smaug 系列:为长时 Agent 循环调优的开源权重模型
💡 工具推荐:AI Token 计数器、AI 数据分析、AI 代码解释器
2026 年 9 月 10 日,Abacus.AI 发布了 Smaug 系列:三款面向企业 Agentic 场景微调的开源权重模型。Smaug Agentic 微调自 Moonshot AI 的 Kimi K3,Smaug Flash 微调自 DeepSeek V4 Flash 0731,Smaug Mini 微调自 Qwen3.8 27B。官方把 Smaug 定义为一种可套用于任意开源底座模型的微调方法,并称它能把长时 Agent 循环的表现提升 15% 至 20%,且不增加成本。三款模型都在 Hugging Face 上以开源权重发布,可在企业自有 VPC 内托管——这让「开源权重」与「仅 API」之间的差别变得具体。
一、Smaug 是什么:它要修的那面墙
先讲它要解决什么问题。Abacus.AI 说,他们在构建「自我改进的 Agent、自动化复杂工作流程」时反复撞到同一个墙:长时 Agent 循环的成本与低效——上下文很大、工具调用反复发生,而提示缓存会随着时间间隔而失效。Smaug 的方法论由此而来:把人工整理的真实 Agent 轨迹,与基于高难度样例的合成数据结合起来做微调。它对多种开源底座模型都做了实验,结果是「一致的显著提升」,方向锁定在 Agent 真正关心的基准上:Agent 编码、真实世界工具使用、自动化,以及长上下文推理与指令遵循。值得注意的机制细节是:训练时把推理 token 从损失中屏蔽,使训练去引导模型的「动作」,同时保留底座模型原有的推理分布。
# All three Smaug models are open-weight and downloadable, and because the
# fine-tune changed no architectural parameters, the base model's serving
# recipes still apply. For Smaug Agentic the project lists vLLM, SGLang and
# TokenSpeed. Here is the download-then-serve path end to end.
huggingface-cli download abacusai/Smaug-Flash --local-dir ./smaug-flash
huggingface-cli download abacusai/Smaug-Mini --local-dir ./smaug-mini
huggingface-cli download abacusai/Smaug-Agentic-2.8T --local-dir ./smaug-agentic
vllm serve abacusai/Smaug-Agentic-2.8T \
--served-model-name smaug-agentic \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--port 8000
# Weights are open; the training data is not disclosed. Smaug Agentic keeps
# the Kimi K3 licence it inherits from its base model, so read the terms
# before commercial deployment. Plan capacity honestly too: the base is a
# 2.8T-total mixture-of-experts model with 104B activated parameters and a
# 1,048,576-token context window, and the model card reports results from a
# dedicated 8 x B300 deployment, so this is not a laptop model.三款模型,三条能力—效率曲线
二、三款模型与各自的底座
三款模型分别站在能力—效率曲线的不同位置。Smaug Flash(底座 DeepSeek V4 Flash 0731)是这一系列的主力,面向企业自我改进型 Agent,要求速度、成本、效率与 Agent 表现同时成立;官方指出 DeepSeek Flash 在处理长上下文的 Agent 工具调用时容易「打转与混乱」,Smaug Flash 正是针对这一点做改进。Smaug Mini(底座 Qwen3.8 27B)面向多模态用例与较小的推理任务,适合需要多模态能力的一次性任务,企业还可以在自己的数据上继续微调。Smaug Agentic(微调自 Kimi K3)是三者中最大的,底座是一个 2.8 万亿总参数、1040 亿激活参数的混合专家模型,上下文长度 1,048,576 token,并带 MoonViT-V2 视觉编码器。
# A fine-tune that "masks reasoning tokens from the loss" steers actions
# without disturbing the base model's reasoning distribution. When you
# evaluate, measure the things that break in long loops: repetition, stalls
# and tool-call validity -- not just a single-shot score.
def run_agent_task(client, task, max_steps=80, stall_patience=4):
history, seen, stall = [], set(), 0
for _ in range(max_steps):
msg = client.chat(model="smaug-agentic",
messages=history, tools=task.tools)
call = msg.tool_calls[0] if msg.tool_calls else None
key = (call.function.name, call.function.arguments) if call else msg.content
if key in seen:
stall += 1
if stall >= stall_patience: # the failure mode Smaug targets
return {"status": "stalled", "history": history}
else:
stall, _ = 0, seen.add(key)
history.append(msg)
return {"status": "completed", "history": history}三、对 Agent 真正重要的基准
把三类基准摊开看,重点看「相对底座」的增量,因为那才是微调预算真正买到的东西。Smaug Flash 在 LiveBench 上总分 77.4(底座 74.2),其中 Agent 编码 61.1 对 46.8,提升 14.3 分;AutomationBench(公开 600 题、严格通过)38.83 对 25.1;NL2repo-bench 73.3 对 54.2。Smaug Mini 在 LiveBench 总分 76.9(底座 75.3),IFBench 82.0,AutomationBench 41.8 对 37.3,JobBench 官方 65 题协议 50.5 对 33.4。Smaug Agentic 在 GPQA Diamond 得 94.1(底座 93.5),AA-LCR 长上下文推理 75.7(74.7),LiveBench Agent 编码 64.6(62.2),而 Terminal-Bench 2.1 为 86.5,低于底座的 88.3。要注意评测协议:官方称除标注「非自家 harness 复现」的分数外,所有模型在同一 harness 下评测,且 LiveBench 各行取自 2026 年 6 月 25 日发布的榜单。
# Agent cost is not per token, it is per completed task. A cheaper model
# that takes 70 percent more steps can cost more than the expensive one.
# Compare the number that survives contact with a finance team.
def cost_per_completed_task(model, tasks=100):
results = [run(msg.client, t) for t in sample_tasks(tasks)]
done = [r for r in results if r["status"] == "completed"]
spent = sum(model.price(r["tokens_in"], r["tokens_out"]) for r in results)
return {
"cost_per_task": round(spent / len(done), 4),
"completion_rate": round(len(done) / len(results), 2),
"steps_per_task": round(sum(r["steps"] for r in done) / len(done), 1),
}
# Abacus.AI claims open-source model costs are typically 10-100x lower than
# frontier models. Verify that on your tasks; the multiplier is a marketing
# number until your own completion rate is in it.官方称长时 Agent 循环提升 15%–20%
四、论点与它的边界
方法论与它的边界同样重要。Abacus.AI 的核心论点写在 CEO Bindu Reddy 的表述里:开源权重模型正在快速逼近前沿闭源模型,但在长时 Agent 循环中仍表现不足;Smaug 系列要解决的正是这个短板,同时仍比闭源模型便宜 10 到 100 倍。这是一条值得当假设而非结论来读的论断,原因有三:第一,三款模型的基准是厂商自建与引用混合,虽然对「非自家复现」的分数做了标注,但这仍不是第三方独立测评;第二,三款模型的底座各不相同,因此相对底座的增量里,很难完全分离「微调方法」与「底座本身」的贡献;第三,也是最关键的一点:基准提升发生在什么硬件、什么任务分布上,未必等于发生在你的场景里。
# Reading the benchmark tables is easier when you know the protocol traps.
# Every row in the Smaug brief is paired against its own base model under
# the same harness -- except the ones marked as reported, not re-run.
CHECKLIST = [
"Pair against the BASE model, not just a frontier column.",
"Note the harness: same questions for every model, or not?",
"Watch for daggered scores: reported, not reproduced by the vendor.",
"Check the LiveBench release date the rows were drawn from.",
"Read the deployment used: 8 x B300 is not a developer laptop.",
"Re-test on YOUR tasks before trusting a +14.3 point agentic jump.",
]
def score_delta(model, base, benchmark):
# The honest comparison for an agent decision is delta over the base,
# because that delta is what your fine-tuning budget actually bought.
return round(model[benchmark] - base[benchmark], 1)五、能做什么、不能做什么
能做什么、不能做什么,需要分清。能做的:三款都是开源权重、可下载,企业在自有 VPC 内托管即可获得数据与托管位置的控制权,官方也提到对安全与隐私敏感的组织可以在自有 GPU 集群上托管 Smaug Agentic;由于微调未改变架构参数,凡是能跑底座模型的地方就能跑 Smaug,官方给出 vLLM、SGLang 与 TokenSpeed 的部署配方。不能做的:训练数据未披露;Smaug Agentic 继承底座 Kimi K3 的许可,商用前必须阅读其条款;容量规划要诚实——其模型卡报告的评测是在专用 8×B300 部署、温度 1.0、推理强度最大的条件下完成的,这不是笔记本级模型。
# The line is a thesis as much as a product: with the right fine-tuning
# methodology, open-weight models can compete with -- and on agentic tasks
# surpass -- frontier models. The mental model to keep:
thesis = {
"claim": "open weights close the gap, then underperform in long agent loops",
"fix": "fine-tune on human-curated real agent traces + synthetic hard cases",
"mechanism": "mask reasoning tokens so actions are steered, reasoning preserved",
"reported_effect": "15-20% improvement on long-running agentic loops, same cost",
"cost_claim": "10-100x cheaper than frontier closed models",
"buyer_question": "does that hold on MY tasks, on MY hardware, at MY volume?",
}
# Treat "reported_effect" as a hypothesis. The brief itself distinguishes
# scores it re-ran from scores it only cites, which is a good sign and also
# exactly why you need your own eval harness.开源权重,可在企业 VPC 内托管
六、怎么自己评测它
最后是如何评测它。不要拿厂商的排行榜当结论,而要用你自己的任务集跑三件事:第一,比「完成每个任务」的成本,而不是每 token 价格——一个更便宜但多花 70% 步数的模型,可能比贵的那个更贵(示例 3);第二,测长时循环的稳定性,重点看重复、卡死与工具调用合法性,而不是单次得分(示例 2);第三,明确你的推理长度与上下文规模,因为 Smaug Agentic 的模型卡显示,在 SciCode 与 AA-LCR 上 p99 推理长度降到基座的约 0.6 倍,而可见答案长度与基座无统计差异——这类细节只有在你自己的评测里才会显形。示例 4 给出一份看榜避坑清单,示例 5 把整个论证整理成一个可检验的心智模型。
📌 常见问题 FAQ
Smaug 是什么?
Abacus.AI 于 2026 年 9 月 10 日发布的系列:三款面向企业 Agentic 用例微调的开源权重模型,以及一套可套用于任意开源底座模型的微调方法。
三款模型分别基于什么?
Smaug Flash 基于 DeepSeek V4 Flash 0731,Smaug Mini 基于 Qwen3.8 27B,Smaug Agentic 微调自 Moonshot AI 的 Kimi K3(2.8 万亿总参数、1040 亿激活、1,048,576 token 上下文)。
官方声称的效果是什么?
该公司称 Smaug 把长时 Agent 循环的表现提升 15% 至 20% 且不增加成本,并称开源模型成本通常比前沿闭源模型低 10 至 100 倍。
权重可以商用吗?
权重以开源形式在 Hugging Face 发布,可下载并可自带 VPC 托管;但 Smaug Agentic 继承底座 Kimi K3 的许可,商用前需遵守该许可条款,且训练数据未披露。
要注意哪些评测陷阱?
基准多为厂商自建或引用、不同模型底座不同、部分分数标注为非自家 harness 复现,且 Smaug Agentic 的评测在 8×B300 专用部署上完成;应改用自己任务集、按「完成每任务成本」与长循环稳定性来评测。
🔧 推荐工具
📚 参考资料
- Abacus.AI — Introducing the Smaug Line: Open-Weight Models for Agentic AI (technical brief and benchmarks)
- Abacus.AI — Smaug launch press release (Sep 10, 2026): three models, 15-20% lift, 10-100x cost claim
- Hugging Face — abacusai models (Smaug-Agentic-2.8T, Smaug-Flash, Smaug-Mini model cards)
- LiveBench — leaderboard including the finetunes filter used for the Smaug rows