信心缺口:Harness《2026 Agent 开发生命周期现状》报告揭示了什么
💡 工具推荐:AI 代码审查、AI Token 计数器、API Mock 生成器
2026 年 9 月 10 日,Harness 发布了《2026 Agent 开发生命周期现状》(The State of Agent DLC 2026)。这份由 Sapio Research 于 2026 年 7 月执行的调研覆盖 700 名技术从业者,受访组织均在千人以上、拥有 100 名以上开发者,并且已在生产、试点或验证环境中部署 AI Agent。报告最重要的发现不是一个百分比,而是一个在所有领域都重复出现的模式:信心落在 75% 附近,而能支撑这份信心的控制手段,落地的组织不到一半,有些甚至不到五分之一。
一、五个领域里的同一个缺口
报告横跨测试、安全、资产清点、成本与回滚,每个领域的形状都一样。77% 的组织相信自己对环境中运行的每个 Agent、MCP 服务器和模型都有完整清点,但只有 44% 真正运行了主动发现工具去验证;74% 相信自己的测试能拦下影响生产的故障,但只有 19% 设有能自动拦截每一次坏发布的闸门;76% 认为能在 15 分钟内关停一个行为异常的 Agent,但只有 33% 真的有即时开关;74% 说能看清每个 Agent 的真实开销,然而 60% 在上一季度仍然超支。最有说服力的一项是安全:75% 声称自家 Agent 端到端安全,而这个群体过去一年发生安全事件的比例是 88%,与全体受访者的 87% 几乎没有差别。
# Gate 1: inventory. You cannot govern agents you cannot enumerate.
# 77% of organisations were confident they had a complete inventory;
# only 44% ran discovery tooling to prove it. Make the proof automatic.
async def discover(org_id: str) -> list[dict]:
found = []
async for ep in control_plane.list_endpoints(org_id): # k8s, VM, laptop
async for proc in ep.processes():
if proc.image in AGENT_IMAGES or proc.cmdline_has("--acp"):
found.append({"host": ep.id, "agent": proc.image,
"version": proc.version, "owner": ep.labels.get("team")})
async for srv in mcp_registry.list_servers(org_id): # MCP servers
found.append({"host": srv.endpoint, "agent": "mcp:" + srv.name,
"version": srv.version, "owner": srv.owner})
return found
def reconcile(found, declared):
# Anything running that nobody declared is a finding, not a surprise.
return {"undeclared": [a for a in found if a["agent"] not in declared]}信心高于控制的模式在五个领域重复出现
二、为什么确定性软件的控制手段会失效
Harness 的 Field CTO 兼研究负责人 Keith Mann 给了一个很准确的类比:云计算和移动互联网都经历过「信心跑在治理前面」的阶段,但最终控制手段追了上来,因为底层系统一旦建好护栏就是可预测的。Agent 不一样——它不会以同样的方式静止。同一份输入跑两次可能得到不同输出,因此为确定性软件设计的控制手段会失效:一次在测试环境里表现完美的检查,到了生产依然可能漏掉东西,因为 Agent 的行为会变。示例 2 把这一点写成工程要求:闸门判定必须针对「变更是否通过固定标准」,而不是针对「我们有没有跑过测试」。
// Gate 2: an eval gate, not an eval report. 74% were confident their
// testing would catch a production-impacting failure; only 19% had a gate
// that automatically blocks every bad release. The difference is "must pass".
export const releaseGate = {
id: "agent-release-gate",
stages: [
{ name: "evals", mustPass: true, minScore: 0.90, dataset: "golden-200" },
{ name: "safety", mustPass: true, maxViolations: 0, suite: "jailbreak+exfil" },
{ name: "cost", mustPass: true, maxUsdPerRun: 0.12 },
{ name: "latency", mustPass: true, p95Ms: 4000 },
],
onFail: "block", // no manual override without a signed waiver
};
// Every promoted change is checked against the same fixed standard.
export async function promote(change, gate = releaseGate) {
const results = await runStages(change, gate.stages);
if (results.some((r) => r.failed && r.mustPass)) {
throw new BlockedRelease(change.id, results); // 81% of orgs do not have this
}
return change;
}三、第一道闸门:可被证明的资产清点
治理的前提是枚举。报告显示信心与证明之间的落差高达 33 个百分点,说明大多数组织的清点来自「应该都在这里」的假设,而不是来自持续运行的发现工具。示例 1 给出一个务实的做法:同时扫描端点上的进程与 MCP 服务器注册表,把运行中的 Agent 汇总成清单,再与声明清单做对账——凡是「在跑但没人声明」的,都视为发现项而不是意外。这一步的工程价值在于把治理从问卷变成数据:你不需要相信清单是完整的,你需要的是清单会自己报告不完整。
# Gate 3: a real kill switch. 76% believed they could disable a
# misbehaving agent in under 15 minutes; only 33% had an instant switch.
# "Instant" means a control the agent cannot argue with at runtime.
import redis, time
r = redis.Redis()
def revoke(agent_id: str, reason: str, by: str) -> None:
# Step 1: flip the flag the runtime checks on every single action.
r.set(f"agent:{agent_id}:revoked", "1", ex=86400)
# Step 2: record who did it, when, and why. Auditable, not improvised.
r.xadd("agent-revocations", {
"agent": agent_id, "reason": reason, "by": by,
"ts": str(time.time()),
})
def may_act(agent_id: str) -> bool:
# Checked before every tool call, not once at session start.
if r.get(f"agent:{agent_id}:revoked"):
raise PermissionError("revoked: %s" % agent_id)
return True75% 说 Agent 安全,但该群体的事故率与整体几乎相同
四、第二、三道闸门:必须通过的评估与真的能用的开关
评估有两种截然不同的形态:报告和闸门。报告告诉你分数,闸门决定能不能上线。示例 2 展示了一个「必须通过」的发布闸门,评估、安全、成本、延迟四关全部为硬性条件,失败即阻断,除非有人签署豁免;这正好对应报告里只有 19% 的组织具备的能力。示例 3 处理的是回滚:76% 的信心对上 33% 的能力,差的就是一个 Agent 无法讨价还价的运行时开关。注意示例里的两个细节——开关在每个工具调用前都被检查,而不是在会话开始时检查一次;关停动作本身也被写进审计流,记录是谁、何时、为何关停。25 分钟的差距,通常就藏在这两处。
# Gate 4: spend per agent, with a ceiling. 74% said they had a complete
# picture of true spend; 60% still overran budget last quarter. Visibility
# without a limit is a dashboard, not a control.
BUDGET = {"research-agent": 900.00, "pr-review-agent": 250.00}
def admit(agent_id: str, est_usd: float, month_spend: dict) -> dict:
cap = BUDGET.get(agent_id)
if cap is None:
return {"allow": False, "why": "no budget line: " + agent_id}
spent = month_spend.get(agent_id, 0.0)
if spent + est_usd > cap:
return {"allow": False, "why": "over cap", "spent": spent, "cap": cap}
return {"allow": True, "remaining": round(cap - spent - est_usd, 2)}
# Route the request down a cheaper tier before refusing outright.
def degrade(est_usd: float) -> str:
return "cheap-tier" if est_usd > 0.05 else "default-tier" 五、第四道闸门:把开销变成上限
成本是报告里最容易被忽视的一条。74% 的组织说能看清每个 Agent 的真实开销,但 60% 仍然超支,这说明「可视」不等于「可控」。示例 4 把预算写成准入判定:没有预算行的 Agent 不予放行,估算开销加已用额度超过上限则拒绝,并在拒绝之前先尝试降级到更便宜的模型档位。这个顺序很重要——先降级、再拒绝,能让治理在绝大多数情况下不表现为「业务被拦住」,而表现为「成本被优化」。既然 58% 的组织报告每 100 次变更的生产事故在增加,把开销和事故率放在同一张看板上,是最省力的起点。
// Agent changes are not code changes. 42% route prompt edits through the
// same pipeline as code, and just 34% have a dedicated configuration
// system for AI behaviour. A schema gives you a reviewable diff instead
// of a one-line prompt tweak that silently changes tool permissions.
{
"id": "pr-review-agent",
"model": { "primary": "gpt-5.6-sol", "fallback": "claude-fable-5.1" },
"prompt": { "ref": "prompts/pr-review.md", "rev": "sha256:41ab..." },
"tools": {
"allow": ["readRepo", "runTests", "commentOnPr"],
"deny": ["pushToMain", "writeSecrets", "deleteBranch"]
},
"budget": { "usdPerMonth": 250, "usdPerRun": 0.12 },
"rollout": { "strategy": "canary", "percent": 10, "watch": "errorRate<1.5%" },
"evaluations": { "gate": "agent-release-gate", "minScore": 0.9 }
}
// Now "we changed the agent" comes with a diff, an owner, and a reason.58% 的组织报告每 100 次变更的生产事故在增加
六、Agent 变更不是代码变更
报告里还有一组关于流程的数据:42% 的组织把提示词修改走与代码相同的流水线,只有 34% 为 AI 行为建有专门的配置系统;只有 53% 的 Agent 相关变更在进入生产前会经过任何标准流水线,37% 的组织连一半都不到;四成以上组织逐案判断一次变更是否可信,而在把 Agent 变更推上生产的组织里,只有 58% 会让每一次变更都对照固定且可重复的标准。Harness 的建议是把 Agent 生命周期当成独立学科:为 Agent 行为单独构建评估、安全、清点与回滚;用固定标准替代临时评审;并对 Agent 采用渐进式发布,而不只是人工闸门。示例 5 给出了一份 Agent 定义 schema,让「我们改了 Agent」这件事产出一个带 owner、带 diff、带原因的变更记录,而不是一行悄悄改变了工具权限的提示词微调。
📌 常见问题 FAQ
Harness 的这份报告调研了谁?
报告基于 Sapio Research 于 2026 年 7 月执行的调研,覆盖美国、英国、法国、德国和印度的 700 名技术从业者,受访组织均超过 1000 名员工、100 名以上开发者、年收入超过 1 亿美元,且已部署 AI Agent。
报告最核心的发现是什么?
信心与控制的系统性落差:在测试、安全、清点、成本、回滚五个领域,对 Agent 的信心都在 75% 左右,而支撑信心的控制手段落地的组织不到一半,有些不到五分之一。
为什么 Agent 需要不同于传统软件的控制?
Agent 不具备确定性:同一输入在不同运行中可能产生不同输出,因此为确定性软件设计的测试与闸门会失效,需要针对 Agent 行为的变化性单独构建评估与控制。
报告建议组织怎么做?
把 Agent 生命周期当作独立学科,为 Agent 行为单独构建评估、安全、清点与回滚;用固定且可重复的标准替代逐案评审;并对 Agent 变更采用金丝雀、蓝绿等渐进式发布方式。
哪些数据最值得管理层关注?
三组:75% 声称 Agent 安全而该群体事故率与整体几乎相同(88% vs 87%);76% 相信能快速关停而只有 33% 有即时开关;58% 的组织报告每 100 次变更的生产事故在增加。