AgenticOps 已进入网络运维:51% 的组织让 Agent 在生产环境直接动手
💡 工具推荐:API 响应时间计算器、API 限流计算器、HTTP 状态码参考
2026 年 9 月 23 日,Cisco 发布了由 Omdia 独立执行的调研报告《Agentic AI 对网络运维的影响》。这份调研访问了 1000 名 IT 与网络运维负责人,受访组织均在 500 人以上,覆盖北美、西欧与亚太。结论很直接:组织已经从「AI 只做建议」的阶段跨过了「AgenticOps」的门槛,也就是由运维人员定方向与护栏、由 AI Agent 跨域感知、推理与执行。超过五分之四的受访者预期在 12 个月内进入 AI 主导的运营模式,四分之三愿意在 NetOps 中授予 Agentic AI 显著自主权,其中近四分之一甚至接受完全没有人工介入的全自主运行。
一、告警的算术已经算不通
报告用一组数字解释了为什么自主化不再是噱头。平均每个组织每天产生约 4100 条监控告警与事件,其中一半以上与网络相关;按这个体量,靠人工清理每日网络告警积压大约需要 100 名 IT 专家。更糟的是效率:近一半的网络告警未经调查就被关闭,而几乎同样比例的受访者表示调查时间被花在了误报上。与此同时复杂度还在上升——92% 的组织表示性能问题通常跨多个域,需要跨 10 个以上工具做关联才能解决;57% 表示现有变更流程跟不上今天要求的速度。这组数字放在一起,得到的不是「AI 很酷」的结论,而是「旧的告警模型已经失效」。
# Intent, not scripts. AgenticOps means operators declare the outcome and
# the guardrail, and agents choose the path inside it. Everything that
# follows is a constraint the runtime enforces, not advice for the agent.
intent:
name: keep-payments-latency
target: "p95_latency_ms <= 40 for service=payments"
scope: [region:ap-northeast, tier:edge]
allowed_actions:
- shift_traffic_weight # within declared bounds only
- scale_edge_pool
- rollback_last_change
forbidden:
- modify_firewall_policy
- change_dns_ttl # blast radius too wide for autonomy
approval_required:
- anything_touching: [region:us-east, tier:core]
observe_only_hours: 72 # shadow mode before the agent may act平均每个组织每天产生约 4100 条监控告警与事件
二、Agentic 已经被部署,不只是被讨论
报告最值得注意的地方在于,自主性不是未来时态。75% 的组织已经把 AI 用于 NetOps,51% 正在运行会在生产中动手的 Agentic AI。它们做的事包括重新路由流量、调整无线参数、隔离可疑终端,以及在没有事先人工批准的情况下端到端解决事件。意愿也跟上了:80% 的组织愿意授予 AI 高度甚至完全自主的角色,其中 24% 接受完全没有人工监督;82% 愿意让 AI 在没有事先批准的情况下做出至少部分生产网络变更;84% 预期在 12 个月内进入 AI 主导的运营模式。如果你的组织还在讨论「要不要让 Agent 动手」,报告给出的答案是你已经落后于多数同行。
// Blast radius is the whole argument. 82% are comfortable letting AI make
// some production changes without prior approval, which is only reasonable
// if the change cannot exceed a bound you chose in advance.
export function admit(change, intent) {
const checks = [
{ name: "scope", ok: within(change.targets, intent.scope) },
{ name: "action", ok: intent.allowed_actions.includes(change.kind) },
{ name: "forbidden", ok: !intent.forbidden.some((f) => matches(change, f)) },
{ name: "budget", ok: change.affectedPct <= 5 }, // <= 5% of traffic
{ name: "window", ok: !inFreeze(change.now) },
];
const failed = checks.filter((c) => !c.ok).map((c) => c.name);
if (failed.length) {
// Pause for a human rather than guessing. 24% would skip this step;
// for anything in scope [tier:core] that is a bad trade.
return { decision: "escalate", reasons: failed, reviewer: intent.owner };
}
return { decision: "allow", rollout: { canaryPct: 5, watch: "p95_latency_ms" } };
}三、把意图写成可执行的护栏
自主不等于放任,报告里 Cisco 的表述是「从运行运维转向编排意图」。示例 1 把这句话写成一份可执行的文件:目标写清楚(支付服务 p95 延迟不超过 40 毫秒),范围写清楚(东北亚区域的边缘层),允许的动作写清楚(在声明边界内调整流量权重、扩缩边缘池、回滚最近一次变更),禁止的动作也写清楚(修改防火墙策略、改 DNS TTL——因为影响面太宽)。还有一个容易被忽略但很重要的字段:observe_only_hours,也就是允许 Agent 动手之前的影子运行期。把意图写成文件的好处是,它同时是给 Agent 的指令、给人类的评审材料,以及运行时判定的依据。
# The alert math is what makes autonomy attractive: ~4,100 events a day,
# more than half of them network related, and roughly 100 specialists
# needed to clear the backlog by hand. Correlation is not optional.
def dedupe(events: list[dict]) -> list[dict]:
# 92% of organisations say performance issues span multiple domains and
# need 10+ tools to resolve. Collapse them into one incident instead.
groups: dict[str, dict] = {}
for e in events:
key = (e["service"], e["region"], e["failure_mode"]) # not the timestamp
g = groups.setdefault(key, {"key": key, "count": 0, "domains": set(),
"first": e["ts"], "last": e["ts"]})
g["count"] += 1
g["domains"].add(e["domain"])
g["last"] = e["ts"]
return sorted(groups.values(), key=lambda g: (-len(g["domains"]), -g["count"]))
# Roughly half of network alerts are closed without investigation today.
# Ranking by domain spread tells you which half was actually safe to skip.80% 愿意授予高度自主,但 69% 要求详细的可解释性
四、影响面就是那场争论的全部
82% 的组织愿意让 AI 在无事先批准下做部分生产变更——这个比例只有在「变更不可能超出你事先选定的边界」时才合理。示例 2 给出准入判定的五个检查项:目标是否在授权范围内、动作类型是否在允许清单、是否命中禁止项、受影响流量是否不超过 5%、当前是否处于冻结窗口。任何一项不过,就升级给人,而不是让 Agent 猜。示例 5 处理的是另一半:渐进式发布。先让 Agent 在 5% 的切片上动手,用 15 分钟观察意图里承诺的那个指标,如果退化就自动回滚。这两个示例合起来回答的正是报告里「信任」的定义:可见性、可解释的上下文,以及能保证确定性结果的护栏。
// The report is unambiguous about the price of autonomy: full observability,
// with tracing, a summarised rationale, and post-action audits, is the
// minimum acceptable standard for 36% of organisations, and 69% require
// detailed explainability for agent-driven actions. So record all of it.
type AgentAction = {
traceId: string;
agent: string;
intent: string; // which declared objective authorised this
tool: string; // e.g. shift_traffic_weight
args: Record<string, unknown>;
rationale: string[]; // the agent's reasoning, in its own words
evidence: string[]; // telemetry ids it read before deciding
predicted: Record<string, number>; // expected p95, error rate, cost
observed?: Record<string, number>; // filled in after the window closes
rollback?: { at: string; by: string; verified: boolean };
};
// "It shows its work, so I can follow the logic" is the difference between
// a force multiplier and a queue of tickets nobody trusts.五、网络本身也在被 Agent 改变
报告里一个容易被工程团队忽略、但运维团队必须正视的发现是流量。根据 Cisco 对直连 AI 网络遥测的聚合分析,AI 流量正以每六个月翻倍的速度增长;而当把 Agentic AI 产生的流量计算在内时,Cisco 的测试发现由 Agent 执行的任务可能产生最多 450% 的额外总流量。也就是说,Agent 在优化服务指标的同时,可能正在给网络本身制造新的压力。示例 5 因此在回滚判据里同时盯着服务与网络两侧。另外,示例 3 处理的是告警本身:把同类事件按服务、区域、失败模式聚合成一个事件而不是成千上万条,并优先处理跨域数量多的那批——因为近一半网络告警今天是被不经调查直接关掉的,排序能力直接决定了哪些被关掉是安全的。
#!/usr/bin/env bash
# Progressive rollout beats a big-bang switch. Let the agent act on a small
# slice, watch the same metric the intent promised, and lose the argument
# automatically if the slice degrades.
set -euo pipefail
intent="$1"; slice="5"; window="15m"
agent-apply --intent "$intent" --slice "$slice"
if ! agent-watch --intent "$intent" --window "$window" --max-regression 2pct; then
echo "regression detected in slice; rolling back"
agent-apply --intent "$intent" --rollback
fi
# Cisco testing found agent-performed tasks can generate up to 450% more
# total network traffic, and AI traffic is on a trajectory to double every
# six months. Watch the network too, not just the service. Cisco 测试发现 Agent 执行任务可产生最多 450% 的额外流量
六、可解释性不是附加项,是入场费
报告对自主权的定价说得很清楚:69% 的组织要求 Agent 驱动的动作具备详细的可解释性;36% 认为「完整可观测性」,即详细追踪、归纳后的决策理由与事后审计,是最低可接受标准;86% 认为最优路径是一个集成的单一平台,而不是再买一个单点工具。示例 4 给出了应当记录的内容:追踪 ID、是哪一个意图授权了这次动作、调用了哪个工具、参数是什么、Agent 自己的推理、它决策前读了哪些遥测、预期的指标值,以及窗口结束后实际观测到的值。Cisco 的 Joe Vaccaro 把这件事概括为「建立在可见性之上的信任:对每一个决策的可见性、对每条建议的可解释上下文、以及能确保确定性结果的护栏」。Room & Board 的高级网络工程师 Mark Rodrigue 说得更直白:能展示推理过程,他才能跟上逻辑、看到证据链,那才是把 Agent 部署到规模上的信任。
📌 常见问题 FAQ
Cisco 这份报告是谁做的?
报告题为《Agentic AI 对网络运维的影响》,由 Omdia 独立执行,调研了 1000 名 IT 与网络运维负责人,受访组织均在 500 人以上,覆盖北美、西欧与亚太。
现在有多少组织已经让 Agent 在生产动手?
51% 的组织正在运行会在生产中采取动作的 Agentic AI,75% 已将 AI 用于网络运维,84% 预期在 12 个月内进入 AI 主导的运营模式。
为什么组织愿意授予 Agent 自主权?
告警体量已经超出人力:平均每个组织每天约 4100 条告警与事件,其中一半以上与网络相关,人工清理积压约需 100 名 IT 专家,而近一半网络告警目前未经调查就被关闭。
报告对 Agent 产生的网络流量有什么发现?
据 Cisco 对直连 AI 网络遥测的聚合分析,AI 流量正以每六个月翻倍的速度增长;Cisco 测试发现由 Agent 执行的任务可能产生最多 450% 的额外总流量。
报告认为自主化的前提条件是什么?
可解释性与可观测性:69% 的组织要求详细可解释性,36% 认为含详细追踪、归纳理由与事后审计的完整可观测性是最低标准,86% 倾向于单一集成平台。