11 亿次 AI 动作之后:企业级 Agent 工作流为什么需要一层控制面

·阅读约11分钟·Evergreen Tools Team

2026 年 9 月 24 日,Workato 在自家大会 World of Workato(WOW)2026 上公布:截至 2026 年 9 月,其客户累计处理的企业 AI 动作已超过 11 亿次,平台用户超过 88.6 万人(数据来自其内部平台统计,属厂商口径)。大会主题是「把意图转化为业务结果」。比这个数字更值得工程团队研究的,是它背后的架构选择:在模型与企业系统之间,放一层中立的「控制与执行」层。

一、这个数字到底在说什么

先把口径讲清楚:11 亿次动作与 88.6 万用户都是 Workato 基于内部平台数据发布的厂商口径,并非第三方审计数字,因此在引用时应当标明来源。但它指向的架构趋势是真实可核查的。Workato 把自己定位为「企业 AI 的控制与执行平台」,其逻辑是:企业里已有超过 1.7 万个需要 AI 去操作的企业系统,如果让每一个 Agent 各自去对接、各自持有凭据、各自定义重试与审计策略,那么规模化带来的不是效率,而是失控。联想、VOIS(沃达丰智能解决方案)与红帽在大会上分享了各自的实践,Anthropic、OpenAI 与 AWS 则说明了「模型 + 控制执行平台」的组合方式。

# Define the business intent once, declare what the agent may do, and let
# the platform decide the path. This is the difference between a workflow
# that a business owner can read and a prompt nobody can audit.

intent: resolve-duplicate-invoices
trigger:
  source: erp.invoices            # 17,000-odd systems is the real integration problem
  where: "status = 'pending' and duplicate_of is not null"
agent:
  role: reconcile
  may:
    - read:  [invoices, vendors, purchase_orders]
    - write: [invoices.status, notes]
    - call:  [email.send_template]     # named tools only, never a generic shell
  must_not:
    - payment.execute                  # money movement is out of scope, by design
limits:
  max_actions_per_item: 6
  max_usd_per_item: 0.05
  on_exceed: escalate_to_human
企业 AI 编排

主题是「把意图变成业务结果」

二、为什么需要单独一层「控制面」

把 AI Agent 直接接到系统记录上,是很多团队第一个版本的做法,也是第一个遇到的问题。原因有三个。第一是权限:如果 Agent 直接用某个人的凭据,它继承的是那个人的全部权限,而不是这一次任务需要的那几个动作。第二是重试:企业系统最怕的不是失败,而是「执行了两次」——一笔发票被重复处理、一条记录被重复创建。第三是审计:业务方问「这单为什么被合并」,如果答案要从提示词历史里翻,那这个系统就不可运营。示例 1 把这三个问题一次性写进意图定义:可读什么、可写什么、只允许调用哪些具名工具、以及一条明确的红线(不允许执行付款)。注意「只允许具名工具」这一点——把工具做成具名清单,而不是给 Agent 一个通用 shell,是最容易被忽略但收益最大的设计。

// A control layer earns its name by deciding, per action, before anything
// runs. Route by cost and capability, then enforce the intent's limits.
// The agent never talks to a system of record directly.

export function planAction(action, intent, ctx) {
  const policy = [
    { name: "scope",   ok: intent.may[action.kind]?.includes(action.target) },
    { name: "forbid",  ok: !intent.must_not.includes(action.op) },
    { name: "budget",  ok: ctx.spentUsd + action.estUsd <= intent.limits.max_usd_per_item },
    { name: "calls",   ok: ctx.calls + 1 <= intent.limits.max_actions_per_item },
    { name: "data",    ok: ctx.fields.every((f) => allowedField(f)) },
  ];
  const failed = policy.filter((p) => !p.ok).map((p) => p.name);
  if (failed.length) return { run: false, escalate: true, reasons: failed };

  // Model routing belongs here, not in the prompt. Cheap work first.
  const model = action.complexity === "low" ? "fast-tier" : intent.agent.model;
  return { run: true, model, idempotencyKey: hash(action, ctx) };
}

三、在动作发生之前做决定

控制层的价值体现在时序上:它必须在动作执行之前判定,而不是在执行之后记录。示例 2 给出准入判定的五个检查项——动作是否在声明的范围内、是否命中禁止项、是否会超出单次成本上限、是否会超出动作次数上限、以及所读字段是否被允许。任何一项不通过,就升级给人,而不是让模型自行决定。同一段代码里还有第二个关键决定:模型路由。把「这个任务复杂度低,走便宜档」这样的判断放在控制层,而不是写进提示词,好处是它可被统一调整、可被度量、可被回滚。示例 1 里的 on_exceed: escalate_to_human 与示例 2 的 escalate 是同一条规则的两端:超限就交给人,而不是让 Agent 赌一把。

# Token and action cost are operational metrics, not finance trivia.
# With 1.1 billion actions reported, a one-cent drift per action is a
# seven-figure line. Meter per intent, per action, and per tenant.

def meter(intent, action, usage, tenant):
    cost = usage.input_tokens * IN_RATE + usage.output_tokens * OUT_RATE
    events.emit("ai.action", {
        "tenant": tenant,
        "intent": intent.name,
        "action": action.op,
        "model": action.model,
        "tokens": usage.input_tokens + usage.output_tokens,
        "usd": round(cost, 6),
        "cache_hit": usage.cached,
    })

def nightly(tenant):
    report = rollup(tenant, window="1d")
    # Guardrail: alert before the invoice, not after.
    if report.usd > BUDGET[tenant] * 0.8:
        notify(owner_of(tenant), {"at": "80% of monthly AI budget", **report})
    # Rewrite the expensive path if the cheap one matched quality.
    for step in report.top_cost_steps(5):
        if step.cheap_tier_win_rate > 0.95:
            propose_routing_change(step)
执行层与控制系统

中立层负责的是风险、成本与规模化

四、11 亿次动作意味着成本是工程指标

当动作量级到达十亿次,成本就不再是财务话题。示例 3 的做法是把 token 与动作成本当作运行指标来采集:按租户、按意图、按动作、按模型分别计量,并记录缓存命中。然后设一条预警规则——到达月度预算的 80% 就通知负责人,而不是等到账单出来才复盘。更重要的是最后那段逻辑:如果某个步骤在便宜档位上的质量胜率超过 95%,就主动提出把路由改过去。这与业界关于「模型路由与推理成本架构」的经验一致:省钱的大部分收益来自把请求放到正确的档位,而不是来自讨价还价。示例 5 则在执行侧给出两条底线:每个有副作用的动作必须幂等(避免重试导致重复扣款),高风险动作必须等待人工批准,并且这次批准本身也要被记录在账本里。

// Every AI action that touches a system of record needs an entry that a
// human can read six months later. This is what makes "who changed this
// invoice, and why" answerable without reading a prompt.

type ActionRecord = {
  at: string;
  tenant: string;
  intent: string;            // resolve-duplicate-invoices
  actor: { kind: "agent"; id: string; model: string; version: string };
  approvedBy?: string;       // present only when a human gated this step
  inputs: { ref: string; hash: string }[];   // hashes, not payloads
  decision: string;          // merged | kept | escalated
  effects: { system: string; object: string; before: unknown; after: unknown }[];
  usd: number;
  idempotencyKey: string;    // replay-safe, and that is not optional
};

// A neutral execution layer is only neutral if the record is portable:
// same schema whether the agent was Anthropic's, OpenAI's, or your own.

五、审计记录要能被业务方读懂

示例 4 给出了动作记录的字段设计,其中有两个取舍值得说明。第一,inputs 存的是引用与哈希,而不是原始载荷——这既满足审计需要,也避免把敏感数据二次散落。第二,approvedBy 只在这步确实经过人工闸门时才存在;它的缺席本身就是有意义的信息,说明这一步是自主完成的。字段里还有 actor.version 与 idempotencyKey:前者回答「是哪个版本的 Agent 做的」,后者回答「这件事有没有被重复执行」。Workato 的 CEO Vijay Tella 在大会上把这层定位为「中立层」,用于降低风险、改善 token 成本并规模化 AI;红帽的 Sherry Gentry-Gasper 给出的建议更直接——先把治理与执行的基础建好,再让采用规模长在上面。示例 4 的注释呼应了这一点:只有当记录格式可移植,无论 Agent 来自哪家厂商都使用同一套 schema,这层才真的中立。

#!/usr/bin/env bash
# Two rules that separate a demo from a production control layer:
# 1) every effectful action is idempotent, so retries cannot double-charge
# 2) risky actions pause for a human, and the pause is recorded, not emailed

set -euo pipefail

run_action() {
  key="$1"; op="$2"; payload="$3"
  if ledger.seen "$key"; then
    echo "replay ignored: $key"       # already applied; do not apply twice
    return 0
  fi
  if [ "$(risk "$op")" = "high" ]; then
    decision=$(await_approval --intent "$INTENT" --op "$op" --timeout 30m)
    ledger.record "$key" "$op" "$payload" "approved_by=$decision"
  fi
  effects.apply "$op" "$payload"
  ledger.record "$key" "$op" "$payload" "applied"
}

# Vendor data point: 886,000+ users touch this layer. At that scale,
# "we usually remember" is not a control. Idempotency and approval are.
规模化的企业系统

一层执行面背后是上万个企业系统

六、自建这层控制面的落地清单

第一,把业务意图写成可评审的定义文件,包含允许动作、禁止动作与限额(示例 1)。第二,把准入判定前置到执行之前,并把模型路由也放在这一层(示例 2)。第三,按租户、意图、动作与模型分维度计量成本,并设置 80% 预警线(示例 3)。第四,为每个动作写可读的审计记录,存哈希而非载荷,记录模型版本(示例 4)。第五,强制幂等与人工审批,并把审批结果写入账本(示例 5)。第六,别忘了把这条链路当作产品来迭代——WOW 从 10 月起将扩展到伦敦、新加坡、阿姆斯特丹、东京、特拉维夫与悉尼,企业 AI 的「控制与执行」正在成为跨区域的基础设施话题,而它既是治理问题,也是交付速度问题。

📌 常见问题 FAQ

Workato 公布的 11 亿次企业 AI 动作是什么口径?

该数字来自 Workato 的内部平台数据,截至 2026 年 9 月,属厂商口径,并非第三方审计数据,引用时应注明来源。

「控制与执行层」和普通的工作流自动化有什么区别?

关键在于时序与边界:控制层在动作执行之前做准入判定(范围、成本、次数、字段),并统一负责模型路由、幂等与审计,而不是在事后记录。

为什么不能让 Agent 直接接入系统记录?

三个主要问题:Agent 会继承个人凭据的全部权限;重试可能导致重复执行(例如发票被处理两次);以及缺乏业务方可读的审计记录。

达到十亿级动作量时,成本该怎么管理?

把 token 与动作成本当作运行指标,按租户、意图、动作与模型分维度计量,设置 80% 预算预警,并对在便宜档位胜率超过 95% 的步骤主动提出路由调整。

哪些客户在 WOW 2026 上分享了实践?

联想、VOIS(沃达丰智能解决方案)与红帽分享了各自的落地经验;Anthropic、OpenAI 与 AWS 则介绍了模型与「控制与执行平台」的组合方式。