可观测性正围绕 Agent 收拢:CloudWatch Omni、OpenObserve v1.0 与 Dynatrace 的 Arize 押注

·阅读约11分钟·Evergreen Tools Team

过去这一周,AI 可观测性市场同时收到三个信号。9 月 23 日,AWS 宣布 Amazon CloudWatch Omni 正式 GA——一个以应用为中心、由 AI 驱动、基于 OpenTelemetry 的可观测性体验,还专门为 Agent 提供了评估驱动的开发流程。9 月 22 日,开源方案 OpenObserve 发布 v1.0,把 Agent 追踪、LLM 监控、评估和会话标注塞进了同一个已经在处理日志、指标、追踪和真实用户监控的平台。同期,Dynatrace 出手收购 Arize。方向很清楚:把 Agent 观测收进统一平台。但整合掩盖了一个更难看的数字。

一周之内,三家公司做了同一个动作

先摆事实。AWS 的 CloudWatch Omni 在 2026 年 9 月 23 日进入 GA,定位是「应用为中心、AI 驱动、基于开放标准、离开控制台交付」:你可以用组织专属 URL 和已有身份 SSO 登录,不必进 AWS 管理控制台。它的 Agent 可观测性覆盖 LangGraph、CrewAI、OpenAI Agents SDK、Vercel AI SDK 和 Strands,对每一次 prompt、模型调用和工具调用做质量评估与实验。OpenObserve v1.0 在 9 月 22 日 GA,主打把 AI 可观测性并入同一个日志/指标/追踪/RUM 平台,支持自托管和云。Dynatrace 则通过收购 Arize 补强 Agent 观测能力。

// Instrument once, in OpenTelemetry, and let the backend be a decision you can
// revisit. The model provider is not the vendor you are locking into here; the
// telemetry format is. Keep it OpenTelemetry and you keep optionality.

const { NodeSDK } = require('@opentelemetry/sdk-node');
const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-http');

const sdk = new NodeSDK({
  traceExporter: new OTLPTraceExporter({
    url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT,
  }),
});

sdk.start();
// Now spans look the same whether they land in CloudWatch Omni, OpenObserve,
// or anything else that speaks OTLP. That is the whole point.
Agent 可观测性整合

一周内三家厂商,同一个方向:把 Agent 观测收进统一平台

真正该看的那个数字

整合很热闹,但缺口更值得抄下来。LangChain 2026 年 6 月对 1300 多名从业者的调查显示:89% 的团队已经实现了某种形式的 Agent 可观测性,但只有 37.3% 跑在线评估。换句话说,绝大多数人「看得见」自己的 Agent,却「打不了分」。Salesforce 2026 年 2 月的报告则显示,企业平均已经运行 12 个 Agent,其中一半处于各自为政的孤岛,只有 54% 的组织对它们有集中治理。追踪是入场券,评估才是分水岭。

// A model call with no span attributes is a log line. Give every agent step the
// same keys so cross-vendor tracing actually joins up.

const { trace } = require('@opentelemetry/api');

async function callModel(prompt, meta) {
  const span = trace.getTracer('agent').startSpan('llm.call');
  span.setAttribute('gen_ai.system', meta.provider);
  span.setAttribute('gen_ai.request.model', meta.model);
  span.setAttribute('agent.step', meta.step);        // plan | act | verify
  span.setAttribute('agent.run_id', meta.runId);
  try {
    const out = await provider.generate(prompt);
    span.setAttribute('gen_ai.usage.input_tokens', out.usage.input);
    span.setAttribute('gen_ai.usage.output_tokens', out.usage.output);
    return out;
  } finally {
    span.end();
  }
}

别把逃生舱口焊死

分析师对 CloudWatch Omni 的评价也很直白:它能减少工具碎片化、加快对 Agent 行为的排查,但「锁定」和「不断上涨的遥测成本」可能限制它的吸引力。对策不是拒绝这些平台,而是把可移植性写进架构里。只要你的遥测是 OpenTelemetry 格式、通过 OTLP 导出,后端就只是一个可以重新审视的决策,而不是一次不可逆的婚姻。真正让你锁死的从来不是模型供应商,而是遥测格式。

# The survey number that should worry you: 89% of teams have some form of agent
# observability, only 37.3% run online evaluations. Tracing is not evaluation.
# Wire an eval into the same pipeline that emits the traces, or you are watching
# a system you cannot grade.

def online_eval(trace_record, rubric):
    score = rubric.grade(trace_record.output)
    emit_metric("agent.eval.score", score, {
        "run_id": trace_record.run_id,
        "step": trace_record.step,
        "model": trace_record.model,
    })
    # Alert on the distribution, not one bad answer.
    return score < rubric.floor
来自 Agent 的遥测数据

OpenTelemetry 是你的逃生舱口,别放弃它

把评估接进发追踪的同一条管道

光有追踪不叫评估。要让评估有意义,它必须和追踪共用同一条管道:每一次模型调用结束时,就地用评分标准打分,把分数作为指标发射出去,并且对「分布」告警,而不是对某一个坏答案告警。这样你才能回答那个真正的问题——这个 Agent 变好了还是变差了,而不是它刚才输出了什么。

// Cost is the second half of observability. Token spend per agent step is easy
// to compute and impossible to argue with. Alert on the slope, not the total.

type StepUsage = { runId: string; step: string; input: number; output: number };

const PRICES = { in: 1.40 / 1e6, out: 4.40 / 1e6 }; // example per-token rates

function stepCost(u: StepUsage) {
  return u.input * PRICES.in + u.output * PRICES.out;
}

function runCost(usages: StepUsage[]) {
  return usages.reduce((sum, u) => sum + stepCost(u), 0);
}

// Chart cost per completed agent run. If it trends up while task success is flat,
// you bought tokens, not outcomes.

成本是可观测性的另一半

可观测性的第二半是钱。按 Agent 步骤计算 token 花费很容易,而且无可辩驳。要看的是坡度而不是总量:把「每个已完成 Agent 运行的成本」画出来,如果它持续上升而任务成功率持平,那你买的是 token,不是结果。这一步不需要新工具,只需要把每步的输入/输出 token 和公开单价乘起来,然后对着趋势告警。

# Before you standardize on any single vendor, do the boring import. Export a
# week of spans and read them yourself. A vendor migration is cheap if your
# telemetry is portable and expensive if it is not.

otel-cli export --endpoint "$OTEL_EXPORTER_OTLP_ENDPOINT" \
  --service agent-gateway --since 7d > spans.jsonl

wc -l spans.jsonl
jq -r '.attributes["gen_ai.system"]' spans.jsonl | sort | uniq -c | sort -rn

# If more than one provider shows up and your spend column cannot explain it,
# you have an observability problem before you have a vendor problem.
评估驱动的开发流程

埋点只是开始,在线评估才是分水岭

迁移很便宜,前提是你早就准备好了

给一个务实的收尾:在你把标准钉死在任何一个厂商之前,先做那件无聊的导入——导出过去一周的 span,自己读一遍。哪些模型供应商出现过、分别调用了多少次、你的花费栏能不能解释它。如果答案是「解释不了」,那你面对的是可观测性问题,而不是供应商问题。可移植的遥测让迁移变得便宜,不可移植的遥测让迁移变得昂贵,而这个选择,今天就在你手里。

📌 常见问题 FAQ

CloudWatch Omni 是什么时候发布的?

AWS 于 2026 年 9 月 23 日宣布 Amazon CloudWatch Omni 正式 GA(通用可用)。它是以应用为中心、由 AI 驱动、基于 OpenTelemetry 的可观测性体验,通过组织专属 URL 与 SSO 使用,无需进入 AWS 管理控制台,并包含专用的 Agent 可观测性与评估驱动开发流程。

OpenObserve v1.0 有什么不同?

2026 年 9 月 22 日发布的 OpenObserve v1.0 主打 AI Observability,把 Agent 追踪、LLM 监控、评估和会话标注并入同一个已在处理日志、指标、追踪和真实用户监控的平台,支持自托管部署与 OpenObserve Cloud。

为什么说「整合」背后有缺口?

因为可见性不等于质量。LangChain 2026 年 6 月对 1300 多位从业者的调查显示,89% 的团队实现了某种 Agent 可观测性,但只有 37.3% 跑在线评估。你可以追踪每一次调用,却仍然无法给 Agent 的表现打分。

分析师对 CloudWatch Omni 的主要担忧是什么?

两点:厂商锁定和不断上涨的遥测成本。应对办法是保持遥测可移植——使用 OpenTelemetry 并通过 OTLP 导出,这样后端只是可重新评估的决策,而不是不可逆的绑定。

我应该先做追踪还是先做评估?

两者共用同一条管道才有意义。先确保每一次模型调用都有结构化 span 属性(供应商、模型、Agent 步骤、运行 ID),然后在这些调用结束时就地打分并作为指标发射,对分数的分布而不是单个答案告警。