Observability Just Consolidated Around Agents: CloudWatch Omni, OpenObserve v1.0, and Dynatrace's Arize Bet
💡 Tool Tip:AI Token Counter, API Response Time Calculator, JSON Formatter
In a single week the AI observability market sent three signals at once. On September 23 AWS took Amazon CloudWatch Omni to general availability — an application-centric, AI-powered, OpenTelemetry-based observability experience with a purpose-built, evaluation-driven workflow for agents. On September 22 the open-source option OpenObserve shipped v1.0, folding agent tracing, LLM monitoring, evaluation, and session annotation into the same platform that already handles logs, metrics, traces, and real user monitoring. Around the same time Dynatrace moved on Arize. The direction is obvious: pull agent observability into one platform. The consolidation is hiding a harder number.
Three Companies, One Move, One Week
Here are the facts. CloudWatch Omni reached GA on September 23, 2026, positioned as app-centric, AI-powered, built on open standards, and delivered off-console: you sign in through an organization-specific URL with SSO for your team instead of the AWS Management Console. Its agent observability covers LangGraph, CrewAI, OpenAI Agents SDK, Vercel AI SDK, and Strands, and it evaluates quality and runs experiments on every prompt, model call, and tool invocation. OpenObserve v1.0 went GA on September 22, bringing AI observability into the same platform as logs, metrics, traces, and RUM for both self-hosted and cloud deployments. Dynatrace, meanwhile, took the acquisition route with Arize. Read the three announcements together and the pattern is hard to miss. Each vendor is arguing the same thing: your agent's traces, its cost, and its quality belong in one place, next to the application telemetry you already collect. What differs is how much of that story you can take with you if you leave.
// Instrument once, in OpenTelemetry, and let the backend be a decision you can
// revisit. The model provider is not the vendor you are locking into here; the
// telemetry format is. Keep it OpenTelemetry and you keep optionality.
const { NodeSDK } = require('@opentelemetry/sdk-node');
const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-http');
const sdk = new NodeSDK({
traceExporter: new OTLPTraceExporter({
url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT,
}),
});
sdk.start();
// Now spans look the same whether they land in CloudWatch Omni, OpenObserve,
// or anything else that speaks OTLP. That is the whole point.Three vendors in one week, one direction: pull agent observability into one platform
The Number That Actually Matters
The consolidation is loud, but the gap is what you should write down. LangChain's June 2026 survey of more than 1,300 professionals found that 89% had implemented some form of agent observability, yet only 37.3% ran online evaluations against it. In other words, most teams can see their agents but cannot grade them. Salesforce's February 2026 Connectivity Benchmark Report found enterprises run 12 agents on average, half of them operating in silos, with only 54% of organizations holding centralized governance. Tracing is table stakes; evaluation is the divide. The uncomfortable reading is that most teams bought visibility and skipped grading. Tracing answers what happened; evaluation answers whether it was good. A dashboard full of spans that no rubric ever scored will tell you an agent ran, not whether it helped.
// A model call with no span attributes is a log line. Give every agent step the
// same keys so cross-vendor tracing actually joins up.
const { trace } = require('@opentelemetry/api');
async function callModel(prompt, meta) {
const span = trace.getTracer('agent').startSpan('llm.call');
span.setAttribute('gen_ai.system', meta.provider);
span.setAttribute('gen_ai.request.model', meta.model);
span.setAttribute('agent.step', meta.step); // plan | act | verify
span.setAttribute('agent.run_id', meta.runId);
try {
const out = await provider.generate(prompt);
span.setAttribute('gen_ai.usage.input_tokens', out.usage.input);
span.setAttribute('gen_ai.usage.output_tokens', out.usage.output);
return out;
} finally {
span.end();
}
}Do Not Weld the Escape Hatch Shut
Analysts are blunt about CloudWatch Omni too: it can reduce tool fragmentation and speed up investigations into agent behavior, but lock-in and rising telemetry costs could limit its appeal. The answer is not to refuse these platforms. It is to write portability into the architecture. As long as your telemetry is OpenTelemetry and leaves through OTLP, the backend is a decision you can revisit rather than a marriage you cannot undo. The thing that locks you in was never the model provider. It is the telemetry format. Concretely, that means three commitments. Export with OTLP rather than a vendor SDK. Keep span attributes on a public convention so a new backend can read them unchanged. And treat any proprietary field name as a cost you pay at migration time.
# The survey number that should worry you: 89% of teams have some form of agent
# observability, only 37.3% run online evaluations. Tracing is not evaluation.
# Wire an eval into the same pipeline that emits the traces, or you are watching
# a system you cannot grade.
def online_eval(trace_record, rubric):
score = rubric.grade(trace_record.output)
emit_metric("agent.eval.score", score, {
"run_id": trace_record.run_id,
"step": trace_record.step,
"model": trace_record.model,
})
# Alert on the distribution, not one bad answer.
return score < rubric.floor
OpenTelemetry is your escape hatch — do not give it up
Wire Evaluation Into the Same Pipeline That Emits Traces
Tracing is not evaluation. For an eval to mean anything it has to share the pipeline: grade the output at the end of each model call, emit the score as a metric, and alert on the distribution rather than on one bad answer. That is what lets you answer the real question — is this agent getting better or worse — instead of the narrow one about what it just printed. The mechanism is mundane and the payoff is large. When the same run id appears on the trace, the eval score, and the cost row, you can finally ask which change made the agent better, and get an answer instead of a hunch.
// Cost is the second half of observability. Token spend per agent step is easy
// to compute and impossible to argue with. Alert on the slope, not the total.
type StepUsage = { runId: string; step: string; input: number; output: number };
const PRICES = { in: 1.40 / 1e6, out: 4.40 / 1e6 }; // example per-token rates
function stepCost(u: StepUsage) {
return u.input * PRICES.in + u.output * PRICES.out;
}
function runCost(usages: StepUsage[]) {
return usages.reduce((sum, u) => sum + stepCost(u), 0);
}
// Chart cost per completed agent run. If it trends up while task success is flat,
// you bought tokens, not outcomes.Cost Is the Other Half of Observability
The second half of observability is money. Token spend per agent step is easy to compute and impossible to argue with. Watch the slope, not the total: chart cost per completed agent run, and if it trends up while task success stays flat, you bought tokens, not outcomes. No new tooling is required — multiply per-step input and output tokens by published rates, then alert on the trend. The arithmetic is deliberately boring. Input and output tokens times the published rate, grouped by agent step and by run, is enough to separate growth from waste, and it needs no new vendor relationship to compute.
# Before you standardize on any single vendor, do the boring import. Export a
# week of spans and read them yourself. A vendor migration is cheap if your
# telemetry is portable and expensive if it is not.
otel-cli export --endpoint "$OTEL_EXPORTER_OTLP_ENDPOINT" \
--service agent-gateway --since 7d > spans.jsonl
wc -l spans.jsonl
jq -r '.attributes["gen_ai.system"]' spans.jsonl | sort | uniq -c | sort -rn
# If more than one provider shows up and your spend column cannot explain it,
# you have an observability problem before you have a vendor problem.Tracing is table stakes; online evaluation is the divide
Migration Is Cheap If You Prepared for It
A practical closing move: before you standardize on any vendor, do the boring import. Export a week of spans and read them yourself. Which model vendors show up, how often each is called, and whether your spend column can explain it. If the answer is that it cannot, you have an observability problem before you have a vendor problem. Portable telemetry makes migration cheap; non-portable telemetry makes it expensive. That choice is in your hands today. That exercise is slower than adopting a platform, and far cheaper than adopting the wrong one. An afternoon spent reading a week of your own spans will tell you more about your observability needs than any comparison chart.
📌 Frequently Asked Questions
When did CloudWatch Omni launch?
AWS announced general availability of Amazon CloudWatch Omni on September 23, 2026. It is an application-centric, AI-powered, OpenTelemetry-based observability experience reached through an organization-specific URL with SSO rather than the AWS Management Console, and it includes a dedicated agent observability and evaluation-driven development workflow.
What is different about OpenObserve v1.0?
OpenObserve v1.0, announced September 22, 2026, headlines AI Observability: agent tracing, LLM monitoring, evaluation, and session annotation in the same platform that already handles logs, metrics, traces, and real user monitoring, for both self-hosted deployments and OpenObserve Cloud.
Why does the consolidation hide a gap?
Because visibility is not quality. LangChain's June 2026 survey of more than 1,300 professionals found 89% had implemented some form of agent observability but only 37.3% ran online evaluations. You can trace every call and still be unable to grade agent performance.
What are the main analyst concerns about CloudWatch Omni?
Two: vendor lock-in and rising telemetry costs. The mitigation is to keep telemetry portable — use OpenTelemetry and export via OTLP — so the backend stays a revisitable decision instead of an irreversible commitment.
Should I do tracing first or evaluation first?
They are only meaningful in the same pipeline. First make sure every model call carries structured span attributes (provider, model, agent step, run id), then grade outputs at the end of those calls and emit scores as metrics, alerting on the distribution rather than a single answer.
🔧 Recommended Tools
📚 Sources
- AWS — Amazon CloudWatch Omni: AI-first observability for agents and applications (What's New)
- AWS News Blog — Introducing Amazon CloudWatch Omni: AI-powered observability for generative AI and agentic workloads
- InfoWorld — AWS launches CloudWatch Omni to unify observability for AI agents and applications
- HPCwire / BigDATAwire — OpenObserve Reaches v1.0, Bringing AI Observability into Same Platform as Logs, Metrics, Traces, and RUM
- CX Today — The Monte Carlo Agent Trust Platform (LangChain 89% / 37.3% and Salesforce 12-agents figures)