Splunk Open-Sourced Token Meter: A Live Cost Gauge for Coding Agents
💡 Tool Tip:AI Token Counter, CSV Statistics Calculator, JSON Formatter
Somewhere this month, an agent processed a fifteen-word user request and burned 60,000 tokens doing it. Sixty-two percent of that was the model re-reading context it had already seen, and 3,000 tokens went to tool schemas the task never used. Nobody wrote a bug. This is the default behavior of agent architectures. In September, Splunk open-sourced Token Meter, a local-first cost gauge that reads the logs your agents already write to disk and puts those numbers in front of you in real time. It does not eliminate the waste, but it makes the waste visible, and that is where every optimization has to start.
Turn a flat invoice into a cost map
1. Why Agent Bills Compound
The chatbot-era mental model was one prompt equals one call. Agents broke that assumption. An agent is a loop: plan, call a tool, read the result, plan again. The API is stateless, so every step re-sends the full accumulated context. Here is a realistic shape. A user asks whether a customer's last three orders are eligible for a refund, about twelve words. The system sends 1,500 context tokens of system instructions, another 3,000 of tool schemas including tools irrelevant to this task, and 2,500 tokens of retrieved customer data, so a twenty-token message becomes 7,000 input tokens before the model produces anything. Then the loop begins: the second call sends 10,000 tokens, the third 13,000, the fourth 17,000 as tool results accumulate. By the fifth call, the model has reprocessed the same context four times over. A five-step task that should cost fewer than 10,000 tokens ends up consuming more than 40,000.
# Token Meter is local-first: it reads the trace logs your agents
# already write to disk, then prices them against public model rates.
# Here is the shape of that math on a captured session record.
RATES = {"gpt-5.6-terra": (3.00, 15.00), "claude-sonnet": (3.00, 15.00)}
def session_cost(records):
total = 0.0
for r in records:
in_rate, out_rate = RATES[r["model"]]
total += r["input_tokens"] / 1_000_000 * in_rate
total += r["output_tokens"] / 1_000_000 * out_rate
return round(total, 4)
# A run that "looked cheap per step" is only cheap until you sum the loop.2. Six Buckets: What Your Tokens Actually Bought
The way to turn a flat invoice into a map is classification. Agent tokenomics sorts every call into six buckets: context tokens (system instructions, conversation history, tool schemas, all riding along on every call), reasoning tokens (planning, chain-of-thought, intermediate steps), retrieval tokens (documents pulled in by RAG), tool tokens (the results tools return, which enter the context window and ride along on every later step), coordination tokens (role prompts, synchronization messages, shared state in multi-agent systems), and governance tokens (validation, safety checks, evaluation, human review triggers). The last bucket is the counterintuitive one: it is often invisible on the bill, so it is never budgeted. Research from the Stanford Digital Economy Lab found that re-sent context can account for 62% of agent inference bills, and the waste inside it is specific: irrelevant schemas, stale history, and stable prefixes that could be cached but are not. The real-world spread is striking. One audit across thirty engineering teams found a 20x gap between the cheapest and most expensive developers doing similar work, and one company went from $87,000 a month to $24,000 after classifying its token spend and changing the architecture, with sprint velocity staying flat.
// Classify tokens into buckets so you know which spend bought progress.
const BUCKETS = ["context", "reasoning", "retrieval", "tool", "coordination", "governance"];
function classify(call) {
const out = Object.fromEntries(BUCKETS.map((b) => [b, 0]));
out.context = call.systemTokens + call.historyTokens + call.toolSchemaTokens;
out.reasoning = call.reasoningTokens;
out.retrieval = call.ragTokens;
out.tool = sum(call.toolResults.map((t) => t.tokens)); // responses ride along
out.coordination = call.multiAgentTokens;
out.governance = call.evalTokens + call.reviewTokens; // usually invisible
return out;
}
// Rule of thumb from published analyses: re-sent context can be the
// single largest line, so attack that bucket first.3. What Token Meter Is: A Local Gauge You Can See
Token Meter is Splunk's open-source, local-first usage and cost dashboard for AI coding agents, originating from Galileo Agent Labs. Its job is prosaic and useful: it reads the trace logs your agent sessions already write to disk and turns them into token usage, estimated cost, context pressure, wait time, output pace, tool calls, and session duration, all in one view. It covers Claude Code, Codex, Cursor, OpenCode, Kiro, and Pi; there is a macOS menu-bar companion, a Linux tray companion, and a Windows extension still in beta. Three properties are worth underlining. It is Python standard library only, and trace analysis needs no API keys. It reports no telemetry, so nothing leaves your machine. And it ships a local MCP server so Codex or Claude can query bounded evidence. After installing, open http://127.0.0.1:8722, run a normal session, and pick it from Sessions.
# Cost per accepted task is the real price of one good result:
# model cost + tool/runtime cost + human review cost.
def cost_per_accepted_task(model, tool, review_rate, minutes_reviewed, hourly):
review = (minutes_reviewed / 60) * hourly
return round(model + tool + review * review_rate, 4)
# A "cheap" $0.02 support answer looks efficient, until 40% need a
# 15-minute human review at $30/hour:
reviewed = cost_per_accepted_task(0.02, 0.00, 0.40, 15, 30) # ~3.02
clean = cost_per_accepted_task(0.02, 0.00, 0.00, 0, 30) # 0.02
print(reviewed, clean) # the blended number sits far above the invoice4. The Two Metrics That Actually Matter
A total spend number tells you nothing about value. You need two metrics that do. The first is token yield: successful sessions per million tokens. It only means something once you define successful, for example the task completed, hallucination stayed under your threshold, and nobody had to escalate to a human, with all conditions required to count. The second, and closer to reality, is cost per accepted task: model cost plus tool and runtime cost plus human review cost. That formula breaks a lot of illusions. A support agent on a cheap model at $0.02 a task looks efficient, but if 40% of its answers need a fifteen-minute human review at $30 an hour, the real cost of those answers is about $3.02, while the clean ones stay at $0.02. The blended number sits far above the invoice.
// Token yield = successful sessions per million tokens.
// You must define "successful" before the number means anything.
function tokenYield(sessions, qualityBar) {
const successful = sessions.filter((s) =>
qualityBar.every((check) => check(s)) // done, no hallucination, no escalation
).length;
const tokens = sessions.reduce((sum, s) => sum + s.totalTokens, 0);
return (successful / tokens) * 1_000_000;
}
// A refund lookup and a summarization are different workflows,
// so each gets its own bar and its own yield. Never blend them.5. Where to Start: Use the Numbers You Already Have
A common mistake is switching models or pruning context first. The right starting point is data you can already pull: every API response gives you total input and output tokens per call. Run at that granularity for a week, find your worst offenders, then add the finer bucket split as tracing improves. Two numbers supply the context. Enterprise AI spending climbed from about $1.2 million per company in 2024 to $7 million in 2026, while per-token API prices fell roughly 280 times over the same period. The unit got cheaper and consumption exploded; reports say Uber burned through its entire 2026 Claude Code budget by April. The problem is not the unit price. It is that nobody knows which step the money went to.
# Guard the runaway. Alert on the distribution, not just the average:
# a handful of very long sessions usually dominates the bill.
def budget_alerts(sessions, monthly_budget, p95_minutes=45):
spend = sum(s["cost"] for s in sessions)
if spend > 0.8 * monthly_budget:
alert("agent budget at 80% - inspect the top spenders")
for s in sessions:
if s["minutes"] > p95_minutes:
alert("session " + s["id"] + " ran long: " + str(s["minutes"]) + " min")
# Splunk's Token Meter streams this live, entirely on your machine:
# no API keys, no telemetry, dashboard at http://127.0.0.1:87226. A Checklist: Turning Visibility into Budget Discipline
Five moves. First, measure before you optimize; do not prune context on instinct. Second, write success as checkable conditions, or token yield is meaningless. Third, fold human review cost into your unit economics, because it is often ten times the invoice. Fourth, alert on the distribution rather than the average: a handful of very long sessions usually dominates the bill, so watch p95 session duration and sessions still open past your cap. Fifth, treat tool responses as a first-class cost, because the call is cheap but its response stays in the context window and is billed on every later step. One idea ties it together: tokens are not a rounding error on the invoice, they are the working capital of intelligence. The question is not how much you spent but which tokens bought progress and which bought nothing at all.
Local-first: nothing leaves the machine, no telemetry
📌 Frequently Asked Questions
Why are agents so much more expensive than chatbots?
Because agents loop, and the API is stateless: each step re-sends the accumulated context, so token use compounds with every call. A very short request can end up billing far more tokens than the useful work needs.
Does Token Meter need API keys or a network connection?
No. It is a local-first tool built on the Python standard library that reads the trace logs your agents already write to disk to estimate cost. It reports no telemetry, nothing leaves your machine, and it includes a local MCP server for bounded evidence queries.
Which coding agents does it support?
The documented runtimes are Claude Code (including Desktop Agent/Cowork), Codex CLI and desktop, Cursor Agent/Composer, OpenCode, Kiro, and Pi, with full support on macOS and Linux and a Windows extension still in beta.
Which token bucket should I optimize first?
Measure first. Sort one week of calls into context, reasoning, retrieval, tool, coordination, and governance tokens, then attack whichever bucket dominates. Published analyses often find re-sent context is the largest single line, but only your own classification tells you for sure.
Why include human review cost in the math?
Because cost per accepted task equals model cost plus tool and runtime cost plus human review cost. A cheap $0.02 model whose answers need fifteen minutes of review 40% of the time at $30 an hour costs about $3.02 for those answers, and the blended figure sits far above the invoice.