The Tool Tax: How Least-Privilege MCP Cuts Agent Token Costs Before They Happen

·10 min read·Evergreen Tools Team

Most AI teams audit cost by looking at model pricing, cache hit rate and context length. Almost none look at the menu of tools riding inside every prompt. In a post on Okta's newsroom, Harish Peri, SVP and GM for AI Security, and Jenna Cline, SVP of Technology, Data and Intelligence, put a name on that cost: the tool tax. Every tool an MCP server exposes is rendered into the model's prompt as a schema, name, description and parameters, on every turn, whether or not the model will ever call it. It does not show up as a rejected call or a security incident. It shows up as a token bill that is harder to explain than it should be.

1. Defining the tool tax

Okta opens with a simple observation: every AI agent's model call includes a menu of every tool it might use, even the ones it will never touch. That is true whether the agent connects to Google Workspace, Slack, or an internal MCP server with a hundred tools bolted on over several years. Every exposed tool is rendered into the prompt as a schema, name, description and parameters, and that happens on every single turn. You pay for the model to reason about that tool whether or not it is ever called. Okta calls this the tool tax, and notes why it has stayed invisible: it does not appear as a rejected call or a security incident. It appears as a token bill that is harder to explain than it should be.

# The tool tax, in one function. Every exposed tool costs tokens on
# every turn, because its schema rides along in the prompt.

def tool_schema_tokens(tools):
    """Roughly: each tool contributes a name, description and parameter
    schema -- every turn, whether or not it is ever called."""
    return sum(len(t.name) + len(t.description) + schema_size(t.params)
               for t in tools)

def monthly_tool_tax(tools, turns_per_conversation, conversations):
    return (tool_schema_tokens(tools)
            * turns_per_conversation
            * conversations)

# Okta's point: this line does not appear as a rejected call or a
# security incident. It appears as a token bill that is hard to explain.
Converging tool entitlements at the identity layer

Shrink the list before you meter it

2. Why after-the-fact rejection recovers nothing

The instinctive fix is a validation layer that blocks unauthorised calls. Okta explains why that saves nothing. Rejecting an unauthorised tool call does not get the tokens back, because they were already spent in the prompt phase, before the model tried anything. The cleanup is worse than the waste: once an agent has been using a tool it should not have access to, walking that back is its own project, because someone has to notice, investigate and revoke the access — and the tokens stay spent regardless. Okta's conclusion follows directly. Scoping upstream means there is nothing to claw back in the first place.

# The failure mode. Rejecting an unauthorised call after the fact
# does not refund anything: the schema was already in the prompt
# before the model tried anything.

def post_hoc_rejection_flow(agent, tool_call):
    tokens_already_spent = True          # paid in the prompt phase
    log_incident(tool_call)              # someone must investigate
    revoke_access(tool_call.tool)        # a project of its own
    return {"refund": None if tokens_already_spent else "n/a",
            "tokens_recovered": 0}

# Okta's conclusion: scoping upstream means there is nothing to claw back.

3. Scope at the identity layer, not the gateway

The mechanism itself is narrow and worth restating precisely. Okta scopes the tool list at the point where an agent connects to an MCP server. An admin configures, in the Okta dashboard, which tools a given identity is entitled to use, and Okta returns only the scoped set rather than everything the server exposes. Three things follow. The shorter list is injected into the prompt on every turn, so per-turn token spend falls before the agent attempts anything. Unauthorised tools are not in the list the model can see at all. And scope is checked again at runtime before any tool call executes. That third point matters: this is not trading security for cost. The same entitlement data both shrinks the prompt and shrinks the blast radius.

# Scoping upstream, at the identity layer. The admin configures which
# tools a given identity is entitled to; Okta returns only the scoped
# set rather than everything the server exposes.

def tools_returned_to_model(identity, mcp_server):
    catalogue = mcp_server.exposed_tools()
    entitled  = identity.entitlements_for(mcp_server)
    return [t for t in catalogue if t.name in entitled]

# Consequences of the shorter list:
#   1. fewer schema tokens per turn, before any call is attempted
#   2. unauthorised tools are absent from what the model can even see
#   3. scope is checked again at runtime before any call executes
Tool count and token cost track nearly linearly

Schema cost tracks tool count almost linearly

4. The numbers, with their caveats attached

The quantitative claim is that in internal modeling across a realistic mix of permission levels, some scenarios cut visible tools by more than 90 percent, and tool-schema cost falls by roughly the same margin. The caveats have to travel with the numbers or the numbers are meaningless. The figures come from Okta internal modeling using Okta product data and public vendor documentation only, and no customer data was used. The model described a single MCP client with access to a catalog of enterprise tools, defined representative user segments — helpdesk read-only, helpdesk operator, app admin, brand and email admin, and super admin — and weighted each by an assumed share of monthly traffic. Because every tool contributes a name, description and parameter schema on every turn, schema token cost tracks tool count nearly linearly, which is why the post reports percentages rather than dollars. Real results will vary with your catalog, permission distribution and model choice.

# What the internal modeling found, with its caveats attached. The
# figures come from Okta internal modeling on product data and public
# vendor documentation only -- no customer data. They are reported as
# percentages, not dollars, because absolute figures vary by
# deployment.

MODELING = {
    "reduction_reported_as": "percentages, not dollars",
    "user_segments": ["helpdesk read-only", "helpdesk operator",
                      "app admin", "brand and email admin", "super admin"],
    "weighting": "assumed share of monthly traffic per segment",
    "headline": "some scenarios cut visible tools by more than 90%",
}

def schema_cost_reduction(tool_count_before, tool_count_after):
    ratio = 1 - tool_count_after / tool_count_before
    return {"tool_count_reduction": ratio,
            "schema_cost_reduction_estimate": ratio}   # near-linear

5. Gateway and identity layer: a division of labour

Okta dedicates a section to the relationship with gateways, and the framing is useful. Gateways cap spend by key, team or group, the finest grain their data supports, and for routing and rate-limiting that is the right layer to own the job. But group-level information cannot tell a gateway what one specific agent, or the person behind it, is actually entitled to touch. So a gateway can only cap the damage after a decision gets expensive, not prevent the decision from being expensive — and capping spend after the fact makes whoever owns the gateway the token police by default, defending limits when teams push back and taking the call when a user hits a wall mid-task. The identity layer filters at per-user and per-agent entitlements rather than group membership. Okta's summary is tight: the gateway meters what gets through, the identity layer shrinks what there is to meter.

# Gateway versus identity layer: complementary, different grains.
# A gateway meters what already happened. An identity layer shrinks
# what there is to meter.

LAYERS = {
    "gateway":  {"grain": "key / team / group",
                 "job":  "rate-limit, route, cap spend after the fact"},
    "identity": {"grain": "per user, per agent",
                 "job":  "decide which tools exist in the prompt at all"},
}

def who_owns_token_policing(has_identity_layer):
    return "nobody, ideally" if has_identity_layer else "whoever owns the gateway"
Tokens spent in the prompt phase are gone

Rejecting a call afterwards saves nothing

6. A workable order of operations

If you want to apply this, sequence matters more than tooling. First, do the arithmetic: multiply tool count by per-turn schema cost, then by turns per conversation and conversations per month, and turn the tool tax into a number you can put in a meeting. Many teams discover it is larger than they assumed. Second, shrink: scope each identity's tool list to its entitled subset and move the scope check to before the call executes rather than after. Third, and only third, meter: once the first two steps are done, gateway spend caps become a backstop rather than the only lever, and the gateway team stops being the token police. The value of that order is that it converts a cost problem into a permissions problem you already have the data to solve.

📌 Frequently Asked Questions

What exactly is the tool tax?

Okta's definition: every AI agent's model call includes a menu of every tool it might use, including the ones it will never touch. Every tool an MCP server exposes is rendered into the prompt as a schema, name, description and parameters, and that happens on every single turn. You pay for the model to reason about that tool whether or not it ever calls it.

Why does rejecting a tool call afterwards not help?

Because rejecting an unauthorised tool call does not get those tokens back. They were already spent in the prompt phase, before the model tried anything. And once an agent has been using a tool it should not have access to, walking that back is its own project: somebody has to notice, investigate and revoke the access. The tokens stay spent regardless.

What is Okta's actual mechanism?

Okta scopes the tool list at the point where an agent connects to an MCP server. Within the Okta dashboard an admin configures which tools a given identity is entitled to use, and Okta returns only the scoped set rather than everything the server exposes. A shorter list is injected into the agent's prompt on every turn, which means less token spend per turn before the agent attempts a call. Unauthorised tools are absent from the list the model sees, and scope is checked again at runtime before any tool call executes.

What did the modeling find, and what are the caveats?

Okta says that in internal modeling across a realistic mix of permission levels, some scenarios cut visible tools by more than 90 percent, with tool-schema cost falling by roughly the same margin. The caveats matter: the figures come from Okta internal modeling using Okta product data and public vendor documentation only, with no customer data. The model defined representative user segments (helpdesk read-only, helpdesk operator, app admin, brand and email admin, super admin) weighted by an assumed share of monthly traffic, and because each tool contributes a name, description and parameter schema to the prompt on every turn, schema token cost tracks tool count nearly linearly. That is why the post reports percentages rather than dollars, and why actual results vary with your catalog, permission distribution and model choice.

Does this compete with your API gateway?

Okta argues it is a division of labour rather than a conflict. Gateways cap spend by key, team or group, which is the finest grain their data supports, and for routing and rate-limiting that is the right layer to own it. But a gateway meters what already happened — tokens in, tokens out, dollars spent — so it can cap the damage after a decision gets expensive, not stop the decision from being expensive. The identity layer filters a tool list at per-user and per-agent resolution, which shrinks what there is to meter in the first place.