Kong AI Gateway 2.2 Goes GA: A Control Plane for MCP Tools, Models and Agent Spend
On September 30, 2026, Kong announced the general availability of Kong AI Gateway 2.2. The announcement describes a standalone platform providing four headline capabilities: unified MCP tool governance, modality-aware cost management, identity-aware AI policies and expanded provider support, plus a new AI-native user experience, advanced cost management and broader model coverage, generally available in Kong Konnect with no beta enablement required. Kong's rationale is worth separating from the feature list: new models, agent architectures, protocols and pricing models are emerging in weeks rather than quarters, and the gap between the pace of AI innovation and the infrastructure enterprises rely on to govern it keeps widening.
1. What shipped in 2.2
Start with the facts. On September 30, 2026, Kong Inc. announced the general availability of Kong AI Gateway 2.2. The announcement's subhead summarises four capabilities: unified MCP tool governance, modality-aware cost management, identity-aware AI policies, and expanded provider support. The body adds that it features a new AI-native user experience, advanced cost management and broader model support, and that it is now generally available in Kong Konnect, Kong's AI Connectivity Platform, with no beta enablement required. Reza Shafii, Senior Vice President of Product, frames the positioning in the release: the promise of agentic AI is not just what an agent can do, but whether an enterprise can confidently put it to work.
# What Kong announced on 2026-09-30: AI Gateway 2.2 is generally
# available in Kong Konnect, with no beta enablement required. The
# announcement frames it as a standalone platform providing four
# headline capabilities.
AI_GATEWAY_2_2 = {
"release": "Kong AI Gateway 2.2",
"status": "GA in Kong Konnect (no beta enablement)",
"highlights": [
"unified MCP tool governance",
"modality-aware cost management",
"identity-aware AI policies",
"expanded provider support",
],
"also": ["new AI-native user experience", "broader model support"],
}Governance cannot live in one part of the stack
2. Why governance moves from a layer to a platform
Kong's argument rests on a speed differential. In its words, AI is outpacing traditional infrastructure: new models, agent architectures, protocols and pricing models are emerging in weeks rather than quarters, creating a growing gap between the pace of AI innovation and the infrastructure enterprises rely on to govern it. When AI connectivity spans models, agents, intelligent tools and context across increasingly diverse environments, governance cannot be limited to a single part of the stack. Kong's answer is to give AI its own platform and release cadence, so enterprises can adopt what is next without compromising security, compliance, cost control, resilience or visibility. There is a practical test buried in that claim: if a governance approach has to keep pace with model releases to be worth having, it cannot sit on a system that ships quarterly.
# The rationale, in Kong's own words: new models, agent architectures,
# protocols and pricing models emerge in weeks rather than quarters,
# creating a growing gap between the pace of AI innovation and the
# infrastructure enterprises rely on to govern it.
def governance_coverage(ai_surfaces, governed_surfaces):
"""Governance limited to one part of the stack leaves the rest
unaccounted for -- and AI connectivity spans models, agents,
intelligent tools and context."""
return {
"covered": [s for s in ai_surfaces if s in governed_surfaces],
"uncovered": [s for s in ai_surfaces if s not in governed_surfaces],
}
AI_SURFACES = ["models", "agents", "intelligent_tools", "context"]3. Unified MCP tool governance: granularity is the real question
The word unified hides most of the substance, so ask about granularity. First, can you enumerate the full tool catalogue a server exposes? If you cannot produce the inventory, no policy built on top of it is trustworthy. Second, can you restrict which tools a caller sees, not merely which it may call? That distinction carries real money, because once a tool's schema is in the prompt the tokens are already spent. Third, is authorisation per-tool or all-or-nothing at the server? Fourth, are invocations logged together with caller identity and outcome? Fifth, can you revoke a single tool without disturbing the rest? If those five questions come back vague, unified governance is a dashboard, not a control.
# Unified MCP tool governance is the capability to test hardest,
# because an MCP server is where tool inventory and authority collide.
# The questions that decide whether "unified" is real:
MCP_GOVERNANCE_QUESTIONS = [
"Can I enumerate the full tool catalogue a server exposes?",
"Can I restrict which tools a given caller sees, not just calls?",
"Is authorisation per-tool, or all-or-nothing at the server?",
"Are tool invocations logged with caller identity and outcome?",
"Can I revoke one tool without disturbing the rest?",
]
def coverage_score(answers):
return sum(1 for a in answers if a) / len(MCP_GOVERNANCE_QUESTIONS)MCP, model and A2A traffic governed in one place
4. Modality-aware cost management: why a total is not enough
The load-bearing words here are modality-aware. Price is not one number: text, image, audio and video are billed differently, input tokens are priced differently from output tokens, and cached tokens follow a third schedule — before you even layer on model family, provider, tenant and caller. A blended total gives you almost nothing to act on. Cost management worth the name breaks spend down across those dimensions and sorts the buckets by size, so finance receives an explainable invoice and engineering receives a prioritised list of what to optimise. Kong lists cost management among the core 2.2 capabilities for a reason: whether you can govern the economics is what decides whether agentic AI is viable in the enterprise.
# Modality-aware cost management matters because price is not a single
# number: text, image, audio and video are billed differently, and so
# are input versus output tokens. A gateway that reports one blended
# figure cannot tell you which workload got expensive.
COST_DIMENSIONS = ["input_tokens", "output_tokens", "cached_tokens",
"modality", "model_family", "provider", "tenant", "caller"]
def explainable_invoice(cost_rows, by):
"""Hand finance a breakdown, not a total."""
buckets = {}
for row in cost_rows:
buckets.setdefault(row[by], 0)
buckets[row[by]] += row.cost
return dict(sorted(buckets.items(), key=lambda kv: -kv[1]))5. Identity-aware policies close the loop
Identity awareness is the capability that makes the other two auditable. A policy is only useful if, after the fact, it can answer three questions without an engineer reconstructing them from scattered logs. Who initiated this call — a human identity or an agent identity? What was used — which tool, which model, which model version, and at what scope? And which policy permitted it, and was it overridden? If answering those takes a war room, the governance is decorative when it matters. Shafii ties this directly to business value: confidence in putting agents to work requires visibility into what is happening, control over what agents can access and execute, and the ability to manage the risks and economics that come with it.
# Identity-aware policies are the part that makes the rest auditable.
# A policy is only useful if you can answer three questions about it
# after the fact -- without asking an engineer to reconstruct them.
AUDIT_TRIO = {
"who": "which human or agent identity initiated this call?",
"what": "which tool, model, model version and scope was used?",
"why": "which policy allowed it, and was it overridden?",
}
def on_policy_violation(event):
return {
"action": "block or escalate per policy", # decide before launch
"record": {"identity": event.who, "surface": event.what,
"policy": event.matched_policy},
"owner": "named human owner for every long-running agent",
}Who called, what they called, and what it cost
6. A short checklist before you standardise
If you are considering making one gateway the policy point for all AI traffic, run three stress tests first. Test it against your dirtiest MCP server — dozens of tools, messy entitlements, some untouched for years — and see whether governance survives contact with the real catalogue. Demand an explainable bill, broken down by modality, model family, tenant and caller, and check that it reconciles with how your finance team already categorises spend. Then take one long-running agent and try to answer the three audit questions: who, what, and under which policy. Finally, ask the reverse question. Kong's own pitch is that AI needs its own platform and release cadence; apply the same standard to your shortlist and ask how long the gap was between the last two capability releases, and when the next one lands.
📌 Frequently Asked Questions
What did Kong announce on September 30, 2026?
Kong Inc. announced the general availability of Kong AI Gateway 2.2, featuring a new AI-native user experience, advanced cost management and broader model support. The announcement says the new standalone platform provides unified MCP tool governance, modality-aware cost management, identity-aware AI policies and expanded provider support, and that 2.2 is now generally available in Kong Konnect, the AI Connectivity Platform, with no beta enablement required.
Why does governance belong at the gateway layer?
Kong's argument is that AI is outpacing traditional infrastructure: new models, agent architectures, protocols and pricing models emerge in weeks rather than quarters, creating a growing gap between the pace of AI innovation and the infrastructure enterprises rely on to govern it. When AI connectivity spans models, agents, intelligent tools and context across increasingly diverse environments, governance can no longer be limited to a single part of the stack.
What does unified MCP tool governance actually imply?
The announcement describes consolidating MCP tool governance into one platform capability. The useful questions are about granularity: can you enumerate the full tool catalogue a server exposes; can you restrict which tools a caller sees, not just which it may call; is authorisation per-tool or all-or-nothing at the server; are invocations logged with caller identity and outcome; and can you revoke one tool without disturbing the rest?
Why does modality-aware cost management matter?
Because price is not a single number. Text, image, audio and video are billed differently, and input tokens are priced differently from output tokens, with cached tokens on a third schedule. A gateway that reports one blended figure cannot tell you which workload became expensive. Useful cost management breaks spend out by input, output and cached tokens, modality, model family, provider, tenant and caller, so finance receives an explainable breakdown instead of a total.
What problem do identity-aware policies solve?
They make the other capabilities auditable. A policy is only useful if it can answer three questions after the fact: who initiated this call, human or agent; what was used, including tool, model, model version and scope; and which policy permitted it and whether it was overridden. Reza Shafii, Senior Vice President of Product at Kong, frames the business stake directly: the promise of agentic AI is not just what an agent can do, but whether an enterprise can confidently put it to work, which requires visibility into what is happening, control over what agents can access and execute, and the ability to manage the risks and economics involved.
🔧 Recommended Tools
📚 Sources
- PR Newswire / Kong Inc. — Kong AI Gateway Expands with New Capabilities, Bringing Enterprise Governance to the Agentic AI Era (2026-09-30)
- Kong — Kong AI Gateway product page
- Kong Blog — Kong AI Gateway 2.0 Is Now GA (2026-09-01)
- PR Newswire / Kong Inc. — Kong AI Gateway Now Supports Agent-to-Agent Traffic (2026-04-14)