The Confidence Gap: What Harness's State of Agent DLC 2026 Says About Agent Governance

·11 min read·Evergreen Tools Team

On September 10, 2026 Harness published The State of Agent DLC 2026, a survey of 700 technology professionals conducted by Sapio Research in July 2026. Every respondent worked at an organisation with more than 1,000 employees, at least 100 developers, and more than $100 million in revenue, and every organisation had already deployed AI agents in production, in pilot, or in a live proof of concept. The headline finding is not any single percentage. It is a pattern that repeats in every domain the survey measured: confidence sits in the mid-70s while the control that would justify it is in place for fewer than half of organisations, and in some cases fewer than one in five.

1. One Gap Across Five Domains

The report spans testing, security, inventory, cost, and rollback, and the shape is identical each time. 77% of organisations are confident they have a complete inventory of every agent, MCP server, and LLM in their environment, but only 44% run active discovery tooling to verify it. 74% are confident their testing would catch a production-impacting failure, but only 19% have a gate that automatically blocks every bad release. 76% believe they could disable a misbehaving agent in under 15 minutes, but only 33% have an instant kill switch. 74% say they have a complete picture of true spend per agent, yet 60% overran their budget last quarter. The most telling number is security: 75% describe their agents as secure end to end, and that group reported security incidents at almost the same rate as everyone else, 88% against 87%.

# Gate 1: inventory. You cannot govern agents you cannot enumerate.
# 77% of organisations were confident they had a complete inventory;
# only 44% ran discovery tooling to prove it. Make the proof automatic.

async def discover(org_id: str) -> list[dict]:
    found = []
    async for ep in control_plane.list_endpoints(org_id):        # k8s, VM, laptop
        async for proc in ep.processes():
            if proc.image in AGENT_IMAGES or proc.cmdline_has("--acp"):
                found.append({"host": ep.id, "agent": proc.image,
                              "version": proc.version, "owner": ep.labels.get("team")})
    async for srv in mcp_registry.list_servers(org_id):           # MCP servers
        found.append({"host": srv.endpoint, "agent": "mcp:" + srv.name,
                      "version": srv.version, "owner": srv.owner})
    return found

def reconcile(found, declared):
    # Anything running that nobody declared is a finding, not a surprise.
    return {"undeclared": [a for a in found if a["agent"] not in declared]}
Engineering leaders in a meeting

The confidence-over-controls pattern repeats across five domains

2. Why Deterministic Controls Fail

Keith Mann, Field CTO and Head of Research at Harness, offers a precise comparison. Cloud and mobile both went through a phase where confidence outran governance, and in both cases the controls eventually caught up because the underlying systems stayed predictable once a guardrail existed. Agents do not hold still that way. The same input can produce different outputs on the next run, which is why controls designed for deterministic software miss. A check that passed cleanly in testing can still miss something in production precisely because agent behaviour varies. Code sample 2 encodes the consequence: a gate must decide whether a change passed a fixed standard, not merely whether testing was performed at some point.

// Gate 2: an eval gate, not an eval report. 74% were confident their
// testing would catch a production-impacting failure; only 19% had a gate
// that automatically blocks every bad release. The difference is "must pass".

export const releaseGate = {
  id: "agent-release-gate",
  stages: [
    { name: "evals",        mustPass: true,  minScore: 0.90, dataset: "golden-200" },
    { name: "safety",       mustPass: true,  maxViolations: 0,  suite: "jailbreak+exfil" },
    { name: "cost",         mustPass: true,  maxUsdPerRun: 0.12 },
    { name: "latency",      mustPass: true,  p95Ms: 4000 },
  ],
  onFail: "block",                 // no manual override without a signed waiver
};

// Every promoted change is checked against the same fixed standard.
export async function promote(change, gate = releaseGate) {
  const results = await runStages(change, gate.stages);
  if (results.some((r) => r.failed && r.mustPass)) {
    throw new BlockedRelease(change.id, results);   // 81% of orgs do not have this
  }
  return change;
}

3. Gate One: Inventory You Can Prove

Governance starts with enumeration, and a 33-point spread between confidence and proof means most inventories are assumptions rather than results. Code sample 1 shows a practical approach: scan running processes across endpoints and the MCP server registry, assemble what is actually executing, then reconcile against the declared list. Anything running that nobody declared is a finding rather than a surprise. The engineering value is that this turns governance from a questionnaire into telemetry. You do not need to believe the inventory is complete; you need an inventory that reports its own incompleteness.

# Gate 3: a real kill switch. 76% believed they could disable a
# misbehaving agent in under 15 minutes; only 33% had an instant switch.
# "Instant" means a control the agent cannot argue with at runtime.

import redis, time

r = redis.Redis()

def revoke(agent_id: str, reason: str, by: str) -> None:
    # Step 1: flip the flag the runtime checks on every single action.
    r.set(f"agent:{agent_id}:revoked", "1", ex=86400)
    # Step 2: record who did it, when, and why. Auditable, not improvised.
    r.xadd("agent-revocations", {
        "agent": agent_id, "reason": reason, "by": by,
        "ts": str(time.time()),
    })

def may_act(agent_id: str) -> bool:
    # Checked before every tool call, not once at session start.
    if r.get(f"agent:{agent_id}:revoked"):
        raise PermissionError("revoked: %s" % agent_id)
    return True
The gap between confidence and control

75% say agents are secure, yet that group's incident rate matched the average

4. Gates Two and Three: Must-Pass Evals and a Switch That Works

There are two very different things called evals: a report and a gate. A report tells you the score; a gate decides whether the change ships. Code sample 2 defines a must-pass release gate where evals, safety, cost, and latency are all hard conditions and a failure blocks promotion unless someone signs a waiver, which is exactly the capability only 19% of organisations have. Code sample 3 addresses rollback, where 76% confidence meets 33% capability. The missing piece is a runtime switch the agent cannot argue with. Note two details in the sample: revocation is checked before every tool call rather than once at session start, and the revocation itself is written to an audit stream recording who did it, when, and why. The fifteen-minute gap usually lives in those two details.

# Gate 4: spend per agent, with a ceiling. 74% said they had a complete
# picture of true spend; 60% still overran budget last quarter. Visibility
# without a limit is a dashboard, not a control.

BUDGET = {"research-agent": 900.00, "pr-review-agent": 250.00}

def admit(agent_id: str, est_usd: float, month_spend: dict) -> dict:
    cap = BUDGET.get(agent_id)
    if cap is None:
        return {"allow": False, "why": "no budget line: " + agent_id}
    spent = month_spend.get(agent_id, 0.0)
    if spent + est_usd > cap:
        return {"allow": False, "why": "over cap", "spent": spent, "cap": cap}
    return {"allow": True, "remaining": round(cap - spent - est_usd, 2)}

# Route the request down a cheaper tier before refusing outright.
def degrade(est_usd: float) -> str:
    return "cheap-tier" if est_usd > 0.05 else "default-tier" 

5. Gate Four: Turning Cost Into a Ceiling

Cost is the most easily ignored finding. Three quarters of organisations say they can see true spend per agent, yet 60% overran budget, which tells you visibility is not control. Code sample 4 turns budget into an admission decision: an agent with no budget line is refused, an estimated cost that would breach the cap is refused, and before refusing, the request is routed to a cheaper tier. The ordering matters. Degrade first and refuse second, and governance shows up as cost optimisation rather than as work being blocked. Given that 58% report more production incidents per 100 changes since deploying agents, putting spend and incident rate on one dashboard is the cheapest place to start.

// Agent changes are not code changes. 42% route prompt edits through the
// same pipeline as code, and just 34% have a dedicated configuration
// system for AI behaviour. A schema gives you a reviewable diff instead
// of a one-line prompt tweak that silently changes tool permissions.

{
  "id": "pr-review-agent",
  "model": { "primary": "gpt-5.6-sol", "fallback": "claude-fable-5.1" },
  "prompt": { "ref": "prompts/pr-review.md", "rev": "sha256:41ab..." },
  "tools": {
    "allow": ["readRepo", "runTests", "commentOnPr"],
    "deny": ["pushToMain", "writeSecrets", "deleteBranch"]
  },
  "budget": { "usdPerMonth": 250, "usdPerRun": 0.12 },
  "rollout": { "strategy": "canary", "percent": 10, "watch": "errorRate<1.5%" },
  "evaluations": { "gate": "agent-release-gate", "minScore": 0.9 }
}

// Now "we changed the agent" comes with a diff, an owner, and a reason.
A production pipeline

58% report more production incidents per 100 changes

6. An Agent Change Is Not a Code Change

The report also measures process. 42% of organisations run prompt edits through the same pipeline as a code change while just 34% have a dedicated configuration system for AI behaviour. Only 53% of agent-related changes pass through any standard pipeline before production, and 37% run less than half of theirs through one. More than four in ten decide case by case whether to trust a change, and among those that promote agent changes to production only 58% check every change against a fixed, repeatable standard. Harness recommends treating the agent lifecycle as its own discipline, building evals, security, inventory, and rollback specifically for agent behaviour, replacing ad hoc reviews with a fixed standard, and adopting progressive rollout for agents rather than relying on manual gates alone. Code sample 5 offers an agent definition schema so that we changed the agent produces a record with an owner, a diff, and a reason, instead of a one-line prompt tweak that quietly altered tool permissions.

📌 Frequently Asked Questions

Who did Harness survey?

The report is based on a survey of 700 technology professionals in the United States, United Kingdom, France, Germany, and India, conducted by Sapio Research in July 2026. Respondents worked at organisations with more than 1,000 employees, 100 or more developers, and over $100 million in revenue, and had already deployed AI agents.

What is the central finding?

A systematic gap between confidence and control. Across testing, security, inventory, cost, and rollback, confidence in agents sits in the mid-70s while the control that would justify it is in place for fewer than half of organisations, and in some cases fewer than one in five.

Why do agents need different controls than deterministic software?

Agents are not deterministic: the same input can produce different outputs from one run to the next, so controls built for deterministic software miss. Evals, gates, and rollback need to be built specifically for that variability.

What does the report recommend?

Treat the agent lifecycle as its own discipline, building evals, security, inventory, and rollback specifically for agent behaviour; replace ad hoc reviews with a fixed, repeatable standard; and adopt progressive rollout such as canary and blue/green for agent changes.

Which numbers should leadership focus on?

Three: 75% call their agents secure but that group reported incidents at 88% versus 87% overall; 76% believe they could disable a misbehaving agent quickly but only 33% have an instant kill switch; and 58% report more production incidents per 100 changes.