56% Pass Rate: What Veracode's 2026 Report Says About AI-Written Code

·12 min read·Evergreen Tools Team

Veracode's 2026 GenAI Code Security Report delivers an unflattering finding: models keep getting better at writing code, and the code is not getting safer. The report covers more than 100 AI models over four years, with 11 new models tested across 80 tasks in the 2026 round. The average security pass rate sits at 56% - barely changed from 55% in the first report - while roughly 44% of tested generation tasks introduced a risky vulnerability. Modern models produce syntactically correct code nearly 100% of the time. Veracode compresses the gap into one line: syntax is solved, security is not. For engineering and security leaders, that means does it run has stopped being a useful acceptance criterion.

Syntax is solved; security is not

Syntax is solved; security is not

1. The Headline: Security Has Not Scaled With Volume

The scope first. More than 100 AI models tested over four years, across languages, weakness classes, and tasks; an average security pass rate of 56%, essentially unchanged from 55% in the first edition; and in the 2026 round, 11 new models evaluated across 80 tasks. Veracode's interpretation is that the failure rate did not move - what moved is the volume of code the failure rate now applies to. In organisations that have adopted AI coding tools, AI authors roughly half of committed code, per data Veracode cites. That is why they reframe GenAI code security as a scale problem rather than a theoretical one: when half the codebase is machine-authored and roughly 44% of generation tasks introduce a known vulnerability, review capacity, remediation queues, and governance all face the same multiplication. The upstream got faster; verification did not.

# 1. Gate by vulnerability class, not by a single pass/fail score
#    Veracode 2026 mean security pass rates by class:
#      cryptographic algorithms  87%
#      SQL injection             83%
#      cross-site scripting      15%
#      log injection             12%
HIGH_SCRUTINY = ("cwe-79", "cwe-117")   # XSS and log injection

def review_policy(finding):
    if finding.cwe in HIGH_SCRUTINY:
        return {"required": ["human_review", "sast_rerun", "exploit_test"]}
    if finding.cwe in ("cwe-89",):          # SQL injection: still verify, cheaply
        return {"required": ["parameterisation_lint"]}
    return {"required": ["sast_rerun"]}

2. Model Selection Is Now a Security Decision

A flat average does not mean a flat field. In the current snapshot GPT-5.5 leads at a 68% security pass rate, and six of the eleven Summer 2026 models cluster between 50% and 53%. Read the leaderboard from both ends: the best available model still fails on nearly one in three security tasks, and the worst fails on one in two. The report also breaks several comfortable assumptions. Models purpose-built for code average a 51% security pass rate; general-purpose models average 52% - being trained to write code faster does not mean writing it safer. Reasoning models average 56% against 51% for non-reasoning models, which suggests internal reasoning helps models catch insecure constructs before output, yet the stronger category still fails on nearly 44% of tasks. Model size barely matters either: large models average 53%, medium and small each average 51%. Veracode's conclusion is pointed - if two models both make developers faster and one consistently produces less vulnerable code, that is not a technical preference, it is a risk management decision.

// 2. Model choice as data, not as folklore
const modelPolicy = {
  default: "gpt-5.5",
  // 2026 GenAI Code Security Report, Summer 2026 dataset:
  securityPassRate: {
    "gpt-5.5": 0.68,        // leads the snapshot
    "six-of-eleven": "0.50-0.53",
    "code-specialised-mean": 0.51,   // vs 0.52 general-purpose
    "reasoning-mean": 0.56,          // vs 0.51 non-reasoning
    "large-mean": 0.53,              // vs 0.51 medium and small
  },
  rule: "A model that fails 1 in 3 security tasks creates less remediation work than one that fails 1 in 2.",
  review: "even the leader needs verification - it still fails nearly 1 in 3",
};
Model selection is now a risk management decision

Model selection is now a risk management decision

3. Language and Weakness Class: Put Review Where It Pays

The most operational numbers in the report are its per-class pass rates. Models do relatively well on SQL injection, averaging 83%, and on cryptographic algorithms at 87%. They fall sharply on cross-site scripting at 15% and log injection at 12%. That spread is not random: some security problems are learnable as repeatable patterns, while others depend on dataflow, application context, and understanding how user input moves through a system - exactly the areas where raw model output is least dependable and verification needs to be strongest. On the language axis, Java shows the clearest improvement trend of any language and remains last by a wide margin, with a mean security pass rate of only 30%. Put those two findings together and the trade-off becomes concrete: gates should reserve human review for XSS, log injection, and Java-heavy changes instead of applying a uniform extra glance to everything.

# 3. Scan the diff an agent produced, before it becomes a pull request
def precommit_gate(diff):
    changed = languages_in(diff)
    findings = []
    for lang in changed:
        findings += sast.scan(diff, language=lang)
    # Java is the most improved language in the 2026 report and still the
    # lowest, with a mean security pass rate of only 30% - keep the gate open
    # for it rather than trusting generated patterns.
    blocking = [f for f in findings if f.severity in ("high", "critical")]
    return {"block": bool(blocking), "findings": findings}

# Roughly 44% of tested AI generation tasks produced code with a known
# vulnerability. A diff-scoped gate is the cheapest way to catch it upstream.

4. Four Things You Can Wire Up This Week

First, gate by weakness class rather than by a single score: require human review and a rescan for CWE-79 and CWE-117, and keep verification cheap for classes that already pass at high rates, such as SQL injection. Second, write model selection down as policy data instead of folklore - store the pass rates in configuration and record the residual risk you accepted, because even the leader needs verification. Third, scan the diff an agent produced rather than the whole repository, so findings surface before a pull request exists. Fourth, instrument remediation itself: which model wrote the code, which model fixed it, whether the same weakness returned within 30 days, and how many human edits the fix required. If AI-assisted fixes keep reintroducing the same CWE, the team is spending review capacity on churn rather than reducing risk.

# 4. Instrument remediation, because "AI fixed it" is not a measurement
REMEDIATION_FIELDS = (
    "finding_id", "cwe", "model_that_wrote_it", "model_that_fixed_it",
    "reintroduced_within_30d", "human_edits", "time_to_fix_minutes",
)

def weekly_report(rows):
    return {
        "opened": len([r for r in rows if r["state"] == "open"]),
        "ai_fixed": len([r for r in rows if r["model_that_fixed_it"]]),
        "ai_reintroduced_30d": len([r for r in rows if r["reintroduced_within_30d"]]),
        "median_human_edits": median([r["human_edits"] for r in rows]),
    }

# If AI-assisted fixes reintroduce the same CWE, review capacity is being
# spent on churn, not on risk reduction.
Review effort belongs where models fail most

Review effort belongs where models fail most

5. Place the Data in the Wider Risk Picture

The urgency of these numbers comes from their timing. Veracode's own analysis cites the 2026 Verizon Data Breach Investigations Report, which found software vulnerabilities are now the top breach entry point at 31%, surpassing stolen credentials, and Veracode's 2026 State of Software Security, which shows security debt affecting 82% of organisations, critical security debt affecting 60%, and high-risk vulnerabilities up 36% year over year. Stack those signals and the question changes shape: not whether AI introduces risk, but how to stop additional risky code from flowing faster into estates that already carry too much unresolved risk. That is why the recommendation is safe enablement rather than blanket restriction. Wrap the model in category-based gates instead of impression-based review, replace AI fixed it with a measurable remediation pipeline, and keep an audit record of every model-selection decision - including the residual risk you consciously accepted when you chose it.

{
  "governance_record": {
    "decision": "select_model_for_code_generation",
    "model": "gpt-5.5",
    "evidence": "2026 GenAI Code Security Report - strongest security pass rate in the Summer 2026 snapshot",
    "accepted_residual_risk": "still fails on nearly one in three security tasks",
    "compensating_controls": ["diff-scoped SAST gate", "CWE-79 and CWE-117 human review", "remediation instrumentation"],
    "reviewed_by": "appsec + platform",
    "review_due": "quarterly",
    "context": {
      "ai_authored_share_of_committed_code": "roughly half in adopting organisations",
      "software_vulnerabilities_as_top_breach_entry_point": "31% (Verizon 2026 DBIR, cited by Veracode)",
      "security_debt": "affects 82% of organisations; critical debt 60%; high-risk vulnerabilities up 36% year over year"
    }
  }
}

📌 Frequently Asked Questions

How many models did Veracode test?

Per the report page: more than 100 AI models tested over four years, with 11 new models evaluated across 80 tasks in the 2026 round. The average security pass rate is 56%.

Is 56% an improvement over the previous report?

Barely. Veracode's blog states the average security pass rate of 56% has barely changed from 55% in the first report - security performance has stayed flat while AI-generated code volume surged.

Which weakness classes are worst?

Cross-site scripting averages a 15% pass rate and log injection 12%, the weakest classes in the report. SQL injection (83%) and cryptographic algorithms (87%) perform comparatively well. XSS and log injection are where human review belongs.

Are coding-specialised models more secure?

No. In the report, models purpose-built for code average a 51% security pass rate while general-purpose models average 52%. Model size barely matters either (large 53%, medium and small 51%), though reasoning models do better at 56% versus 51%.

Is the best model good enough on its own?

No. Veracode notes that even the leader, GPT-5.5 at 68%, still fails on nearly one in three security tasks, and no model in the dataset generates secure code reliably enough to remove the need for verification. Better model choice reduces load; it does not eliminate risk.