AI Writes Working Code 94% of the Time, but Only Half of It Is Secure
On September 29, 2026, Semgrep's security research team ran Claude Opus 5.5 through all 186 tasks of SusVibes. The benchmark takes real vulnerabilities from open-source Python projects: it deletes the code that implemented a feature, including the part where the original developer introduced the vulnerability, and asks an agent to write it again. The result: working code on 93.5 percent of tasks, but only 54.8 percent both working and secure. More tellingly, 53 percent of its solutions were near-identical to the project's real code, meaning a large share of the score came from recall rather than reasoning. Here is how the benchmark scores, what the numbers mean, and how security teams should change how they review AI-generated code.
1. How SusVibes Tests for Secure Code
Start with the scoring, because the strength of the conclusion lives there. SusVibes takes 186 real CVEs from open-source Python projects and turns each into a feature request: it deletes the code that implemented the feature, including the part where the original developer introduced the vulnerability, and asks an agent to write it again. The task description never mentions security beyond a one-line generic reminder. Each solution is scored twice. Functionally, the project's own test suite runs, and the solution passes if it breaks no more tests than the reference implementation. On security, the tests that shipped with the CVE fix run, and the solution passes if it is not vulnerable to the original bug. The final metric counts a task only when the solution is correct and secure, so a solution that works but reintroduces the vulnerability scores zero, as code sample 1 expresses.
# SusVibes takes a real CVE and turns it into a feature request: delete the
# code that implemented the feature (including the vulnerable line), then ask
# the agent to write it again. A task only counts when the result is BOTH
# functionally correct AND secure -- secure-only or working-only scores zero.
def score_task(patch, project):
functional = project.test_suite_passes(patch) # break no more tests
secure = project.cve_fix_tests_pass(patch) # not vulnerable
return int(functional and secure) # the only metric
# Semgrep ran Claude Opus 5.5 with SWE-agent 1.1.0, the canonical prompt,
# the generic one-line security reminder, and a 200-call limit per task.
# Result: working on 93.5% of tasks, working AND secure on 54.8%.186 tasks from real vulnerabilities
2. The Score: 93.5 Percent Working, 54.8 Percent Secure
Read the numbers plainly. Semgrep reports that Opus 5.5 produced working code on 93.5 percent of the 186 tasks, but only 54.8 percent were both working and secure. Of the 174 solutions that worked, 72 still contained the vulnerability the benchmark was built around, which is 41 percent of working solutions shipping the original bug. For context, 30 standard submissions already sit on the public SusVibes v1.0 leaderboard, and the best correct-and-secure score among them is 43.5 percent, achieved by GPT-5.5 with mini-swe-agent, while Claude Opus 4.8 under SWE-agent scored 19.4 percent. At 54.8 percent, Opus 5.5 would place first by roughly 11 points, a large jump over its predecessors, which is exactly why the next section matters. Code sample 2 shows the same numbers as a structured report.
# The headline is a security number split in two. Of 174 solutions that
# worked, 72 still contained the very vulnerability the benchmark was built
# around. That is 41% of working code shipping the original bug -- which no
# amount of green unit tests would catch.
from dataclasses import dataclass
@dataclass
class Run:
tasks: int = 186
worked: int = 174
correct_and_secure: int = 102 # 54.8%
vulnerable_but_working: int = 72 # 41% of worked
def report(r: Run) -> dict:
return {
"functional": round(r.worked / r.tasks, 3), # 0.935
"secure": round(r.correct_and_secure / r.tasks, 3), # 0.548
"bug_shipped": round(r.vulnerable_but_working / r.worked, 3),# 0.41
}3. The Catch: Over Half of the Solutions Were Likely From Memory
This is the crux. SusVibes tasks come from public projects, and the fixes for those CVEs have been public for months or years, so a model trained on public code may have seen the vulnerable version, the fixed version, or both. So Semgrep measured how closely each solution matched the project's real implementation, comparing only the lines each patch adds to non-test files, with whitespace and comments ignored. The finding: 53 percent of Opus's solutions, 93 of the 176 it could judge, were identical or near-identical to the real code, and in the most extreme cases the match was essentially character for character. On one pysaml2 task, Opus reproduced 54 lines of signature-handling code with a similarity of 1.00. On a Django task, it wrote 145 lines of django/utils/http.py in a single step, 99 percent matching the original. It had no network, no repo history and no installed copy; it wrote the code back from memory. By Semgrep's measure, 57 of Opus's 102 secure solutions, or 56 percent, were memorised. Code sample 3 sketches the similarity check.
# Half the score may be memory. SusVibes tasks come from public projects
# whose fixes have been public for months or years, and a model trained on
# public code may have seen the fix. Semgrep measured similarity between each
# added patch line and the project's real implementation, ignoring whitespace
# and comments. 53% of solutions were identical or near-identical.
def similarity(patch_added_lines, reference_lines, threshold=0.80):
score = jaccard_or_diff_ratio(patch_added_lines, reference_lines)
return {"score": round(score, 2), "memorised": score >= threshold}
# 57 of Opus's 102 secure solutions (56%) were memorised, and 36 of those
# reproduced the CVE fix's own lines. On one pysaml2 task the match was 1.00
# across 54 lines; on a Django task, 145 lines at 99% in a single step.
# No network, no repo history -- it wrote the answer back from memory.41% of working solutions still carried the bug
4. What This Means for Security Teams
Bring it down to actions. Semgrep offers five pieces of guidance worth taking one at a time. First, treat passing tests and secure code as separate questions: in this run 41 percent of working solutions still carried the original vulnerability, and a functional review with unit tests will not catch them, because Opus is very good at producing working code the first time. Second, review cryptography, concurrency and access control by hand, because those were the classes where Opus produced no secure solutions at all and where the fix depends on knowledge outside the code in front of it. Third, be most careful with large, cross-cutting security changes: one- to four-line fixes were handled well, while fixes over 50 lines were mostly missed unless remembered. Fourth, do not read a public-CVE benchmark score as a measure of secure-coding ability, and expect those scores to inflate. Fifth, do not count on the model talking about security, since the analysis found discussion and correctness to be unrelated. Code sample 4 turns these into a review router.
# The practical takeaway for a security team pulling AI code into review:
# rank by risk class, because the classes where Opus produced NO secure
# solutions are the ones a human must read. Semgrep also found that whether
# the model DISCUSSED security had nothing to do with getting it right.
HUMAN_REVIEW = {
"cryptography": "always", # model produced no secure solutions
"concurrency": "always",
"access_control": "always", # depends on context outside the file
"large_change": "if > 50 lines",# big fixes were missed unless remembered
"small_fix": "scan + spot", # 1-4 line fixes handled well
}
def route(patch) -> list:
flags = []
if patch.touches_any(HUMAN_REVIEW) and patch.size > 50:
flags.append("human_review")
if patch.size <= 4:
flags.append("sast_scan")
return flags or ["sast_scan", "human_review"]5. The Evaluation Details and the Doubts to Keep
Always read the protocol before the score. Semgrep ran Opus 5.5 under SWE-agent 1.1.0 using the leaderboard's standard protocol: the canonical SusVibes prompt, the generic security reminder, and a 200-call limit per task. It also states its caveats. Its runs used the current SusVibes prompt, which adds an anti-cheating block earlier submissions did not have, and it had to fix harness problems that could plausibly have depressed earlier scores too. Slicing by when each CVE was fixed, memorisation ran at 58 percent for 2014 to 2019, 60 percent for 2020 to 2021, 37 percent for 2022, and 57 percent for 2023 to 2024, with near-verbatim copies appearing across eras, so this is not a problem confined to old vulnerabilities. Code sample 5 collects the caveats.
# Two numbers to keep in mind before trusting any public-CVE benchmark.
# First, memorisation rises as training data grows, so scores inflate over
# time for reasons that have nothing to do with security skill. Second, a
# model that recalls a public fix has no such advantage on YOUR codebase.
BENCHMARK_CAVEATS = [
"fixed task set, public repos, public fixes -> recall inflates over time",
"check how closely a model's patch matches the reference before ranking",
"a vendor harness bug can depress or inflate every model on the board",
"SWE-agent 1.1.0 + 200-call cap + generic security reminder: know the setup",
"secure-coding ability on unseen code is a different claim entirely",
]
def inflate(check_history):
# If a model scores better on OLDER CVEs than newer ones, suspect memory.
return {year: stats for year, stats in check_history.items()}A passing test is not a security review
6. Wiring This Into Your Pipeline
Finally, the engineering. First, route AI-generated code through static analysis by default; Semgrep itself ships a product line for scanning AI-written code in real time. Second, keep functional and security acceptance separate: let tests and review cover behaviour while SAST, dependency and secret scanning cover safety. Third, route review by risk class, forcing human review for cryptography, concurrency, access control and large cross-cutting changes, and letting scanning backstop the one-to-four-line fixes. Fourth, stay clear-eyed about the memory dividend: a fix a model recalls from a public benchmark gives it no such advantage on your own codebase, so someone else's leaderboard score is not a substitute for your own evaluation set. Put together, the productivity gain from AI coding becomes a safe productivity gain rather than technical debt quietly rewritten as vulnerability debt.
📌 Frequently Asked Questions
What is SusVibes?
A benchmark built on real vulnerabilities: it takes 186 real CVEs from open-source Python projects, deletes the code implementing the feature including the vulnerable part, asks an agent to reimplement it, and scores each solution twice, on functional correctness and on security. A task counts only when it is both correct and secure.
How did Claude Opus 5.5 do?
Semgrep reports 93.5 percent functional correctness but only 54.8 percent correct-and-secure. Of 174 solutions that worked, 72 still contained the original vulnerability: 41 percent of working code shipping the bug. That score would place first on the public SusVibes leaderboard, roughly 11 points above the previous best of 43.5 percent.
Why might memorisation inflate the score?
The tasks come from public repositories and their fixes have been public for months or years, so a model trained on public code may have seen the answers. Semgrep measured 53 percent of solutions as identical or near-identical to the real code, and 57 of the 102 secure solutions, or 56 percent, as memorised, with 36 reproducing the CVE fix's own lines.
What does this mean for teams using AI coding agents?
Semgrep's guidance: treat passing tests and secure code as separate questions; review cryptography, concurrency and access control by hand; be careful with large cross-cutting changes over 50 lines; do not read a public-CVE benchmark score as a measure of secure-coding ability; and do not count on the model talking about security as evidence it got security right.
Does the model's security talk help?
Semgrep counted security terms such as vulnerability, sanitise, injection and traversal in Opus's visible reasoning. The third of solutions with the fewest mentions were secure 57 percent of the time and the third with the most were secure 53 percent, so discussing security and getting it right were unrelated.