CLOSEDQUORUM: The First Windows Implant That Lets Four AI Models Vote on Its Next Move

·11 min read·Evergreen Tools Team

On September 22, 2026, Cisco Talos published CLOSEDQUORUM, a Windows implant whose most counterintuitive property is not its encryption or its evasion. It is that after deployment the malware stops waiting for a human. It convenes four commercial large language models — DeepSeek, Qwen, Mistral, and Google Gemini — votes on the next attack action, and executes the winner. Talos says it is the first publicly documented Windows implant to hand its command-and-control decisions to models.

What Talos Actually Found

Start with the confirmed facts. CLOSEDQUORUM is a 64-bit, Go-based Windows implant discovered through CAIRN, a research toolkit Talos released the same day for hunting, classifying, and tracking AI-integrated malware. Its core mechanism is narrow by design: collapse one phase of the attack into a constrained multiple-choice question, let four models answer independently, take the majority, execute the winner. Talos states it has no confirmation of in-the-wild deployment, and artifacts lifted from the binary connected its developer to carding forum posts from 2025. The implant was previously tracked as BALZAK. Talos also framed the find as a template rather than a one-off. It presents CAIRN, the toolkit that surfaced the implant, as an open method for defenders to hunt the same class of malware, and it publishes the indicators so others can reproduce the analysis instead of waiting for the next disclosure.

// The shape of the trick: collapse a phase of the attack into a multiple-choice
// question, then let several independent models vote. LLMs are most reliable
// exactly where the output space is small.

type Action struct {
    Verb   string // "credential_dump" | "inject" | "persist"
    Target string
}

var candidates = []Action{
    {Verb: "credential_dump", Target: "lsass"},
    {Verb: "inject", Target: "explorer.exe"},
    {Verb: "persist", Target: "Run\\Software"},
}

// Ask N models the same constrained question. Majority wins. No operator needed.
func quorum(ask func(Action) bool, acts []Action) Action {
    tally := map[Action]int{}
    for _, a := range acts {
        if ask(a) {
            tally[a]++
        }
    }
    var best Action
    for a, n := range tally {
        if n > tally[best] || (n == tally[best] && a.Verb < best.Verb) {
            best = a
        }
    }
    return best
}
Malware driven by a vote

The first documented Windows implant to hand C2 decisions to a model vote

Why a Quorum, and Why It Matters

The story is not that three models agreed on something. It is where the design moves the attacker's effort. LLMs are most reliable exactly when the output space is small, and CLOSEDQUORUM reduces a phase of the attack to three routes: steal credentials, inject into a process, or establish persistence. That hands tactical reasoning to the models and removes the human operator from the loop entirely. Talos frames it precisely as effort displacement: an expanding portion of the attack chain can now run without operator involvement, so an attacker who loses connectivity, sleeps, or changes time zones no longer interrupts an active campaign. That is the real shift worth internalising. A quorum does not make any single model smarter or more capable; it removes one person from one decision at the exact moment the attacker would otherwise have to make a call, and that changes who has to be awake for a campaign to continue.

// The vote is only half the design. The other half is the transport: results are
// delivered through a chat webhook, so there is no attacker-owned C2 domain to
// blocklist. That is the part your egress policy has to catch instead.

type VoteRecord struct {
    Provider string // "deepseek" | "qwen" | "mistral" | "gemini"
    Chosen   string
    At       int64
}

func report(votes []VoteRecord, hook string) error {
    body := map[string]any{"content": votes, "username": "sync"}
    payload, _ := json.Marshal(body)
    // One POST to a chat endpoint is indistinguishable from normal API traffic
    // unless you are watching which processes are allowed to make it.
    return post(hook, "application/json", payload)
}

The Details Keeping It From Spreading (So Far)

There are real brakes here, and they matter. Each copy needs API keys and a real Discord webhook to function. In the test versions the keys are compiled in at build time; in the public version both are placeholders, so it can neither reach the models nor send anything out. Talos published SHA-256 hashes for six builds from the malware's development. One practical wrinkle: when The Hacker News checked on September 23, CAIRN's bundled rule file did not include a CLOSEDQUORUM rule, so analysts have to add Talos's rule themselves. The practical conclusion is that the implant is currently more warning than weapon. But the missing pieces are configuration, not invention, and that is a thin margin to build a detection strategy on: the distance between a research sample and a live campaign is a build flag and a key.

# You cannot blocklist your way out of a quorum C2. You can watch who is calling
# model providers, from where, and with whose key. Start with the outflow.

# 1) Which hosts in this fleet have ever resolved a model-provider endpoint?
for host in $(cat fleet.txt); do
  dig +short @$host api.deepseek.com api.mistral.ai generativelanguage.googleapis.com \
    | grep -q . && echo "MODEL EGRESS: $host"
done

# 2) Which local processes hold provider credentials they should not?
#    A desktop app should not carry an inference key compiled into it.
sudo lsof -nP | grep -Ei 'deepseek|mistral|generativelanguage' || echo "clean"
Watching outbound API calls

When C2 becomes a model call, egress is the new detection surface

Defending When the Attacker Is a Vote

Classic detection asks who owns the C2 domain. When C2 is a model call, that question dies. The new surface is egress: which process is calling a model provider, with whose key, from which host. A desktop application has no business carrying an inference key compiled into it, and a single process reaching three or more different model vendors inside one session deserves an alert. This is not a silver bullet, but it drags an "unblockable C2" back into the observable world. Add process-level egress logging and the control keeps working even if the attacker rotates to a fifth model vendor tomorrow. Behaviour, not identity, is the part of the picture that stays stable enough to detect.

// Turn the Talos guidance into a rule you can actually run. The point is not the
// hash of one sample; it is the shape: a process that fans out to several model
// vendors in one session is worth an alert.

type EgressEvent struct { PID int; Host string; At int64 }

func flagMultiVendor(window []EgressEvent, providers map[string]bool) bool {
    seen := map[string]bool{}
    for _, e := range window {
        if providers[e.Host] {
            seen[e.Host] = true
        }
    }
    // One process, three or more model vendors, no user-driven feature that needs it.
    return len(seen) >= 3
}

// Pair this with key hygiene: keys should be short-lived, scoped, and never
// compiled into a binary where a security tool can recover them at rest.

Four Things You Can Add to the Pipeline Today

First, inventory egress: which hosts resolve model-provider endpoints, and which processes hold inference keys they should not. Second, key hygiene: keys should be short-lived, scoped, and never compiled into a binary. Third, a detection rule: one process touching three or more model vendors in one session is worth an alert. Fourth, hash verification: Talos shipped SHA-256 hashes for six builds, so hash before you rule. Every one of these is doable today, and every one also catches the implant that ships zero hardcoded IPs. None of the four needs a new product or a new budget line, which is exactly why deferring them until after the next disclosure is hard to justify to anyone who reads this incident as a preview rather than an anecdote.

// Finally, verify what you were told. Talos published SHA-256 hashes for six
// builds from the malware's development, and its CAIRN rule file did not ship a
// CLOSEDQUORUM rule at first. Hash first, then rule.

import hashlib, pathlib

def sha256_file(p: str) -> str:
    h = hashlib.sha256()
    with open(p, "rb") as f:
        for chunk in iter(lambda: f.read(65536), b""):
            h.update(chunk)
    return h.hexdigest()

def match_known_hashes(path: str, known: set[str]) -> bool:
    return sha256_file(path) in known

# Do the same for every build artifact you keep. An implant that never ships a
# single hardcoded IP is exactly the kind that survives until you hash it.
Multi-model routing

Four vendors, one vote, zero human commands

The Signal Behind the Noise

Fit this into the trend and the direction is clear, but the distinction between fact and judgment matters. The fact is that Talos documented this sample, published hashes and IOCs, and explicitly said it could not confirm in-the-wild use. The judgment is that the pattern becomes normal. The sample is a proof of concept, but the technique it demonstrates — narrow the decision space, then outsource the vote — is reproducible by anyone with an API budget. What is worth changing today is not one hash. It is your egress assumption. The proof of concept is cheap to build and cheap to copy. The assumption it breaks — that a quiet network means a quiet attacker — is expensive to keep holding.

📌 Frequently Asked Questions

What is CLOSEDQUORUM?

It is a 64-bit Go-based Windows implant disclosed by Cisco Talos on September 22, 2026. After deployment it convenes four commercial LLMs (DeepSeek, Qwen, Mistral, Google Gemini) that vote on the next attack action, and the binary executes the winner without any human command. Talos says it is the first publicly documented Windows implant to delegate C2 decisions to models.

Did Talos confirm it was found in the wild?

No. Talos explicitly states it has no confirmation of in-the-wild deployment. Artifacts from the binary connected the developer to carding forum posts dating back to 2025, indicating a criminal background. It was previously tracked as BALZAK and renamed CLOSEDQUORUM.

Does it actually work, and what does it need?

Each copy needs API keys and a real Discord webhook. In the test versions the keys are compiled in at build time; in the public version both keys and webhook are placeholders, so it can neither reach the models nor exfiltrate data. That makes it a research sample rather than a plug-and-play weapon.

Why use a vote instead of one model?

Because LLMs are most reliable when the output space is small. A plurality vote across independent models reduces the impact of any single model failing or refusing, while moving tactical reasoning off the human operator so an expanding part of the attack chain can run unattended.

How should I detect this class of threat?

Start with egress rather than a domain blocklist: inventory which hosts resolve model-provider endpoints, which processes hold inference keys, and which process touches three or more model vendors in one session. Also verify the SHA-256 hashes Talos published, and remember CAIRN's rule file did not initially include a CLOSEDQUORUM rule.