Z.ai's ZCode Uploaded Repos Without Consent: Governing Agent Egress
💡 Tool Tip:Regex Tester, gitignore Generator, Base64 Encoder
In September 2026, developers discovered that Z.ai's desktop coding agent, ZCode, had been packaging entire local workspaces, encrypting them, and uploading them to cloud servers without users' consent. The whole repository went along for the ride, .git history included. Reuters reported on September 21 that Z.ai then disabled some of the assistant's features and publicly apologized. The engineering lesson is plain: when a feature like codebase indexing is on by default and the index is built in the cloud, your repository becomes a data-transfer event. And most teams have no idea what the developer tools they use every day are sending out.
Invisible egress is the hardest part to govern
1. What Happened: A Timeline
On September 18, a paying subscriber published packet captures showing that ZCode had packaged and uploaded a logged-in user's entire workspace, including .git history, LFS cache, reflogs, and application configs. According to that technical write-up, roughly 87% of the payload was .git itself, and it landed in object storage on Alibaba Cloud. Developers then raised two harder problems on X and RedNote. First, the Codebase Indexing feature was enabled by default and there was no toggle to turn it off. Second, the privacy policy had not acknowledged it in advance. On September 21, Z.ai said the issue originated from Codebase Indexing, patched the software vulnerability, disabled certain features, open-sourced ZCode with a transparency pledge, and enabled zero-data retention for developers and enterprises.
# 1) Default-deny egress for the agent process, then allowlist only
# the endpoints its job actually requires. MacOS/Linux sketch:
# block everything, permit the model API, deny cloud object storage.
nft add table inet agent
nft 'add chain inet agent output { type filter hook output priority 0; policy drop; }'
nft add rule inet agent output ip daddr 127.0.0.0/8 accept
nft add rule inet agent output tcp dport { 443 } accept # then tighten below
# A code indexer has no business reaching an object-storage bucket.
# If it can POST your repo to one, assume it eventually will.2. Why "Indexing" Means "Sending Everything"
To see why this was almost inevitable, look at how indexing works. For a model to answer questions about your codebase, it needs your code. If embedding and index construction happen in the cloud, the whole repository ships with them. And the whole repository is usually much larger than you picture: beyond source code there is full commit history, reflogs, LFS caches, untracked files, local configuration, and the sensitive files you thought .gitignore had handled but that are still tracked. The problem is therefore not that a vendor quietly sent a little extra. It is that cloud indexing, by design, requires a complete upload. Users' follow-up question cut to the same point: they found the data was encrypted with a backend private key held only by Z.ai, which meant they could neither open their own uploaded files nor independently confirm deletion.
// 2) Audit what an indexer WOULD send before you enable it.
// Enumerate the workspace minus excludes, and weigh .git separately.
import { execSync } from "node:child_process";
import { statSync } from "node:fs";
const EXCLUDES = [".env", "node_modules", "dist", "*.pem", "*.key"];
const files = execSync("git ls-files -co --exclude-standard", { encoding: "utf8" })
.split("\n")
.filter((f) => f && !EXCLUDES.some((e) => f.endsWith(e.replace("*", ""))));
let bytes = 0;
for (const f of files) bytes += statSync(f).size;
const gitBytes = Number(execSync("du -sb .git | cut -f1", { encoding: "utf8" }).trim());
console.log("tracked+untracked files:", files.length);
console.log("workspace MB:", (bytes / 1e6).toFixed(1));
console.log(".git MB:", (gitBytes / 1e6).toFixed(1), "- history travels too");3. The Aftermath: One Candour and One Retraction
There is more than one side to this. Chengming Technology claimed that six of its company coding workspaces had been uploaded by ZCode without consent, including complete source code, database passwords, and employees' personal information. On Monday the company retracted that statement, saying it had the wrong evidence. That same Monday, Z.ai said an independent security assessment by an IT standards think tank affiliated with China's industry ministry and the cybersecurity firm NSFOCUS found that users' code data had been deleted and was not retained by the cloud platform. Read side by side, the pair is a good reminder: in a security incident, an unverified accusation and an unverified clean bill both need their sources labeled, including the technical analysis cited here, which is analysis rather than an official conclusion. In the same period, China's cyber regulator released an updated AI safety framework warning about shutdown resistance, evaluator deception, and sandbox escape, and Z.ai had already delayed the release of GLM-5.3 for safety reasons, the first Chinese lab to do so explicitly.
# 3) Keep indexing local. Embed on your machine, ship nothing.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2") # runs locally, no API
def index_repo(paths):
index = {}
for path in paths:
text = open(path, "r", errors="ignore").read()
index[path] = model.encode(text[:8192]) # vectors stay on disk
return index
# Rule of thumb: if the feature is called "cloud indexing" and it is on
# by default, that is the first setting to turn off.4. Technical Governance: Four Defenses You Can Ship
Whatever the vendor says, your conclusion is the same: you have to govern egress yourself. The first defense is a default-deny egress policy. Give the agent process its own network rule set, allow only the endpoints its job actually needs, and explicitly forbid object storage. A code indexer has no business POSTing your repository to a bucket. The second is audit before you enable. Before switching on any indexing feature, enumerate the files that would actually be sent and weigh .git separately, so it becomes visible that history travels too. The third is local indexing: build vectors with a local embedding model on your own machine, so nothing leaves. The fourth is a secret pre-flight gate: run pattern scans before anything is embedded or uploaded, abort on suspected keys, and then verify that your .gitignore really excludes what you assume. A secret that was ever tracked is history, not configuration, so rotate it rather than deleting it.
# 4) Scan for secrets BEFORE anything gets embedded or uploaded.
# A short, boring gate catches most of the damage.
PATTERNS=('AKIA[0-9A-Z]{16}|ghp_[A-Za-z0-9]{36}|sk-[A-Za-z0-9]{20,}|BEGIN [A-Z ]*PRIVATE KEY')
if grep -RInE "$PATTERNS" --exclude-dir=.git --exclude-dir=node_modules . ; then
echo "secrets found - refusing to build any index" >&2
exit 1
fi
# Then verify your .gitignore actually excludes what you think it does.
# Tracked secrets are history, not configuration; rotate them anyway.5. In Practice: From Firewall to Audit Log
The first snippet is the egress rule: default deny, then allow only what is required. The second is an audit script that enumerates tracked and untracked files, sums their size, and prints .git's bytes separately, so the cost of history is explicit. The third is local indexing, using a local embedding model so vectors stay on disk. The fourth is the secret gate, scanning the repository for common key patterns and refusing to build any index when it hits. The fifth is an egress audit table recording the host, path, bytes sent, whether repo data was included, and the retention policy the vendor claims, with alerts on three events: a new host, a first upload containing repo data, and a retention claim you cannot independently verify. None of this is advanced. All of it converts invisible into visible, which is the precondition for governing anything.
// 5) Log every outbound call the agent makes, and diff it weekly.
// What you cannot see, you cannot govern.
function recordEgress(req, bytesOut, classification) {
return db.egress.insert({
ts: new Date().toISOString(),
host: new URL(req.url).host,
path: new URL(req.url).pathname,
method: req.method,
bytesOut,
containsRepoData: classification.repoData, // true = escalate
tool: process.env.AGENT_NAME,
retentionClaim: classification.retention, // "zero" | "unknown"
});
}
// Alert on: new host, first upload containing repo data, or a tool
// whose retention claim you cannot independently verify.6. A Checklist: Treat Agent Traffic as a Data Flow
Five checks. First, can you list every domain each coding agent, plugin, and MCP server you run talks to? Second, which of their features are on by default and upload data by default? Third, are your indexes and embeddings built locally or in the cloud? Fourth, do you have a secret gate sitting in front of indexing? Fifth, when a vendor says the data was deleted, is there any way for you to verify it independently? Of these, the second is the one most often skipped, because a default is a product decision, not a security promise. ZCode's core flaw was not a weak cipher. It was the combination of on by default, built in the cloud, and no toggle, which left users no choice to make. The reflex worth building in the agent era is this: whether a tool can read your code is one question, and where it sends your code is another. The second one is the one that deserves your signature.
Keep indexing local so data never has to travel
📌 Frequently Asked Questions
What exactly did ZCode upload?
According to packet captures published by developers and the technical write-up on the incident, ZCode packaged and uploaded a logged-in user's entire workspace, including .git history, LFS cache, reflogs, and application configs; roughly 87% of the payload was estimated to be .git itself, sent to Alibaba Cloud object storage.
How did Z.ai respond?
Per Reuters, Z.ai said the issue originated from the Codebase Indexing feature enabled by default, patched the vulnerability, disabled some features, and apologized publicly, while open-sourcing ZCode, enabling zero-data retention for developers and enterprises, and publishing an independent security assessment.
Why would codebase indexing send .git history too?
Because cloud indexing needs the full source to build embeddings and an index. A repository is not just source code: it includes commit history, reflogs, LFS caches, untracked files, and local configs. If the index is built in the cloud, one upload carries all of it.
How do I stop a coding agent from exfiltrating code?
Four layers: a default-deny egress policy for the agent process allowing only necessary endpoints, an audit of the files and bytes that would be sent before enabling any indexer, local embedding and indexing, and a secret pre-flight gate that runs before indexing begins.
A vendor says the data was deleted. How can I verify that?
Look for a verifiable evidence chain, such as an independent third-party assessment or auditable deletion proof, and reconsider the architecture: if the encryption key is held only by the vendor, users cannot independently verify content or deletion. The stronger position is to keep data off the wire in the first place.