Claude Code Deleted 48,000 Files in 103 Seconds. The Control That Failed Was the Filesystem, Not the Prompt
A developer reports that a Claude Code cleanup script followed 614 Windows junctions out of a test mirror and wiped 48,218 live files plus the Git object store. The claim is user reported and unverified, but the failure mode is real, and it tells you where agent blast radius has to be enforced.
On 21 September 2026 a developer posted a detailed account of a Claude Code session on Windows that went badly wrong. Asked to rebuild a test mirror, the agent wrote and ran a Python cleanup script. Between 10:10:31 and 10:12:14 in the evening, 103 seconds in total, the script removed 55,550 files. Only about 7,300 of them were inside the mirror it was meant to clear. The other 48,218 were live working files, and the Git objects, refs and logs directories were emptied too, so the repository could not be used to recover them. When the agent noticed, it told the developer to stop and read, and admitted it had broken something. TechRadar, Cybersecurity News and others picked the story up over the weekend.
Two caveats belong at the top. First, this is a user report with an attached verifier log, not an independent forensic investigation, and Anthropic has not published a response that we could find. Second, nobody has confirmed which permission mode the session ran in. Claude Code asks for approval before shell commands and file changes by default, and only skips those prompts when a user turns on a bypass mode. Treat the precise numbers as claimed, not proven. The mechanism, though, is well understood, reproducible, and not specific to Claude Code, Cursor, Devin or any other agent. That is why it is worth your time.
The mechanism is Windows directory junctions. The test mirror contained 614 of them, pointers that look like ordinary folders but lead somewhere else on the disk. The script walked the tree with os.walk and followlinks set to False, and it checked os.path.islink before deleting. On Windows, islink returns False for a junction, because a junction is not a symbolic link in the sense that check understands. Python added a separate os.path.isjunction function in version 3.12 precisely because of this gap. So the guard that looked correct in review did not fire, the walk went through the junctions into real project folders, and delete did what delete does. The agent did not misread an instruction. It wrote plausible code with a platform specific hole in it and ran it at machine speed.
That is the lesson for security teams. The instinct after an incident like this is to add a sentence to the system prompt, never delete outside the working directory. The agent already believed it was only deleting inside the working directory. A prompt cannot fix a wrong belief about what a path points to. The only controls that would have stopped this are ones the agent cannot reason its way around: an operating system account that has no write access outside the sandbox, a container or virtual machine with only the mirror mounted, a filesystem snapshot taken before any destructive step, or a rule that bulk deletes above a threshold stop and ask a human with the real count in front of them. Each of those works whether the code is right or wrong.
Speed changes the risk calculation as well. A human running the same flawed script might notice their editor tabs going blank and hit Control C after a few hundred files. 103 seconds for 48,000 files is roughly 470 files a second, which is faster than any person watching a terminal can react. When you assess coding agents, the question is not only what can this agent do, but how much can it do before anyone can intervene. That is blast radius, and it should be written into your AI risk assessment as a number, not a feeling.
The Git detail matters for backup design. Many developers treat the local repository as their backup, and it was the first thing to go, because .git lives inside the same tree the agent was allowed to touch. A backup that sits in the agent workspace is not a backup against the agent. Pushed remotes, off device snapshots, Time Machine or Windows File History to a separate volume, and cloud sync with version history all survive this scenario. ISO 27001 Annex A 8.13 on information backup and 8.14 on redundancy are the controls, and the test your auditor should ask about is simple: restore a developer workstation from backup after a simulated mass delete, and time it.
For the rest of the framework mapping, ISO 27001 Annex A 8.2 on privileged access and 8.18 on the use of privileged utility programs apply directly to an agent that can run arbitrary shell commands, and 8.31 on separating development, test and production environments is exactly the line the junctions crossed. ISO 42001 Annex A asks you to document the intended use, operation and monitoring of each AI system, and our ISO 42001 guide shows how to record permission modes and sandbox settings as evidence. For SOC 2, CC6.1 and CC6.3 cover least privilege and CC7.2 covers detecting anomalies, and a sudden spike in delete operations from an agent process is an anomaly your endpoint tooling can already alert on. If you track controls in Vanta, Drata or Secureframe, coding agents belong on the asset list alongside laptops, not in a footnote.
The practical list for this week. Find out which coding agents your developers run and which permission mode each uses, and ban bypass modes on machines holding anything you cannot restore. Run agents in a container, dev container or disposable virtual machine with only the project mounted, and never mount a parent folder that contains junctions or symlinks to real data. Make sure every repository is pushed to a remote before an agent session starts, and that workstations have versioned backups on a separate volume. Add a hook or wrapper that blocks recursive deletes above a set file count without a human confirming. Our policy templates include an AI acceptable use policy you can extend with these rules, and the compliance readiness checklist will show where endpoint backup evidence is missing.
Our view is that this incident will be argued about as a question of whether Claude Code is safe, and that is the wrong argument. Every agent that can run a shell will eventually write a script with a bug in it, just as every engineer does. The difference is that engineers work at human speed inside controls built for human mistakes. Agents need the same controls pushed down into the operating system, the filesystem and the backup layer, where a wrong belief about a path does not get a vote.
Editorial note: AES Tech reviews are independent. Some outbound links are affiliate links and are marked sponsored; they never change our rankings. See our disclosure.
Get the next post by email
One short email when something worth knowing ships. No spam, unsubscribe anytime.
Comments
Moderation policyLoading comments...
Add a comment
Corrections and first-hand experience are the most useful things you can leave. Comments are screened automatically and reviewed by a human; see the moderation policy.