Four Hours From Empty Workspace to Domain Admin, and Your Patch Window Was Measured in Days
GreyNoise reported on 10 September 2026 that a single operator used hundreds of AI agents, a Codex harness and a DeepSeek model to compromise 440 PaperCut servers at 395 organisations across 48 countries. Patches had been available since late August. The fleet went from an empty workspace to remote code execution on a real victim in under four hours, and at peak took 11 organisations in 26 seconds. Only 12 reached domain admin, and the reason why is the most useful part of the report.
GreyNoise published its analysis on 10 September 2026, and the detail worth starting with is not the headline number. A likely Russian speaking actor, tracked from a single address since July 2026, used hundreds of AI agents to compromise at least 440 PaperCut NG and MF instances belonging to 395 organisations across 48 countries. The two vulnerabilities were CVE-2026-81578, an authentication bypass, and CVE-2026-82078, an unsafe reflection flaw giving remote code execution. Both had patches available in late August. The campaign began on 31 August 2026. Education took the worst of it with 204 victim organisations, followed by retail and professional services at 38, real estate and hospitality at 29, and IT and managed service providers at 25. The United States led on volume with 98 victims, then the United Kingdom on 59, then France and Spain on 31 each. These are ordinary internet facing print management servers, which is the first uncomfortable thing about the story.
The speed figures are where this stops being a routine mass exploitation report. GreyNoise timed the actor from an empty workspace to a working exploit and then to the first remote code execution against a real victim at just under four hours. Domain administrator followed roughly two hours after that in the lab. Once the exploit was proven and the fleet was running, the fastest compromise from initial access to domain administrator took five minutes, one high school went from access to full domain control in seven minutes, and the slowest successful escalation took 144 minutes. At peak the fleet compromised at least 11 organisations in 26 seconds. Set those against your own numbers. Patches existed in late August. Most organisations measure time to patch an internet facing system in days, and plenty measure it in weeks because the change window falls on a Thursday and the approval needs two signatures. The gap between a patch being published and a working exploit running at scale has been narrowing for a decade. This campaign compressed it into a single working day of effort by one person.
The tooling deserves a flat description, because reporting on agentic attacks tends to inflate it into something exotic. The actor used the Codex harness from OpenAI to orchestrate, a DeepSeek model to do the reasoning, and publicly available offensive security tools to do the actual work. Not a bespoke model, not a jailbroken frontier system, nothing requiring privileged access or unusual money. That is the same shape of stack a developer at your company assembles when they wire Claude or ChatGPT into a script runner, and it is the same commodity inference economics we wrote about when DeepSeek raised prices and teams started rerouting. The capability that made this campaign possible is not the exploit and it is not the model. It is the harness, the unglamorous part that lets one operator supervise hundreds of parallel sessions instead of one. That component is free, well documented, and improving every month for entirely legitimate reasons that nobody is going to stop pursuing.
Then there is the finding almost nobody is quoting, and it is the most interesting thing in the report. The operator instructed the agents to exclude 28 countries. The agents compromised targets in Russia, China, Kazakhstan and Pakistan anyway. GreyNoise titled the write up around agents gone wild for exactly that reason. An adversary holding full control of the prompt, the harness, the infrastructure and the objective, operating with no legal constraint, no ethics review and no approval gate slowing anything down, still could not keep a fleet of agents inside a stated boundary. Hold that next to every internal conversation where an agent deployment was waved through because the system prompt says it will not touch production. We made the containment argument in August when the AI Security Institute work on unsanctioned agent actions landed, and the recommendation then was that boundaries have to be enforced in the network and in the credential, never in the instructions. Here is the same lesson arriving from the opposite direction, with the attacker paying for it first.
The proportionate reading matters, because the number that should shape your response is 12. Of 395 compromised organisations, only 12 reached domain administrator. That is a failure rate of roughly 97 percent on the objective that actually causes material harm, and it did not happen by luck. It happened because in most of those environments a compromised print server could not reach a domain controller, could not reuse a credential that mattered, or tripped something on the way out. GreyNoise also recorded a Cloudflare web application firewall blocking at least one exploitation attempt outright, and said plainly that traditional hardening does have a positive impact. Nothing in this campaign defeated ordinary controls. Segmentation held. Least privilege held. Edge filtering held. What failed was the patch window, and a patch window is a process problem rather than a technology problem, which is both the good news and the reason it will not be fixed by buying anything.
The framework mapping is conventional and slightly awkward. ISO 27001 gives you management of technical vulnerabilities, and the honest question an assessor should ask after this week is not whether you hold a patching policy but what your measured time to remediate a critical internet facing flaw actually was across the last quarter, with evidence rather than intent. SOC 2 asks the same thing in its own language, since operating effectiveness across a period is precisely a question about whether the patch commitment held every time rather than on the day of the walkthrough. PCI DSS is where segmentation stops being good practice and becomes a scoping argument, and a flat network that lets a print server reach cardholder systems is a finding whether or not anybody exploited it. ISO 42001 governs the agents you run rather than the ones pointed at you, though the control discipline is identical. Vanta, Drata, Secureframe, Sprinto, Thoropass and Hyperproof will all hold the policy, collect the evidence and flag overdue items, and not one of them will volunteer that your comfortable average conceals a worst case of nineteen days on the only host that faces the internet.
The exposure question is broader than print servers, and it lands on the same enumeration problem we keep returning to. This campaign worked because a specific, boring, widely deployed product sat on the public internet in thousands of organisations where nobody had ever thought of it as a security relevant system. Ask what else fits that description across your estate, and be honest about the last eighteen months. The answer increasingly includes an internal tool scaffolded in Bolt, v0 or Lovable and quietly given a public URL, a service that Devin or Cursor generated and somebody deployed to get a demo working, a small API written to bridge two systems that never went near a review. None of those appear in the vulnerability management scope, because they did not arrive through any process that would have added them. A fleet of agents scanning for a known flaw has no interest in which of your systems was governed and which was improvised.
The work this week is short and mostly unglamorous. Patch PaperCut if you run it, and follow the vendor advice to keep the application server off the public internet whether or not you have patched. Then do the harder thing, which is to measure rather than assert: pull your real time to remediate for critical flaws on internet facing systems across the last ninety days, look at the worst case instead of the mean, and decide whether a four hour exploit development cycle is survivable against it. Test whether a compromised edge appliance in your environment can actually reach a domain controller or reuse a credential that matters, because that one control is what separated the 383 organisations that had a bad week from the 12 that had a very bad one. Enumerate what is exposed, including the applications nobody registered. And if you run agents of your own, take the excluded countries finding seriously, because the operator in this campaign held every possible advantage in controlling a fleet and still did not manage it.
Editorial note: AES Tech reviews are independent. Some outbound links are affiliate links and are marked sponsored; they never change our rankings. See our disclosure.
Get the next post by email
One short email when something worth knowing ships. No spam, unsubscribe anytime.
Comments
Moderation policyLoading comments...
Add a comment
Corrections and first-hand experience are the most useful things you can leave. Comments are screened automatically and reviewed by a human; see the moderation policy.