The reconnaissance loop: what it would take for the agent to run the hunt itself
The question
Which configuration would let an autonomous agent conduct the credential hunt on its own, instead of being used as a tool while something else sets the direction?
The thesis
Malware orchestrating trusted AI CLIs is the demonstrated step. The conditional step imagines an agent that receives its target list through content it reads during legitimate work. Every capability the second step needs already exists in enterprise development environments. What separates the two is where the credentials sit relative to the agent's reach, and how reachable its instructions are from untrusted input.
Contents
Claim ledger
Verified fact
- Malicious versions of the Nx build system package were published to npm beginning 26 August 2025 at approximately 22:32 UTC and remained available for just over five hours, during which thousands of developers may have been exposed. [stepsec]
- The malware harvested SSH keys, npm tokens and .gitconfig files, and StepSecurity describes the incident as the first documented case of malware weaponising AI CLI tools — including Claude, Gemini and Amazon Q — for reconnaissance and data exfiltration. [stepsec]
- The Nx maintainers published a security advisory confirming the compromise and stating that a maintainer's npm account was compromised through a token leak. [advisory]
- A second wave on 28 August 2025 used the leaked credentials to turn private organisation repositories public and to fork them. [stepsec]
Technical reading
- The AI CLIs contributed the part of the hunt that is hardest to script: locating credentials inside heterogeneous developer environments. Their filesystem reading is a product feature, and the malware used it as intended, on a target it chose.
Hypothesis
- An agent that reads untrusted content during ordinary work could receive a target list through that content, then conduct the same hunt without an external driver. This is a conditional construction. No such case is documented in this incident.
Open
- The published record does not establish how many of the harvested credentials were actually used for follow-on access, and it does not report whether agent-assisted discovery outperformed the scripted parts of the attack.
The starting state
Take an engineering organisation of ordinary size and ordinary habits. Developers install an AI coding assistant as a command-line tool. The tool runs with the developer’s own identity, reads the working tree, and can execute shell commands, because that is what makes it useful. Continuous integration runners execute build steps from installed dependencies, and those runners hold credentials so that pipelines can publish artefacts and comment on pull requests.
Now add the demonstrated event. On 26 August 2025, malicious versions of the Nx build system package
reached npm and stayed available for just over five hours. [1] Installing one of them
executed a payload that harvested SSH keys, npm tokens and .gitconfig files, and that payload
also invoked the AI assistants installed on the machine, using them to search the filesystem for
credentials and to push what it found to repositories the attacker controlled. [1] The Nx
maintainers confirmed the compromise and traced it to a leaked npm token belonging to a maintainer
account. [2] A second wave two days later used the harvested credentials to make private
organisation repositories public. [1]
That incident is the foundation. Everything that follows is built on top of it and labelled as construction.
1. Capabilities the scenario requires, and where each is demonstrated
| Capability | Status |
|---|---|
| Executing code from a public package during installation | Demonstrated in the incident |
| Harvesting credentials from a developer environment | Demonstrated in the incident |
| Invoking an installed AI assistant non-interactively | Demonstrated in the incident |
| Locating credentials inside a heterogeneous filesystem | Demonstrated in the incident |
| Exfiltrating to attacker-controlled infrastructure over ordinary protocols | Demonstrated in the incident |
| An agent receiving a target list through content it reads during normal work | Hypothesis, with published prompt-injection results as supporting evidence |
Five of six capabilities are documented as already exercised in the wild. The sixth sits between two literatures that rarely meet: prompt injection, where the delivery channel is well studied, and supply-chain credential theft, where the target is well understood. The conditional step combines them.
2. Access and permission conditions
The scenario needs five conditions to hold at once. A long-lived token with repository or registry scope must sit in a file the agent can read, with no consent asked each time it runs. The agent must run as the human, so that any action it takes carries the human’s authority and lands in the human’s audit trail. It must read files, issues, dependencies, documentation or tool output that an outside party can influence. The environment must allow outbound connections to arbitrary endpoints, including paste sites and public repository hosts. And something must run the agent with no human present: a build step, a hook, a scheduled job, a watcher. Remove any one of these and the conditional scenario loses its engine. That is the point of writing it down.
3. The causal mechanism
In the conditional scenario, a piece of content the agent meets during routine work contains an instruction. The agent processes it the way it processes any retrieved text, and the instruction sends it looking for credential-shaped material: environment files, key stores, cloud configuration, connection strings in adjacent repositories. Discovery is followed by collection, and collection by transmission to an endpoint the instruction named.
The cleverness of the instruction matters little. What matters is where the capability sits. The demonstrated incident showed that a filesystem-reading assistant is an effective credential search engine in environments nobody has catalogued. Combine that search capability with a delivery channel for the target list, and the attacker no longer has to write per-environment reconnaissance code, which is the most expensive part of a supply-chain campaign.
4. Failure points, ranked by how cheap they are to fix
The failure points reward attention in the order of what they cost to close. A token that expires in fifteen minutes keeps little of its value on a host that has already been compromised. An agent with its own identity, holding its own scoped claims, leaves a clean audit trail and denies the adversary the human’s reach. An agent granted the working tree rather than the home directory meets fewer credential files. A runner that can reach three hosts cannot publish to a repository the attacker selected. An assistant that requires an interactive session cannot be driven by a post-install hook.
5. Prevention and containment
The controls this atlas already carries map onto the failure points with unusual tidiness. Keeping agent authorisation off the agent itself means tool calls and privileged operations are mediated by a party the agent cannot impersonate. Treating agent-run actions as high-impact and gating them behind a human decision breaks the loop at the moment of collection. Registering and constraining which AI tools may run on developer and runner images reduces the installed search surface, and validating and logging agent inputs and outputs turns an unobserved conversation with untrusted content into a reviewable record. Each of those controls is catalogued, with its own scope and limits, in the control entries linked at the foot of this piece. Short-lived credentials issued to workloads, rather than tokens stored beside the code, remove the prize.
Containment deserves a separate thought, because the demonstrated incident compressed the response window badly. The malicious package was live for roughly five hours, and the follow-on abuse of the credentials began two days after that. [1] Organisations that cannot answer “which packages were installed in this window, on which host, and what credentials were present there” in hours cannot contain this class of event at all.
6. Uncertainties
The published record establishes the mechanism, and it leaves several quantities unmeasured. How many harvested credentials produced further access is not established publicly. Whether the assistant-assisted search outperformed the scripted components has not been reported, and the honest reading is that the assistants widened the search where scripting was inconvenient rather than replacing it. How long the credentials remained valid in practice depends on hygiene that varies per organisation, which is exactly the variable that determines the blast radius.
7. Evidence that limits plausibility
Two limits keep this scenario honest.
The first is the shape of the demonstrated abuse. The malware supplied the driving, the target and the exfiltration channel, and the assistants supplied file reading. A scenario in which the agent generates its own objectives out of a piece of untrusted content is a different construction, and its plausibility rests on prompt-injection results rather than on this incident.
The second is the difficulty of self-directed search. An agent that must decide which of ten thousand files look like credentials works against an unbounded problem, and error rates in autonomous multi-step tasks remain high. The conditional scenario becomes realistic in environments where the credentials sit in predictable places, which describes most organisations and none of the careful ones.
The Atlas verdict
Malware compromised a widely used build package and used installed AI coding assistants to find and exfiltrate developer credentials at scale, inside a window of about five hours.
- The reading we find strongest
- What is new here is the division of labour. The malware brought the objective and the plumbing, and the assistants brought the file-reading capability that a script would have had to write from scratch for every environment it met.
- What usually gets misread
- Coverage framed this as AI agents turning on their users. The agents were invoked as tools by an external process, and their behaviour followed from the access they already held.
- Why it matters for safety, science or philosophy
- Every developer workstation and CI runner with an ambient-credentialed agent installed has added a general-purpose search capability to whatever else runs there, including code that arrived five minutes ago in a dependency update.
- Evidence that would change this conclusion
- A configuration that separates agent identity from human identity, and that denies ambient credentials during package installation, would make the demonstrated step fail. The absence of such failures in audited environments is itself testable evidence.
Sources
- [1] StepSecurity, 's1ngularity: Popular Nx Build System Package Compromised with Data-Stealing Malware'
The write-up that documented the AI CLI abuse and the second wave, and the source for most of the technical detail here.
- [2] GitHub Security Advisory GHSA-cxm3-wv7p-598c (Nx compromise)
Maintainer advisory confirming the compromise and the npm account behind it. Published to the npm registry advisory database.
-
Vendor-tracked identifier for the tampered package, worth keeping for asset queries.