A technical map of operational risk in generative AI systems.
The focus is not the model in isolation. It is the point where language connects to data, identity, tools, permissions, and decisions—and where useful capability can acquire operational authority.
Each entry turns that surface into something reviewable: a failure mode, a control boundary, or a governance question supported by evidence.
Maps data, tools, APIs, repositories, workflows, users, and decisions within reach of the system.
AI Capability
What can it do?
Identifies generation, retrieval, reasoning, transformation, tool invocation, recommendation, and action.
ATLAS reads the stack in both directions: from governance down to implementation, and from capability up to organizational risk.
Observation field
Enterprise copilots, RAG systems, AI Agents, and connected tools—especially when their output can affect data, access, or action.
Method
Start with a concrete failure mode, trace its exposure path, identify the control boundary, and ask what evidence would make the risk reviewable.
Coordinates
ATLAS links risks such as Shadow AI and Prompt Injection to practical controls, accountable owners, and governance decisions.
Featured research briefing
Shadow AI and the New Problem of Delegated Authority
A research briefing on the shift from unmanaged AI usage to agentic systems with delegated authority, and the controls needed to govern both risk domains.
A methodology for deriving a single canonical risk model from fifteen incompatible security and AI governance frameworks, so a scenario-assessment engine can reason over one fixed vocabulary instead of special-casing every source standard.
A new arXiv paper proposes AgenticAI-Supervisor, a simulation environment that decouples environment creation from scalable execution for verifiable agentic RL.
A synthesis of known failure modes in LLM-based agents, covering tool-use errors, planning breakdowns, and reasoning vulnerabilities that compound into systemic security risks.
External intelligence
Atlas News Radar
Recent external signals on AI security, agentic systems and governance.
arXiv:2607.05743v1 Announce Type: new Abstract: AI coding agents now read repositories, call tools, and execute shell commands with limited human oversight, and a fast-growing body of work studies whether the execution layer around them is actually safe. That literature is…
arXiv:2607.05775v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons. Reported benchmark gains often obscure recurring…
arXiv:2607.06008v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external environments. However, most existing benchmarks implicitly assume a monolingual…
arXiv:2607.05744v1 Announce Type: new Abstract: The Model Context Protocol (MCP) is the dominant way coding agents discover and invoke external tools. A server advertises each tool through a tools/list handshake that returns a name, a natural-language description, and a JSON…
arXiv:2607.06223v1 Announce Type: new Abstract: Reinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where the agent must make a sequence of intermediate decisions before receiving a final outcome.…
arXiv:2607.05773v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making. We introduce AgenticAI-Supervisor, an API and UI-driven RL Gym environment that decouples…
arXiv:2607.05804v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic…
arXiv:2607.05518v1 Announce Type: new Abstract: AI agents issue tool calls on the basis of text they cannot verify, so any party who controls part of the context can forge the appearance of authority. I evaluate 15 contemporary language models against eight attack scenarios…
arXiv:2607.05456v1 Announce Type: new Abstract: While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical deficiencies: (i) generated claims are not deterministically grounded in verifiable…
Researchers show how context manipulation can cause agentic browsers to abandon safety guardrails and exfiltrate sensitive credentials. The post ‘BioShocking’ Attack Tricks AI Browsers Into Stealing Credentials appeared first on SecurityWeek .
A database of almost a million passports from around the world was leaked online. Note what happened. A high-value credential—a passport—was used in an ancillary low-value authentication system: ID verification for cannabis dispensaries. And it’s the low-value system that got…
This is a fascinating explotation of how LLMs fall for prompt injection attacks. It turns out that they learn to recognize the style of text in different role/instruction blocks, and not just the tags. Their conclusion: Role tags were a formatting trick that became the security…
How Elastic's security team built an AI agent with RAG against MITRE's CWE and CAPEC catalogues to draft CVE advisories from raw vulnerability reports, including the full prompt and crawler configs.
llm appsec agents
Risk Catalogue
Failure modes that emerge when language crosses into data, identity, tools, permissions, and decision paths.
The trust channel between an MCP client and server can carry malicious content, intercepted requests, or manipulated trust dialog, allowing an attacker to control agent behavior through the communication channel.
An attacker-controlled artifact reaches an AI coding agent through plugin resolution, symlink paths, model cache, or CI trigger, and executes with the agent's privileges.
Classifying data before AI use defines what may be shared, transformed, summarized, retained, or logged.
A technical map, not a risk score
Use the catalogues to frame a review, then use Frameworks to connect findings to owners and evidence. Learning Log records how the map changes as the field develops.
Secure AI Atlas is built to make AI systems legible before capability becomes authority.