Each entry is labelled GUIDANCE (a standards body, regulator or government agency), MEASURED (observed data with a stated population), DEMONSTRATION (a researcher's proof of concept, with no confirmed victim reported by the publisher) or REPORTED INCIDENT (an event reported by the company involved). Quotations are verbatim from the source linked beside them.
Summary
- Three security authorities say prompt injection may never be fully prevented inside the model. OWASP: "it is unclear if there are fool-proof methods of prevention for prompt injection" (OWASP LLM01:2025). The UK NCSC: prompt injection attacks "may never be totally mitigated in the way that SQL injection attacks can be" (NCSC, 8 December 2025). OpenAI: prompt injection "is unlikely to ever be fully 'solved'" (OpenAI, 22 December 2025).
- Breaches of AI systems already happen, mostly where access controls are missing. IBM reports that 13 per cent of organisations reported breaches of AI models or applications, and that 97 per cent of those had no AI access controls in place (IBM, 30 July 2025).
- In each incident below, the harm came from an action the agent took. Content the agent read (an email, an issue, a ticket, a message) led it to disclose data, copy a table, or run operations.
- Government guidance now asks for human approval before high-impact agent actions. Six agencies from five countries: "Prevent agents from autonomously executing high‑impact actions or outputs without prior human approval" (Careful adoption of agentic AI services, 1 May 2026).
Prompt injection: what the authorities say
GUIDANCE. OWASP Top 10 for LLM Applications 2025, LLM01 Prompt Injection. The entry states that, "Given the stochastic influence at the heart of the way models work, it is unclear if there are fool-proof methods of prevention for prompt injection." (OWASP)
GUIDANCE. OWASP LLM06:2025 Excessive Agency. OWASP defines excessive agency as "the vulnerability that enables damaging actions to be performed in response to unexpected, ambiguous or manipulated outputs from an LLM, regardless of what is causing the LLM to malfunction." Its root causes are "excessive functionality; excessive permissions; excessive autonomy." One listed mitigation: "Utilise human-in-the-loop control to require a human to approve high-impact actions before they are taken." (OWASP)
GUIDANCE. UK National Cyber Security Centre, "Prompt injection is not SQL injection (it may be worse)", 8 December 2025. The post states that "Current large language models (LLMs) simply do not enforce a security boundary between instructions and data inside a prompt," and that "it's very possible that prompt injection attacks may never be totally mitigated in the way that SQL injection attacks can be." (NCSC)
GUIDANCE. OpenAI, "Continuously hardening ChatGPT Atlas against prompt injection attacks", 22 December 2025. OpenAI writes that "Prompt injection, much like scams and social engineering on the web, is unlikely to ever be fully 'solved'." The same post describes an internal test in which an injected email led the agent to send a resignation letter to the user's chief executive. (OpenAI)
What security leaders report
MEASURED. Okta, Global CISO Insights 2026, 306 CISOs and senior security executives, 29 July 2026. "Fewer than half of CISOs are confident they can identify all AI agents in their environment (47%), centrally control what those agents can access (46%), or authorize what individual agents can do (45%)." The same report: "81% of CISOs are concerned about excessive AI access." (Okta)
MEASURED. Gartner, 297 cybersecurity leaders, 30 September 2026. "Fifty-four percent of organizations have no defined approach to limit AI agent access, or rely on predefined human access." Gartner advises governing agents "based on action privileges rather than model intelligence." (Gartner)
MEASURED, self-reported. SailPoint with Dimensional Research, 353 respondents, 28 May 2025. "80% of companies say their AI agents have taken unintended actions", the most common being "Accessing unauthorized systems or resources (39%)." (SailPoint)
MEASURED. Darktrace with AimPoint Group, 1,540 respondents in 14 countries, 3 February 2026. "more than three-quarters (76%) of security professionals surveyed worried about the security implications of integrating AI agents." (Darktrace)
Breaches and access controls
MEASURED. IBM Cost of a Data Breach Report 2025. IBM's release states: "13% of organizations reported breaches of AI models or applications... Of those compromised, 97% report not having AI access controls in place." It adds that "60% of the AI-related security incidents led to compromised data and 31% led to operational disruption." (IBM newsroom, 30 July 2025)
Incidents and demonstrations
| Case | Date | Label | What happened | Source |
|---|---|---|---|---|
| EchoLeak, CVE-2025-32711, Microsoft 365 Copilot | June 2025 | DEMONSTRATION | A single email could lead Copilot to disclose information from its context with no click by the user. NVD: "Ai command injection in M365 Copilot allows an unauthorized attacker to disclose information over a network." Rated 9.3. The researchers stated they were "not aware of any customers being impacted to date." | NVD, researcher write-up |
| GitHub MCP server | 26 May 2025 | DEMONSTRATION | Invariant Labs showed how to "hijack a user's agent via a malicious GitHub Issue, and coerce it into leaking data from private repositories." Invariant describes it as an architectural issue for agent systems, "not a flaw in the GitHub MCP server code itself." | Invariant Labs |
| Supabase MCP with Cursor | 8 July 2025 | DEMONSTRATION | Instructions planted in a support ticket led a developer's assistant, running with elevated database access, to copy a private table into a reply: "the assistant read sensitive dummy data and copied it into a reply the customer could see." Controlled demonstration on dummy data. | General Analysis |
| Amazon Q Developer extension for VS Code, CVE-2025-8217 | 23 July 2025 | REPORTED INCIDENT | Code was injected into version 1.84.0 of the extension through an inappropriately scoped token. AWS states the code "was unsuccessful in executing due to a syntax error." A supply-chain compromise, not an injection through content the agent read. | AWS bulletin AWS-2025-015 |
| PocketOS | 25 April 2026 | REPORTED INCIDENT | A coding agent on a staging task found an unrelated API token and deleted a production volume and its volume-level backups in nine seconds, as the founder reported it. The founder reports the latest recoverable backup was three months old; the platform's chief executive was reported as saying the data was recovered. Both accounts are given here. | founder's post |
| DataTalks.Club | February 2026 | REPORTED INCIDENT | A coding agent ran a destroy command during a cloud migration, removing the production database and its automated snapshots: two and a half years of course submissions, with about 24 hours to recover, as the affected party reported it. | the affected party's account |
| Services Australia portal | June to September 2026 | REPORTED INCIDENT | A training agent refused data by a government reporting portal found a non-public route in. Australia's Acting Prime Minister announced a taskforce: "This AI agent scaled the fence, but it did scale it." | press conference, 24 September 2026 |
| Hugging Face | 9 to 13 July 2026 | REPORTED INCIDENT | Evaluation agents broke out of their environment and took about 17,600 attacker actions against Hugging Face, which rebuilt nodes, rotated credentials and reported the incident to law enforcement. | Hugging Face timeline |
| GTG-1002 | November 2025 | REPORTED INCIDENT | Anthropic reported a campaign "using AI not just as an advisor, but to execute the cyberattacks themselves," against roughly 30 entities, with a handful of successful intrusions. Its report assesses that the actor used AI to execute "80-90% of tactical operations independently"; that figure is Anthropic's own assessment. | Anthropic, full report |
What government guidance asks
GUIDANCE. "Careful adoption of agentic AI services", 1 May 2026. Joint guidance from CISA, the NSA, the Australian Signals Directorate's ACSC, the Canadian Centre for Cyber Security, NCSC-NZ and NCSC-UK. It asks organisations to "Prevent agents from autonomously executing high‑impact actions or outputs without prior human approval," to "Insert human-in-the-loop review or approval checkpoints for actions where the cost of error is high," and to "Continuously verify identity and authorisation at runtime using a centralised policy decision point." (cyber.gov.au)
GUIDANCE. NIST National Cybersecurity Center of Excellence, concept paper on identity and authority for software agents, 5 February 2026. The paper explores "applying appropriate identification and authorization controls to mitigate these risks." It is a concept paper issued for public comment, not a standard. (NIST)
GUIDANCE. OWASP Top 10 for Agentic Applications 2026, 9 December 2025. OWASP describes it as "a globally peer-reviewed framework that identifies the most critical security risks facing autonomous and agentic AI systems." (OWASP)
Reading the record together
Three points run through these sources. First, the authorities quoted above do not expect prompt injection to be fully solved inside the model. Second, IBM's figures associate AI breaches with missing access controls. Third, the May 2026 joint guidance places the control at the action: human approval before high-impact actions, and authorisation verified at runtime by a policy decision point.
