Skip to content
Regulayer™Human Control for AI
Book a live demo

Research note

Why AI goes rogue, and why drift is rising.

AI goes rogue for more than one reason. It drifts. It gets talked into it. It follows an instruction hidden in a document. Another agent asks it. This note sets out what the published record shows about each, what it cannot yet show, and what follows for anyone putting AI agents to work.

Drift, counted

About one in five documented AI incidents is drift. And the share is rising.

1 in 5Documented AI and agent incidents where the system departed from its instructions with no attacker involved.
35%Of incidents where the AI itself took the action, drift was the cause.
6% to 32%Drift’s share of documented incidents, 2024 to 2026, as agents reached real systems.
72Distinct forms of drift named in Regulayer™’s drift taxonomy.

Incident figures: an independent coding of 130 AI Incident Database records, 2023 to 2026, commissioned by Regulayer™. Taxonomy: Regulayer™, including forms from Regulayer™’s own research.

The record is growing

362 documented AI incidents in 2025, up from 233 in 2024.

That is the count kept by the AI Incident Database, as reported in the Stanford AI Index 2026. Most early entries concerned what chatbots said. The newer entries increasingly concern what agents did: deleting, buying, entering systems.

1 · It drifts

No one attacks it. Over a conversation or a task, it moves away from what it was told.

  • In the lab
    In the SysBench study, GPT-4o followed its system instructions 84.8% of the time at the first turn of a dependent multi-turn conversation, and 33.7% by the fifth.
  • In production, as reported
    An AI development agent on Replit reportedly deleted a live production database during an active code freeze, despite repeated instructions not to make changes. Incident 1152
  • OpenAI’s Operator agent reportedly made a $31.43 purchase the user had not authorised, when asked only to compare prices. Incident 1028
  • Claude Cowork reportedly deleted a folder of 15 years of family photos when asked to clear temporary files. The files were later recovered. Incident 1441
  • An OpenAI research agent reportedly bypassed repeated access blocks and entered non-public areas of an Australian government statistics service. Incident 1707

2 · It gets talked into it

The conversation itself persuades it to act.

  • Attackers reportedly convinced Meta’s AI support bot to change account details and trigger recovery, leading to takeovers of high-profile Instagram accounts. Incident 1510
  • A car dealer’s chatbot reportedly agreed to sell a Chevrolet Tahoe for $1 after a user steered the conversation. Incident 622

3 · It follows an instruction hidden in what it reads

Planted in a document, a web page or an email, read as if it came from the user.

OWASP ranks prompt injection, including this indirect form, first among risks to applications built on language models (OWASP Top 10 for LLM Applications 2025, LLM01).

  • In the InjecAgent benchmark, a tool-using GPT-4 agent followed injected instructions about one time in four, and nearly twice as often when the attack added a hacking prompt.
  • A website reportedly published material designed to be picked up by chatbots, and testing found chatbots citing it. Incident 1659

4 · Another agent asks it

One agent’s request can carry another’s error into a real action.

As agents hand work to other agents, a request can arrive already shaped by drift, persuasion or a hidden instruction somewhere upstream. The public incident record here is still thin. Regulayer™ analysis.

No attacker required

Systems break their own instructions without anyone trying to make them.

  • In PrivacyLens, GPT-4 and Llama-3-70B agents leaked sensitive information in 25.68% and 38.69% of cases, even when prompted with privacy-enhancing instructions.
  • In SysBench, instruction-following fell with conversation length alone.

The newer production cases share a pattern: an agent with access to real systems, acting beyond what it was asked, with no adversary reported.

What the record cannot tell you yet

The cause, a rate, and the regulated record.

  • The cause, at the time it matters. Reports describe what happened. Why it happened is often unclear, disputed, or established only later.
  • A rate. Incident databases count incidents. None we reviewed counts how many actions AI systems took, so no one can yet state how often an agent acts outside its instructions per action taken.
  • Regulated work. We found no public report of an AI agent acting outside its authority in GxP-regulated manufacturing. That is an absence of reports, not evidence of safety.

What follows

By the time the cause is known, the action has already happened.

A control that waits to understand why is too late for the database, the payment or the photos.

Regulayer™ does not need to know why. Before the agent acts, it checks whether a named person still allows that action. If not, it stops it. Either way it keeps a record you can hand to an inspector.

Every action, whatever its cause, passes a named person’s current authority before it takes effect.

What regulators now ask for

Oversight, records, and incident reports.

The EU AI Act requires human oversight of high-risk AI systems (Article 14), automatic recording of events (Article 12), and reporting of serious incidents (Article 73).

Sources

Figures are quoted as their sources state them. Incident descriptions follow the AI Incident Database, which reports events as “reportedly” or “allegedly”; this note keeps that wording. Regulayer™ supplies the evidence. Your auditor, regulator or court makes the determination. Patent pending.