Skip to content
Regulayer™Human Control for AI
Book a live demo

Learning Center · AI agent security

AI agent security, in plain words.

The terms security teams use for AI agents: what each one means, where it comes from, and the question it leaves open.

Prompt injectionOWASP LLM01:2025

What it meansInstructions hidden in content the model reads, such as an email, a web page, a ticket or a document, that steer it to do something its user never asked for. Direct injection comes from the user; indirect injection rides in on the content.

The question it leaves openOWASP says “it is unclear if there are fool-proof methods of prevention” (source). If the model can be steered, what stands between it and the action?

Excessive agencyOWASP LLM06:2025

What it meansAn AI system that can take damaging actions because it has more functionality, permissions or autonomy than it needs, whatever is causing it to malfunction.

The question it leaves openOWASP’s mitigation is human approval before high-impact actions (source). Whose approval, and is it still current when the action runs?

Least privilegealso “least agency”

What it meansGive each identity only the access its task requires. OWASP’s agentic guidance extends it to “least agency”: only the autonomy the task requires.

The question it leaves openAccess is granted to an account in advance. Is each action still inside what a person allows at the moment it happens?

Blast radius

What it meansHow far the damage spreads when something goes wrong. The May 2026 joint guidance asks organisations to “limit blast radius of agent failure scenarios”.

The question it leaves openAn agent that finds a credential can reach whatever that credential reaches. What limits it to what someone actually authorised?

Non-human identityNHI, agent identity

What it meansService accounts, API keys, tokens and now AI agents: identities that act without a person at the keyboard. NIST’s 2026 concept paper asks how agents should be identified and authorised.

The question it leaves openNIST asks: “How can an agent prove its authority to perform a specific action?” (source). An identity says who the agent is, not who allowed this action.

Runtime authorisationpolicy decision point

What it meansChecking each request as it happens, rather than once at sign-in. The joint guidance asks for identity and authorisation verified continuously at runtime, through a centralised policy decision point for each request.

The question it leaves openIs the check made outside the agent, so the agent cannot skip it, and does it leave a record someone else can verify?

Human in the loop

What it meansA person reviews or approves before the system acts. Most guidance asks for it on high-impact actions.

The question it leaves openA person on every step becomes the bottleneck, and approvals become rubber stamps. Human authority in the loop sets the authority once and checks every action against it.

Kill switchthe stop

What it meansA way to halt an AI system quickly. The EU AI Act asks for a way to interrupt a high-risk system so it comes to a halt in a safe state (Article 14).

The question it leaves openOnce stopped, what stops it starting again? Who pressed it, did they hold the authority, and what ran afterwards?

Revocation

What it meansWithdrawing access. Okta reports half of organisations can revoke agent access in minutes.

The question it leaves openWhat happens to the actions already in flight when the access goes?

Rogue agentOWASP ASI10

What it meansAn agent acting outside what anyone intended. It can drift on its own, be persuaded, follow a hidden instruction, or be asked by another agent.

The question it leaves openDiagnosing the cause takes time. Can the action be stopped without knowing why it was proposed?

Shadow AIunsanctioned agents

What it meansAI tools and agents in use without the security team’s knowledge or approval.

The question it leaves openFinding them is the first step. Once found, which of their actions does anyone have the authority to allow?

Confused deputy

What it meansA program with legitimate access is tricked into using it for someone who has none. Named in the May 2026 joint guidance on agentic AI.

The question it leaves openThe agent holds the access. Was this particular action one that a person with authority allowed?

Fail closed

What it meansWhen a check cannot complete, the action does not run. The opposite, fail open, lets it through.

The question it leaves openDoes the record then say stopped or unknown, never a success it cannot show?

Audit trail

What it meansA record of what a system did. Rules from 21 CFR Part 11 to NYDFS Part 500 to the EU AI Act ask for one.

The question it leaves openCan someone outside the company check it was not changed, without trusting the system that wrote it? Try the verifier.

Next

See the questions answered, on one agent.