Your existing checklist is mostly right

Security teams asked to review an AI agent often assume they need an entirely new discipline. They do not. Most of what makes an agent safe is identity, least privilege, input validation, authorisation on sensitive actions, egress control and audit logging — all of which you already assess.

What is genuinely new is narrow. The agent decides its own actions at runtime, and it makes those decisions partly on the basis of text supplied by someone else. Everything unfamiliar follows from that.

Here is what to ask, in the order that saves the most time.

First: what identity does it act as?

Ask before anything else, because the answer tells you whether the rest of the review is worth starting.

If the agent uses a developer's credentials, a shared service account, or a long-lived key that other systems also use, stop there. Send it back. Nothing you assess afterwards is meaningful, because the agent's effective permissions are unbounded and unattributable.

A good answer is a dedicated identity, with short-lived credentials, used by nothing else.

Second: say its permissions in one sentence

Ask the team to state what the agent can reach without reading from a policy document. If they cannot, the scope is too broad to review, regardless of what the policy says.

Then probe the chain, because this is where the real finding usually is. Can the agent assume another identity? Can anything it writes to be executed by something with higher privileges? Does it hold a configuration permission that indirectly grants more than its direct permissions suggest?

A modest-looking write permission into a path that a pipeline reads and runs is an administrator permission wearing a disguise.

Third: which actions cannot be undone?

Have the team list every action the agent can take that deletes, pays, sends something outside the organisation, or otherwise cannot be reversed. Then ask what sits in front of each one.

Your position should be that irreversible actions require a human approval step or a hard constraint that a persuaded agent cannot talk its way past. This is not distrust of the technology. It is the ordinary asymmetry between the cost of a delay and the cost of an irreversible mistake.

Fourth: what does it read, and who can write to that?

This is the question most likely to be missed, and it is where the distinctive risk lives.

Map everything the agent consumes: documents, tickets, database fields, emails, retrieved web content, output from other systems. For each one, ask who can put text into it. If any of those sources can be influenced by someone outside your trust boundary — and a customer-submitted support ticket qualifies — then that source can carry instructions to your agent.

The team's answer should not be that they filter for malicious instructions. That approach fails, because natural language offers unlimited phrasings. The answer you want is that the agent's permissions are narrow enough that a successfully injected instruction cannot achieve anything significant, and that any consequential action has a check in front of it.

Fifth: how could data get out?

Enumerate the outbound paths, not just the obvious one. The response to the user is obvious. Logs, outbound network calls, writes into records other users can read, and content passed to a third-party service are the ones that get overlooked.

Ask whether outbound destinations are allow-listed. Open egress combined with read access to sensitive data is an exfiltration path, and it does not require a model to exploit.

Sixth: show me a log of one run

Do not accept a description of the logging. Ask to see an actual trace of one real run.

It should let you reconstruct what the agent did and why: the input and its source, what context was retrieved from where, which tools were called with which arguments, what each returned, and the final output — all correlated by one identifier.

If a trace cannot answer “why did it do that”, then after an incident nobody will be able to answer it either, and the first incident is when that matters most.

Seventh: who gets the alert, and who owns it?

Ask what conditions raise an alert, which human receives it, and whether that human is on a roster. Then ask for the runbook, and read it.

A system with no named owner cannot be operated, only tolerated. This question is reasonable to treat as blocking, and asking it early saves everyone a month.

Eighth: who has already attacked it?

Ask whether anyone has deliberately tried to break it, and ask for what they found.

A team that hands you a record of attempted attacks, what succeeded, what they changed and what now prevents it has done the work. A team that says the controls are in place has written intentions. The difference is most of your remaining risk.

What a reasonable approval looks like

You are not being asked to certify that an agent will never misbehave. It will occasionally produce something wrong, because that is the nature of the technology.

You are being asked whether the consequences are contained, visible and owned. Narrow identity, enumerable permissions, human approval on irreversible actions, allow-listed egress, a reconstructable trail, an alert that reaches a person, and evidence that someone already tried to break it.

That is an approvable system. Everything else is a conversation about the model, and the model was never the part that was going to hurt you.

tMinus1 Team
about  the  author

tMinus1 Team

Digital Agency

The tMinus1 team builds digital products and systems for startups and businesses across Australia and globally. Based in Sydney, tMinus1 specialises in UI/UX design, web development, mobile app development, and generative AI services.

Learn about our Editorial policy