Giving an AI agent its own identity: permissions, guardrails and an audit trail
An agent is a production service that improvises
Every other service in your estate does a fixed set of things in a fixed order. An agent decides what to do at runtime, partly on the basis of text it was handed by something or someone else. That one difference is the whole security problem.
It means the usual question — is this code correct — is not sufficient. The new question is what happens when the agent is persuaded to do something reasonable-looking that nobody intended. Four things contain that: its own identity, a scope you can enumerate, guardrails on the actions themselves, and a trail that lets you reconstruct any decision.
Its own identity, not a shared one
Start here, because nothing else works without it. The agent must authenticate as itself, with credentials belonging to no human and no other service.
This matters for three reasons. Blast radius: when the agent is abused, the damage is bounded by what the agent could do rather than by what a developer could do. Attribution: every action in your logs is unambiguously the agent's, so an investigation does not begin with working out who did what. Revocation: you can disable the agent in one action, without locking out a person or breaking four other systems that shared the key.
Prefer short-lived credentials issued at runtime over a stored key. A key in a configuration file is a permanent liability that outlives the project.
A scope you can read aloud
The test for whether permissions are tight enough is whether you can state them in a sentence to someone who has not seen the system.
“It can read these three tables, write to this one queue, and call this one internal service.” That is a scope. “It has read access to the data warehouse” is not, because nobody in the room knows what that includes, including the person who said it.
Two habits help. Enumerate resources explicitly rather than by wildcard, even though it is more work and more maintenance, because the maintenance is the point — adding a resource becomes a decision somebody makes rather than something that happens silently. And separate reads from writes into distinct capabilities, so the agent's ability to look at something does not imply an ability to change it.
Treat every tool as a trust boundary
The agent decides which tool to call and with what arguments, influenced by text it did not author. So each tool must validate its own inputs as though they came from the public internet, because in effect they did.
A tool that takes a query string and interpolates it into a database call is an injection waiting to happen, and the fact that an agent sits in front of it rather than a web form changes nothing about the defence required.
Guardrails on actions, not only on text
Input and output filtering is worth having and it is not sufficient, because it tries to constrain language, which is an endless game. Constrain actions instead, where the rules are finite and checkable.
Useful patterns, roughly in order of value:
Irreversibility needs a human. Anything that deletes, pays, sends externally or cannot be undone goes through an approval step. Not because the agent is untrustworthy, but because the cost of being wrong is asymmetric.
Rate and volume limits per task. An agent that normally reads six records and suddenly reads sixty thousand is either broken or being used against you. Either way you want it stopped, and a limit stops it faster than an alert does.
Allow-list the destinations. If the agent can make outbound calls, enumerate where to. Open egress is the standard route for turning data access into data exfiltration.
Stay inside the task. Check that the action the agent chose belongs to the job it was given. A support agent asked about a refund has no business calling the deployment pipeline, and that check is cheap to write.
A trail that answers “why did it do that?”
Logging that records only outcomes is not enough. When an agent does something surprising, the investigation needs the reasoning path, not just the result.
Record, for every run: the input and where it came from, what context was retrieved and from which source, which tools were selected with which arguments, what each returned, and the final output. Keep them correlated by a single run identifier so one query reconstructs the whole episode.
This is more volume than a typical service log, so set retention deliberately and be careful about what sensitive content you are persisting in the process. A log that becomes its own data-protection problem has not helped you.
Then try to break it
Controls that have not been tested are intentions. Before the agent goes anywhere near a security review, attack it deliberately.
Try instructing it through the data it reads rather than through its prompt. Try persuading it to use a permission it holds for a purpose it was not given. Try to get it to reveal its instructions or its context. Try to steer it off task entirely. Try to make a tool do something the tool's author did not anticipate.
Keep a record of what worked, what you changed, and what now prevents it. That record is the most persuasive document you can take into an approval conversation, because it shows the controls holding against real attempts rather than asserting that they would.
The honest summary
None of this is exotic. It is identity, least privilege, input validation, authorisation on sensitive actions, egress control, audit logging and testing — the same disciplines any production service needs.
What is different is that the agent chooses its own actions at runtime, which means the controls have to sit around the actions rather than inside the logic. Teams that treat an agent as a production service with an unusual control surface ship them. Teams that treat it as a model with an API spend six months in review.
Ready to Unlock the Full Power of AWS?
Let’s talk about your cloud goals — no pressure, no hard sells. We’ll audit your setup, suggest improvements, and help you scale smarter.





