The day people remember is the day they break it

We have run this week enough times to notice something consistent. Engineers are politely engaged while they are building. They are fully awake the moment they successfully attack the thing they just built.

Something changes when a developer types a sentence into a document, watches an agent they wrote obediently follow that sentence instead of its instructions, and realises nobody told them this was possible. That realisation does more for the quality of their next system than any amount of guidance about best practice.

So we put the attack day in the middle of the week rather than at the end, and we have your engineers run the attacks themselves.

Why reading about it does not transfer

Prompt injection is easy to explain and almost impossible to take seriously from a description. It sounds like a curiosity. Of course the model follows instructions in its input, that is what a model does, and surely you just filter for that.

Then an engineer tries to filter for it and discovers within twenty minutes that they cannot, because there is no finite list of ways to phrase an instruction in a natural language. That discovery is the one that redirects their design from filtering text to constraining actions, and it is a discovery that has to be made by hand.

The same is true of the others. Permission escalation sounds like something your cloud team has handled. Then an engineer notices that the harmless-looking write permission their agent holds feeds a pipeline that executes with higher privileges, and the shape of the problem becomes visible.

The five attacks worth running first

You do not need a research agenda. Five attacks cover most of what matters, and all five can be run by an ordinary engineering team against their own work in an afternoon.

Instructions through the data

Put an instruction into something the agent reads rather than into the prompt — a document, a ticket description, a database field, a web page it retrieves. The lesson is that any content the agent consumes is untrusted input, including content from your own systems, because your own systems contain text that users wrote.

Using a permission for the wrong purpose

Take a permission the agent legitimately holds and get it to use that permission for something outside its task. Nothing is exceeded. The agent is simply persuaded to apply a capability to a different end. The lesson is that permissions must be scoped to the task, not to the system, and that someone should check whether a chosen action belongs to the job that was given.

Getting data out

The agent can read something sensitive. Can you get that something to leave — into its output, into a log, into an outbound call, into a record a different user can read? Teams are usually confident about the first path and have not thought about the other three. The lesson is that read access plus any outbound channel equals exfiltration, so egress needs an allow-list.

Extracting the instructions

Get the agent to reveal its system prompt, its tool definitions, or the contents of its context. This matters less than people assume for secrecy and more than people assume for reconnaissance, because it tells an attacker exactly what to aim at. The lesson is that the prompt is not a security boundary and must never contain a credential.

Breaking a tool directly

Forget the model. Call the agent's tools with hostile arguments. Most tools written during a pilot assume their caller is friendly, because during the pilot it was. The lesson is that every tool validates its own inputs as though they arrived from the public internet.

How to run the session so it works

A few things make the difference between a genuine exercise and a demonstration people watch.

They attack their own code. Attacking a deliberately vulnerable sample teaches that samples are vulnerable. Attacking the thing they wrote on Tuesday teaches that they write vulnerable things, which is the transferable lesson.

Nothing counts until it is written down. Each successful attack gets recorded: what was tried, what happened, what the fix is, and how you would know if it happened in production. That document becomes your team's actual standard, and it is more likely to be followed than any policy, because the team wrote it.

Pairs, competitively, with a time limit. Two engineers per pair, forty-five minutes per attack, and the pairs compare results. Competition produces far more creativity here than instruction does, and creativity is exactly the faculty being trained.

Invite the security team to watch this day only. Not the build days. The attack day. It is the fastest way to give your reviewers a concrete sense of agent risk, and it turns the eventual review into a shared conversation rather than an adversarial one.

What your team takes back to its desk

After a session like this, engineers stop proposing to filter inputs and start proposing to narrow permissions. They stop treating the prompt as configuration and start treating it as untrusted. They write tools defensively without being asked.

That shift is the whole return on the week. It is not knowledge of a framework, which will be obsolete within a year. It is a reflex that applies to whatever they build next.

This is the middle day of our AI training workshops, and it is the day clients tell us about afterwards.

tMinus1 Team
about  the  author

tMinus1 Team

Digital Agency

The tMinus1 team builds digital products and systems for startups and businesses across Australia and globally. Based in Sydney, tMinus1 specialises in UI/UX design, web development, mobile app development, and generative AI services.

Learn about our Editorial policy