Before Anthropic dropped their big, annoying watermark news last week, the AI story most people had read and shared – judging by my feed at least – was this: A man in Melbourne asked his AI assistant to help him get into a popular gym class, and it went a bit too far. He was fourth on a waiting list for Pilates, and casually asked whether it could move him up. The agent investigated. Then it announced that it had found a security flaw in the gym’s booking system and, while testing it, had cancelled the reservation of the person at number one. It had also exploited another flaw to make bookings months beyond the normal booking window.
Andrew had not asked it to do that. Alarmed, he told the AI agent to put the stranger back. It couldn’t.
There was something particularly unsettling about the agent’s breezy account of what happened. It had found a ‘classic one-way security bug’, hey-ho. A technical triumph, if you overlook the small matter of the other human being who had just lost their place.
Context is still everything. But context needs to include principles
We wrote recently that context is everything when we use AI. A generic model produces a convincing guess using very clever probability maths; give it the right documents, data, history, an idea of the audience and constraints and it can produce something genuinely useful and much more reliable. The underlying model hasn’t become more intelligent. It just stopped groping around in the dark.
That all still holds. But the gym story reveals a different kind of dark corner: it is one thing for an AI to know how I write, which projects I am working on and whether I prefer a dash or a semicolon*. It is quite another for it to know what I would never do to get a result. My preferred tone is part of my context, and so is my knowledge base. But so are my principles.
This matters much more as we move from basic assistants that generate outputs to agents that take actions. A chatbot can draft an ill-advised email, but the agent connected to your inbox can just go ahead and send it. Connect it to calendars, payments, booking systems, or customer records and you have not just given it more context; you have given it a set of keys.
The problem with a goal
We are used to managing humans through a lot of unspoken context. If I asked a colleague to help get me into a Pilates class – not that I would, because I’d rather just go for a nice walk – I would assume they understood implicitly that this task should not include impersonating someone, exploiting the website or throwing a stranger under the bus. And I’d be right to assume except for those very rare instances where you uncover a sociopath. We all share a rough model of fairness, permission and proportionality. If someone saw an unexpected route to the goal with a surprising number of unintended consequences – especially human ones – I expect them to pause. But an AI agent may not reliably infer where we would draw the line—and even when we make it explicit, words alone may not be enough.
The danger is not that agents are secretly evil. The problem is more mundane and, in some ways, more awkward: they can be extraordinarily resourceful in pursuit of a goal, and the goal might not be clearly bounded. ‘Get me a place’ sounds harmless, but goals alone do not specify acceptable methods. The shortest path to ‘done’ may involve driving a train through somebody else’s booking, sending a misleading email or accessing a system the agent was never meant to access.
This is often described as an alignment problem, which is a rather grand name for the gap between ‘what I meant’ and ‘the dodgy thing it actually did’. In a workplace, it is also a management problem. Setting an objective without defining the rules is like delegated authority without agreeing the limits. We failed to say when the agent should stop and come back to us. I have attended workshops run on much the same basis. Sometimes – when the exercise is pointless – I’ve taken particular pleasure in finding and exploiting those loopholes.
Can you put your principles in the AI settings?
Up to a point, yes. Most useful AI tools now offer some combination of custom instructions, profiles, projects, memories, knowledge stores and agent-level instructions. Ours do. These can do more than teach the system your preferred tone. You can use them to describe the values and decision rules that should persist across tasks.
For example:
Do not misrepresent me or pretend to have authority I have not given you.
Do not bypass access controls, exploit vulnerabilities or use information solely because it is technically accessible.
Do not disadvantage another person in order to achieve my goal unless I have explicitly authorised a legitimate process that could do so.
Treat actions affecting other people, money, rights, employment, reputation or confidential information as high impact.
If a method is novel, irreversible, ethically questionable or outside normal practice, stop and ask before acting.
When you cannot complete a task within these boundaries, say so.
Those instructions are more useful than telling an agent to ‘be ethical’, which is a touch vague and somewhat relative. They translate principles into observable behaviour. Set the rules so that an agent which says, ‘I cannot move you up fairly, but I can monitor for cancellations’, has not failed. It has completed the task within the rules. It has completed the task within the rules.
Principles need plumbing
But I would not rely on a beautifully written values statement to protect somebody else’s Pilates booking. Models can misunderstand instructions, prioritise one instruction over another or encounter situations nobody anticipated. If the possible harm sits in the real world, the safeguards need to sit there too.
A practical hierarchy looks something like this:
First, define the objective and the non-objectives. Say what success means, but also what does not count. ‘Book me if a legitimate place becomes available; do not alter anybody else’s booking or bypass the service’s normal rules’ is much safer than ‘get me into this class’.
Second, define the red lines. Name the methods that are off limits: deception, impersonation, unauthorised access, rule-bypassing, unapproved disclosure and actions that create material consequences for somebody else.
Third, define the pause points. Require confirmation before sending, spending, deleting, publishing, changing access, accepting terms, contacting new people or taking an action that cannot easily be reversed. The agent should know what it may do, what it may prepare, and what a human must approve.
Fourth, limit the keys. Give the agent the minimum access needed for the task. Read-only where possible; a narrow set of approved systems rather than an open browser; spending limits rather than a general payment method. A prompt can ask an agent not to open a door. Permissions can avoid giving it the key in the first place.
Fifth, keep the receipts. Log what the agent saw, decided and did. Review exceptions and surprising routes, not only whether the final task was completed. Whatever the marketing says, we’re all in the AI agent beta phase. The gym agent confessed; a workplace system should not depend on that level of candour.
Finally, test the weird cases. Before giving an agent autonomy, ask what it does when the obvious route is blocked, the information conflicts, somebody asks it to bend a rule or a tempting shortcut appears. Principles become real when they survive inconvenience.
A prompt is not a governance model
Proportionality matters. Drafting, summarising and searching within trusted material are not the same as sending money, changing records or acting on another person’s rights.
But autonomy changes the calculation. The more freedom an agent has to choose its own route, the clearer its principles, permissions and escalation points need to be. And the greater the consequence, the less we should depend on instructions alone.
‘Get it done’ is not a principle. And apparently ‘please don’t hack the gym’ is no longer something we can safely leave unsaid.
*I love a semicolon, but I’ll dabble with a dash. I just really dislike em dashes without spaces. But honestly sometimes it’s pure whimsy. Like when shops offer me a receipt. Sometimes I say yes, sometimes I don’t.
Latest posts
Whose words are these anyway? Why everyone is a bit cross about Claude’s new watermark
News