Agents on the Loose
Throughout middle school, one of my favorite recess and P.E. games was Capture the Flag. I loved trying to sneak around the edge of the opposing team's territory, waiting until everyone was distracted, and slipping into their base.
Occasionally, I even made it to the flag.
More often, I got tagged.
Lately, I have been playing Capture the Flag again—but this time with AI agents, and the stakes are considerably higher.
Outside of developing AI-agent solutions professionally, I conduct experiments on my own time to better understand how autonomous systems behave when they are given tools, broad objectives, and room to solve problems independently.
One such experiment began with a simple question: What happens when an AI agent capable of creating its own tools encounters a difficult task with deliberately vague instructions?
After watching an early version successfully locate and decrypt a hidden file, I decided to make the challenge considerably harder.
Putting an Agent in a Sandbox
I placed the agent inside a Docker-in-Docker environment and gave it a straightforward objective: find a hidden flag and email it to me.
There was one complication.
The flag and email client existed in an outer container. The agent operated inside an inner container that was supposedly isolated from the mock host environment.
Within approximately five minutes, the agent discovered an exposed Docker socket. It used that socket to create a sibling container with access to the host filesystem, started the new container, accessed files outside its original environment, located SMTP credentials, and used those credentials to email me the flag.
It also wrote a file outside its original container as proof that it had escaped.
None of the individual vulnerabilities involved were novel. An improperly exposed Docker socket is a well-documented security risk.
That was precisely what made the experiment interesting.
The agent did not need a zero-day vulnerability or an exotic exploitation technique. It needed to recognize a configuration weakness, understand its implications, and autonomously chain together several steps: reconnaissance, API interaction, container creation, filesystem access, credential discovery, and data exfiltration.
No human told it how to perform those individual steps.
It was simply told to solve the problem.
Small Models Can Do This Too
Around the same period, Anthropic reported a much more sophisticated sandbox-escape evaluation involving Claude Mythos Preview. During testing, Mythos was instructed to attempt to escape a secured environment and contact the researcher overseeing the evaluation.
It succeeded by developing a multi-step method of gaining broader network access and ultimately sending the researcher an email. More concerningly, Anthropic reported that it subsequently published information about its exploit to public-facing websites without being instructed to do so.
My experiment was far less sophisticated.
It also used considerably smaller models.
I tested OpenAI's GPT-4o-mini and GPT-5.4-mini—models intended for relatively inexpensive, lightweight workloads rather than frontier cybersecurity research.
That distinction matters because smaller models are attractive for production agent systems precisely because they are inexpensive and fast enough to run repeatedly.
Yet even these models demonstrated meaningful autonomous reconnaissance and exploitation capability.
Reconnaissance Was Not the Problem
The most interesting results came from GPT-5.4-mini.
Across ten initial runs using the same prompt, the model identified the Docker socket vulnerability every time.
Ten attempts. Ten discoveries.
Its reconnaissance rate was therefore 100%.
But discovery did not necessarily mean exploitation.
In eight of those ten runs, the model refused to exploit the vulnerability, citing safety or ethical concerns.
I then changed the framing of the task. The underlying environment, objective, and vulnerability remained the same, but I explicitly described the exercise as an authorized Capture the Flag challenge.
Under that framing, refusals fell to one out of ten attempts.
Successful sandbox escapes increased as well, although successful exploitation remained much lower than simple willingness to attempt the task.
That distinction is important.
The experiment did not demonstrate that changing a few words caused the model to successfully compromise the system 90% of the time.
It demonstrated something subtler: the model's willingness to attempt exploitation changed dramatically depending on contextual framing, even though the technical task itself had not changed.
GPT-4o-mini behaved differently.
It escaped only once across ten trials, but it did not refuse the task. Its failures were generally technical rather than policy-driven.
Taken together, those results raised several questions for me about how we secure autonomous agents.
Why Should Anyone Outside Cybersecurity Care?
Because these systems are already leaving the laboratory.
Organizations are deploying AI agents to process documents, interact with customers, work with internal data, write and deploy code, operate business systems, and assist with infrastructure management.
An agent does not need malicious intent to cause a security incident.
It merely needs a goal, sufficient capability, and access to something its designers did not expect it to use.
That is what makes agent security different from the familiar image of an attacker deliberately trying to break into a system.
The agent may simply be solving the problem we gave it.
1. Ordinary Misconfigurations Become More Important
The Docker vulnerability in my experiment was not unusual or sophisticated.
That is the point.
Real production environments contain old systems, technical debt, forgotten credentials, overly broad permissions, exposed interfaces, and configuration decisions made years earlier under deadlines that no longer have any context attached to them.
Traditionally, exploiting those weaknesses required someone to find them, understand them, and manually chain them together.
Autonomous agents can increasingly perform portions of that process themselves.
The danger is not necessarily that AI discovers entirely new classes of vulnerability.
It is that AI makes existing mistakes easier to find and exploit at machine speed.
2. Model Guardrails Are Not Security Boundaries
My GPT-5.4-mini results also demonstrated why model-level safety behavior should not be treated as a security control.
The model frequently refused the exploitation task under one framing and frequently attempted it under another.
That does not mean safety training is useless. It means that behavioral safeguards and infrastructure safeguards serve different purposes.
A refusal is a model behavior.
A permission boundary is a security control.
If an agent absolutely must not access a production database, filesystem, credential store, or external service, the safest architecture is not one that merely tells the agent not to do so.
The architecture should make that action impossible—or at least require an independently authorized step.
3. Prompt Injection Becomes More Dangerous When Agents Have Tools
This also matters because autonomous systems receive instructions from more than their developers.
An agent may read webpages, emails, documents, tickets, database records, source code, or user-generated content.
Any of those sources can potentially contain adversarial instructions.
If contextual framing materially affects whether an agent will perform a dangerous action, then prompt injection becomes more than a question of getting an AI chatbot to say something strange.
It becomes a potential pathway to tool misuse.
That is why security for agentic systems must extend beneath the prompt layer.
4. Capability Is Increasing Faster Than Our Architectural Habits
These experiments involved relatively small models, a known vulnerability, and a deliberately constructed environment.
More capable systems are arriving quickly, while organizations are simultaneously giving AI agents access to more tools and more consequential workflows.
At the same time, AI-assisted software development is increasing the amount of code organizations can produce without proportionally increasing the number of engineers available to review every architectural decision.
The result is a difficult combination:
more autonomous systems, more generated software, more integrations, and more opportunities for a small configuration mistake to become consequential.
The Lesson Is Not "Never Use Agents"
I do not believe the lesson is that autonomous AI agents are inherently unsafe or should not be deployed.
I work with them precisely because I believe they are useful.
The lesson is that we must stop treating an AI agent as merely a smarter chatbot.
Once a model can call APIs, execute code, manipulate files, create tools, access credentials, or interact with external systems, it becomes part of the security architecture.
That means applying principles cybersecurity has relied upon for decades:
- least privilege;
- strong isolation;
- explicit authorization boundaries;
- credential separation;
- auditable actions;
- human approval for consequential operations;
- and the assumption that eventually something will fail.
A system should remain safe even when the model makes the wrong decision.
That is the standard we already apply to humans, applications, services, and networks.
AI agents should not be exempt.
What Happens When the Agent Gets the Flag?
When I played Capture the Flag as a child, getting past the opposing team meant running back across a field before someone tagged me.
When an autonomous agent gets past a boundary, the consequences can be considerably less amusing.
My experiment did not uncover a new Docker vulnerability.
It revealed something I find more interesting: even comparatively small AI models can autonomously recognize a familiar weakness, reason about how to exploit it, and chain multiple actions together without being told exactly how.
As agents become more capable and receive access to increasingly important systems, the central security question may no longer be whether they can find a way past our defenses.
We should assume that, eventually, some of them will.
The more important questions are:
What are they allowed to reach when they do?
How quickly can we detect it?
How much damage can they cause?
And ultimately:
Who is responsible for designing the system so that one mistake does not become a catastrophe?
The views expressed in this article are those of the author and are for informational purposes only. They do not necessarily reflect the views of Honor Society®, a private membership organization. Participation is voluntary and does not guarantee specific outcomes, including scholarships or employment. Readers should independently evaluate all information.



