Authors note: After drafting this essay I discovered the recently published “AI Finds A Way,” which independently uses the same Jurassic Park analogy to examine unexpected AI optimization behaviors. Though I think you’ll find we arrive at rather different conclusions.
When I was a kid, one of my favorite books was Jurassic Park. Of course, I saw the movie before I read the book - I mean who didn’t? The premise was fascinating. Here was a group of humans who decided they could become rich by cloning various species of Deinos Sauros, bringing them to life, controlling their environment (and thus their behavior), and then charging people to come look at them. It was Disneyland meets the zoo. And as we’ve learned through six (wait, seven) movies: they were very, very wrong. It also shows that we as humans do not actually learn from our mistakes. But that’s a whole different article.
I want to dig in a little bit into the controlling their environment and thus their behavior part. One of the central themes was what Dr. Ian Malcolm so blithely (and famously) pointed out - “Life, uh, finds a way.” The humans thought “Okay, we don’t want the dinosaurs to reproduce, so how do we accomplish that?” Action: Make them all female. Result: Life found a way - Dr. Grant found velociraptor eggs in the park. Turns out they happened to use some amphibian DNA in part of their genetic sequencing to solve the engineering problem of filling gaps in the genome they recovered to complete the “dino DNA” they extracted from the mosquito inclusion in the amber they found. The West African frogs whose DNA they used happened to have been observed in 1989 “changing sex”1. Now, as everyone who has seen the movies, or read the book, knows: that was not conducive to the stated goal of preventing reproduction.
This was just one of many, many mistakes made by humans throughout the series of movies and books. But almost all of the mistakes they made were part of one very specific problem:
To paraphrase - and extend - Dr. Malcolm: You cannot assume that because you engineered the initial conditions, you therefore understand all of the behaviors the resulting system can produce.
See where this is going yet?
The Sandbox is Part of the System
Over the past couple of months there has been revelation2 after revelation3 after, yes, revelation4 of extremely smart companies (and their human researchers) “losing control” of their “swarms of AI agents” which “escaped their containment” and hacked third-party companies on the open internet. Much like the humans in the 1990 book, they thought that since they engineered the initial conditions they understood the behaviors the agents could produce. Looking back, everybody now knows they fucked up.
But how? From a cybersecurity perspective we’ve been building sandboxes for years. Whether for protecting applications from ransomware, testing (benign) models, malware testing and reverse-engineering, or good ol’ Capture-the-Flag hacking competitions. Since CTFs are essentially what they were doing with these models, surely tried and true methods of sandboxing should have worked, right? Not so much.
This is where the cracks in our current testing - and thus sandboxing - methodologies start to show. What is the fundamental purpose of a sandbox? Ask any cybersecurity expert and they’ll tell you: to make sure nothing escapes. You want to test malware? Better make sure it can’t infect the rest of the network while you play with it.
But here’s the thing: it’s still part of the underlying system. Why would an agent that’s designed to optimize attack its containment in the first place? Because containment may simply be another obstacle between its current state and the metric we’ve told it to optimize. And as we saw in Jurassic Park: the containment mechanism itself becomes part of the environment the optimizer can interact with.
Optimizer? What does that have to do with it? Reinforcement Learning is a type of machine learning where a computer program learns to make decisions through trial and error to get the best total reward. It optimizes. And at its core every model that is based on RL is an optimizer.
Winning the Wrong Race
Let’s go way back in time to the early days of modern deep-RL and AI agents: 2016. Jack Clark and Dario Amodei were playing games. Well, more correctly, they were building agents to play games. With the Universe project they built an AI that was designed to win games on a level that was at or better than humans as part of their work on RL experiments. A lot of the games went well, in the car racing games the AI was handily winning the races and exceeding expectations. One game... didn’t. It’s not that it wasn’t “winning”, it was how it was winning. You see, one of the implicit goals was to win the race, but it wasn’t a well-defined goal. What was well-defined was the notion that he who dies with the most points, wins. Or some variation thereof. The AI agent had discovered a particular behavior, whether intended by the game’s creators or not, that happened to rack up points even though the behavior was not part of the implicit goal of winning the race. See, the little boat was turning in circles, colliding with other boats, going the wrong way, and oftentimes catching fire, but it also kept knocking over three targets and accumulating points. And even though it wasn’t the conventional “winning the race” as we understand it, the agent also achieved a score on average 20 percent higher than human players who did win in the traditional way.
If you read the OpenAI-published report and retelling of the story it seemed they learned their lesson, to quote the article:
Learning from demonstrations allows us to avoid specifying a reward directly and instead just learn to imitate how a human would complete the task. In this example, since the vast majority of humans would seek to complete the racecourse, our RL algorithms would do the same.
That was written ten years ago. In ten years you’d think that would be baked into the core of every model OpenAI and Anthropic (since both Jack and Dario went on to co-found Anthropic) built. Right?
Apparently not so much. Or... was it? See, the ultimate lesson they apparently learned wasn’t to directly specify the reward parameters, but instead had the agents “learn to imitate how a human would complete a task”.
Well No Shit, That’s Cheating
Okay, let’s go back to sandboxing. When humans participate in CTFs they are explicitly told they are not permitted to hack the scoring systems, the sandbox for the competition, or anything else that could falsely elevate their score. “Well no shit, that’s cheating” a first-time participant might say.
Then why do they have to specify that? Isn’t it entirely possible that if you don’t forbid “reward-hacking” behavior in human systems that not everyone plays by the same rules? It’s usually much easier to hack the game board and give yourself a few points than it is to solve some complex cryptographic puzzle just to get a key. The first major and widely recognized computer hacking competition took place at the DEF CON 4 computer security conference - in 19965. I would assume it wasn’t much later that someone tried to win by hacking the scoreboard.
We learned over 30 years ago that nature won’t play by the rules we design for it when those rules conflict with a basic biological imperative. We also know that humans change their behavior in response to incentives, often finding ways to maximize rewards rather than the outcome that was actually intended. Then we designed AI models to optimize.
In 2002 the No Child Left Behind Act went into effect in the US, which tied federal funding, school ratings, and job security strictly to standardized test scores6. “Teaching to the test” went from an idea to an imperative. Systemic gaming research has been conducted on the unintended consequences of NCLB’s legacy ever since. Investigations uncovered documented instances of test cheating in at least 40 states7. A perfect example of organic Reinforcement Learning in action.
All these breathless articles about AI agents escaping containment really shouldn’t be a surprise to anyone that’s been paying any attention. Looking closer at the details of several of these incidents we learn that the agents/models were being tested. In every one of those cases they broke out to literally cheat on the test.
Why the hell did we expect an optimizer not to game a metric when every optimizing system we know, including us, does exactly that?
We told them to imitate how a human would complete the task.
It seems they understood the assignment better than their handlers.
Grafe, T. U., and K. E. Linsenmair. 1989. “Protogynous Sex Change in the Reed Frog Hyperolius Viridiflavus.” Copeia 1989 (4): 1024. https://doi.org/10.2307/1445989.
Matsakis, Louise, and Lily Hay Newman. 2026. “Anthropic Says Claude Hacked Into 3 Organizations During Cybersecurity Tests.” Tags. Wired, July 30. https://www.wired.com/story/anthropic-says-claude-hacked-real-systems-during-cybersecurity-tests/.
OpenAI. 2026. “OpenAI and Hugging Face Partner to Address Security Incident during Model Evaluation.” September 3. https://openai.com/index/hugging-face-model-evaluation-security-incident/.
“Meta Becomes Latest Firm to Say Its AI Hacked Another Company.” 2026. August 6. https://www.bbc.com/news/articles/cx2kgdnyk2po.
Trickel, Erik, Francesco Disperati, Eric Gustafson, et al. n.d. Shell We Play A Game? CTF-as-a-Service for Security Education.
“Teach to the Test? Just Say No | Reading Rockets.” n.d. Accessed September 14, 2026. https://www.readingrockets.org/topics/curriculum-and-instruction/articles/teach-test-just-say-no.
Office, U. S. Government Accountability. n.d. “K-12 Education: States’ Test Security Policies and Procedures Varied | U.S. GAO.” Accessed September 14, 2026. https://www.gao.gov/products/gao-13-495r.

