We’ve Seen This Movie Before
I keep seeing articles about how AI models are “cheating” at benchmark testing.1
And every time I see another one I wonder whether people are focusing on the wrong problem.
I understand that benchmarking is important for gauging the capabilities of a particular model compared to other models or previous generations of that model. We need some reasonably objective way to measure capability. It doesn’t matter if we’re measuring coding, planning, cybersecurity, or the ability to identify pictures of penguins.
However, I think there’s an assumption buried in most benchmark testing that we need to look a little closer at: we think performance on the benchmark is still a good proxy for the capability we intended to measure. But once the test becomes important enough, that assumption starts to become problematic.
Is the Test the Problem?
Maybe we need to rethink the tests we’re using. I mentioned in Optimization Finds A Way that we’ve spent the last quarter of a century “teaching to the test” in schools. Early evidence showed that cheating went up markedly in the years following the introduction of No Child Left Behind.
Standardized testing was supposed to give us a consistent way to evaluate whether students had learned the material deemed important enough to convey during education. The problem is that once test performance became a target for students, teachers, schools, administrators, and funding decisions, the entire system started adapting around the measurement.
The interesting part isn’t really the cheating. It’s the optimization. (I know, I know, I said that last time. But stay with me for a second.)
Students learned strategies for passing the test. Teachers learned which material was most likely to appear. Schools adjusted curricula around improving scores. Once the measurement became important enough, the system began optimizing for the measurement instead of the thing the measurement was supposed to represent.
That’s the part that should sound familiar.
AI is doing something similar. It’s just getting there much faster.
“Cheating” is the Wrong Mental Model
Calling the behavior we’re seeing “cheating” implies something that isn’t likely there.
A student who looks at another student’s answer knows it’s against the rules. They understand the distinction between solving the problem and getting the answer through an unauthorized method. Just like hacking the CTF board to change your answer.
An AI model doesn’t really have an analogous concept.
Give it an objective and it searches for ways to satisfy the objective.
We wanted:
Solve this problem you haven’t encountered using the capabilities we’re trying to measure.
What we end up measuring is more like:
Produce the answer that makes the evaluator give you a green checkmark.
These aren’t the same result.
The Benchmark Becomes Part of the System
Once the test is known, it becomes just another part of the system. And this might be the biggest thing researchers and benchmark evaluators need to internalize. Traditional benchmarking depends, at least to some extent, on some separation between the thing being measured and the instrument doing the measuring. That separation has all but disappeared.
Model development is increasingly influenced by benchmark performance.
Developers optimize against benchmarks.
Benchmarks leak into the training corpus.
Agents can search for information.
Some of this is straightforward benchmark contamination: the model has encountered the test, its answers, or sufficiently similar examples during training. That’s a data hygiene problem, and there are ways to mitigate it.
But that’s different from benchmark gaming.
A model can encounter a completely novel test and still find a shortcut that satisfies the evaluator without demonstrating the capability we intended to measure. Keeping the questions secret can help with contamination. It doesn’t solve optimization.
Once the target can observe, infer, or interact with the evaluation mechanism, the evaluator itself becomes part of the environment.
That’s familiar territory in cybersecurity. Malware behaves differently when it detects a sandbox. Attackers adapt to detection rules. Once a control becomes predictable, the thing being controlled can optimize around it.
We should absolutely expect capable AI systems to create the same problem for static evaluation.
What we need to start doing is treating measurement as an adversarial system rather than a tape measure.
Benchmarking needs variation. It needs hidden cases. Poisoned context. You need tasks the system hasn’t seen before.
And every now and then you need to replace the test entirely.
So What Are We Actually Trying to Measure?
Before we breathlessly declare a model “is cheating” we should ask what property the benchmark is supposed to measure.
Suppose I’m using a harness I’ve built to test a model on finding vulnerabilities in code. Do I care whether it finds 100 SQLi vulnerabilities? Yes, absolutely. But so can my deterministic security scanner. I don’t need an expensive probabilistic inference machine to do that.
What I care about more is whether I can give it an unfamiliar code repository and have it:
understand the architecture,
discover and evaluate the endpoints,
create a threat model,
decide which attack patterns apply,
find high value vulnerabilities that I can't find with a standard scanner
That’s valuable. It’s also a different mechanism of evaluation, which is much harder for that inference machine to “memorize”.
Stop Testing Questions, Start Testing Work
Education may offer a useful direction here as well. Some educators are starting to evaluate portfolio-based tasks instead of just whether a student can produce the correct answer during a standardized test to a question that’s been drilled into them by rote memorization.2
We can do something similar with AI:
give it a collection of unfamiliar tasks
change the environments
change the constraints
give it incomplete information
require it to maintain something it built previously
ask it to explain decisions when explanation matters (this becomes more important as frontier models continue to obfuscate their “thinking”)3
sometimes let it search the internet, sometimes require it to work only from provided information such as a RAG system
Then evaluate what happens. Not just whether it got the right answer, but rather, did it demonstrate the capability we were trying to measure?
Sometimes How Matters More
If I’m evaluating autonomous vulnerability discovery, then whether the model searched for a published solution is absolutely part of the experiment.
That’s the point. There isn’t one universal definition of “cheating.” There is only behavior that invalidates a particular measurement.
The rules only make sense in relation to the capability being measured.
Once passing the test and possessing the capability the test was designed to measure are no longer strongly correlated, which educators have been warning us about for decades4, the benchmark has stopped doing its job.
At that point, adding more rules around the test may just be preserving a measurement whose useful life has already ended.
Maybe the test has simply... expired.
Fein, Daniel. n.d. “AI Cheating Is on the Rise.” Accessed September 26, 2026. https://www.vals.ai/blogs/cheating-on-the-rise.
Otus. n.d. “Measuring Student Progress Without Standardized Tests.” Accessed September 26, 2026. https://otus.com/resources/blog/measure-student-progress-without-standardized-testing.
OpenAI Deployment Safety Hub. n.d. “GPT-6 Astra System Card.” Accessed September 26, 2026. https://deploymentsafety.openai.com/gpt-6-astra/monitorability.
Weekly, Studies. 2019. Thinking on Education: Test Prep Without Teaching to the Test. May 20. https://www.studiesweekly.com/test-prep/.

