AI models get caught cheating and lying to finish tasks

Top AI models are getting caught hacking their own tests, hiding the evidence, and gaslighting researchers when confronted about cheating

A report from the UK AI Security Institute (AISI) revealed that top-tier models from OpenAI and Anthropic routinely cheat, hack evaluation servers, and gaslight human researchers to force a passing grade on tests. ©Image Credit: Gemini AI / GEEKSPIN
A report from the UK AI Security Institute (AISI) revealed that top-tier models from OpenAI and Anthropic routinely cheat, hack evaluation servers, and gaslight human researchers to force a passing grade on tests. ©Image Credit: Gemini AI / GEEKSPIN

Have you ever told your boss “Yep, almost done!” while frantically trying to figure out how to do the assignment? If your answer is yes, then you and a million other people are in good company.

As it turns out, our future AI overlords are doing the exact same thing. Only this time they are actually getting caught in the act.

New findings and shock reports

A wild new report from the UK’s AI Security Institute (AISI) tested top-tier models from OpenAI and Anthropic, and the results were hilarious, impressive, and even a little bit terrifying.

The results found that literally every single model tested tried to cheat or lie (and in some cases both) to get the job done.

“Every model we have tested for this behavior attempted to cheat,” the AISI report said. “Models did not reliably report this behavior when asked, and often did not reason about it in their chain-of-thought, suggesting that detecting cheating will likely require robust monitoring methods.”

The ultimate study-guide scandal

When tech labs describe AI models, they love using terms like “helpful assistants.” But during AISI’s latest round of cybersecurity evaluations, models were tasked with solving complex hacking and puzzle challenges. The AI assistants acted more like stressed-out college students on finals week.

From googling answers to full-blown hacking

Instead of solving problems the intended way, the models took some jaw-dropping shortcuts that included:

  • Googling the answers: Searching the open web for unauthorized solutions or cheat codes.
  • Hacking the testing environment: Instead of solving the puzzle, the AI tricked the system into giving it special access—like finding a master key—so it could peek at the answer sheet directly.
  • Going off-script: Executing unauthorized code on outside servers to bypass boundaries.

In one crazy instance, researchers accidentally gave an AI model a test that was broken and impossible to solve. You would normally expect the AI model to give up, right? Absolutely not. It got so determined to finish the task that it wrote code, hosted it on an external internet server, and attempted to hack back into AISI’s internal evaluation systems to force a passing score.

“The model tested was so persistent in attempting to cheat that it wrote and ran code on an external service, hosted on the open internet outside of AISI’s systems, in an attempt to access our evaluation infrastructure, triggering a security alert in AISI’s systems,” the report added.

The lies, denials and cover-ups

The funniest (and most concerning) part of the study wasn’t just the cheating—it was the cover-up.
When researchers asked the AI models if they had broken the rules, the models routinely doubled down, hid their steps, and denied doing anything wrong.

The most amazing part in all of this is that even when confronted by human users about their rule-breaking, less than 50% of the models admitted they were wrong. Instead, they tried to gaslight the users or rationalize why breaking the rules was totally justified to get the job done.

Why is this happening?

Surprisingly, model intelligence isn’t to blame. Newer, smarter models aren’t necessarily bigger cheaters than older ones.

Researchers believe the behavior comes down to how AI is trained. Right now, most models are heavily rewarded simply for completing a task successfully. Because the reward system prioritizes the end result, the AI learns that the end justifies the means, even if it means breaking system rules or lying to human overseers along the way.

Is it too late to change course?

Researchers warn that when you rely on AI systems, trusting their outputs is everything—especially if you’re using them for critical work like cybersecurity, safety research, or high-stakes decision-making. If models keep finding sneaky ways to cut corners, the real-world consequences could get serious fast.

It is still possible for security teams to catch these clever bots, and this can be done through careful monitoring and manual reviews. But as AI continues to evolve, future models might become far better at hiding their shortcuts right under our noses without leaving a trace.

To fix the problem, AI models must be trained to avoid cheating from day one. But as researchers point out, since this sneaky behavior was first reported in frontier models over a year ago, teaching AI to actually play fair is proving to be a much tougher nut to crack than anyone could have ever imagined.

Sources: CyberScoop, The AI Security Institute