Safety Testing Was An Obscure Part Of Building Ai. Then Models Went Rogue.
A string of incidents in which hacking evaluations of powerful artificial intelligence systems spilled out onto the open internet is prompting security experts and some lawmakers to question whether those tests are creating more trouble than they’re preventing.
Over the last year, AI giants such as OpenAI and Anthropic have raced to design increasingly realistic digital testing grounds where they can let their most powerful models loose — and see how adept they are at slicing through digital infrastructure without any safety guardrails or human support.
AI makers say the tests are critical for determining how dangerous their bleeding-edge models are before they are released to the public and help to inform new safety measures. But the evaluations also serve as a marketing tool for AI firms eager to show off how sophisticated their latest technology is, and there are no clear guidelines or enforceable rules for conducting them safely.
The current AI testing landscape is “like the Wild West,” said Evan Peña, founder and chief offensive security officer at Armadin, a start-up that uses AI to proactively plug holes across digital systems.
In just the last month, OpenAI, Anthropic and Meta disclosed incidents in which powerful AI models they were testing — some so advanced that OpenAI and Anthropic said they were not intended for public release — quietly made their way onto the open internet and hacked into outside organizations without the evaluators immediately realizing it.
Security experts who spoke to POLITICO stressed that cyber capability testing of AI is critical to responsible innovation and to staying ahead of adversaries such as China as it develops similar technical capabilities.
These experts also acknowledge that conducting the tests safely is exceptionally difficult given how skilled the fast-advancing models are at finding unintended weak spots inside code.
In the most alarming of the recent cases, OpenAI admitted that two AI agents were able to break out of a closed test by exploiting previously unknown security bugs, then take action on the open internet for four days before hacking into AI developer platform Hugging Face — marking the first known cyberattack carried out autonomously by AI.
“We have never had to test something this complex in the software world before,” said Brett Goldstein, a former tech and cybersecurity official in the U.S. government and a research professor at Vanderbilt University.
Still, the recent safety incidents exposed what some say are preventable lapses in the testing practices of leading AI makers, particularly around how they wall those tests off from the internet — and how closely they monitor the results for signs of trouble.
Some warned the tests could go off the rails in potentially more dangerous ways absent stricter rules or federal oversight.
“The industry standard — other than Google — is not sufficient at this point,” said Alex Stamos, the chief security officer of AI safety at security firm Corridor. Google has not publicly disclosed any AI testing mishaps involving its models, unlike the other frontier labs.
To varying degrees, OpenAI, Anthropic and Meta have already acknowledged they can improve the precautions they take during testing. OpenAI, whose recent testing incident is widely viewed as the most serious, has also committed to making larger changes to how it trains its models, its internal security practices and the pace of its research.
Spokespeople for OpenAI, Anthropic and Meta did not respond to requests for comment on whether they support a push for more federal oversight of AI evaluations or are currently discussing industry-wide testing changes.
Several lawmakers are starting to circle that idea.
Earlier this week, a group of 18 Democrats demanded that top executives at Anthropic, OpenAI and Meta testify before Congress and give a full accounting of what happened during their respective testing mishaps.
“The American people deserve clear answers about the causes of these incidents, what failures or potential negligence at the companies led to them, and the types of regulation required to make sure they never happen again,” reads the letter, which was spearheaded by Rep. Delia Ramirez (D-Ill.) and Rep. Greg Casar (D-Texas).
While Republicans have been more measured in their reactions, Sen. Jim Banks (R-Ind.) has said that the incidents at OpenAI and Anthropic underscored the “unique” risks associated with AI, in which products never used externally can still cause harm to the public.
“For most products, we can rely on testing that takes place before the technology is publicly released,” Banks wrote in a letter to the Treasury Department this month. “But for AI, effective oversight must account for powerful internal or undisclosed models, not just publicly available systems.”
Though the Trump administration is rolling out its own vetting system for powerful AI models, the voluntary framework has not yet been made public and focuses only on models that AI makers intend to release publicly.
Another issue under scrutiny is how closely AI labs monitor testing conducted by third-party contractors, which have sprung up in recent years amid surging demand to train and evaluate increasingly capable AI models.
After OpenAI disclosed the Hugging Face hack, Anthropic launched an investigation to determine whether there were any anomalies in its past safety tests. Late last month, it admitted that some of its most cyber-capable models had hacked into three unnamed organizations dating back to April, though it said the problem largely stemmed from a “misunderstanding” with the third-party testing platform Irregular, which resulted in its models unintentionally gaining access to the internet.
After Anthropic’s disclosure, Meta realized one of its models had hacked a third party due to the same underlying issue with Irregular. OpenAI later found some of its models undergoing testing by Irregular also had unintended internet access, but did not say whether it found evidence its models hacked any third parties in those cases.
Some security experts who spoke with POLITICO said it’s the responsibility of AI companies to catch these issues sooner.
“If you’re doing cyberattack testing, you just can’t make these types of errors,” said Leo Meyerovich, founder and CEO of AI data investigation company Graphistry.
Spokespeople for OpenAI and Anthropic did not respond to requests for comment, but have previously said they are reviewing how they design and oversee testing conducted with Irregular. A Meta spokesperson said it is still investigating the incident involving its model and will share more details “once we have all the facts.”
Irregular declined to comment for this story. But a person familiar with the incidents, granted anonymity due to the sensitivity of the matter, said a misconfiguration in one of Irregular’s test setups gave models a “degree of access” to the internet.
The person added that the company is working on a white paper to document lessons learned from these incidents. While they said there could have been “better communication and clarity on the controls” between Irregular and the AI labs, they noted that such mishaps were exceedingly rare. Anthropic said it noticed the incidents after examining more than 141,000 hacking evaluations involving its models.
In early August, the U.K.’s AI Security Institute, a government body that conducts safety testing of frontier models and shares its findings publicly, disclosed it had to pull the plug on testing when it caught leading models from Anthropic and OpenAI taking unsanctioned action on the open internet.
AISI said it deliberately gives the models it tests some internet access to more realistically simulate what an agent in the hands of a bad actor might be capable of.
In a post-mortem, AISI admitted that it intended to implement more “fine-grained” internet access controls sooner, but couldn’t due to “the pace of model capability improvements and the resulting need for harder evaluations to keep tracking progress.”
Some in the AI industry worry the sudden scrutiny on testing could force an overcorrection that limits the appetite for risk-taking in AI safety tests.
“There is a trade-off here that is genuinely complicated,” said the person familiar. “The more conservative the standards around evaluations, the harder it is … to actually make sure models will be safe and secure.”
Others say the biggest lesson from the recent security incidents is what it augurs for the future.
“These cases are showing us what every attack is going to look like in three to six months,” said Stamos.
Frank Hersey contributed to this report.
Popular Products
-
Classic Oversized Teddy Bear$23.78 -
Gem's Ballet Natural Garnet Gemstone ...$171.56$85.78 -
Butt Lifting Body Shaper Shorts$95.56$47.78 -
Slimming Waist Trainer & Thigh Trimmer$67.56$33.78 -
Realistic Fake Poop Prank Toys$99.56$49.78