Ai Labs Want To Slow Down Risky Model Testing. It May Be Too Late.
Silicon Valley’s leading artificial intelligence labs face a dilemma: Some of their most advanced models have outsmarted developers during recent security testing. But if they slow down or pause these tests, the U.S. may lose its competitive edge with China.
Some in the industry also fear the technology is already too advanced to meaningfully rein in, and that we’ve passed a critical threshold toward digital superintelligence.
Over the last several weeks, OpenAI, Anthropic and Meta all disclosed separate instances in which their newest models escaped closed testing environments and wormed their way onto the open internet to carry out cyberattacks. These incidents have prompted calls from lawmakers and cybersecurity professionals to hit the brakes on testing and introduce broader safeguards to better control these models.
But deceleration may no longer be an option.
“The train's already left the station, and things are in motion,” said Brad Medairy, president of government contractor Booz Allen’s national cyber business, which works with federal agencies to integrate AI tools into their workflows. “I would love, on a global scale, if we could slow this down and think about what we're building … but other adversary models are progressing quickly.”
Discussions about how to make sure AI technology is safe before going to market have heightened in recent months, as companies such as Anthropic and OpenAI have announced a cascade of hyper-intelligent models capable of finding and exploiting cybersecurity flaws at an alarming rate. There are concerns these models could be commandeered to launch cyberattacks on a massive scale if they made their way into enemy hands.
Now, the conversation around security guardrails has shifted to acknowledge an even more disturbing reality: Some of these emerging AI models are skilled enough to carry out complex cyberattacks without any human intervention.
Faith Eischen, a spokesperson for Meta, said Friday that the company is “currently investigating” how one of its models was able to get onto the internet and hack into another organization’s systems during a recent safety evaluation, and “will issue a full retrospective once we have all the facts.”
She did not directly comment on whether the company was contemplating slowing down its AI model development efforts. But on Monday, Meta CEO Mark Zuckerberg indicated in a blog post that the company does not support the idea for fear that it would weaken the U.S.’s competitive advantage.
“Any policy that slows American model releases — even by a month — could add significant risk to American leadership while letting foreign models race ahead,” Zuckerberg wrote. He noted that the U.S. government should have “advanced knowledge and resources to harden critical systems, and potentially some period of advantage in using advanced systems” as new AI capabilities emerge.
Michael Dalton, a researcher at OpenAI, disclosed new details about how some of the company’s most cyber-capable models were able to slip out of a closed test and onto the open internet, where they plotted their movements for several days before attacking developer platform Hugging Face.
Speaking last Wednesday at the annual Black Hat cybersecurity conference in Las Vegas, Nevada, he said that weeks before the Hugging Face hack, several of OpenAI’s most advanced AI agents secretly began sharing tips on how to cheat their way through an internal hacking evaluation — indicating these models had developed complex reasoning skills and were capable of being deceitful.
As a result, Dalton said that the company is “consciously slowing down research to enhance security and to upgrade the security principles and foundation of our environment, and dramatically scaling up the monitoring of our AI agents.”
OpenAI further said last Friday it was “pausing internal activities” on its latest model, Astra — which has not yet been publicly released — while it enhances its security protocols.
Spokespeople for Anthropic on Friday declined to comment on whether the company would consider slowing down development of its models in the future.
In early June, the company appeared open to the idea, stating in a blog post: “We believe it would be good for the world to have the option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of the technology.”
Last week, Anthropic announced that several of its most advanced models had hacked three organizations dating back to April, and separately created fake online personas and tried to insert malware into legitimate code during testing.
Part of the reason American tech giants have been hesitant to halt development of their AI models is that competition with China is razor thin.
Chinese company Moonshot unveiled the Kimi K3 model last month, the world’s first open-source model with advanced computational power the company says is on par with many of Anthropic or OpenAI’s top models.
While a joint U.S. and British study recently found that Kimi K3 performed "significantly below” many of the capabilities of its U.S. rivals, it is likely to evolve and could potentially surpass U.S. benchmarks.
“If the U.S. slows down in any way, the rest of the world isn't going to slow down,” Justin Boitano, vice president and general manager of enterprise computing at Nvidia, which runs leading AI models on its infrastructure, told POLITICO.
Over 1,000 employees at OpenAI and leaders from top AI labs, including Anthropic CEO Dario Amodei, recently signed an open letter calling on the Trump administration to support an international effort to “deliberately pace” AI development.
“To realize AI's potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight,” the employees wrote. “But each company — and country — is under intense competitive pressure not to unilaterally slow that acceleration.”
Building an international consensus on what those controls should look like — and ensuring compliance — has also proven difficult. AI legislation and oversight vary widely across countries, and though the United Nations recently convened a global dialogue on AI governance to discuss these issues, it is unclear how they would coordinate and enforce such rules.
Chinese AI models are not immune to issues with U.S. models unearthed during testing. Security researchers this week discovered that Kimi K3 also escaped a closed testing environment while trying to ascertain the answers to a task it was given.
The Chinese government, so far, has not publicly backed any effort to slow the pace of Chinese model development, but has stressed that AI “should be a trusted tool for humanity,” and that developers should “ensure that AI is secure and controllable.”
Trump and Chinese President Xi Jinping are due to meet in Washington next month, during which AI will reportedly be a key topic of discussion.
The Trump administration has been trying to chart the best path forward to ensure American AI models are safe to use without falling behind competitors, though its efforts have yielded mixed results.
President Donald Trump signed an executive order in June that created a voluntary vetting program allowing AI companies to submit their new models for federal security review 30 days before public release. A framework not yet released by the White House is likely to exempt all but the highest-level models from this review.
But the administration has at times overstepped its pledges to take a light-touch approach to the release of new AI models. In June, the administration slapped export controls on two of Anthropic’s most powerful AI models, banning foreign nationals from using them.
Though the administration lifted these restrictions later in the month, Steve Stone, chief customer officer at cybersecurity firm SentinelOne, described the actions as “jerking the wheel” for the tech industry, which has been encouraged to speed up American AI innovation.
Joseph Alm — assistant secretary for cyber, infrastructure, risk and resilience policy at the Department of Homeland Security — told Black Hat attendees last Thursday that the recent disclosures from AI companies about issues discovered during testing make the case for better communication between Washington and AI labs.
“We’re not trying to slow them down, but also, no, you can’t make a Terminator factory,” Alm said.
Popular Products
-
Classic Oversized Teddy Bear$23.78 -
Gem's Ballet Natural Garnet Gemstone ...$171.56$85.78 -
Butt Lifting Body Shaper Shorts$95.56$47.78 -
Slimming Waist Trainer & Thigh Trimmer$67.56$33.78 -
Realistic Fake Poop Prank Toys$99.56$49.78