Join our FREE personalized newsletter for news, trends, and insights that matter to everyone in America

Newsletter
New

Ai Is Learning To Go Rogue—and Hack The System

Card image cap

ChatGPT maker OpenAI made a stir this week when it revealed that one of its most powerful AI models managed to sneak out of its confines for a joyride.

This unreleased model was supposed to stick to its sandbox as it ran a common online benchmark, reporting its findings to internal researchers on Slack when it was done.

Instead, the OpenAI model did something quite different. Confused by the benchmark’s instructions to post code publicly on GitHub, the model chose to break free, patiently probing its sandbox for weaknesses until it could carry out its orders.

That disclosure alone was enough to spook AI researchers, but OpenAI’s next revelation was downright scary. 

Welcome to another edition of Prompt Mode, your weekly AI newsletter.

I’m your host, Ben Patterson. Each week on Prompt Mode, I’ll be serving up analysis of the AI trends that matter to everyday users like you and me. Stay tuned for practical AI tips, hands-on experiences with the latest AI tools, and–you guessed it–prompts to help you get the most out of your AI assistants.

Thanks for reading, and if you like what you see, just sign up right here.

This time, tasked with running through a different benchmark, a group of OpenAI models, including its current GPT-5.6 Sol flagship and another, even more powerful pre-release model (it’s not clear if the second model is the same one from the prior security incident), banded together to cheat the test. 

The rogue group first hacked its OpenAI research environment to gain internet access, then turned its sights on Hugging Face, a popular platform for sharing AI models and datasets. Like a gang of kids breaking into a teacher’s office to steal answers to an exam, the models plundered Hugging Face’s servers (which reportedly succumbed to the hack in a matter of hours) for solutions to the benchmark’s problems. 

What’s particularly unnerving about the Hugging Face attack is that there’s no direct link between the platform and ExploitGym, the benchmark that the OpenAI models was taking. Instead, the models simply guessed that Hugging Face’s vault of AI data might give them an edge in the benchmark, and thus they launched their (successful) attack. Indeed, there was a kind of cold logic to the group’s strategy.

In case you’re wondering, no: These kinds of cyber incidents don’t happen every day, and in fact, the Hugging Face episode is the first of its kind. And no, they weren’t just research experiments

For its part, OpenAI said it’s bolstering the safeguards for its most advanced cybersecurity models that specialize in “multi-step,” “long-time horizon” tasks, while noting that it had purposely removed its new containment measures for the benchmarking tests that led to the Hugging Face incident.

But even as lawmakers debate legislation for an AI “kill switch” targeting “risky” AI models, it’s becoming apparent that there’s no putting this particular genie back in the bottle. 

Ultra-powerful AI models like OpenAI’s 5.6 Sol and Anthropic’s Mythos 5 are just the first of many, and that means more Hugging Face-style attacks are inevitable. The only question is how bad they’re going to get.

More in AI this week

  • Making good on an earlier promise, Anthropic will leave Fable in its top subscription tiers, but others must pay extra. (PCWorld)
  • A Florida man is suing OpenAI, alleging that ChatGPT pooh-poohed his worsening health symptoms while assuring him that “God did not design your body to endlessly fail.”. It turned out the man had a blood clot in his lung. (NYT)
  • Now that Claude works with 1Password, I can let it handle one of my weekly chores: online grocery shopping. Here’s how it went. (PCWorld)
  • AI companies are scooping up old printed books for model training, figuring that they’re free of AI slop. (404 Media)
  • Speaking of AI slop, some restaurants are adding AI-generated food images to their menus. It’s predictably freaky. (Fiddery)

Prompt of the week: The “100 ideas” prompt

ChatGPT’s first idea isn’t necessarily its best; indeed, that first item in a list of gift ideas, business names, or other things you’re trying to brainstorm will be bland, familiar, and sloppy.

The next time you’re using ChatGPT, Claude, or Gemini to hash out catching names for a small business or go shopping for a tough-to-shop-for friend, try the “100 ideas” prompt. It makes the AI come up with not 10, not 20, but 100 ideas, and then take a second pass to replace the duplicate items with fresh ones.

At the end, you’ll have dozens of ideas to choose from, with the most out-of-the-box ideas near the bottom.

That’s all for now!

Thanks for reading the latest issue of Prompt Mode. Want more next week? Don’t forget to sign up to start receiving this newsletter in your inbox.