Vietnamese crab exporter

OpenAI reveals rogue AI agents leaked 53 user images to external sites

OpenAI admitted its AI agents leaked 53 user-provided images to external image-hosting sites during training and evaluation. Most of the images have been removed, while the company continues investigating how the agents behaved outside their intended boundaries.

Advertisement
The string of loss-of-control incidents has injected intense momentum into global demands for strict AI regulation.
OpenAI says AI agents leaked 53 user images to external sites.

OpenAI’s AI agents are back in the news again. The AI firm has confirmed that agents accidentally leaked 53 user-submitted images onto external image-hosting websites. Since the July Hugging Face incident and recent reports of AI breaching Australian health sites, rogue AI agents have become a hot topic online. Now, this time, for some, it touched something deeply personal, photos some users may never have expected to see outside their own chats.

advertisement

In a blog post, OpenAI said it has "identified 53 instances to date where user-provided images were posted" via links that were never meant to be publicly discoverable. In simple words, the links were not openly searchable or visible on the public internet, but anyone who had the link could potentially access the image. The company hasn't said whether the images were AI-generated or showed real people, or when they first surfaced online. Most have since been taken down, and OpenAI says it's pushing hosting providers to remove the rest.

How did this happen

OpenAI relies on anonymised user data to help train its models, data that's supposed to be scrubbed of names, metadata, and anything else that could identify users. But that process isn't foolproof, and researchers, speaking to Reuters, warned there's always a chance a stray detail slips through, or that an agent stumbles onto something it shouldn't while doing its work.

Rogue AI-agents are becoming a modern-day problem that keeps resurfacing in new forms. Back in July, OpenAI's AI agents broke out of their intended boundaries and hacked into Hugging Face, a popular AI code repository, while hunting for answers to a test. Since then, a steady drip of similar stories has emerged: OpenAI-linked agents have been spotted poking around U.S. government websites, including the SEC and Census Bureau, and appear to have made a failed attempt to breach a Department of Education civil-rights site.

Researchers at Transluce also linked OpenAI agents to an attempted break-in at an Australian government health portal, and to hijacking an obscure German wiki page to swap tips on dodging the company's own safety guardrails.

Reuters reported that, by one insider's count, OpenAI has now logged roughly two dozen instances of its agents behaving in unexpected or undesirable ways, a number that keeps climbing as the company digs through months of activity logs.

The episodes have rattled the wider AI industry: Anthropic, Google, and Meta have all since reported finding similar rogue-agent behaviour of their own. Even OpenAI CEO Sam Altman and Anthropic's Dario Amodei have publicly urged the industry to slow down, though both companies still shipped new models this week.

- Ends