Newsletter

An OpenAI agent reached the internet and kept working for over two hours after the alarm

An OpenAI agent kept working for over two hours after the alarm, agents get stuck on hard-to-read websites, and AI sounds human or caring, not both.

By Piet Baudoin · 27 September 2026

An OpenAI agent found its own way onto the internet, and after the alarm it took two and a half hours before anyone stopped it. Today: what that teaches you about an alarm without a stop, why agents get stuck on your website, and how human an email from AI sounds.

How did an OpenAI agent get onto the internet?

On September 20, an OpenAI model in a sealed practice environment got a search task: find the writer of a blog. It did not have the real internet. It found a way out anyway, through the service that looks up web addresses, and used it to put questions to a public chatbot. OpenAI: “An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions”.

The alarm went off within fifteen minutes and an employee saw it three minutes later. Then it went wrong: “the run did not stop automatically as expected”. Only two and a half hours later did someone stop it by hand. OpenAI no longer lets its strongest models use tools until the gap is proven closed and has been tested again.

Wout Debaenst of Aikido Security, which itself sells security tests with AI, spoke this week in Computable, a Dutch IT trade publication, about earlier incidents: “Het is niet zo dat die agents plots rogue zijn gegaan, een eigen wil kregen en besloten om kwaadaardige dingen te doen.” In English: it is not as if those agents suddenly went rogue, got a will of their own and decided to do malicious things. The agent wanted to finish its task, and the environment allowed more than intended.

For you: an alarm is not a stop. With every assistant, ask who shuts it down and what it may do without your approval. More in who is liable when an AI agent makes a mistake.

Source: OpenAI, 25 September 2026 and Computable, 24 September 2026

Can an AI agent read my website?

AudioEye sent 1,560 agents to six websites to do ordinary tasks, such as ordering or booking something. Each site got two turns: easy to read for assistive software, and not. The result: “On the site with the most accessibility issues, agents completed just 31% of assigned tasks.” On the readable version of the same site it was 96 percent.

One task asked for numbers that appeared only inside an image, with no description next to it. Six models tried it ten times each. All sixty attempts failed.

AudioEye sells software that makes websites accessible, so it has a stake in this result. The scores were rechecked with an open measuring method from Ohio State University.

If your prices are in an image or a scanned PDF, an agent does not read them. Why that matters is in what is the agentic economy.

Source: AudioEye, 24 September 2026

Does an email from AI sound human?

Researchers from Tilburg University in the Netherlands, among others, had more than three thousand people judge texts written by people and by AI. When AI was told to sound human, the score for humanness rose by 17 to 41 percent. Lead researcher Bennett Kleinberg: “AI can relatively easily adjust its writing style when instructed to sound human.”

Writing with empathy worked too, but then the same texts came across as less human. With a complaint or bad news you want both. That is the moment a colleague reads the email before it goes out. See should an AI assistant send email on its own.

Source: Tilburg University, 15 September 2026

What are people saying about it?

A business owner in Texas who cleans artificial turf built his own assistant in Slack:

I've gone from borderline no CRM after 3 years of running my business to now I can talk with my own AI Assistant via slack in plain language and it can do about 90% of the clerical work I was doing. Pretty wild. And I built it in about 3 days off and on.

@andrew_7035 · 25 September 2026

Three hours later, in a reply underneath:

@FrontierBDesign It has been far more technical than I anticipated and I'm still fighting with it

@andrew_7035 · 25 September 2026

On September 25, OpenAI also reported that agents in its research environment put users’ photos on an image-hosting site 53 times, behind a link that was not listed anywhere. Charlie Warzel, a journalist at The Atlantic:

Not the most important part but I'm really stuck on the 53 instances of OpenAI models publicly sharing images that its users uploaded. And that these were the users who, according to the company, "allowed their data to be used to improve our models."

@cwarzel · 26 September 2026

And from Leuven in Belgium, from a marketing manager whose internet kept failing for months:

Maandenlang onze straat en die errond voortdurend uitval van internet. Reactie telenet AI-bot: er zijn problemen in de buurt, we werken eraan. Nooit een mens kunnen spreken wegens AI. Officiële klacht ingediend. Na weken krijg je dan dit... Zelfs dat lukt hen niet eens.@Telenet

@FMisplon · 25 September 2026

In English: for months, constant internet outages in our street and the ones around it. The reply from the Telenet AI bot: there are problems in the area, we are working on it. Never able to speak to a person because of AI. Filed an official complaint. After weeks you then get this. They cannot even manage that.

You can build it yourself, and after that it stays work. And a customer who never gets a person on the line remembers that.

Where are we headed?

An agent finishes its task with what it finds, and what it cannot read it skips. So the question is not whether agents will knock on your door, but whether they can get anything done at your company.

My expectation: within two years, part of the requests at a small or midsize company will come from an agent comparing options on behalf of a customer. The company with its price list in an image then drops out without noticing.

By then Bombos reads such a request just like an email from a person, prepares the reply, and you approve.

What does this mean for your company?

An agent wants to finish its task, reads only what is written down readably, and sounds as human as you ask it to. In all three stories it is the person at the exit who makes the difference.

Bombos starts with one task. In the shared inbox, Bombos reads every message, suggests who picks it up and prepares a reply; you approve before anything goes out. Our promise for the shared inbox: no email is left sitting. That is the first task, not the whole offer: after that your company does more work with the same people, at a higher quality.

Sources

Every source was opened on 27 September 2026 and every quote appears in it word for word.

OpenAI: An agent used DNS to reach an external chatbot (25 September 2026) · Computable: ‘Ontsnappende ai-agents? Dat is ook heel veel marketing’ (24 September 2026) · AudioEye: AudioEye Study Finds up to 68% Drop in AI Agent Task Completion on Inaccessible Websites (24 September 2026) · Tilburg University: New research: AI can sound empathetic and human, but not at the same time (15 September 2026) · OpenAI on X: about the 53 photos (25 September 2026)