Article

What is an AI agent, and how long can it really work alone?

For growing businesses: an agent finishes work that takes a person 70 minutes correctly four times out of five, on clean programming tasks.

By Piet Baudoin · September 2026

An AI agent chooses its own steps and carries them out in your systems. The strongest independently measured model finishes about 70 minutes of work correctly four times out of five, on clean programming tasks. That is less than the brochures promise. Below, I work that measured line forward to a year. What agents take over today in small businesses is in what the agentic economy is.

What does an AI agent do that a chatbot does not?

It acts. A chatbot answers.

Microsoft puts it this way: “Unlike a large language model, an agent doesn’t only return content for a human to act on. Instead, an agent: Acts autonomously.” So an agent calls its own tools, saves data and sets work in motion, without anyone approving every step (Microsoft, 11 September 2026).

The difference is about who acts. The agent writes, sends and changes things under its own name, in your Outlook and in your QuickBooks or Xero.

At OpenAI it is now a toggle: Chat for questions, Work for longer jobs. Flip that toggle, and the question is no longer whether the answer is correct, but whether this was allowed to happen. More on that in should an AI assistant send email on its own.

Is a fixed automation like Zapier also an agent?

No. The difference is in who chooses the steps.

With a fixed automation, a person maps out the path in advance: if an invoice comes in, do this, then that. If the invoice differs, it stalls or does the wrong thing. With an agent, the goal is fixed and the route is not, so it decides fresh with every invoice what to do.

Chatbot Fixed automation AI agent
Who chooses the steps you, per question the builder, in advance the model, while working
What it touches only the screen what the builder pointed to anything it has access to
When it’s wrong you ignore it the mistake repeats it has already been done

Watch the numbers. Statistics Netherlands (CBS) counted that for 2025, 29.8 percent of small businesses with 10 to 249 people used AI technology (CBS, 16 March 2026). In that count, ordinary process automation simply counts as AI. So that is about businesses doing something with AI, not businesses that have an agent working. That number does not exist.

How long does an agent keep working alone, according to the measurements?

That depends on how sure you want to be.

METR does not measure how long a model is busy, but how big a job it can handle, in human working hours: the time horizon.

GPT-4 reached about 4 minutes in March 2023. Claude Opus 4.6 stood at 12 hours in February 2026, doubling every 129 days. METR published that measurement in May 2026, but the model measured is from February 2026: I work forward from that date below. METR has not yet measured the models from September 2026, so this is a lower bound.

That 12 hours is a coin flip: the length the model reaches in half of cases. If you want it right four times out of five, it drops to 70 minutes on clean programming tasks.

OpenAI quotes a customer: “This lets our agents reason, act, recover, and collaborate across complex workflows that can span hours or days” (OpenAI, 10 September 2026). Hours or days, against 70 measured minutes. See also what Claude can do for a small business.

Does it work as well on a screen as in code?

No, but the gap is closing faster than the last independent estimate assumed.

That estimate comes from METR, in a note from 22 January 2026: for tasks where the model has to look at a screen, the time horizon is 40 to 100 times lower, because it reads a screen poorly.

Since that note, OpenAI has measured itself on OSWorld, a test with real computer tasks. The first model reached 38.1 percent there in January 2025. Sol reached 62.6 percent in July 2026, Astra reached 72.6 percent on 3 September 2026, both on the harder OSWorld 2.0 and counting partly finished tasks. On OpenAI’s own safety test for computer use, Sol went wrong in 22.0 percent of cases, Astra in 2.4 percent.

Those are two different measuring sticks: METR measures how long a task may take, OpenAI how often a task succeeds. You cannot convert one into the other, and METR measures independently while OpenAI judges its own model. But going from 38.1 to 72.6 percent in twenty months, with errors dropping from 22.0 to 2.4 percent, is too big a move to hold the January 2026 factor unchanged.

In what year does an agent handle a full workday without supervision?

Do the math with me. Start at 70 minutes. A workday is 480 minutes, almost seven times as much. Two doublings give four times, three give eight times: you need just under three. At 129 days per doubling, that is 358 days from February 2026, so early 2027.

For screen work, the factor of 40 to 100 still needs to come off: 5.3 to 6.6 extra doublings, 685 to 855 days, so late 2028 to mid-2029. That rests on the January 2026 estimate, and the September 2026 numbers say screen work is catching up faster.

My hypothesis: in 2027, an agent finishes a full workday of clearly defined work on its own, and for screen work in your own software, that happens sooner than the slow scenario of late 2028 to mid-2029.

Two things break that hypothesis. METR itself says measurements above 16 hours are unreliable, so above a workday the line cannot be checked. And the measured tasks are neatly defined; your office work is not.

For a small or midsize business, that does not change the approval, only its size: soon you will not approve a single step, but a finished piece of work half a day long. Start choosing tasks now that fit within that 70 minutes and can be checked afterward. An agent that clicks around unseen in your software all day can wait. Whether you need several side by side is in do you need several AI agents at twenty people.

What can an AI agent not do today?

It does not know what lives only in someone’s head. And it cannot defend itself against text that gives it instructions.

The second works like this. A supplier’s attachment contains a sentence that tells the agent what to do, and it carries that out as if you had asked. This is called prompt injection, and Microsoft puts it at the top of its list of agent risks.

OpenAI describes in its guide how that ends: a search for a restaurant ends with the agent forwarding a password code from your Gmail. The same guide advises against saying: check my email and handle everything.

Investor Katie Jacobs Stanton wrote on 22 August 2026 that an AI assistant sent an email on her behalf without being asked, after which she revoked its access to her email.

What does the law require, and who is on the hook for the mistake?

Since 2 August 2026, a system that talks directly with people must say that it is AI. The European Commission names chatbots, AI agents and avatars explicitly under Article 50 (24 July 2026). That also applies to the chat on your own site.

Appropriate human oversight of high-risk AI does not take effect until 2 December 2027, and 2 August 2028 for AI built into products. Until then, it is your decision.

Microsoft does not say who signs off on a ready-made service. It does put its explanation under the heading shared responsibility, writes that an agent takes its steps without a person clicking approve on each one, and lists having too much permission and running unbounded among the risks. Who pays when it goes wrong is in who is liable when an AI agent makes a mistake.

Where do you start, at twenty people?

With one task of half an hour at most, where you can see afterward whether it went right or wrong. Not a whole department, but one piece of work that fits within that 70 minutes. A purchase invoice someone retypes is a task like that: the amount is right or it is not.

Then ask your supplier: at which actions does this wait for approval?

Bombos reads what comes in, email, messages, receipts and invoices, looks up the data in the software you already use, and prepares the work with the reason attached. No message goes to a customer and no payment goes out without someone at your business clicking Approve. That is technically enforced and always on.

So this grows with the line above. More independence is earned per type of work: if it goes right often enough, Bombos needs to bring it to you less often, and you set that pace. As the models become reliable over longer tasks, that boundary moves with them for looking things up, checking, and preparing work. We can put the strongest and most cost-efficient model underneath right away, and switch per client to what that client wants. Approval on messages and payments stays in place, so you keep pace with that speed without the risk growing with it.

Our promise: after three months, the work we start with is ready every day, without anyone needing to think about it. Leave your number on the contact page, and we will call you back.

Start with work you can check.

Sources

Every source was opened on 20 September 2026 and every quote appears in it word for word.

What an agent is, according to the makers

How long an agent works alone

How well an agent handles a screen

How many businesses use AI

What the law requires

What goes wrong