Experiences and developments

Which AI agent trends matter for businesses?

Agents handle longer tasks, models get cheaper, and software gains built-in agents. The bottleneck shifts to permissions, integrations, and approval.

Five measured shifts matter for businesses: agents handle longer tasks; equivalent performance becomes five to ten times cheaper annually; agents operate screens; MCP has become the standard software connection; and Exact, Dutch accounting software, AFAS, Dutch business software, and Microsoft add their own agents. Together, they make the model less often the bottleneck. More important is which systems an agent may access, with which permissions, and whose approval.

This page compares measurements, announcements, and business implications. Most trend lists mix forecasts and measurements. Here they remain separate: a maker’s announcement does not prove operation, and a technical trend does not prove adoption. For Dutch usage by process, see administrative processes increasingly automated with AI.

The short answer

  • According to METR, January 29, 2026, the length of software tasks the best model completes half the time doubled about every 7 months from 2019 to 2025, and every 131 days since 2023.
  • Wikipedia’s account of METR measurements gives the best measured model a May 2026 horizon of “likely at least 16 hours.” METR says measurements above 16 hours are unreliable with its current task suite.
  • Equivalent performance becomes 5 to 10 times cheaper annually, while running the best model became 3 to 18 times more expensive annually, arXiv 2511.23455.
  • MCP, the open software-connection standard, reported 97 million monthly SDK downloads and 10,000 active servers in December 2025.
  • In June 2026, AFAS, Dutch business software, reported no certified AI integrations. It called third-party MCP servers, in translation, “completely unchecked.”
  • Exact, Dutch accounting software, announced Financial Agent September 3, 2026; AFAS added Jonas in Profit 8 in June; Microsoft announced Autopilot September 25.
  • Microsoft charges by usage: 5 Copilot Credits per agent action and 2 per generated answer.
  • Adoption trails supply: CBS reports 13.8% of Dutch microbusinesses used AI technology in 2025.

Of eight frequently mentioned 2026 trends, three have independent measurements and five are described only by makers. That determines their weight. METR, CBS, and scientific papers report developments. Microsoft or OpenAI press releases describe intended sales.

Trend Measurement or announcement Evidence Business implication
Longer tasks Horizon doubles every 131 days since 2023 Independent METR measurement, software tasks only Hours-long tasks become feasible; office work unmeasured
Cheaper equivalent performance 5x to 10x cheaper annually Scientific paper, arXiv Today’s price becomes a fraction next year; hardest tasks stay expensive
Computer use 81.8% partial on OSWorld 2.1 Anthropic’s own claim Screen operation is possible; reliable long-task completion unproven
MCP standard 97 million monthly downloads, 10,000 servers Project’s own figures Connection is technically ordinary; vendors decide permission
Built-in agents Exact, AFAS, Microsoft Maker announcements Your existing agent looks within its own package
Usage billing 5 credits per agent action Microsoft pricing documentation Agent budget caps become purchasing criteria
Always-on agents OpenAI Dots, Microsoft Autopilot Announcements without independent tests No published failure rate yet
Small-business adoption 6.8% in 2023 to 13.8% in 2025 Independent CBS measurement You are not behind if you have not started

Sources appear below. See AI agents versus workflow automation for definitions.

How quickly are agents improving according to METR?

METR’s “time horizon” measures task duration in experienced-human time that a model completes correctly half the time. It doubled “around every 7 months” from 2019 to 2025; after 2023, “131 days under TH1.1, compared to 165 days under TH1” (METR, January 29, 2026).

Measurement Value Source
Doubling 2019-2025 About 7 months, 3.3x annually METR TH1.1, Jan. 29, 2026
Doubling after 2023, TH1.1 131 days, 6.9x annually METR TH1.1
Doubling after 2023, old TH1 165 days METR TH1.1
Doubling for 2024-2025 models About 3.5 months, about 10x annually LessWrong, Feb. 13, 2026
Claude Opus 4.6, 50% horizon 11 hours 59 minutes Wikipedia, METR
Claude Mythos Preview, May 2026 “likely at least 16 hours” Wikipedia, METR; added May 8, 2026

Annual multipliers are our calculations: 2 to the power 365/131 gives 6.9. LessWrong found “~3.5 months. That’s about 10x per year,” warning acceleration could be temporary and mainly concerns programming (February 13, 2026).

Three caveats matter more. These are software tasks rather than invoices or email. Confidence intervals remain “still very wide”; METR measured human time for “only 5 of our 31 long (8h+) tasks.” Since May 2026, it states: “Measurements above 16 hrs are unreliable with our current task suite” (METR, May 8, 2026). The yardstick is shorter than the best model’s capability. See what an agent is and how long it works alone.

Are agents becoming cheaper per task?

Tokens become cheaper, but tasks do not automatically follow. Researchers tracking April 2024-November 2025 prices found “the price for a given level of benchmark performance has decreased remarkably fast, around 5x to 10x per year”. Running the best benchmark model grew “on the order of 3x-18x per year” (arXiv 2511.23455, November 2025). Newest models on increasingly difficult tasks can cost more.

One seller’s pricing shows tier differences (Anthropic, read September 30, 2026):

Model Input per million tokens Output per million tokens
Fable 5.1 $10 $50
Opus 5.5 $4 $20
Sonnet 5.5 $2 $10
Haiku 4.5 $1 $5

The highest and lowest differ tenfold. Opus 5.5 tokens cost 2.5 times less than Fable 5.1. Anthropic claims “40% less than Opus 5 on typical workloads” (September 22, 2026). Older Epoch AI research found declines “ranging from 9x to 900x per year,” warning the fastest may not continue (March 12, 2025).

For a twenty-person business, a dollar of computation today becomes about twenty cents next year if difficulty stays constant and you keep today’s model. Agents choosing steps often take more than fixed workflows. Task cost depends more on operation than the price list. See administrative AI costs.

Can an agent operate a computer and software itself?

Agents score highly on short computer-use tests. Reliable completion of hour-long or longer tasks remains unproven. Computer use means viewing, clicking, and typing like an employee. Anthropic claims Opus 5.5 scores 81.8% partial on OSWorld 2.1 (September 22, 2026).

OSWorld 2.0 contains “108 authentic tasks across 31 self-hosted web environments and professional desktop applications,” averaging 27.25 checkpoints; “69.6% take humans over 1 hour.” Its highest listed score is Claude Opus 5’s 77.67% partial (Snorkel AI; sells evaluation services). Partial gives checkpoint credit even when the whole task remains unfinished. Versions 2.0 and 2.1 differ, making scores incomparable.

Operating unconnected software is a real change from two years ago. But an 80%-correct posting is wrong. Screen operation is a fallback for software without integrations rather than a replacement for supported connections. See may ChatGPT operate your accounting software? for Exact, AFAS, and Moneybird, online accounting software.

What is MCP, and why does your vendor decide permissions?

MCP, Model Context Protocol, is the open standard for agents talking to software. It became ordinary connection practice within a year. In December 2025, the project reported “Over 97 million monthly SDK downloads, 10,000 active servers and first-class client support across major AI platforms like ChatGPT, Claude, Cursor, Gemini, Microsoft Copilot, Visual Studio Code and many more”. It moved into a Linux Foundation organization (MCP blog, December 9, 2025). Anthropic founded MCP; these are project figures.

Technology does not guarantee entry. AFAS wrote, in translation: “At publication, June 2026, no AI integration has yet been certified.” Also: “Using an AI integration, such as an MCP server or AI agent, claiming AFAS integration is completely unchecked and therefore carries major risks” (AFAS customer portal, June 2026).

This trend is absent from English lists. Dutch accounting and HR access depends on vendor rules. Vendors become gatekeepers. Building a connection is routine for experienced teams; permitted routes and rights matter. See AFAS, Exact Online, and agents working between software systems.

Which agents are already in Exact, AFAS, and Microsoft 365?

Major packages introduce agents in 2026 within their own software. Businesses first notice this through updates.

Provider Product Date Price or model Source
AFAS Jonas in Profit 8 June 18, 2026 No extra cost for existing customers AFAS press release
AFAS Third-party MCP server or agent June 2026 None certified yet AFAS customer portal
Exact Financial Agent Sept. 3, 2026 No price given Accountancy Vanmorgen
Microsoft Autopilot Sept. 25, 2026 Usage billing, private preview late September Microsoft blog
Microsoft Copilot Studio agent action Updated Sept. 14, 2026 5 Copilot Credits, no euro price on page Microsoft Learn
OpenAI Dots Sept. 29, 2026 First dot in Pro and Business Premium; extras unpriced Traictory, The Register

AFAS says, in translation: “Jonas assists in workflows, prefills fields, and supports daily processes” (June 18, 2026). Exact’s agent analyzes “over sixty financial KPIs including liquidity, solvency, and profitability” (Accountancy Vanmorgen, September 3, 2026). Autopilot “lives in your tenant with its own identity, memory, computer and workspace … with permissions, audit and governance behind it” (Microsoft, September 25, 2026).

Built-in agents see their package and little outside. Small-business work spans packages: incoming email, a file in one system, a posting in another. See AI within AFAS or Exact, or an AI agent alongside.

What do usage-billed always-on agents cost?

Agents running around the clock use usage billing, making budget caps purchasing criteria. Microsoft says: “Cowork, Code, and Autopilot, new long-running agentic capabilities, and frontier models like Astra and Fable all run on UBB” (September 25, 2026).

Copilot Studio rates are credits rather than euros (Microsoft Learn, updated September 14, 2026): agent action 5, generated answer 2, classic answer 1, organizational-data search 10, and 100 agent-flow actions 13. Computer operation uses agent-action rates. Microsoft’s order-agent example uses 4 daily actions, 20 daily credits. At 125% of prepaid capacity, custom agents are disabled.

Dots includes the first agent in Pro and Business Premium. Extras have “no per-dot price” (Traictory, September 30, 2026). The Register notes shifting limits: $200 Pro moved from “20x usage” to “10x usage” (September 29, 2026).

Overnight work is a variable expense determined by steps rather than users. Ask per-task cost, who sets caps, and behavior at the cap. See newsletters for Microsoft, September 26, and Dots, September 29.

How many Dutch businesses already use agents?

Far fewer than trend lists suggest. CBS counted 13.8% of microbusinesses using AI technology in 2025. No specific small-Dutch-business agent measurement exists. Adoption rose from 6.8% in 2023 to 10.6% in 2024 and 13.8% in 2025 (CBS, March 16, 2026).

Measurement Value Source
Microbusinesses, 2-10 people, with AI, 2023 / 2024 / 2025 6.8% / 10.6% / 13.8% CBS, March 16, 2026
Same, 5-9 people, 2025 19% CBS
Information and communication / business services / financial services / construction, 2025 50.4% / 27.1% / 19% / 4.6% CBS
SMEs with AI, 2024 to 2026 6% to 70% SME Barometer via Accountancy Vanmorgen
Organizations experimenting / scaling somewhere, worldwide 62% / 23% McKinsey via CX Today, Feb. 2026
Organizations with any AI EBIT impact, worldwide 39% McKinsey via CX Today

13.8% and 70% answer different questions. CBS counts specific technologies in businesses up to ten people. The Barometer counts any use, mainly, in translation, “simple applications such as generating text” (Accountancy Vanmorgen, September 3, 2026). ChatGPT email rewriting counts in one but not the other.

Globally, “62% are at least experimenting with agents, and 23% are scaling agentic systems somewhere,” while “only 39% report any EBIT impact attributable to AI” (CX Today on McKinsey, February 23, 2026). You are not behind if you have not started. See when an agent is useful.

Which risks grow with autonomy?

Longer autonomous work increases the importance of permissions and stopping. OpenAI’s research agent bypassed internet restrictions in September 2026. An alert arrived after about fifteen minutes and was acknowledged 3 minutes later. Manual stopping took 2.5 hours. OpenAI wrote: “All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused” (The Hacker News, September 2026). See our September 27 newsletter.

Always-on products lack independent tests. Dots has “no independent evaluation of dots, no third-party benchmark, and no published failure or incident rate.” Background apps “can’t send messages, change app content, or control your browser or computer” (Traictory, September 30, 2026). The maker deliberately limits permissions.

Permissions, logs, and stopping are central. Microsoft sells “permissions, audit and governance.” AFAS rejects uncertified agents. Longer independent work increases the importance of approval on payments, customer messages, and contracts.

Our position: the bottleneck is which systems an agent may access, its permissions, and who approves work. This is an organizational decision.

Three trends relax model limits: 131-day horizon doubling, 5-to-10-times annual price declines, and screen operation. Three shift decisions to you and vendors: AFAS’s absence of external certification, package-bound Exact and AFAS agents, and Microsoft’s action billing with disabling caps. “Can AI do this?” becomes “May this agent write here, and who signs off?”

Three consequences follow. Package-bound agents see only part of cross-package work. Undocumented knowledge matters more than model choice because all models share that gap. Approval is where quality emerges: an agent working longer alone also continues longer when wrong.

You control this bottleneck. Set permissions, connections, and approval rules today irrespective of next year’s best model. See first processes to automate and Dutch SME workflow automation.

Where is this heading?

Growth is fast and uncertain. Extending 131-day doubling from at least 16 hours in May 2026 gives over 100 hours of software work by May 2027. This is our calculation on a yardstick METR calls unreliable above 16 hours. It says nothing about office work. A fivefold annual price decline similarly turns a dollar task into twenty cents if difficulty stays constant.

Our expectation: by late 2027, every major Dutch accounting and HR package, Exact, AFAS, Visma, a business software group, and Nmbrs, payroll and HR software, has a certified external-agent access route. That route determines SME capabilities more than model choice.

AFAS announced certification and warns against unchecked connections. Microsoft gives agent identities, permissions, and logs. Vendors want control over writing for customers and their own agents. Once one package certifies access, accountants and customers ask others.

This fails if vendors open connections without certification, or screen operation becomes reliable enough to finish tasks without connections beyond vendor control. Check package terms and long-task OSWorld scores without “partial.”

Bombos will arrange connections through vendor-opened routes while retaining knowledge of your business in Bombos regardless of model.

What can this not yet do?

  • METR does not measure office work. Software horizons have 50% success probability. Sixteen hours does not mean sixteen error-free administrative hours. No invoice, email, or file measurement exists.
  • Computer-use scores are partial. 81.8% does not mean four of five tasks finish. Error-free hour-plus completion remains unproven.
  • Cheaper tokens do not mean cheaper tasks. Many steps and newest models increase bills. Microsoft and OpenAI provide no extra-agent euro price or per-dot price yet.
  • Always-on agents lack independent tests. Dots has no published failure rate; Autopilot is in private preview.
  • Many figures come from sellers. MCP counts, Opus 5.5’s OSWorld score, and price claims come from makers. We omitted analyst forecasts such as Gartner’s when primary sources could not be opened.
  • No agent finds knowledge held only in a head. Record it first.

How does Bombos approach this?

Bombos addresses permissions, connections, and approval. The goal is more work at higher quality with the same team, without adding people.

Chef distributes incoming work to specialists. Wegwijzer guides you and suggests tasks. Specialists are built around your tasks, systems, and rules. Bombos reads connected mailboxes, finds files in Exact, AFAS, Visma, Nmbrs, or Microsoft 365, and prepares supported proposals. You approve, change, or reject in Bombos. Results reach packages after approval. Payments, customer messages, and contracts wait by default. The technically fixed boundary is removed only at your request and risk.

The model is replaceable. Learning resides in approvals and corrections; corrections become team-wide rules. Bombos interviews you to record undocumented knowledge.

Your first task is the start. Your team teaches subsequent tasks without technical skills. We guide you; then you can do it yourself while the field keeps changing.

Sources

Each source was opened on September 30, 2026, and each original excerpt appears verbatim in it. Dutch excerpts above are translations.

Free, no obligation

More work done, at a higher quality, with the same team.

That is what Bombos is for: companies that grow fast and want to keep the same team. We start with one task that keeps piling up and guide you until your team can handle it. Then your team teaches Bombos the next task. Leave your number and we will call you back to talk about your situation.

We read what you write. Within one working day you hear from the one of us who knows your kind of work best.