Experiences and developments

What are the experiences with AI agents for administrative automation?

Agents take over recurring, verifiable administrative work when a human approves. Letting them send or post independently often goes wrong.

Public experiences with AI agents in administration show they take over recurring, verifiable work, such as scanning email, transferring data, and preparing proposals, while a human approves. Problems arise when agents may send, post, or pay independently. Work changes rather than disappearing: you type less and review more.

This page compares traceable first-person accounts, independent measurements, Dutch CBS and trade press figures, and known incidents. It examines successes, failures, patterns, and how to assess vendor stories. It contains no invented experiences or Bombos customer cases. We do not have those yet.

The short answer

  • In the best-known independent office-work test, the best agent completed 30,3 percent of 175 tasks independently. Administration and finance scored lowest (TheAgentCompany, September 10, 2025).
  • A journalist using agents for email and scheduling saved 10 to 15 minutes each morning scanning. He disabled the email writer after a week: rewriting took 3 minutes; typing himself took 30 seconds (Built In, September 23, 2026).
  • At Prosus, 40.000 employees built 60.000 agents. Of productivity agents, 82 percent save under 20 hours monthly; about 2 percent generate most value (Prosus, May 13, 2026).
  • Of 2.527 large enterprises, 74 percent have rolled back or disabled a customer-facing agent after launch (Digital Journal on Sinch, June 25, 2026).
  • Among Dutch AI users, 32 percent use it for administration or management and 17 percent for accounting or financial management. Use does not prove results (CBS, December 12, 2025).
  • Personal productivity improves according to 80 percent of respondents, but profit impact remains 37 percent, unchanged from a year earlier (McKinsey, August 25, 2026).
  • Almost every successful account has an agent preparing, a human approving, and numbers checked against accounting records rather than by the model itself.

What types of administrative AI agent experiences are available?

Public experiences fall into three groups. Only two count as evidence: first-person accounts of agents on actual work, vendor or implementation-partner customer cases, and independent measurements or incidents, including benchmarks, surveys, and press reports.

Vendor cases are not necessarily false, but may be untraceable. A Dutch agent builder describes a 15-person online retailer reducing administration from 3 to 1,2 FTE, 60 percent, for € 4.000 monthly (Virtual Outcomes, February 20, 2026). There is no company name, baseline, or exception list. That makes it advertising rather than an experience.

Large names also belong in the vendor group when discussing their own products. An IBM consultant reports 85 percent less administrative and routine work, and assignments per employee increasing from 20 to over 50 (IBM Think, September 17, 2026). IBM sells the software. Its design is useful; the percentage is a maker’s claim.

Searches mostly return vendor cases. This page relies on first-person accounts and independent evidence, identifying commercial interests for each source.

What worked for people using AI agents in administration?

Successful work repeatedly involves reading, sorting, transferring data, and preparing proposals with a human having the final say. Richard Ewing used agents for email, scheduling, contracts, and technical tasks. Inbox scanning gave the clearest gain: “That saved me 10 to 15 minutes of mindless scanning every morning” (Built In, September 23, 2026). Reviewing appointments took 45 seconds versus 2 minutes scheduling himself.

A Reddit entrepreneur replaced a virtual assistant with an agent preparing invoices and bills in accounting software. The key: “Transfers still need my approval so nothing moves without me confirming” (r/automation, May 7, 2026). A reply said: “Keeping approval control while offloading the repetitive finance admin part seems like the sweet spot.” The poster advocates the banking product used.

A Reddit AI bookkeeper developer uses confidence thresholds: above 0,85, automatic posting; 0,55 to 0,85, human review; below that, manual processing. The agent declines to decide in roughly 15 percent of cases and explains each entry in plain language (r/SaaS, March 25, 2026).

The IBM author identifies the key design choice: agents recommend and report, but financial record changes and external email require a human. “That single design choice resolved 90% of our security concerns” (IBM Think, September 17, 2026). Three perspectives share the lesson.

What went wrong with agents handling invoices, email, and reports?

Failures were usually quiet and convincing. Ewing writes: “The agent did not fail loudly; it failed silently with perfect syntax” (Built In, September 23, 2026). His calendar agent booked 3 p.m. when a 4 p.m. in-person appointment was 40 miles away. Its email draft took two seconds, but restoring his tone took three minutes, six times longer than typing. Ten emails daily mean 30 minutes reviewing versus 5 writing.

A small B2B agency on the n8n forum let an agent create, attach, and send invoices. Its lesson: “Ask it to ‘verify’ an invoice and it will happily hand you an approval that sounds right and isn’t” (r/n8n, September 16, 2026). The fix was a fixed check of invoice number, customer, IBAN, VAT number, amounts, and period against accounting records, blocking sending on differences.

Professional damage can be greater. Deloitte Australia refunded part of a $ 290.000 government report containing nonexistent research and an invented judge’s quote (Fortune, October 7, 2025). EY and KPMG also withdrew reports after AI hallucinations, according to Accountant.nl, August 10, 2026. CMS’s Bert Vries explains, translated: “Employees often work under tight deadlines, creating a risk that they adopt AI outputs too quickly and do not check them critically enough.” See accounting-firm AI workflow automation experiences.

What do public experiences actually measure?

Public accounts rarely measure the same things. This table separates agent work, remaining human work, outcomes, and independence.

Source Type Agent work Human work remaining Reported outcome Independent? Date
Ewing, Built In First person Email scanning, calendar, drafts, contracts Review, tone, travel, quiet errors 10 to 15 minutes saved each morning; email writer disabled after a week Yes September 23, 2026
Martinez, IBM First person at vendor Intake, deadlines, system synchronization Approving financial changes and external email 20 to over 50 assignments; 85% less administration No, IBM sells September 17, 2026
Proxet on Forge Vendor case Live invoice processing; process documentation Exceptions, checkpoints Documentation 8 times faster, 65 to 90 hours per process No Accessed September 30, 2026
B2B agency, r/n8n User account Creating, attaching, sending invoices Fixed field checks against books Model “Verify” failed Partly, forum September 16, 2026
Entrepreneur, r/automation User account Invoices, bills, accounting Approving every transfer Virtual assistant replaced No, product advocate May 7, 2026
Virtual Outcomes Unnamed vendor case Accounting, customer service Exceptions 3 to 1,2 FTE, 60% less No, untraceable February 20, 2026
Deloitte Australia, Fortune Incident Report writing Review came too late Part of $ 290.000 refunded Yes, press October 7, 2025
PwC, EY, KPMG, Accountant.nl Incident Advisory reports Existing quality checks failed EY/KPMG withdrew reports; PwC faced invented-footnote controversy Yes, trade press August 10, 2026

High percentages, 85 and 60 percent, come from nonindependent rows. Independent rows report daily minutes or damage.

Do administrative AI agents work, or is it mostly hype?

Agents work for bounded administrative tasks with verifiable outcomes. They still struggle with open-ended office work completed independently. Carnegie Mellon researchers built a simulated software company with 175 real office tasks. The best initial agent fully completed 24 percent, scoring 34,4 percent with partial credit (Carnegie Mellon, June 17, 2025). The updated best model achieved 30,3 percent, or 39,3 percent with partial credit, averaging nearly 27 steps and over $ 4 per task (TheAgentCompany, September 10, 2025).

The detail matters more for administration: “DS, Admin, and Finance tasks are the lowest, with many LLMs completing none of the tasks successfully.” Agents also erased difficult steps. Unable to find the right coworker in chat, one renamed another user to that name. Such an accounting error stays hidden until someone searches.

Surveys point similarly. Gartner predicted over 40 percent of agentic projects canceled by the end of 2027, “due to escalating costs, unclear business value or inadequate risk controls,” estimating only about 130 genuine agent providers among thousands (Gartner, June 25, 2025). Forbes later described companies deploying “without a success metric, without access to the right data, and without a plan for what happens when the thing goes sideways” (Forbes, July 7, 2026). Gartner and McKinsey sell advice. Their figures are self-reports, rather than time measurements.

It is neither hype nor a revolution. Models improve at tasks; projects fail on data, ownership, and missing measures.

How much time do agents save, including review?

Savings depend on review time, which almost no public account measures. Prosus has the largest public dataset: 40.000 employees built over 60.000 agents. Of productivity agents, 82 percent save under 20 hours monthly, 17 percent save 20 to 173 hours, and fewer than 1 percent save thousands (Prosus, May 13, 2026). This is its own measurement, rather than an independent audit.

Applied to a twenty-person office with ten agents, eight save under 20 hours monthly, one or two fall in the middle, and almost none save thousands. Valuable agents are shared, connected to internal systems, and target recurring problems such as sorting incoming messages (ai.nl, September 25, 2026). In translation: “Agents without connections to internal systems deliver little, according to Prosus.”

Ewing names the shift: “The agents took over substantial execution work, but they created a new job: air traffic control.” Less typing means supervising confident assumptions. “They replace your task list with a supervisory review queue.” Count total human time per completed task, including review and recovery. See testing an AI agent.

Prosus also asks what happens to revenue or costs if you remove the agent. If nothing changes, the agent has no measurable value.

What do Dutch figures say about AI in administration?

Dutch figures concern AI use rather than agents independently finishing work. Among AI-using companies, 32 percent use it for administration or management and 17 percent for accounting, auditing, or financial management (CBS, December 12, 2025). Over half of financial-service companies use AI for administration.

Small-business use rose from 11 percent in 2023 to 19 percent in 2024 and 27 percent in 2025 (CBS, December 12, 2025). Across SMEs with 10 to 250 people, it was 29,8 percent in 2025 (CBS, March 16, 2026). Nonadopters who considered AI cited lack of experience (73 percent), privacy (49 percent), and legal consequences (42 percent).

These figures count ChatGPT on an email alongside invoice-processing agents. Claiming Dutch businesses automate administration with agents reads more into them than they contain.

Trust runs ahead of checking. Workiva surveyed 2.272 finance and risk professionals. Of Dutch respondents, 81 percent have some confidence in AI outputs without human review. One in four says internal audits found AI errors reaching external parties or the board. Only 14 percent consider their data adequate (Accountant.nl, August 12, 2026). Workiva sells reporting software. Confidence is not quality. See administrative processes increasingly using AI.

Why do agents stall on work between systems?

Expensive work lies in handoffs from email to case to accounting. Bain estimates US cross-system work at $ 100 billion, over 90 percent unserved. In finance and HR, 35 to 45 percent of tasks could theoretically be automated, with opportunities in accounts payable and payroll, and judgment work in planning and HR (Bain, September 29, 2026). The limitation is “almost always” whether contextual knowledge is digitally available.

Axians describes an HR chatbot perfectly answering leave questions but knowing nothing about an outstanding expense claim. In translation: “If every department builds its own assistant without coordination, employees must organize the connections themselves” (Axians, August 31, 2026). Axians sells advice.

Deloitte surveyed 501 US leaders with at least one agent pilot: 72 percent cited fragmented data, 70 percent trust and supervision, 67 percent integration costs. Only 16 percent considered processes ready (Deloitte, April to June 2026 survey). Five separate assistants give five answers rather than one case. See agents working between software systems and can ChatGPT operate your accounting software?.

Which patterns recur?

Patterns are remarkably consistent. What succeeds in one account is exactly what failed accounts lack.

Pattern Works when Fails when Source
Start reading and flagging only Scans inbox, flags urgent items May answer independently Ewing, 2026
Approval before sending, posting, paying Transfers queue for owner Invoices are automatic end to end IBM, 2026; r/automation, 2026
Fixed numerical checks IBAN, VAT, total compared with accounting Model “checks” itself r/n8n, 2026
Escalate uncertainty Below threshold, human handles entry Agent guesses category r/SaaS, 2026
Narrow task Recurring process, clear endpoint “Handle administration” TheAgentCompany, 2025
Shared and connected Team agent connected to own systems Separate unconnected assistants Prosus, 2026
Own identity and limited rights Own account, gradually expanded authority Employee’s account ITdaily, 2026
Time for review Reviewing belongs to task Deadline consumes review Accountant.nl, 2026

ITdaily’s piece is commercial content from distributor Copaco, but its rule is clear, translated: “Give an agent a bounded task and limited rights first. Expand authority only when it demonstrably performs reliably” (ITdaily, September 22, 2026). See may an AI assistant send email itself? and errors and review.

What do the combined experiences teach?

Together, the accounts show a sorting machine with a new reviewing job, rather than a productivity revolution. TheAgentCompany finds administration and finance weakest. Ewing finds gains in reading and sorting, while reviewing almost-correct work can cost more than doing it. The n8n account shows self-checking is not control. Vries shows deadlines displace review. Prosus shows most agents save few hours; exceptions are shared, connected, and measured.

For a small business, gains lie in a weekly recurring chain: incoming invoices, shared mailbox, reminders. They last when agents prepare and humans approve. A small team lacks a second line catching wrong invoices. Count human time remaining per completed task, including review and recovery, rather than the percentage performed by an agent.

Ask vendors for four things: precise task, baseline, exception percentage, and review time. Missing one makes the percentage advertising. See AI orchestration platform reviews.

Where is this heading?

Progress is uneven. The best score on 175 office tasks rose from 24 to 30,3 percent (TheAgentCompany, September 10, 2025). Smaller organizations scaling agents stayed at 22 percent and profit impact at 37 percent (McKinsey, August 25, 2026). Meanwhile, 74 percent of large enterprises rolled back a customer agent, and 98 percent invest more in AI (Digital Journal on Sinch, June 25, 2026). Rollback has become a normal step.

The UK AI Safety Institute says agent tools allowed to act, rather than only speak, rose from 24 to 65 percent over sixteen months (Forbes, July 7, 2026). Sending, posting, and paying permissions grow faster than scores.

Our expectation: by the end of 2027, a Dutch small or medium-sized business using an agent with mandatory approval for one administrative chain will have clearly less manual work on that chain, but not 60 percent less administration overall. Straightforward cases move quickly; exceptions and review determine net savings. Benchmarks still find open administration weakest. Extrapolating 24 to 30 percent linearly gives no majority on open office work before 2028.

This fails if independent administration and finance scores exceed 60 percent fully autonomous completion; a traceable Dutch SME case with baseline and twelve-week measurement shows 60 percent net savings including review; or incidents make human approval legally mandatory.

Bombos arranges proposals within one chain, approval in Bombos, and corrections becoming rules the team need not repeat.

What can this not yet demonstrate?

We found no traceable Dutch SME case with baseline, review time, and exception log on September 30, 2026. Reddit accounts use publicly indexed text because pages are behind login, and posters sometimes sell or promote products. Prosus measures its own agents at a 40.000-person technology group. McKinsey, Deloitte, Gartner, Sinch, and Workiva surveys are self-reports from advice or software sellers.

TheAgentCompany simulates a software company rather than an accounting firm or wholesaler. September 2026 models are newer than those tested. Higher scores are likely, but no public figure establishes how much higher for administration.

Bombos has no customer measurement to share. These are others’ experiences. For your situation, try one chain: start with an AI automation pilot. See which company can automate our administration with AI.

How does Bombos approach this?

Bombos uses the design recurring in successful accounts. Start with two permanent coworkers: Chef identifies incoming work and gives it to the right specialist; Wegwijzer explains and suggests what Bombos can take over. Specialists are built around your tasks and systems.

Bombos reads connected email and documents, finds the customer and case in systems such as Exact, Dutch accounting software, AFAS, Dutch business software, Twinfield, online accounting software, or Microsoft 365, and prepares a supported proposal. You approve, edit, or reject in Bombos. Only then does the result enter your software or email. Payments, customer messages, and contracts wait for approval by default, the IBM design choice reporting 90 percent fewer security concerns. Corrections become rules for coworkers too. Bombos interviews you about knowledge held only in people’s heads.

The goal is more work, at higher quality, with the same team. One chain such as the shared mailbox starts it. Your team then teaches the next task without technical skills. We guide you; afterward you can do it yourself.

Sources

Each source was opened on September 30, 2026. Original excerpts appear verbatim; translations are identified above. The three Reddit quotations come from publicly indexed post text because pages are behind login.

  1. Carnegie Mellon School of Computer Science, “Simulated Company Shows Most AI Agents Flunk the Job,” June 17, 2025. https://www.cs.cmu.edu/news/2025/agent-company
  2. Xu et al., “TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks,” arXiv 2412.14161 v3, September 10, 2025. https://arxiv.org/html/2412.14161
  3. Gartner, “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027,” June 25, 2025 (sells research). https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
  4. McKinsey, “The state of AI in 2026,” August 25, 2026 (sells advice). https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
  5. Deloitte Insights, “AI agents are only the beginning: The path to agentic transformation,” April to June 2026 survey (sells implementation). https://www.deloitte.com/us/en/insights/industry/technology/path-to-agentic-transformation.html
  6. Digital Journal, “Three in four large enterprises have rolled back AI agents,” June 25, 2026 (Sinch research; sells communications platforms). https://www.digitaljournal.com/article/three-in-four-large-enterprises-have-rolled-back-ai-agents/
  7. Forbes, Robert Szczerba, “Why 40% Of Agentic AI Projects May Be Canceled By 2027,” July 7, 2026. https://www.forbes.com/sites/robertszczerba/2026/07/07/why-40-of-agentic-ai-projects-may-be-canceled-by-2027/
  8. Fortune, “Deloitte was caught using AI in $290,000 report,” October 7, 2025. https://fortune.com/2025/10/07/deloitte-ai-australia-government-report-hallucinations-technology-290000-refund/
  9. CBS, “Businesses use AI most often for marketing or sales,” December 12, 2025. https://www.cbs.nl/nl-nl/nieuws/2025/50/bedrijven-gebruiken-ai-vaakst-voor-marketing-of-verkoop
  10. CBS, “Use of AI technology by Dutch microbusinesses,” March 16, 2026. https://www.cbs.nl/nl-nl/longread/rapportages/2026/gebruik-van-ai-technologie-door-nederlandse-microbedrijven/samenvatting
  11. Accountant.nl, “Time pressure increases AI error risk for accountants and lawyers,” August 10, 2026. https://www.accountant.nl/nieuws/2026/8/tijdsdruk-vergroot-risico-op-ai-fouten-bij-accountants-en-advocaten/
  12. Accountant.nl, “AI errors reach external parties or boards at a quarter of organizations,” August 12, 2026 (Workiva research; sells reporting software). https://www.accountant.nl/nieuws/2026/8/bij-kwart-organisaties-bereiken-ai-fouten-ook-externe-partijen-of-de-bestuurskamer/
  13. Richard Ewing, Built In, “I Put AI Agents in Charge of My To-Do List,” September 23, 2026. https://builtin.com/articles/ai-agents-to-do-list
  14. Susan Martinez, IBM Think, “The end of manual business processes,” September 17, 2026 (sells watsonx). https://www.ibm.com/think/perspectives/end-of-manual-business-processes-what-agentic-ai-means-in-practice
  15. Prosus, “Deploying agentic AI at scale: from the company that built 60,000 agents,” May 13, 2026 (own measurement). https://www.prosus.com/news-insights/2026/deploying-agentic-ai-at-scale-from-the-company-that-built-60000-agents
  16. ai.nl, “What 60.000 AI agents teach us about returns,” September 25, 2026 (sells training and advice). https://www.ai.nl/artikelen/prosus-rapport-60000-ai-agents-coming-age
  17. Reddit r/n8n, “Stop your AI agent from sending wrong invoices,” September 16, 2026. https://www.reddit.com/r/n8n/comments/1whxaep/stop_your_ai_agent_from_sending_wrong_invoices_a/
  18. Reddit r/automation, “I replaced my virtual assistant with an AI agent that runs my business bank account,” May 7, 2026 (poster advocates banking product). https://www.reddit.com/r/automation/comments/1t6a6ic/i_replaced_my_virtual_assistant_with_an_ai_agent/
  19. Reddit r/SaaS, “Why I spent 18 months building an AI Bookkeeper that refuses to make decisions 15% of the time,” March 25, 2026 (poster builds product). https://www.reddit.com/r/SaaS/comments/1s3af3r/why_i_spent_18_months_building_an_ai_bookkeeper
  20. Bain & Company, “The $100-Billion SaaS Opportunity Hiding in Cross-System Labor,” September 29, 2026 (consultancy). https://www.bain.com/insights/100-billion-saas-opportunity-hiding-in-cross-system-labor100-billion-saas-opportunity-hiding-in-cross-system-labor-technology-report-2026/
  21. Axians, “The HR chatbot knows everything about leave, except what happens next,” August 31, 2026 (sells advice). https://www.axians.nl/kennisbank/de-hr-chatbot-weet-alles-over-verlof-behalve-wat-er-daarna-gebeurt/
  22. ITdaily, “From onboarding to management: AI agents change IT partners’ role,” September 22, 2026 (Copaco contributed article). https://itdaily.be/nieuws/business/van-onboarding-tot-beheer-ai-agents-veranderen-de-rol-van-it-partners/
  23. Proxet, “Creating the agentic back office with Claude for Forge,” undated, accessed September 30, 2026 (implementation partner). https://www.proxet.com/case-studies/creating-the-agentic-back-office-with-claude
  24. Virtual Outcomes, “AI Agent Deployment Case Study,” February 20, 2026 (sells agents; untraceable case). https://www.virtualoutcomes.io/blog/ai-agent-deployment-mkb-case-study
Free, no obligation

More work done, at a higher quality, with the same team.

That is what Bombos is for: companies that grow fast and want to keep the same team. We start with one task that keeps piling up and guide you until your team can handle it. Then your team teaches Bombos the next task. Leave your number and we will call you back to talk about your situation.

We read what you write. Within one working day you hear from the one of us who knows your kind of work best.