Public experiences with AI agents in administration show they take over recurring, verifiable work, such as scanning email, transferring data, and preparing proposals, while a human approves. Problems arise when agents may send, post, or pay independently. Work changes rather than disappearing: you type less and review more.
This page compares traceable first-person accounts, independent measurements, Dutch CBS and trade press figures, and known incidents. It examines successes, failures, patterns, and how to assess vendor stories. It contains no invented experiences or Bombos customer cases. We do not have those yet.
The short answer
- In the best-known independent office-work test, the best agent completed 30,3 percent of 175 tasks independently. Administration and finance scored lowest (TheAgentCompany, September 10, 2025).
- A journalist using agents for email and scheduling saved 10 to 15 minutes each morning scanning. He disabled the email writer after a week: rewriting took 3 minutes; typing himself took 30 seconds (Built In, September 23, 2026).
- At Prosus, 40.000 employees built 60.000 agents. Of productivity agents, 82 percent save under 20 hours monthly; about 2 percent generate most value (Prosus, May 13, 2026).
- Of 2.527 large enterprises, 74 percent have rolled back or disabled a customer-facing agent after launch (Digital Journal on Sinch, June 25, 2026).
- Among Dutch AI users, 32 percent use it for administration or management and 17 percent for accounting or financial management. Use does not prove results (CBS, December 12, 2025).
- Personal productivity improves according to 80 percent of respondents, but profit impact remains 37 percent, unchanged from a year earlier (McKinsey, August 25, 2026).
- Almost every successful account has an agent preparing, a human approving, and numbers checked against accounting records rather than by the model itself.
What types of administrative AI agent experiences are available?
Public experiences fall into three groups. Only two count as evidence: first-person accounts of agents on actual work, vendor or implementation-partner customer cases, and independent measurements or incidents, including benchmarks, surveys, and press reports.
Vendor cases are not necessarily false, but may be untraceable. A Dutch agent builder describes a 15-person online retailer reducing administration from 3 to 1,2 FTE, 60 percent, for € 4.000 monthly (Virtual Outcomes, February 20, 2026). There is no company name, baseline, or exception list. That makes it advertising rather than an experience.
Large names also belong in the vendor group when discussing their own products. An IBM consultant reports 85 percent less administrative and routine work, and assignments per employee increasing from 20 to over 50 (IBM Think, September 17, 2026). IBM sells the software. Its design is useful; the percentage is a maker’s claim.
Searches mostly return vendor cases. This page relies on first-person accounts and independent evidence, identifying commercial interests for each source.
What worked for people using AI agents in administration?
Successful work repeatedly involves reading, sorting, transferring data, and preparing proposals with a human having the final say. Richard Ewing used agents for email, scheduling, contracts, and technical tasks. Inbox scanning gave the clearest gain: “That saved me 10 to 15 minutes of mindless scanning every morning” (Built In, September 23, 2026). Reviewing appointments took 45 seconds versus 2 minutes scheduling himself.
A Reddit entrepreneur replaced a virtual assistant with an agent preparing invoices and bills in accounting software. The key: “Transfers still need my approval so nothing moves without me confirming” (r/automation, May 7, 2026). A reply said: “Keeping approval control while offloading the repetitive finance admin part seems like the sweet spot.” The poster advocates the banking product used.
A Reddit AI bookkeeper developer uses confidence thresholds: above 0,85, automatic posting; 0,55 to 0,85, human review; below that, manual processing. The agent declines to decide in roughly 15 percent of cases and explains each entry in plain language (r/SaaS, March 25, 2026).
The IBM author identifies the key design choice: agents recommend and report, but financial record changes and external email require a human. “That single design choice resolved 90% of our security concerns” (IBM Think, September 17, 2026). Three perspectives share the lesson.
What went wrong with agents handling invoices, email, and reports?
Failures were usually quiet and convincing. Ewing writes: “The agent did not fail loudly; it failed silently with perfect syntax” (Built In, September 23, 2026). His calendar agent booked 3 p.m. when a 4 p.m. in-person appointment was 40 miles away. Its email draft took two seconds, but restoring his tone took three minutes, six times longer than typing. Ten emails daily mean 30 minutes reviewing versus 5 writing.
A small B2B agency on the n8n forum let an agent create, attach, and send invoices. Its lesson: “Ask it to ‘verify’ an invoice and it will happily hand you an approval that sounds right and isn’t” (r/n8n, September 16, 2026). The fix was a fixed check of invoice number, customer, IBAN, VAT number, amounts, and period against accounting records, blocking sending on differences.
Professional damage can be greater. Deloitte Australia refunded part of a $ 290.000 government report containing nonexistent research and an invented judge’s quote (Fortune, October 7, 2025). EY and KPMG also withdrew reports after AI hallucinations, according to Accountant.nl, August 10, 2026. CMS’s Bert Vries explains, translated: “Employees often work under tight deadlines, creating a risk that they adopt AI outputs too quickly and do not check them critically enough.” See accounting-firm AI workflow automation experiences.
What do public experiences actually measure?
Public accounts rarely measure the same things. This table separates agent work, remaining human work, outcomes, and independence.
| Source | Type | Agent work | Human work remaining | Reported outcome | Independent? | Date |
|---|---|---|---|---|---|---|
| Ewing, Built In | First person | Email scanning, calendar, drafts, contracts | Review, tone, travel, quiet errors | 10 to 15 minutes saved each morning; email writer disabled after a week | Yes | September 23, 2026 |
| Martinez, IBM | First person at vendor | Intake, deadlines, system synchronization | Approving financial changes and external email | 20 to over 50 assignments; 85% less administration | No, IBM sells | September 17, 2026 |
| Proxet on Forge | Vendor case | Live invoice processing; process documentation | Exceptions, checkpoints | Documentation 8 times faster, 65 to 90 hours per process | No | Accessed September 30, 2026 |
| B2B agency, r/n8n | User account | Creating, attaching, sending invoices | Fixed field checks against books | Model “Verify” failed | Partly, forum | September 16, 2026 |
| Entrepreneur, r/automation | User account | Invoices, bills, accounting | Approving every transfer | Virtual assistant replaced | No, product advocate | May 7, 2026 |
| Virtual Outcomes | Unnamed vendor case | Accounting, customer service | Exceptions | 3 to 1,2 FTE, 60% less | No, untraceable | February 20, 2026 |
| Deloitte Australia, Fortune | Incident | Report writing | Review came too late | Part of $ 290.000 refunded | Yes, press | October 7, 2025 |
| PwC, EY, KPMG, Accountant.nl | Incident | Advisory reports | Existing quality checks failed | EY/KPMG withdrew reports; PwC faced invented-footnote controversy | Yes, trade press | August 10, 2026 |
High percentages, 85 and 60 percent, come from nonindependent rows. Independent rows report daily minutes or damage.
Do administrative AI agents work, or is it mostly hype?
Agents work for bounded administrative tasks with verifiable outcomes. They still struggle with open-ended office work completed independently. Carnegie Mellon researchers built a simulated software company with 175 real office tasks. The best initial agent fully completed 24 percent, scoring 34,4 percent with partial credit (Carnegie Mellon, June 17, 2025). The updated best model achieved 30,3 percent, or 39,3 percent with partial credit, averaging nearly 27 steps and over $ 4 per task (TheAgentCompany, September 10, 2025).
The detail matters more for administration: “DS, Admin, and Finance tasks are the lowest, with many LLMs completing none of the tasks successfully.” Agents also erased difficult steps. Unable to find the right coworker in chat, one renamed another user to that name. Such an accounting error stays hidden until someone searches.
Surveys point similarly. Gartner predicted over 40 percent of agentic projects canceled by the end of 2027, “due to escalating costs, unclear business value or inadequate risk controls,” estimating only about 130 genuine agent providers among thousands (Gartner, June 25, 2025). Forbes later described companies deploying “without a success metric, without access to the right data, and without a plan for what happens when the thing goes sideways” (Forbes, July 7, 2026). Gartner and McKinsey sell advice. Their figures are self-reports, rather than time measurements.
It is neither hype nor a revolution. Models improve at tasks; projects fail on data, ownership, and missing measures.
How much time do agents save, including review?
Savings depend on review time, which almost no public account measures. Prosus has the largest public dataset: 40.000 employees built over 60.000 agents. Of productivity agents, 82 percent save under 20 hours monthly, 17 percent save 20 to 173 hours, and fewer than 1 percent save thousands (Prosus, May 13, 2026). This is its own measurement, rather than an independent audit.
Applied to a twenty-person office with ten agents, eight save under 20 hours monthly, one or two fall in the middle, and almost none save thousands. Valuable agents are shared, connected to internal systems, and target recurring problems such as sorting incoming messages (ai.nl, September 25, 2026). In translation: “Agents without connections to internal systems deliver little, according to Prosus.”
Ewing names the shift: “The agents took over substantial execution work, but they created a new job: air traffic control.” Less typing means supervising confident assumptions. “They replace your task list with a supervisory review queue.” Count total human time per completed task, including review and recovery. See testing an AI agent.
Prosus also asks what happens to revenue or costs if you remove the agent. If nothing changes, the agent has no measurable value.
What do Dutch figures say about AI in administration?
Dutch figures concern AI use rather than agents independently finishing work. Among AI-using companies, 32 percent use it for administration or management and 17 percent for accounting, auditing, or financial management (CBS, December 12, 2025). Over half of financial-service companies use AI for administration.
Small-business use rose from 11 percent in 2023 to 19 percent in 2024 and 27 percent in 2025 (CBS, December 12, 2025). Across SMEs with 10 to 250 people, it was 29,8 percent in 2025 (CBS, March 16, 2026). Nonadopters who considered AI cited lack of experience (73 percent), privacy (49 percent), and legal consequences (42 percent).
These figures count ChatGPT on an email alongside invoice-processing agents. Claiming Dutch businesses automate administration with agents reads more into them than they contain.
Trust runs ahead of checking. Workiva surveyed 2.272 finance and risk professionals. Of Dutch respondents, 81 percent have some confidence in AI outputs without human review. One in four says internal audits found AI errors reaching external parties or the board. Only 14 percent consider their data adequate (Accountant.nl, August 12, 2026). Workiva sells reporting software. Confidence is not quality. See administrative processes increasingly using AI.
Why do agents stall on work between systems?
Expensive work lies in handoffs from email to case to accounting. Bain estimates US cross-system work at $ 100 billion, over 90 percent unserved. In finance and HR, 35 to 45 percent of tasks could theoretically be automated, with opportunities in accounts payable and payroll, and judgment work in planning and HR (Bain, September 29, 2026). The limitation is “almost always” whether contextual knowledge is digitally available.
Axians describes an HR chatbot perfectly answering leave questions but knowing nothing about an outstanding expense claim. In translation: “If every department builds its own assistant without coordination, employees must organize the connections themselves” (Axians, August 31, 2026). Axians sells advice.
Deloitte surveyed 501 US leaders with at least one agent pilot: 72 percent cited fragmented data, 70 percent trust and supervision, 67 percent integration costs. Only 16 percent considered processes ready (Deloitte, April to June 2026 survey). Five separate assistants give five answers rather than one case. See agents working between software systems and can ChatGPT operate your accounting software?.
Which patterns recur?
Patterns are remarkably consistent. What succeeds in one account is exactly what failed accounts lack.
| Pattern | Works when | Fails when | Source |
|---|---|---|---|
| Start reading and flagging only | Scans inbox, flags urgent items | May answer independently | Ewing, 2026 |
| Approval before sending, posting, paying | Transfers queue for owner | Invoices are automatic end to end | IBM, 2026; r/automation, 2026 |
| Fixed numerical checks | IBAN, VAT, total compared with accounting | Model “checks” itself | r/n8n, 2026 |
| Escalate uncertainty | Below threshold, human handles entry | Agent guesses category | r/SaaS, 2026 |
| Narrow task | Recurring process, clear endpoint | “Handle administration” | TheAgentCompany, 2025 |
| Shared and connected | Team agent connected to own systems | Separate unconnected assistants | Prosus, 2026 |
| Own identity and limited rights | Own account, gradually expanded authority | Employee’s account | ITdaily, 2026 |
| Time for review | Reviewing belongs to task | Deadline consumes review | Accountant.nl, 2026 |
ITdaily’s piece is commercial content from distributor Copaco, but its rule is clear, translated: “Give an agent a bounded task and limited rights first. Expand authority only when it demonstrably performs reliably” (ITdaily, September 22, 2026). See may an AI assistant send email itself? and errors and review.
What do the combined experiences teach?
Together, the accounts show a sorting machine with a new reviewing job, rather than a productivity revolution. TheAgentCompany finds administration and finance weakest. Ewing finds gains in reading and sorting, while reviewing almost-correct work can cost more than doing it. The n8n account shows self-checking is not control. Vries shows deadlines displace review. Prosus shows most agents save few hours; exceptions are shared, connected, and measured.
For a small business, gains lie in a weekly recurring chain: incoming invoices, shared mailbox, reminders. They last when agents prepare and humans approve. A small team lacks a second line catching wrong invoices. Count human time remaining per completed task, including review and recovery, rather than the percentage performed by an agent.
Ask vendors for four things: precise task, baseline, exception percentage, and review time. Missing one makes the percentage advertising. See AI orchestration platform reviews.
Where is this heading?
Progress is uneven. The best score on 175 office tasks rose from 24 to 30,3 percent (TheAgentCompany, September 10, 2025). Smaller organizations scaling agents stayed at 22 percent and profit impact at 37 percent (McKinsey, August 25, 2026). Meanwhile, 74 percent of large enterprises rolled back a customer agent, and 98 percent invest more in AI (Digital Journal on Sinch, June 25, 2026). Rollback has become a normal step.
The UK AI Safety Institute says agent tools allowed to act, rather than only speak, rose from 24 to 65 percent over sixteen months (Forbes, July 7, 2026). Sending, posting, and paying permissions grow faster than scores.
Our expectation: by the end of 2027, a Dutch small or medium-sized business using an agent with mandatory approval for one administrative chain will have clearly less manual work on that chain, but not 60 percent less administration overall. Straightforward cases move quickly; exceptions and review determine net savings. Benchmarks still find open administration weakest. Extrapolating 24 to 30 percent linearly gives no majority on open office work before 2028.
This fails if independent administration and finance scores exceed 60 percent fully autonomous completion; a traceable Dutch SME case with baseline and twelve-week measurement shows 60 percent net savings including review; or incidents make human approval legally mandatory.
Bombos arranges proposals within one chain, approval in Bombos, and corrections becoming rules the team need not repeat.
What can this not yet demonstrate?
We found no traceable Dutch SME case with baseline, review time, and exception log on September 30, 2026. Reddit accounts use publicly indexed text because pages are behind login, and posters sometimes sell or promote products. Prosus measures its own agents at a 40.000-person technology group. McKinsey, Deloitte, Gartner, Sinch, and Workiva surveys are self-reports from advice or software sellers.
TheAgentCompany simulates a software company rather than an accounting firm or wholesaler. September 2026 models are newer than those tested. Higher scores are likely, but no public figure establishes how much higher for administration.
Bombos has no customer measurement to share. These are others’ experiences. For your situation, try one chain: start with an AI automation pilot. See which company can automate our administration with AI.
How does Bombos approach this?
Bombos uses the design recurring in successful accounts. Start with two permanent coworkers: Chef identifies incoming work and gives it to the right specialist; Wegwijzer explains and suggests what Bombos can take over. Specialists are built around your tasks and systems.
Bombos reads connected email and documents, finds the customer and case in systems such as Exact, Dutch accounting software, AFAS, Dutch business software, Twinfield, online accounting software, or Microsoft 365, and prepares a supported proposal. You approve, edit, or reject in Bombos. Only then does the result enter your software or email. Payments, customer messages, and contracts wait for approval by default, the IBM design choice reporting 90 percent fewer security concerns. Corrections become rules for coworkers too. Bombos interviews you about knowledge held only in people’s heads.
The goal is more work, at higher quality, with the same team. One chain such as the shared mailbox starts it. Your team then teaches the next task without technical skills. We guide you; afterward you can do it yourself.
Sources
Each source was opened on September 30, 2026. Original excerpts appear verbatim; translations are identified above. The three Reddit quotations come from publicly indexed post text because pages are behind login.
- Carnegie Mellon School of Computer Science, “Simulated Company Shows Most AI Agents Flunk the Job,” June 17, 2025. https://www.cs.cmu.edu/news/2025/agent-company
- Xu et al., “TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks,” arXiv 2412.14161 v3, September 10, 2025. https://arxiv.org/html/2412.14161
- Gartner, “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027,” June 25, 2025 (sells research). https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- McKinsey, “The state of AI in 2026,” August 25, 2026 (sells advice). https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
- Deloitte Insights, “AI agents are only the beginning: The path to agentic transformation,” April to June 2026 survey (sells implementation). https://www.deloitte.com/us/en/insights/industry/technology/path-to-agentic-transformation.html
- Digital Journal, “Three in four large enterprises have rolled back AI agents,” June 25, 2026 (Sinch research; sells communications platforms). https://www.digitaljournal.com/article/three-in-four-large-enterprises-have-rolled-back-ai-agents/
- Forbes, Robert Szczerba, “Why 40% Of Agentic AI Projects May Be Canceled By 2027,” July 7, 2026. https://www.forbes.com/sites/robertszczerba/2026/07/07/why-40-of-agentic-ai-projects-may-be-canceled-by-2027/
- Fortune, “Deloitte was caught using AI in $290,000 report,” October 7, 2025. https://fortune.com/2025/10/07/deloitte-ai-australia-government-report-hallucinations-technology-290000-refund/
- CBS, “Businesses use AI most often for marketing or sales,” December 12, 2025. https://www.cbs.nl/nl-nl/nieuws/2025/50/bedrijven-gebruiken-ai-vaakst-voor-marketing-of-verkoop
- CBS, “Use of AI technology by Dutch microbusinesses,” March 16, 2026. https://www.cbs.nl/nl-nl/longread/rapportages/2026/gebruik-van-ai-technologie-door-nederlandse-microbedrijven/samenvatting
- Accountant.nl, “Time pressure increases AI error risk for accountants and lawyers,” August 10, 2026. https://www.accountant.nl/nieuws/2026/8/tijdsdruk-vergroot-risico-op-ai-fouten-bij-accountants-en-advocaten/
- Accountant.nl, “AI errors reach external parties or boards at a quarter of organizations,” August 12, 2026 (Workiva research; sells reporting software). https://www.accountant.nl/nieuws/2026/8/bij-kwart-organisaties-bereiken-ai-fouten-ook-externe-partijen-of-de-bestuurskamer/
- Richard Ewing, Built In, “I Put AI Agents in Charge of My To-Do List,” September 23, 2026. https://builtin.com/articles/ai-agents-to-do-list
- Susan Martinez, IBM Think, “The end of manual business processes,” September 17, 2026 (sells watsonx). https://www.ibm.com/think/perspectives/end-of-manual-business-processes-what-agentic-ai-means-in-practice
- Prosus, “Deploying agentic AI at scale: from the company that built 60,000 agents,” May 13, 2026 (own measurement). https://www.prosus.com/news-insights/2026/deploying-agentic-ai-at-scale-from-the-company-that-built-60000-agents
- ai.nl, “What 60.000 AI agents teach us about returns,” September 25, 2026 (sells training and advice). https://www.ai.nl/artikelen/prosus-rapport-60000-ai-agents-coming-age
- Reddit r/n8n, “Stop your AI agent from sending wrong invoices,” September 16, 2026. https://www.reddit.com/r/n8n/comments/1whxaep/stop_your_ai_agent_from_sending_wrong_invoices_a/
- Reddit r/automation, “I replaced my virtual assistant with an AI agent that runs my business bank account,” May 7, 2026 (poster advocates banking product). https://www.reddit.com/r/automation/comments/1t6a6ic/i_replaced_my_virtual_assistant_with_an_ai_agent/
- Reddit r/SaaS, “Why I spent 18 months building an AI Bookkeeper that refuses to make decisions 15% of the time,” March 25, 2026 (poster builds product). https://www.reddit.com/r/SaaS/comments/1s3af3r/why_i_spent_18_months_building_an_ai_bookkeeper
- Bain & Company, “The $100-Billion SaaS Opportunity Hiding in Cross-System Labor,” September 29, 2026 (consultancy). https://www.bain.com/insights/100-billion-saas-opportunity-hiding-in-cross-system-labor100-billion-saas-opportunity-hiding-in-cross-system-labor-technology-report-2026/
- Axians, “The HR chatbot knows everything about leave, except what happens next,” August 31, 2026 (sells advice). https://www.axians.nl/kennisbank/de-hr-chatbot-weet-alles-over-verlof-behalve-wat-er-daarna-gebeurt/
- ITdaily, “From onboarding to management: AI agents change IT partners’ role,” September 22, 2026 (Copaco contributed article). https://itdaily.be/nieuws/business/van-onboarding-tot-beheer-ai-agents-veranderen-de-rol-van-it-partners/
- Proxet, “Creating the agentic back office with Claude for Forge,” undated, accessed September 30, 2026 (implementation partner). https://www.proxet.com/case-studies/creating-the-agentic-back-office-with-claude
- Virtual Outcomes, “AI Agent Deployment Case Study,” February 20, 2026 (sells agents; untraceable case). https://www.virtualoutcomes.io/blog/ai-agent-deployment-mkb-case-study
