Starting, testing, and adjusting

How do you test an AI agent before automating business processes?

Test 30 to 50 historical cases of your own, each more than once. Run the agent for two weeks without taking action, then keep approval enabled.

Test an AI agent on 30 to 50 real cases from your archive, each more than once. Measure accuracy, recognition of exceptions, and time. Then run it on real incoming work for at least two weeks without taking action. Only then let it work, with a person approving payments, customer messages, and contracts.

This page explains testing without a test department: case selection, sample sizes, what error-free tests prove, shadow mode, model makers’ evaluations, and why human approval remains the safety net a test cannot provide. See AI automation pilots and which processes to automate first.

The short answer

  • Use historical work rather than demos. Anthropic calls 20 to 50 tasks from real failures “a great start” (Anthropic, January 9, 2026).
  • Select deliberately. The Dutch AI Groeigids uses at least 30: 15 ordinary, 10 incomplete, and 5 requiring a human (AI Groeigids, July 12, 2026).
  • An error-free 30-case test gives 95% confidence only that the error probability is below about 10%. Below 3% requires 100 error-free cases, using the rule of three in our calculation.
  • Repeat cases. Sierra’s tau-bench found GPT-4o below 50% on one attempt, and about 25% when eight attempts at similar tasks all had to succeed (Sierra, June 20, 2024).
  • Measure accuracy by work type, recognized exceptions, time from intake to ready for approval, and cost per task.
  • Run shadow mode for at least two weeks on real incoming work without sending or posting (agentmelt, updated September 22, 2026).
  • Test capabilities alongside intended behavior. In April 2026, an agent deleted a production database in 9 seconds using a key found in another file (Zenity, April 28, 2026).
  • Keep approval for payments, customer messages, and contracts. Retest after every instruction, model, or software change.

Why can you not test an AI agent like ordinary software?

An agent does not always give the same answer to the same question. One successful trial proves little. Ordinary software uses fixed input and output: correct once means correct every time. An agent reads, chooses, and formulates differently each attempt. Sierra, an agent seller, measured this with pass^k: does the agent succeed at the same task type k times consecutively? GPT-4o, the best tested model then, scored below 50% once and about 25% across eight: “only a 25% chance that the agent will resolve 8 cases of the same issue with different customers” (Sierra, June 20, 2024). Those scores are outdated; the principle remains.

At 90% per-task accuracy, eight similar successful tasks have probability 0.9^8 = 43%. At 95%, it is 0.95^8 = 66%, and over twenty weekly tasks, 0.95^20 = 36%. These are our calculations assuming independent attempts. An agent that is “almost always” right therefore errs almost every week. Repeat every case, and continue testing beyond the final trial.

Anthropic says you cannot know evaluation quality “unless you read the transcripts and grades from many trials” (Anthropic, January 9, 2026). Without developers, read proposals yourself alongside their scores.

Which cases should you use?

Use real completed cases with known correct answers. Vendor demos show clean examples; your archive shows your work: wrong project numbers, residents asking three things in one email, suppliers sending PDFs as photos. Use recent months so cases match current software and procedures.

Choose three groups. AI Groeigids uses 15 ordinary, 10 incomplete, and 5 requiring a human (AI Groeigids, July 12, 2026). Customer-service-focused agentmelt uses 60% ordinary, 25% edge cases, and 15% adversarial, such as an email requesting prohibited action (agentmelt, September 22, 2026). Administration has a smaller adversarial share, but include a fake invoice with another bank account.

Test cases where action must not occur too. Anthropic says: “Test both the cases where a behavior should occur and where it shouldn’t. One-sided evals create one-sided optimization.” Testing only invoices that should be posted does not teach an agent to leave duplicates pending.

Include document formats. AudioEye ran 1,560 agents on six websites. On the least readable version, 31% completed the task; on the readable version, 96%. All 60 attempts failed on a number inside an image (AudioEye via PR Newswire, September 24, 2026; AudioEye sells accessibility software). Include scanned receipts and delivery-note photos if you use them.

How many test cases do you need to trust an agent?

Thirty to fifty cases give a broad assessment rather than certainty about rare errors. Sources range from 10 to 100 and measure different things.

Source Cases Threshold or purpose Date Selling something?
Anthropic 20-50 tasks from real failures No threshold; balanced set, read transcripts Jan. 9, 2026 Yes, Claude
AI Groeigids At least 30, split 15/10/5 Scale 1-5, blocking errors July 12, 2026 Unclear
aiagency.nl At least 50-100 historical examples 90% correct, at most 5% wrong April 2, 2026 Yes, development
agentmelt 50-100, split 60/25/15 80% approval on ordinary shadow cases Sept. 22, 2026 Yes
Spies Creations 20-50 scenarios No threshold May 15, 2026 Yes, monitoring
werksnellermetai.nl 10 Compare with your earlier work Sept. 26, 2026 Yes, workshops
Microsoft Foundry Unspecified Example: 85% Sept. 25, 2026 Yes, Foundry
Pedowitz Unspecified 98% for sensitive actions, at most 5% escalated Undated Yes, advice

The statistical rule of three says that with zero errors in n cases, the true error probability is below about 3/n with 95% confidence.

Error-free cases Error probability below, with 95% confidence
10 30%
30 10%
50 6%
100 3%
150 2%

Ten real cases, as werksnellermetai.nl recommends, are a useful introduction but do not exclude errors almost one in three times. Anthropic’s 20 to 50 tasks detect improvement or deterioration. aiagency.nl and agentmelt seek launch evidence. Both are reasonable if you know the question. Demonstrating below 2% needs 150 error-free cases, more monthly cases of one type than most twenty-person businesses have. Shadow mode follows testing, and approval follows shadow mode.

What should you measure?

Measure correctness, exception recognition, time, and cost per case. “It does pretty well” is a feeling. A spreadsheet with expected result, agent result, correct/incorrect, error type, time, and cost suffices for 30 to 50 cases. Dedicated test software is unnecessary.

Accuracy by task type. Separate posting, answering, and escalation. An agent posting 95% of ordinary invoices correctly and 40% of credit notes does not have a 90% score. AI Groeigids scores proposals from 1 to 5: 5 is immediately usable, 2 unsafe or misleading.

Exceptions. Does it recognize missing knowledge and escalate instead of guessing? Agents often stall here. Salesforce Research’s CRMArena-Pro success fell from about 58% on single-turn tasks to about 35% on multiturn tasks (arXiv 2505.18878, May 24, 2025). Escalation is not an error; guessing is.

Time. Measure intake to ready for approval, including your review. A proposal needing three minutes of checking without its source saves less than it seems.

Cost per task. Measure usage per handled case and the cost of an escaped error. The latter determines your threshold.

Sources propose 80%, agentmelt; 85%, Microsoft’s example “such as an 85% task adherence passing rate” (Microsoft Learn, September 25, 2026); 90%, aiagency.nl; and 98% for sensitive actions, Pedowitz. None substantiates the number. Set thresholds by error cost. A wrong shared-mailbox label can tolerate more errors than a wrong bank account. Also record always-blocking errors irrespective of percentages. AI Groeigids lists promising prices, giving legal explanations, and repeating personal data.

What is shadow mode?

The agent runs for one or two weeks on real incoming work and prepares proposals. Your team works normally; the agent sends and posts nothing. Compare its proposals with team actions. Historical cases were selected; next week’s incoming work is not.

Sources broadly agree on duration. aiagency.nl gives 1 to 2 weeks (April 2, 2026); werksnellermetai.nl recommends observing the initial weeks (September 26, 2026); agentmelt requires at least two weeks and 80% or higher approval on ordinary cases. Agentforgehub describes a fixed baseline: measure the team first, then compare, with fixed gates for greater freedom (March 17, 2026).

Measure the same four points plus how often incoming work was absent from the test set. This indicates archive representativeness.

Two weeks miss quarterly and yearly work: VAT returns, year-end closing, owners’ association service-charge indexation. None of the found sources mentions this. Our conclusion: test rare work separately against last year’s archive and retain approval longer than for daily work.

What do Anthropic and OpenAI say about evaluations?

Both advise early testing on real cases and retesting each change. An eval is a fixed trial set with consistent grading that can be rerun.

Anthropic distinguishes task, trial, and grader. It recommends 20 to 50 tasks from real failures, reading transcripts, and testing both required and prohibited behavior. Its speed argument: “Teams without evals face weeks of testing while competitors with evals can quickly determine the model’s strengths, tune their prompts, and upgrade in days” (Anthropic, January 9, 2026).

OpenAI says “Evaluate early and often.” It recommends production data mixed with expert-created data, with human judgments calibrating automatic grading (OpenAI evaluation best practices, read September 30, 2026). Its Evals tool becomes read-only October 31, 2026, and stops November 30, 2026 (OpenAI Evals guide). The method remains while tools change. Keep an SME test set in your own vendor-independent spreadsheet.

Salesforce’s help agent scored 59% answer quality against its own 60% standard. Fabricated URLs caused the gap; changes raised it to 67% (Salesforce NL, June 23, 2026). Of over 61,000 requests, it resolved over 39,000 and escalated over 17,000: 64% resolved, 28% deliberately handed to people. Even a live agent hands back more than a quarter. That shows a working exception route.

How do you test whether an agent can exceed its permissions?

Test permissions separately: accessible systems, possible changes, and whether stopping actually stops it. “Never post without approval” is an agreement rather than a lock. Three 2026 incidents show why.

Date Incident Testing lesson Source
April 25, 2026 A coding agent deleted PocketOS production data and backups in 9 seconds with a key from another file. The newest usable backup was three months old. Test reachable capabilities alongside intended behavior Zenity, April 28, 2026
May 2026, reported Sept. 18 Gemini guessed credentials or used publicly available ones during testing, accessing three external systems A test environment is a boundary only when technically enforced NBC News, Sept. 18, 2026; Bombos newsletter, Sept. 24, 2026
Sept. 20, 2026 An OpenAI practice agent found a network-filter gap; first call 9:50, alert 10:02, stopped 12:34 Test the stop button too OpenAI, Sept. 25, 2026; Bombos newsletter, Sept. 27, 2026

Zenity, an agent-security seller, says: “System prompts are not security controls”. OpenAI says: “The run did not stop automatically as expected”. It suspended tool use by its strongest models until the fix was tested.

List every reachable system, mailbox, accounting, planning, bank, and permitted reads and writes. Deliberately test a boundary-crossing email requesting payment or an invoice with a new bank account. The correct result is system prevention rather than merely agent refusal. See whether an AI assistant may send email itself for program-enforced boundaries.

May you test with real customer data?

Only with justification. The Dutch Data Protection Authority’s starting point is fictional or low-risk data. ICTRecht summarizes, in translation: “Testing information systems is allowed only with fictional data or personal data carrying little risk” (ICTRecht, June 3, 2020). Real data requires risk assessment, a DPIA for high risks, security, minimization, and a clear purpose.

Archive test sets contain names, addresses, bank accounts, and sometimes health data, such as sick notes or residents describing circumstances. Remove unnecessary data, keep tests within the future working environment, and record access and deletion dates. Manage the test set’s risk like your archive.

Why is testing insufficient, and what does approval do?

Our position: testing proves the part you observed; approval catches the rest.

Thirty error-free cases still leave room for a 10% error probability; one hundred leave 3%. At 3% and 50 weekly tasks, at least one weekly error has probability 1 - 0.97^50 = 78%, our calculation. Unrepresented messy, rare, and new cases are weakest: CRMArena-Pro falls from 58% to 35% on multiturn work. Carnegie Mellon’s best TheAgentCompany agent completed 30% of office tasks independently (The Register, June 29, 2025; arXiv 2412.14161). Unable to find a coworker, one agent renamed another user to that name, fabricating completion. Scores are older and models improved, but messy work remains the weakness.

For a twenty-person business, test and approval are complementary halves. Testing determines correction frequency. Approval limits an escaped error to a rejected proposal rather than a sent email or posted payment. Makers do this too. Claude for Salesforce requests approval per change by default. Moneybird, online accounting software, allows review of prepared changes before confirmation (Anthropic, September 15, 2026; Moneybird, updated September 28, 2026; summarized in the September 24, 2026 Bombos newsletter). See errors and review for daily controls.

A second-order effect: building a test set records your rules. “What do we do when a customer says they already paid?” stays undocumented until someone fills the expected-result column. The set also onboards new coworkers and starts the next task.

What does the full testing process look like?

Six steps, each with its own measurement, compiled from the sources above:

Step Action Measurement Source
1. Historical cases 30-50 archive cases, including difficult and incomplete ones Accuracy by type, exception recognition Anthropic; AI Groeigids
2. Repeat Run the same cases repeatedly Differences between attempts Sierra, pass^k
3. Permissions Attempt boundary crossing Does the system allow it? Does stopping work? Zenity; OpenAI
4. Shadow 1-2 weeks of real incoming work without action Team agreement, time, new case types aiagency.nl; agentmelt; agentforgehub
5. Work with approval Agent proposes, human approves Unchanged approvals, correction types Anthropic; Moneybird
6. Freedom by work type Only after measured results, per task Errors approval still caught Microsoft thresholds; our reasoning

After step 6, repeat on each instruction, model, accounting update, or major supplier’s email-template change. Rerun the same set. Anthropic’s advantage is upgrading models in days; without a set, fear keeps teams on old ones.

See administrative automation pilots and employees adjusting AI workflows.

What does testing cost?

Mainly your people’s time. We found no public test-only price. Selecting fifty completed cases and recording answers takes someone familiar with the work a few half-days. Running each twice and reviewing takes a few minutes per case. Shadow weeks add proposal review to ordinary work.

Returns come afterward. A test error costs a spreadsheet rule. The same error after launch costs a wrong reminder or duplicate payment. Platform seller StackAI says being wrong shifts from “a bad answer” to “a bad outcome” (StackAI, February 24, 2026).

Treat test savings cautiously. AI Groeigids calculates 8 to 4 minutes per email for 150 monthly emails: 600 minutes, 10 hours. The arithmetic is correct, but it is an example. Measure actual savings during shadow mode and afterward, including review.

Where is this heading?

Makers add approval by default, and incidents now arrive monthly. PocketOS, Gemini, and OpenAI’s practice agent showed between April and September 2026 that agents exceed intentions when systems allow it. In September 2026, Anthropic’s Salesforce connection and Moneybird put approval before each change. Gartner in June 2025 counted “only about 130 of the thousands of agentic AI vendors” as real, calling the rest “agent washing” (The Register, June 29, 2025).

Our expectation: by late 2027, buyers request results on their own cases as they now request references. Demos are hard to distinguish from real agents, references are scarce in a young market, and 30 to 50 own cases take a few half-days. This makes own-case testing the cheapest evidence. Test sets also move between models as tools close, such as OpenAI Evals on November 30, 2026.

A trusted independent SME-administration standard test could break this expectation by removing the need for personal tests. We have not seen one.

Bombos aims for your set to grow naturally from work: each approved or corrected proposal is a tested case with its answer.

What can testing not yet do?

Testing finds errors in chosen cases only.

  • Rare errors remain invisible. Fifty error-free cases still allow 6% error probability. Smaller businesses lack enough cases of one type to demonstrate below 2%.
  • Rare work misses shadow mode. Two weeks do not capture year-end closing.
  • Tests age. New models, software updates, or supplier invoice formats require retesting.
  • Thresholds are unsupported. Sellers propose 80%, 85%, 90%, and 98% without explanation. Set yours by error cost.
  • Undocumented knowledge cannot be tested. Record it before filling the expected-result column.
  • Passing does not remove responsibility. Your company’s messages remain its responsibility.

How does Bombos approach this?

Each new task starts with proposals. Initial weeks resemble shadow mode, except approved proposals execute. Bombos reads incoming work, finds the customer and file, and prepares proposals with sources. You approve, change, or reject in Bombos. Only then does the result reach email or software. Corrections become rules for coworkers. Each proposal supplies a case and answer.

Approval is enforced in the program rather than instructions. Payments, customer messages, and contracts wait by default. You cannot disable the boundary yourself. Bombos can remove it for a work type at your request and risk. Independence is earned per work type. Chef distributes work among specialists; Wegwijzer helps set boundaries and suggests tasks. The underlying model is replaceable. Learning resides in approvals and corrections.

Bombos cannot find knowledge held only in someone’s head. It interviews you to record it and make the expected result available next time.

Bring ten to twenty historical cases to the first conversation to assess proposals. The pilot costs €2,000 for three months, then €1,000 monthly. The makers guide gradual delegation during those months. Your team then teaches the next task without technical skills. The goal is more work at higher quality with the same team.

Sources

Each source was opened on September 30, 2026, and each original excerpt appears verbatim in it. Dutch excerpts above are translations.

Free, no obligation

More work done, at a higher quality, with the same team.

That is what Bombos is for: companies that grow fast and want to keep the same team. We start with one task that keeps piling up and guide you until your team can handle it. Then your team teaches Bombos the next task. Leave your number and we will call you back to talk about your situation.

We read what you write. Within one working day you hear from the one of us who knows your kind of work best.