Starting, testing, and adjusting

How do you start a pilot for administrative automation with AI?

Choose one recurring task, measure for two weeks, agree one metric, owner, and end date, and record them. A thirteen-week implementation plan.

Start an administrative AI automation pilot by choosing one recurring task, measuring current performance for two weeks, writing down the metric that must change, its assessor, and the decision date, and recording those in your vendor agreement alongside data arrangements.

This is a small-business implementation plan: thirteen weeks, task selection, baseline, team roles, and pilot terms. Whether a pilot makes sense and what it proves is covered in can I start with an AI automation pilot?. For costs, see administrative AI automation prices.

The short answer

  • Three months, thirteen weeks, divides into two weeks baseline/setup, eight weeks real work, two full-volume weeks, and one decision week.
  • Do not use AI on real work in the first two weeks. Establish the baseline first (Xact IT).
  • Choose daily reading/retyping work requiring little judgment and measurable within weeks: shared mailbox or purchase invoices.
  • Agree one primary metric in writing, rather than eight dashboard metrics (Xact IT).
  • Include current workers. Successful buyers sourced initiatives from frontline people rather than central labs (MIT NANDA, July 2025, p.20).
  • Agree at least eight items: success criteria, test set, end date, price, data, working methods/configuration, confidentiality, and afterward. Personal-data processing requires an agreement under GDPR Article 28.
  • Only 1 in 4 task-specific pilots reached production: 20 percent piloted, 5 percent implemented (MIT NANDA, p.6). The report blames approach rather than models.

What does a thirteen-week plan look like?

Five phases: baseline/setup weeks 1/2, fully reviewed proposals weeks 3 to 6, sampling/adjustment weeks 7 to 10, full volume weeks 11/12, decision week 13. This combines detailed online plans for businesses without IT departments.

Week Work Who Complete when
0 Sign one-page task, metric, owner, end date, data agreement Owner/vendor Both signed
1-2 Baseline volume, time, errors, completion time; arrange email/software access and record unwritten knowledge Current worker/vendor Two weeks recorded; access works
3-6 AI proposes; human approves, edits, rejects every item. Weekly ten-minute discussion of errors/causes Current worker Corrections decline
7-10 Expert reviews 50 to 100 proposals; measure again Second coworker/owner Sample discussed, bottlenecks resolved
11-12 All task volume through AI, still requiring approval Affected team Two weeks measured under normal pressure
13 Scale, retain, or stop against agreed metric Owner Written decision and follow-up

The sample comes from Arjun Jaggi’s twelve-week plan, 25-07-2026, with experts reviewing 50 to 100 model outputs in week 8. Three outcomes come from SpotDev: “Scale it. Keep it running at this size. Stop it.”

Durations differ mainly by scope:

Source Duration Phases Decision
Clevertech, updated 27-09-2026 4 to 8 weeks Inventory 1 week, prototype 1-2, measurement 4-6, optimization 1-2 Next process afterward
SpotDev 8 to 12 weeks Baseline/criteria 1-2, build/weekly measure 3-8, full volume 9-12 Scale, retain, stop
Jaggi, 25-07-2026 12 weeks Scope/baseline 1-4, build 5-8, evaluate 9-12 Week 12 executive gate
Summit Trails, 16-07-2026 90 days Baseline 30, instrumentation 15, measure 35, decide 10 Data-backed go/no-go
Xact IT 90 days First 2 weeks measurement only One primary measure
Bombos, 26-09-2026 3 months Map email, record knowledge, approve, review After 3 months

Four to eight weeks can cover email without accounting connections. Twelve to thirteen suits work ending in Exact, Dutch accounting software, AFAS, Dutch business software, or Twinfield, online accounting software, including a closing period and peak. Three months is a safe choice for one task: focused but long enough for exceptions.

Which task should you choose first?

Choose daily reading, searching, and retyping with little judgment and measurable changes within weeks. Often this is email manually entering software: the shared mailbox, invoice lines, or photographed timesheets.

Recurrence creates enough cases. Reading/retyping can be prepared by AI. Limited judgment tests correctness rather than difficult decisions. Quick measurability supports the fixed date. “Pick a problem with clear pain and repeat volume” (Data Pilot).

MIT found “some of the most dramatic cost savings we documented came from back-office automation,” with “often delivered faster payback periods” (July 2025, p.20). Among Dutch AI users, 32 percent use it for administration/management (CBS, 12-12-2025).

See which administrative processes first and where to start: map email in an afternoon with current readers and select the greatest retyping burden. Can you count incoming items and time each within two weeks? If not, choose another task. Uncountable work cannot be assessed.

What should you measure, and why two weeks?

Measure weekly volume, time per item, error frequency, and elapsed time to completion. One week can be distorted by vacation or closing.

Xact IT says: “Spend the first two weeks of the 90-day AI pilot measuring current state only. Do not use the tool yet.” SpotDev says “Nothing is built in the first fortnight.” Summit Trails, 16-07-2026 reserves 30 of 90 days. Of ten Dutch/English pages reviewed, Dutch guides mention measurement but none requires a two-week pre-start baseline.

Measurement runs alongside necessary setup: mailbox/package access and unwritten knowledge collection. No extra calendar time is needed. Real work starts week 3.

Select one success metric: “One primary success metric (…) in writing. One. Not a dashboard of eight.” (Xact IT). Shared mailboxes might measure correct assignment time; invoices, time until ready to post. Keep the other three as safeguards: faster with more errors is failure.

See test design and comparisons.

Who participates?

Three roles: owner/director defining success and deciding, current daily worker approving proposals, and person arranging email/software access. The latter may be an office manager or external administrator. No IT department is required.

Xact IT similarly names problem owner, technical administrator, and daily internal champion. See staff adjusting rules. Successful MIT organizations “Sourced AI initiatives from frontline managers, not central labs” and “Benchmarked tools on operational outcomes, not model benchmarks” (p.20). The actual worker is the pilot’s main participant.

Choose the most knowledgeable coworker rather than the least busy. Approvals/corrections teach company rules. The least experienced reviewer teaches the least experienced person’s rules. None of the reviewed plans makes this choice explicit.

Allow baseline/interview time in weeks 1/2, then review and weekly error discussions. Early work runs alongside AI preparation. Tell staff the pilot measures the task rather than the person. A baseline feeling like appraisal produces poor data.

Since February 2, 2025, businesses must address AI literacy for “their staff and other persons dealing with the operation and use of AI systems on their behalf” (AI Act Article 4). Reviewing proposals teaches when they are correct. Record that explanation occurred.

What should a pilot agreement cover?

At least eight items. Without them, disagreement is common. “Define exactly what constitutes a successful pilot, who assesses it, and what consequences are attached to it” (MKB Juristen).

Item Required agreement Source
Success Metric, assessor, consequences MKB Juristen
Test set/baseline Fixed before start; joint changes only Bosin
Duration/end Fixed date Bosin
Price Full-pilot fixed price and subsequent cost Bosin
Data Yours, no pilot-data training, processing agreement Bosin; GDPR Article 28
Methods/configuration Rules, examples, settings you may take Morgan Lewis
Confidentiality Applies even without continuation MKB Juristen
Afterward Optional continuation or purchase on criteria; handover help MKB Juristen/Morgan Lewis

End date. “Every pilot should end on a date, not an event. Open-ended pilots slide into unpaid production use” (Bosin, 2026). Buyers similarly risk silently becoming subscribers. The test set “will be fixed on or before the Pilot Start Date and may be changed only by mutual written agreement.” Vendors cannot move the threshold halfway through.

Data. Customer email and named invoices involve personal data. Processing “shall be governed by a contract or other legal act” specifying subject, duration, nature, and purpose. The processor “deletes or returns all the personal data to the controller” afterward (GDPR Article 28(3), (3g)). This applies from day one. Also agree no training for others on pilot data.

Methods. Pilot value extends beyond data. “Prompts, prompt libraries, workflows, evaluation datasets, retrieval indexes, embeddings, and guardrails created for the customer’s environment” “can be as important to portability as the underlying data” (Morgan Lewis, 25-02-2026). Mailbox rules and explained exceptions are your knowledge. Ask how you receive them at exit. Handover help belongs “as a contractual obligation, not an informal expectation.”

Why paid? “Charging for a pilot does two things: it filters out prospects who aren’t serious, and it forces the customer to get internal budget approval” (Bosin). Without a budget owner, nobody decides in week 13. See preparing a platform request.

How do you decide at the end?

Three outcomes: expand to a new task, retain current operation, or stop. The agreed person decides on the agreed date and metric.

SpotDev supplies those outcomes. Wavect, 16-07-2026 scores 20 to 24 for scaling, 14 to 19 for fixing/continuing, 0 to 13 for stop/pause. Independent hard stops: unacceptable harm, missing audit trail for consequential actions, or no valid baseline. The last prevents anything beyond feelings in week 13.

The middle outcome matters. A working task unsuitable for expansion still succeeded. Expand when a second task resembles the first in reading/retyping and limited judgment. See daily implementation and management.

Keep the decision short: baseline metric, same metric in weeks 11/12, weekly corrections, unauthorized outbound actions. Four lines and a signature.

Why do pilots fail, and what does that mean for starting?

MIT’s 300 public projects, 52 interviews, and 153 surveys found results “does not seem to be driven by model quality or regulation, but seems to be determined by approach” (July 2025, p.3).

Among task-specific evaluators, one third piloted and one quarter of pilots implemented. General assistants need no connections; task-specific AI must fit actual work and software.

AI type Evaluated Piloted Implemented Pilot to implementation
General assistants, ChatGPT/Copilot 80% 50% 40% 4 in 5
Task-specific or built-in 60% 20% 5% 1 in 4

Source: MIT NANDA, p.6. These are “Preliminary Findings” measuring profit-and-loss effects rather than technical success. Treat as direction rather than Dutch SME fact.

Gartner expects “more than 40% of agentic AI projects will be canceled by end of 2027” due to costs, unclear value, or inadequate controls. Anushree Verma says: “Most agentic AI projects right now are early-stage experiments or proof of concepts that are mostly driven by hype and are often misapplied.” Only “around 130” providers supply genuine agentic functionality (Gartner via MarTech, 25-06-2025).

Our position: pilots rarely fail in week 10. They fail in week 0 when nobody records the metric, assessor, and thirteen-week outcome. MIT emphasizes approach, Xact one written metric, Bosin a date, Wavect a baseline. Together: one page with task, baseline, metric, owner, decision date. An afternoon separates a useful test from an unnoticed subscription. Large businesses have innovation budgets. Twenty-person offices may have one chance to show staff AI works.

Where is this heading?

Dutch AI use grows faster than experience. Businesses with ten or more workers rose from 14 percent in 2023 to 23 percent in 2024, 9 percentage points (CBS 2025). In 2025, 27 percent with 10 to 49 workers used AI. Considering nonusers cited lack of experience at 73 percent (CBS, 12-12-2025). Proper pilots fill that gap.

Gartner’s forecast and MIT’s one-in-four implementation also suggest buyers with open-ended failures will demand terms next time.

Our expectation: by the end of 2027, SME buyers routinely request written metrics, baseline, end date, and return of data/methods, like quotes with delivery dates now. Adoption grows about 9 points yearly; inexperienced buyers meet project cancellations. Article 4 already requires attention to staff learning.

It fails if free open-ended pilots keep winning buyers or further enforcement delays remove documentation pressure. Test Dutch offers at the end of 2027. Bombos records that framework with you at every start.

What can a pilot not yet prove?

Thirteen weeks prove one task under those conditions.

  • Not the second task. Mailbox distribution does not prove monthly invoicing.
  • Not every peak. Year-end, quarterly VAT, or summer vacations may be absent. Record untested peaks.
  • Small businesses lack control groups. SpotDev uses a holdout; twenty-person offices often have one worker, making baseline comparisons vulnerable to volume changes.
  • Approval is not perfection. Blind clicking approves errors. Weeks 7 to 10 sampling checks this.
  • AI cannot find unwritten exceptions. Bombos interviews current workers and records knowledge. Otherwise pilots mainly measure missing information.
  • Most guides sell something. All ten compared providers sell pilots, scans, contracts, or platforms, including this page. Weigh sourced figures above advice. MIT’s roughly 67 percent external-partner implementation versus 33 percent internal is no vendor proof and “does not necessarily prove causation” (p.19).

How does Bombos approach this?

One task for three months costs € 2.000 excluding VAT, then € 1.000 monthly, the same for everyone (costs). About € 154 weekly; € 11.000 in year one. Compare hiring for growth. Bombos costs less than half a day weekly of the employee otherwise hired.

In week one we observe incoming work, such as shared mailbox distribution or retyping purchase invoices. Bombos interviews workers, then proposes with reasons and sources. The knowledgeable coworker approves, edits, or rejects in Bombos. Only then does the result appear in Outlook, Exact, or AFAS. Corrections become team rules.

Customer messages, payments, and contracts await approval by default. Bombos can remove boundaries at your request and risk later; keep them during pilots (errors and review). Personal pilot processing agreements are recorded separately. Discuss stopping and data return before starting. See privacy.

The first task starts growth with the same team at higher quality. After the pilot, your team teaches the next without technical skills. We guide the first; afterward you can do it yourself.

Sources

Each source was opened on September 30, 2026. Each original quotation appears verbatim.

Most sources sell pilots, measurement platforms, IT services, advice, or contracts. CBS, MIT, Jaggi, and legal texts do not. Bombos does.

Free, no obligation

More work done, at a higher quality, with the same team.

That is what Bombos is for: companies that grow fast and want to keep the same team. We start with one task that keeps piling up and guide you until your team can handle it. Then your team teaches Bombos the next task. Leave your number and we will call you back to talk about your situation.

We read what you write. Within one working day you hear from the one of us who knows your kind of work best.