Starting, testing, and adjusting

Can I start with an AI automation pilot?

Yes. Choose the right trial: free software proves something different from a paid pilot using your own email and software, with a stop date agreed first.

Yes, you can start with an AI automation pilot. For a small or medium-sized business, that is the normal starting point. The question is which trial you buy. A free software trial, a proof of concept on sample data, and a paid pilot in your own email and software prove three different things.

This page covers whether a pilot makes sense and what makes a good one. It examines what providers sell as “try it first” in 2026, their prices and terms, what the familiar MIT, Gartner, McKinsey, and RAND failure figures actually measure, costs including your time, and what to check before signing. The weekly implementation plan is in starting a pilot for administrative automation with AI. Testing AI itself is covered in testing an AI agent before you automate.

The short answer

  • Almost every provider offers a trial: Zapier gives 14 days of Professional without a credit card, n8n offers 1.000 executions, Exact Online Boekhouden, Dutch accounting software, gives 30 free days, and Copilot Studio lets you build an agent but not publish it.
  • A free trial proves you understand the interface. Only a pilot using your own email, software, and people proves the work becomes lighter.
  • MIT’s familiar “95 percent fail” measures something different from the headline. Of organizations piloting task-specific AI (20 percent), one quarter reached production (5 percent), against a threshold of noticeable, lasting productivity or profit impact.
  • Gartner expects more than 40 percent of agentic AI projects to be canceled before the end of 2027 because of costs, unclear value, or weak risk controls. That is a prediction.
  • McKinsey (August 2026): 80 percent report personal productivity gains, but the share reporting profit impact remains 37 percent a year later.
  • In the Netherlands, 27 percent of small businesses used AI in 2025, versus 11 percent in 2023. Among those considering AI but not adopting it, 73 percent cite lack of experience.
  • A Dutch consultancy charges € 20.000 for a pilot delivering a first working process. Bombos charges € 2.000 for three months. Allow 30 to 60 hours of your own time in either case.
  • Our position: a pilot succeeds when it produces a decision, even when the answer is no.

What types of AI pilots do providers offer in 2026?

Providers offer four forms of trying first: a free software trial, an accounting software trial including AI, a paid agency audit or proof of concept, and a guided paid pilot on your own work. Only the last two use your data. Only the last uses your daily work.

Trial type Example Provider’s price and terms What it proves
Free software trial Zapier 14 days Professional, no credit card You can operate the software
Free software trial n8n Cloud No credit card, 1.000 executions, Pro features; then from € 20 monthly with annual billing; Community Edition free to self-host You can operate the software
Trial without publishing Copilot Studio Free 30-day trial, extendable; “you can’t publish the agent” You can build an agent and talk to it in a test window
Software package trial Exact Online Boekhouden 30 free days; then € 49 to € 299 monthly, with purchasing, banking, and support agents in every edition The package fits, rather than the agent handling your email stream
Business AI free entry Salesforce Agentforce Free start with Salesforce Foundations; then Flex Credits, € 500 per 100.000; details: “Contact a sales representative” You can use the building tools, rather than know the cost for your work
Paid audit Crux Digits € 2.500 excluding VAT, 1 to 2 weeks What you could automate and its benefits
Paid MVP Crux Digits € 20.000, 4 to 6 weeks, your data, agreed acceptance criteria A first process works
Guided live pilot Bombos € 2.000 for three months, then € 1.000 monthly, excluding VAT, same for every company A recurring task in your own email and software becomes lighter
Subsidized or success-based EDIH, WBSO combined with MIT and SLIM, success fee (Stratalytic, May 26, 2026) Costs partly or fully reimbursed, 30 to 60 hours of your time Depends on the route

Prices were read on providers’ pages on September 30, 2026 (Copilot Studio: August 3, 2026 page). A price does not prove performance. Free trials require your own setup. Paid trials include guidance. Free and guided together is rare outside subsidies. Providers offering both generally recover the cost afterward.

How do a free trial, proof of concept, and pilot differ?

A free trial tests software. A proof of concept tests whether something can work. A pilot tests whether it works here. Dutch AI agency Crux Digits defines them, in translation, as “can this work at all?”, “does this work here, with our data and people?”, and, for an MVP, “will people use this, and is it worth maintaining?” (Crux Digits, August 11, 2026). Crux sells all three, so its definition comes from an interested party, but it makes sense.

For administration, this distinction matters completely. Ten sample invoices prove a language model can read an invoice. You already knew that. You do not know what happens with a supplier placing the project number in the description, a credit note attached to a forwarded email, or a coworker forwarding three emails together on Friday afternoon. Only weeks in your own mailbox reveal that.

A package trial is different again. Exact includes purchasing, banking, and support agents in every Exact Online Boekhouden edition (Exact, 30-09-2026). Thirty days shows how they work inside the package. It does not show what happens to email stuck in a personal inbox that never reaches it. See AI inside AFAS, Dutch business software, or Exact, or an AI agent alongside.

The closer a trial is to real work, the more it costs and proves. A free trial is fine for understanding software. Just do not call it a pilot.

Is it true that 95 percent of AI pilots fail?

No. MIT NANDA’s July 2025 The GenAI Divide says “95% of organizations are getting zero return” (MIT NANDA, July 2025, p. 3). It concerns organizations and returns, rather than collapsing pilots.

The report supplies its own task-specific AI funnel: “Sixty percent of organizations evaluated such tools, but only 20 percent reached pilot stage and just 5 percent reached production.” Of evaluators, 1 in 12 reached production (5 divided by 60). Of actual pilots, 1 in 4 did (5 divided by 20). General chatbots such as ChatGPT and Copilot have a much wider funnel: 50 percent piloted and 40 percent implemented, so 4 in 5 pilots were implemented. The report calls this a roughly 83 percent “pilot-to-implementation rate.”

Two qualifications make the 95-percent figure stricter than it sounds. Success requires noticeable, lasting productivity or profit impact. The method is limited: 300 public initiatives, 52 interviews, and 153 surveys, with the warning “These figures are directionally accurate based on individual interviews rather than official company reporting.” The project producing the report builds agent infrastructure itself.

Its useful distinction is: “External partnerships see twice the success rate of internal builds.” Mid-market companies took an average 90 days from pilot to full adoption; large companies took nine months or more. Also: “Buyers who succeed demand process-specific customization and evaluate tools based on business outcomes rather than software benchmarks.” This guides pilot design rather than arguing against pilots.

Why do so many AI pilots stop according to Gartner, McKinsey, and RAND?

The major studies point to the same reason: nobody agreed which outcome counts, so afterward nobody can say whether it worked. Figures vary because they measure different things.

Source What it measured Main figure Limitation
MIT NANDA, July 2025 300+ initiatives, 52 interviews, 153 surveys 95% no return; task-specific tools: 20% pilot, 5% production Interviews, not peer-reviewed
Gartner, 25-06-2025 Agentic AI forecast More than 40% canceled before end of 2027 Forecast, rather than measurement
Gartner, July 2024 (cited by Computerworld, 04-07-2025) Generative AI forecast At least 30% abandoned after proof of concept before end of 2025 Forecast, different topic from 40%
McKinsey, 25-08-2026 Global survey, mainly large organizations 44% scaling AI; 37% see profit impact; 6% “high performers” Few small businesses
RAND, 13-08-2024 65 interviews with data scientists and researchers “More than 80 percent” is a cited estimate Traditional ML projects, rather than generative AI
CBS, 12-12-2025 Dutch small businesses 27% used AI in 2025 (19% in 2024, 11% in 2023) Provisional 2025 figures

Gartner cites rising costs, unclear value, and weak risk management. Analyst Anushree Verma says: “Most agentic AI projects right now are early-stage experiments or proof of concepts that are mostly driven by hype and are often misapplied” (MarTech on Gartner). It estimates only about 130 of thousands of “agentic” providers are genuine, calling the rest “agent washing.” Keep its figures separate: 30 percent concerns generative AI through 2025; 40 percent concerns agents through 2027.

McKinsey shows the use/results gap clearly. “Eighty percent of respondents report that AI has improved their individual productivity,” but the share seeing profit contribution is “essentially unchanged from a year ago, at 37 percent” (McKinsey, 25-08-2026). Agent scaling rose from 27 to 40 percent at large organizations, staying at 22 percent at smaller ones. Almost three quarters of high performers redesign work itself, versus a quarter of the rest.

RAND’s interviews identify five causes: wrong problem, insufficient or wrong data, chasing new technology, missing infrastructure, and tasks beyond AI’s ability. RAND recommends at least a year on one problem and warns leaders may redirect successful teams too soon.

The four disagree on numbers and agree on the mechanism. None blames the model. They blame pilots without a goal, owner, or decision point.

What does an AI pilot really cost, including your time?

A pilot costs the provider’s invoice plus 30 to 60 hours of your time. At a small business, that time is often the largest item. Dutch consultancy Stratalytic budgets 30 to 60 hours for a 6 to 10-week pilot from a sponsor, someone who knows the data, and someone who knows the work (Stratalytic, 26-05-2026). At an assumed internal rate of € 50, that adds € 1.500 to € 3.000.

  • Free software trial: € 0 invoice, but all setup is yours. Those 30 to 60 hours may increase without help.
  • Crux Digits: € 2.500 audit, reportedly credited if you commission an MVP within 60 days; € 20.000 MVP over 4 to 6 weeks, roughly € 4.000 weekly (Crux Digits).
  • Bombos: € 2.000 for three months, about € 154 weekly over thirteen weeks. Continuing makes year one € 2.000 plus nine times € 1.000, totaling € 11.000 (Bombos).

These buy different things. A € 20.000 MVP is software you subsequently own. A € 2.000 pilot is a task operating in your email for three months with its makers alongside. Compare price to what you want to learn.

Can it be free? Stratalytic names three routes. An EDIH (European Digital Innovation Hub) program has the EU pay for a bounded 4 to 12-week program; you contribute time. Combined WBSO, MIT, and SLIM subsidies reportedly recover 46 percent on a € 30.000 project in its example. Some agencies use success fees where outcomes are directly measurable. Stratalytic sells advice, and we have not checked subsidy terms with RVO. For a roughly € 2.000 pilot, applying usually costs more time than it is worth.

What should you check before signing an AI pilot?

Write down five things: process, metric, threshold, metric owner, and stopping point. Without them, nobody can establish success, exactly where Gartner, McKinsey, and MIT see projects stall.

Agreement Why Source
One named recurring process, such as purchase invoices from mailbox to Exact “Try AI in administration” produces no decision Our reasoning
Baseline of at least two normal weeks No starting point means no comparison Crux Digits, 11-08-2026
A pass threshold that can fail, plus one countermeasure that must not worsen Speed without correctness is no gain Crux Digits, 11-08-2026
Metric owner In translation: “A metric without an owner does not survive the pilot.” Crux Digits, 11-08-2026
Real users and real data A pilot asks whether it works here Crux Digits, 11-08-2026
Stop date and three outcomes: continue, adjust, stop A pilot only allowing continuation is a sales process Our reasoning

Crux’s usable criterion, translated: “the drafted reply is sent without substantial changes in at least 60% of cases.” After three months, you can answer yes or no. You cannot do that with “AI helps us.”

Ask three more questions. What do you receive if you stop, including process descriptions, rules, and action logs? Who approves, and where? A pilot without approval on every payment or customer email is a risk; see errors and review. What happens to your data? Processing personal data requires a processing agreement. Pilot terms are covered in the starting guide, task selection in which administrative processes to automate first.

When is an AI pilot successful?

A pilot succeeds when it produces a decision on the agreed date, including no. That is our position, based on three rarely combined facts.

First, MIT’s threshold is wrong for you. NANDA requires noticeable, lasting productivity or profit impact and counts the 5 percent achieving “millions in value.” A twenty-person administration office needs a smaller test: does this task on real email meet agreed accuracy with approval still enabled, and how much review does it take?

Second, the Dutch buyer is on the favorable side of the divide. MIT sees spending favor sales and marketing while returns more often lie in back-office work: “Budgets favor visible, top-line functions over high-ROI back office” (MIT NANDA, July 2025). Among Dutch AI users, 32 percent already use it for administration or management (CBS, 12-12-2025). Purchase invoices or a shared mailbox sit precisely where major studies find returns.

Third, experience is the obstacle, rather than money. Dutch businesses considering but not adopting AI cite lack of experience (73 percent), privacy (49 percent), and legal consequences (42 percent), according to CBS. A pilot buys experience without moving all administration.

Its value lies in what you know after three months. “We stop because our three largest suppliers send scans, and that must change first” earns its cost. “It looks promising; let’s watch another month” does not. Software development became cheap over the past two years, making trials cheap. Ending them became scarce. A small business risks four unfinished experiments leaving separate side systems behind.

Where is this heading?

Dutch small-business AI use rose from 11 percent in 2023 to 19 percent in 2024 and 27 percent in 2025, around 8 percentage points annually (CBS, 12-12-2025). Extrapolating gives about 35 percent in 2026 and 43 percent in 2027. That is our calculation, rather than a CBS forecast. McKinsey’s smaller organizations stayed at 22 percent scaling agents, while large ones rose from 27 to 40 percent. Use grows faster than operational adoption.

Our expectation: by the end of 2027, a guided paid trial on one recurring task is the usual first step for Dutch small and medium-sized businesses. A free tool trial no longer counts as a pilot. Free trials are everywhere. The remaining obstacle is experience, bought through guidance on your work. As nearly half your sector uses AI, the question becomes what remains when the pilot stops. See AI agent trends for businesses.

This fails if providers give away guidance and recover costs through licenses, if privacy and liability move the obstacle to lawyers, or if the use/adoption gap stays as flat as McKinsey’s smaller-business figures.

By then, Bombos arranges the trial as we believe it should be: one task, fixed price, three months, and a written decision at the end.

What can a pilot not yet do?

Three months cannot show performance through a merger, new accounting package, or new coworker changing everything. It shows a season, rather than every exception.

It cannot prove error-free work. Approval catches mistakes but takes effort. A human checking a hundred proposals daily pays less attention to the hundredth. An honest pilot measures review time and samples results, rather than only counting approvals. See testing an AI agent before you automate.

A pilot cannot find an agreement existing only in a coworker’s head. Bombos interviews you and records it for next time. That coworker’s time belongs in the 30 to 60 hours.

Most figures here concern large organizations. MIT, Gartner, McKinsey, and RAND did not measure a twenty-person Dutch business. CBS measures use, rather than results. No public study gives a success rate for Dutch small and medium-sized business pilots. We have no such figure either.

How does Bombos approach this?

Start with three months for € 2.000, then € 1.000 monthly, the same for every company and excluding VAT (costs). First-task setup is included. The makers guide you throughout.

The pilot uses your work. Bombos reads connected email, finds the customer and case in your software, and prepares a proposal showing its sources. You approve, edit, or reject in Bombos. Only then does the result appear in Exact, AFAS, or your email. Payments, customer messages, and contracts wait for approval by default. Corrections become rules for coworkers too. First tasks include matching invoices to projects and the shared mailbox; see where to start.

The goal is more work at higher quality with the same team, and knowing after three months whether this task delivers that. Your team then teaches Bombos the next task without technical skills. We guide you; afterward you can do it yourself.

Sources

Each source was opened on September 30, 2026. Original excerpts appear verbatim; translations are identified above.

Free, no obligation

More work done, at a higher quality, with the same team.

That is what Bombos is for: companies that grow fast and want to keep the same team. We start with one task that keeps piling up and guide you until your team can handle it. Then your team teaches Bombos the next task. Leave your number and we will call you back to talk about your situation.

We read what you write. Within one working day you hear from the one of us who knows your kind of work best.