Choosing for your business

How do I choose an AI orchestration platform for my business?

Set firm boundaries, weigh seven criteria, give each candidate the same test, and check that you can take your data and rules with you before signing.

Choose an AI orchestration platform by defining your task and firm boundaries first: actions requiring approval, data destinations, and how you leave. Give at most three candidates the same test, including failures and an exit test. Only then score them using weights agreed in advance.

This page provides the complete method: a decision sheet, five exclusion criteria, a weighted 100-point scorecard, a worked example showing when results reverse, ten demonstration questions, an exit test for data and business rules, and fair comparison of different billing units. It is written for a small or medium-sized business without an IT department.

The short answer

  • Before any demo, define who decides and what the platform may do. Gartner puts “authority boundaries” before scoring (Gartner, 28-04-2026).
  • Apply exclusion criteria first. A high total does not override, translated, “a knock-out on data, testability, or exit” (Flowstate, 03-09-2026).
  • Weight criteria by your risks. In our example, one score point on the heaviest criterion changes the total by 5 points.
  • Test permission as behavior rather than a promise. OWASP advises “human-in-the-loop control” for consequential actions and authorization in executing systems, rather than only in the model (OWASP, 2025).
  • Export is not departure. n8n has “no automated way” to migrate Cloud to self-hosting. Credentials must be reentered (n8n, accessed 30-09-2026).
  • From January 12, 2027, in 104 days, EU cloud services may no longer charge switching fees under the Data Act (European Commission, accessed 30-09-2026). Export becomes cheaper; rebuilding does not.
  • Compare costs per correctly completed case: 1.000 cases mean 1.000 n8n executions, but 8.000 Make credits at eight actions per case.
  • No candidate passing boundaries, or scores within measurement uncertainty, makes “do not sign yet” a valid result.

What is an AI orchestration platform, and what are you comparing?

It distributes and monitors work between AI agents, existing software, and people: who does which step, in what order, with what rights, and what happens on failure. See orchestration versus ordinary automation and orchestration versus RPA.

Compare demonstrably correct completion of a whole task. Developer frameworks, workflow tools such as n8n or Make, Microsoft Copilot Studio agent builders, and managed services handling setup all use the platform label. Gartner distinguishes multiple orchestrator categories (Gartner, 28-04-2026).

Compare equal tasks and responsibilities. If A supplies software and B also provides setup and management, record who supplies A’s missing work and its cost. If unsure you need orchestration, start with AI software for administrative tasks.

Where do you start before speaking with vendors?

Use one decision sheet for one task. Record start and end, monthly and peak volume, current human time, systems and data, damaging errors, approval requirements, budget, and internal owner.

For example: match an incoming supplier invoice to case and project, show a proposed entry, process after approval, and demonstrate the entry in accounting software. The destination result proves completion. An agent’s completion message does not.

Include the alternative of doing nothing new: can a built-in feature or ordinary workflow already do this? If so, another platform may be unnecessary.

Our proposed sequence, rather than a proven optimum:

  1. Fix requirements and boundaries.
  2. Shortlist at most three candidates.
  3. Check documentation: export, rights, data flows, pricing rules.
  4. Give each the same test, including failures.
  5. Calculate costs and departure.
  6. Discuss scores once with process owner and administrator.
  7. Record reasoning and stop criteria.

One person may fill both roles in a fifteen-person company. Name a backup. NIST explicitly requires “contingency processes” for third-party AI failures (NIST AI RMF Playbook, GOVERN 6.2).

Which firm boundaries come before scoring?

Five requirements precede scoring. Price or usability points cannot compensate for failure. Your Tech Club similarly separates blocking requirements from scores (20-07-2026). These thresholds are our proposal for businesses without an IT department.

Boundary Required demonstration Accepted by Basis
Task is executable Exact reading/writing in your software, appropriate rights, verifiable confirmation Process owner and application administrator Aekyam, 03-06-2026, OWASP, 2025
No action without approval Missing/refused approval blocks actions after restart or retry; zero unauthorized test actions Authorized decision-maker OWASP, 2025
Data use is known Flows, locations, recipients, retention, subprocessors, processing agreement; no unresolved essential question Data owner and management EDPB, accessed 30-09-2026
Recovery has an owner Work can stop; open cases and completed actions remain visible; manual fallback tested Operational owner NIST
Exit is arranged Written readable data/rule export agreement, demonstrated by limited exit test Management and succeeding administrator Our minimum from n8n, Make, Microsoft documentation

Missing answers mean unknown rather than impossible. Unknown on a boundary means not ready to buy. Require proof or exclude the vendor.

Why zero unauthorized actions rather than model instructions? The UK AI Security Institute reports: “good containment should not depend on the model choosing not to test its boundaries” (AISI, incident discovered 28-07-2026). In 122 runs, 10 contained 19 unauthorized actions. These cybertests deliberately used broad access and disabled filters, so they do not establish an office error rate. They establish the mechanism: instructions alone are no boundary. See errors and review.

Which criteria and weights should you use?

Use seven criteria totaling 100 points, weighted by business impact. Equal weights, such as Your Tech Club’s ten at 10 percent each, hide the difference between extra administration and wrong payments. These are our proposed SME administrative orchestration weights, rather than a Gartner standard or law.

Criterion Weight Question and evidence required Theme source
Correctly complete whole task 25 Which agreed cases succeed? Human time per case? Show input, expected outcome, destination result. NIST
Control and traceability 20 Show refused action, source, rule version, decision-maker, result. How quickly can staff review? OWASP, Flowable, 28-08-2026
Connections and recovery 15 Second system step fails: show safe resumption without duplicates. Aekyam
Exit and rules usable elsewhere 15 Export and have another administrator resume a representative task. Measure missing parts and recovery hours. n8n, Make, Microsoft export documentation
Management and continuity 10 Who handles failures outside office hours? Show recovery report. NIST, CaseFy, 12-03-2026
Total costs and limits 10 Normal, peak, error-loop, review, change, exit costs? What happens at usage limits? n8n, Make
Self-service changes and support 5 Our employee changes, tests, restores a rule. Which help remains paid? CaseFy, Your Tech Club
Total 100

Score 0 to 5: absent, inadequate, usable with much manual work, meets predefined need, demonstrably better, repeatedly better including agreed exceptions. Define what 3 and 5 mean beforehand. Total = sum of (weight × score / 5).

Record confidence separately: D documentation only, V vendor demonstration, E observed in your agreed test, H repeated by another administrator. Labels do not affect scores. A 5 based only on D or V still needs testing. Untested marketing claims remain unknown. See agent testing and staff changing workflows.

What does a worked choice look like, and when does it reverse?

Two fictional candidates passing all boundaries illustrate sensitivity to weights. These are teaching numbers, rather than vendor results.

Criterion Weight A score A points B score B points
Task 25 4 20 3 15
Control 20 4 16 4 16
Connections 15 3 9 4 12
Exit 15 2 6 5 15
Management 10 4 8 3 6
Costs 10 3 6 4 8
Changes 5 4 4 3 3
Total 100 69 75

B wins by 6, mainly on exit. Move 10 weight points from exit to task: task 35, exit 5. A becomes 69 + 10 × (4 - 2) / 5 = 73; B becomes 75 + 10 × (3 - 5) / 5 = 71. A wins. Choosing B values easy switching over task performance. That is valid if recorded explicitly.

One task score point changes the total by 25 / 5 = 5. A 6-point lead based only on a short vendor demo, label V, is unconvincing. Resolve uncertainty through your test before adding decimals. The guides we read give criteria without this worked tradeoff (Flowable, CaseFy, Your Tech Club, Flowstate).

Which questions should you ask during the demo?

Ask for your task, then deliberate failure. Aekyam asks what happens at step 7 of 12 (03-06-2026). Use ten questions:

  1. Run our normal task with our data and software. Which steps belong to humans, agents, your staff, or third parties?
  2. Submit the same request twice. One entry or two? Where can we see that afterward?
  3. Disconnect after an action but before confirmation. How is repeated execution prevented?
  4. Change the case while approval is pending. Is old approval still valid?
  5. Have an unauthorized coworker approve and repeat without valid approval. What actually blocks execution?
  6. Stop while waiting and restart. Are status, reason, decision-maker, and last confirmed step preserved?
  7. Send instructions conflicting with our rules. Which permissions constrain actions regardless of model interpretation?
  8. Reach the usage limit. Stop, queue, or rising costs? Who is notified?
  9. Have our employee change and restore a rule. Are versions linked to affected cases?
  10. Deliver the next section’s exit test. What was measured, costs extra, or cannot leave?

Real users ask “How do you prevent requests from getting stuck waiting forever?” (r/AI_Agents, accessed 30-09-2026) and “Idempotency/replay protection for retries?” for purchase approvals (12-02-2026). That means preventing retries becoming second payments. Record evidence, observed result, open question, and owner. Use test data in an isolated environment.

How do you arrange exit and ownership of data and rules?

Prove three things separately: readable data, understandable business rules, and another operator actually resuming work. Major export features prove at most the first:

  • n8n: “Currently, there is no automated way to migrate from n8n Cloud to self-hosted.” Cloud credentials cannot be exported and require reentry (accessed 30-09-2026).
  • Copilot Studio: “Not all agent components and properties are included.” Dependencies transfer separately; authentication needs reconfiguration (Microsoft, updated 30-04-2026).
  • Make: scenarios export “as a .json file” with modules and field mappings. Connections require recreation; blueprints must be under 2 MB (accessed 30-09-2026).

None promises running on competing products without rebuilding. An n8n buyer asks about “export all my workflows,” “credentials, or execution data” before purchasing (n8n Community, 26-03-2026).

Make ownership concrete. This is a quote-appendix checklist, rather than a legal contract.

Component Vendor question Exit evidence
Customer data and outcomes What leaves, what does not, why? Field descriptions, timestamps, completeness sample
Business rules and corrections May we keep using them through another provider? Current rules, exceptions, versions, originating corrections
Configuration and custom work What is ours versus your generic software? Ownership/use rights, export, documentation
System access How does a successor get authorized access? Connections list, access grants, old rights revoked
Test cases May successors use the reference test? Test data, outcomes, rejection rules, past results
Pending cases What happens to work awaiting approval? Status, last confirmed step, next authorized action
History and retention What remains, where, how long? Readable event log, retention period, confirmed deletion
Exit costs Help, deadlines, prices? Schedule and agreed rates, export/rebuild priced separately

The Data Act has applied since September 12, 2025. Switching fees including data egress end January 12, 2027. SaaS export must use a “commonly used and machine-readable format.” Equivalent switching outcomes apply specifically to IaaS. Vendor intellectual property and trade secrets are excluded (European Commission, accessed 30-09-2026). GDPR requires the processor to “deletes or returns” all personal data at the controller’s choice (EDPB, accessed 30-09-2026). Returning personal data differs from business rules. “Data remains yours” proves none of these steps.

How do you run an exit test before signing?

Transfer one representative task with one recorded correction and one pending approval. Export agreed components. A second administrator reconstructs behavior in a separate environment. Compare outcomes and authority, measuring missing data/rules and work hours. Reauthorizing credentials is appropriate.

Report moving environments within the same product separately from moving products. The first is documented for n8n and Microsoft; the second is not. Set your own acceptance threshold, for example all agreed rules recoverable and no lost pending cases. Recovery-hour limits depend on budget and tolerable downtime. No generally achievable migration duration exists without measurement.

September 30, 2026 to January 12, 2027 is 104 days: October 31, November 30, December 31, January 12. When signing an annual contract now, ask how switching rules apply then. Reconfiguration, contracts, and retesting remain. A free file can mean an expensive transition.

How do you compare platform costs fairly?

Compare correctly completed cases rather than credits or executions. n8n charges “based on monthly workflow executions, regardless of complexity”: Starter €20 monthly, annually billed, for 2.500 executions; Pro €50 for 10.000 (30-09-2026). Make generally charges one credit per action, more for some AI functions. Scenarios can stop when credits run out, with alerts at 75 and 90 percent (accessed 30-09-2026).

Three calculations, using assumptions rather than market averages:

  • Units: 1.000 cases at one execution each = 1.000 n8n executions. At eight Make actions each: 1.000 × 8 = 8.000 credits.
  • Empty starts: checking every five minutes in a 30-day month means 30 × 24 × 60 / 5 = 8.640 starts, rather than completed cases. Ask which empty checks and retries are billed.
  • Human time: 1.000 cases, 2 minutes checking each, €45 hourly internally: 1.000 × 2 / 60 × €45 = €1.500 monthly. At 3.000 cases, review takes 100 monthly hours. Price pages omit that.

Year-one costs = setup + 12 × (subscription + measured usage + management + internal review/recovery) + expected changes. Cost per correct case = period costs / cases demonstrably meeting endpoint criteria. Include failed attempts and recovery. Always record correction minutes and report unauthorized actions separately. Low averages cannot offset those. Request normal volume, peak, and outage-with-retries scenarios. See platform prices.

How much does an error-free trial tell you?

Zero errors in 100 cases does not mean zero risk. For fixed error probability and 100 independent representative cases, the one-sided 95% upper bound is 1 - 0,05^(1/100) = 2,95 percent. You can only say with 95 percent confidence that errors are below 3 percent. At 1.000 monthly cases, that could still mean 29.

Self-selected or highly similar cases weaken even that figure. Use it to restrain false certainty rather than as an approval standard. Do not choose on one flawless demo. Tie autonomy to potential damage. Human-checked proposals tolerate more errors than unapproved payments. See may AI send email itself?.

Model switching differs from platform switching. Multiple supported models do not free your rules or history. “Model-agnostic” says little about exit. Repeat quality tests after model changes and reconstruct the task after platform changes.

Our position: is a good exit test also a management test?

Yes. What a successor cannot reconstruct is also hard for your team to recover after a serious outage. Exports omit setup (n8n, Microsoft, Make). Approval and rights must remain auditable (OWASP). NIST requires third-party contingency processes (NIST). Their common question: can humans read how work operates?

Suppose corrections add 10 monthly rules, each needing 20 minutes to reconstruct at exit. After 6 months: 60 × 20 / 60 = 20 hours; after 24 months: 80 hours. At €75 hourly, €1.500 versus €6.000. These fictional numbers fail if rules expire or transfer automatically. The direction remains: learning increases rule value and exit cost. Agree correction ownership at purchase.

Central orchestration also concentrates outages while simplifying oversight. Agree manual continuation and where staff find pending work. This position is unmeasured; same-product recovery is easier than product migration. Still, inability to demonstrate exit also prevents demonstrating recovery.

Where is this heading?

The pace and exit attention are growing. Gartner published scoring guidance in April 2026. Dutch July/September guides explicitly make exit a blocking criterion. Switching fees disappear January 12, 2027. We found no SME migration-cost or implementation-quality time series.

Our expectation: by the end of 2027, a company documenting rules, tests, and a workbook monthly from its pilot transfers a representative task using at most half the reconstruction hours of a comparable company exporting only at departure. Expensive work recovers rule meaning rather than files. Continuous documentation does it once in small pieces.

It fails if monthly recording costs more than exit savings, standardized exports erase the difference, or operators differ too much for fair comparison. Test with a late-2026 baseline and fourth-quarter 2027 repetition.

Bombos makes each correction-created rule readable, with approver and effective date, turning exit testing into a list rather than a search.

What can this method not do?

It does not choose a winner. We found no independent comparison of complete platforms on identical Dutch SME processes. Results come from your test.

Weights and thresholds are our proposal, neither law nor Gartner-approved nor validated elsewhere. Different risks need different weights, as the sensitivity example shows.

Sources have interests. Gartner sells research; we read only its public abstract. Flowable, Aekyam, CaseFy, Flowstate, and Your Tech Club sell software or implementation. Advice provides criteria rather than product superiority evidence. Forum questions establish questions rather than frequency.

Ten demo questions, exit tests, and three cost scenarios take a fifteen-person company a few days. Otherwise a years-long choice rests on a demo.

This is not legal advice. Official Data Act/GDPR explanations are summarized. Applicability depends on your contract.

How does Bombos approach this?

Hold us to the same boundaries, ten questions, and exit test. We have no completed scorecard with test outcomes to show. You create it with us on your task.

Bombos reads incoming work and finds customers and cases in your software: Exact, Dutch accounting software; AFAS and Visma, business software; Twinfield, Moneybird, and SnelStart, accounting packages; Nmbrs, payroll software; Syntess, software for technical installation businesses; Microsoft 365; or Google Workspace. It prepares supported proposals. You approve, edit, or reject in Bombos. Only the result goes into your package. Payments, customer messages, and contracts await approval by default. That technical boundary meets the second requirement. You cannot disable it. Bombos can remove it at your request and risk.

Chef distributes work, Wegwijzer guides you, and specialists fit your tasks and rules. Corrections become rules for coworkers too. Models are replaceable. Bombos interviews you to record knowledge held only in people’s heads.

The goal is more work at higher quality with the same team. Your team then teaches the next task without technical skills. See which platform to start with now and working with existing software.

Sources

Each source was opened on September 30, 2026. Original excerpts appear verbatim; translated excerpts are identified above.

Free, no obligation

More work done, at a higher quality, with the same team.

That is what Bombos is for: companies that grow fast and want to keep the same team. We start with one task that keeps piling up and guide you until your team can handle it. Then your team teaches Bombos the next task. Leave your number and we will call you back to talk about your situation.

We read what you write. Within one working day you hear from the one of us who knows your kind of work best.