Keep traditional workflow automation while fixed rules handle the work. Add a fixed AI step where free text needs interpretation. Use an AI agent only for exceptions whose investigation differs by case. Compare costs per correctly completed file rather than per execution.
This page covers selection rather than definitions: volume, variation, exceptions, task costs, and error profiles, with a table, actual n8n, Zapier, Make, and model billing units, and a complete break-even calculation. See the difference between an agent and workflow automation.
The short answer
- There are three options: fixed workflow, fixed workflow with one AI step, and bounded agent choosing subsequent steps (Anthropic, December 19, 2024).
- Dutch packages already mix them. AFAS, Dutch business software, allows administrators to, in translation, “include AI instructions in the workflow.” Exact, Dutch accounting software, combines AI posting proposals with configured approval flows (AFAS Help, Exact, read September 30, 2026).
- n8n bills full workflow executions; Zapier counts successful actions as tasks. Zapier Agents bills activities: 1,500 monthly for $400 annually (n8n, Zapier, September 30, 2026).
- Five reasoning rounds consume five times the model usage. Our example uses $0.009 per round: $450 rather than $90 for 10,000 files.
- One extra review minute per file at €45 hourly costs €1,875 monthly across 2,500 files, exceeding model usage in almost every scenario.
- Fixed workflows fail too. Microsoft documents duplicate emails and records in Power Automate (Microsoft Learn, August 7, 2026).
- AutomationBench’s leading agent fully completes 44.75% of business tasks at $1.14 each (Zapier, September 30, 2026).
- Our 2,500-file, 20%-exception scenario breaks even from 152 monthly exceptions. Four-minute review plus 1% additional missed errors turns it into a loss.
What is the difference in two sentences?
Workflow automation follows predesigned permitted paths. An agent’s model selects the next step during work, within granted permissions. A workflow can branch and contain a language model without being an agent.
Anthropic says: “Workflows are systems where LLMs and tools are orchestrated through predefined code paths” (December 19, 2024; sells models). Microsoft’s product named “agent flow” is described as “Agent flows are deterministic” (Microsoft Learn, August 3, 2026). Subscription labels do not establish who determines the route. Check who chooses steps and what they may do.
Email classification is an AI step in a fixed route. Agent freedom becomes necessary when investigation changes with findings. See question 49 and what an AI agent is for fuller explanations and autonomous work duration.
When does each approach win on volume, variation, exceptions, cost, and errors?
Fixed workflows suit many similar cases with stable rules. AI steps suit variable input format with fixed follow-up. Agents suit case-dependent investigation only when added value is measured. This is our purchasing synthesis rather than a measured ranking. Designs can coexist: agents call workflows and vice versa.
| Criterion | Fixed workflow without AI | Fixed workflow with AI step | Bounded agent |
|---|---|---|---|
| Volume | Strong with many similar stable-rule cases | Strong if one interpretation step suffices; count model calls | Defensible at high volume when value and review allow |
| Variation | Known differences become branches | Format or language varies, path stays fixed | Investigation changes with findings |
| Exceptions | Rule or human queue | AI converts free text; rules validate | Bounded investigation; missing policy remains human work |
| File costs | Development, platform, management, remaining manual work | Plus model usage and quality review | Plus variable investigation, repeated rounds, broader monitoring |
| Errors | Wrong rule, missed trigger, duplicate or partial execution | Plus misinterpretation and wrong fields | Plus wrong route, false completion, unwanted actions |
| Checks | Results and exceptions | AI output and final result | Also actions, boundaries, and stopping |
| Maintenance | Rules, integrations, changed agreements | Plus examples, instructions, regression tests | Plus permissions, investigation routes, tools, evaluations |
| Switching | Keep if demonstrably good and cheap | First test the bottleneck step | Only after measuring value above the AI step |
Basis: Anthropic architecture; n8n, Zapier, Make billing; Microsoft Learn and OWASP error types. All read September 30, 2026.
Without inventing a customer: daily export of approved orders remains a workflow. Converting free-text orders to proposals with fixed checks is an AI step. Comparing an unusual order with quotes, delivery notes, and correspondence may require varying investigation. Only that category is an agent candidate. Test those exact cases to establish cheaper, better work.
Does volume decide between agents and workflows?
Volume multiplies effects rather than choosing the design. Ten thousand identical actions favor rules. Ten thousand diverse files may contain considerable investigation. Exception workload decides: counts by type, current minutes, and remaining review.
A €0.10 improvement per file gives €50 monthly at 500 files and €5,000 at 50,000. At unchanged error rates, affected files also increase a hundredfold. Volume magnifies good and bad choices.
Exception percentage alone is insufficient. Distinguish:
- Known deviations, such as amounts above thresholds: rules.
- Unfamiliar wording, such as project IDs in descriptions: usually a fixed AI step.
- Investigation branching with findings, such as missing project IDs requiring quotes, email, and schedules: possible agents.
- Missing authority or policy, such as new supplier bank accounts: humans or process changes.
Count each category first. Only the third is agent work. See an invoice without a project number.
What does an agent cost per task compared with a workflow?
Providers bill different units, none equivalent to a completed file. Compare period costs divided by demonstrably correct completions. Published September 30, 2026 rates:
| Offer | Published amount | Unit and limit | Source |
|---|---|---|---|
| n8n Starter | €20 monthly, billed annually | 2,500 monthly executions | Pricing |
| n8n Pro | €50 monthly, billed annually | 10,000 monthly executions | Pricing |
| Zapier Agents Pro | $400 yearly, displayed $33.33 monthly | 1,500 activities rather than files | Pricing |
| Make Free | $0 | Up to 1,000 monthly credits | Pricing |
| Make Core | Displayed $9 monthly at 10,000 credits | Monthly versus annual payment unclear in extracted page | Pricing |
| Anthropic Sonnet 4.6, model only | $3 per million input tokens, $15 output | Excludes platform and setup | Claude pricing |
This is a rate card rather than a ranking. See total administrative automation costs. VAT, setup, existing licenses, management, and human review are excluded. No currency conversion was made.
n8n says: “An execution is a single run of your entire workflow.” Zapier: “Each successful action in a Zap counts as a separate task.” Make: “Non-AI apps: 1 operation equals 1 credit”; AI model usage is additional (Make Help). Zapier Agents counts actions, web use, and connected-knowledge searches as activities. One file can mean one execution, ten tasks, or thirty activities.
Utilization changes allocation. n8n’s €20 means €0.008 per execution at 2,500, but €0.08 at 250. These are allocated subscription costs rather than marginal rates.
Why does an agent consume more model usage than an AI step?
It reasons again after each tool response. Each round calls the model. Using Sonnet 4.6 prices with assumed 2,000 input and 200 output tokens, without caching or discounts: (2,000 × $3 + 200 × $15) / 1,000,000 = $0.009 per round. One AI round across 10,000 files costs $90; five equal agent rounds cost $450 in model usage alone. Context often grows, making later rounds more expensive.
An n8n user asked whether a thousand agent actions still count as one cloud execution and about hidden costs. The reply: one run is one execution regardless of actions, with model usage separate (n8n Community, December 10, 2025). Another asked “Ai chat model run twice and incurring cost twice” (July 15, 2025). Count calls in execution traces rather than design blocks.
Human review usually matters more. At assumed €45 hourly, one minute costs €0.75. An extra minute across 2,500 files costs €1,875 monthly, several times the example’s $90 model cost. Choosing by model-call price targets the wrong budget line.
How many exceptions make an agent layer pay off?
In this scenario, 152 monthly exceptions. All counts, times, and amounts are chosen inputs rather than market averages, Bombos results, or customer outcomes. Use the formula with your own data.
There are 2,500 monthly invoice files. Existing workflows handle 2,000 standard cases; people investigate 500. We compare only exceptions, keeping the standard route identical.
| Assumption | Value |
|---|---|
| Monthly exceptions, E | 500, 20% |
| Current manual time, t0 | 8 minutes each |
| Internal hourly rate, w | €45 |
| Added fixed monthly costs, F | €600: €300 setup depreciation, €300 management and licenses |
| Variable AI and platform usage, a | €0.08 each |
| Proposal review, t1 | 2 minutes each |
| Still requiring manual investigation, q | 10%, 50 cases |
| Additional manual time, t2 | 6 minutes |
- Current: 500 × 8/60 × €45 = €3,000 monthly.
- New: €600 + 500 × €0.08 = €640; review 500 × 2/60 × €45 = €750; investigation 50 × 6/60 × €45 = €225. Total €1,615 monthly.
- Difference: €1,385 monthly, or €6.00 versus €3.23 per exception.
Formula: net benefit = E × [w × (t0 - t1 - q × t2) / 60 - a] - F - L, where L is additional escaped-error damage. This scenario gives €45 × (8 - 2 - 0.6)/60 - €0.08 = €3.97 per exception. Break-even: €600 / €3.97 = 151.1, so 152 exceptions, about 6% of 2,500 files.
| Change | New costs | Net benefit |
|---|---|---|
| Baseline | €1,615 | €1,385 |
| Review 4 rather than 2 minutes | €2,365 | €635 |
| Double AI and platform usage | €1,655 | €1,345 |
| 1 percentage point extra missed errors at €250 | €2,865 | €135 |
| 4-minute review and extra errors | €3,615 | -€615 |
| Only 100 monthly exceptions | €803 | -€203 versus €600 manual work |
Double AI costs add €40. Two review minutes add €750. One percent extra missed errors adds €1,250. Review and escaped errors reverse the outcome.
Which errors do workflows and agents make?
Both fail differently. Workflows mainly have technical and rule errors. Agents add wrong routes, false completion, and unwanted actions. Compare error profiles.
Microsoft documents “you might see the results of the flow being duplicated,” including emails and list items. Failures include connection problems, expired tokens, licenses, and triggers (Microsoft Learn, August 7, 2026). A rule can repeatedly choose the wrong customer.
Flowmoat says workflows “fail loudly” (July 14, 2026); SOIS says agent failures are, in translation, “usually silent” (September 4, 2026). Both sell automation and provide no incident figures. Distinguish error types: interrupted runs are visible; wrongly linked files appear technically successful in either design.
Count wrong decisions, nonexecution, duplicate execution, false success, and unauthorized action separately. Models add “Indirect prompt injections occur when an LLM accepts input from external sources, such as websites or files” (OWASP LLM01:2025). Fixed AI steps share this risk. More permissions enlarge consequences.
Ten independent steps with 99% success each and no recovery produce 0.99^10 = 90.4% complete success. This applies equally to workflows and agents.
How reliably do agents complete business tasks today?
The leading agent completes fewer than half on the strictest public measurement. AutomationBench 1.0.6 has over 600 tasks across six business functions and 47 simulated apps. It checks system end states without an AI jury or human oversight, with at most 50 steps per task. Sonnet 5.5 led September 30, 2026 at 44.75%, measured $1.14 per task (Zapier). Zapier sells agents and automation. These are simulations rather than Dutch-office fieldwork.
Convincing answers differ from completed files. Check actual accounting results. Full autonomy differs from human-approved proposals.
The READY preprint reports GPT-5.4 and Sonnet 5 autonomous scores of 72.8% and 72.5% across 375 clinical audit cases. Reaching 76% target reliability required human escalation of 39.2% versus 29.6%, assuming 90% human success (Chatrath et al., September 2, 2026). That is 32% more review for a 0.3-point score difference. Measure both quality and review workload.
Verified Tool Calls found pre-retry checks reduced duplicate actions in 300 simulated runs (Mansoor, Phadke, Rana, July 31, 2026). One model and a deliberately weak comparator support measuring duplicate effects rather than proving reliable agents. Workflow restarts also must not email or pay twice.
Can an agent work on top of existing automation?
Yes, usually sensibly: retain standard workflows and put agents on the remaining queue. Agents can call workflows as tools; workflows can invoke an agent for a step. Feasibility depends on systems, permissions, and handoffs. See your own software.
Microsoft describes converting Power Automate flows to Copilot Studio agent flows as one-way because billing changes (August 3, 2026). Migration can change charges without improving work.
Compare existing or improved fixed workflow, workflow with one AI step, and investigative agent on identical files, inputs, sources, permissions, and completion definitions.
| Measure | Count |
|---|---|
| Correct completions | All agreed conditions, no prohibited side effects |
| Cost per correct file | All period costs divided by correct completions |
| Coverage | Correct completions divided by incoming work, plus pending work |
| Human time | Investigation, review, correction, management by exception type |
| Turnaround | Median and outliers, including approval waiting |
| Error profile | Five error types separately |
Zero errors in 100 cases still allows 2.95% true error probability, a one-sided 95% upper bound from our calculation; zero in 300 allows 0.99%. Same-supplier invoices are not independent, making practical bounds wider. See agent testing.
A justified question is not failure. If bank-account changes require confirmation, escalation is correct. Separate justified escalation, unnecessary escalation, wrongful autonomous processing, and errors despite approval. Otherwise providers inflate automation rates by forcing doubtful cases through.
What does this mean for an SME?
Our position: buy less investigation for stranded files and prove whole-process quality.
Billing units never count completed files. READY shows one-third review differences at similar scores. Microsoft documents duplication and expired connections in both designs. AFAS and Exact already combine AI and workflows.
The economic unit is the resolved exception. The quality unit is the correctly completed file. A twenty-person office cannot afford reasoning through every file or silently wrong workflows. It does have a daily investigation queue.
Start with existing paid package features, as in AI or traditional administrative software. Measure remaining exceptions by type. Consider agents when branching investigation makes the formula positive.
- Remaining work gets harder. Halving manual files does not halve hours.
- Queues shift to approval. Two hundred good proposals have little value if the only authorized coworker sees them Friday. Plan review capacity.
- Better input can remove agent need. Dropping from 500 to 100 exceptions through supplier project IDs makes the scenario unprofitable. Better forms can be cheaper than smarter software.
Where is this heading?
No series measures when agents economically overtake workflows. Dates claimed as facts are invented. See agent trends for observed developments. Benchmarks, rankings, and promises measure different things. We do know the top AutomationBench score is 44.75%, and review can differ one-third at similar scores.
Our expectation: on September 30, 2027, in bounded administrative processes with working standard routes, fixed execution plus a bounded investigative agent will beat all-route agents on cost per correct file in most fair equal-quality comparisons.
Halving model usage saves €20 monthly in the scenario. Halving two-minute review saves €375. Gains lie in reviewable proposals for narrow queues. Stable judgments can later become fixed rules. Mature designs move both ways.
This fails if full agents win at least half of preregistered process comparisons, or review and missed errors fall enough to remove extra-step costs. Bombos aims to retain sources, meaningful approval, and corrections through model changes.
What can agents and workflows not yet do?
Agents cannot invent undocumented business knowledge. Workflows cannot handle cases without designed rules.
- Missing policy remains human work. If sources are absent, asking is correct.
- Approval does not prove review. Reviewers can miss wrong reasoning. Measure post-approval errors. See errors and review.
- Language is an attack route. Agents and AI workflow steps reading external text face indirect prompt injection.
- Models do not repair connections. Expired tokens and missing licenses need operational fixes.
- No universal threshold exists. No independent employee, document-volume, or exception-rate cutoff establishes agent value. Our 152 is illustrative.
- Vendor figures differ. Exact claims 73% faster and 98% accuracy (September 30, 2026) without a test set. Recognition and speed differ from Zapier’s completion score.
- We have not measured this comparison ourselves. No controlled Bombos versus workflow versus AI-step test exists on identical unusual invoices. Figures remain examples.
See AI orchestration versus RPA for screen operation.
How does Bombos approach this?
Bombos puts AI agents on exceptions left by packages and workflows, retaining standard routes. The goal is more work at higher quality with the same team.
Chef recognizes and routes jobs to specialists built around your tasks, systems, and rules. They find customers, projects, and files in Exact or AFAS and prepare supported proposals. Wegwijzer explains, helps set limits, and suggests additional work.
You approve, change, or reject in Bombos; results then appear in software. Payments, customer messages, and contracts wait for technically enforced approval by default. You cannot disable it; Bombos can at your request and risk. Corrections become coworker rules, moving investigation into fixed rules.
Bombos interviews you to record undocumented agreements. After our guidance on the first task, your team teaches the next without technical skills.
Sources
Each source was opened on September 30, 2026, and each original excerpt appears verbatim in it. Dutch excerpts above are translations.
- Anthropic, Building effective agents, December 19, 2024. Sells models.
- Microsoft Learn, Agent flows overview, updated August 3, 2026. Sells Copilot Studio.
- Microsoft Learn, Troubleshoot issues with triggers, updated August 7, 2026. Sells Power Automate.
- AFAS Help, AI applications at AFAS, no visible date. Sells business software.
- Exact, Automatic invoice processing, no visible date. Sells accounting software.
- n8n, Plans and Pricing, rates read September 30, 2026.
- Zapier, Plans & Pricing, rates read September 30, 2026.
- Make, Pricing, rates read September 30, 2026.
- Make Help Center, Credits, no visible date.
- Anthropic, Claude pricing, rates read September 30, 2026.
- Zapier, AutomationBench leaderboard, version 1.0.6, September 30, 2026. Sells agents and automation.
- Chatrath et al., READY or Not: Reliable Enterprise Agent Deployment, preprint, September 2, 2026.
- Mansoor, Phadke, Rana, Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures, preprint, July 31, 2026.
- OWASP GenAI Security Project, LLM01:2025 Prompt Injection, 2025 edition.
- Flowmoat, AI agents vs automation: when you actually need each, July 14, 2026. Sells implementation.
- SOIS, AI agents versus workflow automation, updated September 4, 2026. Sells a platform.
- n8n Community, Costs and limitations of AI Agents (on cloud n8n), December 10, 2025.
- n8n Community, Ai chat model run twice and incurring cost twice, July 15, 2025.
