Tasks and systems

Which AI agents automate work between software systems?

Agents in Copilot Studio, Zapier, Make, n8n, and UiPath connect tasks across systems. Learn how MCP, screen operation, and controls work.

Agents in Microsoft Copilot Studio, Zapier Agents, Make, n8n, and UiPath, and assistants such as Claude with connected tools, can automate work between software systems. They retrieve information, choose next steps, and make changes through integrations, MCP, or screen operation within configured permissions and approval rules.

Suitability depends on the whole task: reading email, finding the customer, checking agreements, and recording results. This page compares providers, access routes, and current reliability evidence. Available connections differ from completed tasks. See when an agent is useful.

The short answer

  • Zapier Agents claims access to over 9,000 apps. This does not prove full support for every action. Zapier, read September 30, 2026
  • Microsoft lists 1,400 prebuilt data connectors. Individual actions must be added to agents. Microsoft, read September 30, 2026
  • Moneybird, online accounting software, offers official read-only and read/write MCP endpoints. Writing excludes deletion. Moneybird, September 28, 2026
  • Computer use operates websites and Windows apps with virtual mouse and keyboard, without direct connections. Microsoft, July 3, 2026
  • AutomationBench 1.0.6’s leading configuration scores 44.75% fully successful tasks in simulated apps without human clarification. Zapier, September 30, 2026
  • Partial computer-task success is not workflow completion. Fable 5.1 scores 77.9% partial and 41.7% strict on the same OSWorld 2.0 test. Anthropic, September 2026

What must an agent do between email, CRM, and accounting?

It must connect request meaning to the right data and make a verifiable change. CRM is your customer system. Accounting includes receivables and invoices. Customer names or IDs can differ between them.

Design example, not measured customer case: a customer emails a changed price agreement. Hours are in Excel, the project in CRM, and the invoice belongs in accounting.

Step Cross-system action Evidence
1. Read Read email and attachment Original email and relevant passage
2. Identify Find customer and project in CRM IDs rather than a name guess
3. Compare Match hours with rate and effective date Calculation and agreement per invoice line
4. Propose Present changes and deviations Specific proposal for approval
5. Process Create invoice and update file after approval Invoice ID and confirmed file change
6. Read back Check actual stored results Saved amounts, status, and references

This is our proposed design, aligned with the monthly invoicing round. It does not prove every provider supports every step in your packages.

Steps 5 and 6 are distinct. An agent saying it created an invoice is not evidence. Correct software readback is. AutomationBench likewise judges system end states. Shepard and Salimans, April 21, 2026

Which providers supply cross-system agents?

Microsoft, Zapier, Make, n8n, and UiPath provide tools for agents across systems. Assistants with vendor MCP tools offer another route. Features below come from vendor documentation read September 30, 2026. The final column is our assessment.

Provider or route Documented capabilities What to establish
Copilot Studio Connector actions, agent flows, web and desktop computer use. Connectors, computer use Available actions and whose permissions apply
Zapier Agents Tasks across over 9,000 apps with connected business knowledge. Product Required writing alongside searches
Make AI Agents Modules, scenarios, MCP tools, and other agents as tools; new agent app described as open beta. Documentation Fixed configuration versus agent choice
n8n AI Agent Per-action approval; rejection prevents execution. Documentation Workflow builder, permissions, recovery owner
UiPath agents with Maestro Orchestrate Coordinates agents, robots, and people; pause, resume, and correct processes. Product Whether process scale justifies setup
Claude with Moneybird MCP Find administration and create or update data through official server. Documentation Connected email and CRM, recurring-work initiation and monitoring

This is not a ranking. Providers sell development environments, executing assistants, or whole-process orchestration. Model names alone say little. Models, tools, business rules, and controls make agents useful.

Connector counts differ in meaning. An app can have dozens of actions while missing your posting. See connections with existing systems.

How do APIs, MCP, and computer use differ?

APIs provide agreed read/write actions. MCP standardizes exposing tools to assistants. Computer use operates screens. These are access routes for the same task.

An API might receive customer IDs and invoice lines. A connector is a prebuilt connection to it. Microsoft describes connector tools as individual actions or operations. Connector existence is only the start. Microsoft, read September 30, 2026

MCP, Model Context Protocol, covers tool discovery and calls. Documentation distinguishes tools/list and tools/call. MCP servers can use APIs, so they are not competing choices. MCP architecture, July 28, 2026 version

Moneybird’s September 28 documentation lists read-only and read/write endpoints, with “deletion is not supported.” For predictable continuous automation, it recommends API or CLI, command-line interface. MCP suits human-assistant collaboration. Moneybird, September 28, 2026

Our conclusion: choose access per action. An agent can interpret email, look up customers through MCP, and save approved results through fixed posting functions. See AI with your software.

When is screen operation needed?

Use it when an action exists on screen but lacks a usable direct connection. Microsoft documents selecting buttons, menus, and fields on websites and Windows desktop apps. Microsoft, July 3, 2026

For example, an order portal may require a reference and confirmation download. This design example still requires correct customer recognition and verification that saving succeeded.

For recurring postings, we prefer explicit write actions returning recognizable result IDs. Use screens for missing steps rather than the entire administration when only the final action lacks integration.

Microsoft recommends separate machines, restricted permissions, and allowed websites. An agent should not automatically inherit everything a director can do on their computer. Microsoft, July 3, 2026

How reliable are cross-app agents?

API access does not automatically provide complete reliability. AutomationBench checks workflows in 47 simulated apps across six functions and about 500 endpoints. Agents find their own actions. Zapier, read September 30, 2026

Version 1.0.6 lists Claude Sonnet 5.5 at maximum reasoning with default fallback models at 44.75%; similarly configured Opus 5.5 scores 42.47%. Fallbacks handle refused steps, so scores cover configurations. All checked end conditions must hold. No clarification or human approval is allowed. Zapier sells automation. AutomationBench, September 30, 2026

These percentages do not predict your invoicing or grade Zapier Agents. Configured tasks with controlled input and people differ from unknown environments.

The researchers check “whether the correct data ended up in the right systems”. Good wording cannot replace that. Shepard and Salimans, April 21, 2026

Our purchasing criterion: vendors must show saved results in every package, including what remained unchanged. Chat-only demonstrations prove part of the task.

What do computer-use benchmarks say about complete workflows?

Benchmarks measure partial progress, full completion, or intermediate steps. Percentages without this distinction mislead.

OSWorld 2.0 includes 108 longer tasks. Its paper says 64.8% need at least two apps or services. Median human duration is about 1.6 hours. This is relevant to system boundaries but broader and harder than one bounded administrative action. XLANG Lab, June 28, 2026

Model and setup Partial score Strict successes Source
Fable 5, OSWorld 2.0, August 2026 task version 72.9% 36.1% Anthropic, September 2026
Fable 5.1, same setup 77.9% 41.7% Anthropic, September 2026
Opus 5.5, OSWorld 2.1 81.8% Not in announcement Anthropic, September 22, 2026

All figures come from the maker. 77.9% versus 41.7% shows how incomplete work changes scoring, rather than 77.9% correct invoices.

Anthropic warns changed task files prevent comparison with earlier OSWorld 2.0 results. Version 2.1 differs again. This is not a growth curve. June’s paper explains difficulty; old scores do not judge current capability.

Why do 99%-correct steps not guarantee reliable agents?

All necessary steps must succeed. Assuming independent equally reliable steps, chain success equals per-step success raised to step count.

Assumed step reliability Steps Fully error-free chain
98% 10 0.98¹⁰ = 81.7%
99% 20 0.99²⁰ = 81.8%
99.5% 20 0.995²⁰ = 90.5%

These are our chosen-assumption calculations rather than measurements. Wrong customer selection can affect every later step. Recovery can improve results.

Our position: assess cross-system agents mainly on proven completion and recovery rather than integration counts. Catalogs show entry points. AutomationBench demonstrates end-state checks as a separate requirement.

If employees reconstruct every completed file, you have added a review process. Targeted review places source, decision, and saved outcome together. This is a design requirement rather than a savings promise.

How do you prevent duplicate postings and unwanted actions?

Permission, retries, and recovery solve different problems.

Permission: block intended changes before execution without approval. n8n pauses before tool calls. Approval executes; rejection cancels. Model instructions to be careful differ from execution barriers. n8n, read September 30, 2026

MCP does not fully arrange approval. Its specification requires no fixed interface and recommends human control. Check actual assistant configuration. MCP Tools, July 28, 2026

Retries: if accounting saves an invoice but confirmation is lost, another creation request may duplicate it. Our advice: unique references and checking existing results before retrying.

Recovery: if an invoice exists but CRM is unavailable, resume the missing file update without reinvoicing. Save confirmed steps. UiPath documents resuming and correcting without repeating completed work. UiPath, read September 30, 2026

Who handles stalled work belongs in preventing stranded tasks. Technically, failures must not erase knowledge of completed actions.

What proof should vendors show?

Demonstrate reads, writes, and failed-step behavior for your systems. Package logos do not answer that.

Test normal work and difficult variants: similar customer names, missing agreements, duplicate email, or interrupted writing. These are proposed tests rather than universal vendor standards.

Report full correct completion, justified human handoff, and incorrect execution separately. Safe handoff is valuable but not autonomous completion. Count review and recovery.

Define completion beforehand: one draft invoice for the correct debtor, correct amounts, file reference, and no unapproved customer email. See testing before automation.

Where is this heading?

Access and screen operation improve while measures change. Anthropic’s September 28 OSWorld 2.1 comparison gives Sonnet 5 and 5.5 partial scores of 57.0% and 80.1%, a 23.1-point generation difference. This is maker-measured, rather than annual growth or full completion. Anthropic, September 28, 2026

Our expectation: by late 2027, recurring administrative agents more often combine fixed integrations with screen operation for missing actions, while approval and result checks remain separately configured. Make already combines scenarios and MCP; Microsoft combines connectors and screens. Better models support this without rebuilding reliable steps.

Connection setup becomes less distinctive. Value shifts to applicable agreements, leading data, and actual completion. This weakens if vendors restrict access or maintenance and recovery exceed saved work.

Bombos aims to improve execution while preserving approvals, corrections, and rules across replaceable models.

What can this not yet do?

Incomplete or contradictory files do not yield certain truth. Two undated rates need clarification. Wrong customer linking stays wrong even if every later action technically succeeds.

Benchmarks do not establish error rates in your Exact accounting software, AFAS business software, or CRM environment. Tasks, data, and settings differ. These results do not replace your own trial. The provider table compares documentation rather than products we tested.

MCP does not guarantee safe execution. Incoming documents can contain others’ instructions. Maintain separation between content and permitted actions. Anthropic reports some approval bypasses in September safety tests. These are test setups rather than table-product failure rates. Anthropic, September 2026

Bombos lacks public measured traces supporting a full-chain success rate here. We give no proprietary result figure. See errors and review.

How does Bombos approach this?

Bombos builds AI agents for cross-system work at Dutch SMEs, targeting more work at higher quality with the same team. Start with one recurring task, configured around your systems and rules.

Chef recognizes and distributes work. Wegwijzer guides you, helps set boundaries, and suggests tasks.

Bombos reads connected email and documents, finds customers and files, and prepares sourced proposals. You approve, change, or reject in Bombos. After approval, it writes results such as Exact postings, AFAS data, or Microsoft 365 documents.

Payments, customer messages, and contracts require technically enforced approval by default. You cannot disable it yourself. Bombos can do so at your request and risk.

Bombos interviews staff to record undocumented knowledge. Corrections become coworker rules. We configure the first task together; your team teaches the next without technical skills. Approvals and corrections persist through model changes.

Sources

Each source was opened on September 30, 2026.

Free, no obligation

More work done, at a higher quality, with the same team.

That is what Bombos is for: companies that grow fast and want to keep the same team. We start with one task that keeps piling up and guide you until your team can handle it. Then your team teaches Bombos the next task. Leave your number and we will call you back to talk about your situation.

We read what you write. Within one working day you hear from the one of us who knows your kind of work best.