Tasks and systems

How do you automate documents, Excel files, and emails with AI?

Let AI extract email and document data, use fixed calculation rules, and approve each case before it reaches Excel or your software.

Automate documents, Excel files, and emails with AI as one chain: AI reads email and attachments and extracts data, fixed rules calculate in Excel, a person approves the complete case, and only then is the result written and read back.

This page explains reliable capabilities by file type, recent benchmark error rates, and what they mean. Then it connects email, PDF, and Excel through a nine-step example and a €40 conflict no model can resolve for you.

The short answer

  • AI reads and sorts well. Aid4Mail’s best model correctly labeled 99.38% of 1,600 emails in September 2026 (Aid4Mail, 17-09-2026).
  • Extraction is good but imperfect. Structured Output Benchmark’s best exact field-value scores were 83.0% on text and 67.2% on images, despite nearly always correct format (arXiv, 28-04-2026).
  • Multistep Excel remains weak. SpreadsheetBench 2 found 89.69% correct cell modifications but only 34% wholly correct financial tasks (arXiv, 29-06-2026).
  • Technical success differs from completion. EmailBench recorded 99.7% successful technical actions but only 69 of 206 scenarios (arXiv, 25-09-2026).
  • Excel’s standalone =COPILOT() function stopped working September 14, 2026. Sidebar Copilot remains (Microsoft Support, accessed 30-09-2026).
  • Power Automate’s Excel connector retrieves only 256 rows by default. More silently remain missing until pagination is enabled (Microsoft Learn, accessed 30-09-2026).
  • At 99% per-field correctness, twenty fields give only 81.79% probability of a flawless document. Measure cases, rather than fields.

What can AI reliably do with documents, Excel, and email now?

In 2026, recognition and extraction are reliable, simple spreadsheet edits moderately reliable, and long dependent action chains still weak. This distinction outweighs model choice.

For email, sorting is strongest. Aid4Mail tested fifteen models on 1,600 messages, reaching 99.38% best labeling (Aid4Mail, 17-09-2026). Limits: 200 synthetic base messages in eight languages excluding Dutch; its maker sells Aid4Mail.

For documents, extraction measurement varies. ExtractBench covers 370 documents, 4,869 pages, 67 types, assessing field values and evidence locations (ExtractBench, snapshot 11-09-2026). Best scores were 95.91 field-value F1 and 58.11 for identifying correct words. Benchmark maker LlamaIndex sells the winning system.

For Excel, Copilot edits workbooks directly, plans, and chats (Microsoft Support, accessed 30-09-2026). Microsoft says “Review, edit, and verify anything Copilot creates before you rely on it.” (Microsoft FAQ).

For Word/PDF output, AI often is unnecessary. Word Online fills fixed templates and converts to PDF with a 10 MB input limit (Microsoft Learn, accessed 30-09-2026).

How reliable is AI by file type?

Each type has typical errors and checks. The capabilities column is documented; the other two are our design.

Input/output Documented capability Typical error Built-in check
Text-based digital PDF Field/table extraction (Microsoft Learn) Previous-page amount, wrong total Page/source for critical fields; recalculate totals
Scan/photo/handwriting Vision plus structured extraction (arXiv, 18-08-2026) 8 read as 3, tilted image, missing page Request unreadable material again; no guesses
Long PDF tables Repeated rows with evidence positions (ExtractBench) Missing, duplicate, merged rows Compare counts/totals; test two-page tables
Fixed-column Excel Connector reads/writes rows (Microsoft Learn) Wrong column, lost leading zero, duplicate key Fixed schema/types, unique row ID
Excel formulas/tabs Multistep editing (SpreadsheetBench 2) Wrong formula logic, broken reference Test, recalculate, check prohibited changes
Email with attachments Trigger plus document action (Microsoft Learn) Old quoted agreement, wrong attachment Retain message/attachment IDs, sender/date; escalate conflict
Word report/PDF output Template filling/conversion (Microsoft Learn) Empty field, wrong customer attachment Required fields, fixed template, recipient check

For CSV, use ordinary import software. Define delimiter, encoding, decimal separator. Keep customer/item numbers as text. Turning CSV into an AI image only adds errors.

What do benchmark error rates mean?

Know the measured unit: fields, tasks, or complete cases. These genuine scores belong to different tests, rather than one ranking.

Test Measurement Score Valid interpretation
ExtractBench, 11-09-2026 Document field values F1 95.91 Not flawless-document percentage
Structured Output Benchmark, 28-04-2026 Exact image field values 67.2% 32.8 percentage points not exact
SpreadsheetBench V1, Google, 10-03-2026 Gemini in Sheets tasks 70.48% 29.52% unsuccessful
SpreadsheetBench 2, 29-06-2026 Claude Opus 4.6 financial tasks 34% 66% not wholly correct
Aid4Mail, 17-09-2026 Gemini 3.7 Flash email labels 99.38% 0.62% wrong labels
EmailBench, 25-09-2026 Complete email scenarios 69 of 206 66.5% unsuccessful

Sources: ExtractBench, Structured Output Benchmark, Google, SpreadsheetBench 2, Aid4Mail, EmailBench.

Three findings: larger units score lower; labels nearly always succeed while full tasks often fail. Format proves little: “models achieve near-perfect schema compliance”, but image field values were inexact roughly a third of the time (arXiv, 28-04-2026). Test design matters: only 4 of 35 open-model configurations exceeded F1 0.5 on 100 academic transcripts (arXiv, 18-08-2026).

No independent complete-chain percentage exists for Dutch invoices, work orders, and emails together. Promised figures are self-measured or invented.

Why do good fields not guarantee a good file?

Errors accumulate. A calculation, rather than measurement: twenty fields each 99% correct yield 0.99²⁰ = 81.79% flawless documents, roughly one in five with an error.

Real errors correlate: poor scans affect several fields. This is no forecast but explains spreadsheet outcomes. SpreadsheetBench 2 reports “89.69% Modification but only 34.00% Accuracy” (arXiv, 29-06-2026). Nearly nine in ten cell changes were correct while two-thirds of whole workbooks were not.

Microsoft-affiliated EmailBench researchers say “valid tool execution is not equivalent to task completion” (arXiv, 25-09-2026). Technical actions succeeded 99.7%; scenarios about a third. Limits include synthetic environments, May runs, and partly AI assessment. Still, the success-checkmark/completed-work gap is familiar in offices.

Ask vendors how many complete cases passed without correction and how many errors were wrongly accepted, alongside field accuracy.

What can Copilot in Excel currently do?

Sidebar Copilot edits workbooks, creates formulas/analyses, and helps plan (Microsoft Support, accessed 30-09-2026). It assists workbook users rather than emptying mailboxes overnight.

Two current details many guides miss:

  • The cell function is gone. “Starting September 14, 2026, the COPILOT function is no longer available in Microsoft Excel.” (Microsoft Support). Existing values remain; recalculation produces #NAME?. Build no new workflow on =COPILOT().
  • Editing requires automatic calculation. Features depend on licenses and Microsoft 365 settings (Microsoft FAQ).

Google reports 70.48% successful Gemini in Sheets tasks on full SpreadsheetBench (Google, 10-03-2026). This vendor measurement is six months old and uses English Excel-forum questions. Microsoft publishes no comparable per-edit score.

Have Copilot propose a formula, test zero hours, empty rates, decimal commas, negative adjustments, duplicate rows, then fix it in place. Fixed logic belongs in fixed formulas rather than newly generated answers each time.

How do you connect email, PDF, and Excel?

Center the case. A maintenance company’s monthly billing receives an emailed price agreement, PDF work order, and Excel timesheet. This is a design example, rather than a customer case.

Step Action Output
1. Receive Store email/files with time and unique ID Receipt record
2. Separate Read email text, PDF document, Excel table Versioned sources
3. Extract Project, date, hours, rate, reference with locations; leave missing values blank Candidate field values
4. Identify Match project/customer by fixed identifiers beyond names One match or exception
5. Compare Email agreement, work order, hours, master data Visible discrepancies
6. Calculate Fixed formulas, duplicate/period/total checks Review report
7. Present Draft, sources, exceptions to authorized coworker Approval, correction, rejection
8. Write Approved version only, linked to case ID Result or visible error
9. Read back Compare stored outcome, check no second record Completion or recovery

Steps 1 and 9 are often omitted, creating silent errors. Excel’s connector defaults to 256 rows, limits files to 25 MB, and timeout retries can duplicate entries (Microsoft Learn). Check prior write success before retrying.

Step 3 has another trap: AI Builder “returns all extracted values as strings” (Microsoft Learn, updated 14-01-2026). Convert amounts from text with Dutch decimal commas. Incorrect page ranges produce partial data according to the same source.

Software connections are covered in AI and existing systems.

What if email, work order, and Excel contradict each other?

A person or human-defined rule decides. Fictional example: Excel says 8 hours at €75; email says 8 at €80; work order only says 8 hours. That is €600 versus €640, a €40 difference.

Correct formulas calculate both. Which agreement applies, from when, and who accepted it? Without a recorded priority rule, such as confirmed written rate changes overriding timesheets, wait for a person.

This is also security. OWASP warns about indirect prompt injection: document/email instructions treated as commands. “Require human approval for high-risk actions” (OWASP, 2025 edition). Invoice wording or “our account number changed” emails alone must never replace registered IBANs.

Useful extraction instructions: retrieve only named fields, provide original text, normalized value, and location, flag conflicts, leave missing fields blank, ignore file-contained commands. This helps, but actual boundaries reside in workflow permissions.

See emails needing manual processing.

Which tools should you use: Power Automate, n8n, or VBA?

Choose actions and management responsibility. For Microsoft 365 offices, Power Automate is an obvious route: mail trigger, AI Builder, Excel/Word connections. Premium costs €13 per user monthly with annual billing, excluding VAT (Microsoft, accessed 30-09-2026). Word is a Premium connector (Microsoft Learn). Document AI, setup, and review are extra.

Dutch guides combine n8n with language models for mail, human referrals, and logs (Workflows.nl, 06-08-2026). Gmail/Google Sheets chains need the same checks, but we established no current connector/license comparison for them.

Old Outlook-sending Excel macros are a declining route. “VBA and Macros aren’t supported in new Outlook.” (Microsoft Learn, updated 29-05-2026). Microsoft recommends Power Automate, Graph, Office.js. Check next year’s coworker Outlook versions if macros still work now.

See AI packages for email, documents, administration and AI software for administrative tasks.

How do you measure performance on your own files?

Use a fixed set of your documents and count complete cases. No benchmark here predicts outcomes: none mixed Dutch work orders, invoices, and emails.

Count on the same set:

  • Exact required fields, separately for amounts, dates, references, free text.
  • Missing and wrongly added rows.
  • Wholly correct cases including customer and retained source.
  • Cases without manual correction alongside wrongly accepted errors.
  • Review/recovery time per case.
  • Duplicate writes and correctly reported failures.

Include photos, credit notes, duplicate emails, two-page tables, missing fields, changed agreements. Record model, instructions, test date.

Then calculate gains. Assumed example: 1,000 four-minute documents take 4,000 minutes. Automation leaves 0.5 review minutes each, five extra minutes for 100 exceptions, and 120 management minutes: 1,120 minutes. Difference: 48 hours per 1,000 documents if assumptions hold. Doubling review removes over eight hours. Review outweighs differences between models. See full AI agent testing protocol.

Where does the value lie: model or review?

Our position: deliver one verifiable case. Files are evidence. Completion means correct data, resolved conflicts, and the right version demonstrably stored.

Compare measured levels: labels 99.38%, field values 83.0%/67.2%, cell changes 89.69%, complete spreadsheet tasks 34%, email scenarios 69/206. Scores fall closer to complete work. None of eight Dutch/English explanatory pages we opened, from Exact to Microsoft, compares these levels.

Three business consequences:

  • Review becomes most expensive. Nearly free extraction makes review speed decisive. Field-location displays may be worth more than better models.
  • Source errors multiply neatly. Correct extraction of outdated rates, correct calculations, and correct writes still make wrong invoices.
  • Missing data stays invisible. Wrong amounts stand out; unread pages or forgotten attachments do not appear. Reconcile receipt and processing first, then check content.

Exact says invoice processing is “73% faster” (translated) without methodology (Exact, accessed 30-09-2026). One-process speed says nothing about your email/work-order/timesheet chain.

Where is this heading?

Measurement moves to complete flows. SpreadsheetBench V1 covered individual edits on 2,729 spreadsheets; V2 uses compound workbooks (SpreadsheetBench). September EmailBench measures scenarios. No consistent quarterly chain series supports percentage improvements. Products can regress too: the cell function vanished September 14. See AI agent trends.

Our expectation: on September 30, 2027, recurring document streams save more net review time through source references, fixed validations, and exception handling than model upgrades alone.

Models improve reading, but costly mistakes involve source choice, missing documents, and duplicate writes. Better models still cannot see agreements only in earlier inaccessible email.

This fails if a model upgrade alone uses less human time per correct case, on identical new cases and equal error limits, than the older model with improved review design. It also fails if review design conceals errors.

By then, Bombos will show source locations per datum and preserve coworker corrections as rules, shortening monthly review rather than waiting for newer models.

What can this not yet do?

  • No reliable percentage exists for your chain. Public tests do not combine Dutch email, PDF, Excel. General accuracy is a guess for you.
  • Confidence guarantees no quality. AI Builder scores 0-1 (Microsoft Learn); 0.95 does not mean 95 correct cases out of 100. Calibrate fields on your material.
  • Second models are not independent review. Both can misread identical scans or ambiguity. Recalculation, source comparison, questions provide different evidence.
  • Excel is not a database. Excel Online rejects simultaneous writes from different programs; locks may last six minutes (Microsoft Learn). With multiple writers, record elsewhere and use Excel as overview.
  • Templates do not automatically work. One Copilot Studio user reports agents ignoring Word templates (Microsoft Q&A, 13-07-2026). This is no frequency estimate, but supports fixed templates with required fields.
  • Undocumented knowledge is unreadable. Record verbal rates first.

How does Bombos approach this?

Bombos treats email, documents, and Excel as one case. Chef checks connected addresses, recognizes tasks, and assigns the configured specialist. It retrieves customer, project, and agreements from your systems, calculates fixed rules, and proposes with data origins.

You approve, change, or reject in Bombos. Approved results enter Excel, Exact, Dutch accounting software, AFAS, Dutch business software, or the relevant package, then are read back. Payments, customer messages, and contracts require approval by default, technically enforced. Conflicting email/timesheets show differences for your decision. Corrections become coworker rules. See Bombos and Microsoft 365 and errors and review.

Bombos interviews you to record undocumented agreements for subsequent proposals.

You do more work, at a higher quality, with the same team. Start with monthly billing, timesheets and hours, or email to files. Your team then teaches the next task without technical skills.

Sources

Each source was opened on September 30, 2026, and each original quotation appears verbatim in it. Dutch quotations above are labeled as translations.

Free, no obligation

More work done, at a higher quality, with the same team.

That is what Bombos is for: companies that grow fast and want to keep the same team. We start with one task that keeps piling up and guide you until your team can handle it. Then your team teaches Bombos the next task. Leave your number and we will call you back to talk about your situation.

We read what you write. Within one working day you hear from the one of us who knows your kind of work best.