Automate documents, Excel files, and emails with AI as one chain: AI reads email and attachments and extracts data, fixed rules calculate in Excel, a person approves the complete case, and only then is the result written and read back.
This page explains reliable capabilities by file type, recent benchmark error rates, and what they mean. Then it connects email, PDF, and Excel through a nine-step example and a €40 conflict no model can resolve for you.
The short answer
- AI reads and sorts well. Aid4Mail’s best model correctly labeled 99.38% of 1,600 emails in September 2026 (Aid4Mail, 17-09-2026).
- Extraction is good but imperfect. Structured Output Benchmark’s best exact field-value scores were 83.0% on text and 67.2% on images, despite nearly always correct format (arXiv, 28-04-2026).
- Multistep Excel remains weak. SpreadsheetBench 2 found 89.69% correct cell modifications but only 34% wholly correct financial tasks (arXiv, 29-06-2026).
- Technical success differs from completion. EmailBench recorded 99.7% successful technical actions but only 69 of 206 scenarios (arXiv, 25-09-2026).
- Excel’s standalone
=COPILOT()function stopped working September 14, 2026. Sidebar Copilot remains (Microsoft Support, accessed 30-09-2026). - Power Automate’s Excel connector retrieves only 256 rows by default. More silently remain missing until pagination is enabled (Microsoft Learn, accessed 30-09-2026).
- At 99% per-field correctness, twenty fields give only 81.79% probability of a flawless document. Measure cases, rather than fields.
What can AI reliably do with documents, Excel, and email now?
In 2026, recognition and extraction are reliable, simple spreadsheet edits moderately reliable, and long dependent action chains still weak. This distinction outweighs model choice.
For email, sorting is strongest. Aid4Mail tested fifteen models on 1,600 messages, reaching 99.38% best labeling (Aid4Mail, 17-09-2026). Limits: 200 synthetic base messages in eight languages excluding Dutch; its maker sells Aid4Mail.
For documents, extraction measurement varies. ExtractBench covers 370 documents, 4,869 pages, 67 types, assessing field values and evidence locations (ExtractBench, snapshot 11-09-2026). Best scores were 95.91 field-value F1 and 58.11 for identifying correct words. Benchmark maker LlamaIndex sells the winning system.
For Excel, Copilot edits workbooks directly, plans, and chats (Microsoft Support, accessed 30-09-2026). Microsoft says “Review, edit, and verify anything Copilot creates before you rely on it.” (Microsoft FAQ).
For Word/PDF output, AI often is unnecessary. Word Online fills fixed templates and converts to PDF with a 10 MB input limit (Microsoft Learn, accessed 30-09-2026).
How reliable is AI by file type?
Each type has typical errors and checks. The capabilities column is documented; the other two are our design.
| Input/output | Documented capability | Typical error | Built-in check |
|---|---|---|---|
| Text-based digital PDF | Field/table extraction (Microsoft Learn) | Previous-page amount, wrong total | Page/source for critical fields; recalculate totals |
| Scan/photo/handwriting | Vision plus structured extraction (arXiv, 18-08-2026) | 8 read as 3, tilted image, missing page | Request unreadable material again; no guesses |
| Long PDF tables | Repeated rows with evidence positions (ExtractBench) | Missing, duplicate, merged rows | Compare counts/totals; test two-page tables |
| Fixed-column Excel | Connector reads/writes rows (Microsoft Learn) | Wrong column, lost leading zero, duplicate key | Fixed schema/types, unique row ID |
| Excel formulas/tabs | Multistep editing (SpreadsheetBench 2) | Wrong formula logic, broken reference | Test, recalculate, check prohibited changes |
| Email with attachments | Trigger plus document action (Microsoft Learn) | Old quoted agreement, wrong attachment | Retain message/attachment IDs, sender/date; escalate conflict |
| Word report/PDF output | Template filling/conversion (Microsoft Learn) | Empty field, wrong customer attachment | Required fields, fixed template, recipient check |
For CSV, use ordinary import software. Define delimiter, encoding, decimal separator. Keep customer/item numbers as text. Turning CSV into an AI image only adds errors.
What do benchmark error rates mean?
Know the measured unit: fields, tasks, or complete cases. These genuine scores belong to different tests, rather than one ranking.
| Test | Measurement | Score | Valid interpretation |
|---|---|---|---|
| ExtractBench, 11-09-2026 | Document field values | F1 95.91 | Not flawless-document percentage |
| Structured Output Benchmark, 28-04-2026 | Exact image field values | 67.2% | 32.8 percentage points not exact |
| SpreadsheetBench V1, Google, 10-03-2026 | Gemini in Sheets tasks | 70.48% | 29.52% unsuccessful |
| SpreadsheetBench 2, 29-06-2026 | Claude Opus 4.6 financial tasks | 34% | 66% not wholly correct |
| Aid4Mail, 17-09-2026 | Gemini 3.7 Flash email labels | 99.38% | 0.62% wrong labels |
| EmailBench, 25-09-2026 | Complete email scenarios | 69 of 206 | 66.5% unsuccessful |
Sources: ExtractBench, Structured Output Benchmark, Google, SpreadsheetBench 2, Aid4Mail, EmailBench.
Three findings: larger units score lower; labels nearly always succeed while full tasks often fail. Format proves little: “models achieve near-perfect schema compliance”, but image field values were inexact roughly a third of the time (arXiv, 28-04-2026). Test design matters: only 4 of 35 open-model configurations exceeded F1 0.5 on 100 academic transcripts (arXiv, 18-08-2026).
No independent complete-chain percentage exists for Dutch invoices, work orders, and emails together. Promised figures are self-measured or invented.
Why do good fields not guarantee a good file?
Errors accumulate. A calculation, rather than measurement: twenty fields each 99% correct yield 0.99²⁰ = 81.79% flawless documents, roughly one in five with an error.
Real errors correlate: poor scans affect several fields. This is no forecast but explains spreadsheet outcomes. SpreadsheetBench 2 reports “89.69% Modification but only 34.00% Accuracy” (arXiv, 29-06-2026). Nearly nine in ten cell changes were correct while two-thirds of whole workbooks were not.
Microsoft-affiliated EmailBench researchers say “valid tool execution is not equivalent to task completion” (arXiv, 25-09-2026). Technical actions succeeded 99.7%; scenarios about a third. Limits include synthetic environments, May runs, and partly AI assessment. Still, the success-checkmark/completed-work gap is familiar in offices.
Ask vendors how many complete cases passed without correction and how many errors were wrongly accepted, alongside field accuracy.
What can Copilot in Excel currently do?
Sidebar Copilot edits workbooks, creates formulas/analyses, and helps plan (Microsoft Support, accessed 30-09-2026). It assists workbook users rather than emptying mailboxes overnight.
Two current details many guides miss:
- The cell function is gone. “Starting September 14, 2026, the COPILOT function is no longer available in Microsoft Excel.” (Microsoft Support). Existing values remain; recalculation produces
#NAME?. Build no new workflow on=COPILOT(). - Editing requires automatic calculation. Features depend on licenses and Microsoft 365 settings (Microsoft FAQ).
Google reports 70.48% successful Gemini in Sheets tasks on full SpreadsheetBench (Google, 10-03-2026). This vendor measurement is six months old and uses English Excel-forum questions. Microsoft publishes no comparable per-edit score.
Have Copilot propose a formula, test zero hours, empty rates, decimal commas, negative adjustments, duplicate rows, then fix it in place. Fixed logic belongs in fixed formulas rather than newly generated answers each time.
How do you connect email, PDF, and Excel?
Center the case. A maintenance company’s monthly billing receives an emailed price agreement, PDF work order, and Excel timesheet. This is a design example, rather than a customer case.
| Step | Action | Output |
|---|---|---|
| 1. Receive | Store email/files with time and unique ID | Receipt record |
| 2. Separate | Read email text, PDF document, Excel table | Versioned sources |
| 3. Extract | Project, date, hours, rate, reference with locations; leave missing values blank | Candidate field values |
| 4. Identify | Match project/customer by fixed identifiers beyond names | One match or exception |
| 5. Compare | Email agreement, work order, hours, master data | Visible discrepancies |
| 6. Calculate | Fixed formulas, duplicate/period/total checks | Review report |
| 7. Present | Draft, sources, exceptions to authorized coworker | Approval, correction, rejection |
| 8. Write | Approved version only, linked to case ID | Result or visible error |
| 9. Read back | Compare stored outcome, check no second record | Completion or recovery |
Steps 1 and 9 are often omitted, creating silent errors. Excel’s connector defaults to 256 rows, limits files to 25 MB, and timeout retries can duplicate entries (Microsoft Learn). Check prior write success before retrying.
Step 3 has another trap: AI Builder “returns all extracted values as strings” (Microsoft Learn, updated 14-01-2026). Convert amounts from text with Dutch decimal commas. Incorrect page ranges produce partial data according to the same source.
Software connections are covered in AI and existing systems.
What if email, work order, and Excel contradict each other?
A person or human-defined rule decides. Fictional example: Excel says 8 hours at €75; email says 8 at €80; work order only says 8 hours. That is €600 versus €640, a €40 difference.
Correct formulas calculate both. Which agreement applies, from when, and who accepted it? Without a recorded priority rule, such as confirmed written rate changes overriding timesheets, wait for a person.
This is also security. OWASP warns about indirect prompt injection: document/email instructions treated as commands. “Require human approval for high-risk actions” (OWASP, 2025 edition). Invoice wording or “our account number changed” emails alone must never replace registered IBANs.
Useful extraction instructions: retrieve only named fields, provide original text, normalized value, and location, flag conflicts, leave missing fields blank, ignore file-contained commands. This helps, but actual boundaries reside in workflow permissions.
See emails needing manual processing.
Which tools should you use: Power Automate, n8n, or VBA?
Choose actions and management responsibility. For Microsoft 365 offices, Power Automate is an obvious route: mail trigger, AI Builder, Excel/Word connections. Premium costs €13 per user monthly with annual billing, excluding VAT (Microsoft, accessed 30-09-2026). Word is a Premium connector (Microsoft Learn). Document AI, setup, and review are extra.
Dutch guides combine n8n with language models for mail, human referrals, and logs (Workflows.nl, 06-08-2026). Gmail/Google Sheets chains need the same checks, but we established no current connector/license comparison for them.
Old Outlook-sending Excel macros are a declining route. “VBA and Macros aren’t supported in new Outlook.” (Microsoft Learn, updated 29-05-2026). Microsoft recommends Power Automate, Graph, Office.js. Check next year’s coworker Outlook versions if macros still work now.
See AI packages for email, documents, administration and AI software for administrative tasks.
How do you measure performance on your own files?
Use a fixed set of your documents and count complete cases. No benchmark here predicts outcomes: none mixed Dutch work orders, invoices, and emails.
Count on the same set:
- Exact required fields, separately for amounts, dates, references, free text.
- Missing and wrongly added rows.
- Wholly correct cases including customer and retained source.
- Cases without manual correction alongside wrongly accepted errors.
- Review/recovery time per case.
- Duplicate writes and correctly reported failures.
Include photos, credit notes, duplicate emails, two-page tables, missing fields, changed agreements. Record model, instructions, test date.
Then calculate gains. Assumed example: 1,000 four-minute documents take 4,000 minutes. Automation leaves 0.5 review minutes each, five extra minutes for 100 exceptions, and 120 management minutes: 1,120 minutes. Difference: 48 hours per 1,000 documents if assumptions hold. Doubling review removes over eight hours. Review outweighs differences between models. See full AI agent testing protocol.
Where does the value lie: model or review?
Our position: deliver one verifiable case. Files are evidence. Completion means correct data, resolved conflicts, and the right version demonstrably stored.
Compare measured levels: labels 99.38%, field values 83.0%/67.2%, cell changes 89.69%, complete spreadsheet tasks 34%, email scenarios 69/206. Scores fall closer to complete work. None of eight Dutch/English explanatory pages we opened, from Exact to Microsoft, compares these levels.
Three business consequences:
- Review becomes most expensive. Nearly free extraction makes review speed decisive. Field-location displays may be worth more than better models.
- Source errors multiply neatly. Correct extraction of outdated rates, correct calculations, and correct writes still make wrong invoices.
- Missing data stays invisible. Wrong amounts stand out; unread pages or forgotten attachments do not appear. Reconcile receipt and processing first, then check content.
Exact says invoice processing is “73% faster” (translated) without methodology (Exact, accessed 30-09-2026). One-process speed says nothing about your email/work-order/timesheet chain.
Where is this heading?
Measurement moves to complete flows. SpreadsheetBench V1 covered individual edits on 2,729 spreadsheets; V2 uses compound workbooks (SpreadsheetBench). September EmailBench measures scenarios. No consistent quarterly chain series supports percentage improvements. Products can regress too: the cell function vanished September 14. See AI agent trends.
Our expectation: on September 30, 2027, recurring document streams save more net review time through source references, fixed validations, and exception handling than model upgrades alone.
Models improve reading, but costly mistakes involve source choice, missing documents, and duplicate writes. Better models still cannot see agreements only in earlier inaccessible email.
This fails if a model upgrade alone uses less human time per correct case, on identical new cases and equal error limits, than the older model with improved review design. It also fails if review design conceals errors.
By then, Bombos will show source locations per datum and preserve coworker corrections as rules, shortening monthly review rather than waiting for newer models.
What can this not yet do?
- No reliable percentage exists for your chain. Public tests do not combine Dutch email, PDF, Excel. General accuracy is a guess for you.
- Confidence guarantees no quality. AI Builder scores 0-1 (Microsoft Learn); 0.95 does not mean 95 correct cases out of 100. Calibrate fields on your material.
- Second models are not independent review. Both can misread identical scans or ambiguity. Recalculation, source comparison, questions provide different evidence.
- Excel is not a database. Excel Online rejects simultaneous writes from different programs; locks may last six minutes (Microsoft Learn). With multiple writers, record elsewhere and use Excel as overview.
- Templates do not automatically work. One Copilot Studio user reports agents ignoring Word templates (Microsoft Q&A, 13-07-2026). This is no frequency estimate, but supports fixed templates with required fields.
- Undocumented knowledge is unreadable. Record verbal rates first.
How does Bombos approach this?
Bombos treats email, documents, and Excel as one case. Chef checks connected addresses, recognizes tasks, and assigns the configured specialist. It retrieves customer, project, and agreements from your systems, calculates fixed rules, and proposes with data origins.
You approve, change, or reject in Bombos. Approved results enter Excel, Exact, Dutch accounting software, AFAS, Dutch business software, or the relevant package, then are read back. Payments, customer messages, and contracts require approval by default, technically enforced. Conflicting email/timesheets show differences for your decision. Corrections become coworker rules. See Bombos and Microsoft 365 and errors and review.
Bombos interviews you to record undocumented agreements for subsequent proposals.
You do more work, at a higher quality, with the same team. Start with monthly billing, timesheets and hours, or email to files. Your team then teaches the next task without technical skills.
Sources
Each source was opened on September 30, 2026, and each original quotation appears verbatim in it. Dutch quotations above are labeled as translations.
- ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction, LlamaIndex, configuration snapshot 11-09-2026
- SpreadsheetBench, V1/V2 project, ongoing
- Gemini in Google Sheets just achieved state-of-the-art performance, Google, 10-03-2026
- Excel Online (Business), Connectors, Microsoft Learn, current
- Using flows with Excel, Microsoft Learn, updated 04-03-2025
- Use a document processing model in Power Automate, Microsoft Learn, updated 14-01-2026
- Get started with Copilot in Excel, Microsoft Support, current
- Frequently asked questions about Copilot in Excel, Microsoft Support, current
- COPILOT Function, Microsoft Support, ended 14-09-2026
- SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows, arXiv, 29-06-2026
- EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks, arXiv, 25-09-2026
- The Structured Output Benchmark, arXiv, 28-04-2026
- Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application, arXiv, 18-08-2026
- Automatic invoice processing, Exact, undated
- Building an n8n/Claude AI customer-service bot, Workflows.nl, 06-08-2026
- LLM01:2025 Prompt Injection, OWASP, 2025 edition
- VBA, Macros, and Custom Flows Alternatives, Microsoft Learn, updated 29-05-2026
- Word Online (Business), Connectors, Microsoft Learn, current
- AI Email Classification Benchmark: 15 Models Tested Inside Aid4Mail, Fookes Software, 17-09-2026
- Power Automate pricing, Microsoft, accessed 30-09-2026
- Copilot Studio: Unreliable Word Template Population, Microsoft Q&A, 13-07-2026
