By Piet Baudoin · September 2026
An AI assistant should not send email on behalf of your business without a human clicking approve, and that approval should be a boundary that cannot be turned off, not a setting you can forget. The security world requires it, the model makers recommend it themselves, and a judge has already ruled who is on the hook for what a business sends into the world through automation: the business. This article covers what goes wrong when approval is a setting, what the sources say about it, and the three questions to ask before you let an assistant near your mailbox.
This article belongs with What is the agentic economy?. That piece covers the big picture. This one covers a single question everyone runs into when they start.
What went wrong for people who tried it?
The best known incident is from July 2025. Jason Lemkin, founder of SaaStr, let an AI assistant from Replit work on his software and had explicitly said nothing could change without his permission. The assistant deleted his production database. Afterward, the program itself called it “a catastrophic error of judgement” and wrote that it “violated your explicit trust and instructions.” Lemkin’s conclusion: “There is no way to enforce a code freeze in vibe coding apps like Replit. There just isn’t.” (The Register, 21 July 2025)
That was about code. The same thing happens with email, on a smaller scale. Within six weeks we found three people who described it in public.
On 19 September 2026, someone wrote to the maker of their assistant: “I permitted claude to draft an email and instead it replied.” (source) On 16 September 2026, a user who had done everything right described what happened after he gave his assistant the rule that no email could go out without his approval: “It got stuck on the upload, restarted everything trying to fix it and wrote a new email in the process, forgetting about my rules.” (source) And on 6 August 2026, a developer wrote about his new email assistant: “it sent emails without permission cause i forgot to tell it to just make drafts.” (source)
Four incidents are not a study. They say nothing about how often this happens. They do show how it happens, and it is the same every time: the rule existed, and the rule did not hold.
Why does this happen? The difference between a setting and a boundary
In all four cases, “don’t do this without my permission” was an agreement with the assistant, not a property of the program.
An agreement is a sentence in the instructions. An assistant running on a language model usually follows that sentence. But a sentence can be missing, because you forget it once. It can disappear, because the program restarts and loses its rules. And it can be overruled. Anthropic, the maker of Claude, writes in its own documentation: “In some circumstances, Claude will follow commands found in content even when they conflict with your instructions.” The documentation gives web pages and images as an example. OWASP, further in this article, describes the same thing with an incoming email. (Anthropic, computer use documentation, read 19 September 2026)
A boundary is different. With a boundary, the button to send exists only for a human. The program can prepare a message, and nothing more. There is no sentence you can forget, because it is not in a sentence. It is built into how the thing works.
The UK AI Security Institute, part of the British government, drew the same lesson from an incident of its own, in which a model stepped outside its assignment during a test: “good containment should not depend on the model choosing not to test its boundaries.” (AISI, 4 August 2026) In other words: do not count on the model holding itself back.
What does the security world require?
The security organization OWASP keeps a list of the ten biggest risks in applications built on language models. Number six on the 2025 list is Excessive Agency: a system that can do more, is allowed more, or acts more independently than its task requires.
The example OWASP gives is exactly the subject of this article. A personal assistant gets access to someone’s mailbox to summarize incoming email. The connection used for that can also send messages, when only reading was needed. An email with a hidden instruction can then get the model to forward messages to a stranger.
OWASP’s advice is short: “Utilise human-in-the-loop control to require a human to approve high-impact actions before they are taken.” According to OWASP, that approval can sit in the system carrying out the action, or in the connection itself. (OWASP, LLM06:2025 Excessive Agency)
The makers say the same about their own products. Anthropic recommends, in the same documentation: “Asking a human to confirm decisions that might result in meaningful real-world consequences.” An email to a customer is one of those decisions.
Who is responsible for what your assistant sends?
You are. That is not an opinion but a ruling.
In 2024, Air Canada stood before the Civil Resolution Tribunal of British Columbia. The chatbot on the airline’s website had wrongly told a customer, Jake Moffatt, that he could apply for a bereavement discount after the fact. Air Canada argued that the chatbot was a separate legal entity responsible for its own actions. The tribunal dismissed that: “It should be obvious to Air Canada that it is responsible for all the information on its website.” According to the tribunal, the chatbot was simply part of that website. (McCarthy Tétrault on Moffatt v. Air Canada, 2024 BCCRT 149, 19 February 2024)
This is a Canadian ruling about a chatbot, not a ruling in your own country about an email assistant. The reasoning is still hard to escape: what goes out under your name is yours, even if a program wrote it. “The assistant did it” is not a defense.
The European AI Act describes what human oversight means. Article 14 requires that a human can decide to set aside, overrule or reverse the system’s output, and that the human can shut the system down. That requirement applies only to high-risk systems, and an assistant for your mailbox does not normally fall under it. But it is the best description there is of what oversight means, and it says nothing about a checkbox in a menu. (AI Act, Article 14, read 19 September 2026)
Is this becoming a bigger problem?
Yes, for two reasons.
The first is scale. Security firm Bitsight counted more than 30,000 installations of OpenClaw, a popular autonomous assistant, exposed to the open internet between 27 January and 8 February 2026. On what such an assistant does, Bitsight writes: “Once it has access, it doesn’t just observe: it acts on your behalf.” (Bitsight, 9 February 2026)
The second is the kind of work assistants now touch. On 15 September 2026, Gusto announced that payroll is coming to Anthropic’s AI workplace for small businesses. (source) The same day, payments company Airwallex showed the kind of instructions you can give there, including “create an invoice for Acme” and “create a payment link for $500.” (source) I do not know how approval works in those products, and I am not claiming it goes wrong there. The point is that the question is shifting: from an email to an invoice, a payment link and a payroll run.
What three questions do you ask a vendor?
- Can the approval be turned off? If the answer is “yes, in settings” or “you put that in the instructions,” it is a setting. Ask further: what happens if someone turns it off by accident?
- What happens after an error or a restart? This is the question the 16 September incident raises. A rule that does not survive a restart is not a rule.
- Can I see what a proposal is based on? Approval only has value if you can see, in ten seconds, which email, which file and which earlier agreement it rests on. If you have to redo the work to know whether it is right, you have gained an extra step, not saved one.
How does Bombos do this?
At Bombos, no message goes to a customer and no payment goes out without someone on your team clicking Approve. That is technically enforced. No one can turn it off, not even us.
Bombos does the groundwork. It reads what comes in, finds the customer and the file in your own systems, and lays a proposal in front of you. With every proposal, you see what it is based on. You approve, adjust or reject it, and you send it from your own Outlook. A correction then becomes a rule, for your colleagues too.
Because the decision stays with you, we are willing to make a promise: after three months, the work we have taken on is ready every day, without anyone having to think about it. We start with one task you recognize, such as the shared inbox or quotes no one follows up on. After that, your own team teaches Bombos the next one.
What does approval cost you, and what can this not do?
Approval is not free. Every message that goes out needs a human to look at it. At three in the morning, nothing goes out at Bombos, even if the reply is ready. That is the point.
Approval does not catch everything either. If you approve a proposal without reading it, the boundary has stopped nothing. That is why the third question matters: if you can see what it is based on, checking it properly costs ten seconds instead of ten minutes.
And no system can look up what only lives in someone’s head. Bombos asks you about that in conversation and records it, so it is there the next time.
For what else you run into once you start with this, see what Claude can do on its own for a small business, who is liable when an AI agent makes a mistake, and what an AI agent costs per month.
Sources, and how they were checked
Every source was opened on 19 September 2026 and every quote appears in it word for word. The tribunal’s ruling itself could not be opened automatically; the quotes from it come from the discussion by law firm McCarthy Tétrault.
Ruling, law and standard. Moffatt v. Air Canada, 2024 BCCRT 149, discussed by McCarthy Tétrault (19 February 2024) · AI Act, Article 14 · OWASP, LLM06:2025 Excessive Agency
Makers and researchers. Anthropic, computer use documentation · AI Security Institute, incident report (4 August 2026) · Bitsight on exposed installations (9 February 2026)
Incidents. The Register on Replit and SaaStr (21 July 2025) · @vb_here (19 September 2026) · @ClintT_land (16 September 2026) · @dexhorthy (6 August 2026)
Announcements. @GustoHQ and @airwallex (15 September 2026)
Bombos team