An agent that drafts text is a demo. An agent that sends invoices, messages clients, and books meetings for a real business, while its owner is asleep, is a liability with a login. We shipped one at HoneyBook. This talk is the chain of problems it forced, where every fix opened the next one. "Ask before acting" can't live in the prompt, because a prompt can be talked out of anything. So a gate outside the LLM decides what runs, what waits, and what never happens. The gate defers, and now you hold a half-finished thought for eighteen hours across redeploys. Then the owner answers "sure", mid-conversation. Which pending action does that bind to? Is the sender really the owner? They approved one action out of three: do the other two run? Is a yes still a yes when the world changed since you asked? None of these problems live in the model. They live in the architecture around it. That architecture is the talk: how a business stays in control of an agent acting while nobody's watching.

Principal Engineer @ Honeybook