BlogHow-to

How to choose work automation you can trust

Can you see it work, can you stop it, does it wait for your sign-off. Three questions that make AI automation safe to trust with real work.

Every automation pitch ends the same way: hand it your work. Before you hand over anything that touches your clients, your books, or your license, you are owed answers to three short questions, and every one of them is checkable in an afternoon.

Can you see it work. Can you stop it. Does it wait for your sign-off. A tool that clears all three can be trusted with real work in real accounts. A tool that fails any one of them is asking for faith, and your business does not run on faith.

Model names, benchmarks, and demo footage are missing from that list on purpose. Those tell you what a tool can do on a good day. The three questions tell you what happens on a bad one.

Can you see it work?

Visibility means you can watch the work happen, step by step, in the actual apps and files where it happens, and then read a report that includes the failures. "Three of five carriers quoted" is a trustworthy sentence precisely because it admits the two that did not. Be suspicious of any summary that only ever says done.

The practical test: hand over one small real task and watch the entire run. In Agent FM, the session is visible while it works, and every report ends with what could not be checked, so silence never gets mistaken for success.

Watching earns something subtler than confidence, too: calibration. Five minutes of observing how a tool handles a login wall or a missing field tells you what class of task it deserves next.

Can you stop it?

Two kinds of stopping matter. The brake: you can pause or halt a run mid-task and see exactly how far it got. And the boundary: a limit you write into the ask itself, in plain language, that holds on every run including the hundredth. The second kind matters more, because it scales; you cannot personally brake a routine that runs at 6am, but a boundary holds whether you are watching or asleep.

Boundaries look like this. The ask below does real work in a real credentialing profile, and it pauses at the one step that legally matters.

Re-attest [provider, e.g. Dr. Adaeze Okafor] on CAQH ProView. Review every section, fix stale entries using the credentialing sheet, and pause at the final attestation step so I can look before you submit. Then sweep the payer portals for anything expiring in the next [window, e.g. 90 days] and log every date in the credentialing sheet. Queue the renewal filings but submit nothing until I approve. Tell me every expiration you caught and any portal you couldn't get into.
Use caseRe-attest CAQH and sweep the payer portals for expirationsRe-attest [provider]'s CAQH profile and sweep the payer portals for anything expiring.View →

And sometimes the boundary is absolute. A mortgage desk can let the rate sheets be watched daily while keeping one action entirely human, with one plain sentence.

Each morning, pull the day's rate sheets from [lenders, e.g. UWM and Flagstar] and log the benchmark scenarios we track on the Rates tab. Compare against yesterday and flag any move bigger than [threshold, e.g. an eighth]. Brief me in three lines max. Read only, never touch a lock, locks are mine. Tell me if any portal wouldn't show a sheet.
Use caseWatch the lender rate sheets and flag real movesPull today's rate sheets from our lenders and flag moves bigger than [threshold].View →

If a tool cannot honor a sentence like "locks are mine," it is not automation, it is a coworker you cannot fire mid-mistake.

Does it wait for your sign-off?

The third question separates automation you supervise from automation you apologize for. Work that stays inside your own files can run to completion. Work that goes outward, an email, a filing, a posted reply, a statement to an owner, should arrive as a draft and wait. The sign-off is not friction; it is the moment the work becomes yours.

In practice the sign-off is a batch, not a ceremony. Drafts stack in one place, each next to its context, and the pass takes minutes: approve, sharpen, hold. You do not have to touch everything; nothing outward happens without a human having had the chance to.

Prep this month's owner statements for [portfolio, e.g. the 12 single family owners]. Pull each statement and check it before it goes anywhere: maintenance charges against work orders, management fees at the contract rate, any negative balance explained. Flag anomalies in the review sheet with a one line explanation each. Draft the send-out emails but send nothing until I sign off. Tell me how many statements are clean and walk me through each flag.
Use casePrep the monthly owner statements and flag the anomaliesPrep owner statements for [portfolio]: pull from AppFolio, check the charges, flag anomalies.View →

Trust is earned per task

Start small and start real: one chore, watched end to end, report read closely. Widen the handoff as the evidence accumulates, and keep the approval beat even after you relax, because the OK costs seconds and the alternative can cost a client. If a run ever surprises you, stop it, read how far it got, and tighten the ask by a sentence. And write the boundaries down even when they feel obvious: the sentence that says read only is doing quiet work on every single run, and it is a sentence you write once.

See also: How to ask an AI assistant for real work for how to write the boundaries and report-backs into the ask itself.

Three questions, all checkable with a single real chore. Agent FM is built to answer yes to each: visible runs, stops you write yourself, and drafts that always wait for you. Get early access on macOS or Windows; when access opens, test those questions with one real chore.

Related