You test an AI agent by running it through a two-week supervised trial that widens its freedom in stages: shadow mode, then draft mode, then act-with-limits, with a set review cadence and clear kill criteria. You never hand a new agent the keys on day one, for the same reason you do not hand a new employee the company card in week one.
Key takeaways
- Trust is earned in stages, not granted on day one.
- The Trust Ladder: shadow mode, draft mode, then act-with-limits.
- Set a review cadence: daily checks in week one, proper reviews at the fortnight marks.
- Define kill criteria upfront, before you are emotionally invested.
- This staged trial is what "done properly" looks like, and it is how a serious build operates.
Call it the Trust Ladder. Each rung gives the agent a little more rope, only after it has proven the last one. It is the testing discipline that sits inside how to get started.
Rung 1: Shadow mode
The agent does the work but touches nothing. It watches real inputs and produces its output where only you can see it, side by side with what actually happened or what you would have done. You are checking one thing: does it get the task right when the stakes are zero? A few days here surfaces the obvious misses cheaply, and it lets you judge the agent's quality without any risk at all, which is the safest possible way to begin.
Rung 2: Draft mode
Now the agent's output goes into the real workflow, but as a draft awaiting your approval. Emails are drafted, not sent. CRM updates are proposed, not saved. You approve, edit, or reject each one, and every correction sharpens the SOP. This is the heart of the trial, usually a week or more, where most of the trust gets built, and it is the same human in the loop pattern you will run permanently for anything sensitive.
Rung 3: Act with limits
Once the drafts are consistently right, you let the agent act on its own, within tight limits. It can send the routine reply but not the sensitive one. It can update the record but never delete one. It operates freely inside a fenced area, and anything beyond the fence still comes to you. Destructive powers stay permanently off, not merely delayed, which is what keeps this stage safe even as you widen the agent's freedom.
Review cadence
Set a rhythm. Daily quick checks in the first week, then a proper review at the end of week one and week two: what did it get right, what did it miss, what needs tightening. Testing is not a single pass, it is a habit. Even a trusted agent gets a periodic look, because tools drift and edge cases evolve, so building the review rhythm now sets you up to keep every agent reliable long after the trial ends.
Kill criteria
Decide upfront what would make you pull the agent, before you are emotionally invested. Repeated errors on the same task after correction. Any breach of a limit. Output you would be embarrassed to see reach a customer. Write these down in advance, so the decision is a checklist, not a gut-wrench. A clear kill switch is what makes it safe to experiment, because you know exactly the line at which you would stop, and that certainty lets you trial boldly rather than nervously.
Why this protocol matters
This staged trial is what "done properly" looks like, and it is the difference between an agent you trust and one you nervously babysit. It is exactly how a serious build operates: nothing gets real responsibility until it has earned it on the ladder. Skip the ladder and you either over-trust and get burned, or under-trust and never delegate. The ladder gives you a third option: confident delegation, earned in two weeks, which is precisely the maturity you want to see in any implementation partner.
What "good enough to trust" actually means
A common question during testing is how good is good enough, and the honest answer is not "perfect." A human employee makes mistakes too; you trust them when their work is reliably good and their errors are caught before they cause harm. The same standard applies to an agent: it is ready to act with limits when its drafts are consistently right, its misses are rare and minor, and your guardrails would catch anything serious. Holding out for flawlessness is the wrong bar, because it never arrives and it stops you delegating; the right bar is "reliably good, with a safety net," which is exactly the standard you would apply to a trusted junior.
Common testing mistakes
Two mistakes undermine most agent trials. The first is skipping shadow mode and going straight to draft or live, which throws away the cheapest, safest way to judge quality. The second is never widening the rope, leaving a proven agent stuck queuing everything forever, so you never actually reclaim the time. The fix for both is to follow the ladder deliberately: start in shadow, move to draft, and once the agent has earned it, let it act within limits. Testing is not just about catching failure; it is also about recognising success and granting the freedom that turns a supervised experiment into a genuine time-saver.
Testing is how you delegate boldly
It is worth reframing what testing is really for. It is not a hoop to jump through before the fun begins; it is the very thing that lets you delegate more than you otherwise would dare. Because you know an agent will spend two weeks proving itself in shadow and draft mode before it can act, you can afford to try it on tasks you would never hand a person cold. The staged trial converts a scary "will this go wrong" into a calm "let us see, safely," which is what allows a cautious owner to end up delegating boldly. Owners who skip testing stay timid, because they never build the evidence that would let them trust; owners who test well become confident, because they have watched the agent earn it.
Writing down what you observed
A small habit makes testing far more useful: keep a short log during the trial. Note the tasks the agent handled well, the ones it fumbled, and what you changed in the instructions each time. This does two things. It turns a vague impression into evidence you can weigh against your kill criteria, so the go or no-go decision is grounded rather than emotional. And it captures the refinements to the SOP, so the agent keeps improving and you have a record of how it was tuned. The log need not be elaborate, a few lines a day is enough, but it is the difference between a trial you can reason about and one you are just squinting at.



