Automate the task that passes five tests: it is frequent, rule-describable, low blast radius, measurable, and genuinely hated. Tasks that tick all five give you 80% of the benefit for 20% of the effort. Everything else can wait. Call it the First Five test, and run every candidate task through it before you build a thing.

Key takeaways

  • The First Five test: frequent, rule-describable, low blast radius, measurable, and hated.
  • A task must pass all five to be a great first automation; any one alone misleads you.
  • The sweet spot is the overlap, the frequent, simple, safe, measurable, hated task.
  • Inbox triage, invoice chasing, research, follow-ups, and reporting usually pass.
  • The test is also about what to leave alone, not just what to grab.

The mistake owners make is automating the interesting task instead of the valuable one. Here is how to tell them apart, and it pairs with the prioritisation grid in what to delegate to AI first.

The First Five test

A task is a great first automation only if it passes all five checks. Frequent: it happens daily or weekly, not twice a year, because frequency is where the hours hide. Rule-describable: you can explain how it is done in plain English, because if you cannot describe it, you cannot hand it over. Low blast radius: if it goes wrong and you catch it, nothing catastrophic happens, so you have a safe place to build trust. Measurable: you can tell whether the agent did it well, because no feedback signal means no way to improve. Hated: you would genuinely rather never do it again, because reclaiming a task you loathe feels like a pay rise. Five yeses and you have found your first automation.

Why all five matter together

Any one on its own misleads you. A frequent task that is not rule-describable will frustrate you. A rule-describable task with a huge blast radius is a scary place to start. A hated task you do once a year is not worth the setup. The magic is in the overlap: the frequent, simple, safe, measurable, hated task. That is the 80/20 sweet spot, and there is usually one screaming for attention in every business. The test works precisely because it forces all five conditions at once, filtering out the tasks that look appealing on one axis but fail on another.

The usual winners

Run the test and the same tasks tend to pass for most owners: inbox triage, invoice chasing, prospect research, meeting follow-ups, weekly reporting. They are frequent, describable, safe, measurable, and widely loathed, which is exactly why they are the classic first builds. If you are unsure where to begin, start with whichever of these makes you groan the loudest, because the "hated" axis is a surprisingly reliable guide to where the relief will be greatest.

Frequency: where the hours hide

Of the five tests, frequency deserves special attention, because it is the one that determines how much time you actually save. Automating a task you do fifty times a week is transformative; automating one you do twice a year is a rounding error, however clever it feels. When you rank your candidate tasks, weight frequency heavily, because a boring daily task almost always beats an interesting rare one on hours reclaimed. This is the single biggest reason the "exciting" task is so often the wrong first choice: it is usually rare.

Blast radius: your safety margin

Low blast radius is what makes a first automation safe to experiment with. You want a task where, if the agent gets it wrong and you catch it in a draft, nothing bad happens, no client offended, no money moved, no record destroyed. That safety margin is what lets you run the supervised trial with confidence and learn how the agent behaves before trusting it with anything weightier. Starting with a low-blast-radius task is not timidity; it is the sensible way to build the trust that lets you delegate bigger things later.

Naming the framework helps

There is a reason to give this a name and use it out loud: the First Five test turns a vague instinct into a repeatable filter your whole team can apply. Instead of arguing about what to automate, you run each candidate through five clear questions and the answer becomes obvious. Named frameworks like this also make the decision defensible and consistent over time, so the tenth task you automate is chosen the same disciplined way as the first. It is a small habit that keeps your whole automation programme pointed at value rather than novelty.

What fails the test (and that is fine)

Closing your biggest deal fails on rule-describable and blast radius. A once-a-year board report fails on frequent. A delicate client apology fails on blast radius and description. These are not automation targets, and forcing them is how people get burned. Let them stay human. The First Five test is as much about what to leave alone as what to grab, which is the honest and safe way to think about it, and it echoes what to keep human.

Turning the test into an order

Passing the First Five test tells you a task is a good candidate, but when several pass, you need an order, and the order is simple: rank the passing tasks by hours saved, and start at the top. The heaviest, most frequent, most hated task that clears all five checks is your first build, because it delivers the biggest, most obvious relief and the strongest proof that this works. Then work down the list one at a time, proving each over a fortnight before starting the next. Resist reordering by what sounds impressive; order by hours reclaimed, because momentum in an automation programme comes from felt relief, not clever demos, and nothing builds confidence like getting your worst weekly chore off your plate first.

Measurable is the test people skip

Of the five checks, "measurable" is the one owners most often overlook, and skipping it quietly undermines everything. If you cannot tell whether the agent did the task well, you cannot refine it, you cannot trust it, and you cannot prove it was worth doing. So before you automate a task, decide how you will judge success: hours saved, error rate, response time, whatever fits. Capturing a simple before figure takes minutes and turns a vague "this feels helpful" into a hard "this saves me four hours a week," which is exactly the evidence that justifies expanding. Measurement is not bureaucracy here; it is the feedback loop that makes the whole thing improve.