For agency owners and operators: digital, creative, marketing, or development shops. Written by someone who ran a 30-person agency for twelve years and now builds AI systems for a living. I know exactly which hours you are losing, because I lost them too.
Every agency owner I talk to right now is somewhere on the same journey: convinced AI should be saving the team time, unsure where to start, and quietly suspicious of the last pilot that impressed everyone in the demo and then died in a drawer. The pattern behind most of those dead pilots is identical. The task was chosen by excitement instead of by fit. Someone automated the flashy thing, or the founder's pet annoyance, or the hardest problem in the building, because those are the ones that come to mind first.
The right first automation is almost never the most impressive one. It is the most repeatable one, and repeatability can be scored. This guide is the triage I run before any agency automation project: an inventory of where the candidates hide, six criteria to score them, and the rules that keep the winner from becoming another drawer pilot.
Two ground rules before the method. First: one task. Not a "transformation", not a toolkit. One task, done properly, producing proof. The second automation is easier to choose, cheaper to build, and better trusted after the first one visibly works. Second: measure the baseline now. Before anything is built, know how long the task takes today and how often it happens. Without the baseline you will never be able to say whether it worked, and "we cannot tell" is how budgets stop.
Where the candidates hide in an agency
Walk through a normal week and list every task that fits this description: same kind of input, same kind of thinking, same kind of output, over and over. In most agencies the list clusters in eight places.
Intake and triage. Every inquiry that arrives gets read, assessed, and answered, or worse, does not. Deciding which inquiries deserve real attention is pattern work: fit, budget signals, red flags, urgency.
Briefs and scoping. Every brief gets checked, or should get checked, for the same missing pieces: undefined deliverables, hidden functionality, absent content plans, impossible timelines. Agencies that skip this check pay for it in every project. It is pure pattern recognition.
Proposals and their guardrails. The assumptions, exclusions, and boundary language that should wrap every quote, written fresh each time, or copied from the last proposal and quietly drifting.
Reporting and client updates. The weekly and monthly translation of "what happened" into client language. Data in, narrative out, same structure every cycle. For many agencies this is the single largest block of automatable hours.
Research. Prospect research before a pitch, competitor scans, audience background, "find out about this company before the call". Hours of tab-hopping producing a summary someone could have specified.
First drafts. Not final creative: first drafts of the recurring formats. Post variations from an approved concept, meta descriptions, alt text, case-study skeletons from project notes, internal briefs from call transcripts.
QA and checklists. Pre-launch checks, brand-compliance passes, link checks, spec verification. Judgment-light, attention-heavy, error-prone precisely because it is boring.
Admin and money. Time-entry summaries, invoice preparation, the overdue-invoice chase and its delicate tone decisions, contract boilerplate checks.
Write your own version. Ten to fifteen tasks is typical. Include who does each one and your gut estimate of hours per month. That list is the raw material; now we score it.
The FIRST-6 score
Rate every candidate 0, 1, or 2 on six criteria. Total per task: 0 to 12.
1. Frequency. How often does it happen? Daily or with every project (2), weekly-ish (1), monthly or rarely (0). Automation pays per repetition; a quarterly task almost never earns its build cost first.
2. Pattern strength. Does the task have the same shape every time: same inputs, same steps, same output structure? Nearly identical each time (2), a common core with variations (1), reinvented per instance (0).
3. Judgment load. How much genuinely contextual human judgment does it need? Little, rules and patterns cover it (2); a human should review, but the heavy lifting is mechanical (1); the judgment IS the task (0). Be honest here. This criterion kills more bad automation ideas than any other.
4. Example supply. Do you have real past examples, inputs and good outputs, lying in your email, drive, and project tools? Dozens easily gathered (2), a few with digging (1), none, or all confidential beyond use (0). Examples are what the tool gets built and tested against; no examples means you are speculating.
5. Pain-to-hours. Combine the monthly hours with how much the team hates it. Big hours AND actively dreaded (2), meaningful hours or real dread (1), neither (0). Dread matters beyond the math: the first automation must produce felt relief, because felt relief is what buys you the team's enthusiasm for automation number two.
6. Blast radius. If the tool gets it wrong and nobody catches it, what happens? Internal inconvenience only (2), awkward but recoverable client moment (1), real client, money, or reputation damage (0). Your first automation should fail privately, not publicly.
Reading the scores. Anything at 9 or above is a genuine first candidate. Six to eight: second wave, often good once a review step is designed in. Under six: leave it, and notice why it scored low, because the reasons are instructive. The tasks agency owners most want to automate (final creative, client relationships, strategic recommendations) reliably score 0 to 2 on judgment load and blast radius. That is not AI skepticism. That is the triage doing its job.
A worked example from the list above. Inquiry triage typically scores: frequency 2 (every week, every inquiry), pattern 2 (same assessment each time), judgment 1 (human confirms, machine pre-sorts), examples 2 (your inbox is full of them, with known outcomes), pain 1 to 2 (rarely loved, often skipped, and skipping costs), blast radius 2 (an internal recommendation a human reviews). Ten to eleven out of twelve, which is why "assess this inquiry against our fit criteria" is so often the right first tool for a service business. Run the same arithmetic on your own list. Your winner may differ, but it will be defensible, which drawer pilots never were.
The pilot rules
The triage picked the task. These five rules keep the pilot alive.
Assistive first, autonomous maybe never. Version one produces a draft, a score, or a recommendation that a human reviews. It does not act alone. This is not timidity; it is sequencing. The human-review phase is where you learn the tool's failure modes cheaply, and where the team learns to trust it honestly. Some tools should graduate to acting alone; many are permanently better as assistants. Let the review phase decide, not the ambition.
The machine judges, the code enforces. Wherever the task involves numbers, thresholds, or rules (scores, limits, calculations), those live in ordinary code, not in the AI's discretion. AI interprets and drafts; deterministic logic decides and enforces. Tools built this way behave consistently. Tools that leave the math to the model behave like moods.
One owner, thirty days. One named person uses the tool for a month as part of their real workflow. Not "the team can try it", which means nobody does. At day thirty, compare against the baseline you measured: hours saved, quality of outputs, how often the human overrode it and why.
Overrides are data. Every time the reviewer corrects the tool, that correction is the most valuable artifact the pilot produces. It either tunes the tool or teaches you the task had more judgment in it than the triage scored. A tool nobody ever overrides is suspicious in the other direction: check whether anyone is actually reviewing.
Kill or scale, out loud. At day thirty: keep it and make it standard, fix the one thing the pilot exposed, or kill it and say why. All three are wins. The drawer is the only failure. And when it works, the second task is already on your scored list, and the second buying decision takes a tenth of the deliberation, because now you know what the process feels like when it is done right.
What not to automate first, and what that list is really telling you
The recurring low scorers, and the reason each fails the triage: client relationships (judgment IS the task, blast radius maximal), final creative quality (pattern strength near zero; if it were patterned, it would not be the product), strategy and recommendations (context-heavy judgment, and clients are paying specifically for the human owning it), anything client-facing without a review step (blast radius), and hiring and people decisions (judgment, ethics, and law all in one).
Notice what the low scorers share: they are the work your clients actually pay a premium for. The triage, run honestly, delivers a conclusion that surprises exactly no one who has run an agency and everyone who sells AI transformation: automate the work around the craft, never the craft. The hours you reclaim from triage, reporting, scoping checks, and first drafts do not replace the judgment work. They fund it. That is the entire economic argument for agency AI in one sentence, and it is also why the drawer pilots died: they aimed at the craft and missed, instead of aiming at the grind and hitting.
The whole method on one page
List every task with the same shape every time (intake, scoping, proposals, reporting, research, first drafts, QA, admin) with hours and owners. Score each on the FIRST-6: frequency, pattern strength, judgment load, example supply, pain-to-hours, blast radius. Take a 9 or above, measure its baseline, and pilot it assistive-first with one owner for thirty days: numbers in code, overrides logged, verdict out loud. Keep the craft human; automate the grind around it. One task, one month, one visible win, and the second automation will choose itself.
