The AI Quality Bar

Marketers credit AI with raising the quality of their output, and in the same breath name AI slop as their biggest fear. Both instincts are right, because they describe the same tool with and without a quality system around it. This is the quality system: how the four failure modes differ, the five checks that catch them, and why review stopped being the tax on production and became the product.

By Jordi Buskermolen8 min read
ai-systemsagencyplaybook
The AI Quality Bar

How to use AI heavily and never ship slop - the review standards that separate AI-leveraged agencies from AI-damaged ones

For agency owners, studio leads, and independents whose delivery is now AI-assisted end to end. The industry tells a strange double story. AI has moved from curiosity into daily marketing practice, and most of the people running those teams will tell you it lifted their output - and in the same breath they name "AI slop", generic, hollow, confidently wrong AI content, as the thing that worries them most. In ACAM's 2026 Australian AI in Marketing Benchmark - 126 CMOs and senior marketing leaders across 12 industries, run with Kantar - AI slop ranked as the number one operational concern. Note who was surveyed: those are your clients, not your peers. Both instincts are right, because they describe the same tool with and without a quality system around it. This is the quality system.

First, name the enemy precisely, because "slop" gets used lazily. AI-assisted work fails in four distinct ways, and each needs a different check:

Generic - technically correct, completely interchangeable. The strategy deck that could be about any client in any industry. The post that reads like the average of everything ever written on its topic - because, mechanically, it is. Genericness is AI's default gravity: the model produces the most probable version, and the most probable version of anything is the one with no client, no place, no scar, and no point of view in it.

Fabricated - specifics that were invented to sound right. Statistics without sources, citations to papers that don't exist, features a product doesn't have, quotes nobody said. The most dangerous failure because it wears the costume of the most valuable content: concrete detail.

Hollow - structurally complete, substantively empty. Every section present, every transition smooth, nothing actually decided or claimed. AI is exceptional at producing the shape of finished work, which is exactly what makes hollow work slip past tired reviewers: it pattern-matches to done.

Off-voice - competent, but not yours and not the client's. The brand that sounds like every brand; the em-dash-studded, "delve"-flavored register that readers now recognize on sight and silently discount. Off-voice is the mildest failure and the most corrosive, because it leaks in one paragraph at a time until the agency's output has no fingerprint.

Underneath all four, the economics just inverted, and the whole conversation depends on grasping it: AI collapsed the cost of producing work and thereby moved all of the value into what survives review. When drafts are nearly free, nobody pays for drafts - clients pay for the judgment that separated the shipped version from the forty possible versions. Review used to be the tax on production; it is now the product. Agencies that staff and price accordingly thrive on AI; agencies that treat review as the old spot-check ritual ship slop at scale and discover that their clients, who also have AI, can produce their own slop for free.

The bar itself - five checks per deliverable

A quality bar is a short, written, per-deliverable-type standard - not a mood, not "senior eyes", not a vibe. Five checks, adaptable per work type:

1. The specificity check - against genericness. Could this deliverable belong to a different client? Read it hunting for the client's actuals: their numbers, their market, their constraint, their words from the brief. A strategy with no facts that are theirs is a template wearing their logo. The operational version: every deliverable must contain a set number of client-specific anchors (facts, quotes, data points) that no other client's version could contain - counted, not felt.

2. The verification check - against fabrication. Every fact, figure, name, citation, and claim traced to a source outside the AI conversation. Not "does it look plausible" - plausible is precisely what fabrications are. The rule that makes it practical: the drafter marks every specific as they write ([verified: source] or [check]), so the reviewer inherits a map instead of a minefield. Unverifiable specifics get sourced, softened, or cut - never shipped on confidence.

3. The position check - against hollowness. What does this deliverable actually claim, recommend, or decide - in one sentence? If the reviewer can't extract it, there's nothing inside the shape. Every client-facing piece must survive the question "so what should they do?" with a real answer, including the uncomfortable ones AI structurally avoids: this option over that one, stop doing X, we were wrong about Y.

4. The voice check - against the fingerprint fade. Read one paragraph aloud against the client's (or agency's) voice reference - the short document of register, phrases-we-use, phrases-we-never-use that every account should have anyway. AI-flavored tells get hunted explicitly: the em-dash cascades, the "it's not just X, it's Y" constructions, the adjective stacking, whatever this quarter's model-isms are. Keep the tell-list current; it drifts with the models.

5. The stakes check - proportional scrutiny. Not every deliverable needs all four checks at full depth; every deliverable needs a decision about depth. Internal notes: light. Client-facing narrative: standard. Anything published, legal-adjacent, numbers-bearing, or reputation-carrying: full verification plus second reviewer. Writing the tiers down is what prevents the quiet failure mode where deadline pressure silently demotes everything to "light".

The named human - review as ownership, not activity

This is the single most important sentence in any AI quality system, and it isn't about AI: every deliverable has one named person whose professional signature it carries. Not "the team reviewed it" - a name. That person read every word, ran the five checks at the deliverable's tier, and would defend the work in the client's boardroom without mentioning what drafted it. The name goes in the project tool next to the deliverable, visibly.

Why naming does what process alone can't: diffuse review responsibility fails against AI output specifically, because AI work looks finished - it pattern-matches to reviewed - so everyone assumes someone else's pass was the real one. A named signature makes the assumption impossible. It also produces the correct emotional relationship to the tool: the drafter treats the AI as their extremely fast junior whose work they own, not a colleague whose work they forward. And it answers the client-trust question before it's asked - the anti-slop clause of your AI policy ("no deliverable reaches a client without named human review") is only enforceable if the name exists.

A corollary keeps naming honest: reviewing your own AI-assisted work counts only at the lowest tier. Above it, the reviewer and the prompter are different people - not because the drafter is careless, but because the drafter has already read the AI's framing forty times and can no longer see it.

Upstream of review - the quality moves that happen before drafting

Even a good review system loses to bad inputs. Three upstream disciplines cut the slop rate before the first token:

Feed it the client, not the topic. The genericness failure is mostly an input failure: a prompt containing only the topic gets the internet's average; a prompt containing the brief, the client's data, their voice reference, and the last deliverable gets their draft. The rule of thumb that changes everything: the context you give should be longer than the output you want.

Demand positions in the prompt. Hollowness is also upstream: ask for "an overview" and receive one; ask for "a recommendation with the strongest counter-argument addressed" and the shape arrives with a spine. Prompts that require the AI to choose produce drafts that require the human only to check the choice.

And keep the human's scar in the loop. The one input AI cannot supply is the thing your client is paying for: what you learned the last three times this went wrong. The strongest AI-assisted deliverables are drafted after the senior spends ten minutes dictating the actual insight - the position, the trap to avoid, the thing this client specifically needs to hear - and the AI builds around it. Insight in, structure out beats structure out, insight never.

Making it stick - the system around the standard

A quality bar that lives in a document and not in the week is decoration. The installation, briefly: the five checks become a per-type checklist (one page: blog post, strategy deck, campaign copy, code, report - each with its tier and its anchor-count); the checklist lives in the workflow, not the wiki - a template field, a PR requirement, a pre-send gate; overrides are data - every time a reviewer catches a fabrication or guts a hollow section, it's logged in one line, because the log is your prompt-improvement backlog and your training material; the tell-list and the checklists get a quarterly refresh, since model behavior drifts and so do the failure patterns; and the bar gets priced in - review is the product now, so proposals include it as named senior time rather than hiding it in margin, which also gives the agency its cleanest answer to "AI should make this cheaper": it made the drafting cheaper; you're paying for what we didn't let through.

The one-page version

Name the four failures - generic, fabricated, hollow, off-voice - because each needs its own check. Run the bar per deliverable: specificity counted in client-anchors, every fact traced outside the AI conversation, a real position extractable in one sentence, the voice read aloud against a reference, and a written tier deciding how deep - with deadline pressure never silently demoting the tier. Put one name on every deliverable, a person who'd defend it in the boardroom; above the lowest tier, prompter and reviewer are different people. Fix quality upstream too: feed the client not the topic, demand positions in the prompt, and put the senior's scar in before drafting instead of hoping it appears after. Then install it - checklists in the workflow, catches logged, tells refreshed quarterly, review priced as the product it now is.

AI collapsed the cost of drafts and moved all the value into what survives review. The agencies that understand this are not using AI less than the slop-shops. They're using it more - behind a bar.

Want more of this?

I write regularly on LinkedIn about what I'm building and learning: agency growth, AI development, product judgment, and the messy reality behind making things work.

Follow on LinkedIn