Named in July. Gone in August.

Last month I closed the first field report with a sentence: I will publish what the instrument finds, including the runs that show nothing interesting, because a measurement you only publish when it is flattering is not a measurement. This run is the first test of that sentence. In July, one provider named me. In August, with nothing changed on my side, it did not.

By Jordi Buskermolen5 min read
ai-systemsfield-report
Named in July. Gone in August.

Second run of a recurring measurement. Run date: August 1, 2026.

Last month I closed the first field report with a sentence I knew would get tested: I will publish what the instrument finds, including the runs that show nothing interesting, because a measurement you only publish when it is flattering is not a measurement.

This run is the test. It is the unflattering one.

In July, one provider named me. My own published GPT, found on one question, attributed correctly, with a citation pointing at the source. It was the instrument's one flattering result, and I was careful at the time not to make much of it.

In August, same question, same provider, same GPT still live: gone. Zero tracked commercial entities appeared anywhere in this run. Nothing about my content changed between the two measurements.

That is not a malfunction. That is the July article's own argument, arriving on schedule. A single appearance in a high-variance system is a draw, not a ranking. I wrote that sentence about other companies. This month it is about me.

The instrument, briefly

The method is unchanged and the July piece describes it in full. A frozen, versioned panel of questions, fired once a month at OpenAI, Anthropic and Perplexity through their APIs, with web search available, every question asked twice in fresh contexts, every raw answer stored verbatim. One standing note that matters for reading any of this: the measurement runs through the provider APIs, and API behavior differs from the consumer apps.

This run: 132 responses, no errors, nothing truncated. Cost: $13.15, against roughly $9 in July. The increase is ten percent more questions plus longer, search-heavier answers. Both numbers stated plainly, because a report built on receipts has to show its own.

The first month-over-month reading

August is the first run with a previous run behind it, which means the instrument can finally say "changed" instead of only "is".

Nearly everything changed. Twenty-one of the twenty-two questions shifted month over month, in the entities named, the agreement between passes, or both.

The flagship buyer question, who builds custom AI tools for marketing agencies, is the clearest receipt. Against July, the August answers added 28 company names that had not appeared before, and dropped 63 that had. More names vanished from the answers than survived in them. A company that took its July appearance as a ranking lost its position without doing anything. A company that appeared this month may be gone in September, for reasons equally unrelated to anything it did.

That churn number is the July variance argument with a month of history attached. Two passes in a single run already showed the lottery. Two runs a month apart show it does not settle.

What stayed stable, reported because it did

The honest surprise of this run is what did not move.

Search behavior barely changed. OpenAI searched on 68 percent of responses in July and 66 percent in August. Anthropic went from 75 to 77. Perplexity searched on every response, both months. And at the question level, the pattern holds almost exactly: the six questions OpenAI answered purely from memory, skipping search on both passes, are the same six in July and in August, to the question. Anthropic's from-memory set kept its core and shifted at the edges, one question out, two in.

Put the two findings next to each other and the actual shape of the system appears. The machinery is consistent. The output is a lottery. Whether a model searches, when it searches, how it behaves: stable, predictable, almost boring. Which names come out the other end: churn, at a scale that removes most of last month's answers.

That distinction is the finding, and it is the one a dashboard structurally cannot show you. A score summarizes the output. The output is the unstable half.

The panel grew, on record

The panel moved from v1 to v2 this run: twenty-two questions, up from twenty.

Two of the original questions were replaced, because they measured nothing useful. One returned freelance-marketplace listicles regardless of phrasing. One returned the AI vendors' own pricing pages. Neither told me anything about visibility, so they went. Two questions were added in their place, and two more joined the panel.

The bookkeeping matters more than the edit. The eighteen unchanged questions carry their July baseline forward. The changed and added ones start their history this month and are never compared against what they replaced. Every run records which panel version it asked. An instrument that quietly edits its own questions and keeps the old trend lines is not measuring anything, so the versioning is the part I would defend hardest.

What this run cannot say

Between the two runs, my own site was restructured, partly informed by what the July measurement showed. This run shows no effect from that work.

None was expected. Indexing lag means any effect belongs to September or October, not to a run fired days after the changes. I am saying this before a reader can ask it, because the alternative reading, that the August run somehow tests the site update, would contaminate the next two reports. It does not test it. The freeze exists so that the runs that can are clean.

What this does not mean

Same section as July, because the caveats did not expire.

One provider dropping one entity is not evidence the provider is broken, and my disappearance is not a grievance. It is a data point behaving exactly as the variance predicted. Two runs are still only two runs. The month-over-month churn is one interval, not a trend, and September could plausibly look like either of its predecessors.

And the stability finding cuts both ways. Consistent machinery means the system can be studied. It does not mean the output can be gamed. Which questions trigger search stayed fixed this month, and that is worth knowing. It is not a lever.

What happens next

The panel stays frozen at v2. Run 3 fires September 1. The findings publish either way.

Last month I wrote that a measurement you only publish when it is flattering is not a measurement. This was the month that sentence cost something, which is the only way to find out if it was true.

Want more of this?

I write regularly on LinkedIn about what I'm building and learning: agency growth, AI development, product judgment, and the messy reality behind making things work.

Follow on LinkedIn