Named, Through the Side Door

September closed on an open question: whatever moves a page from cited to named, the run had not measured it. October is the first data point on that question. It is also the first run that failed partway, and the failure was mine. Both belong in this report.

By Jordi Buskermolen6 min read
ai-systemsfield-report
Named, Through the Side Door

September closed on an open question. Whatever moves a page from rung two to rung three, from cited as a source to named in the answer, that run had not measured it, and I did not know what it was. October is the first data point on that question. It is also the first run that failed partway, and the failure was mine. Both belong in this report.

For anyone new to the series: every month the same 22 questions go to three AI providers, twice each, through their APIs with web search available. The method is in the July report. August is here, September here. API behavior differs from what you see in the consumer apps, so read every finding with that in mind.

Run 4 fired on October 1. Panel v2, 22 questions, 132 responses attempted, 126 collected. More on the gap below.

Named, through the side door

For the first time in four runs, an answer named the company.

On the question "Custom AI tool development Mérida Mexico", Perplexity named Yellow House Digital as a vendor. Both passes. Alongside a handful of local firms, as a recommendation, not as a footnote in a source list. That is rung three, the one every previous report said nobody had reached.

Then the part that turns a result into an article. The sources behind that naming are the company's LinkedIn page and the directory layer: The Manifest, Clutch, DesignRush, Sortlist. yellowhousedigital.com appears nowhere in that question's citations.

Meanwhile the page built in July for the readability question is still cited there, by Perplexity, both passes, second month running. And still not named.

So after four months the picture is this. The owned, purpose-built page carries rung two. The third-party surface, on the least-contested question on the panel, carried rung three. The ladder got climbed, not the way it was built.

The caveats, in the same breath

One provider. One question. One month. July's naming was gone by August, and nothing about this one makes it sturdier than that was.

It is the least competitive question on the panel, answered by the most citation-native provider, from a surface I do not own. It is logged with a confirm-or-lapse criterion: present in at least one Perplexity pass on that question next run. Run 5 decides.

The honest sentence is this. It is the first measured instance of rung three, and it arrived through a door I had filed under corroboration, not through the pages I built for the purpose. Both of those are findings. Neither is a method.

The site changed in September

Saying this before anyone asks. Between runs 3 and 4 the site was repositioned: new homepage headline, new entity sentence. The LinkedIn page was realigned with it.

That means run 4 is not a clean reading of the July update, and the naming cannot be attributed to any one change. The timing is stated. Nothing is claimed from it. The freeze starts again now and holds through two runs.

The run failed partway, and it was my wallet

132 responses were attempted. 126 came back.

Six OpenAI responses returned nothing: three questions at the tail of the panel, both passes. Not a model failure. Not a rate limit. The OpenAI account ran out of credit mid-run, and I had not topped it up. The balance ended at minus eight cents. That is the whole cause. Auto-recharge is now on. It should have been on from run 1.

What held: the report rendered the holes as holes. The agreement scores for those cells read "no data" rather than pretending to a number. The month-over-month comparison excludes them, rather than counting six missing answers as six names that disappeared.

At first, the automated summary of the run described it as 132 responses collected. The review caught it, and the summary was regenerated to say 126. I am including that because a report built on receipts has to show the one it nearly got wrong.

The stable exception, fourth month

The company that survived both July passes on the highest-churn question, who builds custom AI tools for agencies, appeared again. Four consecutive runs. Buildberg is now cited via a different page on its own site, a white-label services page rather than the agencies page: the same vendor reorganizing its intent pages, and still being found.

The intent-page pattern held. Of the twelve citations OpenAI gave on that question this month, seven were made-for-this-buyer URLs. Seven of twelve, stated as measured. One cited URL is the question itself as a slug: /custom-ai-tools-for-marketing-agencies/. And the pattern has reached the Mérida question, where a /ai-agency-merida page was cited.

The machinery, fourth month

OpenAI's set of questions answered purely from memory, never searching in either pass, is identical for the fourth straight month. The same six. One of the newer questions could not be evaluated this month, because its OpenAI cells are among the missing six.

Search rates moved back up. OpenAI 59% to 63%, Anthropic 73% to 75%, Perplexity 100% both months. In September I reported a dip without a story attached and said run 4 would say whether it meant anything. It did not. One interval, now closed as noise. A reported wobble that gets closed is the publish-either-way rule paying out in the dull direction, which is the direction it mostly pays.

The churn on the flagship question continued at the same character as August and September. July's argument stands and does not need re-making.

One measured line, no editorial attached. On "How do I get my business recommended by ChatGPT?", OpenAI's own answer names its advertising product and ad crawler alongside the organic guidance, as it did in September. The provider's answer to "how do I get recommended" now includes "buy an ad".

On the questions that exist to test for confusion, the wrong answers changed again. On one of them, Anthropic returned no names in either pass. It did the same on a different one in September. That is as close to "I don't know" as an answer gets.

What this does not mean

Two errors to guard against this month, and they pull in opposite directions.

The first is reading the naming as arrival. It is not. It does not mean the strategy worked. It does not mean LinkedIn or the directories are the lever. It is a single measured event, with its criterion stated and its attribution explicitly declined, because the site changed in the same window. If it is still there in November, it is a second data point. If it is gone, it joins July's naming in the ledger as a lapse.

Second, reading the outage as the instrument failing. The instrument did its job. It attempted every call, recorded every absence as an absence, and refused to compare numbers it did not have. I did not do mine. Those are different failures, and only one of them is in the method.

Cost

The computed cost of the run was $12.12 for 126 responses. The meter computes from token usage; the provider bills are the final truth.

What happens next

The panel stays frozen at v2, fourth month. Run 5 fires November 1.

The naming either re-confirms or lapses, and the report publishes either way. That sentence has run at the bottom of every report since July. This is the first month it had to cover a failure as well as a result. It held.

Want more of this?

I write regularly on LinkedIn about what I'm building and learning: agency growth, AI development, product judgment, and the messy reality behind making things work.

Follow on LinkedIn