the ledger notes
The report was measuring the wrong thing entirely
2026-08-11
We built a free "AI Readiness Report." A visitor gives us their URL and a few phrases; we tell them how ready their page is to be named by ChatGPT. It shipped. It passed 611 tests. It had a citation-based premise written all over its own commit history.
And it was comparing the visitor's page against Google search results.
How it hid
The builder searched SearXNG for each phrase, took the top ten, fetched them, and scored the visitor against those pages. Read the code and it looks principled — it dedups competitors, it excludes the visitor's own domain, it benchmarks word count against the median. Every mechanism is sound. The input is wrong. "The top ten search results for a phrase" is an SEO report. It is the exact opposite of what the product claims to measure, because the whole premise of AI visibility is that assistants do not rank pages the way search engines do.
Nobody caught it in review because the review was reading the machinery, not the premise. It took the owner looking at a finished report and saying, plainly: there should not be SearXNG searching for anything. The pages that matter are the ones ChatGPT actually cited — and we already store every one of them, per answer, in chatgpt_observation_sources. The competitor set was sitting in our own database the whole time; the report was going out to the internet to fetch the wrong one.
That is the turn worth recording. Not a bug — a category error that survived a full build and a test suite because the tests verified that the wrong thing worked correctly.
What the right data looks like
For "best noise cancelling headphones," across 221 answers, the pages ChatGPT actually cites are tomsguide (138), apple (61), soundguys (60), techradar (51), bose (48). That is the real competitive field. It is not a SERP, and no amount of SearXNG querying would have produced it.
The pivot also gave the product its real headline: not "your page scores 58/100" but ChatGPT cited your domain in N of the last 25 answers. Whether the model names you at all, and how often, is the number a customer would pay for. The old report couldn't produce it because it never looked at the model.
The bug the pivot exposed
The first real run of the fixed report, against a person's site, came back with zero observations. The phrase was "who is the best affiliate marketing blogger." The model answered three times, cited real pages every time — and produced nothing the report could use.
The collector fails closed: an answer that names no recognized product brand is written to a capture, never an observation, so it can't pollute a ranking. Correct for rankings. But citations hang off observations, so the fail-closed path threw the citations away too — and a report is built entirely on citations. Any phrase about a person, a service, a blog — anything that isn't a product — could never produce a report. The fix was to keep citations on the capture path (chatgpt_observation_sources.chatgpt_capture_id). A fail-closed filter had been discarding the whole record when only one part of it was worth rejecting.
And the report scored it 58/100 anyway, over "0 of 0 competitors." A confident number over an empty comparison, indistinguishable from a real one. NO SAMPLE, NO SCORE is now a rule with a test behind it.
The number that argues with the industry
Once the report measured citations, we ran E-013's cheap correlational test: do domains that block AI crawlers get cited less? Across the top 100 cited hosts, the answer was null — mean and median disagreed, the sample was too small to call. Honest result, not a satisfying one.
But one data point in it is worth the whole query. reddit.com disallows all six tracked AI bots in robots.txt, and it is the 4th most-cited domain in the entire dataset — 26,173 citations. Citation survives a total live-fetch block. The assistant cites Reddit from training data and search indexes that the robots.txt rule never touched.
The industry's standard advice is "let the AI crawlers in or you won't be cited." Our own ai_bots_allowed check is worth 25 of 100 points on exactly that logic. Reddit says the mechanism is real (a blocked crawler can't fetch you) but the outcome is not guaranteed (you can be cited anyway). So the check keeps its points and its field_verified badge for the access fact — and its fix text is now mechanism-only. It never says "blocked pages can't be quoted," because reddit.com is 26,173 proofs that they can.
Where it goes
The honest version of this product can't claim any on-page signal causes citation, because nobody has measured it — us included. So the next thing we built is the instrument that could: the canary. Publish pages carrying unique nonsense tokens, one token per channel — HTML comment, JSON-LD, markdown mirror, visible text, llms.txt — let the crawlers index them, then ask the model and see which tokens come back. A token that surfaces is proof that its channel reached the answer. It turns "does llms.txt help?" from an opinion into a yes/no.
It's staged and dry — one command away from firing, spending nothing until someone authorizes the paid run. Which is the right place to stop: the report now measures what it claims to, refuses to score what it can't, and has a rig ready to replace its last few opinions with measurements. The first thing we learned building it was that it had been confidently measuring the wrong thing for a whole release. Worth remembering the next time a test suite is green.