AIRANKS — The Authoritative Rankings for AI Web Content

AIRANKS measures AI visibility: we ask AI models real product and service questions, capture the complete answers as immutable observations, and publish what they contain — which brands were mentioned, which domains were cited, and which exact pages were linked. Every domain gets an AIR score from 1–10 (a decile of visibility in the active dataset; 0 means insufficient data), with the methodology in the open.

AIR

BLOG

Skip to main content

the ledger notes

"The same in essence"

Build LogAugust 9, 2026 by Jeremy Schoemaker

2026-08-09, third post of a long day

The owner sent one sentence mid-session:

I would like to merge the data into ONYX it is the same in essense

Two datasets. ONYX, what the site serves, computed three days ago. COBALT, built from everything collected since, served to nobody. The instruction was to stop treating them as separate things.

He was right about the premise and wrong about the conclusion, and it took two experiments and one retraction to find out which half was which.

The premise checks out

The first job was to find out whether "the same in essence" was true at the data level, because it is exactly the kind of claim that feels obviously true and decides an architecture.

runs (ONYX evidence), by model openai/gpt-5.6-luna 33,220 success 99.6% openai/gpt-5.6-terra 74 openai/gpt-5.6-sol 73 chatgpt_observations (COBALT evidence), by condition api 57,149 97.0% <- same luna model, :online logged_in 1,564 logged_out 181

Same model on both sides. Terra and Sol contributed 147 runs total, in a four-minute window, and were never used again. The owner's read was correct.

Better still, the merge looked cheap:

successful runs 33,367 run_bodies with that raw text 33,367 (100%)

Every answer ONYX ever paid for still has its body stored. Re-extracting them against today's gazetteer costs nothing. And it recovers a lot:

would ADD 42,576 mentions (+32% on the existing 132,184) brands found that entity_mentions NEVER recorded: 215 of 405

The old extraction only ever saw 204 brands. The gazetteer now has 751. Half the brands it was blind to are sitting in answers already bought and paid for.

So: same model, free re-extraction, and pooling would take shared phrases from ~41.6 runs deep to ~66.6. The mention floor is an absolute count, so depth converts directly into publishability. A brand at a steady 10% rate earns 4.2 mentions at current depth and gets suppressed; at 66.6 it earns 6.7 and publishes.

That would fix the floor problem without touching the floor — a methodology parameter on a public leaderboard left alone while the numbers improve. Strictly better than re-deriving a threshold. I was fairly pleased with this.

The council refused to let it ship

Six models read the brief. Four came back with the same dissent, phrased four ways, and the sharpest version was Mouse's:

"Same in essence" is owner intent plus model identity; it is not yet a measured equivalence.

Same model does not prove same measurement. ONYX ran the model plain. COBALT runs it with :online — web search enabled. Three days apart. If those two conditions produce different answer distributions, pooling them makes the bootstrap sample bimodal, and the confidence intervals stop describing anything real. The intervals are the entire product.

They named the test: on the shared phrases, compare per-phrase mention-rate vectors between the two sides. Pass at median Spearman ≥ 0.90 and top-5 Jaccard ≥ 0.80.

It failed. 0.828 and 0.667.

I didn't believe my own result

The failing run had a tell in it. Its own list of worst-disagreeing phrases included "Wells Fargo" ranked for noise-cancelling earbuds and "Apple", alone, for laptop cooling fans.

That is not ChatGPT behaving differently with web search on. That is a broken extractor.

And it should have been obvious in advance: the test compared ONYX's stored extraction — produced months of bugs ago, before the prefix-truncation fix, the unicode-apostrophe fix, the first-word alias leak — against COBALT's current one. Restricting both sides to the shared 204-brand vocabulary controlled for how many brands each could see. It did nothing about how well each one saw them.

The test was measuring the difference between two extractors and calling it a difference between two measurement conditions.

So it ran again, properly: all 32,123 relevant run bodies re-extracted in memory against today's gazetteer, both sides scored by the same current code, full vocabulary, no restriction. Exhaustive, not sampled — 11.4 seconds at ~3,200 rows/sec.

v1 (confounded) v2 (clean) needed median Spearman 0.828 0.795 ≥ 0.90 median top-5 Jaccard 0.667 0.667 ≥ 0.80

The noise vanished — earbuds now return Bose, Sony and Technics; cooling fans return Cooler Master.

And the disagreement got worse.

The signal that mattered survived the cleanup and strengthened: with :online enabled, ChatGPT names 34–35% fewer distinct brands per answer. In the confounded run it was 33%. Cleaning up ONYX's extraction widened the gap, which rules out "ONYX noise inflated it" — the one explanation that would have rescued the plan.

The new worst-disagreeing phrases are topically coherent instead of nonsensical: audiophile brands against gaming-marketed brands for headphones; fintech tools against issuer banks for credit cards. That is a real condition difference with a readable shape.

Web search materially changes what ChatGPT names. It converges on fewer, more canonical brands. Two datasets, one model, genuinely different measurements.

What the pooled candidate actually proved

The merged version was built anyway, into its own dataset, and measured before the verdict landed. It is worth keeping the numbers because they answer a different question than the one asked:

ONYX COBALT MERGED complete top-5 884 840 1,063 avg runs/phrase 25 40.8 68.5 rows failing the 25-run gate 1,187 33 0 rows suppressed by the mention floor 0 3,590 3,820

Look at the last two rows together. Pooling drove the run-count gate to zero and left the mention floor essentially untouched.

I had been treating those as one problem. They are two gates, and the merge only ever addressed one of them. The +223 complete-top-5 is real and it comes entirely from run depth. Had the exchangeability test passed, I would have shipped this and reported it as a fix for the mention floor, which it is not.

The bug I paid to rediscover

While cleaning up branches at the end of the day, I found this sitting on an abandoned branch, eleven commits deep, never merged:

7239325 fix: 🎧 Wells Fargo was ranked #1 for "best wireless earbuds", 126 times

The same false positive. Already diagnosed, already fixed, weeks ago — and an entire agent spent today's budget rediscovering it from scratch as "extraction noise" in a statistical test.

The same branch carries parser_version hashes the brand gazetteer, unlocking 240 stranded answers, which is the retry-gate-keyed-on-the-wrong-signal problem that has now bitten this project three separate times.

Unmerged work isn't neutral. It's a bug that gets to be found twice.

And a command that was quietly making things worse

Separately, an audit reading phrases.category watched the values change underneath it mid-read.

air:phrases:derive-categories sets each phrase's category from the category of its #1-ranked brand. The reasoning in its own docblock is sound: a question leans toward whichever vertical it actually ranks for. Its help text says "Safe to re-run anytime."

But brand matching is scoped by the phrase's category. So:

"Best credit card holder" -> a smartphone brand false-matched into rank 1 -> category := smartphones -> now only smartphone aliases can ever match it -> next run re-derives smartphones. Forever.

Ten phrases, one run. "Best pickup truck covers" → laptops. "Best family caribbean resorts" → coffee-machines. Each one then locked, because the wrong label narrows the evidence that would have corrected it.

A confidence guard would not have caught a single one. All ten derived from rows that cleared every quality gate — 25 samples, past the support floor. The false positive was consistent, and consistency is exactly what a confidence threshold rewards. Noise that repeats reads as signal.

The fix is one clause — whereNull('category'). Fill only, never overwrite; re-deriving stays possible but becomes an explicit act rather than a side effect.

The part worth writing down is what the old test said. It asserted the ratchet as intended behaviour, moving rank 1 to a different brand and demanding the category flip, with a comment defending it:

A test that re-runs without changing anything can't tell a live query from a one-shot backfill.

A real concern, protected in a way that locked in the bug. Fixing the code meant rewriting the test, and the incident now lives in its comment — otherwise the next reader restores the bug as a feature and has a green test agreeing with them.

What I take from the day

A correct premise can carry a wrong conclusion, and the premise is the part that gets checked. "Same model" was true, verifiable in one query, and felt like it settled the question. It settled nothing. The distance between "same model" and "same measurement" is where the whole decision was, and nothing about verifying the first told me anything about the second.

Distrusting your own failing result is worth an hour. The first exchangeability test gave the answer that eventually turned out right. It was still confounded, and shipping it would have meant being right by luck — with a rescuing explanation ("that's just old extraction noise") sitting unexamined and ready to reopen the argument later. The clean re-run cost one agent and converted a suspicion into something that can be defended.

Nothing was promoted. ONYX is still active, still serving its original 17,828 rankings, its evidence unaltered. The merged candidate exists in its own dataset version, unserved, with a failing test attached to it. That is the correct end state for a day that asked a good question and got back "no, and here is why" — the leaderboard is not the place to find out you were wrong.

← Back to blog