AIRANKS — The Authoritative Rankings for AI Web Content

AIRANKS measures AI visibility: we ask AI models real product and service questions, capture the complete answers as immutable observations, and publish what they contain — which brands were mentioned, which domains were cited, and which exact pages were linked. Every domain gets an AIR score from 1–10 (a decile of visibility in the active dataset; 0 means insufficient data), with the methodology in the open.

AIR

BLOG

Skip to main content

the ledger notes

The council ran the experiment

Build LogAugust 8, 2026 by Jeremy Schoemaker

The first real meeting of the Matrix Council settled a question that had been argued twice without data, refuted a proposal from inside the council, and — while nobody was looking — turned up Wells Fargo ranked first for "best wireless earbuds."

The question

Should we hold one browser window open and keep typing questions into it, instead of launching a fresh Chromium for every single question?

The owner's instinct was that one window would be faster and less alarming: "that would alarm me more than if we just spammed our stuff in 1 window... it made sense when we needed to trigger a new proxy."

Both halves of that turned out to be right, including the archaeology. Per-question teardown was introduced to rotate the proxy exit. Once the collector logged in, the account cookie became the identity, and tearing down the browser stopped buying anything at all.

Three rounds, and nobody rubber-stamped

Round 1: five members, five dissents. Seraph refused to sign for the right reason — "I have not first-hand re-run any capture, config, or code path cited here; everything labelled VERIFIED rests on the brief." That is the council working. The chair's job is to fix exactly that.

Round 2 added two pieces of evidence the chair gathered by hand: a Pro account, and a question the agenda had never contained.

The Pro account. Captured with a fail-closed check that refuses to save a session under a "pro" filename if the page reports planType: free. It reported pro, and probing it produced something that falsified a standing entry in the project's own constitution:

Model GPT-5.6 Sol Effort Instant

Free shows a bare "ChatGPT", which is why all 642 stored observations are unattributed. The constitution recorded "Can we label the model? No — VERIFIED." A Pro account falsifies it. Two members retracted on the spot. And "Sol" turns out to be real — absent from OpenRouter because it is a ChatGPT-product routing tier, exactly as an earlier council had inferred.

Also, on the free page but not the Pro page: soft_rate_limit_modal_trigger_message_count: 5. First hard evidence the limiter is tier-scaled, carefully labelled CONTESTED rather than VERIFIED, because it does not prove the "Too many requests" modal we actually hit is that same mechanism.

The question that wasn't on the agenda. Why are we in Temporary Chat at all? The line ?temporary-chat=true sits in the collector with no comment; the justification lives only in a spec paragraph. But the mission is what a human sees, and almost no human runs with memory off. Mouse put it better than the brief did:

"we may be indexing a clean room and calling it the city."

Then the chair ran the experiment

Three arms, same Pro account so tier was held constant, real page.on('request') counts. Nobody in this project had ever counted requests — every previous figure was inferred from counting <link> tags in a saved page.

arm requests/question seconds answered COLD — fresh browser + context per question 467.3 24.0 6/6 ARM 3 — persistent profile + full reload 468.6 23.6 5/5 WARM — one window, in-app "New chat" 68.3 24.2 6/6

Three things fell out of that table.

The council's estimates were an order of magnitude low. Rounds 1 and 2 argued over 30–50 requests per question. The real number is 467. The <link> tags were only what the initial HTML declares, not what the SPA fetches after boot.

Wall clock is identical. 24.0 versus 24.2 seconds. The time is dominated by ChatGPT generating the answer, so the cold boot was ~400 wasted requests per question buying exactly zero speed.

Arm 3 was refuted — and it was Trinity's own proposal. A persistent on-disk profile with a full navigation per question costs 468.6 requests, statistically identical to cold. So the saving does not come from browser reuse, and it does not come from HTTP cache. It comes only from not navigating. That distinction was invisible to every argument made in the first two rounds, and it would have stayed invisible without the measurement.

The load-bearing test

Batching is only safe if an in-app new conversation isolates as well as a fresh incognito context does. If it doesn't, batching corrupts the dataset while looking like a speedup — the worst available outcome.

So the warm arm asked the same control phrase at position 1 and again at position 6, with marathon headphones, laptops and cooling mattress pads in between.

Position 1: Soundcore Space Q45, Sony WH-CH720N, Soundcore Space One Pro. Position 6: Anker Soundcore Space Q45, Sony WH-CH720N, Sennheiser ACCENTUM Plus.

No mattresses. No laptops. No marathons. Same top two products, brand Jaccard 0.60 — ordinary ChatGPT nondeterminism, not contamination.

What every member retracted

  • Trinity: "cold incognito per question is necessary to guarantee isolation" → warm is safe. And "30–50 requests" → 467.3.
  • Mouse: "every Nx claim is a model, not a measurement" → verified. "Cold boot is dead."
  • Morpheus: model attribution unavailable → available on Pro.
  • Tank: "lower boot time may be offset by rate-limiting" → there is no throughput gain to offset; the win is request volume, not speed.

Four retractions in one round. The doctrine that rewards changing your mind on evidence appears to work.

The ruling

Trinity voted consensus. Morpheus, Mouse and Tank dissented — but none of them dissented on the question asked. All three objected to the same two subordinate items, so the chair ruled consensus on the session model and productive deadlock on the rest:

  • Rotation N is unmeasured. Trinity picked 50; Mouse objected that it came "from a hat," and Mouse is right. Nothing in the data shows degradation across six questions. Named experiment: a soak run logging requests, heap and modal state per question until something moves.
  • Temporary Chat versus normal chat is unmeasured, and it is a mission-validity question rather than a performance one. Note the isolation test above was run inside temporary chat, so it says nothing about memory-on behaviour.

And then, sideways, Wells Fargo

While the council deliberated, a parsing question led somewhere else entirely. Checking how many rankings carry a brand from the wrong category turned up one that could not possibly be legitimate: a credit-card brand ranked on headphone phrases, 126 times.

Traced end to end:

sentence : | **Best overall for Android** | **Sony WF-1000XM5** | Great sound, ANC, multipoint... | alias : "WF" -> Wells Fargo (credit-cards), requires_category = NULL method : gazetteer result : Wells Fargo, final_rank #1, "best wireless earbuds"

Sony's entire true-wireless line is WF-*. An unguarded two-letter alias matched inside the model number of the most-recommended earbuds on earth, and Wells Fargo came first for wireless earbuds one hundred and twenty-six times.

The interesting part is not the bug. It is that the detector built to catch this cannot see it. air:brands:audit-aliases reports two collisions, both "MacBook" — because it checks aliases against phrase vocabulary, and "WF" never appears in the phrase "best wireless earbuds." It appears in the answer.

That is the same shape as the rate-limit modal from this morning: a real signal, sitting in plain text, invisible to the instrument built to find it. Twice in one day.

Five aliases with demonstrated cross-category damage are now category-scoped, in the gazetteer rather than only in the table so a re-seed cannot undo it. Deliberately not swept by length — the codebase's own doctrine is explicit that the predicate is "is this an ordinary English word," not "is it short," and a length sweep would flag BMW, HP, JBL, AWS and OXO, all of which are safe.

And the edit moved parser_version from 857ab3d3a59a to fcca6eeec261 — which is this morning's fix working on the very next gazetteer change. Before today, editing brand data moved nothing, and every capture that edit would have rescued stayed unreachable forever.

← Back to blog