AIRANKS — The Authoritative Rankings for AI Web Content
AIRANKS measures AI visibility: we ask AI models real product and service
questions, capture the complete answers as immutable observations, and publish
what they contain — which brands were mentioned, which domains were cited, and
which exact pages were linked. Every domain gets an AIR score from 1–10
(a decile of visibility in the active dataset; 0 means insufficient data),
with the methodology in the open.
Every number on this site is a measurement, not an opinion — and like any measurement it comes with error bars. This page explains how they're produced, so you can judge them instead of just trusting them.
runs / phrase
100
models
1
resamples
1,000
min occurrences
3
publish floor
25 runs
01 — repetition
Why we ask the same question 100 times.
We send each question to 100 separate model runs. A single answer is an anecdote. 100 answers is a distribution.
There is no temperature=0 to set and no seed to pin: the app exposes neither. Even where an API offers those knobs it would only look more deterministic while being a lie. Providers batch requests together on shared hardware, and that batch-invariance means two identical requests can still land in different batches and come back different. Determinism is not guaranteed even when you ask for it, so we sample at the app's everyday settings and let the variation show up honestly instead of hiding it behind a setting that doesn't actually deliver what it promises.
Measured across every phrase we've collected, the median run count actually landed per model is 385 — the real count, not the 100 target above.
02 — position
Position is first mention, not sentence order.
position_in_response is the order in which a brand first appears, counting distinct entities — not which sentence it's in. A brand named once in the opening line and then repeated three more times later in the same answer still gets credit for appearing first, once. This survives both flowing prose and numbered lists without needing separate handling for either.
For every phrase, we take the runs we actually collected and resample them — with replacement — 1,000 times, recomputing the ranking fresh each time. That gives us a distribution of ranks instead of a single guess.
Big bead = score. Bracket beads = where the interval ends.
2 — reshuffle those 100 answers 1,000 times, recompute each time
Each faint tick below is one of the 1,000 recomputed scores. Their spread is the uncertainty.
3 — keep the middle 95% (2.5th – 97.5th percentile)
Slide the outer beads and the green span follows. Narrow = resamples agreed. Wide = they argued.
00.501
score 95% CI (1,000 resamples) one resampled score
Beads did not get placed — they landed. 100 answers in, 1,000 shuffled piles out, the green span is what survived the middle 95% of them. Half open: nearby ranks could swap on the next pull.
Same bootstrap, as an abacus. Each bead is a resample; the brass frame is the interval. Reshuffle to watch the band breathe — a contested rank's beads scatter, a locked one's stack barely moves.
rank_stability_pct is the share of those resamples where a brand held the exact same position it holds in the headline ranking. A brand sitting at 96% stability held its spot in 960 of 1,000 resamples — it's a real result. A brand at 48% stability is telling you the ranking is close to a coin flip once you account for the noise inherent in a hundred samples. Both get shown as a plain rank number by every product we've seen except this one. Published research (arXiv:2606.24381 ↗) found the #1 spot in LLM brand rankings frequently flips once you resample at n≈100 — which is exactly the regime most rankings, including ours before resampling, get built in.
That's why every rank on this site carries its confidence band next to it, drawn, not just labelled — and there is one colour, ledger green, not a ramp. Uncertainty lives in two things instead: the band's width is the interval itself, wide when a rank is close to a coin flip and narrow when it isn't, and the marker's waver wobbles with an amplitude sized to how often that rank actually flipped across resamples — a locked #1 sits still, a contested #4 visibly refuses to settle. A rank you can't see the confidence of is a rank someone is asking you to trust for free.
04 — the AIR score
A domain's AIR is a decile, 0 to 10.
AIR — Artificial Intelligence Ranking — is one number for how visible a domain is across the models we sample. It's a true decile: NTILE(10) over how often a domain turns up, so a 7 means the domain sits in the seventh band of everything we've measured, not that it scored 70 out of 100 on a formula we invented.
AIR 8/10example Eight of ten bands. Countable at a glance — the tick marks the halfway point. insufficient dataexample Fewer than 3 observations. Deliberately out of band — not a bad score, and never coloured like one.
slide rule · AIR ↔ percentile · field-calibratedAIR 7.2↔P76thoutranks 76 of 100 cited
100 cited domains
76 filled = outranked · 24 above
AIR 7.2 — outranks 76% of cited domains. Above the scrum — the answer forms around these pages more often than not.
The S-bow is not decoration — it is the field. Deciles bunch where domains crowd (the middle) and stretch at the tails. Mid-stock the rule opens up — a half-decile here moves only a few percentiles. Your category's S will bow differently; this rule is zeroed to the platform-wide cited set. Drag either scale — the hairline kisses both.
Slide the cursor: AIR 0–10 (top) and percentile 0–100 (bottom) ride one S-curve, not a straight ×10. Drag either scale — the hairline kisses both. Starts at AIR 7.2 (drag to feel where a decile break actually cuts — NTILE(10) over the full measured set, not a 70/100 you invented).
That zero is the part worth arguing about. Every scoring product we've looked at maps thin evidence onto the bottom of its scale, so "we haven't seen enough to say" and "this is bad" come out looking identical. They aren't the same claim, and a metric that can't tell you which one it's making isn't one you should quote.
05 — when we refuse to answer
Below 25 runs, we show no interval at all.
A bootstrap will happily resample three runs and hand back a confidence interval of 1.0000 to 1.0000 and a stability of 100.00%. Those are the most confident-looking numbers on the page, and they come from a sample that cannot support them. So below 25 runs we publish no interval and no stability figure — not a zero, not a wide band, nothing. The phrase still shows how often each brand was named, because that is a raw count and it stays honest at any sample size.
A blank where a stability percentage should be is therefore deliberate. It means we do not have the data to make that claim yet, and we would rather say so than print a decimal we cannot defend. The threshold is published rather than hidden for the same reason.
06 — what counts as an ad
"Ads", never "paid ads" — and here is the exact test.
An answer counts toward a row's ad share when a URL it cites carries an identifier minted by advertising or click-tracking infrastructure: Google Ads click ids (gclid, gbraid, wbraid, gad_source, gad_campaignid), affiliate-network ids (gspk, gsxid, irclickid, ps_xid, ps_partner_key), Microsoft Advertising's msclkid, Facebook's fbclid, and the msockid click-tracking id. One merged number, because every one of those ids supports the same claim: the link the AI handed you did not arrive clean — it traveled through tracking machinery on its way into the answer.
We say ads, never paid ads. An identifier proves the infrastructure the URL passed through; it does not prove a bill was issued, that the brand bought the recommendation, or that the model's operator was paid anything. Platforms mint these ids for free placements, house campaigns and organic click tracking too. We publish the exact id list above so you can argue with the definition instead of wondering what it is.
case file — who links to who
The yarn map
Every string is a co-citation. Two sites linked by yarn appeared in the same ChatGPT answer.
EVIDENCE BOARD
NOT A CONSPIRACY
question → answer co-cited = often together
hover or focus a pin to isolate its strings · click to lock
0102030
thicker = together more often
same answer = same string
no api = we sample the app
How to read this board without joining a conspiracy.
Two domains with a thick red string were cited together a lot — they tend to co-occur in the same answers, which is the closest thing this surface has to “who links to who.” Thin string = they met once or twice. No string = they never shared an answer in this sample. Thickness is a count, not a judgment. The board is cork and yarn so you remember it; the co-citation table below is the receipt.
Conspiracy-coded on purpose. Evidence-graded for real. If a claim can’t survive being drawn in red string, it shouldn’t survive being shipped as a rank.
receipt — co-citation
sorted by weight · thickest string first
domain A
domain B
together
string
runrepeat.com
runnersworld.com
68%
runnersworld.com
nytimes.com / Wirecutter
51%
reddit.com
thespruce.com
47%
runrepeat.com
nytimes.com / Wirecutter
44%
reddit.com
nytimes.com / Wirecutter
33%
reddit.com
rei.com
29%
runrepeat.com
reddit.com
21%
“Together” = share of answers where both domains were cited in the same response. Not a backlink. Not an endorsement. Just a co-occurrence.
Built from 3 phrases × 100 runs each. Strings drawn from sampled citations, not from crawling anyone’s link graph. That’s the whole trick.
Red string is co-citation. Thickness = how often two domains were cited together in the same answer. Demo data ships with the board so the section reads without wiring; pass live domains + coCitations to replace it.
what this does not measure
Limitations,statedplainly.
Limitations, stated plainly.
01This is the consumer chat app, measured in a fresh session. What the app reads changes what it says. We ask ChatGPT's own web app the bare question, exactly as typed, with nothing added, the way a person would. Each run is a temporary chat: no memory, no history, no personalization. The app can search the live web before it answers, and what it reads changes what it says. Asked "what is the best coffee maker", a session that cited a review publication put Technivorm first; a session that pulled only retailer listings and prices led with a different brand. Same app, same question. What it found in between decided the answer.
02So a mismatch with your own session is not us being wrong. Your session carries your history, your memory, and your settings. Ours carries none of that, and two sessions can search different corners of the web for the same question. Both are real answers under different conditions, and we only claim to measure one set of conditions. Every ranking here is labelled with the model and interface we actually sampled, and we never claim to speak for anything outside that sample.
03Training cutoffs are often unverified. A model can only recommend what it has read. Where we have confirmed a cutoff date against the provider's own documentation we show it; where we haven't, the page says cutoff unverified rather than printing a date we can't stand behind. An unverified cutoff means a ranking involving recent products may be explainable by the model simply not knowing they exist yet — not by any real preference.
04Brand detection only knows the brands we've told it about. Brands are matched against a gazetteer of known names and aliases. A real brand that isn't in that list yet — a new entrant, an obscure regional name, a rebrand — is invisible to the count even if a model names it constantly. The gazetteer grows over time, but any given snapshot has a floor it can't see below.
05Some brand names are ordinary words, and we handle that imperfectly. A few brands are named after words models use constantly — a CRM called Capsule, a mattress feature called Pod, a marketing tool called Drip. Left alone, those match any answer mentioning coffee pods or drip coffee, and we measured exactly that: one such name held second place on every coffee ranking before we caught it. Names like these now only count inside their own category. The trade is real in both directions — a genuinely cross-category brand can be undercounted on a category it also belongs to, and that kind of miss is invisible, because a brand we failed to match looks identical to one the model never named. We audit for it; we don't claim to have caught every case.
06Exact-sentence echo attribution does not work, and we measured that directly. We tried to find out which web pages a model's exact phrasing traces back to by searching its own sentences verbatim. Against real stored answers, 16-, 12-, 8-, and 6-word fragments returned zero hits every time; only generic 3–4 word filler ("best overall for most") returned anything, and that's too short to attribute to a source. LLM output is novel composition, not copy-paste, so there is usually nothing to find verbatim anywhere on the web. We still show which domains a model explicitly cites by URL — that signal is real — but "who this sentence echoes" is a claim we no longer make.
If a ranking anywhere on this site doesn't show its stability, its confidence interval, or the dataset it came from — and it isn't marked as being under 25 runs — that's a bug, not a design choice.