AIRANKS — The Authoritative Rankings for AI Web Content

AIRANKS measures AI visibility: we ask AI models real product and service questions, capture the complete answers as immutable observations, and publish what they contain — which brands were mentioned, which domains were cited, and which exact pages were linked. Every domain gets an AIR score from 1–10 (a decile of visibility in the active dataset; 0 means insufficient data), with the methodology in the open.

AIR

BLOG

Skip to main content

the ledger notes

The JSON-LD product that research rewrote before we built it

Build LogAugust 12, 2026 by Jeremy Schoemaker

Today was supposed to be about shipping the API and the toolbar. It was, mostly — 90-odd agents across three workflows built both. But the turn happened in a side quest: we asked whether AIR should sell JSON-LD generation, and the research came back with answers that rewrote the product before a line of it existed.

Going in, the pitch wrote itself. We crawl every domain's AI-files, we know what AI says about each site, we detect missing schema markup — so: generate the markup, sell it as "get cited by AI more", deliver it as a paste-one-tag JS snippet like every analytics product ever. Three assumptions. Six Haiku researchers, a completeness critic, and four follow-up agents later, two of the three were dead.

The citation-lift pitch died by difference-in-differences. The vendor claims floating around this space — "3.2x more AI citations with structured data" — failed an adjudication pass outright (no controls, no disclosed n, one was circular). The only study with a real methodology, Ahrefs' DiD, measures the lift at 0 to +2.4%. That's not a product pitch, that's a rounding error with good marketing. We sell measured AI visibility with confidence intervals; shipping a product whose core claim our own evidence standard rejects would be self-immolation. So the offer reframed: gap-closure and hygiene — "here's the markup your AI-cited competitors have and you don't" — with the lift claim explicitly not made. The critic even ordered a dogfood test: run our own tracer against our own markup and publish the number even if it's null.

The JS-snippet delivery died in one sentence. GPTBot, ClaudeBot, and PerplexityBot do not execute JavaScript. A tag-manager snippet injecting JSON-LD would be invisible to precisely the crawlers whose attention we sell. The easiest-to-build, most-requested delivery mechanism in the category is, for our category, a placebo. Copy-paste server-side blocks or an API. That sentence saved a sprint.

And one landmine got fenced before anyone stepped on it. The obvious "enrichment" — publish AIR scores as aggregateRating markup — is, under Google's July 2026 fake-review policy, indistinguishable from fake review markup. Manual-action territory, for us or any customer we generated it for. It is now a banned-types list with a tripwire test (JsonLdPolicyTest) asserting zero rating types on every public route, and the confidence intervals ship as additionalProperty instead.

The product survived — as a report upsell reusing the crawler, the LLM domain summaries, and ~50 lines of spatie/schema-org, with pricing anchored to competitors ($49 pack / $99mo API / $299–499mo agency, 10x under Schema App) and kill-criteria pre-registered (<5% attach at 90 days and it dies). But it survived different: honest pitch, no snippet, rating markup banned.

Two smaller turns from the same day, for the record:

  • An adversarial reviewer caught our LLM spend cap failing open: Cache::increment() on Laravel's database cache store never creates a missing row, so the daily counter stayed at zero forever and the cap never tripped. Green tests (they run the array driver, which self-initializes), unlimited spend in prod. One Cache::add($key, 0, ttl) before the increment.
  • The same review machinery mutation-tested its way to a dispatch path that could be deleted with 796 tests staying green — the "immediately hydrate on first sight" promise, the toolbar's whole first-visit story, was wired but unproven. It has direct Queue::assertPushed coverage now.

The pattern in all of it: the cheap measurement beat the expensive build, three times in one day. The research workflow cost about 1.2M tokens and a half hour. Building the citation-lever snippet product it killed would have cost weeks, and the aggregateRating "enrichment" could have cost the domain.

← Back to blog