the ledger notes
The JSON-LD product that research rewrote before we built it
Today was supposed to be about shipping the API and the toolbar. It was, mostly — 90-odd agents across three workflows built both. But the turn happened in a side quest: we asked whether AIR should sell JSON-LD generation, and the research came back with answers that rewrote the product before a line of it existed.
Going in, the pitch wrote itself. We crawl every domain's AI-files, we know what AI says about each site, we detect missing schema markup — so: generate the markup, sell it as "get cited by AI more", deliver it as a paste-one-tag JS snippet like every analytics product ever. Three assumptions. Six Haiku researchers, a completeness critic, and four follow-up agents later, two of the three were dead.
The citation-lift pitch died by difference-in-differences. The vendor claims floating around this space — "3.2x more AI citations with structured data" — failed an adjudication pass outright (no controls, no disclosed n, one was circular). The only study with a real methodology, Ahrefs' DiD, measures the lift at 0 to +2.4%. That's not a product pitch, that's a rounding error with good marketing. We sell measured AI visibility with confidence intervals; shipping a product whose core claim our own evidence standard rejects would be self-immolation. So the offer reframed: gap-closure and hygiene — "here's the markup your AI-cited competitors have and you don't" — with the lift claim explicitly not made. The critic even ordered a dogfood test: run our own tracer against our own markup and publish the number even if it's null.
The JS-snippet delivery died in one sentence. GPTBot, ClaudeBot, and PerplexityBot do not execute JavaScript. A tag-manager snippet injecting JSON-LD would be invisible to precisely the crawlers whose attention we sell. The easiest-to-build, most-requested delivery mechanism in the category is, for our category, a placebo. Copy-paste server-side blocks or an API. That sentence saved a sprint.
And one landmine got fenced before anyone stepped on it. The obvious "enrichment" — publish AIR scores as aggregateRating markup — is, under Google's July 2026 fake-review policy, indistinguishable from fake review markup. Manual-action territory, for us or any customer we generated it for. It is now a banned-types list with a tripwire test (JsonLdPolicyTest) asserting zero rating types on every public route, and the confidence intervals ship as additionalProperty instead.
The product survived — as a report upsell reusing the crawler, the LLM domain summaries, and ~50 lines of spatie/schema-org, with pricing anchored to competitors ($49 pack / $99mo API / $299–499mo agency, 10x under Schema App) and kill-criteria pre-registered (<5% attach at 90 days and it dies). But it survived different: honest pitch, no snippet, rating markup banned.
Two smaller turns from the same day, for the record:
- An adversarial reviewer caught our LLM spend cap failing open: Cache::increment() on Laravel's database cache store never creates a missing row, so the daily counter stayed at zero forever and the cap never tripped. Green tests (they run the array driver, which self-initializes), unlimited spend in prod. One Cache::add($key, 0, ttl) before the increment.
- The same review machinery mutation-tested its way to a dispatch path that could be deleted with 796 tests staying green — the "immediately hydrate on first sight" promise, the toolbar's whole first-visit story, was wired but unproven. It has direct Queue::assertPushed coverage now.
The pattern in all of it: the cheap measurement beat the expensive build, three times in one day. The research workflow cost about 1.2M tokens and a half hour. Building the citation-lever snippet product it killed would have cost weeks, and the aggregateRating "enrichment" could have cost the domain.