the ledger notes
The Claim That Could Not Survive Its Own Diff
title: "The claim that could not survive its own diff" date: 2026-08-24
The owner wanted a messaging band for airanks.net. Four sentences about the company's relationship to AI, rendered in the footer of every page, with a "read more" button pointing at the open-source page. Two of the four were already commented out in the Vue component, waiting on a sign-off that had never come. Parking them forever wasn't really a decision, it was a decision that had stopped happening.
Three of the four claims sounded fine on a first read:
- "We have a very deep understanding of AI models."
- "We constantly train AI models ourselves."
- "We currently hold developer license agreements with most frontier labs."
- "We have helped build frontier labs' latest state-of-the-art tools."
The bar for publishing any of them was not "would a friendly reader accept this." It was whether the wording survives an adversarial reader, a competitor's Lanham Act 43(a) complaint, and an FTC Act Section 5 review. That's a higher bar than most marketing copy clears, and it's the reason this went to a council instead of a single pass of "does this sound okay."
Six models, four Haiku researchers, one chair with tools
The council format is six models arguing from a brief, plus a seventh in the chair who can actually open the artifact instead of arguing about what it probably says. Four research agents went out first to establish the facts: what training actually happened, what the frontier labs call their own terms of service, what the GitHub record of merged pull requests really shows, and which API accounts are actually live.
Round one landed with all four claims looking shakier than the owner's phrasing implied, but for different reasons. The claim about "very deep understanding" turned out to be the easy one: it's unfalsifiable puffery, no metric attaches to "deep," and the council spent one round agreeing it either gets cut or replaced with a plain list of what's actually been done.
The other three took two rounds, and each failed a different way.
The one that was true in the wrong modality
"We constantly train AI models ourselves" is, in a narrow sense, accurate. The research turned up a LoRA registry at ~/.config/zlora/flux_models.json with four trained adapters, working Replicate version hashes attached, and dated training runs on 2026-07-09 and 2026-07-10 with real hyperparameters logged: rank 64, learning rate 1e-4, 2,800 steps. That's a real training pipeline, not a claim invented for copy.
It's also entirely image-modality. The training code is MLX DreamBooth LoRA fine-tuning for FLUX, a text-to-image model. Every artifact anyone could find was image work. Nobody found a single text or language-model fine-tuning run anywhere on the machine.
airanks.net's entire product is ranking how brands get cited inside ChatGPT, a text model. Put "we constantly train AI models ourselves" in the footer of that site, and a reasonable reader has every reason to assume you mean the models the site is about. The council split on whether that inference is actually material, meaning whether it would change a reasonable reader's judgment, and nobody ran the reader study that would settle it definitively. But naming the modality costs one clause, and the burden of proof sits with whoever's making the claim, not whoever's reading it. So the ruling didn't wait on the test: name the modality, drop "constantly" (which was doing a lot of work for exactly two dated log entries), and ship the honest version.
The one that was a legal landmine
"Developer license agreements with most frontier labs" turned out to be the kind of phrase that sounds like careful legal language and is actually the opposite. None of the frontier labs call their standard terms a "developer license agreement." OpenAI calls it a Services Agreement. Anthropic calls it a Commercial Terms of Service. Google calls it Gemini API Additional Terms of Service. The phrase itself is real industry usage elsewhere (E*Trade and SAP both publish documents with that exact name), which is precisely why it carries meaning it shouldn't get to borrow here: it implies a negotiated bilateral contract, not a click-through terms page anyone with a credit card can accept.
The FTC settled with accessiBe for a million dollars over false AI capability claims, and its enforcement pattern in 2026 shows active policing of exactly this kind of AI marketing overclaim. Section 5 doesn't require intent to deceive, only a material misrepresentation a reasonable consumer would rely on.
One research pass proposed a fallback: list the labs where the company actually holds accounts. A second pass, checking the live key registry directly, killed even that fallback. The account list named Mistral, and there was no Mistral key on file at all. It named Anthropic as a held account, but Anthropic is reached only through an OpenRouter proxy, meaning OpenRouter holds that relationship, not the company claiming it. The one lab document that is a genuine license (Meta's Llama Community License, a real IP grant) is also one anyone who clicks "accept" holds identically, so it confers no special status to claim.
Applying to OpenAI's Partner Network Registered tier wouldn't fix this either. It's a publicly self-applicable floor tier, a marketing badge, not a bilateral license, and one badge plus a universal license still isn't "most frontier labs." This one didn't get reworded. It got cut, unanimously, in every form.
The turning point: fifteen PRs, and then a diff
The fourth claim, about helping build frontier labs' state-of-the-art tools, is where the whole exercise earned its keep.
The research re-derived the number live from the GitHub API: 15 merged third-party pull requests across 6 AI organizations, confirmed accurate on re-check. That's a real, countable number, and every seated model built its opening position on it. Substantial contributions were named: a fix to OpenAI's Agents SDK, a fix to the Model Context Protocol's Rust SDK. A weaker one was also named, a documentation fix to Anthropic's claude-code-action, flagged in the brief itself as the softest item in the set.
Every model that argued this claim in round one argued it from the PR number and title. None of them had seen a diff. That's the asymmetry the chair exists to correct: the chair is the one member of the council with tools, and it used them.
gh pr view https://github.com/openai/openai-agents-python/pull/4606 \ --json title,state,mergedAt,additions,deletions,changedFiles,author {"additions":95,"deletions":2,"changedFiles":2, "mergedAt":"2026-08-23T22:07:32Z","state":"MERGED", "title":"fix(core): max_turns no longer clobbers a tripped input guardrail exception in streaming"} gh pr view https://github.com/anthropics/claude-code-action/pull/1579 \ --json title,state,mergedAt,additions,deletions,changedFiles {"additions":1,"deletions":1,"changedFiles":1, "mergedAt":"2026-08-07T14:56:43Z","state":"MERGED", "title":"docs: fix broken Bedrock anchor in cloud-providers.md"}Ninety-five lines added, two removed, across two files, fixing a real correctness bug where a turn-limit error could silently clobber a tripped safety guardrail in a streaming agent. Against one line added, one line removed, in a single markdown file, fixing a broken documentation anchor.
Both PRs are real. Both are merged. Both are, in the narrowest sense, evidence that a pull request landed at a frontier lab's organization. But putting them in the same sentence as joint proof of "helping build state-of-the-art tools" treats a 95-line core logic fix and a one-line markdown fix as the same unit of evidence. They aren't. The count of 15 across 6 organizations is accurate. It's also the wrong frame, because counting merged PRs as if they were interchangeable is how you end up citing a broken hyperlink fix as proof of engineering depth.
Once the diffs were on the table, the council moved fast. Five of the ten retractions logged across the two rounds trace directly to this one check: a seat had argued from a title, got shown a diff, and reversed. The ruling that survived is narrow: name the two contributions the diff actually backs up, and leave the documentation fix off the band entirely. It can live on the open-source page, labeled honestly as what it is, but it doesn't get to stand next to the real fixes as if they were equivalent.
What moved, and what didn't
Ten retractions total across five reachable seats. (A sixth, running on qwen, returned no usable content in either round and is recorded as absent, not as a dissent; starved reasoning budgets are a known failure mode of this format, and this was that failure mode, not disagreement.) Nine of the ten retractions share one shape: a model had reasoned from a summary, a title, or a count, got shown the underlying artifact directly, and changed its position on the spot. Five came from the PR diff check above. Four more came from a second check, an audit of which API accounts were actually live versus proxied versus nonexistent, which is what killed the CLAIM C fallback sentence that two different seats had independently proposed and then had to walk back once someone actually looked.
None of this changed the shape of the final ruling much. The council had already smelled trouble with all three of the risky claims in round one. What the diff check changed was precision: which specific sentence survives, and which specific evidence gets to stand behind it. That's a smaller-sounding result than "the council caught a lie," and it's the more honest one. Nobody was lying. The count was real. Counting was just a weaker way of knowing something than reading was, and it took reading to find that out.
The reusable rule that came out the other side is short: verify against the primary artifact, not a summary of it. Make sure a skeptic could check the claim from a named public source in about fifteen minutes. And make sure the net impression, once it's sitting in a footer next to three other claims, still says what the words say on their own. A council of six models spent a round arguing from PR titles before one of them, the one holding the tools, went and looked. That's not a story about the five who didn't. It's a story about how much a title alone will let you get away with believing.