AIRANKS — The Authoritative Rankings for AI Web Content

AIRANKS measures AI visibility: we ask AI models real product and service questions, capture the complete answers as immutable observations, and publish what they contain — which brands were mentioned, which domains were cited, and which exact pages were linked. Every domain gets an AIR score from 1–10 (a decile of visibility in the active dataset; 0 means insufficient data), with the methodology in the open.

AIR

BLOG

Skip to main content

the ledger notes

Three liars in one night

Build LogAugust 13, 2026 by Jeremy Schoemaker

2026-08-13, overnight batch. Eleven jobs ran unattended through two workflows — six from the overnight survey, five from a todo-cycle — and all eleven landed. That part is not the story. The story is that three times during the night, the evidence said something was broken or missing, and three times the evidence was lying. Each lie was killed by a measurement that cost less than a minute. Believing any one of them would have cost hours.

Lie #1: "Eight tests fail on clean main"

Mid-run, the acceptance reviewer did exactly the right thing: it didn't take the implementer's word that 8 test failures were pre-existing, it checked out detached clean main and ran them there. All 8 failed. Identical files, identical line numbers, identical messages. That is the gold-standard regression check, and it said main was broken.

Main was not broken. Those tests — recovery, reparse, self-test, end-to-end — hit the real MinIO bucket and a shared database. Eleven agents in eleven worktrees were running suites at the same time, consuming and polluting each other's fixtures. The reviewer's "clean baseline" run was sharing the bucket with the fleet, so the baseline was dirty in exactly the way the branch runs were.

The tell, in hindsight: the assertion values mutated. 0 is identical to 1 failed one run; the same assertion failed the next run with 2 is identical to 1. A real regression fails the same way every time. Zero means a sibling ate your fixture; two means a sibling left residue.

The kill: wait ten minutes for the fleet to quiet down, re-run solo. 16/17, then 17/17. We had a mutex serializing "full suite" runs — but targeted --filter runs skipped the lock, and those were plenty to poison everything. A lock that doesn't cover every run touching the shared resource is a lock that works until the night you need it.

Lie #2: "Three merges failed"

Landing the branches, a zsh loop merged four of them, each logging to the same file:

redis-hit-rate-r2 merge=0 (eval):2: file exists: /tmp/mx.out observability-docs-r2 merge=1 gpg-backup-proposal-r2 merge=1 readme-test-count merge=1

Three "failed merges." Except zsh had noclobber set, the second redirect refused to overwrite the first iteration's file, and the merge commands never ran at all. The merge=1 was the shell failing to open a file, wearing a git conflict's clothes. The suite after the loop even passed — green on a tree that was silently missing three branches.

The kill: count the artifacts, not the statuses. git log --oneline showed one new merge commit where there should have been four. Re-ran with >|, all three merged clean, zero actual conflicts. The exit code you're gating on can belong to the shell's plumbing rather than your command, and the only reason this was visible at all was an echo "merge=$?" after every step.

Lie #3: "The toolbar badge needs building"

The survey listed it with evidence: the store listing describes an opt-in score badge, the session log says the spec was drafted at 13:14, and the "toolbar five" feature commit didn't include it. Solid chain of reasoning. The job's first instruction, though, was verify the premise before implementing — and the agent found the badge fully implemented in commits from earlier that afternoon, wired to the options page, off by default, exactly as specced. It returned a verified no-op instead of a duplicate implementation, and the reviewer accepted the no-op as the correct deliverable.

A survey is a snapshot of beliefs. On a day with eight deploys before dinner, beliefs about what exists go stale in hours.

The shape of the night

Three different layers — test infrastructure, shell plumbing, work inventory — produced false negatives within four hours of each other. None were exotic. All three were defeated by the same move: re-run the smallest thing that would fail if the claim were true, and count what actually happened. Solo test run. Merge-commit count. Grep the source for the feature.

The eleven jobs were the work. The three lies were the night.

(Also filed: a bonus fourth liar the next morning — the docs pass described the new FAQ page as "25 Q&A" because the QA probe it grew from had 25 questions. The page has 11. Even the documentation of the night tried to lie about the night.)

← Back to blog