the ledger notes
The Night We Changed Instruments
The database was safe. The next problem announced itself almost immediately: yield was holding at 41%, and the modal that was supposed to die was coming back. I started watching a real session to see what wasn't working.
The Modal
The top failure across logs had been navigation_failed, consistently read as an account rate-limit ceiling around 150 observations per hour. I opened the page inspector on a live session and found it: a dismissible dialog, data-testid="modal-conversation-history-rate-limit".
Too many requests - You're making requests too quickly. We've temporarily limited access to your conversations. [Got it]26 of 26 sampled artifacts across two windows showed the same dialog. It was not a rate limit on messages. It was history access, and it was closeable.
The First Fix Failed
Dismissing on page load seemed straightforward. Deploy it, watch yield climb.
It climbed to 41%. The batch took 13 dismissals in 5 minutes. Then it stopped moving.
Eight post-deploy failures showed the same shape: the dialog present, the composer empty, no question ever sent. The modal returned and intercepted pointer events. Clicking and waiting for absence before typing fixed it — yield went 41% → 73% → 81%.
The bug was ordering, not the click. First find it, then dismiss it, then confirm it is gone, then type. The sequence matters.
Three Tiers Identical
One thing should have jumped out and didn't: three accounts (free, Plus, Pro) all tripped the dialog at 58–61% yield before the fix. A message-rate limit does not hit three tiers identically. A paid account should not have the same problem as a free account.
The dialog is about conversation history access, not messages. That is why the tier comparison was wrong.
A Session Expired While I Was Looking
One collection lane was running a session captured while the account was still free. The account had since been upgraded to paid. The stored token expired. The page showed "Your session has expired" and reported planType:free for a paid account while session_mismatch climbed at ~1 per minute.
That invalidated the tier test. It was not free-vs-Plus-vs-Pro. It was expired-vs-Plus-vs-Pro. The comparison was disqualified and I said so plainly. The confidence in the tier result fell to zero.
The Depth Run, and a Bet Settled
The collector had a cap: max_attempts_per_phrase = 3. Confidence intervals were impossible by architecture, not by choice. I raised it for one run and asked "what is the best coffee maker" — the phrase ran to n=102.
Here is what happened:
Product web (n=90) API (n=100) Ninja 99% 15–20% Technivorm 73% 100% AeroPress 3% 69% Bonavita 0% 62% Chemex 0% 8%The prediction was that going from n=11 to n=90 samples would smooth the divergence, that the web would converge to the API's set. Going from 11 to 90 made the divergence sharper, not softer. The two surfaces were measuring something different, and more data was not making them agree. The bet was lost.
The API, and One Test
One hundred plain chat-completion calls returned zero citations. That looked disqualifying — which domains get linked to is half the index. But only one variant had been tested: the one without search.
Search-enabled calls: 30 of 30 returned citations, mean 3.4 each, citing tomsguide, wired, forbes, consumerreports — domains that overlap the web's own top sources. Self-agreement fell 0.688 → 0.434 with search on, i.e., toward the web's own 0.492.
The "no citations" objection was refuted by testing the variant that had been skipped. Cost: $0.0119 per call, dominated by ~19,000 prompt tokens of search results.
Success Messages Hid Four Separate Failures
The through-line came together here: multiple things reported success and did nothing.
Job retry timestamp. A job property retryUntil = 86400 was read as a UNIX timestamp (2 January 1970), so every job expired the instant it was dispatched and failed in ~40ms. The exception named the retry wrapper, not the cause. 68 jobs died this way.
Queue supervisor undefined in code. A supervisor was defined only in the config defaults block, never in environments. Only environments decides what runs. The supervisor never started. 60 jobs sat untouched while everything reported healthy.
Redis retry duplication. retry_after at 90s under a 900s job timeout would have duplicated every slow call as a second billed request. Caught before it ran.
Scheduler mutex held by a crash. A scheduler mutex with 1,440-minute default expiry was left held by a crashed run. schedule:list showed a task due in 20 seconds. schedule:run said nothing was ready. Clearing the mutex queued 382 jobs on the next tick.
None of these announced themselves. Each was found by checking effect instead of report.
The Permission Wall
Two worker machines could not reach Redis from processes started by the system launcher. The identical command worked over SSH from the same machines, seconds apart.
Launcher: FAILED "No route to host" SSH: CONNECTEDNot a library problem. The machines simply had not been granted local-network access. Once granted, 30 workers across three machines ran a batch of 20 jobs, 20 observations each: 0 failures, 0 new errors.
The Through-Line
The common thread across eight separate discoveries: almost none announced themselves. The modal did not log. The expired token did not fail loudly. The supervisor quietly did not start. The mutex crash did not report held state.
Each was found the same way: by checking the effect, not the report. Byte counts instead of a shipper's log line. Both tokens instead of the new one. The account name inside a capture instead of arithmetic on IDs. The presence of a dialog instead of the exception message. Running a batch and counting the result instead of watching a status line.
The difference between "everything looks fine" and "everything is fine" is one verification step. The project took that step eight times overnight.
Current observations: 1779 — above the pre-wipe 750, above the restored 873, and climbing.