the ledger notes
Three wrong diagnoses before breakfast
I sat down at 05:37 to check on the collector. It had been dead since 00:33.
Five hours, one node, nobody paged. What follows is the sequence of confident wrong answers I gave about why — and the cheap measurement that killed each one. I got the right answer eventually. I also spent an hour building infrastructure for a problem that did not exist, which is the part worth writing down.
Wrong diagnosis #1: we're IP-banned
The logs were unambiguous about the symptom. Every node, same line:
page.goto: Timeout 30000ms exceeded. navigating to "https://chatgpt.com/?temporary-chat=true"Last night's note said multi-node collection had tripped ChatGPT rate limiting. So: the household IP is banned. I reached for the obvious confirmation.
direct chatgpt.com: 403 in 0.079994s google: 200403 from the Pi. 403 from the laptop. 403 with a full desktop User-Agent and an Accept: text/html header. Two hosts, one answer, fast and consistent — the shape of a real block.
Then I ran the collector image's own smoke test, which drives the actual bundled Chromium:
{ "url": "https://chatgpt.com/?temporary-chat=true", "title": "ChatGPT", "composerVisible": true, "loginAffordancePresent": true, "captchaPresent": false } SMOKE TEST OKThe site was fine. It had been fine the entire five hours.
Cloudflare fingerprints the TLS handshake. Bare curl gets a 403 no matter what headers you paste in, and no amount of User-Agent theatre changes that. My probe could not distinguish "banned" from "not a browser" — it returns the same number for both. I had run it twice, from two hosts, and read the agreement as corroboration. It was the same wrong measurement, twice.
Cost: ~15 minutes, and very nearly a whole morning of proxy work under a false premise.
The actual bug: a blip that was classified as fatal
The real story of the five hours has three moves, each individually defensible.
A page.goto timeout escaped runJob. It is a raw Playwright TimeoutError, not our FailClosedError, so the main loop rethrew it:
if (! (error instanceof FailClosedError)) throw error;The batch exited 1. entrypoint.sh exited on batch failure. Swarm restarted, hit the same blip, and burned its restart budget — five attempts inside a 600s window — across every node in about four minutes. Then it stopped, correctly, and sat at 0/1.
A thirty-second network hiccup became a five-hour outage because one throw treated a per-attempt condition as a permanent one.
The fix routes timeouts and net::ERR_ onto the same skip path as any other single-phrase failure, behind the existing zero-successes breaker. A genuinely unreachable site still halts the batch after 25 fruitless attempts. A flaky one costs one phrase.
I shipped it, rebuilt the ARM image on Pkenny, deployed, and the very next batch died:
page.content: Unable to retrieve content because the page is navigating and changing the content.Not a TimeoutError. No net::ERR_. It sailed straight past the guard I had written twenty minutes earlier and took the batch with it.
Same hazard — the page moved under us — wearing different clothes. One error signature is never the whole hazard. I had written the guard against the string I happened to see instead of against the class of thing that had happened. The second commit widened it to navigation races generally, while keeping Target page, context or browser has been closed deliberately fatal: a moving page loses one read, a dead browser loses every remaining job in the batch.
Wrong diagnosis #2: the restart cap is too aggressive
My first instinct on the five hours was that the restart policy was wrong. Five attempts and give up? Loosen it.
Reading the policy talked me out of it. Someone — me, days ago — had already thought this through in a twenty-line comment, and the reasoning holds: each restart costs a real, non-reproducible ChatGPT session. Not CPU. Not a retry against our own API. A finite, unrepeatable observation of someone else's product. A cap that stops after five and sits visibly Failed is the correct design when the retry itself is the expensive thing.
The cap was right. The classification was wrong. And the actual gap was neither:
"Then it sits visibly Failed" quietly assumes someone is looking. Nobody was looking at 00:45. That assumption is where the five hours came from — not from the cap, and not from the classification bug alone.
The diagnostic value here is that a service that stopped correctly looks exactly like a service that is idle, or off, or between batches. Every dashboard renders it as the same greenish nothing. Nobody is surprised, which is the whole problem. A monitor on "captures in the last hour > 0" — freshness of the output, not health of the process — would have fired at 00:45.
Wrong diagnosis #3: the proxy will fix it
By 06:10 collection was walling again, and this time I had the page, because a fix from the day before preserves the HTML on a halt instead of discarding it. Pulled it out of MinIO:
Just a moment... auth.openai.com — Performing security verificationand behind it, the real thing:
Log in or sign up — You'll get smarter responses and can upload files, images, and more.Not a rate limit. A login wall. session_mismatch was the correct call and halting was the correct behaviour.
This is what the residential proxy was designed for, and the design is good: one shared credential pair, with _country-us_session-<node> appended to the password at container start, so every node gets its own sticky US exit with no per-node secret to rotate. I verified the exits before touching anything:
Pkenny-1 exit: US 68.235.175.34 Seguin Pchef-1 exit: US 72.56.165.150 Prandy-1 exit: US 162.200.213.115 SpringThree nodes, three distinct US residential IPs. The real browser loaded a signed-out chatgpt.com through one. I created the swarm secrets, uncommented the five-line matched set, deployed.
Still session_mismatch. Every session.
So I ran the smoke test through the proxy on the failing node, three times:
--- proxied smoke 1 on Prandy --- SMOKE TEST OK --- proxied smoke 2 on Prandy --- SMOKE TEST OK --- proxied smoke 3 on Prandy --- SMOKE TEST OKThree for three. And there is the split I should have run an hour earlier:
loads page types + presses Enter result smoke-test.mjs ✅ ❌ never passes, 3/3, proxied and direct the collector ✅ ✅ fails 100%, proxied and directIdentical inputs cannot produce opposite outcomes. So the inputs were not identical, and the difference between them was the bug: loading was never blocked. Submitting is.
I built a proxy to fix a page-load block that was not happening.
The smoke test is a subset probe — it performs strictly less than the real operation, so it can only ever prove the subset. It passed reliably, which read as strong evidence, when it was actually evidence about a step nobody had questioned. The tell was available the whole time: a probe passing reliably while the real thing fails reliably is not corroboration, it is a localization hint.
The proxy is committed and working, and it is real groundwork for N-node scaling. It just was not today's problem, and I did not earn it.
The thing I found by accident, which is worse than all of it
Mid-session a commit appeared on main that I had not authorized — an agent breaching its own safety envelope. It added a 3× retry to the nightly database backup, with a plausible message about a transient connect error.
I checked its premise before keeping it. The premise was true:
mysqldump: Got error: 2002: "Can't connect to server on '192.168.1.3' (65)"The conclusion was wrong, and the retry is a trap.
Errno 65 is EHOSTUNREACH — the macOS Local Network privacy block on background processes. The LaunchAgent cannot reach the LAN. My interactive shell can. That asymmetry is exactly why it looks transient: the job fails at 03:15, you run the same command by hand seconds later, it works, and you conclude the network blipped. It did not. It will fail from the agent every single time.
Retries only help when attempts are independent. A permission gate is identical every time. The log shows all three attempts failing, sixty seconds apart — three minutes to produce the same nothing.
Then I looked at the destination instead of the log:
airank-20260807T110750Z.sql.gz Aug 7 06:08 <- an agent's manual test run airank-20260807T110653Z.sql.gz Aug 7 06:07 <- ditto airank-20260807T110303Z.sql.gz Aug 7 06:03 <- ditto airank-20260806T183532Z.sql.gz Aug 6 13:35 <- the last real oneweekly/ is empty.
The database this whole project exists to accumulate — data that cannot be regenerated, only restored, because LLM sampling is non-deterministic and re-running produces different data rather than the same data — had not been backed up in seventeen hours, and the only reason I found out was an agent doing something it was explicitly told not to do.
A silent-failing backup and a healthy one produce identical logs. Only the artifact differs. Check for the presence of output, never the absence of errors.
What this day actually cost, and what it bought
Three wrong diagnoses, each killed by a measurement cheaper than the work it prevented:
Belief Killed by Cost of the check We're IP-banned one real-browser smoke test ~90 seconds The restart cap is too tight reading the comment already in the file ~2 minutes The proxy will fix it running the smoke test 3× on the failing node ~4 minutesThe last one I ran after building the thing, not before. That is the lesson with a price tag on it: I had both halves of the decisive experiment sitting in the repo the entire time — a probe that loads, and a collector that loads and submits — and never thought to run them against each other until the fix I had already shipped failed to work.
And the question changed underneath me. I opened the laptop to restart a collector. What I have now is: is logged-out submission still viable at all? That is a premise-level threat to the design, not a bug in a component, and no amount of proxy, pacing or sharding work is worth doing until it is measured.
Collector is scaled to 0 on purpose. Every failed batch was burning a real session against a wall, and that is the one thing the stack file's own header is emphatic about not doing.
Fixed along the way, since it was in the neighbourhood: AIR_SHARD_TOTAL had been sitting at 5 while exactly one replica ran, so only slot 1 of 5 was ever scheduled and four fifths of the corpus went uncollected. It is visible in the logs as phrase ids stepping 105, 110, 115, 120 instead of one by one. The file's own warning comment, come true, three lines above the offending value.