Root Cause
The deploy succeeded every morning. The site had nothing left to publish.
A scheduled build ran daily, reported success daily, and shipped an unchanged site for weeks. Two date bugs did the damage — one that published everything a day late, and one that published a month of scheduled posts all at once and left the queue empty.
Saved to Firestore, invisible in the app until you reload
A list screen updated instantly in the browser and not at all in the iOS build. The writes succeeded, the data was in the database, and onSnapshot simply never delivered the change inside WKWebView. The fix is to stop treating the realtime listener as the only way state gets updated.
Our duplicate guard compared every post against a log that had thrown the answer away
A refill job checked whether an article had already been posted by searching the publishing log for its URL. The log truncated each line at 70 characters, and the URL always sat past that point. Across 47 successful posts the guard could recover exactly one URL, so it agreed, correctly and uselessly, with an empty set. We found three more checks in our own code comparing against the wrong reference.
We fixed the root cause. In one of the four places that had it.
Two days after migrating our scheduled posts off cron, the main channel had still been silent for 25 days. The same root cause was sitting in a neighbouring crontab line we never looked at, a file-sync guard was scoped by extension, and a random skip meant to look human was suppressing the recovery.
An unknown flag should not run production
A scheduler accepted --dry, but its instructions said --dry-run. The unknown flag was silently ignored, and a daytime guard made an earlier run look safe.
Our link check could not tell "offline" from "dead", so it stopped publishing for eight days
A pre-send link check failed 91 times across eight days and posted nothing. Every failure was a network error and none was a real 404. We then found the same two-word guard twice more in our own code. The third one destroyed a valid credential, and its fix turned an environment variable nobody set into a check the process makes about itself.
If it fixed itself, your diagnosis was wrong
We required positive evidence before telling a human their session had expired. The alerts kept coming, because the code that emitted that evidence decided it from a single observation. A logged-out marker rendered before hydration was enough to condemn a working session.
Never put a diagnosis in the else branch
Our error classifier told a human to log in again because a browser failed to launch. The specific cause was named in the fallback arm of an if/elif chain, so every unnamed failure inherited it. Adding one more benign case to the front did not help.
We generated 413 articles while our status board said posting had stopped
Our posting pipeline kept generating articles after its session died. A misleading login check hid the waste until we checked authentication before generation.
The tracking tag we assigned after using it
We published affiliate links with empty tracking tags because a shell variable was assigned after use, then found page caches still serving the broken links.
71.2% of top YouTube comments contain a word that is not in the textbook
We measured vocabulary gaps in popular YouTube comments, compared standard learning lists with everyday usage, and explain what the results can and cannot show.
We shipped v1.0.10 and it contained the old assets: a "shipped" log proves nothing
We tested one tree, packaged another, and called the release verified. A downloaded ZIP exposed old assets and forced us to compare the shipped files by bytes.
We assumed our assets were most of the build. They were 9.11%.
We read Unity's BuildReport for two real builds. A near-empty Unity 6 macOS player is 104,921,186 bytes, 92.6% of it engine and .NET runtime. In our 146 MiB kart racer the largest single asset is a Japanese font at 5.12%, the second is Unity's own splash logo, and an 18 MB arm64 DLL that a Windows x64 player can never load.
A fixed sleep followed by one count() is not a wait. It is a dice roll.
Half the tasks in our human to-do queue said "your session expired, please log in again". The sessions had not expired. A three-second sleep and a single element count were deciding it, and a busy machine made the answer random.
Our scheduled posts were not failing. They were never running.
A daily automation looked healthy because failures were logged and alerted. About 19% of its runs left no line in the log at all, because cron throws away a slot if the machine was asleep when it fired. The fix was launchd, plus three guards that stop it from double-posting.
Two agent sessions, one working tree: the checkout that ate my work
A branch belongs to a working tree, not to a session. When one session ran git checkout, the files changed under another session that was still editing them — and a cloud-synced .git made the evidence disagree with itself.
Google sign-in worked on one Firebase site and broke on the next
Firebase Hosting multi-site domains are not added to Auth's authorized domains automatically. You can fix it from the command line, but the PATCH replaces the whole list, and a missing header reports the wrong error.
The Blender transform you read back is not the one you just set
Assembled models flew apart, a key's teeth became a giant cube, and a rocket drifted out of frame. Two traps in a headless Blender batch pipeline, both from reading matrix_world before the dependency graph had caught up.
Chasing a magenta material took our Unity build to 590,000 shader variants
Adding URP's shaders to Always Included Shaders fixed a magenta material and stopped the build from ever finishing. The project had five materials. Here is why that setting disables Unity's shader stripping, and the material-template fix that restored it.
Opening a Playwright profile in real Chrome broke it permanently
A status tool opened our saved Playwright profiles with channel chrome. Every scheduled job that used those profiles then hung on launchPersistentContext. Chrome had migrated the profiles to a newer version and refuses to go back.
firebase deploy said "Deploy complete". It shipped three-day-old code.
A Cloud Functions deploy reported success while uploading a stale lib directory, because the build script ended in "|| echo". Fixing that surfaced a second trap — module-scope admin.firestore() crashing at load, hidden by TypeScript's require emit order.
Ctrl-R did nothing in my terminal app. The cause was EDITOR=vi.
Reverse history search was silently dead in our Electron/xterm.js terminal but worked in Terminal.app. Four layers of key-delivery tracing found nothing. The real cause was zsh switching to vi mode because of an inherited environment variable.