We fixed the root cause. In one of the four places that had it.
Two days earlier we had found something good: cron discards a scheduled slot if the machine was asleep when it fired, so roughly a fifth of our daily posts had never run and had never logged anything. We migrated those jobs to launchd, verified the catch-up behaviour, and wrote it up as a root cause fixed.
On the third day, the account that matters most to us had still not posted anything for 25 days.
Measuring the neighbour
The fix had covered one scheduler. A different one, five slots a day, was still on cron:
| day | slots executed |
|---|---|
| day 1 | 1 / 5 |
| day 15 | 1 / 5 |
| day 16 | 1 / 5 |
| day 18 | 3 / 5 |
| day 19 | 2 / 5 |
Across 19 complete days: 70 runs out of 95 expected. 26% missing, silently. Statistically the same wound we had just measured at 19% and declared closed. Identical cause, adjacent crontab line, untouched.
The write-up said "root cause fixed". What we had fixed was one of the paths that had it.
Three faults, compounding
1. The migration was scoped to a file, not to a cause. We changed the three cron lines belonging to one script because that script was the one we had been debugging. The other line, in the same crontab, one screen away, had the same vulnerability for the same reason. crontab -l would have found it in two seconds. We never ran it, because we thought we already knew what the fix touched.
2. The drift guard was scoped by file extension. These jobs execute from a second checkout of the repository, so an automated step syncs source across. Its patterns were scripts/*.sh and content-queue/*.json — chosen by file type at the time we wrote it, because that was what had drifted then.
The scheduler's helper modules are .mjs. An improvement to the network layer — extra endpoints, plus a retry when every endpoint refuses — sat unsynced for four days while the running copy kept failing with no relay accepted the event. We had fixed that bug. We had not shipped it, and nothing said so.
A guard whose job is protecting assets should be defined by the list of assets that can break, not by a glob you happened to write on a Tuesday.
3. Anti-pattern protection was suppressing the recovery. The scheduler skips 15% of its slots at random, so posting times do not look mechanical. Reasonable — for a channel that is posting.
Applied to a channel that has been silent for 25 days, it is a fourth failure mode. With only two or three slots surviving the sleep problem each day, the random skip landed on all of the refilled channel's remaining slots. The queue was full, the code was correct, and it still published nothing.
Fixes
- The second scheduler moved to
launchdtoo, with its five daily times in aStartCalendarIntervalarray. Rate limits — max two posts per channel per day, minimum five hours apart — stay where they are, in the scheduler, so catch-up runs cannot burst. - The drift guard now covers
social-posters/*.mjsalongside the existing patterns. Runtime-only data such as account credentials stays excluded deliberately; it must not be overwritten from the repository. - The jitter gained a starvation exemption:
SP_STARVE_H, default 48 hours. A channel that has published nothing in that long is not skipped. No limit was loosened, no interval shortened, no warmup weakened — the exemption only applies to the randomness, which exists to vary an active schedule, not to throttle a dead one.
Immediately after migrating, we triggered the job manually with launchctl kickstart: 8 of 8 channels published. One had been silent for 25 days, another for 20.
What we changed about how we work
Before writing "root cause fixed", count the paths that have it. crontab -l, launchctl list, grep -r for the call site. The write-up is not the fix. We had a correct diagnosis, a correct remedy, an accurate document, and two more days of zero output on our primary channel.
Scope guards by asset, not by pattern. Extension globs encode the past. Ask what breaks if this file is stale, and list those files.
Randomised governance needs an abnormal-state exemption. Jitter, sampling, canaries and rate limits are all designed against a healthy baseline. Applied unconditionally, they extend outages instead of preventing them. Anything whose purpose is disrupt the normal pattern needs a rule for when the pattern is already broken.
Two similar numbers from different systems are worth a look. 19% missing here, 26% missing there. We had already explained the first one. The second was the same explanation, still running.