Our link check could not tell "offline" from "dead", so it stopped publishing for eight days
We publish to a handful of small social accounts from a scheduler that runs five times a day. Before each post goes out it runs a pre-send check: pull every URL in the text and request it, so we never publish a link that 404s. We added that check after shipping a post whose link was broken, and it has caught real mistakes since.
Then the machine that runs the scheduler lost its network connection for about eight days. When we came back and read the log, the check had done its job perfectly and the result was that nothing had been published at all.
What the log actually said
Between 2026-09-15 and 2026-09-22:
| Scheduled runs | 5 per day for 8 days |
| Posts published | 0 |
LINK_PREFLIGHT_FAILED |
91 |
of those, fetch failed |
42 |
of those, This operation was aborted |
49 |
| of those, a real HTTP error (4xx/5xx) | 0 |
Spread across every brand we post from: 24, 13, 13, 12, 12, 9, 8. Not one account escaped, because the cause had nothing to do with any account.
Every single failure was the local machine being unable to reach the internet. Zero were a URL that had actually gone missing. The check was not wrong about anything it observed. It simply had no way to say "I could not perform this check" and so it said the only thing it knew how to say, which was "this link is dead".
Downstream, "dead link" meant "do not publish, and do not consume the queue item". That is the correct handling for a dead link. For an unreachable network it is a silent eight-day outage.
The part that stings
We had already fixed this. Nine days earlier, a URL that returns 200 in 0.04 seconds was reported as This operation was aborted, and one scheduled slot was lost. We diagnosed it correctly as a transient network hiccup and added a retry: one extra attempt, 1.5 seconds later. The code even carried a comment explaining that HTTP 4xx and 5xx are real deaths and should not be retried, which is right.
That fix was correct in kind and wrong in scale. It was calibrated against the failure we had in front of us, a blip lasting under a second. The failure that actually arrived lasted eight days. One retry 1.5 seconds later is, against an outage, indistinguishable from no retry at all.
There is a more general shape here. When you fix a transient failure, the tempting move is to make the retry just long enough to cover the instance you observed. But the instance you observed is a sample of one from a distribution you have not looked at, and outages are not distributed like blips.
What we changed
Two things, and the second matters more than the first.
Retry three times with exponential backoff. Attempts at 0s, 1s and 3s, each with a 10 second timeout. This covers a network that is briefly wobbling rather than absent.
Stop calling an unreachable network a dead link. After the last attempt fails with a network exception, the result is no longer DEAD. It is a distinct verdict:
NETWORK <url> unreachable after 3 tries (fetch failed) - retry later, not dead
An HTTP response that is not 2xx still returns DEAD on the first try, with no retry, because a 404 is an answer and answers do not need repeating.
Both paths still refuse to publish. We did not loosen the guard, and we would not: publishing a link you could not verify is worse than publishing nothing. What changed is that the log now distinguishes a wall from a corpse, which is the difference between "your network is down" and "someone deleted that page", and those two require completely different humans to do completely different things.
We verified all three branches by hand before trusting it:
| Input | Verdict | Time |
|---|---|---|
| A live page | passes | fast |
| A URL that really 404s | DEAD status=404 |
one attempt |
| A hostname that does not resolve | NETWORK |
3.04s (1s + 2s of backoff) |
Then we ran one real pass of the scheduler and confirmed nine live posts, each by opening its public URL.
The same shape, twice more, in our own code
We went looking for other guards with a two-word vocabulary on the day we wrote this. We found two, and one of them was worse than ours.
The posting job for a Japanese writing platform. It ran a login check before generating an article, and treated any non-zero exit as "not logged in". While the machine was asleep on 15 and 16 September it reported five runs as needing a human to log in again. The session had never expired. That code now separates the exit codes into network_down, undetermined and not_logged_in, and only the last one asks a human for anything.
The YouTube uploader, found today. This one does not merely stop. Its token refresh caught every failure as one category, "stored token unusable", and responded by renaming the credential file to .revoked and falling back to a browser consent flow. On 14 September the machine was offline, the refresh raised a TransportError, and a valid, unexpired token was moved aside by our own code. A human then had to sit through a browser re-consent that nothing had actually required.
The fallback made it worse. Under cron there is no terminal, so the browser consent flow waited on a stdin that would never produce anything. It hung, and the next scheduled run hung behind it, four deep, for eight days with zero uploads. There was a guard against exactly this, an environment variable that suppresses the browser path in non-interactive runs. We grepped for it: read in one place, set in none. A guard nobody sets is decoration.
That is the part worth taking from the third case. A guard with two words stops working and says it is protecting you. A guard with two words and a destructive fallback will reach for the fallback, and a still-valid credential is the thing it destroys.
That one is now fixed, in two parts. Moving the credential aside is limited to a RefreshError whose message contains invalid_grant, which is the only answer that actually means the grant is gone. A TransportError, or any exception we have not classified, now exits with NETWORK_DOWN and does not touch the file at all.
The second part is the one the grep argued for. Whether it is safe to open a browser is no longer read from an environment variable; it is decided by sys.stdin.isatty() and sys.stdout.isatty(), with the variable kept only as an override. The old design asked every caller to remember to opt into safety, and the grep is the measurement of how well that worked: one read, zero writes. Asking the process what it can see is something the process can always answer, and there is nobody to forget.
We did not take that as done until three cases were run against it: a TransportError injected during refresh leaves the token byte-identical and exits non-zero, an invalid_grant still parks the file as .revoked, and a run with stdin closed exits immediately with NEEDS_REAUTH instead of hanging.
What we would tell ourselves
A guard that fails closed needs a vocabulary of at least three words: passed, failed, and could not check. With two words, "could not check" gets filed under "failed", and the system stops doing its job while reporting that it is protecting you. The logs are full, the exit codes are clean, and the output is zero.
The check we should add next is the one that would have told us on day one: an alert when a scheduled job produces zero output for a full day. We had eight days of evidence and nobody looking at it, which is its own failure and not the link checker's fault.