Never put a diagnosis in the else branch
The day before, we had fixed a false alarm in a scheduled job by adding a new benign outcome to its error classifier. A specific condition that used to be reported as a failure was now recognised, logged quietly, and retried.
The next evening, the same job filed the same wrong alert again.
What actually happened
The job's log line:
ERROR TimeoutError: browserType.launchPersistentContext: Timeout 30000ms exceeded.
The browser did not start. The saved profile was never opened. Nothing about the login state was observed, because nothing got far enough to observe it.
The alert it produced said: "your session expired, please log in again."
We measured the session directly later that evening. It was valid. A human following that instruction would have spent ten minutes logging into an account that was already logged in, and the job would have failed identically the next day.
The shape of the bug
if succeeded:
...
elif inconclusive: # <- added yesterday
...
else:
file_task("session expired — please log in again")
else means everything I have not named yet. It is, by construction, the set of outcomes the author did not think about. Attaching a specific, confident, actionable diagnosis to that set means every unanticipated failure — present and future — arrives at the human wearing a label that was picked for a completely different fault.
This is why the previous day's fix did not help. It added one more elif in front. The residual arm still catches everything else, and "everything else" is not a fixed set that shrinks as you enumerate cases. It is whatever breaks next.
Reading three days of raw logs, we found three distinct faults sitting in that else: browser launch timeouts, the outer subprocess timing out, and uncaught exceptions. All three were being reported as an expired session.
The fix
A branch that asserts a specific cause must test for positive evidence of that cause.
The tool that actually inspects the session prints an unambiguous line when it determines the user is logged out. So the classifier now requires that line:
if 'Not logged in.' in output:
file_task("session expired — please log in again")
else:
log('ERROR', tail=output[-400:]) # named nothing, consumed nothing, retried next slot
Everything else is recorded as ERROR with the tail of the output, does not consume the pending work item, and is retried on the next scheduled run. Transient faults — and launch timeouts usually are transient — resolve themselves without anyone being interrupted.
The second failure mode, which is the opposite one
Swallowing everything unnamed is also a bug, and we had already been bitten by it: a scheduler that dropped roughly a fifth of its runs without logging anything at all, which nobody noticed for weeks precisely because there was no alert.
So silence cannot be the answer either. The classifier has a third stage:
if
ERRORrepeats for three consecutive days, file a task — titled "investigate repeated failures", with the log tail attached.
The trigger for involving a human is "this has not resolved itself", not "I have decided the cause is X". The task tells the person what is observed, not what to do about it, because at that point we genuinely do not know.
Verifying it without touching the network
We replaced the underlying tool with a fake that can act out five endings — success, genuine logout, inconclusive, launch timeout, unhandled exception — and ran the classifier against a fake HOME, then inspected the resulting task queue and work queue.
| case | tasks filed, before | after |
|---|---|---|
| success | 0, item consumed | 0, item consumed |
| genuine logout | 1 | 1 (kept) |
| inconclusive | 0 | 0 |
| launch timeout | 1 (wrong) | 0 |
| unhandled exception | 1 (wrong) | 0 |
| ERROR three days running | — | 1, titled "investigate" |
The row that matters most is the second one. It is easy to remove false alarms by making the classifier quieter; the test exists to prove the real case still gets through.
Rules we now follow
else means "unknown". It does not mean "cause X". If a branch names a cause, its condition must contain evidence of that cause.
When a classifier misfires, do not fix it by adding one more case to the front. Enumerate what is currently landing in the residual arm. Three days of logs was enough here, and it showed three different faults where we expected one.
Log lines are claims, not evidence — until you read the code that emits them. We promoted Not logged in. to a branch condition on the strength of the wording. That is only sound because we then went and checked what the emitting code tests before printing it. (It turned out to have its own bug, but that is a separate post.)
A wrong alert costs more than the time it wastes. It costs credibility. Once a queue contains alerts people have learned to ignore, the genuinely urgent item in the same queue gets ignored with them.