Root CauseWhat broke, why, and the fix.

If it fixed itself, your diagnosis was wrong

· playwright, spa, debugging, testing
ⓘ Operated by TechAthletes. Every post here is a bug we hit in our own work — symptom, root cause, fix. Nothing is sponsored and we are not paid to mention any tool.

Three days, three brands, three tasks in the human queue: "the login session for this account has expired, please sign in again."

All three accounts posted successfully afterwards. Nobody had signed in. One of them recovered 24 minutes after being declared dead.

An expired authentication cookie does not resurrect itself. When a fault clears with no intervention, the name you gave that fault has been falsified. Not "probably wrong" — falsified, by the system's own behaviour.

The awkward part: we had fixed this exact class of false alarm two days earlier.

Why the previous fix did not hold

The earlier fix was sound in principle. Instead of naming a cause in the fallback branch of the error classifier, we required positive evidence: only file the task when the tool prints its explicit Not logged in. line.

That correctly stopped browser-launch timeouts and unhandled exceptions from being misreported. It did nothing about a case we had not considered — that the code printing Not logged in. could itself be wrong.

We had promoted a log line to evidence without auditing what the emitter tests before printing it.

The emitter

ensureLoggedIn() raced two locators: a marker that only exists when logged in (IN), and one that only exists on the login screen (OUT). Whichever resolved first decided the outcome.

if (winner === 'out') return { state: 'loggedOut' };   // no grace, no reload

The IN side had already been given patience — wait for the condition, retry, do not conclude from one look. The OUT side was still terminal on first sight.

That asymmetry is the bug. The site is a heavy single-page app, and it frequently paints logged-out markup before the session hydrates. The logged-out shell is static and cheap; restoring the session takes a round trip. On a loaded machine — and these jobs run at fixed times that collide with other scheduled work — OUT wins the race against a session that is perfectly valid.

We had fixed "do not decide from a single observation" on one branch and left the other branch deciding from a single observation. Half a fix looks like a whole fix in the diff.

The fix

Seeing OUT is now a suspicion, not a verdict:

  1. OUT appears → give IN a grace window (12s by default) to show up anyway.
  2. Still nothing → reload and require it again.
  3. Only a result that reproduces across every attempt is called loggedOut.

With one exception that stays immediate: a redirect to the site's dedicated login flow URL. A live session is never sent there, so that redirect is real evidence rather than a rendering artefact.

Testing a race deterministically

You cannot ask a third-party site to render its logged-out shell first. So we did not involve it at all. The tests drive ensureLoggedIn() with a fake page object whose locators resolve on a schedule we control — node --test, no network.

Written before the fix, confirmed red, then:

case before after
OUT first, IN arrives late (our false alarm) loggedOut loggedIn
genuinely expired loggedOut from one load loggedOut, reproduced after reload
logged in normally loggedIn loggedIn, same speed
neither marker ever appears unknown unknown
redirect to login flow loggedOut loggedOut

Row three matters as much as row one: the common path must not get slower to fix the rare one. It did not — the grace window only opens when OUT has appeared.

Then we ran the detector against the three real saved profiles, in detect-only mode with no posting: 3.1s, 4.7s, 2.1s, all loggedIn. The three tasks sitting in the human queue were pointing at sessions that were working while the tasks were open.

What to carry away

Self-recovery is disconfirming evidence, and it is free. It costs nothing to check what happened after an alert fired. We ran the check because the recovery times were absurd, but it should be routine: before working a task that blames a specific cause, look at whether the system has already contradicted it.

A log line is a claim. Before you make one the condition of a branch — especially a branch that generates human work — read the code that emits it and satisfy yourself about what it actually tests. "Require positive evidence" is only as good as the evidence's provenance.

Fix both sides of a symmetric decision. We taught the healthy path to be patient and left the failure path trigger-happy. If a state machine can be wrong by concluding too early, every terminal transition needs the same grace, not just the one you were staring at when you found the bug.

Weigh false alarms by what they displace, not by their own cost. Each of these tasks was ten minutes of nobody's time. The real damage was to the queue: a genuine legal-compliance task had been sitting in it for 29 days, under a growing stack of phantoms that looked equally urgent.

Get new posts by email

We email you only when a new post goes up here. You can unsubscribe at any time.

Privacy policy