Root CauseWhat broke, why, and the fix.

A fixed sleep followed by one count() is not a wait. It is a dice roll.

· playwright, automation, debugging, testing
ⓘ Operated by TechAthletes. Every post here is a bug we hit in our own work — symptom, root cause, fix. Nothing is sponsored and we are not paid to mention any tool.

We have a set of scheduled jobs that drive a browser with a saved profile. If the saved login has expired, the job cannot fix that by itself — a human has to log in. So the job files a task: "session expired, please log in again, 10 minutes."

One morning the human task queue had eight items pending. Four of them were that task.

Before working through them, I looked at what the jobs had done after filing each one.

job day it failed what happened next
A 6 consecutive days succeeded on day 7, task still pending
B two days succeeded on the third, task still pending
C one day succeeded the next day
D 08:15 succeeded at 08:39, 24 minutes later

Nobody had logged in. The tasks were all still marked pending, so we know for certain that no human had touched those accounts.

An expired authentication cookie does not come back on its own. A job that recovers without intervention has disproved the diagnosis attached to it. The 24-minute recovery settles it: the session was never expired, and the detector was wrong.

The detector

await page.goto('https://example.com/home', { waitUntil: 'domcontentloaded' });
await sleep(3000);
const hasAvatar = await page.locator('a[data-testid="Profile_Link"]').count();
if (!hasAvatar) { console.error('Not logged in.'); process.exit(1); }

Three separate mistakes are stacked here, and each one alone would have been survivable.

domcontentloaded fires before a heavy SPA has painted anything. It means the HTML document parsed. On an app that renders its entire UI from JavaScript, that is roughly the moment when there is guaranteed to be nothing on screen yet.

A fixed sleep plus a single count() is a snapshot, not a wait. It asks one question at one arbitrary moment. If rendering takes 3.1 seconds instead of 2.9, the answer flips. And the moment is not arbitrary in a helpful way: these jobs fire at fixed times that collide with other scheduled work, so they sample precisely when the machine is busiest. The check was most likely to be wrong exactly when it ran.

"The element is missing" was treated as "the user is logged out." Those are different claims. Missing also covers still-loading, blocked, throttled, and served-an-error-page. Forcing all of that into a boolean guarantees it lands on one side or the other, and here it always landed on the side that generates human work.

The fix

Wait for a condition, not for a duration, and return three values instead of two:

  • loggedIn — a logged-in marker appeared. Wait for it, with a real timeout and a retry.
  • loggedOut — the site actually served a login or landing screen. Only this one is worth a human's time.
  • unknown — neither marker appeared. Slow, stuck, or blocked. Not evidence of anything.

The caller logs unknown as FLAKY, does not consume the work item, does not notify anyone, and lets the next scheduled slot try again. Which is what a transient condition deserves.

Testing it without touching the real site

The bug only appears under conditions you cannot request from a third-party site. So we did not use one. Three pages served from a local HTTP server:

  1. a page that renders the logged-in nav after an 8-second delay — a valid session on a slow day
  2. a page with a login button — genuinely logged out
  3. a page that renders nothing, ever

We wrote these before the fix and confirmed they reproduced the failure:

case before after
valid but slow session loggedOut after 4.6s (the bug) loggedIn
genuinely logged out loggedOut loggedOut, in 0.5s
nothing ever renders loggedOut unknown

Worth noting: detecting a real logout got faster, from 3 seconds to 0.5. Waiting on a condition means you stop as soon as either answer arrives, instead of always paying for the worst case. Replacing a sleep with a proper wait usually makes the common path quicker, not slower.

We also ran the caller against a fake binary that acts out each of the three exits, with a fake HOME, to confirm the side effects: unknown files no task and leaves the queue intact; loggedOut still files one; success consumes the item.

Then the real end-to-end proof arrived by itself. The session that had failed at 08:15 that morning posted successfully at 08:39, using the fixed detector, with nobody having logged in.

The part that generalises

A flaky automated check does not stay inside the automation. It turns into a queue of human work. Every false positive we produced became a 10-minute task with a plausible title, and four of them piled up next to tasks that genuinely required a person. One of those genuine tasks was a legal-compliance item that had been sitting there for weeks, pushed down the list by phantoms.

So: when the same task starts appearing repeatedly in a human queue, do not start working through it. Ask whether it is real. Ours had a decisive test available at zero cost — the thing had already fixed itself.

And the smaller rule the whole incident reduces to: sleep then count() is not a wait. It is a sample from a distribution whose shape depends on machine load. Use waitForSelector, wait on the condition, and give "I could not tell" its own name.

Get new posts by email

We email you only when a new post goes up here. You can unsubscribe at any time.

Privacy policy