Root CauseWhat broke, why, and the fix.

Our duplicate guard compared every post against a log that had thrown the answer away

· reliability, debugging, automation, verification
ⓘ Operated by TechAthletes. Every post here is a bug we hit in our own work — symptom, root cause, fix. Nothing is sponsored and we are not paid to mention any tool.

We post once a day to a few small accounts from a queue file. A separate job tops that queue back up from our article index, and it has one job it must not get wrong: never queue an article that has already been posted. Publishing the same thing twice to the same account is the kind of thing that gets an account flagged.

The guard for that is three lines. Read the publishing log, pull every article path out of it, and skip any article whose path is already there.

const postedLog = await fs.readFile(LOGF(), 'utf8').catch(() => '');
for (const m of postedLog.matchAll(/\/a\/([a-z0-9-]+)\//g)) used.add(m[1]);
...
.filter((a) => !used.has(a.slug))

It ran every day. It never threw. It never queued a duplicate that the log said was a duplicate. And it had been letting already-published articles back into the queue for weeks.

What the log actually contained

The writer on the other side of that read is one line in the runner:

log(f"POSTED  {item['text'][:70].replace(chr(10), ' ')} {attempts_note}")

Seventy characters of the post body. Our post bodies open with a sentence about the paper, then the title, then the link. The URL is never inside the first seventy characters. It is structurally impossible for it to be there.

So we measured what the guard could actually see:

POSTED lines article paths recoverable
Log A 32 1
Log B 15 0
Total 47 1

One. And that one only because a single post happened to be short enough that the link landed inside the window.

The guard was not broken in the sense of throwing an error or comparing incorrectly. It compared correctly. It searched a corpus, found nothing, and concluded "not posted yet" — which is exactly what it should conclude when a corpus contains nothing. The corpus was the problem. The answer had been truncated away before the guard ever looked.

What it cost

We checked the queue against the account's live timeline, which is the only record that cannot be truncated by us:

Items sitting in the two queues 27
Items already live on the timeline 7

Seven posts scheduled to go out a second time. They had not gone out twice yet only because the queue is consumed at one item per day and the cycle had not come back around.

It compounded with a second defect we found the same night. When a run failed, the runner rotated the queue head to the back of the queue so the next day would try different content. The comment above that function said it also avoided double-posting on partial success. It did the opposite: rotation keeps the text in the queue, so a run that actually succeeded while being recorded as a failure would put its own content back in line. We had two of those. One of them is on the timeline twice — once with the first four characters missing, because the typing raced the editor, and once complete, seven hours later.

The fix, and why the first half is not the interesting half

The first half is trivial: log the URL.

m = re.search(r'https?://[^\s]+', item['text'])
article_url = m.group(0) if m else ''
...
log(f"POSTED  {item['text'][:70]} {attempts_note} {article_url}")

The second half is the part we would have skipped if we had not asked what the guard had already missed. Fixing the writer only protects future posts. Every article published before today was still invisible to the guard, which meant the same job would keep re-queueing months of back catalogue.

We could not reconstruct those from our own records, because our own records are what lost the data. So we read them off the live timeline instead — the same place we had just used to find the seven duplicates. One detail was worth the trip: the rendered post breaks the URL across several lines, so collapsing whitespace produces …/a/code-as-w orlds-agentic… and matches nothing. Stripping whitespace entirely reassembles the path.

That recovered 29 published article paths, which we wrote back into the log so the existing guard could see them. We then ran the refill job and watched it add seven articles, none of which were among the 29.

Three more of these, in our own code

Once we had the shape, we recognised it.

A login check that compared against markup arriving before the session did. Our poster decided whether a session was still valid by waiting for a signed-out marker or a signed-in marker and taking whichever appeared first. On a loaded machine the signed-out markup renders before the session hydrates, so the check matched it and reported an expired session. The evidence that this was a false match and not a real expiry: the same profile posted successfully twenty-four minutes later with nobody having logged in. An expired credential does not heal itself. We now give the signed-in marker a grace period after the signed-out one appears, and only call it expired when every attempt agrees.

A bug report built on comparing against a display string. A link in one of our posts looked broken, so we requested the URL and got a 404. The URL we requested was the one shown in the post — and the platform truncates long URLs for display, with an ellipsis. Requesting the actual href returned 200. We had confirmed a defect that did not exist and were about to send someone looking for the code that truncated our links. There is no such code. (There was a real defect nearby, of a different kind, which we only found after abandoning the first one.)

A shipping gate that byte-compared the wrong tree. We added a check that the files about to ship match the files we inspected. It passed. It was comparing against a second clone of the repository, not the tree the build ran from. The check agreed, and what shipped was the version with the old assets in it.

Each of those is the same sentence. Agreement is not correctness. A comparison tells you that two things are the same; it tells you nothing about whether the thing you compared against is the thing you meant. When the reference is wrong, a match does not catch the error — it certifies it, in a log line that reads like success.

What we would tell ourselves

When a check compares against a stored record, the check has two failure modes and we habitually only test one. We test that it catches a match. We rarely test that the record can still contain the thing we are matching on — that the field survives whatever wrote it, at the length it actually has, in the format the reader expects.

The cheapest version of that test is to ask the guard what it can see, and count. Not "does the guard work", but "how many of the things it is supposed to know about are actually in its corpus". Forty-seven posts, one recoverable path, is a number that took one command to produce and would have ended this weeks earlier.

We wrote a new guard the same night, for the truncation problem: read the editor's contents back and compare them against the text we intended to send, and refuse to send when the head does not match. We tested it against a fake editor object, it passed all five cases, and we were satisfied.

Which is the same mistake, one level up. So before the next scheduled run we opened the real editor, typed a real queued post into it, and compared — 238 characters in one language, 374 in another, both matching. Only then was it a check.

Root Cause is written by TechAthletes. We run the agents behind pipelines like this side by side in Agent Tile, our macOS app for keeping many Claude Code sessions in view at once.

Get new posts by email

We email you only when a new post goes up here. You can unsubscribe at any time.

Privacy policy