Root CauseWhat broke, why, and the fix.

Our scheduled posts were not failing. They were never running.

· macos, cron, launchd, automation, debugging
ⓘ Operated by TechAthletes. Every post here is a bug we hit in our own work — symptom, root cause, fix. Nothing is sponsored and we are not paid to mention any tool.

We run a few automations on a fixed daily schedule from a MacBook. Each one appends a line to its own log: POSTED, or FAILED with the reason. Failures raise a desktop notification and file a task for a human.

For weeks the logs looked fine. Failures were rare, and every failure had been dealt with.

Then someone asked a question nobody had asked before: not "how many failed", but "how many ran".

The measurement

Three jobs, one per day each, over 18 days:

job scheduled posted failed days with no line at all
A 08:05 14 1 3
B 12:35 10 1 7
C 21:05 15 2 0

Roughly 19% of all expected runs produced no evidence of having happened. Not an error, not a stack trace, not a partial write. Nothing.

The distribution is the clue. Job C, at 21:05, never missed. Job B, at 12:35, missed 39% of its days. The midday slot is exactly when this particular laptop is most likely to be closed.

The cause

cron on macOS does not catch up. If the machine is asleep or powered off at the moment a job is due, that occurrence is simply discarded. There is no queue, no backlog, no "run it when we wake". The slot is gone, and because nothing ran, nothing logged.

This is documented behaviour, not a bug. It is also completely invisible to monitoring that is built around failures, because a job that never started cannot report a failure. You can reproduce it in a minute: confirm sleep is enabled with pmset -g custom, schedule a cron job for two minutes out, close the lid, and open it later. The log stays empty.

Our alerting was built on the assumption that absence of failure means success. On a laptop, absence can also mean absence.

The fix

launchd has the behaviour we assumed cron had. From man launchd.plist(5): with StartCalendarInterval, if the machine is asleep at the fire time, the job runs once shortly after it wakes. Missed occurrences are coalesced into a single catch-up run.

<key>StartCalendarInterval</key>
<dict>
  <key>Hour</key><integer>12</integer>
  <key>Minute</key><integer>35</integer>
</dict>

Migrating is the easy part. The three guards around it are what actually matter, and we needed all of them.

1. Remove the cron line. Actually remove it. If both schedulers are live at the same minute, both will decide independently that today's work has not been done yet, and both will do it. For a job that posts to a social platform, a duplicate post is worse than a missed one.

2. Make the runner idempotent for the day. launchd can legitimately start a job more than once in a day — that is the whole point of the catch-up behaviour. The runner has to decide for itself whether today's work is already done. Reading its own log is enough:

today = datetime.date.today().strftime('%Y-%m-%d')
if any(l.startswith(today) and ' POSTED ' in l for l in open(logfile)):
    sys.exit(0)   # already done today

Note that this checks for POSTED, not for "any line". A day that only contains FAILED should still proceed, otherwise a single early failure locks the job out until midnight and blocks manual retries.

3. Make the installer idempotent, and make the old installer delegate. Ours had an install-cron.sh that people re-ran occasionally. Left alone, it would have quietly reinstated the cron line and walked us into problem 1 weeks later. It now installs the launchd job instead of the cron line.

Verifying without side effects

Testing a scheduler is awkward when the job it runs is externally visible. You want to prove the whole path works — launchd → environment → PATH → node → permissions — without publishing anything.

The same-day guard makes this free. Pick a job that has already run today and kick it manually:

launchctl kickstart -p "gui/$UID/com.example.job"
launchctl print "gui/$UID/com.example.job" | grep -E 'runs =|last exit'

We want runs = 1, last exit code = 0, a SKIP already posted today line in the log, and no change to the work queue. That exercises every layer except the side effect.

One more check is worth doing right after launchctl bootstrap: confirm runs = 0. If you load a job whose scheduled time has already passed today, you want to know whether it fires immediately, before you find out with a real post.

What we took from it

Monitoring failures is not monitoring. A failure handler only sees runs that started. Any fault that prevents starting — a sleeping machine, an unloaded job, a deleted crontab line, a full disk at boot — is invisible to it, and those faults are not rare.

The cheap fix is to compare successes against expected runs on a schedule, rather than waiting to be told about errors. We could do that retroactively here only because every run wrote one dated line, which made 18 days of history countable with grep. That log format was an accident. It is now a rule.

And when a periodic job runs on hardware that sleeps, check what your scheduler does with a missed occurrence before you trust it. cron discards. launchd catches up. systemd timers have Persistent=true. They are not interchangeable, and the difference only shows up as a hole in a log that nobody is counting.

Get new posts by email

We email you only when a new post goes up here. You can unsubscribe at any time.

Privacy policy