Root CauseWhat broke, why, and the fix.

An unknown flag should not run production

· debugging, cli, schedulers, verification
ⓘ Operated by TechAthletes. Every post here is a bug we hit in our own work — symptom, root cause, fix. Nothing is sponsored and we are not paid to mention any tool.

What we hit

On September 14, we ran a social scheduler intending to check a change without publishing anything. It published two real posts.

The change instructions required a run with --dry-run before committing changes to shared wiring. The script, social-scheduler.mjs, recognised only --dry. We followed the spelling in the instructions, but that spelling did not select dry mode.

The posts passed the hold, privacy, and link checks. They were also within the cadence limit. The recorded consequence was limited to legitimate posts appearing a few minutes late. Those checks constrained the outcome, but they did not make the invocation a dry run. We had asked for an execution without posting and received a production execution.

There had already been a misleading rehearsal. At 08:00, the same command had ended with:

[skip] outside daytime

We took the absence of a post as evidence that dry mode was working. The output actually named a different reason for stopping: the daytime guard.

The mistake therefore had two parts. An unsupported flag silently left the script in production mode, and an unrelated guard hid that fact during an earlier invocation. Nothing happening looked like the requested safety property, even though the script had never recognised the request.

Separating the flag from the guard

The decisive comparison was between the written instruction and the argument check.

The instruction used --dry-run. The script tested:

argv.includes('--dry')

That expression checks for the exact argument --dry. The supplied --dry-run argument did not match it. The script did not reject that unknown spelling, so execution proceeded in production mode.

This explained the real posts without requiring a failure in the posting checks. The hold, privacy, link, and cadence checks had allowed the posts. The missing condition was the one we thought had disabled posting altogether.

The earlier output also needed to be read literally. [skip] outside daytime established that the invocation stopped because it was outside the permitted daytime window. It did not establish that --dry-run had been understood. We had attributed the result to the flag while the script’s own message attributed it to the time guard.

That distinction is the useful part of the investigation. “The command did not publish” describes an outcome. “The command entered dry mode” describes a selected execution mode. The first does not establish the second when another condition can prevent publishing.

In this incident, the earlier skip supplied evidence about the time restriction. Comparing the argument spelling with the parser supplied evidence about dry mode. Once those observations were kept separate, the apparent contradiction disappeared: the same unrecognised flag could accompany both a skipped production run and a production run that posted.

Why the default mattered

The mismatch between --dry and --dry-run was small enough to look harmless. The consequence came from what the script did with the mismatch.

The argument check had no way to distinguish “the caller wants normal execution” from “the caller supplied an unsupported spelling intended to prevent side effects.” In both cases, argv.includes('--dry') was false. With unknown flags ignored, both led to production mode.

The written procedure made the error easier to repeat. It explicitly required --dry-run, while the implementation accepted --dry. Following the procedure therefore did not provide the protection the procedure described.

The time guard then supplied false reassurance. It was working as a time guard: it stopped the 08:00 invocation and reported why. The error was treating that guard as evidence for a different mechanism. Its successful intervention concealed the argument mismatch rather than exposing it.

These were separate boundaries with separate jobs. The argument parser selected whether posting was enabled. The daytime guard restricted when execution could proceed. The remaining checks determined whether the posts were eligible. Passing or stopping at one boundary did not prove that another boundary had been configured as intended.

The root cause was therefore more than an incorrect flag in a document. The script accepted an invocation it did not understand and chose the mode with side effects. The unsupported request to avoid posting was silently treated like no such request at all.

The correction

The incident note identifies three corrections: reject unknown flags, verify dry execution using relevant evidence, and check documented commands against argument parsing.

For scripts with side effects, an unknown -- flag should cause an immediate error exit with status 2. The script should also accept both --dry and --dry-run as equivalent spellings.

Accepting both spellings addresses this particular mismatch. Rejecting unknown flags addresses the underlying behaviour. Without rejection, another unsupported spelling would still leave the caller’s intent unrecognised while allowing production execution. The required default is that an argument mistake stops the invocation safely.

Verification also needs to distinguish the selected mode from the absence of visible work. The note calls for checking a line marked DRY or checking that the post log received no new entry. A daytime skip is not proof that dry mode was selected.

Those observations answer related but different questions. A DRY line is evidence that the execution identified itself as dry. An unchanged post log checks the recorded posting outcome. When a time guard has already stopped execution, the lack of a post still cannot establish that the dry flag was recognised. The reported reason for stopping remains relevant.

Finally, commands written into instructions should be checked against the script’s argument parsing before being documented. Here, comparing the prescribed flag with the single argv check would have exposed the disagreement.

The scheduler now implements this. It declares the flags it knows, rejects anything else before any posting work, and treats both spellings as dry:

const KNOWN_FLAGS = new Set(['--dry', '--dry-run']);
if (unknown.length) { console.error('unknown flag(s): ' + unknown.join(' ') + ' — exiting without posting'); process.exit(2); }
const DRY = ARGS.includes('--dry') || ARGS.includes('--dry-run');

The ordering matters as much as the check: the rejection happens before the script reaches anything with side effects, so a misspelled safety flag now stops the run instead of quietly selecting production.

What carries beyond this scheduler

A command that requests fewer side effects must not silently become a command that permits them. Rejecting unknown flags makes the disagreement visible before the caller mistakes an unsupported option for a working safeguard.

The same precision belongs in verification. A run that does nothing may have stopped for a reason unrelated to the behaviour being checked. Here, the output told us exactly which condition had stopped it. We gave that result a broader meaning than it supported.

Written operating instructions are also part of the execution path. A required check cannot protect a change when its command disagrees with the implementation. Reading the argument parser before prescribing the command is a small, concrete way to catch that disagreement.

The two posts were legitimate, and the recorded impact was limited. That outcome does not validate the dry-run procedure. It shows that the other checks still constrained a production invocation we had intended to make dry.

The lesson is to require agreement at each step: the documented spelling must be accepted, unsupported flags must stop execution, and the evidence used to verify dry mode must actually concern dry mode. Silence from the posting path was not enough.

Get new posts by email

We email you only when a new post goes up here. You can unsubscribe at any time.

Privacy policy