Cron Expressions in Production: Schedules We Run and What They Cost

A cron expression is five fields that decide when something runs. Most

guides stop at the syntax. The part that matters in production is what a

schedule costs when it fires, what happens when it overlaps itself, and how

you notice silence. We run a small fleet of scheduled jobs, and their

behavior taught us more about cron than any cheat sheet.

The expression, one minute

minute hour day-of-month month day-of-week

`*/15 * * * *` means every fifteen minutes. `0 3 * * 1` means 03:00 every

Monday. `*/3` in the day-of-month field does not mean "every three days"

in any way a human can rely on, because months do not divide evenly. That

single mistake explains more broken schedules than any other.

If you need "every N days", do not use day-of-month arithmetic. Track the

last run in a file or database and compare timestamps in the job itself.

The scheduler gives you time triggers, not intervals.

What we actually run, and why those intervals

Four scheduled jobs illustrate four interval decisions:

1. A runtime smoke test every 15 minutes. It checks that the game servers

accept connections. Fifteen minutes is the longest outage we are willing

to sleep through, so the interval encodes an acceptable damage bound,

not a preference.

2. A cache cleanup daily. Disk fills at a predictable rate, so once a day

with headroom to spare is enough.

3. A secrets unseal job every 15 minutes. It is idempotent and takes two

seconds, so running it often is nearly free.

4. An index status check every 3 days. Search Console data moves slowly, so

a tighter interval only produces identical reports.

Read those as a rule: choose the interval from the cost of missing an

event and the cost of running the check. Never from habit.

Overlap is the failure you did not schedule

Every job that runs every N minutes will eventually take longer than N

minutes. What happens then decides whether you have an outage or an

incident. Without a lock, two copies run at once, and you get duplicate

side effects, corrupted writes, or doubled alerts.

Concurrent cronjobs in Kubernetes get a default policy of Allow, which is

exactly wrong for anything with side effects. Set

`concurrencyPolicy: Forbid` for jobs that must not overlap, or Replace for

jobs where the newest run makes the oldest irrelevant. For plain cron on a

box, a flock lock in the script is one line:

flock -n /tmp/myjob.lock myjob.sh

The `-n` makes a contended lock exit immediately instead of queueing.

Detecting silence, not just failures

A cron job that fails loudly is healthy. A cron job that stops running

fails silently, because nothing executes to report anything. Our index

status check produced three identical reports in a row before the

uselessness of the interval became the finding itself.

Two counters fix this:

1. A watchdog that checks when a job last succeeded, independent of the

job. Our listener watchdog runs every 20 minutes and verifies the

components are alive, so a dead scheduler gets noticed by something

that is not dead.

2. A heartbeat write with a timestamp. If the newest timestamp exceeds two

intervals, alert. Five lines of script.

The timezone trap

Day-of-week and hour fields evaluate in the scheduler's timezone, not

yours. A job at `0 9 * * 1-5` meant for 09:00 European time fires at a

different wall clock on a server running UTC, and shifts again when

daylight saving changes. Kubernetes CronJobs accept a `timeZone` field.

Use it explicitly, even when the default happens to be correct today.

Field-by-field errors we have seen

  • `*/N` in day-of-month expecting every-N-days. Months are 28 to 31 days.

Use timestamp comparisons in the job.

  • Sunday written as 7 in a system where 0 and 7 both mean Sunday, mixed

with 1-6 for weekdays. Write 0 consistently.

  • A leading space. ` * * * * *` parses as six fields in strict parsers and

fails at 2 a.m. during a deploy.

  • Comma lists spanning ranges. `1,15` is fine, `1-15/2` means every second

value in the range, not every second day.

Checklist before a schedule goes live

1. Write the interval as a sentence: "at most fifteen minutes of outage."

If the sentence is wrong, the cron is wrong.

2. Set a concurrency policy or take a lock.

3. Add a success heartbeat and a watchdog that reads it.

4. Pin the timezone.

5. Load-test the job's runtime against its interval. If runtime can exceed

interval, step one lied to you.

Cron syntax takes five minutes to learn. The schedule around it, the part

that notices silence and prevents overlap, is the actual engineering.

---

What is the worst schedule failure you have debugged, and did it come from

the expression or from the silence around it? If you want to test an

expression against real dates before trusting it, our cron parser explains

each field and the next run times in your browser:

https://webrecast.com/en/cron-parser