A cron expression is five fields that decide when something runs. Most
guides stop at the syntax. The part that matters in production is what a
schedule costs when it fires, what happens when it overlaps itself, and how
you notice silence. We run a small fleet of scheduled jobs, and their
behavior taught us more about cron than any cheat sheet.
The expression, one minute
minute hour day-of-month month day-of-week
`*/15 * * * *` means every fifteen minutes. `0 3 * * 1` means 03:00 every
Monday. `*/3` in the day-of-month field does not mean "every three days"
in any way a human can rely on, because months do not divide evenly. That
single mistake explains more broken schedules than any other.
If you need "every N days", do not use day-of-month arithmetic. Track the
last run in a file or database and compare timestamps in the job itself.
The scheduler gives you time triggers, not intervals.
What we actually run, and why those intervals
Four scheduled jobs illustrate four interval decisions:
1. A runtime smoke test every 15 minutes. It checks that the game servers
accept connections. Fifteen minutes is the longest outage we are willing
to sleep through, so the interval encodes an acceptable damage bound,
not a preference.
2. A cache cleanup daily. Disk fills at a predictable rate, so once a day
with headroom to spare is enough.
3. A secrets unseal job every 15 minutes. It is idempotent and takes two
seconds, so running it often is nearly free.
4. An index status check every 3 days. Search Console data moves slowly, so
a tighter interval only produces identical reports.
Read those as a rule: choose the interval from the cost of missing an
event and the cost of running the check. Never from habit.
Overlap is the failure you did not schedule
Every job that runs every N minutes will eventually take longer than N
minutes. What happens then decides whether you have an outage or an
incident. Without a lock, two copies run at once, and you get duplicate
side effects, corrupted writes, or doubled alerts.
Concurrent cronjobs in Kubernetes get a default policy of Allow, which is
exactly wrong for anything with side effects. Set
`concurrencyPolicy: Forbid` for jobs that must not overlap, or Replace for
jobs where the newest run makes the oldest irrelevant. For plain cron on a
box, a flock lock in the script is one line:
flock -n /tmp/myjob.lock myjob.sh
The `-n` makes a contended lock exit immediately instead of queueing.
Detecting silence, not just failures
A cron job that fails loudly is healthy. A cron job that stops running
fails silently, because nothing executes to report anything. Our index
status check produced three identical reports in a row before the
uselessness of the interval became the finding itself.
Two counters fix this:
1. A watchdog that checks when a job last succeeded, independent of the
job. Our listener watchdog runs every 20 minutes and verifies the
components are alive, so a dead scheduler gets noticed by something
that is not dead.
2. A heartbeat write with a timestamp. If the newest timestamp exceeds two
intervals, alert. Five lines of script.
The timezone trap
Day-of-week and hour fields evaluate in the scheduler's timezone, not
yours. A job at `0 9 * * 1-5` meant for 09:00 European time fires at a
different wall clock on a server running UTC, and shifts again when
daylight saving changes. Kubernetes CronJobs accept a `timeZone` field.
Use it explicitly, even when the default happens to be correct today.
Field-by-field errors we have seen
- `*/N` in day-of-month expecting every-N-days. Months are 28 to 31 days.
Use timestamp comparisons in the job.
- Sunday written as 7 in a system where 0 and 7 both mean Sunday, mixed
with 1-6 for weekdays. Write 0 consistently.
- A leading space. ` * * * * *` parses as six fields in strict parsers and
fails at 2 a.m. during a deploy.
- Comma lists spanning ranges. `1,15` is fine, `1-15/2` means every second
value in the range, not every second day.
Checklist before a schedule goes live
1. Write the interval as a sentence: "at most fifteen minutes of outage."
If the sentence is wrong, the cron is wrong.
2. Set a concurrency policy or take a lock.
3. Add a success heartbeat and a watchdog that reads it.
4. Pin the timezone.
5. Load-test the job's runtime against its interval. If runtime can exceed
interval, step one lied to you.
Cron syntax takes five minutes to learn. The schedule around it, the part
that notices silence and prevents overlap, is the actual engineering.
---
What is the worst schedule failure you have debugged, and did it come from
the expression or from the silence around it? If you want to test an
expression against real dates before trusting it, our cron parser explains
each field and the next run times in your browser:
https://webrecast.com/en/cron-parser