Flink
SeniorFlink checkpointing: why it silently fails, and how to tune the interval and timeout
A checkpoint that quietly times out under load doesn't crash the job - it just leaves you with a much larger recovery window than you think you have.
Checkpoint interval and checkpoint timeout sound like one setting. They control two separate risks, and mixing them up is how a job ends up with a false sense of safety.
Interval vs timeout: two different jobs
checkpointing.interval sets how often a checkpoint is attempted - shorter means less data
to replay on recovery, at the cost of more frequent overhead. checkpointing.timeout sets how
long a single checkpoint attempt is allowed to run before Flink gives up on it - too short and a
perfectly healthy checkpoint under momentary load gets killed for no real reason; too long and a
struggling job keeps attempting checkpoints that were never going to finish, burning resources on
retries instead of actual processing.
The silent failure: checkpoints expiring under backpressure
Under backpressure, a checkpoint barrier has to travel through every operator's input buffer before the checkpoint can complete - if those buffers are full (which is what backpressure means), the barrier queues up behind them and the checkpoint takes far longer than usual. Nothing crashes; the job keeps processing records. But checkpoints keep expiring, quietly, until someone notices the job's actual recovery point is minutes or hours stale instead of seconds.
A reasonable starting point, and the metric that tells you it's wrong
A common starting pair: interval in the tens-of-seconds to low-minutes range, timeout set to several times your typical checkpoint duration - not your best-case duration. Then watch checkpoint duration itself as a trend, not a point-in-time number.
See what this looks like on your own pipeline.
Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get
concrete, approve-before-apply recommendations back - not another rule of thumb.