Flink

Junior

How do I read a Flink metrics export, and which numbers actually matter?

Most of what Flink's metrics API returns is noise for a first pass. A short list of counters answers almost every “is this job okay” question.

Read in:
busyTimeMsPerSecond numRecordsInPerSecond checkpoint duration
The metrics with the highest signal for a first-pass health check - illustrative, not measured.

Flink exposes metrics through a REST API (or JMX, or a reporter of your choice) as flat JSON: one entry per metric name per task/operator/job, no built-in prioritization of which ones matter.

The shape of a metrics export

A pull from /jobs/<id>/vertices/<vertex-id>/metrics returns an array of {"id": "...", "value": "..."} objects - flat, no nesting, no indication of which ones are worth looking at first. That's on you to know going in, since the API itself treats a GC pause counter and your actual throughput counter as equally important.

The handful that matter for a first pass

busyTimeMsPerSecond - how much of each second a task spent doing real work, the closest thing Flink has to a utilization metric. numRecordsInPerSecond / numRecordsOutPerSecond - actual throughput, and a growing gap between "in" and "out" across the pipeline is a backlog forming. Checkpoint duration (from the checkpointing REST endpoint, not the per-task metrics) - a slow or growing checkpoint duration is often the earliest warning sign of a problem, well before throughput visibly drops.

What's technically available but usually noise

JVM heap and GC metrics are real and occasionally the actual answer, but they're a second-pass tool - check them once busyTimeMsPerSecond or checkpoint duration has already told you something is wrong, not as your first stop. Starting there on every investigation means wading through dozens of JVM counters before reaching the two or three that usually explain what's happening.

This is the same reason opti-pipe's Flink rules only ever look at a short, fixed list of metrics - not because the export doesn't have more, but because more isn't more useful past a certain point for answering "is this job okay."

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back - not another rule of thumb.