Flink
JuniorHow do I read a Flink metrics export, and which numbers actually matter?
Most of what Flink's metrics API returns is noise for a first pass. A short list of counters answers almost every “is this job okay” question.
Flink exposes metrics through a REST API (or JMX, or a reporter of your choice) as flat JSON: one entry per metric name per task/operator/job, no built-in prioritization of which ones matter.
The shape of a metrics export
A pull from /jobs/<id>/vertices/<vertex-id>/metrics returns an array of
{"id": "...", "value": "..."} objects - flat, no nesting, no indication of which ones are
worth looking at first. That's on you to know going in, since the API itself treats a GC pause counter
and your actual throughput counter as equally important.
The handful that matter for a first pass
busyTimeMsPerSecond - how much of each second a task spent doing real work, the closest
thing Flink has to a utilization metric. numRecordsInPerSecond /
numRecordsOutPerSecond - actual throughput, and a growing gap between "in" and "out" across
the pipeline is a backlog forming. Checkpoint duration (from the checkpointing REST endpoint, not the
per-task metrics) - a slow or growing checkpoint duration is often the earliest warning sign of a
problem, well before throughput visibly drops.
What's technically available but usually noise
JVM heap and GC metrics are real and occasionally the actual answer, but they're a second-pass tool -
check them once busyTimeMsPerSecond or checkpoint duration has already told you something
is wrong, not as your first stop. Starting there on every investigation means wading through dozens of
JVM counters before reaching the two or three that usually explain what's happening.
This is the same reason opti-pipe's Flink rules only ever look at a short, fixed list of metrics - not because the export doesn't have more, but because more isn't more useful past a certain point for answering "is this job okay."
See what this looks like on your own pipeline.
Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get
concrete, approve-before-apply recommendations back - not another rule of thumb.