Spark
Mid-levelWhy is my Spark cluster utilization low even though the job succeeds?
“It finished without errors” and “we used the cluster we're paying for” are two different claims - only one of them shows up in a green job status.
Job success only means every task completed without throwing. It says nothing about whether half your executors sat idle while it did.
The most common cause: fewer partitions than cores
If a stage has 200 partitions and your cluster has 400 cores available, half your cluster gets zero tasks for that stage's entire duration - full price, zero work. This shows up constantly on jobs that inherited a shuffle-partition count from a much smaller cluster and never revisited it after scaling up.
The second most common cause: one stage dominates wall time
A UDF-heavy transform or a poorly-pushed-down filter can make one stage take 80% of the job's wall clock while only touching a fraction of the data's partitions - the rest of the cluster idles waiting for that one stage to clear. Dynamic allocation helps trim cost here (it releases idle executors), but it doesn't fix the underlying shape; the job is still bottlenecked on one stage regardless of how many executors are technically available.
The metric that actually separates these two
In the Spark UI's Executors tab, compare each executor's task time against the job's total elapsed time. A ratio well under 1 across most executors, for most of the job's duration, points to the partitions-vs-cores mismatch. A ratio near 1 for most of the job but a visible cliff during one specific stage points to the single-stage bottleneck instead - different fixes for what looks like the same symptom on a cost dashboard.
See what this looks like on your own pipeline.
Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get
concrete, approve-before-apply recommendations back - not another rule of thumb.