Spark

Junior

Why did my Spark job suddenly get slower with no code changes?

Nine times out of ten it isn't the code. It's the data, the cluster, or a config that quietly stopped matching either one.

Read in:
Data grew, config didn't Join no longer balanced Cluster config drifted
Roughly ranked by how often each is the actual cause — not measured across a real dataset.

You didn't touch the job. The DAG is identical. The last deploy was three weeks ago. And yet last night's run took 40 minutes instead of 12. This is one of the most common tickets in data engineering, and it almost always traces back to one of three things.

1. The data grew, but the config didn't

Shuffle partition counts, executor memory, and broadcast join thresholds are usually set once, during initial development, against whatever data volume existed at the time. Six months later the same table has 4x the rows and the same spark.sql.shuffle.partitions=200 is now producing partitions too large to fit comfortably in memory, forcing spills to disk on every shuffle stage. Nothing in the code changed; the assumption the config encoded did.

How to check: compare input row counts and shuffle read/write bytes between a recent slow run and an older fast one, in the Spark UI's Stages tab. A large jump there, with a roughly proportional slowdown, points straight at this.

2. A join that used to be balanced isn't anymore

Data skew doesn't announce itself — it just means a handful of partitions end up doing 10-100x the work of the rest, so wall-clock time is dictated by the slowest task, not the average one. A key distribution that was roughly uniform at launch can drift as usage patterns change (one customer, one region, one event type starts dominating volume).

How to check: in the Spark UI's Stage detail view, sort tasks by duration. A stage where the p50 task takes 4 seconds and the max task takes 6 minutes is skewed, full stop.

3. The cluster isn't the cluster you tested on

Autoscaling clusters, spot-instance pools, and shared multi-tenant clusters can all quietly hand your job fewer or slower executors than it got last time, especially under contention from other jobs. This shows up as slower wall-clock time with identical per-task CPU time — the job is doing the same amount of work, just waiting longer for resources to do it in.

How to check: compare the number of executors actually allocated (not requested) across runs, and look at task scheduler delay, not just task duration.

The fastest way to tell these apart: you need two things side by side — your current config, and real metrics from the run that actually happened. That's the whole premise behind opti-pipe's rule engine: it reads your event log's actual task timing, shuffle volume, and memory usage and checks it against your configured shuffle partitions / executor memory / instance count, instead of you eyeballing the Spark UI and guessing which of the three above it is.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back — not another rule of thumb.