Infra

Senior

What's the right cluster size and autoscaling policy for this workload?

The right cluster size isn't a number you pick once — it's a policy you set from how your workload actually behaves, then revisit as that behavior changes.

Read in:
Steady-state baseline Normal variance Occasional peak
A workload's utilization curve — the shape that should set your floor and ceiling.

"How big should my cluster be" is the wrong question. The right one is "what's the utilization curve this workload actually produces, and what policy fits that curve" — a batch job with a sharp hourly spike needs a completely different answer than a steady-state streaming job.

Set the floor from steady-state load, not peak

Your autoscaling minimum should cover baseline load with headroom for normal variance — not the peak you occasionally see. Sizing the floor to peak load means paying for capacity you use a small fraction of the time; that's exactly the "fixed cluster sizing" cost problem autoscaling exists to solve, just recreated with extra steps.

Set the ceiling from a real cost conversation, not "however high it needs to go"

An unbounded autoscaling maximum protects availability but removes any cost ceiling — a runaway query, a retry storm, or an unexpected traffic spike can scale a cluster (and its bill) far past what anyone intended. Set the max from an explicit "what's the worst case we're willing to pay for" number, not from technical capability alone.

Tune scale-up sensitivity to your job's shape

A batch workload with a sharp, predictable spike (an hourly ETL run, say) benefits from aggressive scale-up — better to over-provision briefly than have the job queue behind slow scale-up. A steady-state streaming job benefits from the opposite: slower, more conservative scale-up avoids thrashing (scaling up, then immediately back down) on normal minute-to-minute variance, which itself has a cost in provisioning/deprovisioning overhead.

The number that should actually drive this decision: average CPU/memory utilization across executors during real runs, not during a synthetic load test. opti-pipe's instance-count rule compares your currently configured executor count against utilization actually observed in your event log, and re-fires with an adjusted number if a later upload shows utilization has drifted since the last time you approved a change — it doesn't assume your workload's shape is static any more than it should.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back — not another rule of thumb.