Cost

Mid-level

How do I cut my Databricks/EMR/Glue bill without breaking anything?

Not every cost lever carries the same risk. Here's the order that gets you real savings before you touch anything that could actually break a pipeline.

Read in:
Right-size instance counts Turn on autoscaling Move eligible jobs to spot Reduce shuffle partitions Switch instance families
Ordered by risk, not by size of the saving — see the article for why.

Cloud data infra bills rarely spike from one big mistake — they creep up from a dozen small ones that were reasonable at the time and never revisited. The good news: most of the cheap fixes are also the safe ones.

Low risk, do these first

  • Right-size instance counts against actual utilization, not the peak you provisioned for once. If average CPU utilization across a job's executors sits under 40%, you're very likely over-provisioned — this is close to a free lunch since it doesn't change job logic at all.
  • Turn on autoscaling if you haven't. Fixed cluster sizing means paying for peak capacity around the clock. Most managed Spark platforms (Databricks, EMR) support this natively with minimal config.
  • Move eligible batch workloads to spot/preemptible instances. A retryable, checkpointed batch job (not a long stateful streaming job) is usually a safe fit, often at 60-90% off on-demand pricing.

Medium risk, test first

  • Reduce shuffle partition count if it's set defensively high. Fewer, larger partitions mean less task-scheduling overhead, but going too far reintroduces the memory-pressure problems the high count was probably raised to avoid — verify against GC time and spill bytes after changing, not just the wall-clock number.
  • Consolidate small, frequent jobs into fewer, larger runs. Cluster startup/warmup time is fixed overhead paid every run; a job that runs every 5 minutes pays that overhead 12x more often than the same work batched hourly.

Higher risk, needs real validation

  • Switching instance families (e.g. memory-optimized to general-purpose) changes the memory-per-core ratio your job was implicitly tuned against — a cost win on paper can become a performance or OOM regression in practice.
  • Aggressive TTL/lifecycle changes on intermediate storage save on storage cost but can silently break anything downstream that assumed that data would still be there.

Why the ordering matters: the "low risk" tier above is where opti-pipe's rules focus first — instance count and shuffle partitions sized against your actual measured utilization, shown as a specific recommendation with a real dollar estimate, not a percentage you have to translate into cost yourself. Every suggestion is something you approve, not something that auto-applies.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back — not another rule of thumb.