Cost
Mid-levelHow do I cut my Databricks/EMR/Glue bill without breaking anything?
Not every cost lever carries the same risk. Here's the order that gets you real savings before you touch anything that could actually break a pipeline.
Cloud data infra bills rarely spike from one big mistake — they creep up from a dozen small ones that were reasonable at the time and never revisited. The good news: most of the cheap fixes are also the safe ones.
Low risk, do these first
- Right-size instance counts against actual utilization, not the peak you provisioned for once. If average CPU utilization across a job's executors sits under 40%, you're very likely over-provisioned — this is close to a free lunch since it doesn't change job logic at all.
- Turn on autoscaling if you haven't. Fixed cluster sizing means paying for peak capacity around the clock. Most managed Spark platforms (Databricks, EMR) support this natively with minimal config.
- Move eligible batch workloads to spot/preemptible instances. A retryable, checkpointed batch job (not a long stateful streaming job) is usually a safe fit, often at 60-90% off on-demand pricing.
Medium risk, test first
- Reduce shuffle partition count if it's set defensively high. Fewer, larger partitions mean less task-scheduling overhead, but going too far reintroduces the memory-pressure problems the high count was probably raised to avoid — verify against GC time and spill bytes after changing, not just the wall-clock number.
- Consolidate small, frequent jobs into fewer, larger runs. Cluster startup/warmup time is fixed overhead paid every run; a job that runs every 5 minutes pays that overhead 12x more often than the same work batched hourly.
Higher risk, needs real validation
- Switching instance families (e.g. memory-optimized to general-purpose) changes the memory-per-core ratio your job was implicitly tuned against — a cost win on paper can become a performance or OOM regression in practice.
- Aggressive TTL/lifecycle changes on intermediate storage save on storage cost but can silently break anything downstream that assumed that data would still be there.
Why the ordering matters: the "low risk" tier above is where opti-pipe's rules focus first — instance count and shuffle partitions sized against your actual measured utilization, shown as a specific recommendation with a real dollar estimate, not a percentage you have to translate into cost yourself. Every suggestion is something you approve, not something that auto-applies.
See what this looks like on your own pipeline.
Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get
concrete, approve-before-apply recommendations back — not another rule of thumb.