Two numbers from your Spark UI — peak heap utilization and GC pause time — checked against the same real threshold rules opti-pipe's engine runs. Not a generic pricing calculator: the exact logic behind part of the same engine that coaches your pipeline toward its healthiest config, live in your browser.
Step 1
Find both in the Spark UI's Executors tab (or the Databricks/EMR equivalent) for a recent run — see the FAQ below if you're not sure where to look.
This checks the same two thresholds as opti-pipe's real memory rules: peak heap under 60% with no GC pressure flags overprovisioning (reduce ~20%); peak heap over 90% with GC pauses above 15% of runtime flags OOM risk (increase ~25%). The rest of the rule engine — shuffle partitions, cluster sizing, executor cores, instance family, schedule cadence — needs a real event log, which is what the full dashboard reads, and it's what lets opti-pipe track a pipeline over time and coach it toward its best config instead of a one-off check.
This only checks memory, on numbers you typed in. It can't see shuffle behavior, cluster sizing, or schedule cadence — for those, and for exact numbers instead of what you estimated from the UI, upload a real Spark event log, dbt run-results file, or Flink job-metrics export. Free, no signup.
Run the full analysis →Questions
Open the Spark UI for your application, go to the Executors tab, and look at the peak JVM heap usage against the allocated executor memory for the run. Databricks shows the same figure on a job run's Spark UI under the Executors tab; EMR and standalone clusters expose it the same way via the History Server.
The most common pattern is memory allocated well beyond what a job actually uses — peak heap utilization staying under roughly 60% of allocated executor memory, run after run, with no corresponding spike in garbage collection time. That capacity is paid for every billing cycle whether it's used or not.
Roughly 60–90% peak heap utilization with garbage collection under about 15% of runtime is a reasonable healthy band. Below that range usually means overprovisioned memory being paid for unused; above it, especially combined with rising GC time, is an early warning sign for an out-of-memory failure.
No. This checker runs entirely in your browser against two numbers you type in — nothing is uploaded or sent anywhere. For a full analysis against your real Spark event log, dbt run-results file, or Flink job-metrics export, opti-pipe's dashboard is also free with no signup required.
Upload a real Spark event log, dbt run-results file, or Flink job-metrics export and see the exact rule that fires — real numbers, nothing auto-applied. Track it over a few runs and it starts coaching you toward that pipeline's healthiest config.