Ops

Junior

Why does my pipeline work fine in dev but fail or slow down in prod?

Dev doesn't lie to you exactly — it just answers a different question than the one that matters in production: "does this work," not "does this work at 50x the data."

Read in:
Data volume & shape Cluster contention Config drift
Roughly ranked by how often each explains a dev-to-prod gap — not measured.

This is less a single bug and more a category of bug, and almost every instance of it traces back to one of a few specific ways dev and prod diverge.

Data volume, obviously — but also data shape

The most obvious gap is volume: a dev sample of 10,000 rows doesn't exercise shuffle spill, skew, or memory pressure the way 500 million rows does. Less obvious: data shape also differs — a dev sample is often taken uniformly or recently, which can accidentally avoid the exact skewed key distribution or null-heavy column that production data has. Volume differences show up as slowness; shape differences show up as jobs that fail outright on data patterns dev never contained.

Cluster contention that dev environments don't have

A dev cluster is usually dedicated to one person running one job. Production clusters are frequently shared, multi-tenant, or subject to noisy-neighbor effects from other jobs competing for the same resources. A job that gets full, uncontended executor allocation in dev can see meaningfully degraded performance in prod purely from resource contention that has nothing to do with the job's own logic.

Config drift between environments

Executor memory, shuffle partitions, and instance counts are frequently hand-tuned once for a dev-scale test, then never revisited for prod-scale deployment — or the reverse, prod config gets copied into dev where it's needlessly oversized. Either way, the config that "worked" in one environment was never actually validated against the other environment's real characteristics.

How to actually close this gap before deploying

The most reliable approach: extrapolate, don't guess. Take real dev-run metrics — task duration, shuffle volume, memory usage per row processed — and scale them linearly (or with your own known non-linear factors, like shuffle cost scaling worse than linearly with row count) to estimated production volume, before the first production run. This turns "we'll find out in prod" into a testable prediction you can validate against what actually happens.

Worth being direct about: this specific "extrapolate dev metrics to prod scale" workflow isn't something opti-pipe does today — it analyzes the metrics from a run you've already had, real dev or real prod, rather than projecting forward from a smaller one. It's a genuinely useful feature we don't have yet, not one we're claiming.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back — not another rule of thumb.