Spark

Junior

The small files problem: why your Spark job spends more time on I/O than compute

A job can be I/O-bound and look compute-bound on every dashboard that only tracks CPU, because the bottleneck is opening files, not reading bytes.

Read in:
File-open overhead Task-scheduling overhead Actual read time
Where time goes on a small-files-affected stage, roughly - illustrative, not measured.

Below a certain file size, the fixed overhead of opening a file - metadata lookup, task scheduling, codec setup - costs more than reading its actual contents.

How you end up with thousands of tiny files

The usual sources: partitioning a write by a high-cardinality column (each partition value gets its own small file), streaming jobs committing small micro-batches on a short trigger interval, or a shuffle stage with far more output partitions than the data actually needs. None of these are mistakes exactly - they're often the right call for the job writing the data - but they leave a mess for whatever reads it next.

Why it costs more than the byte count suggests

Each file, however small, still needs a full task: schedule it, open the file (a real network round trip on S3 or GCS), read a footer or header, then close it. At 10,000 tiny files, that per-file overhead can dominate total stage time even though the actual data barely fills a few executors' memory - the job looks compute-bound on a CPU graph because the CPU is busy doing scheduling and I/O bookkeeping, not because it's doing real work.

# rough average file size for a given output path avg_file_size = total_bytes_written / file_count # well under ~128MB (Parquet) is worth investigating

The fix that actually sticks

coalesce() or repartition() right before the write, targeting roughly 128MB-1GB per output file depending on format and downstream engine. A one-time fix on the write side is cheaper than every downstream job paying the file-open tax repeatedly - compacting after the fact works too, but it's a recurring cost that a correctly-sized write avoids entirely.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back - not another rule of thumb.