Spark

Junior

How do I read a Spark event log without opening the Spark UI?

The Spark UI is just a renderer for this file. Once you know which event types actually matter, you can get the same numbers faster - or feed them straight into automation.

Read in:
SparkListenerTaskEnd SparkListenerStageCompleted SparkListenerJobEnd
Which event types carry the metrics you actually need, roughly in that order - not a measured breakdown.

A Spark event log is a newline-delimited JSON file - one event object per line, written in real time as the job runs. The UI you're used to clicking through reads the exact same file you can.

Where the file lives, and what's in it

Set spark.eventLog.enabled=true and spark.eventLog.dir to somewhere durable (local disk, S3, HDFS) and Spark writes one line of JSON per event as the application runs - no extra tooling required. Every line has an "Event" field naming its type: SparkListenerApplicationStart, SparkListenerJobStart, SparkListenerStageSubmitted, SparkListenerTaskEnd, SparkListenerStageCompleted, and a dozen others most jobs never need.

# top 10 longest tasks, straight from the event log cat app-events.log | jq -r 'select(.Event == "SparkListenerTaskEnd") | [."Task Info"."Task ID", ."Task Metrics"."Executor Run Time"] | @tsv' | sort -k2 -n -r | head -10

The three event types that carry almost everything you need

SparkListenerTaskEnd is the one that matters most: its "Task Metrics" object has executor run time, memoryBytesSpilled, diskBytesSpilled, and shuffle read/write byte counts - per task, which is the granularity that actually shows skew. SparkListenerStageCompleted gives you the stage-level summary without re-aggregating every task yourself, and SparkListenerJobEnd tells you whether the whole thing actually succeeded. Almost every diagnostic question - is this job spilling, is one task way slower than the rest, did this stage even finish - is answerable from just those three.

Where this breaks down, and why the UI still exists

Two honest limits: the event log is written after the fact, so it's no good for a job that's still running (the UI's live view reads in-memory state, not the file). And it gets big - a job with 50,000 tasks produces roughly one JSON line per task just for TaskEnd events, so json.load()-ing the whole file into memory is the wrong approach past a few hundred MB; stream it line by line instead.

This is literally what opti-pipe's rule engine does with it: streams through task-end events rather than loading the file whole, extracts the numeric metrics above, and never touches anything that would tell it about your actual SQL or DataFrame code - the event log doesn't carry that anyway, only the operations Spark's own DAG scheduler produced.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back - not another rule of thumb.