Spark

Mid-level

How do I deploy a Spark job - client mode vs cluster mode?

spark-submit has exactly one flag that decides where the driver runs - and getting it wrong doesn't fail loudly, it just makes your job's survival depend on something it shouldn't.

Read in:
Cluster mode Client mode Doesn't matter here
Roughly how deploy mode should map to job type - not a measured breakdown.

Every Spark deployment runs a driver process and some number of executors - client mode and cluster mode only disagree about where the driver lives. That one difference has bigger consequences than it sounds like.

What --deploy-mode actually moves

In client mode, the driver runs on the machine that issued spark-submit - your laptop, a notebook server, a CI runner - and only the executors run on the cluster. In cluster mode, the cluster manager (YARN, Kubernetes, standalone) launches the driver as its own process on a cluster node, so the machine that ran spark-submit can disconnect entirely once the job is accepted.

# the flag that decides everything below spark-submit --deploy-mode cluster --master yarn --class com.example.Job app.jar

Why client mode is the wrong default for anything unattended

The driver coordinates every task and holds the job's final state - if it dies, the job dies, no matter how healthy the executors are. In client mode, that means a laptop going to sleep, a notebook kernel restarting, or a CI runner's job timeout all kill a Spark job that was otherwise running fine. It's not a theoretical edge case; it's the single most common way a "why did this job just disappear with no error" ticket starts. Client mode is genuinely the right choice for interactive work - a notebook where you want driver-side output streamed back to you live - just not for anything that needs to survive after you close the laptop.

The tell: if the answer to "does this job need to keep running if I disconnect" is yes, the answer to --deploy-mode is cluster. There's no in-between setting.

What changes operationally once you switch

Cluster mode moves where driver logs live - they're on a cluster node, not your terminal, so you read them via yarn logs -applicationId ... or kubectl logs rather than watching stdout scroll by. It also changes how you get the job's exit status back: client mode returns it directly to the shell that ran spark-submit; cluster mode requires polling the cluster manager (yarn application -status, or the Spark UI/history server) since the submitting process exits as soon as the job is accepted, not when it finishes. Neither is harder, but a deploy script written assuming client-mode's synchronous behavior will misreport success in cluster mode if it isn't updated to poll instead.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back - not another rule of thumb.