Spark
Mid-levelHow do I deploy a Spark job - client mode vs cluster mode?
spark-submit has exactly one flag that decides where the driver runs - and getting it wrong doesn't fail loudly, it just makes your job's survival depend on something it shouldn't.
Every Spark deployment runs a driver process and some number of executors - client mode and cluster mode only disagree about where the driver lives. That one difference has bigger consequences than it sounds like.
What --deploy-mode actually moves
In client mode, the driver runs on the machine that issued spark-submit - your
laptop, a notebook server, a CI runner - and only the executors run on the cluster. In cluster mode,
the cluster manager (YARN, Kubernetes, standalone) launches the driver as its own process on a cluster node,
so the machine that ran spark-submit can disconnect entirely once the job is accepted.
Why client mode is the wrong default for anything unattended
The driver coordinates every task and holds the job's final state - if it dies, the job dies, no matter how healthy the executors are. In client mode, that means a laptop going to sleep, a notebook kernel restarting, or a CI runner's job timeout all kill a Spark job that was otherwise running fine. It's not a theoretical edge case; it's the single most common way a "why did this job just disappear with no error" ticket starts. Client mode is genuinely the right choice for interactive work - a notebook where you want driver-side output streamed back to you live - just not for anything that needs to survive after you close the laptop.
The tell: if the answer to "does this job need to keep running if I disconnect" is yes, the
answer to --deploy-mode is cluster. There's no in-between setting.
What changes operationally once you switch
Cluster mode moves where driver logs live - they're on a cluster node, not your terminal, so you read
them via yarn logs -applicationId ... or kubectl logs rather than watching stdout
scroll by. It also changes how you get the job's exit status back: client mode returns it directly to the
shell that ran spark-submit; cluster mode requires polling the cluster manager
(yarn application -status, or the Spark UI/history server) since the submitting process exits
as soon as the job is accepted, not when it finishes. Neither is harder, but a deploy script written
assuming client-mode's synchronous behavior will misreport success in cluster mode if it isn't updated to
poll instead.
See what this looks like on your own pipeline.
Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get
concrete, approve-before-apply recommendations back - not another rule of thumb.