Flink
Mid-levelHow do I deploy a Flink job - session mode vs application mode?
Flink gives you two genuinely different ways to run a job, and they trade off in exactly one dimension: whether a bad job can take other jobs down with it.
A session cluster is a long-lived JobManager that accepts multiple jobs, submitted one at a time or concurrently. An application cluster is spun up fresh for exactly one job's main() and torn down when it exits. Same Flink runtime underneath, very different failure isolation.
What each mode actually gives you
In session mode, you submit jars to an already-running JobManager, which schedules them onto
whatever TaskManagers are registered - fast to iterate against (no cluster startup per submission), and a
natural fit for a shared dev/staging environment where spinning up infrastructure per job is wasteful. In
application mode, the JobManager itself runs your job's main() - the cluster exists only
for this one job, from startup to shutdown, and nothing else can be scheduled onto it.
The failure-isolation tradeoff that actually matters
A session cluster's JobManager is a shared resource: a job whose main() leaks memory,
registers too many classes, or otherwise misbehaves at the JobManager level can degrade or crash every other
job on that same session - not just its own. Application mode gives every job its own JobManager, so a
misbehaving job's blast radius stops at that job. This is almost never about raw performance (task-level
execution is the same either way) - it's entirely about whether one team's bad deploy can page another
team.
The practical default: session mode for interactive development and ad hoc jobs where the overhead of a per-job cluster isn't worth it; application mode for anything running unattended in production, specifically because of the isolation, not because it's faster.
Where this still needs judgment
Application mode's per-job cluster startup adds real latency before a job actually starts processing - fine for a long-running streaming job where startup time is negligible against total runtime, more noticeable for short batch-style Flink jobs run frequently, where cluster startup can be a meaningful fraction of total job time. There's no universal answer here; it depends on how the job's runtime compares to its own startup cost, which is worth actually measuring rather than assuming.
See what this looks like on your own pipeline.
Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get
concrete, approve-before-apply recommendations back - not another rule of thumb.