Flink

Mid-level

How do I deploy a Flink job - session mode vs application mode?

Flink gives you two genuinely different ways to run a job, and they trade off in exactly one dimension: whether a bad job can take other jobs down with it.

Read in:
Application mode Session mode (dev) Session (multi-tenant)
Roughly how these modes tend to get used in practice - not a measured breakdown.

A session cluster is a long-lived JobManager that accepts multiple jobs, submitted one at a time or concurrently. An application cluster is spun up fresh for exactly one job's main() and torn down when it exits. Same Flink runtime underneath, very different failure isolation.

What each mode actually gives you

In session mode, you submit jars to an already-running JobManager, which schedules them onto whatever TaskManagers are registered - fast to iterate against (no cluster startup per submission), and a natural fit for a shared dev/staging environment where spinning up infrastructure per job is wasteful. In application mode, the JobManager itself runs your job's main() - the cluster exists only for this one job, from startup to shutdown, and nothing else can be scheduled onto it.

# same jar, two different deployment shapes flink run -m yarn-session -yid application_123 job.jar # session mode flink run-application -t yarn-application job.jar # application mode

The failure-isolation tradeoff that actually matters

A session cluster's JobManager is a shared resource: a job whose main() leaks memory, registers too many classes, or otherwise misbehaves at the JobManager level can degrade or crash every other job on that same session - not just its own. Application mode gives every job its own JobManager, so a misbehaving job's blast radius stops at that job. This is almost never about raw performance (task-level execution is the same either way) - it's entirely about whether one team's bad deploy can page another team.

The practical default: session mode for interactive development and ad hoc jobs where the overhead of a per-job cluster isn't worth it; application mode for anything running unattended in production, specifically because of the isolation, not because it's faster.

Where this still needs judgment

Application mode's per-job cluster startup adds real latency before a job actually starts processing - fine for a long-running streaming job where startup time is negligible against total runtime, more noticeable for short batch-style Flink jobs run frequently, where cluster startup can be a meaningful fraction of total job time. There's no universal answer here; it depends on how the job's runtime compares to its own startup cost, which is worth actually measuring rather than assuming.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back - not another rule of thumb.