Flink
SeniorHow do I deploy a Flink cluster on Kubernetes using the native integration?
Flink doesn't just run inside Kubernetes - with the native integration, it talks to the Kubernetes API directly to manage its own TaskManager pods. That's a meaningfully different deployment than "Flink in a container."
There are two real ways to run Flink on Kubernetes: the standalone deployment (Flink has no idea it's on K8s - you manage replicas like any other stateless deployment) and the native integration (Flink's own ResourceManager talks to the Kubernetes API to request and release TaskManager pods as jobs need them).
What "native" actually means here
With the standalone deployment, TaskManagers are just pods in a Kubernetes Deployment -
Kubernetes restarts them if they die, but Flink itself has no ability to ask for more or fewer of them based
on job demand. With the native integration, Flink's ResourceManager calls the Kubernetes API directly
(create/delete pods) to scale TaskManagers to match what a job actually needs - closer in spirit to how
Flink's YARN integration has always worked, just talking to a different scheduler.
The RBAC this actually requires
Because Flink is creating and deleting pods itself, its service account needs real permissions in the namespace it runs in - not just the read-only access a normal application pod would have:
That's a meaningfully bigger permission surface than most application workloads need, and worth flagging to whoever owns cluster security policy before it shows up as a surprise in a review - it's inherent to how the native integration works, not a misconfiguration.
Where this still leaves a choice
The Flink Kubernetes Operator (a separate, newer project) sits on top of either deployment mode and
adds Kubernetes-native lifecycle management - kubectl apply a FlinkDeployment
custom resource instead of running flink run-application by hand, with the operator handling
upgrades, savepoint-based redeploys, and reconciliation. It's not a third deployment mode so much as a
management layer - worth adopting once you have more than a couple of Flink jobs to operate, less clearly
worth the added moving part for a single job.
Outside opti-pipe's scope: the rule engine reads whatever metrics a running or completed Flink job exports - it doesn't care whether that job's TaskManagers were provisioned by the native integration, a standalone Deployment, or the Kubernetes Operator, since none of that changes what the metrics mean.
See what this looks like on your own pipeline.
Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get
concrete, approve-before-apply recommendations back - not another rule of thumb.