Flink

Senior

How do I deploy a Flink cluster on Kubernetes using the native integration?

Flink doesn't just run inside Kubernetes - with the native integration, it talks to the Kubernetes API directly to manage its own TaskManager pods. That's a meaningfully different deployment than "Flink in a container."

Read in:
Native integration Standalone Deployment Kubernetes Operator
Roughly how these three approaches compare in how much Kubernetes-awareness Flink itself has - not a measured breakdown.

There are two real ways to run Flink on Kubernetes: the standalone deployment (Flink has no idea it's on K8s - you manage replicas like any other stateless deployment) and the native integration (Flink's own ResourceManager talks to the Kubernetes API to request and release TaskManager pods as jobs need them).

What "native" actually means here

With the standalone deployment, TaskManagers are just pods in a Kubernetes Deployment - Kubernetes restarts them if they die, but Flink itself has no ability to ask for more or fewer of them based on job demand. With the native integration, Flink's ResourceManager calls the Kubernetes API directly (create/delete pods) to scale TaskManagers to match what a job actually needs - closer in spirit to how Flink's YARN integration has always worked, just talking to a different scheduler.

# starts a cluster that manages its own TaskManager pods ./bin/flink run-application \ --target kubernetes-application \ -Dkubernetes.cluster-id=my-flink-app \ -Dkubernetes.container.image=my-flink:1.20 \ local:///opt/flink/usrlib/job.jar

The RBAC this actually requires

Because Flink is creating and deleting pods itself, its service account needs real permissions in the namespace it runs in - not just the read-only access a normal application pod would have:

# minimum viable RBAC for the native integration, one namespace rules: - apiGroups: [""] resources: ["pods", "services", "configmaps"] verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]

That's a meaningfully bigger permission surface than most application workloads need, and worth flagging to whoever owns cluster security policy before it shows up as a surprise in a review - it's inherent to how the native integration works, not a misconfiguration.

Where this still leaves a choice

The Flink Kubernetes Operator (a separate, newer project) sits on top of either deployment mode and adds Kubernetes-native lifecycle management - kubectl apply a FlinkDeployment custom resource instead of running flink run-application by hand, with the operator handling upgrades, savepoint-based redeploys, and reconciliation. It's not a third deployment mode so much as a management layer - worth adopting once you have more than a couple of Flink jobs to operate, less clearly worth the added moving part for a single job.

Outside opti-pipe's scope: the rule engine reads whatever metrics a running or completed Flink job exports - it doesn't care whether that job's TaskManagers were provisioned by the native integration, a standalone Deployment, or the Kubernetes Operator, since none of that changes what the metrics mean.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back - not another rule of thumb.