opti-pipe reads one real file per import — a genuine artifact your own Spark, dbt, or Flink run already produces (or, for Flink, two REST responses saved together), not something you fill out by hand. This page is just "where is it," for all three.
Spark
Event logging is opt-in — most clusters already have it on for their own History Server, but if yours doesn't, turning it on is one config line.
Check for these two lines in spark-defaults.conf, your SparkConf, or
your job submit command. If they're already there (they often are, since the History Server
needs them too), skip to step 2.
This has to be set before the run you want to analyze — it can't retroactively log a run that already finished.
Spark writes one file per application, named after the application ID
(application_1753318365000_0142-shaped) inside spark.eventLog.dir.
Where that actually lands depends on how you're running Spark:
Whatever local path or shared filesystem you set spark.eventLog.dir to — just
ls it after the job completes.
Written to the cluster's local HDFS by default and optionally mirrored to S3 if you configured
maximizeResourceAllocation/log mirroring — check your cluster's
spark.eventLog.dir setting in the EMR console's Spark configuration for the exact
location, or pull it via aws s3 cp if it's mirrored there.
Open the job run's Spark UI, then the Environment tab to confirm event logging is on for the cluster, or download the run's event log directly from the cluster's Event Log tab in the workspace UI.
Whatever path or object-store URI you passed to spark.eventLog.dir in your submit
command or Helm values — same rule as local/standalone, just wherever that mount or bucket is.
One file, not a directory. If your log ends up split into rolling segments (Spark's
spark.eventLog.rolling.enabled), concatenate the segments for a single application
back into one file before uploading — the parser reads one JSON-lines file per import, one JSON
object per line, same as it comes off disk in the non-rolling case.
On the dashboard's "Your pipeline" tab, click + Add your pipeline, pick Spark,
fill in your cluster's actual spark.executor.memory /
spark.executor.instances / spark.sql.shuffle.partitions (so
recommendations compare against your real current settings, not defaults), and choose that
file. Your data is anonymized by design: only task timing, record counts, GC time, and
heap-usage percentages are read from it — never your SQL, DataFrame code, or file paths.
Spark's event log has no CPU-utilization signal at all, so that one metric stays honestly
blank rather than guessed.
dbt
target/run_results.jsondbt writes this file automatically on every invocation — there's nothing to configure or enable, it's already sitting on disk after your last run.
Any invocation that actually executes something writes the file — dbt run,
dbt build, or dbt test. A plain dbt compile also writes
one, but with no execution timing in it, so it's not useful here.
target/ directoryBy default that's <dbt project root>/target/run_results.json — the same
directory dbt already uses for compiled SQL and manifests. If your project sets a custom
target-path in dbt_project.yml, look there instead.
Open the run in the dbt Cloud UI and download run_results.json from that run's
"Artifacts" tab — same file, same format, dbt Cloud just also keeps a copy for you.
On the dashboard's "Your pipeline" tab, click + Add your pipeline, pick dbt, and
choose that file. Your data is anonymized by design: only timing
(elapsed_time, per-model execution_time) and row-count numbers are
read from it — never your SQL, model names, or file paths.
Flink
Unlike Spark's event log or dbt's run_results.json, Flink doesn't write one file
that has everything opti-pipe reads. Its real, stable interface for this is the
monitoring REST API
the JobManager already exposes — this is two real responses from that API, saved together as
one file, not a native Flink export.
The job ID is in the Flink Web UI's URL when you're looking at a running job
(/#/job/<job-id>/overview), or from flink list on the command line.
The JobManager's REST address is wherever its web UI is reachable — often
http://localhost:8081 locally, or your cluster's own JobManager endpoint.
Both endpoints are plain GET requests against the JobManager — no auth beyond
whatever your cluster already requires. jq combines them into the one shape
opti-pipe's importer expects.
On the dashboard's "Your pipeline" tab, click + Add your pipeline, pick Flink, and choose that file. Your data is anonymized by design: only run duration, record counts, checkpoint duration, and a backpressure ratio are read from it — never job names, SQL, or file paths. TaskManager-level JVM heap/GC numbers live under a third REST endpoint this doesn't read yet, and no rules use the checkpoint/backpressure numbers yet either — the integration shipped ahead of Flink-specific rules, on purpose.
Everything after this is the dashboard's own upload form — no more digging through cluster configs.