How-to, not a pitch

Where to find the file opti-pipe reads

opti-pipe reads one real file per import — a genuine artifact your own Spark, dbt, or Flink run already produces (or, for Flink, two REST responses saved together), not something you fill out by hand. This page is just "where is it," for all three.

Spark

Finding (or enabling) a Spark event log

Event logging is opt-in — most clusters already have it on for their own History Server, but if yours doesn't, turning it on is one config line.

  1. 1. Make sure event logging is on

    Check for these two lines in spark-defaults.conf, your SparkConf, or your job submit command. If they're already there (they often are, since the History Server needs them too), skip to step 2.

    spark.eventLog.enabled true spark.eventLog.dir <a path or URI your driver can write to> # e.g. file:/tmp/spark-events or s3a://your-bucket/spark-events

    This has to be set before the run you want to analyze — it can't retroactively log a run that already finished.

  2. 2. Locate the file after the run finishes

    Spark writes one file per application, named after the application ID (application_1753318365000_0142-shaped) inside spark.eventLog.dir. Where that actually lands depends on how you're running Spark:

Local / standalone cluster

Whatever local path or shared filesystem you set spark.eventLog.dir to — just ls it after the job completes.

Amazon EMR

Written to the cluster's local HDFS by default and optionally mirrored to S3 if you configured maximizeResourceAllocation/log mirroring — check your cluster's spark.eventLog.dir setting in the EMR console's Spark configuration for the exact location, or pull it via aws s3 cp if it's mirrored there.

Databricks

Open the job run's Spark UI, then the Environment tab to confirm event logging is on for the cluster, or download the run's event log directly from the cluster's Event Log tab in the workspace UI.

Kubernetes / spark-submit

Whatever path or object-store URI you passed to spark.eventLog.dir in your submit command or Helm values — same rule as local/standalone, just wherever that mount or bucket is.

One file, not a directory. If your log ends up split into rolling segments (Spark's spark.eventLog.rolling.enabled), concatenate the segments for a single application back into one file before uploading — the parser reads one JSON-lines file per import, one JSON object per line, same as it comes off disk in the non-rolling case.

  1. 3. Upload it

    On the dashboard's "Your pipeline" tab, click + Add your pipeline, pick Spark, fill in your cluster's actual spark.executor.memory / spark.executor.instances / spark.sql.shuffle.partitions (so recommendations compare against your real current settings, not defaults), and choose that file. Your data is anonymized by design: only task timing, record counts, GC time, and heap-usage percentages are read from it — never your SQL, DataFrame code, or file paths. Spark's event log has no CPU-utilization signal at all, so that one metric stays honestly blank rather than guessed.

dbt

Finding target/run_results.json

dbt writes this file automatically on every invocation — there's nothing to configure or enable, it's already sitting on disk after your last run.

  1. 1. Run dbt normally

    Any invocation that actually executes something writes the file — dbt run, dbt build, or dbt test. A plain dbt compile also writes one, but with no execution timing in it, so it's not useful here.

  2. 2. Look in your project's target/ directory

    By default that's <dbt project root>/target/run_results.json — the same directory dbt already uses for compiled SQL and manifests. If your project sets a custom target-path in dbt_project.yml, look there instead.

    # from your dbt project root, after `dbt run` ls target/run_results.json
  3. 3. Running on dbt Cloud instead of locally?

    Open the run in the dbt Cloud UI and download run_results.json from that run's "Artifacts" tab — same file, same format, dbt Cloud just also keeps a copy for you.

  4. 4. Upload it

    On the dashboard's "Your pipeline" tab, click + Add your pipeline, pick dbt, and choose that file. Your data is anonymized by design: only timing (elapsed_time, per-model execution_time) and row-count numbers are read from it — never your SQL, model names, or file paths.

That's the whole file-finding part.

Everything after this is the dashboard's own upload form — no more digging through cluster configs.