dbt
JuniorWhat's actually inside dbt's run_results.json, and why it's worth reading directly?
Every dbt run writes this file to target/, whether or not you ever open it. It already contains the answer to “which model is slow.”
No dbt Cloud account, no warehouse query history, no extra instrumentation - target/run_results.json already has per-model timing for every run you've done since you last cleaned the target directory.
The shape of the file: one result object per node
The results array has one entry per model, test, or seed that ran, each with a
unique_id, a status ("success", "error", "skipped"), an execution_time
in seconds, a thread_id, and a timing array breaking that time into compile and
execute phases with real timestamps.
Thread ID tells you about concurrency, not just identity
dbt runs models in parallel up to whatever --threads you passed. If your slowest models all
share a thread_id and ran back-to-back rather than overlapping with other work, that's a DAG-shape
problem - a dependency chain forcing serial execution - not a per-model slowness problem. Worth checking
before you spend an afternoon optimizing SQL that was never the bottleneck.
What it can't tell you
Timing, not cost. run_results.json has no bytes-scanned or credits-used figures - for
Snowflake or BigQuery compute cost you still need the warehouse's own query history, joined back to a
dbt run via query comment or query ID if your adapter logs one. That's a real gap, not something to
paper over.
This is exactly the file opti-pipe's dbt rules read - no dbt Cloud account or warehouse
credentials needed, just the file your last dbt run already wrote to disk.
See what this looks like on your own pipeline.
Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get
concrete, approve-before-apply recommendations back - not another rule of thumb.