dbt

Junior

How do I deploy dbt to actually run on a schedule in production?

dbt itself has no scheduler - dbt run just runs once and exits. Getting it to run automatically, reliably, on a schedule is entirely a decision you make outside of dbt.

Read in:
dbt Cloud Airflow / Dagster Plain cron + CI
Roughly how much scheduling/orchestration infrastructure each option leaves you to build yourself, least to most - not a measured breakdown.

Running `dbt build` from your laptop proves the project works. Getting it to run every morning at 6am, retry on transient warehouse errors, and alert someone when it doesn't - that's deployment, and dbt is deliberately unopinionated about how you do it.

The three real options, and what each one owns

dbt Cloud owns the scheduler, run history, and retry logic entirely - you configure a schedule in its UI and it handles the rest, at the cost of running inside dbt Labs' infrastructure rather than yours. A self-hosted orchestrator (Airflow, Dagster, Prefect) treats a dbt run as one task in a larger DAG - the right choice when dbt is one step among several (extract, load, dbt transform, downstream export) that need to be sequenced and depend on each other, not just triggered on a clock. Plain cron plus a CI runner (a scheduled GitHub Actions workflow, a cron job on a VM) is the simplest option and completely adequate for a project with no cross-tool dependencies - just don't reach for it once you actually need retries, alerting, or DAG-level dependencies, which it doesn't give you for free.

# the whole "scheduler" here - dbt itself does nothing extra 0 6 * * * cd /opt/dbt_project && dbt build >> /var/log/dbt.log 2>&1

What you're responsible for no matter which you pick

Two things dbt genuinely doesn't handle for you in any of these setups. First, credentials - profiles.yml (or dbt Cloud's connection config) still needs real warehouse credentials available wherever the run actually executes, following the same env_var()-not-hardcoded pattern regardless of scheduler. Second, failure notification - a failed dbt build exits non-zero, but nothing pages anyone unless the orchestrator around it is wired to check that exit code and alert. dbt Cloud has this built in; Airflow/Dagster need it configured as part of the DAG; plain cron needs it bolted on explicitly (a wrapper script that checks $? and calls out to Slack/PagerDuty, or similar).

The mistake this section exists to prevent: a cron job that silently stops running (a redeployed VM without the crontab, a CI workflow file accidentally disabled) produces no error at all - no failed run, just no run. Freshness checks on the tables it should have updated (see the sources.yml freshness article) are what actually catch that, not the scheduler itself.

Where opti-pipe fits, and where it doesn't

None of these scheduling choices change what opti-pipe reads - the rule engine looks at run_results.json, which every option here produces the same way, since it's dbt's own output regardless of what triggered the run. What it can't tell you is whether the run happened at all, or on schedule - that's entirely the deployment layer's job, and worth verifying independently (a dashboard, a freshness check, an actual alert) rather than assuming quiet means fine.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back - not another rule of thumb.