dbt

Junior

How do I configure dbt sources.yml so stale data fails loud instead of silent?

A model can run cleanly against a table that stopped updating three days ago. sources.yml is where you tell dbt to actually check for that, instead of trusting that upstream is fine.

Read in:
No source declared at all Declared, no freshness Freshness never run
How most teams' source coverage actually looks in practice, roughly in that order - not a measured breakdown.

dbt models can read from any table without declaring it as a source - but only a declared source gets lineage tracking, documentation, and freshness checks. Without those, a stalled upstream load just looks like a normal, successful run.

Declaring a source before you can check its freshness

A source is a YAML declaration pointing at a table dbt didn't build - typically raw data landed by Fivetran, Airbyte, or a custom loader. Nothing about declaring it changes how your models query it; a declared source just gives dbt something to attach documentation, tests, and freshness config to.

# models/staging/sources.yml sources: - name: raw_orders database: analytics schema: raw tables: - name: orders loaded_at_field: _fivetran_synced freshness: warn_after: {count: 6, period: hour} error_after: {count: 24, period: hour}

loaded_at_field is the column dbt checks the max value of - almost always a loader-provided sync timestamp, not a business date column (order date and "when this row landed" are different things, and freshness cares about the second one).

warn_after vs error_after, and actually running the check

warn_after and error_after are independent thresholds, not a two-stage escalation of the same number - you can set a warn threshold and no error threshold at all if you just want visibility. None of this runs as part of a normal dbt run or dbt build, though - freshness is its own command:

# checks every source with a freshness block, writes sources.json dbt source freshness

That has to be scheduled explicitly - a cron step, a separate Airflow task, whatever runs before the models that depend on that source. A freshness block that's configured but never invoked catches nothing; it's not a passive constraint dbt enforces on its own.

Where this still leaves a gap

Freshness only tells you the table got a row recently - it says nothing about whether that row's values are sane, or whether the load only wrote half the expected records. A source can pass freshness and still be wrong. And dbt source freshness exiting non-zero doesn't stop downstream models from running unless your orchestrator is explicitly wired to check its exit code and halt - dbt itself won't do that for you.

Not something opti-pipe checks: freshness and source config live entirely in sources.yml, which the rule engine never reads - it only looks at what run_results.json reports about execution time and status after a run already happened.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back - not another rule of thumb.