dbt

Mid-level

dbt incremental models: when “incremental” quietly becomes a full refresh

The whole point of an incremental model is processing only new or changed rows. A few common triggers quietly turn that into a full rebuild, every single run, with no error raised.

Read in:
Schema / on_schema_change Manual --full-refresh flag Target table dropped
Common triggers for an unintended full rebuild, roughly by how often each one is the cause - illustrative, not measured.

dbt decides whether to run a model incrementally or as a full refresh based on a few conditions - and more than one of them can flip silently after the model's already been working fine for months.

How dbt decides incremental vs full refresh

Beyond an explicit --full-refresh flag, dbt falls back to a full rebuild if the target table doesn't exist yet, or - depending on your on_schema_change config - if the model's column set has changed since the table was created. on_schema_change: "fail" stops the run outright on a mismatch; "ignore" (the default) or a misconfigured "append_new_columns"/"sync_all_columns" can instead silently trigger a full rebuild depending on what actually changed.

Why it's quiet: no error, just a much longer run

A full refresh isn't a failure state from dbt's point of view - the run still shows green. What changes is execution_time in run_results.json jumping well past its normal range, and a warehouse bill line that doesn't match "we only added a day's worth of rows." Nobody's paged for a slow success.

How to catch it before the bill does

Compare each incremental model's execution_time against its own trailing average, not a fixed threshold - a model that normally takes 40 seconds jumping to 25 minutes is the signal, regardless of what "acceptable" looks like for a different model in the same project.

A model this happens to isn't broken - its config is just doing exactly what you told it to under a condition you didn't expect to hit. The fix is almost always in on_schema_change or in whatever upstream change altered the column set, not in the model's SQL itself.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back - not another rule of thumb.