dbt

Junior

How do I configure dbt's profiles.yml so credentials never end up in git?

dbt splits "what to build" from "where to connect" into two files on purpose - one committed, one that should never be. Here's how that split is supposed to work.

Read in:
Hardcoded in profiles.yml Committed .env file Leaked via CI logs
Common ways warehouse credentials end up somewhere they shouldn't, roughly in that order - not a measured breakdown.

dbt_project.yml describes your project and lives in git. profiles.yml describes how to connect to a warehouse and, by default, lives outside any repo entirely. Conflating the two is the single most common way credentials end up committed.

Where profiles.yml lives, and why that's not an accident

By default dbt looks for profiles.yml in ~/.dbt/ - not in your project directory, and not something dbt init puts under version control. You can point it elsewhere with the DBT_PROFILES_DIR environment variable (useful in CI, where there's no home directory to speak of), but the default location is the whole point: it keeps connection details physically separate from the code that gets pushed to GitHub.

# ~/.dbt/profiles.yml - never inside the dbt project itself my_project: target: dev outputs: dev: type: snowflake account: "{{ env_var('DBT_ACCOUNT') }}" user: "{{ env_var('DBT_USER') }}" password: "{{ env_var('DBT_PASSWORD') }}" role: transformer database: analytics warehouse: transforming schema: dbt_dev threads: 4

The my_project: key at the top has to match the profile: value in your dbt_project.yml - that's the only link between the two files, and it's a name, not a secret.

env_var() is the whole trick

Every value that shouldn't be readable by anyone who can cat the file goes through env_var('SOME_NAME') instead of being typed in literally. Locally, that means exporting the variable in your shell profile or a local .env file that's in .gitignore (direnv or dotenv work fine for this). In CI, it means setting the same names as encrypted secrets in whatever runs your dbt jobs - GitHub Actions secrets, Airflow connections, or your orchestrator's equivalent - never as plaintext in a workflow file.

The part people skip: env_var() takes an optional second argument as a default - env_var('DBT_THREADS', '4') - which is fine for non-secret settings, but leaving it off a password variable means a missing secret fails loudly with "env var required but not provided" instead of silently connecting as some other user. Don't give secrets defaults.

Where this still goes wrong

Two real gaps this doesn't close by itself. First, nothing stops a teammate from adding a second target block with a literal password "just for now" - profiles.yml isn't linted by dbt itself, so this needs a code-review habit, not just a template. Second, dbt debug only validates the target you're currently pointed at; a typo in a target you rarely use (a one-off staging output, say) can sit broken for months until someone finally runs against it. Neither of those is a dbt problem to fix - opti-pipe doesn't touch credentials or connection config at all, only the run_results.json a run produces after it already connected successfully.

See what this looks like on your own pipeline.

Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get concrete, approve-before-apply recommendations back - not another rule of thumb.