dbt
JuniorHow do I configure dbt's profiles.yml so credentials never end up in git?
dbt splits "what to build" from "where to connect" into two files on purpose - one committed, one that should never be. Here's how that split is supposed to work.
dbt_project.yml describes your project and lives in git. profiles.yml describes how to connect to a warehouse and, by default, lives outside any repo entirely. Conflating the two is the single most common way credentials end up committed.
Where profiles.yml lives, and why that's not an accident
By default dbt looks for profiles.yml in ~/.dbt/ - not in your project directory,
and not something dbt init puts under version control. You can point it elsewhere with the
DBT_PROFILES_DIR environment variable (useful in CI, where there's no home directory to speak
of), but the default location is the whole point: it keeps connection details physically separate from the
code that gets pushed to GitHub.
The my_project: key at the top has to match the profile: value in your
dbt_project.yml - that's the only link between the two files, and it's a name, not a secret.
env_var() is the whole trick
Every value that shouldn't be readable by anyone who can cat the file goes through
env_var('SOME_NAME') instead of being typed in literally. Locally, that means exporting the
variable in your shell profile or a local .env file that's in .gitignore
(direnv or dotenv work fine for this). In CI, it means setting the same names as
encrypted secrets in whatever runs your dbt jobs - GitHub Actions secrets, Airflow connections, or your
orchestrator's equivalent - never as plaintext in a workflow file.
The part people skip: env_var() takes an optional second argument as a default -
env_var('DBT_THREADS', '4') - which is fine for non-secret settings, but leaving it off a
password variable means a missing secret fails loudly with "env var required but not provided" instead
of silently connecting as some other user. Don't give secrets defaults.
Where this still goes wrong
Two real gaps this doesn't close by itself. First, nothing stops a teammate from adding a second
target block with a literal password "just for now" - profiles.yml isn't linted by dbt itself,
so this needs a code-review habit, not just a template. Second, dbt debug only validates the
target you're currently pointed at; a typo in a target you rarely use (a one-off staging output,
say) can sit broken for months until someone finally runs against it. Neither of those is a dbt problem to
fix - opti-pipe doesn't touch credentials or connection config at all, only the run_results.json
a run produces after it already connected successfully.
See what this looks like on your own pipeline.
Upload a real Spark event log, dbt run_results.json, or Flink metrics export and get
concrete, approve-before-apply recommendations back - not another rule of thumb.