# Database migrations (/docs/operations/migrations)



Ductor manages its PostgreSQL schema with [goose](https://github.com/pressly/goose).
Migrations are goose-style numbered SQL files under
`infrastructure/storage/postgres/cqrs/migrations/` — hundreds of them, with the
highest version in the 450s as of writing — embedded in the binary, so the same
`ductor` you deploy carries the exact schema it expects.

## Run migrations [#run-migrations]

The `migrate` subcommand takes one positional argument:

```bash
ductor migrate up       # apply all pending migrations (default)
ductor migrate status   # list applied and pending migrations
ductor migrate down     # roll back the most recent migration
```

It uses `database.url` / `DUCTOR_DATABASE_URL`. `up` is the default, so bare
`ductor migrate` applies pending migrations. `down` exists for local iteration
only — production schema changes are **forward-only** (see [Forward-only
rollouts in production](#forward-only-rollouts-in-production) below).

## Auto-migrate vs explicit [#auto-migrate-vs-explicit]

Ductor can migrate at startup, but the default is **off**:

| Setting                                                  | Default | Behavior                                                    |
| -------------------------------------------------------- | ------- | ----------------------------------------------------------- |
| `database.auto_migrate` (`DUCTOR_DATABASE_AUTO_MIGRATE`) | `false` | When `true`, `serve` applies pending migrations before boot |

Regardless of that flag, `serve` always checks migration status at startup and
logs a warning if anything is pending:

```
pending migrations detected — run `ductor migrate up` or set database.auto_migrate=true
```

**Recommended by environment:**

* **Local / `ductor dev`** — auto-migrate is convenient; the `dev` command turns
  it on for you, and the dev compose sets `DUCTOR_DATABASE_AUTO_MIGRATE: "true"`.
* **Production** — keep it off and run `ductor migrate up` as an explicit,
  observable deploy step. This makes each schema change a deliberate action you
  can gate and log independently of the app rollout. Schema changes are
  forward-only here: recover from a bad one by rolling the *application* back
  and shipping a forward fix-up migration, never `ductor migrate down` (see
  [Forward-only rollouts in production](#forward-only-rollouts-in-production)).

## Migrations under Helm [#migrations-under-helm]

The Helm chart runs migrations as a **retained-on-failure pre-install /
pre-upgrade Job** (`<release>-migrate`) whose container executes `ductor migrate
up`. So on Kubernetes the schema is migrated *before* new pods roll — and the app
keeps `database.auto_migrate` off. A failed migration blocks the release and the
Job is **kept** so you can inspect it:

```bash
kubectl logs job/ductor-migrate
kubectl describe job/ductor-migrate
```

```mermaid
flowchart TD
    Job["pre-upgrade Job<br/>ductor migrate up"] --> Result{"migration result"}
    Result -->|success| Roll["new pods roll"]
    Result -->|failure| Block["release blocked<br/>Job retained for inspection"]
```

Toggle with `migrations.enabled` (default `true`). GitOps operators that run
migrations separately may set it `false`, but must run `ductor migrate up`
before application pods start. Migration rollback is operator-controlled
database recovery, not an automatic Helm rollback.

## Gate rollouts on schema readiness [#gate-rollouts-on-schema-readiness]

`ductor schema status` verifies the schema is fully applied and exits non-zero
until it is — a clean CI/CD gate:

```bash
ductor schema status --database-url "$DUCTOR_DATABASE_URL" --timeout 15s
# prints: schema=ready
```

## Forward-only rollouts in production [#forward-only-rollouts-in-production]

In production the schema only ever moves forward. The Helm migrate Job runs
`ductor migrate up` and never `migrate down`; treat every applied migration as
permanent history. The `down` command exists for local iteration, but there is
no tested down path in production and running it breaks the forward-only
invariant.

This works because of a **one-release backward-compatibility contract**: every
schema change must be compatible with the code **one release back**. During a
rolling upgrade the old and new pod images both run against the same database,
and the migrate hook applies the new schema *before* the new pods are healthy —
so a migration the previous image can't tolerate takes down the still-running
old pods. Keeping each change compatible one release back is exactly what makes
an image rollback safe.

### Recovering from a bad migration [#recovering-from-a-bad-migration]

Because the previous image is compatible with the newer schema, the safe
recovery is to roll the **application** back, not the schema:

<Steps>
  <Step>
    **Roll pods back** with `helm rollback <release> <previous-revision>`. This
    reverts the *image*, not the schema. The old image runs correctly against the
    ahead-of-code database, and the old chart's migrate Job re-runs `migrate up`,
    which is a **no-op** against the already-newer schema. A Helm rollback does not
    revert migrations — and it doesn't need to.
  </Step>

  <Step>
    **Ship a forward fix-up migration** (release N+1) that corrects the bad change,
    and roll it out normally. Never reach for `ductor migrate down`.
  </Step>
</Steps>

<Callout type="warn" title="helm rollback never reverts the schema">
  A `helm rollback` reverts only the running image. The one-release
  compatibility guarantee is what makes that safe — the old image tolerates the
  newer schema. For anything involving data loss, restore from a
  [backup](/docs/operations/backup-restore) rather than attempting a down path.
</Callout>

## Writing migrations that roll out safely [#writing-migrations-that-roll-out-safely]

Backward compatibility forces **expand/contract** for any destructive change:

* **Release N (expand)** — add the new column/table/constraint and stop reading
  or writing the old column in code. Ship it and let it roll out everywhere.
* **Release N+1 (contract)** — drop the now-unused column. Because release N
  already stopped using it and is fully deployed, no running pod references it.

Dropping a column in the *same* release that stopped using it is the classic
break: the previous image is still `SELECT`ing that column when the migrate hook
removes it, and every old pod errors until the rollout completes.

`scripts/check-migration-safety.sh` (run in `make custom-lint` / `make ci-fast`)
enforces this on newly added migrations. Two rules catch the expand/contract
mistakes above:

<Callout type="warn" title="DROP COLUMN needs an explicit annotation">
  A newly added migration containing `ALTER TABLE ... DROP COLUMN` is **rejected**
  unless it asserts the column was already unused as of the prior release. Once
  release N is fully rolled out, annotate the contract migration — the tag is the
  prior release, and the comment must appear verbatim:

  ```sql
  -- +goose Up
  -- migration-safety: drop-column-verified-unused <prior-release-tag>
  ALTER TABLE pools DROP COLUMN legacy_weight;
  ```
</Callout>

A migration that both `ADD CONSTRAINT ... NOT VALID` **and** `VALIDATE
CONSTRAINT` the same constraint in one file is also **rejected**. goose wraps a
migration in a single transaction, so the `ACCESS EXCLUSIVE` lock from `ADD
CONSTRAINT` is held through the full-table `VALIDATE` scan — negating the split
that `NOT VALID` exists to buy. Add the constraint `NOT VALID` in one migration
and `VALIDATE CONSTRAINT` it in a follow-up. The same header also rejects other
blocking DDL on hot tables (multi-column `ALTER`, non-literal `ADD COLUMN`
defaults, non-`CONCURRENTLY` index builds, unbounded backfills).

<Callout title="Historical migrations are grandfathered">
  A numeric ratchet (`MIGRATION_SAFETY_RATCHET_BELOW`) treats migrations below a
  version cutoff as historical and skips them, so the rules apply only to newly
  added files — pre-existing migrations that predate a rule don't retroactively
  fail the check.
</Callout>

## Typical deploy sequence [#typical-deploy-sequence]

<Steps>
  <Step>
    **Apply the schema** — or let the Helm hook do it.

    ```bash
    ductor migrate up
    ```
  </Step>

  <Step>
    **Confirm** it landed.

    ```bash
    ductor schema status
    ```
  </Step>

  <Step>
    **Roll the application.**

    ```bash
    ductor serve
    ```
  </Step>
</Steps>
