dorfteich/docs/operations/restore-runbook.md
Claude Fable 5 4f6596e8a2
Some checks failed
CI / Lint, typecheck, test (pull_request) Successful in 6m48s
CI / Build container images (pull_request) Successful in 1m46s
CI / Auth e2e pack (pull_request) Successful in 8m22s
CI / Import/export fidelity gate (pull_request) Successful in 1m8s
CD / Deploy to Test (push) Blocked by required conditions
CD / Smoke tests against Test (push) Blocked by required conditions
CD / Promote to Int (push) Blocked by required conditions
CI / Auth e2e pack (push) Blocked by required conditions
CI / Import/export fidelity gate (push) Blocked by required conditions
CI / Build container images (push) Blocked by required conditions
CD / Build and push images (push) Has been cancelled
CI / Lint, typecheck, test (push) Has been cancelled
#288: reset schema before pg_restore — partitioned tables broke --clean
Since #224 read_events is partitioned; the dump carries per-partition
primary keys as own entries, and pg_restore --clean emitted DROP
CONSTRAINT against inherited constraints, which PostgreSQL refuses. The
restore then reported FAILED although the content was restored. Dropping
and recreating the public schema first makes every --clean drop a no-op
and the restore faithful: objects created after the backup no longer
survive. Verified in the isolated environment of #220 (set
20260731-132200, exit 0, readyz green).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AUtYMxwTCMHG9mVHnwbFg8
2026-07-31 17:00:05 +02:00

90 lines
4.4 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Restore runbook (ADR 0015, issue #87)
How to restore a Dorfteich stage from a nightly backup set — manually in an
incident, automatically as the monthly drill. The same procedure doubles as
the **Prod relocation procedure**: restore the latest set on the new host.
A restore set is one backup id `YYYYMMDD-HHMMSS`: `db-<id>.dump`
(`pg_dump -Fc`) plus `files-<id>.tar.gz` (uploads + plugins volumes),
written nightly by the `backup` sidecar onto the `backups` volume, with
`status.json` describing the last run (deploy/monitoring.md).
## Manual restore (incident / relocation)
On the stage host, from the stage directory (`/srv/DOCKER/dorfteich-<stage>/`):
1. **Pick the set.** `docker compose exec backup ls /backups` — usually the
id in `status.json``lastSuccess.backupId`.
2. **Run the automated runbook:** `./restore.sh <backup-id>`
(`deploy/backup/restore.sh`). It stops `web`/`api`/`collab` (the db stays
up), resets the `public` schema and replays the dump with `pg_restore`
(#288 — objects created after the backup do not survive) and unpacks
the volume archive through the backup sidecar image, starts the stack,
and polls `/readyz`.
3. **Verify:** `/readyz` fully green, spot-check one page and one uploaded
file in the browser.
Consistency model (ADR 0015): the volume archive is taken minutes after the
dump — a page referencing a file uploaded in between shows a missing image,
never corruption.
**Relocation to a new host:** provision the stage directory (compose +
`.env`, deploy/stages.md), start only `db` and `backup`
(`docker compose up -d db backup`), copy the set into the backups volume
(`docker run --rm -v <src> -v <project>_backups:/backups …`), then steps 23.
## In-app restore (issue #103)
With a Nextcloud target configured, Site Admins can restore without shell
access: _Admin → System → Backups → Restore_ lists local and remote sets;
after a type-to-confirm prompt the backup sidecar orchestrates the whole
restore (maintenance mode → download + verify → terminate connections →
`pg_restore` + volume extract → api restart). Progress lands in
`restore-status.json` next to `status.json`; the public
`GET /api/v1/backup/restore-status` endpoint keeps answering while
everything else serves 503 `maintenance_mode`. This runbook stays the
disaster path for when the app itself is gone.
## Fetching a set from Nextcloud (total loss)
When the host is gone but the off-host copies exist, rebuild from the
Nextcloud bundle alone:
1. Download `dorfteich-backup-<id>.tar.gz` from the configured Nextcloud
folder (browser or `curl -u <user>:<app-password> -O
https://cloud.example.com/remote.php/dav/files/<user>/<folder>/dorfteich-backup-<id>.tar.gz`).
2. Unpack it: `tar -xzf dorfteich-backup-<id>.tar.gz``db-<id>.dump`,
`files-<id>.tar.gz`, and a `manifest.json` describing the set.
3. Copy the two artifacts into the (fresh) stack's backups volume — see
_Relocation to a new host_ above — and run `./restore.sh <id>`.
## Automated monthly drill (`.gitea/workflows/drill.yml`)
Runs on the 1st of each month — and on demand by pushing a `drill-*` tag
(`git tag drill-$(date +%s) && git push origin --tags`; Gitea 1.22 has no
workflow-dispatch button yet). It executes `deploy/backup/drill.sh`, which
- reads the drilled stage's backups volume **read-only** (stage volumes are
never touched — everything scratch lives under a unique
`dorfteich-drill-<timestamp>` prefix and is removed afterwards),
- restores the latest successful set into a throwaway Postgres + volumes
using the same backup-image code path as `restore.sh`,
- boots the api image against the result and checks: readyz database +
migrations ok, ≥ 1 user and live page, ≥ 1 rendered page in the content
cache, a public API request answers 200, and one media file's bytes on
the volume match its database row,
- reports the outcome as a comment on the pinned **Restore drills** issue
(#98), then tears the scratch environment down (also on failure).
Pre-go-live the drill restores the **Test** stage's set
(`DRILL_SOURCE_VOLUME: dorfteich-test_backups`); at go-live (#89) point it
at the Prod backups volume. Manual invocation on the stage host:
```sh
SOURCE_VOLUME=dorfteich-test_backups sh deploy/backup/drill.sh
```
A drill failure means the current backup set is **not restorable** — treat
it like a failed backup: check the sidecar logs and `status.json`, fix, and
re-run the drill the same day.