dorfteich/docs/operations/restore-runbook.md
Claude Fable 5 69e7b1406b
Some checks failed
CD / Build and push images (push) Successful in 1m9s
CI / Lint, typecheck, test (push) Failing after 1m11s
CI / Auth e2e pack (push) Has been skipped
CI / Import/export fidelity gate (push) Has been skipped
CI / Build container images (push) Has been skipped
CD / Deploy to Test (push) Successful in 9s
CD / Smoke tests against Test (push) Successful in 1m13s
CD / Promote to Int (push) Successful in 11s
Move the stage directories to /srv/DOCKER (host consolidation)
ONE consolidated every Docker stack under /srv/DOCKER (BASEL
convention, tracked in stwaidele/infrastructure-one); the three
Dorfteich stages follow. Deploy targets in cd.yml (test/int) and
prod-deploy.yml plus the docs now point at /srv/DOCKER/dorfteich-<stage>.
Data lives in named volumes keyed by the unchanged compose project
name, so the directory move carries no data migration. The go-live
checklist's uptime-kuma path already lives under /srv/DOCKER — the doc
just catches up.

Deliberately committed together with the host-side move: this commit
must not deploy before the directories exist at the new path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 12:43:29 +02:00

89 lines
4.3 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Restore runbook (ADR 0015, issue #87)
How to restore a Dorfteich stage from a nightly backup set — manually in an
incident, automatically as the monthly drill. The same procedure doubles as
the **Prod relocation procedure**: restore the latest set on the new host.
A restore set is one backup id `YYYYMMDD-HHMMSS`: `db-<id>.dump`
(`pg_dump -Fc`) plus `files-<id>.tar.gz` (uploads + plugins volumes),
written nightly by the `backup` sidecar onto the `backups` volume, with
`status.json` describing the last run (deploy/monitoring.md).
## Manual restore (incident / relocation)
On the stage host, from the stage directory (`/srv/DOCKER/dorfteich-<stage>/`):
1. **Pick the set.** `docker compose exec backup ls /backups` — usually the
id in `status.json``lastSuccess.backupId`.
2. **Run the automated runbook:** `./restore.sh <backup-id>`
(`deploy/backup/restore.sh`). It stops `web`/`api`/`collab` (the db stays
up), replays the dump with `pg_restore --clean --if-exists` and unpacks
the volume archive through the backup sidecar image, starts the stack,
and polls `/readyz`.
3. **Verify:** `/readyz` fully green, spot-check one page and one uploaded
file in the browser.
Consistency model (ADR 0015): the volume archive is taken minutes after the
dump — a page referencing a file uploaded in between shows a missing image,
never corruption.
**Relocation to a new host:** provision the stage directory (compose +
`.env`, deploy/stages.md), start only `db` and `backup`
(`docker compose up -d db backup`), copy the set into the backups volume
(`docker run --rm -v <src> -v <project>_backups:/backups …`), then steps 23.
## In-app restore (issue #103)
With a Nextcloud target configured, Site Admins can restore without shell
access: _Admin → System → Backups → Restore_ lists local and remote sets;
after a type-to-confirm prompt the backup sidecar orchestrates the whole
restore (maintenance mode → download + verify → terminate connections →
`pg_restore` + volume extract → api restart). Progress lands in
`restore-status.json` next to `status.json`; the public
`GET /api/v1/backup/restore-status` endpoint keeps answering while
everything else serves 503 `maintenance_mode`. This runbook stays the
disaster path for when the app itself is gone.
## Fetching a set from Nextcloud (total loss)
When the host is gone but the off-host copies exist, rebuild from the
Nextcloud bundle alone:
1. Download `dorfteich-backup-<id>.tar.gz` from the configured Nextcloud
folder (browser or `curl -u <user>:<app-password> -O
https://cloud.example.com/remote.php/dav/files/<user>/<folder>/dorfteich-backup-<id>.tar.gz`).
2. Unpack it: `tar -xzf dorfteich-backup-<id>.tar.gz``db-<id>.dump`,
`files-<id>.tar.gz`, and a `manifest.json` describing the set.
3. Copy the two artifacts into the (fresh) stack's backups volume — see
_Relocation to a new host_ above — and run `./restore.sh <id>`.
## Automated monthly drill (`.gitea/workflows/drill.yml`)
Runs on the 1st of each month — and on demand by pushing a `drill-*` tag
(`git tag drill-$(date +%s) && git push origin --tags`; Gitea 1.22 has no
workflow-dispatch button yet). It executes `deploy/backup/drill.sh`, which
- reads the drilled stage's backups volume **read-only** (stage volumes are
never touched — everything scratch lives under a unique
`dorfteich-drill-<timestamp>` prefix and is removed afterwards),
- restores the latest successful set into a throwaway Postgres + volumes
using the same backup-image code path as `restore.sh`,
- boots the api image against the result and checks: readyz database +
migrations ok, ≥ 1 user and live page, ≥ 1 rendered page in the content
cache, a public API request answers 200, and one media file's bytes on
the volume match its database row,
- reports the outcome as a comment on the pinned **Restore drills** issue
(#98), then tears the scratch environment down (also on failure).
Pre-go-live the drill restores the **Test** stage's set
(`DRILL_SOURCE_VOLUME: dorfteich-test_backups`); at go-live (#89) point it
at the Prod backups volume. Manual invocation on the stage host:
```sh
SOURCE_VOLUME=dorfteich-test_backups sh deploy/backup/drill.sh
```
A drill failure means the current backup set is **not restorable** — treat
it like a failed backup: check the sidecar logs and `status.json`, fix, and
re-run the drill the same day.