New scheduled workflow (monthly + on demand) runs deploy/backup/drill.sh: it reads the drilled stage's backups volume strictly read-only, restores the latest successful set into a throwaway Postgres and volumes under a unique drill prefix via the backup image's restore path, boots the api against the result, and verifies readyz (database + migrations), row counts, rendered content in the page cache, a public API request, and a media byte-check against the attachments table — then tears everything down, also on failure. Each run reports its outcome as a comment on the pinned "Restore drills" issue (#98). docs/operations/restore-runbook.md carries the manual procedure, which doubles as the Prod relocation path; pre-go-live the drill restores the Test set (switch the source volume at go-live, #89 — off-host fetch from the BASEL mirror stays with #84). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
3.0 KiB
Restore runbook (ADR 0015, issue #87)
How to restore a Dorfteich stage from a nightly backup set — manually in an incident, automatically as the monthly drill. The same procedure doubles as the Prod relocation procedure: restore the latest set on the new host.
A restore set is one backup id YYYYMMDD-HHMMSS: db-<id>.dump
(pg_dump -Fc) plus files-<id>.tar.gz (uploads + plugins volumes),
written nightly by the backup sidecar onto the backups volume, with
status.json describing the last run (deploy/monitoring.md).
Manual restore (incident / relocation)
On the stage host, from the stage directory (/home/DOCKER/dorfteich-<stage>/):
- Pick the set.
docker compose exec backup ls /backups— usually the id instatus.json→lastSuccess.backupId. - Run the automated runbook:
./restore.sh <backup-id>(deploy/backup/restore.sh). It stopsweb/api/collab(the db stays up), replays the dump withpg_restore --clean --if-existsand unpacks the volume archive through the backup sidecar image, starts the stack, and polls/readyz. - Verify:
/readyzfully green, spot-check one page and one uploaded file in the browser.
Consistency model (ADR 0015): the volume archive is taken minutes after the dump — a page referencing a file uploaded in between shows a missing image, never corruption.
Relocation to a new host: provision the stage directory (compose +
.env, deploy/stages.md), start only db and backup
(docker compose up -d db backup), copy the set into the backups volume
(docker run --rm -v <src> -v <project>_backups:/backups …), then steps 2–3.
Automated monthly drill (.gitea/workflows/drill.yml)
Runs on the 1st of each month (and on demand via Run workflow): it
executes deploy/backup/drill.sh, which
- reads the drilled stage's backups volume read-only (stage volumes are
never touched — everything scratch lives under a unique
dorfteich-drill-<timestamp>prefix and is removed afterwards), - restores the latest successful set into a throwaway Postgres + volumes
using the same backup-image code path as
restore.sh, - boots the api image against the result and checks: readyz database + migrations ok, ≥ 1 user and live page, ≥ 1 rendered page in the content cache, a public API request answers 200, and one media file's bytes on the volume match its database row,
- reports the outcome as a comment on the pinned Restore drills issue (#98), then tears the scratch environment down (also on failure).
Pre-go-live the drill restores the Test stage's set
(DRILL_SOURCE_VOLUME: dorfteich-test_backups); at go-live (#89) point it
at the Prod backups volume. Manual invocation on the stage host:
SOURCE_VOLUME=dorfteich-test_backups sh deploy/backup/drill.sh
A drill failure means the current backup set is not restorable — treat
it like a failed backup: check the sidecar logs and status.json, fix, and
re-run the drill the same day.