Adds docs/operations/update-runbook.md (obtain, verify by digest, back up, apply, verify, roll back) with the migration behaviour stated explicitly: a failed migration rolls back its own transaction but is recorded in _prisma_migrations and blocks every further migrate deploy (P3009) — including a re-deployed old image — until migrate resolve --rolled-back; semantically irreversible migrations have exactly one way back, the pre-update backup set. No rolling updates on a compose stage. Rehearsed in the isolated environment of #220: regular update to a v2 image set, then a deliberate failed-update (P3018 division by zero, schema change proven rolled back) with image-rollback-alone shown insufficient and the documented recovery executed. Protocol: docs/vs-nfd/98-update-rollback-protokoll.md. ADR 0024 decisions 5+6 recorded as executed; operations handbook and restore runbook updated; plan checkbox P1-3 ticked. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AUtYMxwTCMHG9mVHnwbFg8
4.6 KiB
Restore runbook (ADR 0015, issue #87)
How to restore a Dorfteich stage from a nightly backup set — manually in an incident, automatically as the monthly drill. The same procedure doubles as the Prod relocation procedure: restore the latest set on the new host.
A restore set is one backup id YYYYMMDD-HHMMSS: db-<id>.dump
(pg_dump -Fc) plus files-<id>.tar.gz (uploads + plugins volumes),
written nightly by the backup sidecar onto the backups volume, with
status.json describing the last run (deploy/monitoring.md).
Manual restore (incident / relocation)
On the stage host, from the stage directory (/srv/DOCKER/dorfteich-<stage>/):
- Pick the set.
docker compose exec backup ls /backups— usually the id instatus.json→lastSuccess.backupId. - Run the automated runbook:
./restore.sh <backup-id>(deploy/backup/restore.sh). It stopsweb/api/collab(the db stays up), resets thepublicschema and replays the dump withpg_restore(#288 — objects created after the backup do not survive) and unpacks the volume archive through the backup sidecar image, starts the stack, and polls/readyz. - Verify:
/readyzfully green, spot-check one page and one uploaded file in the browser.
Consistency model (ADR 0015): the volume archive is taken minutes after the dump — a page referencing a file uploaded in between shows a missing image, never corruption.
Restore is also the designated way back from an update whose migrations
succeeded but must be undone — the pre-update backup set is step 3 of
docs/operations/update-runbook.md (#221).
Relocation to a new host: provision the stage directory (compose +
.env, deploy/stages.md), start only db and backup
(docker compose up -d db backup), copy the set into the backups volume
(docker run --rm -v <src> -v <project>_backups:/backups …), then steps 2–3.
In-app restore (issue #103)
With a Nextcloud target configured, Site Admins can restore without shell
access: Admin → System → Backups → Restore lists local and remote sets;
after a type-to-confirm prompt the backup sidecar orchestrates the whole
restore (maintenance mode → download + verify → terminate connections →
pg_restore + volume extract → api restart). Progress lands in
restore-status.json next to status.json; the public
GET /api/v1/backup/restore-status endpoint keeps answering while
everything else serves 503 maintenance_mode. This runbook stays the
disaster path for when the app itself is gone.
Fetching a set from Nextcloud (total loss)
When the host is gone but the off-host copies exist, rebuild from the Nextcloud bundle alone:
- Download
dorfteich-backup-<id>.tar.gzfrom the configured Nextcloud folder (browser orcurl -u <user>:<app-password> -O https://cloud.example.com/remote.php/dav/files/<user>/<folder>/dorfteich-backup-<id>.tar.gz). - Unpack it:
tar -xzf dorfteich-backup-<id>.tar.gz→db-<id>.dump,files-<id>.tar.gz, and amanifest.jsondescribing the set. - Copy the two artifacts into the (fresh) stack's backups volume — see
Relocation to a new host above — and run
./restore.sh <id>.
Automated monthly drill (.gitea/workflows/drill.yml)
Runs on the 1st of each month — and on demand by pushing a drill-* tag
(git tag drill-$(date +%s) && git push origin --tags; Gitea 1.22 has no
workflow-dispatch button yet). It executes deploy/backup/drill.sh, which
- reads the drilled stage's backups volume read-only (stage volumes are
never touched — everything scratch lives under a unique
dorfteich-drill-<timestamp>prefix and is removed afterwards), - restores the latest successful set into a throwaway Postgres + volumes
using the same backup-image code path as
restore.sh, - boots the api image against the result and checks: readyz database + migrations ok, ≥ 1 user and live page, ≥ 1 rendered page in the content cache, a public API request answers 200, and one media file's bytes on the volume match its database row,
- reports the outcome as a comment on the pinned Restore drills issue (#98), then tears the scratch environment down (also on failure).
Pre-go-live the drill restores the Test stage's set
(DRILL_SOURCE_VOLUME: dorfteich-test_backups); at go-live (#89) point it
at the Prod backups volume. Manual invocation on the stage host:
SOURCE_VOLUME=dorfteich-test_backups sh deploy/backup/drill.sh
A drill failure means the current backup set is not restorable — treat
it like a failed backup: check the sidecar logs and status.json, fix, and
re-run the drill the same day.