All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m45s
CD / Build and push images (push) Successful in 3m49s
CI / Build container images (push) Has been skipped
CD / Deploy to Test (push) Successful in 9s
CD / Smoke tests against Test (push) Successful in 1m18s
CD / Promote to Int (push) Successful in 11s
CI / Auth e2e pack (push) Successful in 5m35s
CI / Import/export fidelity gate (push) Successful in 47s
Off-host backups for every self-hoster, configured entirely in the admin UI — supersedes the host-specific mirror plan behind #84. shared: - webdav.ts (new package entry like token-crypto): minimal WebDAV client with basic auth — PROPFIND (tolerant multistatus parser), MKCOL, PUT (streamed), GET, DELETE; Nextcloud DAV path derived from the plain server URL, explicit DAV bases pass through - backup-status.ts: additive remote-upload status in status.json, the restore-status.json contract (running/succeeded/failed + staleness bound), the backup_command/backup_maintenance NOTIFY channels, and the one-bundle-per-set naming (dorfteich-backup-<id>.tar.gz) - backup-set.ts moved here from apps/backup (api lists local sets) backup sidecar: - reads the backup.* instance settings directly from the database (admin changes apply next run; local retention row overrides the env) and the app password from the secret store - after each successful set: bundle dump + files archive + manifest into ONE self-contained tar.gz, upload via WebDAV per schedule (off/daily/weekly; manual runs always upload), prune remote bundles — never the newest — and record the outcome in status.json; upload failures alert via a new backupUploadFailed mail (de+en) - command listener on backup_command (run / restore) with a serial queue against the nightly timer - restore orchestrator: restore-status.json → maintenance NOTIFY → grace → (remote: download + manifest-verify bundle) → terminate other DB connections → shared perform-restore path (same code as restore.sh) → final status + maintenance exit api: - MaintenanceGuard (global, registered before the setup gate): 503 maintenance_mode while restore-status says running; health endpoints and the new public GET /backup/restore-status stay exempt; a stale running state (crashed sidecar) unblocks after 30 min - MaintenanceStateService watches the file and restarts the api after a successful restore (fresh caches, migrate-on-start for older dumps); main.ts refuses to touch the database while a restore runs — a container restarting mid-restore must not race pg_restore with migrate deploy - worker sweeps (conversion, mail outbox, scheduler) catch transient database failures instead of dying on an unhandled rejection — the restore's connection termination crashed the api in verification - backup admin endpoints under /admin/system/backup: settings (live connection test before save, password write-only into the secret store), nextcloud/test, sets (local via the ro backups mount + remote via WebDAV), run + restore (type-to-confirm backstop, source validation) — commands travel as NOTIFY payloads; audit actions backup.settings_changed/run_triggered/restore_requested - readyz: new warning-level backup_remote check while a target is configured (26 h daily / 170 h weekly bound) collab: - maintenance listener: on enter, persist + close every live session and refuse new connections until exit (failsafe timeout 30 min) — no in-memory document may write pre-restore content back afterwards web: - Admin → System backup section: status card with remote facts and a "Back up now" button, the Nextcloud settings form with test button, and the restore picker (local + remote sets, type-to-confirm) - global maintenance screen: any 503 maintenance_mode flips the SPA to a status page polling the exempt endpoint, reloading when the instance returns Verified end-to-end against a live stack (fresh DB, native api + sidecar, fake WebDAV server): configure → test → manual backup → bundle upload → readyz/sets/status surfaces → remote restore with maintenance gate, marker rollback and api restart; suites: shared 21, backup 9, collab 11, api 58 files green, lint + i18n:check + typecheck clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
89 lines
4.3 KiB
Markdown
89 lines
4.3 KiB
Markdown
# Restore runbook (ADR 0015, issue #87)
|
||
|
||
How to restore a Dorfteich stage from a nightly backup set — manually in an
|
||
incident, automatically as the monthly drill. The same procedure doubles as
|
||
the **Prod relocation procedure**: restore the latest set on the new host.
|
||
|
||
A restore set is one backup id `YYYYMMDD-HHMMSS`: `db-<id>.dump`
|
||
(`pg_dump -Fc`) plus `files-<id>.tar.gz` (uploads + plugins volumes),
|
||
written nightly by the `backup` sidecar onto the `backups` volume, with
|
||
`status.json` describing the last run (deploy/monitoring.md).
|
||
|
||
## Manual restore (incident / relocation)
|
||
|
||
On the stage host, from the stage directory (`/home/DOCKER/dorfteich-<stage>/`):
|
||
|
||
1. **Pick the set.** `docker compose exec backup ls /backups` — usually the
|
||
id in `status.json` → `lastSuccess.backupId`.
|
||
2. **Run the automated runbook:** `./restore.sh <backup-id>`
|
||
(`deploy/backup/restore.sh`). It stops `web`/`api`/`collab` (the db stays
|
||
up), replays the dump with `pg_restore --clean --if-exists` and unpacks
|
||
the volume archive through the backup sidecar image, starts the stack,
|
||
and polls `/readyz`.
|
||
3. **Verify:** `/readyz` fully green, spot-check one page and one uploaded
|
||
file in the browser.
|
||
|
||
Consistency model (ADR 0015): the volume archive is taken minutes after the
|
||
dump — a page referencing a file uploaded in between shows a missing image,
|
||
never corruption.
|
||
|
||
**Relocation to a new host:** provision the stage directory (compose +
|
||
`.env`, deploy/stages.md), start only `db` and `backup`
|
||
(`docker compose up -d db backup`), copy the set into the backups volume
|
||
(`docker run --rm -v <src> -v <project>_backups:/backups …`), then steps 2–3.
|
||
|
||
## In-app restore (issue #103)
|
||
|
||
With a Nextcloud target configured, Site Admins can restore without shell
|
||
access: _Admin → System → Backups → Restore_ lists local and remote sets;
|
||
after a type-to-confirm prompt the backup sidecar orchestrates the whole
|
||
restore (maintenance mode → download + verify → terminate connections →
|
||
`pg_restore` + volume extract → api restart). Progress lands in
|
||
`restore-status.json` next to `status.json`; the public
|
||
`GET /api/v1/backup/restore-status` endpoint keeps answering while
|
||
everything else serves 503 `maintenance_mode`. This runbook stays the
|
||
disaster path for when the app itself is gone.
|
||
|
||
## Fetching a set from Nextcloud (total loss)
|
||
|
||
When the host is gone but the off-host copies exist, rebuild from the
|
||
Nextcloud bundle alone:
|
||
|
||
1. Download `dorfteich-backup-<id>.tar.gz` from the configured Nextcloud
|
||
folder (browser or `curl -u <user>:<app-password> -O
|
||
https://cloud.example.com/remote.php/dav/files/<user>/<folder>/dorfteich-backup-<id>.tar.gz`).
|
||
2. Unpack it: `tar -xzf dorfteich-backup-<id>.tar.gz` → `db-<id>.dump`,
|
||
`files-<id>.tar.gz`, and a `manifest.json` describing the set.
|
||
3. Copy the two artifacts into the (fresh) stack's backups volume — see
|
||
_Relocation to a new host_ above — and run `./restore.sh <id>`.
|
||
|
||
## Automated monthly drill (`.gitea/workflows/drill.yml`)
|
||
|
||
Runs on the 1st of each month — and on demand by pushing a `drill-*` tag
|
||
(`git tag drill-$(date +%s) && git push origin --tags`; Gitea 1.22 has no
|
||
workflow-dispatch button yet). It executes `deploy/backup/drill.sh`, which
|
||
|
||
- reads the drilled stage's backups volume **read-only** (stage volumes are
|
||
never touched — everything scratch lives under a unique
|
||
`dorfteich-drill-<timestamp>` prefix and is removed afterwards),
|
||
- restores the latest successful set into a throwaway Postgres + volumes
|
||
using the same backup-image code path as `restore.sh`,
|
||
- boots the api image against the result and checks: readyz database +
|
||
migrations ok, ≥ 1 user and live page, ≥ 1 rendered page in the content
|
||
cache, a public API request answers 200, and one media file's bytes on
|
||
the volume match its database row,
|
||
- reports the outcome as a comment on the pinned **Restore drills** issue
|
||
(#98), then tears the scratch environment down (also on failure).
|
||
|
||
Pre-go-live the drill restores the **Test** stage's set
|
||
(`DRILL_SOURCE_VOLUME: dorfteich-test_backups`); at go-live (#89) point it
|
||
at the Prod backups volume. Manual invocation on the stage host:
|
||
|
||
```sh
|
||
SOURCE_VOLUME=dorfteich-test_backups sh deploy/backup/drill.sh
|
||
```
|
||
|
||
A drill failure means the current backup set is **not restorable** — treat
|
||
it like a failed backup: check the sidecar logs and `status.json`, fix, and
|
||
re-run the drill the same day.
|