Off-host backups for every self-hoster, configured entirely in the admin UI — supersedes the host-specific mirror plan behind #84. shared: - webdav.ts (new package entry like token-crypto): minimal WebDAV client with basic auth — PROPFIND (tolerant multistatus parser), MKCOL, PUT (streamed), GET, DELETE; Nextcloud DAV path derived from the plain server URL, explicit DAV bases pass through - backup-status.ts: additive remote-upload status in status.json, the restore-status.json contract (running/succeeded/failed + staleness bound), the backup_command/backup_maintenance NOTIFY channels, and the one-bundle-per-set naming (dorfteich-backup-<id>.tar.gz) - backup-set.ts moved here from apps/backup (api lists local sets) backup sidecar: - reads the backup.* instance settings directly from the database (admin changes apply next run; local retention row overrides the env) and the app password from the secret store - after each successful set: bundle dump + files archive + manifest into ONE self-contained tar.gz, upload via WebDAV per schedule (off/daily/weekly; manual runs always upload), prune remote bundles — never the newest — and record the outcome in status.json; upload failures alert via a new backupUploadFailed mail (de+en) - command listener on backup_command (run / restore) with a serial queue against the nightly timer - restore orchestrator: restore-status.json → maintenance NOTIFY → grace → (remote: download + manifest-verify bundle) → terminate other DB connections → shared perform-restore path (same code as restore.sh) → final status + maintenance exit api: - MaintenanceGuard (global, registered before the setup gate): 503 maintenance_mode while restore-status says running; health endpoints and the new public GET /backup/restore-status stay exempt; a stale running state (crashed sidecar) unblocks after 30 min - MaintenanceStateService watches the file and restarts the api after a successful restore (fresh caches, migrate-on-start for older dumps); main.ts refuses to touch the database while a restore runs — a container restarting mid-restore must not race pg_restore with migrate deploy - worker sweeps (conversion, mail outbox, scheduler) catch transient database failures instead of dying on an unhandled rejection — the restore's connection termination crashed the api in verification - backup admin endpoints under /admin/system/backup: settings (live connection test before save, password write-only into the secret store), nextcloud/test, sets (local via the ro backups mount + remote via WebDAV), run + restore (type-to-confirm backstop, source validation) — commands travel as NOTIFY payloads; audit actions backup.settings_changed/run_triggered/restore_requested - readyz: new warning-level backup_remote check while a target is configured (26 h daily / 170 h weekly bound) collab: - maintenance listener: on enter, persist + close every live session and refuse new connections until exit (failsafe timeout 30 min) — no in-memory document may write pre-restore content back afterwards web: - Admin → System backup section: status card with remote facts and a "Back up now" button, the Nextcloud settings form with test button, and the restore picker (local + remote sets, type-to-confirm) - global maintenance screen: any 503 maintenance_mode flips the SPA to a status page polling the exempt endpoint, reloading when the instance returns Verified end-to-end against a live stack (fresh DB, native api + sidecar, fake WebDAV server): configure → test → manual backup → bundle upload → readyz/sets/status surfaces → remote restore with maintenance gate, marker rollback and api restart; suites: shared 21, backup 9, collab 11, api 58 files green, lint + i18n:check + typecheck clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
4.3 KiB
Restore runbook (ADR 0015, issue #87)
How to restore a Dorfteich stage from a nightly backup set — manually in an incident, automatically as the monthly drill. The same procedure doubles as the Prod relocation procedure: restore the latest set on the new host.
A restore set is one backup id YYYYMMDD-HHMMSS: db-<id>.dump
(pg_dump -Fc) plus files-<id>.tar.gz (uploads + plugins volumes),
written nightly by the backup sidecar onto the backups volume, with
status.json describing the last run (deploy/monitoring.md).
Manual restore (incident / relocation)
On the stage host, from the stage directory (/home/DOCKER/dorfteich-<stage>/):
- Pick the set.
docker compose exec backup ls /backups— usually the id instatus.json→lastSuccess.backupId. - Run the automated runbook:
./restore.sh <backup-id>(deploy/backup/restore.sh). It stopsweb/api/collab(the db stays up), replays the dump withpg_restore --clean --if-existsand unpacks the volume archive through the backup sidecar image, starts the stack, and polls/readyz. - Verify:
/readyzfully green, spot-check one page and one uploaded file in the browser.
Consistency model (ADR 0015): the volume archive is taken minutes after the dump — a page referencing a file uploaded in between shows a missing image, never corruption.
Relocation to a new host: provision the stage directory (compose +
.env, deploy/stages.md), start only db and backup
(docker compose up -d db backup), copy the set into the backups volume
(docker run --rm -v <src> -v <project>_backups:/backups …), then steps 2–3.
In-app restore (issue #103)
With a Nextcloud target configured, Site Admins can restore without shell
access: Admin → System → Backups → Restore lists local and remote sets;
after a type-to-confirm prompt the backup sidecar orchestrates the whole
restore (maintenance mode → download + verify → terminate connections →
pg_restore + volume extract → api restart). Progress lands in
restore-status.json next to status.json; the public
GET /api/v1/backup/restore-status endpoint keeps answering while
everything else serves 503 maintenance_mode. This runbook stays the
disaster path for when the app itself is gone.
Fetching a set from Nextcloud (total loss)
When the host is gone but the off-host copies exist, rebuild from the Nextcloud bundle alone:
- Download
dorfteich-backup-<id>.tar.gzfrom the configured Nextcloud folder (browser orcurl -u <user>:<app-password> -O https://cloud.example.com/remote.php/dav/files/<user>/<folder>/dorfteich-backup-<id>.tar.gz). - Unpack it:
tar -xzf dorfteich-backup-<id>.tar.gz→db-<id>.dump,files-<id>.tar.gz, and amanifest.jsondescribing the set. - Copy the two artifacts into the (fresh) stack's backups volume — see
Relocation to a new host above — and run
./restore.sh <id>.
Automated monthly drill (.gitea/workflows/drill.yml)
Runs on the 1st of each month — and on demand by pushing a drill-* tag
(git tag drill-$(date +%s) && git push origin --tags; Gitea 1.22 has no
workflow-dispatch button yet). It executes deploy/backup/drill.sh, which
- reads the drilled stage's backups volume read-only (stage volumes are
never touched — everything scratch lives under a unique
dorfteich-drill-<timestamp>prefix and is removed afterwards), - restores the latest successful set into a throwaway Postgres + volumes
using the same backup-image code path as
restore.sh, - boots the api image against the result and checks: readyz database + migrations ok, ≥ 1 user and live page, ≥ 1 rendered page in the content cache, a public API request answers 200, and one media file's bytes on the volume match its database row,
- reports the outcome as a comment on the pinned Restore drills issue (#98), then tears the scratch environment down (also on failure).
Pre-go-live the drill restores the Test stage's set
(DRILL_SOURCE_VOLUME: dorfteich-test_backups); at go-live (#89) point it
at the Prod backups volume. Manual invocation on the stage host:
SOURCE_VOLUME=dorfteich-test_backups sh deploy/backup/drill.sh
A drill failure means the current backup set is not restorable — treat
it like a failed backup: check the sidecar logs and status.json, fix, and
re-run the drill the same day.