All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m45s
CD / Build and push images (push) Successful in 3m49s
CI / Build container images (push) Has been skipped
CD / Deploy to Test (push) Successful in 9s
CD / Smoke tests against Test (push) Successful in 1m18s
CD / Promote to Int (push) Successful in 11s
CI / Auth e2e pack (push) Successful in 5m35s
CI / Import/export fidelity gate (push) Successful in 47s
Off-host backups for every self-hoster, configured entirely in the admin UI — supersedes the host-specific mirror plan behind #84. shared: - webdav.ts (new package entry like token-crypto): minimal WebDAV client with basic auth — PROPFIND (tolerant multistatus parser), MKCOL, PUT (streamed), GET, DELETE; Nextcloud DAV path derived from the plain server URL, explicit DAV bases pass through - backup-status.ts: additive remote-upload status in status.json, the restore-status.json contract (running/succeeded/failed + staleness bound), the backup_command/backup_maintenance NOTIFY channels, and the one-bundle-per-set naming (dorfteich-backup-<id>.tar.gz) - backup-set.ts moved here from apps/backup (api lists local sets) backup sidecar: - reads the backup.* instance settings directly from the database (admin changes apply next run; local retention row overrides the env) and the app password from the secret store - after each successful set: bundle dump + files archive + manifest into ONE self-contained tar.gz, upload via WebDAV per schedule (off/daily/weekly; manual runs always upload), prune remote bundles — never the newest — and record the outcome in status.json; upload failures alert via a new backupUploadFailed mail (de+en) - command listener on backup_command (run / restore) with a serial queue against the nightly timer - restore orchestrator: restore-status.json → maintenance NOTIFY → grace → (remote: download + manifest-verify bundle) → terminate other DB connections → shared perform-restore path (same code as restore.sh) → final status + maintenance exit api: - MaintenanceGuard (global, registered before the setup gate): 503 maintenance_mode while restore-status says running; health endpoints and the new public GET /backup/restore-status stay exempt; a stale running state (crashed sidecar) unblocks after 30 min - MaintenanceStateService watches the file and restarts the api after a successful restore (fresh caches, migrate-on-start for older dumps); main.ts refuses to touch the database while a restore runs — a container restarting mid-restore must not race pg_restore with migrate deploy - worker sweeps (conversion, mail outbox, scheduler) catch transient database failures instead of dying on an unhandled rejection — the restore's connection termination crashed the api in verification - backup admin endpoints under /admin/system/backup: settings (live connection test before save, password write-only into the secret store), nextcloud/test, sets (local via the ro backups mount + remote via WebDAV), run + restore (type-to-confirm backstop, source validation) — commands travel as NOTIFY payloads; audit actions backup.settings_changed/run_triggered/restore_requested - readyz: new warning-level backup_remote check while a target is configured (26 h daily / 170 h weekly bound) collab: - maintenance listener: on enter, persist + close every live session and refuse new connections until exit (failsafe timeout 30 min) — no in-memory document may write pre-restore content back afterwards web: - Admin → System backup section: status card with remote facts and a "Back up now" button, the Nextcloud settings form with test button, and the restore picker (local + remote sets, type-to-confirm) - global maintenance screen: any 503 maintenance_mode flips the SPA to a status page polling the exempt endpoint, reloading when the instance returns Verified end-to-end against a live stack (fresh DB, native api + sidecar, fake WebDAV server): configure → test → manual backup → bundle upload → readyz/sets/status surfaces → remote restore with maintenance gate, marker rollback and api restart; suites: shared 21, backup 9, collab 11, api 58 files green, lint + i18n:check + typecheck clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
117 lines
6.9 KiB
Markdown
117 lines
6.9 KiB
Markdown
# Operations concept
|
|
|
|
Pragmatic monitoring (kickoff decision): health checks, external uptime
|
|
monitoring, structured logs, backup alerting — no dedicated metrics stack.
|
|
|
|
## Health & monitoring
|
|
|
|
- **Health endpoints**: `api` exposes `/healthz` (liveness: process up) and
|
|
`/readyz` (readiness: DB reachable + migrations applied = hard failures
|
|
→ 503; converter/renderer reachability and backup freshness < 26 h are
|
|
warning-level → overall `status: degraded`, still HTTP 200). `collab`
|
|
exposes `/healthz` (process + DB). `web` serves a static `/healthz`.
|
|
- **Docker healthchecks** use the liveness endpoints only — never `readyz` —
|
|
so a degraded instance is not restart-looped (`restart: unless-stopped`
|
|
restarts on process exit, an `unhealthy` mark just shows in
|
|
`docker compose ps`).
|
|
- **External uptime monitoring**: the operator's existing Uptime-Kuma runs
|
|
the monitor set defined in `deploy/monitoring.md` (web healthz, api
|
|
readyz for down, a keyword monitor on the readyz body for degraded, and
|
|
a collab WebSocket check), with notification on failure for Prod.
|
|
Test/Int get reduced monitors (no paging).
|
|
- **Backup alerting**: the backup sidecar writes
|
|
`backups/status.json` after every run; `readyz` reads it (issue #85) and
|
|
degrades when the last successful backup is older than 26 h, which
|
|
surfaces through Uptime-Kuma without extra tooling. Additionally the
|
|
sidecar sends a failure e-mail via the instance SMTP.
|
|
|
|
## Logging
|
|
|
|
- All services log **structured JSON to stdout** (pino); Docker's json-file
|
|
driver with rotation (`max-size: 10m`, `max-file: 5`).
|
|
- Log content rules: request logs with method/route/status/duration/user id
|
|
(no request bodies), auth events (login success/failure, permission
|
|
denials), admin actions (grants, plugin installs, quota changes) as an
|
|
**audit trail**, collab session open/close. Never log passwords, tokens,
|
|
session ids, or page content.
|
|
- **Persistent audit trail** (issue #86): the auth events and admin actions
|
|
additionally land as rows in `audit_log` via `AuditService` — queryable in
|
|
the Site-Admin System panel (filter by actor/action/time, paginated).
|
|
Content activity (pages, files, exports, labels) stays log-only by design.
|
|
- Reading logs = `docker compose logs` / `docker logs` on the host; no
|
|
central log stack at this scale (revisit if a second Prod host appears).
|
|
|
|
## Backup & restore (operational view of ADR 0015)
|
|
|
|
- Nightly at 03:00 stage-local time (sidecar `backup`, issue #83; env
|
|
`BACKUP_TIME`/`TZ`): `pg_dump -Fc` → uploads/plugins volume archive (one
|
|
tar, same backup id `YYYYMMDD-HHMMSS`) → optional Nextcloud upload
|
|
(issue #103, see below) → prune (`BACKUP_RETENTION_DAYS`, 30 default /
|
|
7 Test+Int; a Site-Admin setting overrides the env; the newest complete
|
|
set always survives) → `status.json` on the `backups` volume → on
|
|
failure a mail directly via the instance SMTP to `BACKUP_MAIL_TO`.
|
|
Mirror to BASEL is issue #84.
|
|
- **Off-host copies** (issue #103): with a Nextcloud target configured in
|
|
the admin UI (WebDAV base URL + username + folder in instance settings,
|
|
app password in the secret store), each successful set is bundled into
|
|
ONE `dorfteich-backup-<id>.tar.gz` (dump + files archive + manifest) and
|
|
uploaded per schedule (daily/weekly/manual); remote prune mirrors the
|
|
local guarantee. The api and the sidecar talk over pg NOTIFY/LISTEN
|
|
(`backup_command`); upload failures mail like run failures, staleness
|
|
shows up as the `backup_remote` readyz check.
|
|
- **On-demand run**: `docker compose run --rm -e BACKUP_RUN_ONCE=1 backup`
|
|
(exit code = outcome); list sets with `docker compose exec backup ls /backups`.
|
|
- **Restore runbook** (also the Prod-relocation procedure) — automated by
|
|
`deploy/backup/restore.sh <backup-id>`, run from the stage directory:
|
|
1. stop the app services (`web`, `api`, `collab`; the db stays up),
|
|
2. restore DB: `pg_restore --clean --if-exists` into the `db` container,
|
|
3. restore volume: unpack the matching uploads/plugins archive,
|
|
4. `docker compose up -d`, verify `/readyz`, spot-check a page + a file.
|
|
- **Drills** (issue #87): monthly automated restore of the latest backup
|
|
set into a scratch environment via `.gitea/workflows/drill.yml` →
|
|
`deploy/backup/drill.sh` (sanity checks + report on the pinned "Restore
|
|
drills" issue); manual procedure and relocation notes in
|
|
`docs/operations/restore-runbook.md`. Quarterly manual full-runbook
|
|
drill on Test.
|
|
- **Admin UI** (issue #103): Site Admin triggers on-demand backups
|
|
("Back up now") and restores a local or Nextcloud set in-app — a
|
|
type-to-confirm prompt, then the sidecar orchestrates: maintenance mode
|
|
(global 503 with an exempt status endpoint), collab sessions closed via
|
|
the `backup_maintenance` NOTIFY channel, connections terminated,
|
|
`pg_restore` + volume extract, api restart. `restore.sh` remains the
|
|
disaster fallback when the app itself is gone.
|
|
|
|
## Maintenance jobs (in-app scheduler, `jobs` table)
|
|
|
|
| Job | Cadence | Purpose |
|
|
| --------------------- | ----------------------- | -------------------------------------------------- |
|
|
| trash purge | daily | delete pages/ponds past trash retention (ADR 0013) |
|
|
| version thinning | daily | auto-version retention policy (ADR 0013) |
|
|
| update-log compaction | hourly, idle pages only | bound Yjs log growth |
|
|
| quota reconciliation | nightly | recompute `pond_usage`, report drift |
|
|
| orphan file sweep | nightly | volume ↔ DB consistency (ADR 0011) |
|
|
| mail outbox retry | every minute | e-mail delivery with backoff |
|
|
|
|
Job outcomes are visible in the Site Admin UI (last run, status) — that
|
|
panel is the operator's single glance for instance health.
|
|
|
|
## Update strategy
|
|
|
|
- **Own stages**: pipeline-driven (see `deployment.md`); Prod only via
|
|
approved release tags.
|
|
- **Self-hosters**: semver releases; `docker compose pull && up -d`;
|
|
migrations auto-apply; release notes flag manual steps and `migration`
|
|
label. Supported downgrade window: one minor release.
|
|
- **Base image / dependency hygiene**: monthly dependency-update story
|
|
(renovate-style batch PR); security advisories for pinned images tracked
|
|
in the release checklist.
|
|
|
|
## Capacity & limits (initial values, instance-tunable)
|
|
|
|
| Limit | Default |
|
|
| ------------------------------- | ------------------------------------------------ |
|
|
| max page document size | 5 MiB Yjs state |
|
|
| max upload size | 25 MiB (quota ladder, ADR 0011) |
|
|
| collab connections per instance | 500 concurrent |
|
|
| rate limits | login 10/min/IP, signup 5/h/IP, API 100/min/user |
|