dorfteich/docs/architecture/operations.md
Claude Fable 5 d95c18e9e8
All checks were successful
CD / Build and push images (push) Successful in 1m5s
CD / Deploy to Test (push) Successful in 10s
CD / Smoke tests against Test (push) Successful in 1m7s
CD / Promote to Int (push) Successful in 11s
CI / Lint, typecheck, test (push) Successful in 3m18s
CI / Build container images (push) Has been skipped
CI / Auth e2e pack (push) Successful in 5m14s
CI / Import/export fidelity gate (push) Successful in 45s
Automate the monthly restore drill with a scratch-stack workflow (#87)
New scheduled workflow (monthly + on demand) runs deploy/backup/drill.sh:
it reads the drilled stage's backups volume strictly read-only, restores
the latest successful set into a throwaway Postgres and volumes under a
unique drill prefix via the backup image's restore path, boots the api
against the result, and verifies readyz (database + migrations), row
counts, rendered content in the page cache, a public API request, and a
media byte-check against the attachments table — then tears everything
down, also on failure. Each run reports its outcome as a comment on the
pinned "Restore drills" issue (#98). docs/operations/restore-runbook.md
carries the manual procedure, which doubles as the Prod relocation path;
pre-go-live the drill restores the Test set (switch the source volume at
go-live, #89 — off-host fetch from the BASEL mirror stays with #84).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
2026-07-11 20:58:06 +02:00

5.9 KiB

Operations concept

Pragmatic monitoring (kickoff decision): health checks, external uptime monitoring, structured logs, backup alerting — no dedicated metrics stack.

Health & monitoring

  • Health endpoints: api exposes /healthz (liveness: process up) and /readyz (readiness: DB reachable + migrations applied = hard failures → 503; converter/renderer reachability and backup freshness < 26 h are warning-level → overall status: degraded, still HTTP 200). collab exposes /healthz (process + DB). web serves a static /healthz.
  • Docker healthchecks use the liveness endpoints only — never readyz — so a degraded instance is not restart-looped (restart: unless-stopped restarts on process exit, an unhealthy mark just shows in docker compose ps).
  • External uptime monitoring: the operator's existing Uptime-Kuma runs the monitor set defined in deploy/monitoring.md (web healthz, api readyz for down, a keyword monitor on the readyz body for degraded, and a collab WebSocket check), with notification on failure for Prod. Test/Int get reduced monitors (no paging).
  • Backup alerting: the backup sidecar writes backups/status.json after every run; readyz reads it (issue #85) and degrades when the last successful backup is older than 26 h, which surfaces through Uptime-Kuma without extra tooling. Additionally the sidecar sends a failure e-mail via the instance SMTP.

Logging

  • All services log structured JSON to stdout (pino); Docker's json-file driver with rotation (max-size: 10m, max-file: 5).
  • Log content rules: request logs with method/route/status/duration/user id (no request bodies), auth events (login success/failure, permission denials), admin actions (grants, plugin installs, quota changes) as an audit trail, collab session open/close. Never log passwords, tokens, session ids, or page content.
  • Persistent audit trail (issue #86): the auth events and admin actions additionally land as rows in audit_log via AuditService — queryable in the Site-Admin System panel (filter by actor/action/time, paginated). Content activity (pages, files, exports, labels) stays log-only by design.
  • Reading logs = docker compose logs / docker logs on the host; no central log stack at this scale (revisit if a second Prod host appears).

Backup & restore (operational view of ADR 0015)

  • Nightly at 03:00 stage-local time (sidecar backup, issue #83; env BACKUP_TIME/TZ): pg_dump -Fc → uploads/plugins volume archive (one tar, same backup id YYYYMMDD-HHMMSS) → prune (BACKUP_RETENTION_DAYS, 30 default / 7 Test+Int; the newest complete set always survives) → status.json on the backups volume → on failure a mail directly via the instance SMTP to BACKUP_MAIL_TO. Mirror to BASEL is issue #84.
  • On-demand run: docker compose run --rm -e BACKUP_RUN_ONCE=1 backup (exit code = outcome); list sets with docker compose exec backup ls /backups.
  • Restore runbook (also the Prod-relocation procedure) — automated by deploy/backup/restore.sh <backup-id>, run from the stage directory:
    1. stop the app services (web, api, collab; the db stays up),
    2. restore DB: pg_restore --clean --if-exists into the db container,
    3. restore volume: unpack the matching uploads/plugins archive,
    4. docker compose up -d, verify /readyz, spot-check a page + a file.
  • Drills (issue #87): monthly automated restore of the latest backup set into a scratch environment via .gitea/workflows/drill.ymldeploy/backup/drill.sh (sanity checks + report on the pinned "Restore drills" issue); manual procedure and relocation notes in docs/operations/restore-runbook.md. Quarterly manual full-runbook drill on Test.
  • Admin UI: Site Admin can download the latest dump/archive and trigger an on-demand backup run (ADR 0015 — restore stays CLI-only).

Maintenance jobs (in-app scheduler, jobs table)

Job Cadence Purpose
trash purge daily delete pages/ponds past trash retention (ADR 0013)
version thinning daily auto-version retention policy (ADR 0013)
update-log compaction hourly, idle pages only bound Yjs log growth
quota reconciliation nightly recompute pond_usage, report drift
orphan file sweep nightly volume ↔ DB consistency (ADR 0011)
mail outbox retry every minute e-mail delivery with backoff

Job outcomes are visible in the Site Admin UI (last run, status) — that panel is the operator's single glance for instance health.

Update strategy

  • Own stages: pipeline-driven (see deployment.md); Prod only via approved release tags.
  • Self-hosters: semver releases; docker compose pull && up -d; migrations auto-apply; release notes flag manual steps and migration label. Supported downgrade window: one minor release.
  • Base image / dependency hygiene: monthly dependency-update story (renovate-style batch PR); security advisories for pinned images tracked in the release checklist.

Capacity & limits (initial values, instance-tunable)

Limit Default
max page document size 5 MiB Yjs state
max upload size 25 MiB (quota ladder, ADR 0011)
collab connections per instance 500 concurrent
rate limits login 10/min/IP, signup 5/h/IP, API 100/min/user