dorfteich/docs/architecture/operations.md
Claude Fable 5 52192eb05f
All checks were successful
CD / Build and push images (push) Successful in 3m51s
CI / Lint, typecheck, test (push) Successful in 4m5s
CD / Deploy to Test (push) Successful in 11s
CI / Build container images (push) Has been skipped
CD / Smoke tests against Test (push) Successful in 1m11s
CD / Promote to Int (push) Successful in 12s
CI / Auth e2e pack (push) Successful in 5m52s
CI / Import/export fidelity gate (push) Successful in 47s
Backup mirror to BASEL: rsync of the sets after every successful run (#84)
The operator-level extra beside the admin-configured Nextcloud target
(#103), unblocked now that the ONE→BASEL tunnel is stable again.

- sidecar: optional mirror step (mirror.ts) driven purely by env —
  BACKUP_MIRROR_TARGET (rsync-over-ssh), BACKUP_MIRROR_SSH_KEY (private
  key on the secrets volume, never in image or repo),
  BACKUP_MIRROR_SSH_PORT. Runs after the prune of every successful run,
  so --delete aligns the remote retention with the local one (the
  newest-complete-set guarantee carries over). Only set files travel
  (db-*.dump, files-*.tar.gz); status files and bundles stay local.
  Host key pinned via accept-new into .mirror_known_hosts on the backups
  volume; fixed remote modes (dirs 750, files 640, symbolic --chmod —
  octal needs rsync ≥ 3, macOS dev machines ship 2.6.9). rsync +
  openssh-client added to the sidecar image.
- status: additive `mirror` block in status.json (outcome, transferred
  count, lastSuccessAt carried across failures) — shown on the admin
  backup card; failures alert via a new backupMirrorFailed mail (de+en)
  while the local run still counts as succeeded.
- deploy/backup-basel.md: complete BASEL-side walkthrough — dedicated
  user dorfteich-backup with a /home/ home and a bash login shell,
  explicitly avoiding the Debian backup-user (UID 34) pitfalls
  (nologin shell rejects rsync sessions, /var/backups home), key
  placement through the api container onto the secrets volume, .env
  values, on-demand verification.
- tests: rsync-arg/stats-parsing units plus an integration suite against
  the real rsync binary (local target; skips where rsync is absent) —
  transfer, idempotent re-run (0 files), retention alignment, failure
  path carrying lastSuccessAt.

Verified live against the real BASEL host from a native sidecar run:
initial transfer, host-key pinning, retention alignment after a local
prune, idempotency, and the failure path (surfaced in status.json while
the local run stayed green). BASEL side provisioned per the doc.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
2026-07-12 12:20:32 +02:00

121 lines
7.2 KiB
Markdown

# Operations concept
Pragmatic monitoring (kickoff decision): health checks, external uptime
monitoring, structured logs, backup alerting — no dedicated metrics stack.
## Health & monitoring
- **Health endpoints**: `api` exposes `/healthz` (liveness: process up) and
`/readyz` (readiness: DB reachable + migrations applied = hard failures
→ 503; converter/renderer reachability and backup freshness < 26 h are
warning-level overall `status: degraded`, still HTTP 200). `collab`
exposes `/healthz` (process + DB). `web` serves a static `/healthz`.
- **Docker healthchecks** use the liveness endpoints only never `readyz`
so a degraded instance is not restart-looped (`restart: unless-stopped`
restarts on process exit, an `unhealthy` mark just shows in
`docker compose ps`).
- **External uptime monitoring**: the operator's existing Uptime-Kuma runs
the monitor set defined in `deploy/monitoring.md` (web healthz, api
readyz for down, a keyword monitor on the readyz body for degraded, and
a collab WebSocket check), with notification on failure for Prod.
Test/Int get reduced monitors (no paging).
- **Backup alerting**: the backup sidecar writes
`backups/status.json` after every run; `readyz` reads it (issue #85) and
degrades when the last successful backup is older than 26 h, which
surfaces through Uptime-Kuma without extra tooling. Additionally the
sidecar sends a failure e-mail via the instance SMTP.
## Logging
- All services log **structured JSON to stdout** (pino); Docker's json-file
driver with rotation (`max-size: 10m`, `max-file: 5`).
- Log content rules: request logs with method/route/status/duration/user id
(no request bodies), auth events (login success/failure, permission
denials), admin actions (grants, plugin installs, quota changes) as an
**audit trail**, collab session open/close. Never log passwords, tokens,
session ids, or page content.
- **Persistent audit trail** (issue #86): the auth events and admin actions
additionally land as rows in `audit_log` via `AuditService` queryable in
the Site-Admin System panel (filter by actor/action/time, paginated).
Content activity (pages, files, exports, labels) stays log-only by design.
- Reading logs = `docker compose logs` / `docker logs` on the host; no
central log stack at this scale (revisit if a second Prod host appears).
## Backup & restore (operational view of ADR 0015)
- Nightly at 03:00 stage-local time (sidecar `backup`, issue #83; env
`BACKUP_TIME`/`TZ`): `pg_dump -Fc` uploads/plugins volume archive (one
tar, same backup id `YYYYMMDD-HHMMSS`) optional Nextcloud upload
(issue #103, see below) prune (`BACKUP_RETENTION_DAYS`, 30 default /
7 Test+Int; a Site-Admin setting overrides the env; the newest complete
set always survives) `status.json` on the `backups` volume on
failure a mail directly via the instance SMTP to `BACKUP_MAIL_TO`.
- **Private mirror** (issue #84): with `BACKUP_MIRROR_TARGET` +
`BACKUP_MIRROR_SSH_KEY` set (Prod: BASEL over WireGuard), every
successful run rsyncs the set files after the prune (`--delete` aligns
the remote retention); outcome in `status.json → mirror`, failures mail
like run failures. Setup: `deploy/backup-basel.md`.
- **Off-host copies** (issue #103): with a Nextcloud target configured in
the admin UI (WebDAV base URL + username + folder in instance settings,
app password in the secret store), each successful set is bundled into
ONE `dorfteich-backup-<id>.tar.gz` (dump + files archive + manifest) and
uploaded per schedule (daily/weekly/manual); remote prune mirrors the
local guarantee. The api and the sidecar talk over pg NOTIFY/LISTEN
(`backup_command`); upload failures mail like run failures, staleness
shows up as the `backup_remote` readyz check.
- **On-demand run**: `docker compose run --rm -e BACKUP_RUN_ONCE=1 backup`
(exit code = outcome); list sets with `docker compose exec backup ls /backups`.
- **Restore runbook** (also the Prod-relocation procedure) automated by
`deploy/backup/restore.sh <backup-id>`, run from the stage directory:
1. stop the app services (`web`, `api`, `collab`; the db stays up),
2. restore DB: `pg_restore --clean --if-exists` into the `db` container,
3. restore volume: unpack the matching uploads/plugins archive,
4. `docker compose up -d`, verify `/readyz`, spot-check a page + a file.
- **Drills** (issue #87): monthly automated restore of the latest backup
set into a scratch environment via `.gitea/workflows/drill.yml`
`deploy/backup/drill.sh` (sanity checks + report on the pinned "Restore
drills" issue); manual procedure and relocation notes in
`docs/operations/restore-runbook.md`. Quarterly manual full-runbook
drill on Test.
- **Admin UI** (issue #103): Site Admin triggers on-demand backups
("Back up now") and restores a local or Nextcloud set in-app a
type-to-confirm prompt, then the sidecar orchestrates: maintenance mode
(global 503 with an exempt status endpoint), collab sessions closed via
the `backup_maintenance` NOTIFY channel, connections terminated,
`pg_restore` + volume extract, api restart. `restore.sh` remains the
disaster fallback when the app itself is gone.
## Maintenance jobs (in-app scheduler, `jobs` table)
| Job | Cadence | Purpose |
| --------------------- | ----------------------- | -------------------------------------------------- |
| trash purge | daily | delete pages/ponds past trash retention (ADR 0013) |
| version thinning | daily | auto-version retention policy (ADR 0013) |
| update-log compaction | hourly, idle pages only | bound Yjs log growth |
| quota reconciliation | nightly | recompute `pond_usage`, report drift |
| orphan file sweep | nightly | volume DB consistency (ADR 0011) |
| mail outbox retry | every minute | e-mail delivery with backoff |
Job outcomes are visible in the Site Admin UI (last run, status) that
panel is the operator's single glance for instance health.
## Update strategy
- **Own stages**: pipeline-driven (see `deployment.md`); Prod only via
approved release tags.
- **Self-hosters**: semver releases; `docker compose pull && up -d`;
migrations auto-apply; release notes flag manual steps and `migration`
label. Supported downgrade window: one minor release.
- **Base image / dependency hygiene**: monthly dependency-update story
(renovate-style batch PR); security advisories for pinned images tracked
in the release checklist.
## Capacity & limits (initial values, instance-tunable)
| Limit | Default |
| ------------------------------- | ------------------------------------------------ |
| max page document size | 5 MiB Yjs state |
| max upload size | 25 MiB (quota ladder, ADR 0011) |
| collab connections per instance | 500 concurrent |
| rate limits | login 10/min/IP, signup 5/h/IP, API 100/min/user |