dorfteich/docs/architecture/operations.md
Claude Fable 5 b16d23297e Scaffold pnpm monorepo with lint, format, and test tooling
pnpm workspace with apps/web, apps/api, apps/collab, and
packages/shared; strict TypeScript base config, repo-wide ESLint (flat)
+ Prettier, Vitest per package, and root scripts lint/typecheck/test/
build. @dorfteich/shared ships a first health-response helper consumed
by apps/api to prove workspace linking. Existing markdown docs are
reformatted once by the new Prettier setup.

Closes #1

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-04 19:06:27 +02:00

4.7 KiB

Operations concept

Pragmatic monitoring (kickoff decision): health checks, external uptime monitoring, structured logs, backup alerting — no dedicated metrics stack.

Health & monitoring

  • Health endpoints: api exposes /healthz (liveness: process up) and /readyz (readiness: DB reachable, migrations applied, converters reachable, backup freshness < 26 h). collab exposes /healthz (process + DB). web serves a static /healthz.
  • Docker healthchecks on every service (compose healthcheck:), so docker compose ps and restarts reflect real state; restart: unless-stopped everywhere.
  • External uptime monitoring: the operator's existing Uptime-Kuma monitors https://dorfteich.online/healthz (web), /api/v1/healthz, and a WebSocket check on /collab, with notification on failure. Test/Int get web-check-only monitors (no paging).
  • Backup alerting: the backup sidecar writes backups/status.json after every run; readyz degrades when the last successful backup is older than 26 h, which surfaces through Uptime-Kuma without extra tooling. Additionally the sidecar sends a failure e-mail via the instance SMTP.

Logging

  • All services log structured JSON to stdout (pino); Docker's json-file driver with rotation (max-size: 10m, max-file: 5).
  • Log content rules: request logs with method/route/status/duration/user id (no request bodies), auth events (login success/failure, permission denials), admin actions (grants, plugin installs, quota changes) as an audit trail, collab session open/close. Never log passwords, tokens, session ids, or page content.
  • Reading logs = docker compose logs / docker logs on the host; no central log stack at this scale (revisit if a second Prod host appears).

Backup & restore (operational view of ADR 0015)

  • Nightly at 03:00 stage-local time: pg_dump -Fc → uploads/plugins volume archive → prune (> 30 days Prod, > 7 days Test/Int) → rsync mirror to BASEL /home/RAID/BACKUPS/dorfteich-prod/ (Prod only, dedicated dorfteich-backup user).
  • Restore runbook (also the Prod-relocation procedure):
    1. docker compose down (keep volumes),
    2. restore DB: pg_restore --clean --if-exists into the db container,
    3. restore volume: unpack the matching uploads archive,
    4. docker compose up -d, verify /readyz, spot-check a page + a file.
  • Drills: monthly automated restore of the latest Prod dump into a scratch database on Test with a row-count sanity report; quarterly manual full-runbook drill on Test.
  • Admin UI: Site Admin can download the latest dump/archive and trigger an on-demand backup run (ADR 0015 — restore stays CLI-only).

Maintenance jobs (in-app scheduler, jobs table)

Job Cadence Purpose
trash purge daily delete pages/ponds past trash retention (ADR 0013)
version thinning daily auto-version retention policy (ADR 0013)
update-log compaction hourly, idle pages only bound Yjs log growth
quota reconciliation nightly recompute pond_usage, report drift
orphan file sweep nightly volume ↔ DB consistency (ADR 0011)
mail outbox retry every minute e-mail delivery with backoff

Job outcomes are visible in the Site Admin UI (last run, status) — that panel is the operator's single glance for instance health.

Update strategy

  • Own stages: pipeline-driven (see deployment.md); Prod only via approved release tags.
  • Self-hosters: semver releases; docker compose pull && up -d; migrations auto-apply; release notes flag manual steps and migration label. Supported downgrade window: one minor release.
  • Base image / dependency hygiene: monthly dependency-update story (renovate-style batch PR); security advisories for pinned images tracked in the release checklist.

Capacity & limits (initial values, instance-tunable)

Limit Default
max page document size 5 MiB Yjs state
max upload size 25 MiB (quota ladder, ADR 0011)
collab connections per instance 500 concurrent
rate limits login 10/min/IP, signup 5/h/IP, API 100/min/user