dorfteich/docs/architecture/operations.md
Claude Fable 5 8dbff86537
All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m9s
CI / Build container images (push) Has been skipped
CD / Build and push images (push) Successful in 3m47s
CD / Deploy to Test (push) Successful in 8s
CD / Smoke tests against Test (push) Successful in 1m10s
CD / Promote to Int (push) Successful in 9s
CI / Auth e2e pack (push) Successful in 5m25s
CI / Import/export fidelity gate (push) Successful in 45s
Add backup sidecar: nightly dump, volume archive, prune, status, restore (#83)
New apps/backup service (ADR 0015): nightly pg_dump -Fc plus one tar of the
uploads/plugins volumes as a consistent restore set on a new backups volume,
retention prune that never removes the newest complete set, atomic
status.json for the readiness/admin consumers (#85/#86), and a failure mail
sent directly via nodemailer (the api may be the broken part) with de/en
texts in the shared mails catalog. BACKUP_RUN_ONCE=1 gives the on-demand
path; deploy/backup/restore.sh automates the documented restore runbook.
The pure secret-store helpers moved to @dorfteich/shared so the sidecar
resolves the wizard-written SMTP relay exactly like the api.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
2026-07-11 17:29:15 +02:00

92 lines
5.2 KiB
Markdown

# Operations concept
Pragmatic monitoring (kickoff decision): health checks, external uptime
monitoring, structured logs, backup alerting — no dedicated metrics stack.
## Health & monitoring
- **Health endpoints**: `api` exposes `/healthz` (liveness: process up) and
`/readyz` (readiness: DB reachable, migrations applied, converters
reachable, backup freshness < 26 h). `collab` exposes `/healthz`
(process + DB). `web` serves a static `/healthz`.
- **Docker healthchecks** on every service (compose `healthcheck:`), so
`docker compose ps` and restarts reflect real state;
`restart: unless-stopped` everywhere.
- **External uptime monitoring**: the operator's existing Uptime-Kuma
monitors `https://dorfteich.online/healthz` (web), `/api/v1/healthz`, and
a WebSocket check on `/collab`, with notification on failure. Test/Int
get web-check-only monitors (no paging).
- **Backup alerting**: the backup sidecar writes
`backups/status.json` after every run; `readyz` degrades when the last
successful backup is older than 26 h, which surfaces through Uptime-Kuma
without extra tooling. Additionally the sidecar sends a failure e-mail
via the instance SMTP.
## Logging
- All services log **structured JSON to stdout** (pino); Docker's json-file
driver with rotation (`max-size: 10m`, `max-file: 5`).
- Log content rules: request logs with method/route/status/duration/user id
(no request bodies), auth events (login success/failure, permission
denials), admin actions (grants, plugin installs, quota changes) as an
**audit trail**, collab session open/close. Never log passwords, tokens,
session ids, or page content.
- Reading logs = `docker compose logs` / `docker logs` on the host; no
central log stack at this scale (revisit if a second Prod host appears).
## Backup & restore (operational view of ADR 0015)
- Nightly at 03:00 stage-local time (sidecar `backup`, issue #83; env
`BACKUP_TIME`/`TZ`): `pg_dump -Fc` uploads/plugins volume archive (one
tar, same backup id `YYYYMMDD-HHMMSS`) prune (`BACKUP_RETENTION_DAYS`,
30 default / 7 Test+Int; the newest complete set always survives)
`status.json` on the `backups` volume on failure a mail directly via
the instance SMTP to `BACKUP_MAIL_TO`. Mirror to BASEL is issue #84.
- **On-demand run**: `docker compose run --rm -e BACKUP_RUN_ONCE=1 backup`
(exit code = outcome); list sets with `docker compose exec backup ls /backups`.
- **Restore runbook** (also the Prod-relocation procedure) automated by
`deploy/backup/restore.sh <backup-id>`, run from the stage directory:
1. stop the app services (`web`, `api`, `collab`; the db stays up),
2. restore DB: `pg_restore --clean --if-exists` into the `db` container,
3. restore volume: unpack the matching uploads/plugins archive,
4. `docker compose up -d`, verify `/readyz`, spot-check a page + a file.
- **Drills**: monthly automated restore of the latest Prod dump into a
scratch database on Test with a row-count sanity report; quarterly manual
full-runbook drill on Test.
- **Admin UI**: Site Admin can download the latest dump/archive and trigger
an on-demand backup run (ADR 0015 restore stays CLI-only).
## Maintenance jobs (in-app scheduler, `jobs` table)
| Job | Cadence | Purpose |
| --------------------- | ----------------------- | -------------------------------------------------- |
| trash purge | daily | delete pages/ponds past trash retention (ADR 0013) |
| version thinning | daily | auto-version retention policy (ADR 0013) |
| update-log compaction | hourly, idle pages only | bound Yjs log growth |
| quota reconciliation | nightly | recompute `pond_usage`, report drift |
| orphan file sweep | nightly | volume DB consistency (ADR 0011) |
| mail outbox retry | every minute | e-mail delivery with backoff |
Job outcomes are visible in the Site Admin UI (last run, status) that
panel is the operator's single glance for instance health.
## Update strategy
- **Own stages**: pipeline-driven (see `deployment.md`); Prod only via
approved release tags.
- **Self-hosters**: semver releases; `docker compose pull && up -d`;
migrations auto-apply; release notes flag manual steps and `migration`
label. Supported downgrade window: one minor release.
- **Base image / dependency hygiene**: monthly dependency-update story
(renovate-style batch PR); security advisories for pinned images tracked
in the release checklist.
## Capacity & limits (initial values, instance-tunable)
| Limit | Default |
| ------------------------------- | ------------------------------------------------ |
| max page document size | 5 MiB Yjs state |
| max upload size | 25 MiB (quota ladder, ADR 0011) |
| collab connections per instance | 500 concurrent |
| rate limits | login 10/min/IP, signup 5/h/IP, API 100/min/user |