Initial deliverable of the architecture phase: 16 ADRs (stack, CRDT collaboration, plugin sandbox, import/export, backups, CI/CD), data model, permission model, real-time collaboration and plugin concepts, deployment/operations/security documentation, and the milestone roadmap that the implementation issues are derived from. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
4.1 KiB
4.1 KiB
Operations concept
Pragmatic monitoring (kickoff decision): health checks, external uptime monitoring, structured logs, backup alerting — no dedicated metrics stack.
Health & monitoring
- Health endpoints:
apiexposes/healthz(liveness: process up) and/readyz(readiness: DB reachable, migrations applied, converters reachable, backup freshness < 26 h).collabexposes/healthz(process + DB).webserves a static/healthz. - Docker healthchecks on every service (compose
healthcheck:), sodocker compose psand restarts reflect real state;restart: unless-stoppedeverywhere. - External uptime monitoring: the operator's existing Uptime-Kuma
monitors
https://dorfteich.online/healthz(web),/api/v1/healthz, and a WebSocket check on/collab, with notification on failure. Test/Int get web-check-only monitors (no paging). - Backup alerting: the backup sidecar writes
backups/status.jsonafter every run;readyzdegrades when the last successful backup is older than 26 h, which surfaces through Uptime-Kuma without extra tooling. Additionally the sidecar sends a failure e-mail via the instance SMTP.
Logging
- All services log structured JSON to stdout (pino); Docker's json-file
driver with rotation (
max-size: 10m,max-file: 5). - Log content rules: request logs with method/route/status/duration/user id (no request bodies), auth events (login success/failure, permission denials), admin actions (grants, plugin installs, quota changes) as an audit trail, collab session open/close. Never log passwords, tokens, session ids, or page content.
- Reading logs =
docker compose logs/docker logson the host; no central log stack at this scale (revisit if a second Prod host appears).
Backup & restore (operational view of ADR 0015)
- Nightly at 03:00 stage-local time:
pg_dump -Fc→ uploads/plugins volume archive → prune (> 30 days Prod, > 7 days Test/Int) → rsync mirror to BASEL/home/RAID/BACKUPS/dorfteich-prod/(Prod only, dedicateddorfteich-backupuser). - Restore runbook (also the Prod-relocation procedure):
docker compose down(keep volumes),- restore DB:
pg_restore --clean --if-existsinto thedbcontainer, - restore volume: unpack the matching uploads archive,
docker compose up -d, verify/readyz, spot-check a page + a file.
- Drills: monthly automated restore of the latest Prod dump into a scratch database on Test with a row-count sanity report; quarterly manual full-runbook drill on Test.
- Admin UI: Site Admin can download the latest dump/archive and trigger an on-demand backup run (ADR 0015 — restore stays CLI-only).
Maintenance jobs (in-app scheduler, jobs table)
| Job | Cadence | Purpose |
|---|---|---|
| trash purge | daily | delete pages/ponds past trash retention (ADR 0013) |
| version thinning | daily | auto-version retention policy (ADR 0013) |
| update-log compaction | hourly, idle pages only | bound Yjs log growth |
| quota reconciliation | nightly | recompute pond_usage, report drift |
| orphan file sweep | nightly | volume ↔ DB consistency (ADR 0011) |
| mail outbox retry | every minute | e-mail delivery with backoff |
Job outcomes are visible in the Site Admin UI (last run, status) — that panel is the operator's single glance for instance health.
Update strategy
- Own stages: pipeline-driven (see
deployment.md); Prod only via approved release tags. - Self-hosters: semver releases;
docker compose pull && up -d; migrations auto-apply; release notes flag manual steps andmigrationlabel. Supported downgrade window: one minor release. - Base image / dependency hygiene: monthly dependency-update story (renovate-style batch PR); security advisories for pinned images tracked in the release checklist.
Capacity & limits (initial values, instance-tunable)
| Limit | Default |
|---|---|
| max page document size | 5 MiB Yjs state |
| max upload size | 25 MiB (quota ladder, ADR 0011) |
| collab connections per instance | 500 concurrent |
| rate limits | login 10/min/IP, signup 5/h/IP, API 100/min/user |