All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m9s
CI / Build container images (push) Has been skipped
CD / Build and push images (push) Successful in 3m47s
CD / Deploy to Test (push) Successful in 8s
CD / Smoke tests against Test (push) Successful in 1m10s
CD / Promote to Int (push) Successful in 9s
CI / Auth e2e pack (push) Successful in 5m25s
CI / Import/export fidelity gate (push) Successful in 45s
New apps/backup service (ADR 0015): nightly pg_dump -Fc plus one tar of the uploads/plugins volumes as a consistent restore set on a new backups volume, retention prune that never removes the newest complete set, atomic status.json for the readiness/admin consumers (#85/#86), and a failure mail sent directly via nodemailer (the api may be the broken part) with de/en texts in the shared mails catalog. BACKUP_RUN_ONCE=1 gives the on-demand path; deploy/backup/restore.sh automates the documented restore runbook. The pure secret-store helpers moved to @dorfteich/shared so the sidecar resolves the wizard-written SMTP relay exactly like the api. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
5.2 KiB
5.2 KiB
Operations concept
Pragmatic monitoring (kickoff decision): health checks, external uptime monitoring, structured logs, backup alerting — no dedicated metrics stack.
Health & monitoring
- Health endpoints:
apiexposes/healthz(liveness: process up) and/readyz(readiness: DB reachable, migrations applied, converters reachable, backup freshness < 26 h).collabexposes/healthz(process + DB).webserves a static/healthz. - Docker healthchecks on every service (compose
healthcheck:), sodocker compose psand restarts reflect real state;restart: unless-stoppedeverywhere. - External uptime monitoring: the operator's existing Uptime-Kuma
monitors
https://dorfteich.online/healthz(web),/api/v1/healthz, and a WebSocket check on/collab, with notification on failure. Test/Int get web-check-only monitors (no paging). - Backup alerting: the backup sidecar writes
backups/status.jsonafter every run;readyzdegrades when the last successful backup is older than 26 h, which surfaces through Uptime-Kuma without extra tooling. Additionally the sidecar sends a failure e-mail via the instance SMTP.
Logging
- All services log structured JSON to stdout (pino); Docker's json-file
driver with rotation (
max-size: 10m,max-file: 5). - Log content rules: request logs with method/route/status/duration/user id (no request bodies), auth events (login success/failure, permission denials), admin actions (grants, plugin installs, quota changes) as an audit trail, collab session open/close. Never log passwords, tokens, session ids, or page content.
- Reading logs =
docker compose logs/docker logson the host; no central log stack at this scale (revisit if a second Prod host appears).
Backup & restore (operational view of ADR 0015)
- Nightly at 03:00 stage-local time (sidecar
backup, issue #83; envBACKUP_TIME/TZ):pg_dump -Fc→ uploads/plugins volume archive (one tar, same backup idYYYYMMDD-HHMMSS) → prune (BACKUP_RETENTION_DAYS, 30 default / 7 Test+Int; the newest complete set always survives) →status.jsonon thebackupsvolume → on failure a mail directly via the instance SMTP toBACKUP_MAIL_TO. Mirror to BASEL is issue #84. - On-demand run:
docker compose run --rm -e BACKUP_RUN_ONCE=1 backup(exit code = outcome); list sets withdocker compose exec backup ls /backups. - Restore runbook (also the Prod-relocation procedure) — automated by
deploy/backup/restore.sh <backup-id>, run from the stage directory:- stop the app services (
web,api,collab; the db stays up), - restore DB:
pg_restore --clean --if-existsinto thedbcontainer, - restore volume: unpack the matching uploads/plugins archive,
docker compose up -d, verify/readyz, spot-check a page + a file.
- stop the app services (
- Drills: monthly automated restore of the latest Prod dump into a scratch database on Test with a row-count sanity report; quarterly manual full-runbook drill on Test.
- Admin UI: Site Admin can download the latest dump/archive and trigger an on-demand backup run (ADR 0015 — restore stays CLI-only).
Maintenance jobs (in-app scheduler, jobs table)
| Job | Cadence | Purpose |
|---|---|---|
| trash purge | daily | delete pages/ponds past trash retention (ADR 0013) |
| version thinning | daily | auto-version retention policy (ADR 0013) |
| update-log compaction | hourly, idle pages only | bound Yjs log growth |
| quota reconciliation | nightly | recompute pond_usage, report drift |
| orphan file sweep | nightly | volume ↔ DB consistency (ADR 0011) |
| mail outbox retry | every minute | e-mail delivery with backoff |
Job outcomes are visible in the Site Admin UI (last run, status) — that panel is the operator's single glance for instance health.
Update strategy
- Own stages: pipeline-driven (see
deployment.md); Prod only via approved release tags. - Self-hosters: semver releases;
docker compose pull && up -d; migrations auto-apply; release notes flag manual steps andmigrationlabel. Supported downgrade window: one minor release. - Base image / dependency hygiene: monthly dependency-update story (renovate-style batch PR); security advisories for pinned images tracked in the release checklist.
Capacity & limits (initial values, instance-tunable)
| Limit | Default |
|---|---|
| max page document size | 5 MiB Yjs state |
| max upload size | 25 MiB (quota ladder, ADR 0011) |
| collab connections per instance | 500 concurrent |
| rate limits | login 10/min/IP, signup 5/h/IP, API 100/min/user |