# Operations concept Pragmatic monitoring (kickoff decision): health checks, external uptime monitoring, structured logs, backup alerting — no dedicated metrics stack. ## Health & monitoring - **Health endpoints**: `api` exposes `/healthz` (liveness: process up) and `/readyz` (readiness: DB reachable + migrations applied = hard failures → 503; converter/renderer reachability and backup freshness < 26 h are warning-level → overall `status: degraded`, still HTTP 200). `collab` exposes `/healthz` (process + DB). `web` serves a static `/healthz`. - **Docker healthchecks** use the liveness endpoints only — never `readyz` — so a degraded instance is not restart-looped (`restart: unless-stopped` restarts on process exit, an `unhealthy` mark just shows in `docker compose ps`). - **External uptime monitoring**: the operator's existing Uptime-Kuma runs the monitor set defined in `deploy/monitoring.md` (web healthz, api readyz for down, a keyword monitor on the readyz body for degraded, and a collab WebSocket check), with notification on failure for Prod. Test/Int get reduced monitors (no paging). - **Backup alerting**: the backup sidecar writes `backups/status.json` after every run; `readyz` reads it (issue #85) and degrades when the last successful backup is older than 26 h, which surfaces through Uptime-Kuma without extra tooling. Additionally the sidecar sends a failure e-mail via the instance SMTP. ## Logging - All services log **structured JSON to stdout** (pino); Docker's json-file driver with rotation (`max-size: 10m`, `max-file: 5`). - Log content rules: request logs with method/route/status/duration/user id (no request bodies), auth events (login success/failure, permission denials), admin actions (grants, plugin installs, quota changes) as an **audit trail**, collab session open/close. Never log passwords, tokens, session ids, or page content. - Reading logs = `docker compose logs` / `docker logs` on the host; no central log stack at this scale (revisit if a second Prod host appears). ## Backup & restore (operational view of ADR 0015) - Nightly at 03:00 stage-local time (sidecar `backup`, issue #83; env `BACKUP_TIME`/`TZ`): `pg_dump -Fc` → uploads/plugins volume archive (one tar, same backup id `YYYYMMDD-HHMMSS`) → prune (`BACKUP_RETENTION_DAYS`, 30 default / 7 Test+Int; the newest complete set always survives) → `status.json` on the `backups` volume → on failure a mail directly via the instance SMTP to `BACKUP_MAIL_TO`. Mirror to BASEL is issue #84. - **On-demand run**: `docker compose run --rm -e BACKUP_RUN_ONCE=1 backup` (exit code = outcome); list sets with `docker compose exec backup ls /backups`. - **Restore runbook** (also the Prod-relocation procedure) — automated by `deploy/backup/restore.sh `, run from the stage directory: 1. stop the app services (`web`, `api`, `collab`; the db stays up), 2. restore DB: `pg_restore --clean --if-exists` into the `db` container, 3. restore volume: unpack the matching uploads/plugins archive, 4. `docker compose up -d`, verify `/readyz`, spot-check a page + a file. - **Drills**: monthly automated restore of the latest Prod dump into a scratch database on Test with a row-count sanity report; quarterly manual full-runbook drill on Test. - **Admin UI**: Site Admin can download the latest dump/archive and trigger an on-demand backup run (ADR 0015 — restore stays CLI-only). ## Maintenance jobs (in-app scheduler, `jobs` table) | Job | Cadence | Purpose | | --------------------- | ----------------------- | -------------------------------------------------- | | trash purge | daily | delete pages/ponds past trash retention (ADR 0013) | | version thinning | daily | auto-version retention policy (ADR 0013) | | update-log compaction | hourly, idle pages only | bound Yjs log growth | | quota reconciliation | nightly | recompute `pond_usage`, report drift | | orphan file sweep | nightly | volume ↔ DB consistency (ADR 0011) | | mail outbox retry | every minute | e-mail delivery with backoff | Job outcomes are visible in the Site Admin UI (last run, status) — that panel is the operator's single glance for instance health. ## Update strategy - **Own stages**: pipeline-driven (see `deployment.md`); Prod only via approved release tags. - **Self-hosters**: semver releases; `docker compose pull && up -d`; migrations auto-apply; release notes flag manual steps and `migration` label. Supported downgrade window: one minor release. - **Base image / dependency hygiene**: monthly dependency-update story (renovate-style batch PR); security advisories for pinned images tracked in the release checklist. ## Capacity & limits (initial values, instance-tunable) | Limit | Default | | ------------------------------- | ------------------------------------------------ | | max page document size | 5 MiB Yjs state | | max upload size | 25 MiB (quota ladder, ADR 0011) | | collab connections per instance | 500 concurrent | | rate limits | login 10/min/IP, signup 5/h/IP, API 100/min/user |