All checks were successful
CI / Lint, typecheck, test (pull_request) Successful in 4m52s
CI / Build container images (pull_request) Successful in 3m54s
CI / Auth e2e pack (pull_request) Successful in 8m4s
CI / Import/export fidelity gate (pull_request) Successful in 56s
CD / Build and push images (push) Successful in 19s
CD / Deploy to Test (push) Successful in 13s
CD / Smoke tests against Test (push) Successful in 1m14s
CD / Promote to Int (push) Successful in 11s
CI / Lint, typecheck, test (push) Successful in 5m0s
CI / Build container images (push) Has been skipped
CI / Auth e2e pack (push) Successful in 7m41s
CI / Import/export fidelity gate (push) Successful in 56s
BACKUP_ALLOWED_TARGETS (comma-separated destination hosts) constrains where backups may go, enforced twice: the api rejects settings writes and connection tests towards non-allowlisted hosts with admin-visible error codes and resolves a non-allowlisted configured target to null, and the sidecar enforces the same policy at the point of egress for the WebDAV upload and the rsync mirror alike (shared policy helpers in packages/shared/src/backup-target-policy.ts). BREAKING: the empty default disables every remote target - backups stay local only, the VS-NfD reference configuration (ADR 0026). Existing deployments with a remote target must list its host or uploads and mirror stop. The admin UI distinguishes unavailable-by-policy from unconfigured (i18n de+en) and shows the permitted hosts. Refs #192 (ADR 0026) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0168Ph5uBmHm8X28CSVpbpnJ
132 lines
7.9 KiB
Markdown
132 lines
7.9 KiB
Markdown
# Operations concept
|
|
|
|
Pragmatic monitoring (kickoff decision): health checks, external uptime
|
|
monitoring, structured logs, backup alerting — no dedicated metrics stack.
|
|
|
|
## Health & monitoring
|
|
|
|
- **Health endpoints**: `api` exposes `/healthz` (liveness: process up) and
|
|
`/readyz` (readiness: DB reachable + migrations applied = hard failures
|
|
→ 503; converter/renderer reachability and backup freshness < 26 h are
|
|
warning-level → overall `status: degraded`, still HTTP 200). `collab`
|
|
exposes `/healthz` (process + DB). `web` serves a static `/healthz`.
|
|
- **Docker healthchecks** use the liveness endpoints only — never `readyz` —
|
|
so a degraded instance is not restart-looped (`restart: unless-stopped`
|
|
restarts on process exit, an `unhealthy` mark just shows in
|
|
`docker compose ps`).
|
|
- **External uptime monitoring**: the operator's existing Uptime-Kuma runs
|
|
the monitor set defined in `deploy/monitoring.md` (web healthz, api
|
|
readyz for down, a keyword monitor on the readyz body for degraded, and
|
|
a collab WebSocket check), with notification on failure for Prod.
|
|
Test/Int get reduced monitors (no paging).
|
|
- **Backup alerting**: the backup sidecar writes
|
|
`backups/status.json` after every run; `readyz` reads it (issue #85) and
|
|
degrades when the last successful backup is older than 26 h, which
|
|
surfaces through Uptime-Kuma without extra tooling. Additionally the
|
|
sidecar sends a failure e-mail via the instance SMTP.
|
|
|
|
## Logging
|
|
|
|
- All services log **structured JSON to stdout** (pino); Docker's json-file
|
|
driver with rotation (`max-size: 10m`, `max-file: 5`).
|
|
- Log content rules: request logs with method/route/status/duration/user id
|
|
(no request bodies), auth events (login success/failure, permission
|
|
denials), admin actions (grants, plugin installs, quota changes) as an
|
|
**audit trail**, collab session open/close. Never log passwords, tokens,
|
|
session ids, or page content.
|
|
- **Persistent audit trail** (issue #86): the auth events and admin actions
|
|
additionally land as rows in `audit_log` via `AuditService` — queryable in
|
|
the Site-Admin System panel (filter by actor/action/time, paginated).
|
|
Content activity (pages, files, exports, labels) stays log-only by design.
|
|
- Reading logs = `docker compose logs` / `docker logs` on the host; no
|
|
central log stack at this scale (revisit if a second Prod host appears).
|
|
|
|
## Backup & restore (operational view of ADR 0015)
|
|
|
|
- Nightly at 03:00 stage-local time (sidecar `backup`, issue #83; env
|
|
`BACKUP_TIME`/`TZ`): `pg_dump -Fc` → uploads/plugins volume archive (one
|
|
tar, same backup id `YYYYMMDD-HHMMSS`) → optional Nextcloud upload
|
|
(issue #103, see below) → prune (`BACKUP_RETENTION_DAYS`, 30 default /
|
|
7 Test+Int; a Site-Admin setting overrides the env; the newest complete
|
|
set always survives) → `status.json` on the `backups` volume → on
|
|
failure a mail directly via the instance SMTP to `BACKUP_MAIL_TO`.
|
|
- **Private mirror** (issue #84): with `BACKUP_MIRROR_TARGET` +
|
|
`BACKUP_MIRROR_SSH_KEY` set (Prod: BASEL over WireGuard), every
|
|
successful run rsyncs the set files after the prune (`--delete` aligns
|
|
the remote retention); outcome in `status.json → mirror`, failures mail
|
|
like run failures. Setup: `deploy/backup-basel.md`.
|
|
- **Off-host copies** (issue #103): with a Nextcloud target configured in
|
|
the admin UI (WebDAV base URL + username + folder in instance settings,
|
|
app password in the secret store), each successful set is bundled into
|
|
ONE `dorfteich-backup-<id>.tar.gz` (dump + files archive + manifest) and
|
|
uploaded per schedule (daily/weekly/manual); remote prune mirrors the
|
|
local guarantee. The api and the sidecar talk over pg NOTIFY/LISTEN
|
|
(`backup_command`); upload failures mail like run failures, staleness
|
|
shows up as the `backup_remote` readyz check.
|
|
- **Target policy** (issue #192, ADR 0026): `BACKUP_ALLOWED_TARGETS` is a
|
|
deploy-level, comma-separated allowlist of permissible destination
|
|
hosts, enforced by the api (settings writes, connection test, admin
|
|
view) and by the sidecar at the point of egress (WebDAV upload and
|
|
rsync mirror alike). **Empty — the default — disables every remote
|
|
target: backups stay local only, which is the VS-NfD reference
|
|
configuration.** Existing deployments with a remote target must list
|
|
its host, or uploads and mirror stop (called out in the release notes).
|
|
The admin UI shows "unavailable by policy" as distinct from
|
|
"not configured". Deliberately not a runtime setting: a compromised
|
|
Site-Admin account cannot widen it.
|
|
- **On-demand run**: `docker compose run --rm -e BACKUP_RUN_ONCE=1 backup`
|
|
(exit code = outcome); list sets with `docker compose exec backup ls /backups`.
|
|
- **Restore runbook** (also the Prod-relocation procedure) — automated by
|
|
`deploy/backup/restore.sh <backup-id>`, run from the stage directory:
|
|
1. stop the app services (`web`, `api`, `collab`; the db stays up),
|
|
2. restore DB: `pg_restore --clean --if-exists` into the `db` container,
|
|
3. restore volume: unpack the matching uploads/plugins archive,
|
|
4. `docker compose up -d`, verify `/readyz`, spot-check a page + a file.
|
|
- **Drills** (issue #87): monthly automated restore of the latest backup
|
|
set into a scratch environment via `.gitea/workflows/drill.yml` →
|
|
`deploy/backup/drill.sh` (sanity checks + report on the pinned "Restore
|
|
drills" issue); manual procedure and relocation notes in
|
|
`docs/operations/restore-runbook.md`. Quarterly manual full-runbook
|
|
drill on Test.
|
|
- **Admin UI** (issue #103): Site Admin triggers on-demand backups
|
|
("Back up now") and restores a local or Nextcloud set in-app — a
|
|
type-to-confirm prompt, then the sidecar orchestrates: maintenance mode
|
|
(global 503 with an exempt status endpoint), collab sessions closed via
|
|
the `backup_maintenance` NOTIFY channel, connections terminated,
|
|
`pg_restore` + volume extract, api restart. `restore.sh` remains the
|
|
disaster fallback when the app itself is gone.
|
|
|
|
## Maintenance jobs (in-app scheduler, `jobs` table)
|
|
|
|
| Job | Cadence | Purpose |
|
|
| --------------------- | ----------------------- | -------------------------------------------------- |
|
|
| trash purge | daily | delete pages/ponds past trash retention (ADR 0013) |
|
|
| version thinning | daily | auto-version retention policy (ADR 0013) |
|
|
| update-log compaction | hourly, idle pages only | bound Yjs log growth |
|
|
| quota reconciliation | nightly | recompute `pond_usage`, report drift |
|
|
| orphan file sweep | nightly | volume ↔ DB consistency (ADR 0011) |
|
|
| mail outbox retry | every minute | e-mail delivery with backoff |
|
|
|
|
Job outcomes are visible in the Site Admin UI (last run, status) — that
|
|
panel is the operator's single glance for instance health.
|
|
|
|
## Update strategy
|
|
|
|
- **Own stages**: pipeline-driven (see `deployment.md`); Prod only via
|
|
approved release tags.
|
|
- **Self-hosters**: semver releases; `docker compose pull && up -d`;
|
|
migrations auto-apply; release notes flag manual steps and `migration`
|
|
label. Supported downgrade window: one minor release.
|
|
- **Base image / dependency hygiene**: monthly dependency-update story
|
|
(renovate-style batch PR); security advisories for pinned images tracked
|
|
in the release checklist.
|
|
|
|
## Capacity & limits (initial values, instance-tunable)
|
|
|
|
| Limit | Default |
|
|
| ------------------------------- | ------------------------------------------------ |
|
|
| max page document size | 5 MiB Yjs state |
|
|
| max upload size | 25 MiB (quota ladder, ADR 0011) |
|
|
| collab connections per instance | 500 concurrent |
|
|
| rate limits | login 10/min/IP, signup 5/h/IP, API 100/min/user |
|