dorfteich/docs/architecture/operations.md
Claude Fable 5 0bc36aa58c
Some checks failed
CI / Lint, typecheck, test (pull_request) Successful in 5m1s
CI / Build container images (pull_request) Successful in 2m48s
CI / Auth e2e pack (pull_request) Failing after 3m12s
CI / Import/export fidelity gate (pull_request) Has been skipped
#194: orphan-file sweep, drop the unused Attachment.deletedAt
Nightly sweep with two directions: attachments still unclaimed (pageId
null) after a 24 h grace period - claimed by no collab persist, page
upload, or import - are reclaimed (row, file, quota released); files on
the uploads volume without a database row (drift after a crashed
upload) are removed once older than the grace period. The grace period
protects the paste-then-insert window.

Deliberate deviation from the issue's content-reference idea, documented
in schema comment and operations.md: claimed attachments whose page
content no longer embeds them are NOT auto-deleted. The page attachments
panel lists claimed files as user-managed objects (inserting into the
document is optional there), so 'not embedded' is not 'unused' - an
auto-delete would destroy panel assets. Humans clean those up in the
panel or the pond file manager, which flags orphans already.

Attachment.deletedAt is removed by migration - deletion is hard
everywhere (sweep, purge, manual), there is no soft-delete state; the
never-true deletedAt:null filters in files/export queries went with it.

Refs #194

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0168Ph5uBmHm8X28CSVpbpnJ
2026-07-30 13:19:58 +02:00

9.3 KiB

Operations concept

Pragmatic monitoring (kickoff decision): health checks, external uptime monitoring, structured logs, backup alerting — no dedicated metrics stack.

Health & monitoring

  • Health endpoints: api exposes /healthz (liveness: process up) and /readyz (readiness: DB reachable + migrations applied = hard failures → 503; converter/renderer reachability and backup freshness < 26 h are warning-level → overall status: degraded, still HTTP 200). collab exposes /healthz (process + DB). web serves a static /healthz.
  • Docker healthchecks use the liveness endpoints only — never readyz — so a degraded instance is not restart-looped (restart: unless-stopped restarts on process exit, an unhealthy mark just shows in docker compose ps).
  • External uptime monitoring: the operator's existing Uptime-Kuma runs the monitor set defined in deploy/monitoring.md (web healthz, api readyz for down, a keyword monitor on the readyz body for degraded, and a collab WebSocket check), with notification on failure for Prod. Test/Int get reduced monitors (no paging).
  • Backup alerting: the backup sidecar writes backups/status.json after every run; readyz reads it (issue #85) and degrades when the last successful backup is older than 26 h, which surfaces through Uptime-Kuma without extra tooling. Additionally the sidecar sends a failure e-mail via the instance SMTP.

Logging

  • All services log structured JSON to stdout (pino); Docker's json-file driver with rotation (max-size: 10m, max-file: 5).
  • Log content rules: request logs with method/route/status/duration/user id (no request bodies), auth events (login success/failure, permission denials), admin actions (grants, plugin installs, quota changes) as an audit trail, collab session open/close. Never log passwords, tokens, session ids, or page content.
  • Persistent audit trail (issue #86): the auth events and admin actions additionally land as rows in audit_log via AuditService — queryable in the Site-Admin System panel (filter by actor/action/time, paginated). Content activity (pages, files, exports, labels) stays log-only by design.
  • Reading logs = docker compose logs / docker logs on the host; no central log stack at this scale (revisit if a second Prod host appears).

Backup & restore (operational view of ADR 0015)

  • Nightly at 03:00 stage-local time (sidecar backup, issue #83; env BACKUP_TIME/TZ): pg_dump -Fc → uploads/plugins volume archive (one tar, same backup id YYYYMMDD-HHMMSS) → optional Nextcloud upload (issue #103, see below) → prune (BACKUP_RETENTION_DAYS, 30 default / 7 Test+Int; a Site-Admin setting overrides the env; the newest complete set always survives) → status.json on the backups volume → on failure a mail directly via the instance SMTP to BACKUP_MAIL_TO.
  • Private mirror (issue #84): with BACKUP_MIRROR_TARGET + BACKUP_MIRROR_SSH_KEY set (Prod: BASEL over WireGuard), every successful run rsyncs the set files after the prune (--delete aligns the remote retention); outcome in status.json → mirror, failures mail like run failures. Setup: deploy/backup-basel.md.
  • Off-host copies (issue #103): with a Nextcloud target configured in the admin UI (WebDAV base URL + username + folder in instance settings, app password in the secret store), each successful set is bundled into ONE dorfteich-backup-<id>.tar.gz (dump + files archive + manifest) and uploaded per schedule (daily/weekly/manual); remote prune mirrors the local guarantee. The api and the sidecar talk over pg NOTIFY/LISTEN (backup_command); upload failures mail like run failures, staleness shows up as the backup_remote readyz check.
  • Target policy (issue #192, ADR 0026): BACKUP_ALLOWED_TARGETS is a deploy-level, comma-separated allowlist of permissible destination hosts, enforced by the api (settings writes, connection test, admin view) and by the sidecar at the point of egress (WebDAV upload and rsync mirror alike). Empty — the default — disables every remote target: backups stay local only, which is the VS-NfD reference configuration. Existing deployments with a remote target must list its host, or uploads and mirror stop (called out in the release notes). The admin UI shows "unavailable by policy" as distinct from "not configured". Deliberately not a runtime setting: a compromised Site-Admin account cannot widen it.
  • On-demand run: docker compose run --rm -e BACKUP_RUN_ONCE=1 backup (exit code = outcome); list sets with docker compose exec backup ls /backups.
  • Restore runbook (also the Prod-relocation procedure) — automated by deploy/backup/restore.sh <backup-id>, run from the stage directory:
    1. stop the app services (web, api, collab; the db stays up),
    2. restore DB: pg_restore --clean --if-exists into the db container,
    3. restore volume: unpack the matching uploads/plugins archive,
    4. docker compose up -d, verify /readyz, spot-check a page + a file.
  • Drills (issue #87): monthly automated restore of the latest backup set into a scratch environment via .gitea/workflows/drill.ymldeploy/backup/drill.sh (sanity checks + report on the pinned "Restore drills" issue); manual procedure and relocation notes in docs/operations/restore-runbook.md. Quarterly manual full-runbook drill on Test.
  • Admin UI (issue #103): Site Admin triggers on-demand backups ("Back up now") and restores a local or Nextcloud set in-app — a type-to-confirm prompt, then the sidecar orchestrates: maintenance mode (global 503 with an exempt status endpoint), collab sessions closed via the backup_maintenance NOTIFY channel, connections terminated, pg_restore + volume extract, api restart. restore.sh remains the disaster fallback when the app itself is gone.

Maintenance jobs (in-app scheduler, jobs table)

Job Cadence Purpose
trash purge daily delete pages/ponds past trash retention (ADR 0013)
version thinning daily auto-version retention policy (ADR 0013)
update-log compaction hourly, idle pages only bound Yjs log growth
quota reconciliation nightly recompute pond_usage, report drift
orphan file sweep nightly volume ↔ DB consistency (ADR 0011)
mail outbox retry every minute e-mail delivery with backoff

Job outcomes are visible in the Site Admin UI (last run, status) — that panel is the operator's single glance for instance health.

Orphan file sweep (issue #194): nightly, two directions. Attachments still unclaimed (pageId null) after a 24 h grace period — claimed by no collab persist, page upload, or import — are reclaimed (row, file, quota); the grace period protects the paste-then-insert window. Files on the uploads volume without a database row (drift after a crashed upload) are removed once older than the grace period. Claimed attachments whose page content no longer embeds them are deliberately NOT auto-deleted: the page attachments panel lists them as user-managed objects (insert is optional there), so cleanup of those is a human decision in the panel or the pond file manager. Attachment deletion is hard everywhere — the unused deleted_at column was removed with #194.

Pond purge (issue #193): a trashed pond past the trash retention is removed with everything it holds — pages (cascading versions, comments, content cache incl. the search vector, update log, mentions, label assignments, favorites, outgoing links, open collab sessions), attachments (rows and files on disk), labels, grants, usage counters, watches, pond quota overrides, pond-plugin opt-ins, and conversion jobs. A Site Admin can purge a trashed pond immediately (DELETE /ponds/:id/purge); both paths record a pond.purged audit event. Known residues by design: page_links.target_slug in OTHER ponds' pages (#235) and backups within their retention.

Update strategy

  • Own stages: pipeline-driven (see deployment.md); Prod only via approved release tags.
  • Self-hosters: semver releases; docker compose pull && up -d; migrations auto-apply; release notes flag manual steps and migration label. Supported downgrade window: one minor release.
  • Base image / dependency hygiene: monthly dependency-update story (renovate-style batch PR); security advisories for pinned images tracked in the release checklist.

Capacity & limits (initial values, instance-tunable)

Limit Default
max page document size 5 MiB Yjs state
max upload size 25 MiB (quota ladder, ADR 0011)
collab connections per instance 500 concurrent
rate limits login 10/min/IP, signup 5/h/IP, API 100/min/user