Deletion now actually deletes: a trashed pond past the trash retention (same clock as pages, extended trash-purge job) or purged manually via DELETE /ponds/:id/purge (Site-Admin-only, like pond restore) is removed with everything it holds. Files go first (idempotent rm, resumable on a crash), then one transaction ordered around the FK actions: attachments and labels (Restrict) precede the pond; the page delete cascades versions, comments, content cache incl. the search vector, update log, mentions, label assignments, favorites, outgoing links and open collab sessions; the pond delete cascades grants, usage counters (that is the quota correction), pond-plugin opt-ins and conversion jobs; polymorphic watches and pond quota overrides are deleted explicitly. A purge racing a restore or another purge is a no-op; both paths record a pond.purged audit event. Known residues by design, documented in operations.md: target_slug in other ponds' page links (#235) and backups within their retention. Refs #193 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0168Ph5uBmHm8X28CSVpbpnJ
8.5 KiB
Operations concept
Pragmatic monitoring (kickoff decision): health checks, external uptime monitoring, structured logs, backup alerting — no dedicated metrics stack.
Health & monitoring
- Health endpoints:
apiexposes/healthz(liveness: process up) and/readyz(readiness: DB reachable + migrations applied = hard failures → 503; converter/renderer reachability and backup freshness < 26 h are warning-level → overallstatus: degraded, still HTTP 200).collabexposes/healthz(process + DB).webserves a static/healthz. - Docker healthchecks use the liveness endpoints only — never
readyz— so a degraded instance is not restart-looped (restart: unless-stoppedrestarts on process exit, anunhealthymark just shows indocker compose ps). - External uptime monitoring: the operator's existing Uptime-Kuma runs
the monitor set defined in
deploy/monitoring.md(web healthz, api readyz for down, a keyword monitor on the readyz body for degraded, and a collab WebSocket check), with notification on failure for Prod. Test/Int get reduced monitors (no paging). - Backup alerting: the backup sidecar writes
backups/status.jsonafter every run;readyzreads it (issue #85) and degrades when the last successful backup is older than 26 h, which surfaces through Uptime-Kuma without extra tooling. Additionally the sidecar sends a failure e-mail via the instance SMTP.
Logging
- All services log structured JSON to stdout (pino); Docker's json-file
driver with rotation (
max-size: 10m,max-file: 5). - Log content rules: request logs with method/route/status/duration/user id (no request bodies), auth events (login success/failure, permission denials), admin actions (grants, plugin installs, quota changes) as an audit trail, collab session open/close. Never log passwords, tokens, session ids, or page content.
- Persistent audit trail (issue #86): the auth events and admin actions
additionally land as rows in
audit_logviaAuditService— queryable in the Site-Admin System panel (filter by actor/action/time, paginated). Content activity (pages, files, exports, labels) stays log-only by design. - Reading logs =
docker compose logs/docker logson the host; no central log stack at this scale (revisit if a second Prod host appears).
Backup & restore (operational view of ADR 0015)
- Nightly at 03:00 stage-local time (sidecar
backup, issue #83; envBACKUP_TIME/TZ):pg_dump -Fc→ uploads/plugins volume archive (one tar, same backup idYYYYMMDD-HHMMSS) → optional Nextcloud upload (issue #103, see below) → prune (BACKUP_RETENTION_DAYS, 30 default / 7 Test+Int; a Site-Admin setting overrides the env; the newest complete set always survives) →status.jsonon thebackupsvolume → on failure a mail directly via the instance SMTP toBACKUP_MAIL_TO. - Private mirror (issue #84): with
BACKUP_MIRROR_TARGET+BACKUP_MIRROR_SSH_KEYset (Prod: BASEL over WireGuard), every successful run rsyncs the set files after the prune (--deletealigns the remote retention); outcome instatus.json → mirror, failures mail like run failures. Setup:deploy/backup-basel.md. - Off-host copies (issue #103): with a Nextcloud target configured in
the admin UI (WebDAV base URL + username + folder in instance settings,
app password in the secret store), each successful set is bundled into
ONE
dorfteich-backup-<id>.tar.gz(dump + files archive + manifest) and uploaded per schedule (daily/weekly/manual); remote prune mirrors the local guarantee. The api and the sidecar talk over pg NOTIFY/LISTEN (backup_command); upload failures mail like run failures, staleness shows up as thebackup_remotereadyz check. - Target policy (issue #192, ADR 0026):
BACKUP_ALLOWED_TARGETSis a deploy-level, comma-separated allowlist of permissible destination hosts, enforced by the api (settings writes, connection test, admin view) and by the sidecar at the point of egress (WebDAV upload and rsync mirror alike). Empty — the default — disables every remote target: backups stay local only, which is the VS-NfD reference configuration. Existing deployments with a remote target must list its host, or uploads and mirror stop (called out in the release notes). The admin UI shows "unavailable by policy" as distinct from "not configured". Deliberately not a runtime setting: a compromised Site-Admin account cannot widen it. - On-demand run:
docker compose run --rm -e BACKUP_RUN_ONCE=1 backup(exit code = outcome); list sets withdocker compose exec backup ls /backups. - Restore runbook (also the Prod-relocation procedure) — automated by
deploy/backup/restore.sh <backup-id>, run from the stage directory:- stop the app services (
web,api,collab; the db stays up), - restore DB:
pg_restore --clean --if-existsinto thedbcontainer, - restore volume: unpack the matching uploads/plugins archive,
docker compose up -d, verify/readyz, spot-check a page + a file.
- stop the app services (
- Drills (issue #87): monthly automated restore of the latest backup
set into a scratch environment via
.gitea/workflows/drill.yml→deploy/backup/drill.sh(sanity checks + report on the pinned "Restore drills" issue); manual procedure and relocation notes indocs/operations/restore-runbook.md. Quarterly manual full-runbook drill on Test. - Admin UI (issue #103): Site Admin triggers on-demand backups
("Back up now") and restores a local or Nextcloud set in-app — a
type-to-confirm prompt, then the sidecar orchestrates: maintenance mode
(global 503 with an exempt status endpoint), collab sessions closed via
the
backup_maintenanceNOTIFY channel, connections terminated,pg_restore+ volume extract, api restart.restore.shremains the disaster fallback when the app itself is gone.
Maintenance jobs (in-app scheduler, jobs table)
| Job | Cadence | Purpose |
|---|---|---|
| trash purge | daily | delete pages/ponds past trash retention (ADR 0013) |
| version thinning | daily | auto-version retention policy (ADR 0013) |
| update-log compaction | hourly, idle pages only | bound Yjs log growth |
| quota reconciliation | nightly | recompute pond_usage, report drift |
| orphan file sweep | nightly | volume ↔ DB consistency (ADR 0011) |
| mail outbox retry | every minute | e-mail delivery with backoff |
Job outcomes are visible in the Site Admin UI (last run, status) — that panel is the operator's single glance for instance health.
Pond purge (issue #193): a trashed pond past the trash retention is
removed with everything it holds — pages (cascading versions, comments,
content cache incl. the search vector, update log, mentions, label
assignments, favorites, outgoing links, open collab sessions),
attachments (rows and files on disk), labels, grants, usage counters,
watches, pond quota overrides, pond-plugin opt-ins, and conversion jobs.
A Site Admin can purge a trashed pond immediately
(DELETE /ponds/:id/purge); both paths record a pond.purged audit
event. Known residues by design: page_links.target_slug in OTHER
ponds' pages (#235) and backups within their retention.
Update strategy
- Own stages: pipeline-driven (see
deployment.md); Prod only via approved release tags. - Self-hosters: semver releases;
docker compose pull && up -d; migrations auto-apply; release notes flag manual steps andmigrationlabel. Supported downgrade window: one minor release. - Base image / dependency hygiene: monthly dependency-update story (renovate-style batch PR); security advisories for pinned images tracked in the release checklist.
Capacity & limits (initial values, instance-tunable)
| Limit | Default |
|---|---|
| max page document size | 5 MiB Yjs state |
| max upload size | 25 MiB (quota ladder, ADR 0011) |
| collab connections per instance | 500 concurrent |
| rate limits | login 10/min/IP, signup 5/h/IP, API 100/min/user |