All checks were successful
CD / Build and push images (push) Successful in 3m51s
CI / Lint, typecheck, test (push) Successful in 4m5s
CD / Deploy to Test (push) Successful in 11s
CI / Build container images (push) Has been skipped
CD / Smoke tests against Test (push) Successful in 1m11s
CD / Promote to Int (push) Successful in 12s
CI / Auth e2e pack (push) Successful in 5m52s
CI / Import/export fidelity gate (push) Successful in 47s
The operator-level extra beside the admin-configured Nextcloud target (#103), unblocked now that the ONE→BASEL tunnel is stable again. - sidecar: optional mirror step (mirror.ts) driven purely by env — BACKUP_MIRROR_TARGET (rsync-over-ssh), BACKUP_MIRROR_SSH_KEY (private key on the secrets volume, never in image or repo), BACKUP_MIRROR_SSH_PORT. Runs after the prune of every successful run, so --delete aligns the remote retention with the local one (the newest-complete-set guarantee carries over). Only set files travel (db-*.dump, files-*.tar.gz); status files and bundles stay local. Host key pinned via accept-new into .mirror_known_hosts on the backups volume; fixed remote modes (dirs 750, files 640, symbolic --chmod — octal needs rsync ≥ 3, macOS dev machines ship 2.6.9). rsync + openssh-client added to the sidecar image. - status: additive `mirror` block in status.json (outcome, transferred count, lastSuccessAt carried across failures) — shown on the admin backup card; failures alert via a new backupMirrorFailed mail (de+en) while the local run still counts as succeeded. - deploy/backup-basel.md: complete BASEL-side walkthrough — dedicated user dorfteich-backup with a /home/ home and a bash login shell, explicitly avoiding the Debian backup-user (UID 34) pitfalls (nologin shell rejects rsync sessions, /var/backups home), key placement through the api container onto the secrets volume, .env values, on-demand verification. - tests: rsync-arg/stats-parsing units plus an integration suite against the real rsync binary (local target; skips where rsync is absent) — transfer, idempotent re-run (0 files), retention alignment, failure path carrying lastSuccessAt. Verified live against the real BASEL host from a native sidecar run: initial transfer, host-key pinning, retention alignment after a local prune, idempotency, and the failure path (surfaced in status.json while the local run stayed green). BASEL side provisioned per the doc. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
7.2 KiB
7.2 KiB
Operations concept
Pragmatic monitoring (kickoff decision): health checks, external uptime monitoring, structured logs, backup alerting — no dedicated metrics stack.
Health & monitoring
- Health endpoints:
apiexposes/healthz(liveness: process up) and/readyz(readiness: DB reachable + migrations applied = hard failures → 503; converter/renderer reachability and backup freshness < 26 h are warning-level → overallstatus: degraded, still HTTP 200).collabexposes/healthz(process + DB).webserves a static/healthz. - Docker healthchecks use the liveness endpoints only — never
readyz— so a degraded instance is not restart-looped (restart: unless-stoppedrestarts on process exit, anunhealthymark just shows indocker compose ps). - External uptime monitoring: the operator's existing Uptime-Kuma runs
the monitor set defined in
deploy/monitoring.md(web healthz, api readyz for down, a keyword monitor on the readyz body for degraded, and a collab WebSocket check), with notification on failure for Prod. Test/Int get reduced monitors (no paging). - Backup alerting: the backup sidecar writes
backups/status.jsonafter every run;readyzreads it (issue #85) and degrades when the last successful backup is older than 26 h, which surfaces through Uptime-Kuma without extra tooling. Additionally the sidecar sends a failure e-mail via the instance SMTP.
Logging
- All services log structured JSON to stdout (pino); Docker's json-file
driver with rotation (
max-size: 10m,max-file: 5). - Log content rules: request logs with method/route/status/duration/user id (no request bodies), auth events (login success/failure, permission denials), admin actions (grants, plugin installs, quota changes) as an audit trail, collab session open/close. Never log passwords, tokens, session ids, or page content.
- Persistent audit trail (issue #86): the auth events and admin actions
additionally land as rows in
audit_logviaAuditService— queryable in the Site-Admin System panel (filter by actor/action/time, paginated). Content activity (pages, files, exports, labels) stays log-only by design. - Reading logs =
docker compose logs/docker logson the host; no central log stack at this scale (revisit if a second Prod host appears).
Backup & restore (operational view of ADR 0015)
- Nightly at 03:00 stage-local time (sidecar
backup, issue #83; envBACKUP_TIME/TZ):pg_dump -Fc→ uploads/plugins volume archive (one tar, same backup idYYYYMMDD-HHMMSS) → optional Nextcloud upload (issue #103, see below) → prune (BACKUP_RETENTION_DAYS, 30 default / 7 Test+Int; a Site-Admin setting overrides the env; the newest complete set always survives) →status.jsonon thebackupsvolume → on failure a mail directly via the instance SMTP toBACKUP_MAIL_TO. - Private mirror (issue #84): with
BACKUP_MIRROR_TARGET+BACKUP_MIRROR_SSH_KEYset (Prod: BASEL over WireGuard), every successful run rsyncs the set files after the prune (--deletealigns the remote retention); outcome instatus.json → mirror, failures mail like run failures. Setup:deploy/backup-basel.md. - Off-host copies (issue #103): with a Nextcloud target configured in
the admin UI (WebDAV base URL + username + folder in instance settings,
app password in the secret store), each successful set is bundled into
ONE
dorfteich-backup-<id>.tar.gz(dump + files archive + manifest) and uploaded per schedule (daily/weekly/manual); remote prune mirrors the local guarantee. The api and the sidecar talk over pg NOTIFY/LISTEN (backup_command); upload failures mail like run failures, staleness shows up as thebackup_remotereadyz check. - On-demand run:
docker compose run --rm -e BACKUP_RUN_ONCE=1 backup(exit code = outcome); list sets withdocker compose exec backup ls /backups. - Restore runbook (also the Prod-relocation procedure) — automated by
deploy/backup/restore.sh <backup-id>, run from the stage directory:- stop the app services (
web,api,collab; the db stays up), - restore DB:
pg_restore --clean --if-existsinto thedbcontainer, - restore volume: unpack the matching uploads/plugins archive,
docker compose up -d, verify/readyz, spot-check a page + a file.
- stop the app services (
- Drills (issue #87): monthly automated restore of the latest backup
set into a scratch environment via
.gitea/workflows/drill.yml→deploy/backup/drill.sh(sanity checks + report on the pinned "Restore drills" issue); manual procedure and relocation notes indocs/operations/restore-runbook.md. Quarterly manual full-runbook drill on Test. - Admin UI (issue #103): Site Admin triggers on-demand backups
("Back up now") and restores a local or Nextcloud set in-app — a
type-to-confirm prompt, then the sidecar orchestrates: maintenance mode
(global 503 with an exempt status endpoint), collab sessions closed via
the
backup_maintenanceNOTIFY channel, connections terminated,pg_restore+ volume extract, api restart.restore.shremains the disaster fallback when the app itself is gone.
Maintenance jobs (in-app scheduler, jobs table)
| Job | Cadence | Purpose |
|---|---|---|
| trash purge | daily | delete pages/ponds past trash retention (ADR 0013) |
| version thinning | daily | auto-version retention policy (ADR 0013) |
| update-log compaction | hourly, idle pages only | bound Yjs log growth |
| quota reconciliation | nightly | recompute pond_usage, report drift |
| orphan file sweep | nightly | volume ↔ DB consistency (ADR 0011) |
| mail outbox retry | every minute | e-mail delivery with backoff |
Job outcomes are visible in the Site Admin UI (last run, status) — that panel is the operator's single glance for instance health.
Update strategy
- Own stages: pipeline-driven (see
deployment.md); Prod only via approved release tags. - Self-hosters: semver releases;
docker compose pull && up -d; migrations auto-apply; release notes flag manual steps andmigrationlabel. Supported downgrade window: one minor release. - Base image / dependency hygiene: monthly dependency-update story (renovate-style batch PR); security advisories for pinned images tracked in the release checklist.
Capacity & limits (initial values, instance-tunable)
| Limit | Default |
|---|---|
| max page document size | 5 MiB Yjs state |
| max upload size | 25 MiB (quota ladder, ADR 0011) |
| collab connections per instance | 500 concurrent |
| rate limits | login 10/min/IP, signup 5/h/IP, API 100/min/user |