dorfteich/docs/architecture/operations.md
Claude Fable 5 74970f6073
Some checks failed
CI / Lint, typecheck, test (pull_request) Successful in 6m12s
CI / Build container images (pull_request) Successful in 3m4s
CI / Auth e2e pack (pull_request) Successful in 8m35s
CI / Import/export fidelity gate (pull_request) Successful in 1m2s
CI / Import/export fidelity gate (push) Blocked by required conditions
CD / Build and push images (push) Successful in 29s
CD / Deploy to Test (push) Successful in 12s
CD / Smoke tests against Test (push) Successful in 1m35s
CD / Promote to Int (push) Successful in 12s
CI / Lint, typecheck, test (push) Successful in 6m10s
CI / Build container images (push) Has been skipped
CI / Auth e2e pack (push) Has been cancelled
#199: SHA-256 integrity hashes for attachments
Every upload stores the SHA-256 of its bytes, computed from the
in-memory buffer that is written — never by re-reading disk. Every
download re-hashes the stored object BEFORE the first byte leaves
(memory bounded by the max_file_bytes quota that gated the upload) and
fails closed on mismatch with attachment_integrity_failure; the
mismatch lands in the audit trail as file.integrity_failed with both
hashes. Detection of payload manipulation is the one integrity duty
par. 52 VSA leaves with the application — only it knows what the file
should be.

Pre-#199 rows are hashed by a bounded, idempotent backfill that rides
the existing nightly orphan-file-sweep job (no new scheduler job, job
fence untouched); unreadable files are logged and retried, never
silently skipped, and null-hash rows are served unverified only until
the backfill reaches them. Operator runbook note in security.md
(restore from backup, re-download, audit entry carries both hashes).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0168Ph5uBmHm8X28CSVpbpnJ
2026-07-31 04:32:29 +02:00

195 lines
12 KiB
Markdown

# Operations concept
Pragmatic monitoring (kickoff decision): health checks, external uptime
monitoring, structured logs, backup alerting — no dedicated metrics stack.
## Health & monitoring
- **Health endpoints**: `api` exposes `/healthz` (liveness: process up) and
`/readyz` (readiness: DB reachable + migrations applied = hard failures
→ 503; converter/renderer reachability and backup freshness < 26 h are
warning-level overall `status: degraded`, still HTTP 200). `collab`
exposes `/healthz` (process + DB). `web` serves a static `/healthz`.
- **Docker healthchecks** use the liveness endpoints only never `readyz`
so a degraded instance is not restart-looped (`restart: unless-stopped`
restarts on process exit, an `unhealthy` mark just shows in
`docker compose ps`).
- **External uptime monitoring**: the operator's existing Uptime-Kuma runs
the monitor set defined in `deploy/monitoring.md` (web healthz, api
readyz for down, a keyword monitor on the readyz body for degraded, and
a collab WebSocket check), with notification on failure for Prod.
Test/Int get reduced monitors (no paging).
- **Backup alerting**: the backup sidecar writes
`backups/status.json` after every run; `readyz` reads it (issue #85) and
degrades when the last successful backup is older than 26 h, which
surfaces through Uptime-Kuma without extra tooling. Additionally the
sidecar sends a failure e-mail via the instance SMTP.
## Logging
- All services log **structured JSON to stdout** (pino); Docker's json-file
driver with rotation (`max-size: 10m`, `max-file: 5`).
- Log content rules: request logs with method/route/status/duration/user id
(no request bodies), auth events (login success/failure, permission
denials), admin actions (grants, plugin installs, quota changes) as an
**audit trail**, collab session open/close. Never log passwords, tokens,
session ids, or page content.
- **Persistent audit trail** (issue #86): the auth events and admin actions
additionally land as rows in `audit_log` via `AuditService` queryable in
the Site-Admin System panel (filter by actor/action/time, paginated).
Content activity (pages, files, exports, labels) stays log-only by design.
- Reading logs = `docker compose logs` / `docker logs` on the host; no
central log stack at this scale (revisit if a second Prod host appears).
## Backup & restore (operational view of ADR 0015)
- Nightly at 03:00 stage-local time (sidecar `backup`, issue #83; env
`BACKUP_TIME`/`TZ`): `pg_dump -Fc` uploads/plugins volume archive (one
tar, same backup id `YYYYMMDD-HHMMSS`) optional Nextcloud upload
(issue #103, see below) prune (`BACKUP_RETENTION_DAYS`, 30 default /
7 Test+Int; a Site-Admin setting overrides the env; the newest complete
set always survives) `status.json` on the `backups` volume on
failure a mail directly via the instance SMTP to `BACKUP_MAIL_TO`.
- **Private mirror** (issue #84): with `BACKUP_MIRROR_TARGET` +
`BACKUP_MIRROR_SSH_KEY` set (Prod: BASEL over WireGuard), every
successful run rsyncs the set files after the prune (`--delete` aligns
the remote retention); outcome in `status.json → mirror`, failures mail
like run failures. Setup: `deploy/backup-basel.md`.
- **Off-host copies** (issue #103): with a Nextcloud target configured in
the admin UI (WebDAV base URL + username + folder in instance settings,
app password in the secret store), each successful set is bundled into
ONE `dorfteich-backup-<id>.tar.gz` (dump + files archive + manifest) and
uploaded per schedule (daily/weekly/manual); remote prune mirrors the
local guarantee. The api and the sidecar talk over pg NOTIFY/LISTEN
(`backup_command`); upload failures mail like run failures, staleness
shows up as the `backup_remote` readyz check.
- **Target policy** (issue #192, ADR 0026): `BACKUP_ALLOWED_TARGETS` is a
deploy-level, comma-separated allowlist of permissible destination
hosts, enforced by the api (settings writes, connection test, admin
view) and by the sidecar at the point of egress (WebDAV upload and
rsync mirror alike). **Empty the default disables every remote
target: backups stay local only, which is the VS-NfD reference
configuration.** Existing deployments with a remote target must list
its host, or uploads and mirror stop (called out in the release notes).
The admin UI shows "unavailable by policy" as distinct from
"not configured". Deliberately not a runtime setting: a compromised
Site-Admin account cannot widen it.
- **On-demand run**: `docker compose run --rm -e BACKUP_RUN_ONCE=1 backup`
(exit code = outcome); list sets with `docker compose exec backup ls /backups`.
- **Restore runbook** (also the Prod-relocation procedure) automated by
`deploy/backup/restore.sh <backup-id>`, run from the stage directory:
1. stop the app services (`web`, `api`, `collab`; the db stays up),
2. restore DB: `pg_restore --clean --if-exists` into the `db` container,
3. restore volume: unpack the matching uploads/plugins archive,
4. `docker compose up -d`, verify `/readyz`, spot-check a page + a file.
- **Drills** (issue #87): monthly automated restore of the latest backup
set into a scratch environment via `.gitea/workflows/drill.yml`
`deploy/backup/drill.sh` (sanity checks + report on the pinned "Restore
drills" issue); manual procedure and relocation notes in
`docs/operations/restore-runbook.md`. Quarterly manual full-runbook
drill on Test.
- **Admin UI** (issue #103): Site Admin triggers on-demand backups
("Back up now") and restores a local or Nextcloud set in-app a
type-to-confirm prompt, then the sidecar orchestrates: maintenance mode
(global 503 with an exempt status endpoint), collab sessions closed via
the `backup_maintenance` NOTIFY channel, connections terminated,
`pg_restore` + volume extract, api restart. `restore.sh` remains the
disaster fallback when the app itself is gone.
## Maintenance jobs (in-app scheduler, `jobs` table)
| Job | Cadence | Purpose |
| ------------------------ | ----------------------- | --------------------------------------------------- |
| trash purge | daily | delete pages/ponds past trash retention (ADR 0013) |
| version thinning | daily | auto-version retention policy (ADR 0013) |
| update-log compaction | hourly, idle pages only | bound Yjs log growth |
| quota reconciliation | nightly | recompute `pond_usage`, report drift |
| orphan file sweep | nightly | volume DB consistency (ADR 0011) |
| audit retention | daily | prune `audit_log` past `audit.retentionDays` (#196) |
| conversion payload prune | daily | null finished conversion jobs' raw bytes (#233) |
| mail outbox retention | daily | delete sent/failed outbox rows past period (#234) |
| mail outbox retry | every minute | e-mail delivery with backoff |
Job outcomes are visible in the Site Admin UI (last run, status) that
panel is the operator's single glance for instance health.
Orphan file sweep (issue #194): nightly, two directions. Attachments
still unclaimed (`pageId` null) after a 24 h grace period claimed by
no collab persist, page upload, or import are reclaimed (row, file,
quota); the grace period protects the paste-then-insert window. Files on
the uploads volume without a database row (drift after a crashed upload)
are removed once older than the grace period. Claimed attachments whose
page content no longer embeds them are deliberately NOT auto-deleted:
the page attachments panel lists them as user-managed objects (insert is
optional there), so cleanup of those is a human decision in the panel or
the pond file manager. Attachment deletion is hard everywhere the
unused `deleted_at` column was removed with #194. Since #199 the same
nightly run also backfills SHA-256 integrity hashes for attachments that
predate the column (bounded batch per night, idempotent; unreadable
files are logged and retried, see security.md §Content & upload
security).
Conversion payload prune (issue #233): daily. Nulls the raw `input` and
`result` bytes of import/export conversion jobs that finished (succeeded
or failed) more than `conversion.payloadRetentionDays` (default 30) ago
the bytes are transient, not the durable copy an Attachment is, so a
deleted page cannot live on inside its last export. The row itself
survives for status/audit purposes. Pending or running jobs including
a crashed RUNNING row the worker's stale-lock recovery will re-run
keep their payload. Data-export downloads additionally expire much
earlier through their own link TTL (#68).
Pond purge (issue #193): a trashed pond past the trash retention is
removed with everything it holds pages (cascading versions, comments,
content cache incl. the search vector, update log, mentions, label
assignments, favorites, outgoing links, open collab sessions),
attachments (rows and files on disk), labels, grants, usage counters,
watches, pond quota overrides, pond-plugin opt-ins, and conversion jobs.
A Site Admin can purge a trashed pond immediately
(`DELETE /ponds/:id/purge`); both paths record a `pond.purged` audit
event. Known residues by design: `page_links.target_slug` in OTHER
ponds' pages (#235) and backups within their retention.
Decision on `page_links` rows pointing at a purged page (issue #235,
recorded on #231): they are KEPT, as an accepted residue. The row is
only the index of a wikilink whose text including the slug, which
carries the page title remains visible in the linking page's own
content, governed by that page's permissions and written by an author
who could read the target at the time. Deleting the index row would
remove nothing the system still shows (content, content cache, and the
linking page's search index all keep the text) while breaking the
deliberate phantom-link behaviour: a page recreated under the same slug
resolves those links again. An operator answering for deletion refers
to this paragraph: the residue is the linking author's own document,
not a copy of the purged page.
## Update strategy
- **Own stages**: pipeline-driven (see `deployment.md`); Prod only via
approved release tags.
- **Self-hosters**: semver releases; `docker compose pull && up -d`;
migrations auto-apply; release notes flag manual steps and `migration`
label. Supported downgrade window: one minor release.
- **Base image / dependency hygiene**: monthly dependency-update story
(renovate-style batch PR); security advisories for pinned images tracked
in the release checklist.
- **Toolchain pin (issue #236)**: `.node-version` is the single
authoritative Node version. CI/CD select Node exclusively via
`node-version-file`, every Dockerfile pins `node:<version>-alpine`, and
the `engines.node` floor in `package.json` states the same version
(open-ended upwards a newer local Node keeps working; reproducibility
rests on the images and CI, not the laptop). An early CI step fails on
any drift between those places. Raising Node (e.g. for a security fix):
update `.node-version`, all four Dockerfiles and the engines floor in
ONE commit and let CI confirm. pnpm is pinned the same way via
`packageManager`.
## Capacity & limits (initial values, instance-tunable)
| Limit | Default |
| ------------------------------- | ------------------------------------------------ |
| max page document size | 5 MiB Yjs state |
| max upload size | 25 MiB (quota ladder, ADR 0011) |
| collab connections per instance | 500 concurrent |
| rate limits | login 10/min/IP, signup 5/h/IP, API 100/min/user |