Convert read_events to monthly RANGE partitions on occurred_at, with a DEFAULT partition as safety net: a lagging maintenance job must never turn the trail's hard-failure semantics into an outage for classified reads. The dedup unique pair (#223) moves to per-partition indexes (PostgreSQL cannot carry it on the parent); a bucket spanning a month boundary may record one duplicate — over-recording is acceptable, gaps are not. New daily job read-trail-maintenance (job-count fence 9 -> 10) creates months ahead — each with its dedup index — and applies the trail's own retention readTrail.retentionDays (default 365, deliberately independent of audit.retentionDays): whole expired months are DROPped without scanning, remainders deleted by range, every run audited as read_trail.pruned (catalogue v1.2; the fence regex now admits an underscore namespace). Site-Admin query path GET /admin/system/read-events answers "who read page X" and "what did user Y read" within a period — API-only by design, documented. Growth measured and documented in data-model.md: ~1 MB per 1000 events including indexes. Tests: retention pruning + audited deletion + admin queries on the shared database; the partitioned shape, per-partition P2002 dedup, months-ahead creation and DROP-based pruning against a fresh database built by the real migration chain. Refs #224. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AUtYMxwTCMHG9mVHnwbFg8
218 lines
13 KiB
Markdown
218 lines
13 KiB
Markdown
# Operations concept
|
|
|
|
Pragmatic monitoring (kickoff decision): health checks, external uptime
|
|
monitoring, structured logs, backup alerting — no dedicated metrics stack.
|
|
|
|
## Health & monitoring
|
|
|
|
- **Health endpoints**: `api` exposes `/healthz` (liveness: process up) and
|
|
`/readyz` (readiness: DB reachable + migrations applied = hard failures
|
|
→ 503; converter/renderer reachability and backup freshness < 26 h are
|
|
warning-level → overall `status: degraded`, still HTTP 200). `collab`
|
|
exposes `/healthz` (process + DB). `web` serves a static `/healthz`.
|
|
- **Docker healthchecks** use the liveness endpoints only — never `readyz` —
|
|
so a degraded instance is not restart-looped (`restart: unless-stopped`
|
|
restarts on process exit, an `unhealthy` mark just shows in
|
|
`docker compose ps`).
|
|
- **External uptime monitoring**: the operator's existing Uptime-Kuma runs
|
|
the monitor set defined in `deploy/monitoring.md` (web healthz, api
|
|
readyz for down, a keyword monitor on the readyz body for degraded, and
|
|
a collab WebSocket check), with notification on failure for Prod.
|
|
Test/Int get reduced monitors (no paging).
|
|
- **Backup alerting**: the backup sidecar writes
|
|
`backups/status.json` after every run; `readyz` reads it (issue #85) and
|
|
degrades when the last successful backup is older than 26 h, which
|
|
surfaces through Uptime-Kuma without extra tooling. Additionally the
|
|
sidecar sends a failure e-mail via the instance SMTP.
|
|
|
|
## Logging
|
|
|
|
- All services log **structured JSON to stdout** (pino); Docker's json-file
|
|
driver with rotation (`max-size: 10m`, `max-file: 5`).
|
|
- Log content rules: request logs with method/route/status/duration/user id
|
|
(no request bodies), auth events (login success/failure, permission
|
|
denials), admin actions (grants, plugin installs, quota changes) as an
|
|
**audit trail**, collab session open/close. Never log passwords, tokens,
|
|
session ids, or page content.
|
|
- **Persistent audit trail** (issue #86): the auth events and admin actions
|
|
additionally land as rows in `audit_log` via `AuditService` — queryable in
|
|
the Site-Admin System panel (filter by actor/action/time, paginated).
|
|
Content activity (pages, files, exports, labels) stays log-only by design.
|
|
- Reading logs = `docker compose logs` / `docker logs` on the host; no
|
|
central log stack at this scale (revisit if a second Prod host appears).
|
|
|
|
## Backup & restore (operational view of ADR 0015)
|
|
|
|
- Nightly at 03:00 stage-local time (sidecar `backup`, issue #83; env
|
|
`BACKUP_TIME`/`TZ`): `pg_dump -Fc` → uploads/plugins volume archive (one
|
|
tar, same backup id `YYYYMMDD-HHMMSS`) → optional Nextcloud upload
|
|
(issue #103, see below) → prune (`BACKUP_RETENTION_DAYS`, 30 default /
|
|
7 Test+Int; a Site-Admin setting overrides the env; the newest complete
|
|
set always survives) → `status.json` on the `backups` volume → on
|
|
failure a mail directly via the instance SMTP to `BACKUP_MAIL_TO`.
|
|
- **Private mirror** (issue #84): with `BACKUP_MIRROR_TARGET` +
|
|
`BACKUP_MIRROR_SSH_KEY` set (Prod: BASEL over WireGuard), every
|
|
successful run rsyncs the set files after the prune (`--delete` aligns
|
|
the remote retention); outcome in `status.json → mirror`, failures mail
|
|
like run failures. Setup: `deploy/backup-basel.md`.
|
|
- **Off-host copies** (issue #103): with a Nextcloud target configured in
|
|
the admin UI (WebDAV base URL + username + folder in instance settings,
|
|
app password in the secret store), each successful set is bundled into
|
|
ONE `dorfteich-backup-<id>.tar.gz` (dump + files archive + manifest) and
|
|
uploaded per schedule (daily/weekly/manual); remote prune mirrors the
|
|
local guarantee. The api and the sidecar talk over pg NOTIFY/LISTEN
|
|
(`backup_command`); upload failures mail like run failures, staleness
|
|
shows up as the `backup_remote` readyz check.
|
|
- **Target policy** (issue #192, ADR 0026): `BACKUP_ALLOWED_TARGETS` is a
|
|
deploy-level, comma-separated allowlist of permissible destination
|
|
hosts, enforced by the api (settings writes, connection test, admin
|
|
view) and by the sidecar at the point of egress (WebDAV upload and
|
|
rsync mirror alike). **Empty — the default — disables every remote
|
|
target: backups stay local only, which is the VS-NfD reference
|
|
configuration.** Existing deployments with a remote target must list
|
|
its host, or uploads and mirror stop (called out in the release notes).
|
|
The admin UI shows "unavailable by policy" as distinct from
|
|
"not configured". Deliberately not a runtime setting: a compromised
|
|
Site-Admin account cannot widen it.
|
|
- **On-demand run**: `docker compose run --rm -e BACKUP_RUN_ONCE=1 backup`
|
|
(exit code = outcome); list sets with `docker compose exec backup ls /backups`.
|
|
- **Restore runbook** (also the Prod-relocation procedure) — automated by
|
|
`deploy/backup/restore.sh <backup-id>`, run from the stage directory:
|
|
1. stop the app services (`web`, `api`, `collab`; the db stays up),
|
|
2. restore DB: `pg_restore --clean --if-exists` into the `db` container,
|
|
3. restore volume: unpack the matching uploads/plugins archive,
|
|
4. `docker compose up -d`, verify `/readyz`, spot-check a page + a file.
|
|
- **Drills** (issue #87): monthly automated restore of the latest backup
|
|
set into a scratch environment via `.gitea/workflows/drill.yml` →
|
|
`deploy/backup/drill.sh` (sanity checks + report on the pinned "Restore
|
|
drills" issue); manual procedure and relocation notes in
|
|
`docs/operations/restore-runbook.md`. Quarterly manual full-runbook
|
|
drill on Test.
|
|
- **Admin UI** (issue #103): Site Admin triggers on-demand backups
|
|
("Back up now") and restores a local or Nextcloud set in-app — a
|
|
type-to-confirm prompt, then the sidecar orchestrates: maintenance mode
|
|
(global 503 with an exempt status endpoint), collab sessions closed via
|
|
the `backup_maintenance` NOTIFY channel, connections terminated,
|
|
`pg_restore` + volume extract, api restart. `restore.sh` remains the
|
|
disaster fallback when the app itself is gone.
|
|
|
|
## Maintenance jobs (in-app scheduler, `jobs` table)
|
|
|
|
| Job | Cadence | Purpose |
|
|
| ------------------------ | ----------------------- | --------------------------------------------------- |
|
|
| trash purge | daily | delete pages/ponds past trash retention (ADR 0013) |
|
|
| version thinning | daily | auto-version retention policy (ADR 0013) |
|
|
| update-log compaction | hourly, idle pages only | bound Yjs log growth |
|
|
| quota reconciliation | nightly | recompute `pond_usage`, report drift |
|
|
| orphan file sweep | nightly | volume ↔ DB consistency (ADR 0011) |
|
|
| audit retention | daily | prune `audit_log` past `audit.retentionDays` (#196) |
|
|
| conversion payload prune | daily | null finished conversion jobs' raw bytes (#233) |
|
|
| mail outbox retention | daily | delete sent/failed outbox rows past period (#234) |
|
|
| mail outbox retry | every minute | e-mail delivery with backoff |
|
|
| read-trail maintenance | daily | partition months ahead + prune `read_events` (#224) |
|
|
|
|
Job outcomes are visible in the Site Admin UI (last run, status) — that
|
|
panel is the operator's single glance for instance health.
|
|
|
|
Orphan file sweep (issue #194): nightly, two directions. Attachments
|
|
still unclaimed (`pageId` null) after a 24 h grace period — claimed by
|
|
no collab persist, page upload, or import — are reclaimed (row, file,
|
|
quota); the grace period protects the paste-then-insert window. Files on
|
|
the uploads volume without a database row (drift after a crashed upload)
|
|
are removed once older than the grace period. Claimed attachments whose
|
|
page content no longer embeds them are deliberately NOT auto-deleted:
|
|
the page attachments panel lists them as user-managed objects (insert is
|
|
optional there), so cleanup of those is a human decision in the panel or
|
|
the pond file manager. Attachment deletion is hard everywhere — the
|
|
unused `deleted_at` column was removed with #194. Since #199 the same
|
|
nightly run also backfills SHA-256 integrity hashes for attachments that
|
|
predate the column (bounded batch per night, idempotent; unreadable
|
|
files are logged and retried, see security.md §Content & upload
|
|
security).
|
|
|
|
Conversion payload prune (issue #233): daily. Nulls the raw `input` and
|
|
`result` bytes of import/export conversion jobs that finished (succeeded
|
|
or failed) more than `conversion.payloadRetentionDays` (default 30) ago —
|
|
the bytes are transient, not the durable copy an Attachment is, so a
|
|
deleted page cannot live on inside its last export. The row itself
|
|
survives for status/audit purposes. Pending or running jobs — including
|
|
a crashed RUNNING row the worker's stale-lock recovery will re-run —
|
|
keep their payload. Data-export downloads additionally expire much
|
|
earlier through their own link TTL (#68).
|
|
|
|
Pond purge (issue #193): a trashed pond past the trash retention is
|
|
removed with everything it holds — pages (cascading versions, comments,
|
|
content cache incl. the search vector, update log, mentions, label
|
|
assignments, favorites, outgoing links, open collab sessions),
|
|
attachments (rows and files on disk), labels, grants, usage counters,
|
|
watches, pond quota overrides, pond-plugin opt-ins, and conversion jobs.
|
|
A Site Admin can purge a trashed pond immediately
|
|
(`DELETE /ponds/:id/purge`); both paths record a `pond.purged` audit
|
|
event. Known residues by design: `page_links.target_slug` in OTHER
|
|
ponds' pages (#235) and backups within their retention.
|
|
|
|
Decision on `page_links` rows pointing at a purged page (issue #235,
|
|
recorded on #231): they are KEPT, as an accepted residue. The row is
|
|
only the index of a wikilink whose text — including the slug, which
|
|
carries the page title — remains visible in the linking page's own
|
|
content, governed by that page's permissions and written by an author
|
|
who could read the target at the time. Deleting the index row would
|
|
remove nothing the system still shows (content, content cache, and the
|
|
linking page's search index all keep the text) while breaking the
|
|
deliberate phantom-link behaviour: a page recreated under the same slug
|
|
resolves those links again. An operator answering for deletion refers
|
|
to this paragraph: the residue is the linking author's own document,
|
|
not a copy of the purged page.
|
|
|
|
## Update strategy
|
|
|
|
- **Own stages**: pipeline-driven (see `deployment.md`); Prod only via
|
|
approved release tags.
|
|
- **Self-hosters**: semver releases; `docker compose pull && up -d`;
|
|
migrations auto-apply; release notes flag manual steps and `migration`
|
|
label. Supported downgrade window: one minor release.
|
|
- **Base image / dependency hygiene**: monthly dependency-update story
|
|
(renovate-style batch PR); security advisories for pinned images tracked
|
|
in the release checklist.
|
|
- **Toolchain pin (issue #236)**: `.node-version` is the single
|
|
authoritative Node version. CI/CD select Node exclusively via
|
|
`node-version-file`, every Dockerfile pins `node:<version>-alpine`, and
|
|
the `engines.node` floor in `package.json` states the same version
|
|
(open-ended upwards — a newer local Node keeps working; reproducibility
|
|
rests on the images and CI, not the laptop). An early CI step fails on
|
|
any drift between those places. Raising Node (e.g. for a security fix):
|
|
update `.node-version`, all four Dockerfiles and the engines floor in
|
|
ONE commit and let CI confirm. pnpm is pinned the same way via
|
|
`packageManager`.
|
|
|
|
## Capacity & limits (initial values, instance-tunable)
|
|
|
|
| Limit | Default |
|
|
| ------------------------------- | ------------------------------------------------ |
|
|
| max page document size | 5 MiB Yjs state |
|
|
| max upload size | 25 MiB (quota ladder, ADR 0011) |
|
|
| collab connections per instance | 500 concurrent |
|
|
| rate limits | login 10/min/IP, signup 5/h/IP, API 100/min/user |
|
|
|
|
## Classified attachment downloads (issue #212, ADR 0022)
|
|
|
|
An attachment is an opaque binary — the application cannot write the
|
|
VS-NfD marking into arbitrary file formats. The marking therefore lives
|
|
**around** the file:
|
|
|
|
- **Filename prefix `VS-NfD_`** on every download whose effective
|
|
classification is `vs_nfd` (single source:
|
|
`classificationFilenamePrefix()` in `@dorfteich/shared`).
|
|
- **Effective classification**: the linked page's level. An attachment
|
|
whose `pageId` is not (yet) set — paste-then-insert, pond-level files —
|
|
**fails closed** to the highest level of any live page in its pond.
|
|
- **Containing archive**: the pond export ZIP states each media file's
|
|
level in `manifest.json` and adds a sibling
|
|
`<file>.classification.txt` companion carrying the full marking for
|
|
classified media.
|
|
|
|
Residual risk, deliberately documented rather than hidden (recorded on
|
|
issue #231): the file's own **content** carries no marking — a user who
|
|
renames the file has an unmarked classified binary. Marking file contents
|
|
would require rewriting arbitrary formats, which ADR 0019 rules out.
|