pnpm workspace with apps/web, apps/api, apps/collab, and packages/shared; strict TypeScript base config, repo-wide ESLint (flat) + Prettier, Vitest per package, and root scripts lint/typecheck/test/ build. @dorfteich/shared ships a first health-response helper consumed by apps/api to prove workspace linking. Existing markdown docs are reformatted once by the new Prettier setup. Closes #1 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
87 lines
4.7 KiB
Markdown
87 lines
4.7 KiB
Markdown
# Operations concept
|
|
|
|
Pragmatic monitoring (kickoff decision): health checks, external uptime
|
|
monitoring, structured logs, backup alerting — no dedicated metrics stack.
|
|
|
|
## Health & monitoring
|
|
|
|
- **Health endpoints**: `api` exposes `/healthz` (liveness: process up) and
|
|
`/readyz` (readiness: DB reachable, migrations applied, converters
|
|
reachable, backup freshness < 26 h). `collab` exposes `/healthz`
|
|
(process + DB). `web` serves a static `/healthz`.
|
|
- **Docker healthchecks** on every service (compose `healthcheck:`), so
|
|
`docker compose ps` and restarts reflect real state;
|
|
`restart: unless-stopped` everywhere.
|
|
- **External uptime monitoring**: the operator's existing Uptime-Kuma
|
|
monitors `https://dorfteich.online/healthz` (web), `/api/v1/healthz`, and
|
|
a WebSocket check on `/collab`, with notification on failure. Test/Int
|
|
get web-check-only monitors (no paging).
|
|
- **Backup alerting**: the backup sidecar writes
|
|
`backups/status.json` after every run; `readyz` degrades when the last
|
|
successful backup is older than 26 h, which surfaces through Uptime-Kuma
|
|
without extra tooling. Additionally the sidecar sends a failure e-mail
|
|
via the instance SMTP.
|
|
|
|
## Logging
|
|
|
|
- All services log **structured JSON to stdout** (pino); Docker's json-file
|
|
driver with rotation (`max-size: 10m`, `max-file: 5`).
|
|
- Log content rules: request logs with method/route/status/duration/user id
|
|
(no request bodies), auth events (login success/failure, permission
|
|
denials), admin actions (grants, plugin installs, quota changes) as an
|
|
**audit trail**, collab session open/close. Never log passwords, tokens,
|
|
session ids, or page content.
|
|
- Reading logs = `docker compose logs` / `docker logs` on the host; no
|
|
central log stack at this scale (revisit if a second Prod host appears).
|
|
|
|
## Backup & restore (operational view of ADR 0015)
|
|
|
|
- Nightly at 03:00 stage-local time: `pg_dump -Fc` → uploads/plugins volume
|
|
archive → prune (> 30 days Prod, > 7 days Test/Int) → rsync mirror to
|
|
BASEL `/home/RAID/BACKUPS/dorfteich-prod/` (Prod only, dedicated
|
|
`dorfteich-backup` user).
|
|
- **Restore runbook** (also the Prod-relocation procedure):
|
|
1. `docker compose down` (keep volumes),
|
|
2. restore DB: `pg_restore --clean --if-exists` into the `db` container,
|
|
3. restore volume: unpack the matching uploads archive,
|
|
4. `docker compose up -d`, verify `/readyz`, spot-check a page + a file.
|
|
- **Drills**: monthly automated restore of the latest Prod dump into a
|
|
scratch database on Test with a row-count sanity report; quarterly manual
|
|
full-runbook drill on Test.
|
|
- **Admin UI**: Site Admin can download the latest dump/archive and trigger
|
|
an on-demand backup run (ADR 0015 — restore stays CLI-only).
|
|
|
|
## Maintenance jobs (in-app scheduler, `jobs` table)
|
|
|
|
| Job | Cadence | Purpose |
|
|
| --------------------- | ----------------------- | -------------------------------------------------- |
|
|
| trash purge | daily | delete pages/ponds past trash retention (ADR 0013) |
|
|
| version thinning | daily | auto-version retention policy (ADR 0013) |
|
|
| update-log compaction | hourly, idle pages only | bound Yjs log growth |
|
|
| quota reconciliation | nightly | recompute `pond_usage`, report drift |
|
|
| orphan file sweep | nightly | volume ↔ DB consistency (ADR 0011) |
|
|
| mail outbox retry | every minute | e-mail delivery with backoff |
|
|
|
|
Job outcomes are visible in the Site Admin UI (last run, status) — that
|
|
panel is the operator's single glance for instance health.
|
|
|
|
## Update strategy
|
|
|
|
- **Own stages**: pipeline-driven (see `deployment.md`); Prod only via
|
|
approved release tags.
|
|
- **Self-hosters**: semver releases; `docker compose pull && up -d`;
|
|
migrations auto-apply; release notes flag manual steps and `migration`
|
|
label. Supported downgrade window: one minor release.
|
|
- **Base image / dependency hygiene**: monthly dependency-update story
|
|
(renovate-style batch PR); security advisories for pinned images tracked
|
|
in the release checklist.
|
|
|
|
## Capacity & limits (initial values, instance-tunable)
|
|
|
|
| Limit | Default |
|
|
| ------------------------------- | ------------------------------------------------ |
|
|
| max page document size | 5 MiB Yjs state |
|
|
| max upload size | 25 MiB (quota ladder, ADR 0011) |
|
|
| collab connections per instance | 500 concurrent |
|
|
| rate limits | login 10/min/IP, signup 5/h/IP, API 100/min/user |
|