dorfteich/deploy/monitoring.md
Claude Fable 5 5cef359b8f
All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m45s
CD / Build and push images (push) Successful in 3m49s
CI / Build container images (push) Has been skipped
CD / Deploy to Test (push) Successful in 9s
CD / Smoke tests against Test (push) Successful in 1m18s
CD / Promote to Int (push) Successful in 11s
CI / Auth e2e pack (push) Successful in 5m35s
CI / Import/export fidelity gate (push) Successful in 47s
Nextcloud backup target: admin-configured, manual + scheduled uploads, in-app restore (#103)
Off-host backups for every self-hoster, configured entirely in the admin
UI — supersedes the host-specific mirror plan behind #84.

shared:
- webdav.ts (new package entry like token-crypto): minimal WebDAV client
  with basic auth — PROPFIND (tolerant multistatus parser), MKCOL, PUT
  (streamed), GET, DELETE; Nextcloud DAV path derived from the plain
  server URL, explicit DAV bases pass through
- backup-status.ts: additive remote-upload status in status.json, the
  restore-status.json contract (running/succeeded/failed + staleness
  bound), the backup_command/backup_maintenance NOTIFY channels, and the
  one-bundle-per-set naming (dorfteich-backup-<id>.tar.gz)
- backup-set.ts moved here from apps/backup (api lists local sets)

backup sidecar:
- reads the backup.* instance settings directly from the database (admin
  changes apply next run; local retention row overrides the env) and the
  app password from the secret store
- after each successful set: bundle dump + files archive + manifest into
  ONE self-contained tar.gz, upload via WebDAV per schedule
  (off/daily/weekly; manual runs always upload), prune remote bundles —
  never the newest — and record the outcome in status.json; upload
  failures alert via a new backupUploadFailed mail (de+en)
- command listener on backup_command (run / restore) with a serial queue
  against the nightly timer
- restore orchestrator: restore-status.json → maintenance NOTIFY →
  grace → (remote: download + manifest-verify bundle) → terminate other
  DB connections → shared perform-restore path (same code as restore.sh)
  → final status + maintenance exit

api:
- MaintenanceGuard (global, registered before the setup gate): 503
  maintenance_mode while restore-status says running; health endpoints
  and the new public GET /backup/restore-status stay exempt; a stale
  running state (crashed sidecar) unblocks after 30 min
- MaintenanceStateService watches the file and restarts the api after a
  successful restore (fresh caches, migrate-on-start for older dumps);
  main.ts refuses to touch the database while a restore runs — a
  container restarting mid-restore must not race pg_restore with
  migrate deploy
- worker sweeps (conversion, mail outbox, scheduler) catch transient
  database failures instead of dying on an unhandled rejection — the
  restore's connection termination crashed the api in verification
- backup admin endpoints under /admin/system/backup: settings (live
  connection test before save, password write-only into the secret
  store), nextcloud/test, sets (local via the ro backups mount + remote
  via WebDAV), run + restore (type-to-confirm backstop, source
  validation) — commands travel as NOTIFY payloads; audit actions
  backup.settings_changed/run_triggered/restore_requested
- readyz: new warning-level backup_remote check while a target is
  configured (26 h daily / 170 h weekly bound)

collab:
- maintenance listener: on enter, persist + close every live session and
  refuse new connections until exit (failsafe timeout 30 min) — no
  in-memory document may write pre-restore content back afterwards

web:
- Admin → System backup section: status card with remote facts and a
  "Back up now" button, the Nextcloud settings form with test button,
  and the restore picker (local + remote sets, type-to-confirm)
- global maintenance screen: any 503 maintenance_mode flips the SPA to a
  status page polling the exempt endpoint, reloading when the instance
  returns

Verified end-to-end against a live stack (fresh DB, native api + sidecar,
fake WebDAV server): configure → test → manual backup → bundle upload →
readyz/sets/status surfaces → remote restore with maintenance gate,
marker rollback and api restart; suites: shared 21, backup 9, collab 11,
api 58 files green, lint + i18n:check + typecheck clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
2026-07-12 10:39:18 +02:00

68 lines
3.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Uptime monitoring (issue #85)
Pragmatic monitoring per operations.md: no metrics stack — the health
endpoints are the single integration point, watched by the operator's
existing **Uptime-Kuma** instance.
## Semantics: down vs. degraded
`GET /api/v1/readyz` returns:
| HTTP | `status` | Meaning | Reaction |
| ---- | ---------- | ------------------------------------------------------------------- | --------------------------------------------- |
| 200 | `ok` | all checks green | — |
| 200 | `degraded` | a warning-level check: converter/renderer down, backup stale (26 h) | alert, fix without urgency — users are served |
| 503 | `unready` | hard failure: database unreachable or migrations pending | page — the instance cannot serve |
Checks enumerated in the body: `database`, `migrations` (hard),
`converter`, `renderer`, `backup` (warning-level; `backup` reads the
sidecar's `status.json` and warns when the last successful backup is
older than 26 h — ADR 0015). With a Nextcloud backup target configured
(issue #103) a `backup_remote` check appears too: it warns when the last
successful off-host upload is older than its schedule allows (26 h daily,
170 h weekly; manual-only schedules are never stale).
During an **in-app restore** (issue #103) every application route answers
503 `maintenance_mode` for a few minutes and the api restarts itself once
afterwards; `healthz`/`readyz` and `GET /api/v1/backup/restore-status`
keep answering throughout. A short monitor blip around a restore is
expected.
**Degraded never restarts containers**: the Docker healthchecks use the
liveness endpoints only (api `/api/v1/healthz`, web `/healthz`, collab
`/healthz`), never `readyz`. The collab `/healthz` includes its database
probe (issue #33) — Compose marks the container `unhealthy` then, but
`restart: unless-stopped` only restarts on process exit, so an unhealthy
container is visible in `docker compose ps`, not restart-looped.
## Monitor set (per stage)
Four monitors per stage in Uptime-Kuma; suggested interval 60 s, retries 2.
| # | Monitor | Type | Target (Test example) | Alert condition |
| --- | ---------------------- | ----------------- | ------------------------------------------------- | -------------------------------------- |
| 1 | `<stage> web` | HTTP(s) | `https://test.dorfteich.cloud/healthz` | non-2xx |
| 2 | `<stage> api ready` | HTTP(s) | `https://test.dorfteich.cloud/api/v1/readyz` | non-2xx (= unready/down) |
| 3 | `<stage> api degraded` | HTTP(s) Keyword | same URL, keyword `"status":"ok"` must be present | body says `degraded` while HTTP is 200 |
| 4 | `<stage> collab` | WebSocket | `wss://test.dorfteich.cloud/collab` | connect failure |
Monitor 3 is what catches `degraded` (stale backup, dead sidecars) —
monitor 2 alone only sees hard failures. If the Kuma version has no
WebSocket monitor type, an HTTP monitor on
`https://<stage>/collab/healthz` (200, includes the DB probe) is the
fallback for monitor 4.
Stages: `test.dorfteich.cloud`, `int.dorfteich.cloud` (web-check-only is
acceptable per operations.md — at minimum monitors 12, no paging),
`dorfteich.online` (Prod, full set + notification; set up and verified at
go-live, checklist item in #89).
Setting up the monitors is an operator action in Uptime-Kuma (owner:
Stefan); this file is the definition they follow.
## Related
- operations.md §Health & monitoring (endpoint semantics)
- ADR 0015 (backup freshness feeds the `backup` check)
- deploy/backup/restore.sh (what to do when the backup check warns)