All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m45s
CD / Build and push images (push) Successful in 3m49s
CI / Build container images (push) Has been skipped
CD / Deploy to Test (push) Successful in 9s
CD / Smoke tests against Test (push) Successful in 1m18s
CD / Promote to Int (push) Successful in 11s
CI / Auth e2e pack (push) Successful in 5m35s
CI / Import/export fidelity gate (push) Successful in 47s
Off-host backups for every self-hoster, configured entirely in the admin UI — supersedes the host-specific mirror plan behind #84. shared: - webdav.ts (new package entry like token-crypto): minimal WebDAV client with basic auth — PROPFIND (tolerant multistatus parser), MKCOL, PUT (streamed), GET, DELETE; Nextcloud DAV path derived from the plain server URL, explicit DAV bases pass through - backup-status.ts: additive remote-upload status in status.json, the restore-status.json contract (running/succeeded/failed + staleness bound), the backup_command/backup_maintenance NOTIFY channels, and the one-bundle-per-set naming (dorfteich-backup-<id>.tar.gz) - backup-set.ts moved here from apps/backup (api lists local sets) backup sidecar: - reads the backup.* instance settings directly from the database (admin changes apply next run; local retention row overrides the env) and the app password from the secret store - after each successful set: bundle dump + files archive + manifest into ONE self-contained tar.gz, upload via WebDAV per schedule (off/daily/weekly; manual runs always upload), prune remote bundles — never the newest — and record the outcome in status.json; upload failures alert via a new backupUploadFailed mail (de+en) - command listener on backup_command (run / restore) with a serial queue against the nightly timer - restore orchestrator: restore-status.json → maintenance NOTIFY → grace → (remote: download + manifest-verify bundle) → terminate other DB connections → shared perform-restore path (same code as restore.sh) → final status + maintenance exit api: - MaintenanceGuard (global, registered before the setup gate): 503 maintenance_mode while restore-status says running; health endpoints and the new public GET /backup/restore-status stay exempt; a stale running state (crashed sidecar) unblocks after 30 min - MaintenanceStateService watches the file and restarts the api after a successful restore (fresh caches, migrate-on-start for older dumps); main.ts refuses to touch the database while a restore runs — a container restarting mid-restore must not race pg_restore with migrate deploy - worker sweeps (conversion, mail outbox, scheduler) catch transient database failures instead of dying on an unhandled rejection — the restore's connection termination crashed the api in verification - backup admin endpoints under /admin/system/backup: settings (live connection test before save, password write-only into the secret store), nextcloud/test, sets (local via the ro backups mount + remote via WebDAV), run + restore (type-to-confirm backstop, source validation) — commands travel as NOTIFY payloads; audit actions backup.settings_changed/run_triggered/restore_requested - readyz: new warning-level backup_remote check while a target is configured (26 h daily / 170 h weekly bound) collab: - maintenance listener: on enter, persist + close every live session and refuse new connections until exit (failsafe timeout 30 min) — no in-memory document may write pre-restore content back afterwards web: - Admin → System backup section: status card with remote facts and a "Back up now" button, the Nextcloud settings form with test button, and the restore picker (local + remote sets, type-to-confirm) - global maintenance screen: any 503 maintenance_mode flips the SPA to a status page polling the exempt endpoint, reloading when the instance returns Verified end-to-end against a live stack (fresh DB, native api + sidecar, fake WebDAV server): configure → test → manual backup → bundle upload → readyz/sets/status surfaces → remote restore with maintenance gate, marker rollback and api restart; suites: shared 21, backup 9, collab 11, api 58 files green, lint + i18n:check + typecheck clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
68 lines
3.9 KiB
Markdown
68 lines
3.9 KiB
Markdown
# Uptime monitoring (issue #85)
|
||
|
||
Pragmatic monitoring per operations.md: no metrics stack — the health
|
||
endpoints are the single integration point, watched by the operator's
|
||
existing **Uptime-Kuma** instance.
|
||
|
||
## Semantics: down vs. degraded
|
||
|
||
`GET /api/v1/readyz` returns:
|
||
|
||
| HTTP | `status` | Meaning | Reaction |
|
||
| ---- | ---------- | ------------------------------------------------------------------- | --------------------------------------------- |
|
||
| 200 | `ok` | all checks green | — |
|
||
| 200 | `degraded` | a warning-level check: converter/renderer down, backup stale (26 h) | alert, fix without urgency — users are served |
|
||
| 503 | `unready` | hard failure: database unreachable or migrations pending | page — the instance cannot serve |
|
||
|
||
Checks enumerated in the body: `database`, `migrations` (hard),
|
||
`converter`, `renderer`, `backup` (warning-level; `backup` reads the
|
||
sidecar's `status.json` and warns when the last successful backup is
|
||
older than 26 h — ADR 0015). With a Nextcloud backup target configured
|
||
(issue #103) a `backup_remote` check appears too: it warns when the last
|
||
successful off-host upload is older than its schedule allows (26 h daily,
|
||
170 h weekly; manual-only schedules are never stale).
|
||
|
||
During an **in-app restore** (issue #103) every application route answers
|
||
503 `maintenance_mode` for a few minutes and the api restarts itself once
|
||
afterwards; `healthz`/`readyz` and `GET /api/v1/backup/restore-status`
|
||
keep answering throughout. A short monitor blip around a restore is
|
||
expected.
|
||
|
||
**Degraded never restarts containers**: the Docker healthchecks use the
|
||
liveness endpoints only (api `/api/v1/healthz`, web `/healthz`, collab
|
||
`/healthz`), never `readyz`. The collab `/healthz` includes its database
|
||
probe (issue #33) — Compose marks the container `unhealthy` then, but
|
||
`restart: unless-stopped` only restarts on process exit, so an unhealthy
|
||
container is visible in `docker compose ps`, not restart-looped.
|
||
|
||
## Monitor set (per stage)
|
||
|
||
Four monitors per stage in Uptime-Kuma; suggested interval 60 s, retries 2.
|
||
|
||
| # | Monitor | Type | Target (Test example) | Alert condition |
|
||
| --- | ---------------------- | ----------------- | ------------------------------------------------- | -------------------------------------- |
|
||
| 1 | `<stage> web` | HTTP(s) | `https://test.dorfteich.cloud/healthz` | non-2xx |
|
||
| 2 | `<stage> api ready` | HTTP(s) | `https://test.dorfteich.cloud/api/v1/readyz` | non-2xx (= unready/down) |
|
||
| 3 | `<stage> api degraded` | HTTP(s) – Keyword | same URL, keyword `"status":"ok"` must be present | body says `degraded` while HTTP is 200 |
|
||
| 4 | `<stage> collab` | WebSocket | `wss://test.dorfteich.cloud/collab` | connect failure |
|
||
|
||
Monitor 3 is what catches `degraded` (stale backup, dead sidecars) —
|
||
monitor 2 alone only sees hard failures. If the Kuma version has no
|
||
WebSocket monitor type, an HTTP monitor on
|
||
`https://<stage>/collab/healthz` (200, includes the DB probe) is the
|
||
fallback for monitor 4.
|
||
|
||
Stages: `test.dorfteich.cloud`, `int.dorfteich.cloud` (web-check-only is
|
||
acceptable per operations.md — at minimum monitors 1–2, no paging),
|
||
`dorfteich.online` (Prod, full set + notification; set up and verified at
|
||
go-live, checklist item in #89).
|
||
|
||
Setting up the monitors is an operator action in Uptime-Kuma (owner:
|
||
Stefan); this file is the definition they follow.
|
||
|
||
## Related
|
||
|
||
- operations.md §Health & monitoring (endpoint semantics)
|
||
- ADR 0015 (backup freshness feeds the `backup` check)
|
||
- deploy/backup/restore.sh (what to do when the backup check warns)
|