dorfteich/apps/api/src/health/backup-freshness.ts
Claude Fable 5 5cef359b8f
All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m45s
CD / Build and push images (push) Successful in 3m49s
CI / Build container images (push) Has been skipped
CD / Deploy to Test (push) Successful in 9s
CD / Smoke tests against Test (push) Successful in 1m18s
CD / Promote to Int (push) Successful in 11s
CI / Auth e2e pack (push) Successful in 5m35s
CI / Import/export fidelity gate (push) Successful in 47s
Nextcloud backup target: admin-configured, manual + scheduled uploads, in-app restore (#103)
Off-host backups for every self-hoster, configured entirely in the admin
UI — supersedes the host-specific mirror plan behind #84.

shared:
- webdav.ts (new package entry like token-crypto): minimal WebDAV client
  with basic auth — PROPFIND (tolerant multistatus parser), MKCOL, PUT
  (streamed), GET, DELETE; Nextcloud DAV path derived from the plain
  server URL, explicit DAV bases pass through
- backup-status.ts: additive remote-upload status in status.json, the
  restore-status.json contract (running/succeeded/failed + staleness
  bound), the backup_command/backup_maintenance NOTIFY channels, and the
  one-bundle-per-set naming (dorfteich-backup-<id>.tar.gz)
- backup-set.ts moved here from apps/backup (api lists local sets)

backup sidecar:
- reads the backup.* instance settings directly from the database (admin
  changes apply next run; local retention row overrides the env) and the
  app password from the secret store
- after each successful set: bundle dump + files archive + manifest into
  ONE self-contained tar.gz, upload via WebDAV per schedule
  (off/daily/weekly; manual runs always upload), prune remote bundles —
  never the newest — and record the outcome in status.json; upload
  failures alert via a new backupUploadFailed mail (de+en)
- command listener on backup_command (run / restore) with a serial queue
  against the nightly timer
- restore orchestrator: restore-status.json → maintenance NOTIFY →
  grace → (remote: download + manifest-verify bundle) → terminate other
  DB connections → shared perform-restore path (same code as restore.sh)
  → final status + maintenance exit

api:
- MaintenanceGuard (global, registered before the setup gate): 503
  maintenance_mode while restore-status says running; health endpoints
  and the new public GET /backup/restore-status stay exempt; a stale
  running state (crashed sidecar) unblocks after 30 min
- MaintenanceStateService watches the file and restarts the api after a
  successful restore (fresh caches, migrate-on-start for older dumps);
  main.ts refuses to touch the database while a restore runs — a
  container restarting mid-restore must not race pg_restore with
  migrate deploy
- worker sweeps (conversion, mail outbox, scheduler) catch transient
  database failures instead of dying on an unhandled rejection — the
  restore's connection termination crashed the api in verification
- backup admin endpoints under /admin/system/backup: settings (live
  connection test before save, password write-only into the secret
  store), nextcloud/test, sets (local via the ro backups mount + remote
  via WebDAV), run + restore (type-to-confirm backstop, source
  validation) — commands travel as NOTIFY payloads; audit actions
  backup.settings_changed/run_triggered/restore_requested
- readyz: new warning-level backup_remote check while a target is
  configured (26 h daily / 170 h weekly bound)

collab:
- maintenance listener: on enter, persist + close every live session and
  refuse new connections until exit (failsafe timeout 30 min) — no
  in-memory document may write pre-restore content back afterwards

web:
- Admin → System backup section: status card with remote facts and a
  "Back up now" button, the Nextcloud settings form with test button,
  and the restore picker (local + remote sets, type-to-confirm)
- global maintenance screen: any 503 maintenance_mode flips the SPA to a
  status page polling the exempt endpoint, reloading when the instance
  returns

Verified end-to-end against a live stack (fresh DB, native api + sidecar,
fake WebDAV server): configure → test → manual backup → bundle upload →
readyz/sets/status surfaces → remote restore with maintenance gate,
marker rollback and api restart; suites: shared 21, backup 9, collab 11,
api 58 files green, lint + i18n:check + typecheck clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
2026-07-12 10:39:18 +02:00

122 lines
4.0 KiB
TypeScript

import { existsSync, readFileSync } from 'node:fs';
import { join } from 'node:path';
import {
BACKUP_FRESH_MAX_AGE_HOURS,
BACKUP_REMOTE_WEEKLY_MAX_AGE_HOURS,
BACKUP_STATUS_FILE,
type BackupStatus,
} from '@dorfteich/shared';
import type { ReadinessCheck } from './readiness.service';
/**
* Backup freshness for readyz (issue #85, operations.md §Health): reads the
* sidecar's `status.json` from the shared backups volume. Always
* warning-level — a stale or missing backup degrades the instance so
* monitors alert, but the api keeps serving (503 is reserved for the
* database/migrations hard failures).
*/
export function backupFreshnessCheck(backupsDir: string, now: Date): ReadinessCheck {
const path = join(backupsDir, BACKUP_STATUS_FILE);
if (!existsSync(path)) {
return { name: 'backup', status: 'warn', detail: 'no backup status recorded yet' };
}
let status: BackupStatus;
try {
status = JSON.parse(readFileSync(path, 'utf8')) as BackupStatus;
} catch {
return { name: 'backup', status: 'warn', detail: 'backup status file is unreadable' };
}
if (!status.lastSuccess) {
const error = status.lastRun?.error;
return {
name: 'backup',
status: 'warn',
detail: error ? `no successful backup yet — last run: ${error}` : 'no successful backup yet',
};
}
const ageHours = (now.getTime() - new Date(status.lastSuccess.finishedAt).getTime()) / 3_600_000;
if (!Number.isFinite(ageHours) || ageHours > BACKUP_FRESH_MAX_AGE_HOURS) {
return {
name: 'backup',
status: 'warn',
detail: `last successful backup ${status.lastSuccess.backupId} is ${Math.round(ageHours)} h old (max ${BACKUP_FRESH_MAX_AGE_HOURS} h)`,
};
}
// Fresh success; still surface a failed newer run so operators see it
// before the freshness window runs out.
if (status.lastRun.outcome === 'failed') {
return {
name: 'backup',
status: 'ok',
detail: `fresh, but the last run failed: ${status.lastRun.error ?? 'unknown error'}`,
};
}
return { name: 'backup', status: 'ok' };
}
/**
* Off-host copy freshness (issue #103) — evaluated only while a Nextcloud
* target is configured. Like the local check it is always warning-level.
* With schedule `off` (manual uploads only) there is no cadence to hold the
* instance to, so the check reports ok with a detail.
*/
export function backupRemoteFreshnessCheck(
backupsDir: string,
now: Date,
schedule: 'off' | 'daily' | 'weekly',
): ReadinessCheck {
const name = 'backup_remote';
if (schedule === 'off') {
return { name, status: 'ok', detail: 'manual uploads only (schedule off)' };
}
const maxAgeHours =
schedule === 'weekly' ? BACKUP_REMOTE_WEEKLY_MAX_AGE_HOURS : BACKUP_FRESH_MAX_AGE_HOURS;
const path = join(backupsDir, BACKUP_STATUS_FILE);
let status: BackupStatus | null = null;
if (existsSync(path)) {
try {
status = JSON.parse(readFileSync(path, 'utf8')) as BackupStatus;
} catch {
status = null;
}
}
if (!status) {
return { name, status: 'warn', detail: 'no backup status recorded yet' };
}
if (!status.remote?.lastSuccessfulUpload) {
const error = status.remote?.lastUpload?.error;
return {
name,
status: 'warn',
detail: error
? `no successful off-host upload yet — last attempt: ${error}`
: 'no successful off-host upload yet',
};
}
const upload = status.remote.lastSuccessfulUpload;
const ageHours = (now.getTime() - new Date(upload.finishedAt).getTime()) / 3_600_000;
if (!Number.isFinite(ageHours) || ageHours > maxAgeHours) {
return {
name,
status: 'warn',
detail: `last off-host copy ${upload.backupId} is ${Math.round(ageHours)} h old (max ${maxAgeHours} h)`,
};
}
if (status.remote.lastUpload.outcome === 'failed') {
return {
name,
status: 'ok',
detail: `fresh, but the last upload failed: ${status.remote.lastUpload.error ?? 'unknown error'}`,
};
}
return { name, status: 'ok' };
}