All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m45s
CD / Build and push images (push) Successful in 3m49s
CI / Build container images (push) Has been skipped
CD / Deploy to Test (push) Successful in 9s
CD / Smoke tests against Test (push) Successful in 1m18s
CD / Promote to Int (push) Successful in 11s
CI / Auth e2e pack (push) Successful in 5m35s
CI / Import/export fidelity gate (push) Successful in 47s
Off-host backups for every self-hoster, configured entirely in the admin UI — supersedes the host-specific mirror plan behind #84. shared: - webdav.ts (new package entry like token-crypto): minimal WebDAV client with basic auth — PROPFIND (tolerant multistatus parser), MKCOL, PUT (streamed), GET, DELETE; Nextcloud DAV path derived from the plain server URL, explicit DAV bases pass through - backup-status.ts: additive remote-upload status in status.json, the restore-status.json contract (running/succeeded/failed + staleness bound), the backup_command/backup_maintenance NOTIFY channels, and the one-bundle-per-set naming (dorfteich-backup-<id>.tar.gz) - backup-set.ts moved here from apps/backup (api lists local sets) backup sidecar: - reads the backup.* instance settings directly from the database (admin changes apply next run; local retention row overrides the env) and the app password from the secret store - after each successful set: bundle dump + files archive + manifest into ONE self-contained tar.gz, upload via WebDAV per schedule (off/daily/weekly; manual runs always upload), prune remote bundles — never the newest — and record the outcome in status.json; upload failures alert via a new backupUploadFailed mail (de+en) - command listener on backup_command (run / restore) with a serial queue against the nightly timer - restore orchestrator: restore-status.json → maintenance NOTIFY → grace → (remote: download + manifest-verify bundle) → terminate other DB connections → shared perform-restore path (same code as restore.sh) → final status + maintenance exit api: - MaintenanceGuard (global, registered before the setup gate): 503 maintenance_mode while restore-status says running; health endpoints and the new public GET /backup/restore-status stay exempt; a stale running state (crashed sidecar) unblocks after 30 min - MaintenanceStateService watches the file and restarts the api after a successful restore (fresh caches, migrate-on-start for older dumps); main.ts refuses to touch the database while a restore runs — a container restarting mid-restore must not race pg_restore with migrate deploy - worker sweeps (conversion, mail outbox, scheduler) catch transient database failures instead of dying on an unhandled rejection — the restore's connection termination crashed the api in verification - backup admin endpoints under /admin/system/backup: settings (live connection test before save, password write-only into the secret store), nextcloud/test, sets (local via the ro backups mount + remote via WebDAV), run + restore (type-to-confirm backstop, source validation) — commands travel as NOTIFY payloads; audit actions backup.settings_changed/run_triggered/restore_requested - readyz: new warning-level backup_remote check while a target is configured (26 h daily / 170 h weekly bound) collab: - maintenance listener: on enter, persist + close every live session and refuse new connections until exit (failsafe timeout 30 min) — no in-memory document may write pre-restore content back afterwards web: - Admin → System backup section: status card with remote facts and a "Back up now" button, the Nextcloud settings form with test button, and the restore picker (local + remote sets, type-to-confirm) - global maintenance screen: any 503 maintenance_mode flips the SPA to a status page polling the exempt endpoint, reloading when the instance returns Verified end-to-end against a live stack (fresh DB, native api + sidecar, fake WebDAV server): configure → test → manual backup → bundle upload → readyz/sets/status surfaces → remote restore with maintenance gate, marker rollback and api restart; suites: shared 21, backup 9, collab 11, api 58 files green, lint + i18n:check + typecheck clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
148 lines
5.7 KiB
TypeScript
148 lines
5.7 KiB
TypeScript
import { Injectable } from '@nestjs/common';
|
|
|
|
import { BackupTargetService } from '../backup/backup-target.service';
|
|
import { AppConfig } from '../config/app-config.service';
|
|
import { PrismaService } from '../prisma/prisma.service';
|
|
import { InstanceSettingsService } from '../settings/instance-settings.service';
|
|
import { backupFreshnessCheck, backupRemoteFreshnessCheck } from './backup-freshness';
|
|
|
|
export interface ReadinessCheck {
|
|
/** `warn` reports a degraded-but-serving dependency: the instance still
|
|
* works, only some feature is unavailable. It never flips overall
|
|
* readiness to `unready` (that is reserved for `failed`). */
|
|
name: string;
|
|
status: 'ok' | 'warn' | 'failed';
|
|
detail?: string;
|
|
}
|
|
|
|
export interface ReadinessReport {
|
|
/**
|
|
* `degraded` = at least one warning-level check (converter/renderer down,
|
|
* backup stale) while the instance still serves; monitors alert on it via
|
|
* the body, but only `unready` turns into HTTP 503 (issue #85,
|
|
* operations.md §Health).
|
|
*/
|
|
status: 'ok' | 'degraded' | 'unready';
|
|
checks: ReadinessCheck[];
|
|
}
|
|
|
|
/** Converter reachability is a warning, not a failure — the pandoc probe is
|
|
* given a short budget so readyz stays fast even when the sidecar is down. */
|
|
const CONVERTER_PROBE_TIMEOUT_MS = 2000;
|
|
|
|
@Injectable()
|
|
export class ReadinessService {
|
|
constructor(
|
|
private readonly prisma: PrismaService,
|
|
private readonly config: AppConfig,
|
|
private readonly backupTarget: BackupTargetService,
|
|
private readonly settings: InstanceSettingsService,
|
|
) {}
|
|
|
|
/**
|
|
* Readiness = the api can do real work: database reachable and all
|
|
* migrations applied — those two are the hard failures behind HTTP 503.
|
|
* Everything else (converter, renderer, backup freshness) is
|
|
* warning-level: the feature degrades, the instance stays ready, and the
|
|
* overall status says `degraded` so monitors can alert on the body.
|
|
*/
|
|
async report(): Promise<ReadinessReport> {
|
|
const checks: ReadinessCheck[] = [
|
|
await this.databaseReachable(),
|
|
await this.migrationsApplied(),
|
|
await this.converterReachable(),
|
|
await this.rendererReachable(),
|
|
backupFreshnessCheck(this.config.env.BACKUPS_DIR, new Date()),
|
|
...(await this.remoteBackupCheck()),
|
|
];
|
|
const status = checks.some((c) => c.status === 'failed')
|
|
? 'unready'
|
|
: checks.some((c) => c.status === 'warn')
|
|
? 'degraded'
|
|
: 'ok';
|
|
return { status, checks };
|
|
}
|
|
|
|
/** Off-host copy freshness (issue #103) — only while a target is
|
|
* configured; the check itself never talks to the network, it reads the
|
|
* sidecar's status.json. A database hiccup here must not break readyz. */
|
|
private async remoteBackupCheck(): Promise<ReadinessCheck[]> {
|
|
try {
|
|
if (!(await this.backupTarget.resolveTarget())) return [];
|
|
const schedule = await this.settings.get('backup.nextcloud.uploadSchedule');
|
|
return [backupRemoteFreshnessCheck(this.config.env.BACKUPS_DIR, new Date(), schedule)];
|
|
} catch {
|
|
return [];
|
|
}
|
|
}
|
|
|
|
private async converterReachable(): Promise<ReadinessCheck> {
|
|
const controller = new AbortController();
|
|
const timer = setTimeout(() => controller.abort(), CONVERTER_PROBE_TIMEOUT_MS);
|
|
try {
|
|
const response = await fetch(`${this.config.env.PANDOC_URL}/version`, {
|
|
signal: controller.signal,
|
|
});
|
|
return response.ok
|
|
? { name: 'converter', status: 'ok' }
|
|
: { name: 'converter', status: 'warn', detail: `pandoc returned ${response.status}` };
|
|
} catch (error) {
|
|
return { name: 'converter', status: 'warn', detail: shortMessage(error) };
|
|
} finally {
|
|
clearTimeout(timer);
|
|
}
|
|
}
|
|
|
|
/** The Gotenberg PDF renderer (issue #67) — warning-level like the converter:
|
|
* PDF export degrades when it is down, but the instance stays ready. */
|
|
private async rendererReachable(): Promise<ReadinessCheck> {
|
|
const controller = new AbortController();
|
|
const timer = setTimeout(() => controller.abort(), CONVERTER_PROBE_TIMEOUT_MS);
|
|
try {
|
|
const response = await fetch(`${this.config.env.GOTENBERG_URL}/health`, {
|
|
signal: controller.signal,
|
|
});
|
|
return response.ok
|
|
? { name: 'renderer', status: 'ok' }
|
|
: { name: 'renderer', status: 'warn', detail: `gotenberg returned ${response.status}` };
|
|
} catch (error) {
|
|
return { name: 'renderer', status: 'warn', detail: shortMessage(error) };
|
|
} finally {
|
|
clearTimeout(timer);
|
|
}
|
|
}
|
|
|
|
private async databaseReachable(): Promise<ReadinessCheck> {
|
|
try {
|
|
await this.prisma.$queryRaw`SELECT 1`;
|
|
return { name: 'database', status: 'ok' };
|
|
} catch (error) {
|
|
return { name: 'database', status: 'failed', detail: shortMessage(error) };
|
|
}
|
|
}
|
|
|
|
private async migrationsApplied(): Promise<ReadinessCheck> {
|
|
try {
|
|
// `prisma migrate deploy` records every migration here; an entry
|
|
// without finished_at is pending or failed.
|
|
const rows = await this.prisma.$queryRaw<{ pending: bigint }[]>`
|
|
SELECT count(*)::bigint AS pending
|
|
FROM _prisma_migrations
|
|
WHERE finished_at IS NULL AND rolled_back_at IS NULL
|
|
`;
|
|
const pending = Number(rows[0]?.pending ?? 0);
|
|
return pending === 0
|
|
? { name: 'migrations', status: 'ok' }
|
|
: { name: 'migrations', status: 'failed', detail: `${pending} migration(s) pending` };
|
|
} catch (error) {
|
|
return { name: 'migrations', status: 'failed', detail: shortMessage(error) };
|
|
}
|
|
}
|
|
}
|
|
|
|
function shortMessage(error: unknown): string {
|
|
const message = error instanceof Error ? error.message : String(error);
|
|
// Keep readiness output single-line and free of connection strings.
|
|
return message.split('\n').filter(Boolean).slice(-1)[0]?.slice(0, 200) ?? 'unknown error';
|
|
}
|