All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m45s
CD / Build and push images (push) Successful in 3m49s
CI / Build container images (push) Has been skipped
CD / Deploy to Test (push) Successful in 9s
CD / Smoke tests against Test (push) Successful in 1m18s
CD / Promote to Int (push) Successful in 11s
CI / Auth e2e pack (push) Successful in 5m35s
CI / Import/export fidelity gate (push) Successful in 47s
Off-host backups for every self-hoster, configured entirely in the admin UI — supersedes the host-specific mirror plan behind #84. shared: - webdav.ts (new package entry like token-crypto): minimal WebDAV client with basic auth — PROPFIND (tolerant multistatus parser), MKCOL, PUT (streamed), GET, DELETE; Nextcloud DAV path derived from the plain server URL, explicit DAV bases pass through - backup-status.ts: additive remote-upload status in status.json, the restore-status.json contract (running/succeeded/failed + staleness bound), the backup_command/backup_maintenance NOTIFY channels, and the one-bundle-per-set naming (dorfteich-backup-<id>.tar.gz) - backup-set.ts moved here from apps/backup (api lists local sets) backup sidecar: - reads the backup.* instance settings directly from the database (admin changes apply next run; local retention row overrides the env) and the app password from the secret store - after each successful set: bundle dump + files archive + manifest into ONE self-contained tar.gz, upload via WebDAV per schedule (off/daily/weekly; manual runs always upload), prune remote bundles — never the newest — and record the outcome in status.json; upload failures alert via a new backupUploadFailed mail (de+en) - command listener on backup_command (run / restore) with a serial queue against the nightly timer - restore orchestrator: restore-status.json → maintenance NOTIFY → grace → (remote: download + manifest-verify bundle) → terminate other DB connections → shared perform-restore path (same code as restore.sh) → final status + maintenance exit api: - MaintenanceGuard (global, registered before the setup gate): 503 maintenance_mode while restore-status says running; health endpoints and the new public GET /backup/restore-status stay exempt; a stale running state (crashed sidecar) unblocks after 30 min - MaintenanceStateService watches the file and restarts the api after a successful restore (fresh caches, migrate-on-start for older dumps); main.ts refuses to touch the database while a restore runs — a container restarting mid-restore must not race pg_restore with migrate deploy - worker sweeps (conversion, mail outbox, scheduler) catch transient database failures instead of dying on an unhandled rejection — the restore's connection termination crashed the api in verification - backup admin endpoints under /admin/system/backup: settings (live connection test before save, password write-only into the secret store), nextcloud/test, sets (local via the ro backups mount + remote via WebDAV), run + restore (type-to-confirm backstop, source validation) — commands travel as NOTIFY payloads; audit actions backup.settings_changed/run_triggered/restore_requested - readyz: new warning-level backup_remote check while a target is configured (26 h daily / 170 h weekly bound) collab: - maintenance listener: on enter, persist + close every live session and refuse new connections until exit (failsafe timeout 30 min) — no in-memory document may write pre-restore content back afterwards web: - Admin → System backup section: status card with remote facts and a "Back up now" button, the Nextcloud settings form with test button, and the restore picker (local + remote sets, type-to-confirm) - global maintenance screen: any 503 maintenance_mode flips the SPA to a status page polling the exempt endpoint, reloading when the instance returns Verified end-to-end against a live stack (fresh DB, native api + sidecar, fake WebDAV server): configure → test → manual backup → bundle upload → readyz/sets/status surfaces → remote restore with maintenance gate, marker rollback and api restart; suites: shared 21, backup 9, collab 11, api 58 files green, lint + i18n:check + typecheck clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
111 lines
4.3 KiB
TypeScript
111 lines
4.3 KiB
TypeScript
import { existsSync } from 'node:fs';
|
|
import { join } from 'node:path';
|
|
import { setTimeout as sleep } from 'node:timers/promises';
|
|
|
|
import type { BackupCommand, MaintenanceEvent, RestoreStatus } from '@dorfteich/shared';
|
|
|
|
import { archiveFileName, dumpFileName } from './backup-set.js';
|
|
import type { RemoteLogger } from './remote.js';
|
|
import { writeRestoreStatus } from './status.js';
|
|
|
|
/**
|
|
* The in-app restore, orchestrated by the sidecar (issue #103):
|
|
*
|
|
* 1. `restore-status.json` → `running` — the api's maintenance gate flips
|
|
* to 503 for everything but health and the status endpoint.
|
|
* 2. NOTIFY maintenance `enter` — the collab server persists + closes every
|
|
* live session and refuses new ones, so no in-memory document writes
|
|
* pre-restore content back afterwards.
|
|
* 3. Grace period, then terminate all other database connections.
|
|
* 4. Fetch the set (remote source: download + unpack the bundle), verify
|
|
* both artifacts exist, `pg_restore --clean` + volume archive extract —
|
|
* the exact path of the operator's `restore.sh`.
|
|
* 5. `restore-status.json` → final state; NOTIFY `exit`. The api restarts
|
|
* itself on `succeeded` (fresh caches, migrate-on-start for older dumps).
|
|
*
|
|
* Every failure lands in `restore-status.json` — the admin watches that
|
|
* file through the exempt status endpoint, so it must always resolve.
|
|
*/
|
|
|
|
export interface RestoreOrchestratorDeps {
|
|
backupsDir: string;
|
|
now(): Date;
|
|
/** NOTIFY on a short-lived connection (maintenance enter/exit). */
|
|
notifyMaintenance(event: MaintenanceEvent): Promise<void>;
|
|
/** Kills every other DB connection right before pg_restore. */
|
|
terminateOtherConnections(): Promise<void>;
|
|
/** Downloads + unpacks the remote bundle into backupsDir (remote.ts). */
|
|
fetchRemoteSet(backupId: string): Promise<void>;
|
|
/** pg_restore + volume extract — shared with the restore.js CLI. */
|
|
restoreSet(backupId: string): Promise<void>;
|
|
/** Milliseconds between maintenance enter and connection termination. */
|
|
graceMs?: number;
|
|
log: RemoteLogger;
|
|
}
|
|
|
|
export async function orchestrateRestore(
|
|
deps: RestoreOrchestratorDeps,
|
|
command: Extract<BackupCommand, { kind: 'restore' }>,
|
|
): Promise<RestoreStatus> {
|
|
const startedAt = deps.now().toISOString();
|
|
const base: Omit<RestoreStatus, 'state' | 'finishedAt'> = {
|
|
schemaVersion: 1,
|
|
backupId: command.backupId,
|
|
source: command.source,
|
|
requestedBy: command.requestedBy,
|
|
startedAt,
|
|
};
|
|
await writeRestoreStatus(deps.backupsDir, { ...base, state: 'running', finishedAt: null });
|
|
deps.log.info(
|
|
{ backupId: command.backupId, source: command.source, requestedBy: command.requestedBy },
|
|
'restore started — instance entering maintenance mode',
|
|
);
|
|
|
|
let entered = false;
|
|
try {
|
|
await deps.notifyMaintenance({ phase: 'enter' });
|
|
entered = true;
|
|
await sleep(deps.graceMs ?? 5000);
|
|
|
|
if (command.source === 'remote') {
|
|
await deps.fetchRemoteSet(command.backupId);
|
|
}
|
|
for (const file of [dumpFileName(command.backupId), archiveFileName(command.backupId)]) {
|
|
if (!existsSync(join(deps.backupsDir, file))) {
|
|
throw new Error(`restore set is incomplete — ${file} not found`);
|
|
}
|
|
}
|
|
|
|
await deps.terminateOtherConnections();
|
|
await deps.restoreSet(command.backupId);
|
|
|
|
const status: RestoreStatus = {
|
|
...base,
|
|
state: 'succeeded',
|
|
finishedAt: deps.now().toISOString(),
|
|
};
|
|
await writeRestoreStatus(deps.backupsDir, status);
|
|
deps.log.info({ backupId: command.backupId }, 'restore succeeded — api will restart');
|
|
return status;
|
|
} catch (error) {
|
|
const message = error instanceof Error ? error.message : String(error);
|
|
const status: RestoreStatus = {
|
|
...base,
|
|
state: 'failed',
|
|
finishedAt: deps.now().toISOString(),
|
|
error: message,
|
|
};
|
|
// Best effort — if even this write fails the api's staleness bound
|
|
// (RESTORE_STALE_MAX_AGE_MINUTES) unblocks the instance eventually.
|
|
await writeRestoreStatus(deps.backupsDir, status).catch(() => undefined);
|
|
deps.log.error({ backupId: command.backupId, error: message }, 'restore failed');
|
|
return status;
|
|
} finally {
|
|
if (entered) {
|
|
await deps.notifyMaintenance({ phase: 'exit' }).catch((error: unknown) => {
|
|
deps.log.warn({ error: String(error) }, 'maintenance exit notify failed');
|
|
});
|
|
}
|
|
}
|
|
}
|