All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m45s
CD / Build and push images (push) Successful in 3m49s
CI / Build container images (push) Has been skipped
CD / Deploy to Test (push) Successful in 9s
CD / Smoke tests against Test (push) Successful in 1m18s
CD / Promote to Int (push) Successful in 11s
CI / Auth e2e pack (push) Successful in 5m35s
CI / Import/export fidelity gate (push) Successful in 47s
Off-host backups for every self-hoster, configured entirely in the admin UI — supersedes the host-specific mirror plan behind #84. shared: - webdav.ts (new package entry like token-crypto): minimal WebDAV client with basic auth — PROPFIND (tolerant multistatus parser), MKCOL, PUT (streamed), GET, DELETE; Nextcloud DAV path derived from the plain server URL, explicit DAV bases pass through - backup-status.ts: additive remote-upload status in status.json, the restore-status.json contract (running/succeeded/failed + staleness bound), the backup_command/backup_maintenance NOTIFY channels, and the one-bundle-per-set naming (dorfteich-backup-<id>.tar.gz) - backup-set.ts moved here from apps/backup (api lists local sets) backup sidecar: - reads the backup.* instance settings directly from the database (admin changes apply next run; local retention row overrides the env) and the app password from the secret store - after each successful set: bundle dump + files archive + manifest into ONE self-contained tar.gz, upload via WebDAV per schedule (off/daily/weekly; manual runs always upload), prune remote bundles — never the newest — and record the outcome in status.json; upload failures alert via a new backupUploadFailed mail (de+en) - command listener on backup_command (run / restore) with a serial queue against the nightly timer - restore orchestrator: restore-status.json → maintenance NOTIFY → grace → (remote: download + manifest-verify bundle) → terminate other DB connections → shared perform-restore path (same code as restore.sh) → final status + maintenance exit api: - MaintenanceGuard (global, registered before the setup gate): 503 maintenance_mode while restore-status says running; health endpoints and the new public GET /backup/restore-status stay exempt; a stale running state (crashed sidecar) unblocks after 30 min - MaintenanceStateService watches the file and restarts the api after a successful restore (fresh caches, migrate-on-start for older dumps); main.ts refuses to touch the database while a restore runs — a container restarting mid-restore must not race pg_restore with migrate deploy - worker sweeps (conversion, mail outbox, scheduler) catch transient database failures instead of dying on an unhandled rejection — the restore's connection termination crashed the api in verification - backup admin endpoints under /admin/system/backup: settings (live connection test before save, password write-only into the secret store), nextcloud/test, sets (local via the ro backups mount + remote via WebDAV), run + restore (type-to-confirm backstop, source validation) — commands travel as NOTIFY payloads; audit actions backup.settings_changed/run_triggered/restore_requested - readyz: new warning-level backup_remote check while a target is configured (26 h daily / 170 h weekly bound) collab: - maintenance listener: on enter, persist + close every live session and refuse new connections until exit (failsafe timeout 30 min) — no in-memory document may write pre-restore content back afterwards web: - Admin → System backup section: status card with remote facts and a "Back up now" button, the Nextcloud settings form with test button, and the restore picker (local + remote sets, type-to-confirm) - global maintenance screen: any 503 maintenance_mode flips the SPA to a status page polling the exempt endpoint, reloading when the instance returns Verified end-to-end against a live stack (fresh DB, native api + sidecar, fake WebDAV server): configure → test → manual backup → bundle upload → readyz/sets/status surfaces → remote restore with maintenance gate, marker rollback and api restart; suites: shared 21, backup 9, collab 11, api 58 files green, lint + i18n:check + typecheck clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
101 lines
4.1 KiB
TypeScript
101 lines
4.1 KiB
TypeScript
import { collabEnvSchema, parseEnv } from '@dorfteich/shared';
|
|
import { Client } from 'pg';
|
|
|
|
import { createAccessListener } from './access-listener.js';
|
|
import { createPool, pingDatabase } from './db.js';
|
|
import { createLogger } from './logger.js';
|
|
import { createMaintenanceListener, type MaintenanceListener } from './maintenance-listener.js';
|
|
import { PostgresPagePersistence } from './persistence.js';
|
|
import { createRestoreListener } from './restore-listener.js';
|
|
import { closeDocumentConnections, createCollabServer } from './server.js';
|
|
import { PostgresSessionRegistry } from './session-registry.js';
|
|
import { PostgresVersionStore } from './version-store.js';
|
|
|
|
/**
|
|
* Entry point of the collaboration server (ADR 0003). Validates the
|
|
* environment, opens a database pool for the health probe, and starts the
|
|
* Hocuspocus server. A clean shutdown closes both so the container stops fast.
|
|
*/
|
|
async function bootstrap(): Promise<void> {
|
|
// Validate configuration up front; crash with a readable list on problems.
|
|
const env = parseEnv(collabEnvSchema, process.env);
|
|
const logger = createLogger(env);
|
|
const pool = createPool(env.DATABASE_URL);
|
|
|
|
// Created before the server (the server's hooks call markOpen/markClosed);
|
|
// its heartbeat gets the server's open-document accessor at start() below,
|
|
// which sidesteps the mutual reference between the two.
|
|
const sessionRegistry = new PostgresSessionRegistry({ pool, logger });
|
|
const versionStore = new PostgresVersionStore({ pool, logger });
|
|
|
|
// Persists + closes all sessions before the backup sidecar replaces the
|
|
// database, and refuses new ones until the restore is over (issue #103).
|
|
// The close action is bound after the server exists (mutual reference,
|
|
// same trick as the session registry above).
|
|
let closeAllConnections = (): void => {};
|
|
const maintenanceListener: MaintenanceListener = createMaintenanceListener({
|
|
createClient: () => new Client({ connectionString: env.DATABASE_URL }),
|
|
closeAllConnections: () => closeAllConnections(),
|
|
logger,
|
|
});
|
|
|
|
const server = createCollabServer({
|
|
version: env.APP_VERSION,
|
|
logger,
|
|
tokenSecret: env.COLLAB_TOKEN_SECRET,
|
|
pingDatabase: () => pingDatabase(pool),
|
|
persistence: new PostgresPagePersistence(pool),
|
|
sessionRegistry,
|
|
versionStore,
|
|
isMaintenanceActive: () => maintenanceListener.isActive(),
|
|
});
|
|
closeAllConnections = () => {
|
|
for (const documentName of server.hocuspocus.documents.keys()) {
|
|
closeDocumentConnections(server.hocuspocus, documentName);
|
|
}
|
|
};
|
|
|
|
// Terminate live sessions when access to a pond is revoked (issue #39). The
|
|
// listener owns a dedicated connection because `LISTEN` is connection-bound
|
|
// and cannot be served from the pool.
|
|
const accessListener = createAccessListener({
|
|
createClient: () => new Client({ connectionString: env.DATABASE_URL }),
|
|
pool,
|
|
openDocumentNames: () => [...server.hocuspocus.documents.keys()],
|
|
closeConnections: (documentName) => closeDocumentConnections(server.hocuspocus, documentName),
|
|
logger,
|
|
});
|
|
|
|
// Applies api-requested restores to the live document (issue #42).
|
|
const restoreListener = createRestoreListener({
|
|
createClient: () => new Client({ connectionString: env.DATABASE_URL }),
|
|
pool,
|
|
openDirectConnection: (documentName) =>
|
|
server.hocuspocus.openDirectConnection(documentName, { userId: 'restore', mode: 'rw' }),
|
|
logger,
|
|
});
|
|
|
|
await server.listen(env.PORT);
|
|
await accessListener.start();
|
|
await restoreListener.start();
|
|
await maintenanceListener.start();
|
|
sessionRegistry.start(() => [...server.hocuspocus.documents.keys()]);
|
|
logger.info({ event: 'listen', port: env.PORT }, 'collaboration server listening');
|
|
|
|
const shutdown = (signal: NodeJS.Signals): void => {
|
|
logger.info({ event: 'shutdown', signal }, 'shutting down');
|
|
sessionRegistry.stop();
|
|
void Promise.allSettled([
|
|
accessListener.stop(),
|
|
restoreListener.stop(),
|
|
maintenanceListener.stop(),
|
|
server.destroy(),
|
|
pool.end(),
|
|
]).then(() => process.exit(0));
|
|
};
|
|
process.on('SIGTERM', () => shutdown('SIGTERM'));
|
|
process.on('SIGINT', () => shutdown('SIGINT'));
|
|
}
|
|
|
|
void bootstrap();
|