All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m45s
CD / Build and push images (push) Successful in 3m49s
CI / Build container images (push) Has been skipped
CD / Deploy to Test (push) Successful in 9s
CD / Smoke tests against Test (push) Successful in 1m18s
CD / Promote to Int (push) Successful in 11s
CI / Auth e2e pack (push) Successful in 5m35s
CI / Import/export fidelity gate (push) Successful in 47s
Off-host backups for every self-hoster, configured entirely in the admin UI — supersedes the host-specific mirror plan behind #84. shared: - webdav.ts (new package entry like token-crypto): minimal WebDAV client with basic auth — PROPFIND (tolerant multistatus parser), MKCOL, PUT (streamed), GET, DELETE; Nextcloud DAV path derived from the plain server URL, explicit DAV bases pass through - backup-status.ts: additive remote-upload status in status.json, the restore-status.json contract (running/succeeded/failed + staleness bound), the backup_command/backup_maintenance NOTIFY channels, and the one-bundle-per-set naming (dorfteich-backup-<id>.tar.gz) - backup-set.ts moved here from apps/backup (api lists local sets) backup sidecar: - reads the backup.* instance settings directly from the database (admin changes apply next run; local retention row overrides the env) and the app password from the secret store - after each successful set: bundle dump + files archive + manifest into ONE self-contained tar.gz, upload via WebDAV per schedule (off/daily/weekly; manual runs always upload), prune remote bundles — never the newest — and record the outcome in status.json; upload failures alert via a new backupUploadFailed mail (de+en) - command listener on backup_command (run / restore) with a serial queue against the nightly timer - restore orchestrator: restore-status.json → maintenance NOTIFY → grace → (remote: download + manifest-verify bundle) → terminate other DB connections → shared perform-restore path (same code as restore.sh) → final status + maintenance exit api: - MaintenanceGuard (global, registered before the setup gate): 503 maintenance_mode while restore-status says running; health endpoints and the new public GET /backup/restore-status stay exempt; a stale running state (crashed sidecar) unblocks after 30 min - MaintenanceStateService watches the file and restarts the api after a successful restore (fresh caches, migrate-on-start for older dumps); main.ts refuses to touch the database while a restore runs — a container restarting mid-restore must not race pg_restore with migrate deploy - worker sweeps (conversion, mail outbox, scheduler) catch transient database failures instead of dying on an unhandled rejection — the restore's connection termination crashed the api in verification - backup admin endpoints under /admin/system/backup: settings (live connection test before save, password write-only into the secret store), nextcloud/test, sets (local via the ro backups mount + remote via WebDAV), run + restore (type-to-confirm backstop, source validation) — commands travel as NOTIFY payloads; audit actions backup.settings_changed/run_triggered/restore_requested - readyz: new warning-level backup_remote check while a target is configured (26 h daily / 170 h weekly bound) collab: - maintenance listener: on enter, persist + close every live session and refuse new connections until exit (failsafe timeout 30 min) — no in-memory document may write pre-restore content back afterwards web: - Admin → System backup section: status card with remote facts and a "Back up now" button, the Nextcloud settings form with test button, and the restore picker (local + remote sets, type-to-confirm) - global maintenance screen: any 503 maintenance_mode flips the SPA to a status page polling the exempt endpoint, reloading when the instance returns Verified end-to-end against a live stack (fresh DB, native api + sidecar, fake WebDAV server): configure → test → manual backup → bundle upload → readyz/sets/status surfaces → remote restore with maintenance gate, marker rollback and api restart; suites: shared 21, backup 9, collab 11, api 58 files green, lint + i18n:check + typecheck clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
159 lines
5.7 KiB
TypeScript
159 lines
5.7 KiB
TypeScript
import { Injectable, OnModuleDestroy, OnModuleInit } from '@nestjs/common';
|
|
import { Prisma } from '@prisma/client';
|
|
import { PinoLogger } from 'nestjs-pino';
|
|
|
|
import { ClockService } from '../common/clock.service';
|
|
import { AppConfig } from '../config/app-config.service';
|
|
import { PrismaService } from '../prisma/prisma.service';
|
|
|
|
export interface JobDefinition {
|
|
/** Stable identity — also the `jobs` table primary key. */
|
|
name: string;
|
|
cadenceSeconds: number;
|
|
run: () => Promise<void>;
|
|
}
|
|
|
|
/** How often the scheduler checks which registered jobs are due. */
|
|
const TICK_MS = 60_000;
|
|
/** A job stuck `RUNNING` past this (crashed process, never released) is
|
|
* treated as available again rather than blocked forever. */
|
|
const STALE_LOCK_MS = 60 * 60_000;
|
|
|
|
/**
|
|
* Generic maintenance-job scheduler (issue #31; data-model.md/`jobs`,
|
|
* operations.md's job table). Every later maintenance job (version
|
|
* thinning, compaction, quota reconciliation, orphan sweep, …) registers
|
|
* here instead of growing its own timer loop.
|
|
*
|
|
* Due-ness and the run-mutex both live in the `jobs` table, not in
|
|
* memory: `lastRunAt` is what makes the schedule survive an api restart,
|
|
* and claiming a due job is a single atomic
|
|
* `UPDATE ... WHERE status != 'RUNNING'` — the same "row as mutex"
|
|
* technique works whether it's this process racing itself (two ticks
|
|
* overlapping because a job ran long) or, defensively, two processes.
|
|
*/
|
|
@Injectable()
|
|
export class SchedulerService implements OnModuleInit, OnModuleDestroy {
|
|
private readonly jobs = new Map<string, JobDefinition>();
|
|
private timer: NodeJS.Timeout | undefined;
|
|
|
|
constructor(
|
|
private readonly prisma: PrismaService,
|
|
private readonly clock: ClockService,
|
|
private readonly config: AppConfig,
|
|
private readonly logger: PinoLogger,
|
|
) {
|
|
this.logger.setContext(SchedulerService.name);
|
|
}
|
|
|
|
register(job: JobDefinition): void {
|
|
this.jobs.set(job.name, job);
|
|
}
|
|
|
|
/** Every job registered in this process (issue #86 admin panel). */
|
|
definitions(): JobDefinition[] {
|
|
return [...this.jobs.values()];
|
|
}
|
|
|
|
/**
|
|
* Manual trigger from the admin panel (issue #86): runs `name` now,
|
|
* regardless of cadence — only the run-mutex still applies, so a job
|
|
* already running elsewhere reports `already_running` instead of
|
|
* doubling up.
|
|
*/
|
|
async runNow(name: string): Promise<'succeeded' | 'failed' | 'already_running'> {
|
|
const job = this.jobs.get(name);
|
|
if (!job) throw new Error(`unknown job: ${name}`);
|
|
return this.claimAndRun(job, { force: true });
|
|
}
|
|
|
|
onModuleInit(): void {
|
|
if (this.config.env.NODE_ENV === 'test') return; // tests drive jobs directly
|
|
// A tick hitting a transient database failure (outage, or the backup
|
|
// sidecar terminating connections mid-restore, #103) must retry on the
|
|
// next tick, never crash the api via an unhandled rejection.
|
|
this.timer = setInterval(
|
|
() =>
|
|
void this.tick().catch((error: unknown) => {
|
|
this.logger.warn({ err: error }, 'scheduler tick failed; retrying on the next tick');
|
|
}),
|
|
TICK_MS,
|
|
);
|
|
this.timer.unref();
|
|
}
|
|
|
|
onModuleDestroy(): void {
|
|
if (this.timer) clearInterval(this.timer);
|
|
}
|
|
|
|
/** One scheduling pass over every registered job; public for tests/manual triggers. */
|
|
async tick(): Promise<void> {
|
|
for (const job of this.jobs.values()) {
|
|
await this.runIfDue(job);
|
|
}
|
|
}
|
|
|
|
/** Runs `job` now if due and not already running elsewhere; a no-op otherwise. */
|
|
async runIfDue(job: JobDefinition): Promise<void> {
|
|
await this.claimAndRun(job, { force: false });
|
|
}
|
|
|
|
private async claimAndRun(
|
|
job: JobDefinition,
|
|
{ force }: { force: boolean },
|
|
): Promise<'succeeded' | 'failed' | 'already_running'> {
|
|
try {
|
|
await this.prisma.job.upsert({
|
|
where: { name: job.name },
|
|
create: { name: job.name, cadenceSeconds: job.cadenceSeconds },
|
|
update: {},
|
|
});
|
|
} catch (error) {
|
|
// Two overlapping ticks can both try to create the row for a
|
|
// never-seen-before job at once; whichever loses just means the row
|
|
// already exists now, which is exactly what this call wants anyway.
|
|
const isDuplicate =
|
|
error instanceof Prisma.PrismaClientKnownRequestError && error.code === 'P2002';
|
|
if (!isDuplicate) throw error;
|
|
}
|
|
|
|
const now = this.clock.now();
|
|
const dueBefore = new Date(now.getTime() - job.cadenceSeconds * 1000);
|
|
const staleLockBefore = new Date(now.getTime() - STALE_LOCK_MS);
|
|
|
|
const claim = await this.prisma.job.updateMany({
|
|
where: {
|
|
name: job.name,
|
|
AND: [
|
|
// A manual trigger skips the due check, never the run-mutex.
|
|
...(force ? [] : [{ OR: [{ lastRunAt: null }, { lastRunAt: { lte: dueBefore } }] }]),
|
|
{ OR: [{ status: { not: 'RUNNING' } }, { lockedAt: { lte: staleLockBefore } }] },
|
|
],
|
|
},
|
|
data: { status: 'RUNNING', lockedAt: now, lastRunAt: now },
|
|
});
|
|
if (claim.count === 0) return 'already_running'; // or, unforced, simply not due
|
|
|
|
try {
|
|
await job.run();
|
|
await this.prisma.job.update({
|
|
where: { name: job.name },
|
|
data: { status: 'IDLE', lastError: null, lastDurationMs: this.sinceMs(now) },
|
|
});
|
|
return 'succeeded';
|
|
} catch (error) {
|
|
const message = error instanceof Error ? error.message.slice(0, 500) : String(error);
|
|
this.logger.error({ job: job.name, err: error }, 'maintenance job failed');
|
|
await this.prisma.job.update({
|
|
where: { name: job.name },
|
|
data: { status: 'FAILED', lastError: message, lastDurationMs: this.sinceMs(now) },
|
|
});
|
|
return 'failed';
|
|
}
|
|
}
|
|
|
|
private sinceMs(start: Date): number {
|
|
return Math.max(0, this.clock.now().getTime() - start.getTime());
|
|
}
|
|
}
|