dorfteich/apps/api/src/admin/system-admin.service.ts
Claude Fable 5 5cef359b8f
All checks were successful
CI / Lint, typecheck, test (push) Successful in 3m45s
CD / Build and push images (push) Successful in 3m49s
CI / Build container images (push) Has been skipped
CD / Deploy to Test (push) Successful in 9s
CD / Smoke tests against Test (push) Successful in 1m18s
CD / Promote to Int (push) Successful in 11s
CI / Auth e2e pack (push) Successful in 5m35s
CI / Import/export fidelity gate (push) Successful in 47s
Nextcloud backup target: admin-configured, manual + scheduled uploads, in-app restore (#103)
Off-host backups for every self-hoster, configured entirely in the admin
UI — supersedes the host-specific mirror plan behind #84.

shared:
- webdav.ts (new package entry like token-crypto): minimal WebDAV client
  with basic auth — PROPFIND (tolerant multistatus parser), MKCOL, PUT
  (streamed), GET, DELETE; Nextcloud DAV path derived from the plain
  server URL, explicit DAV bases pass through
- backup-status.ts: additive remote-upload status in status.json, the
  restore-status.json contract (running/succeeded/failed + staleness
  bound), the backup_command/backup_maintenance NOTIFY channels, and the
  one-bundle-per-set naming (dorfteich-backup-<id>.tar.gz)
- backup-set.ts moved here from apps/backup (api lists local sets)

backup sidecar:
- reads the backup.* instance settings directly from the database (admin
  changes apply next run; local retention row overrides the env) and the
  app password from the secret store
- after each successful set: bundle dump + files archive + manifest into
  ONE self-contained tar.gz, upload via WebDAV per schedule
  (off/daily/weekly; manual runs always upload), prune remote bundles —
  never the newest — and record the outcome in status.json; upload
  failures alert via a new backupUploadFailed mail (de+en)
- command listener on backup_command (run / restore) with a serial queue
  against the nightly timer
- restore orchestrator: restore-status.json → maintenance NOTIFY →
  grace → (remote: download + manifest-verify bundle) → terminate other
  DB connections → shared perform-restore path (same code as restore.sh)
  → final status + maintenance exit

api:
- MaintenanceGuard (global, registered before the setup gate): 503
  maintenance_mode while restore-status says running; health endpoints
  and the new public GET /backup/restore-status stay exempt; a stale
  running state (crashed sidecar) unblocks after 30 min
- MaintenanceStateService watches the file and restarts the api after a
  successful restore (fresh caches, migrate-on-start for older dumps);
  main.ts refuses to touch the database while a restore runs — a
  container restarting mid-restore must not race pg_restore with
  migrate deploy
- worker sweeps (conversion, mail outbox, scheduler) catch transient
  database failures instead of dying on an unhandled rejection — the
  restore's connection termination crashed the api in verification
- backup admin endpoints under /admin/system/backup: settings (live
  connection test before save, password write-only into the secret
  store), nextcloud/test, sets (local via the ro backups mount + remote
  via WebDAV), run + restore (type-to-confirm backstop, source
  validation) — commands travel as NOTIFY payloads; audit actions
  backup.settings_changed/run_triggered/restore_requested
- readyz: new warning-level backup_remote check while a target is
  configured (26 h daily / 170 h weekly bound)

collab:
- maintenance listener: on enter, persist + close every live session and
  refuse new connections until exit (failsafe timeout 30 min) — no
  in-memory document may write pre-restore content back afterwards

web:
- Admin → System backup section: status card with remote facts and a
  "Back up now" button, the Nextcloud settings form with test button,
  and the restore picker (local + remote sets, type-to-confirm)
- global maintenance screen: any 503 maintenance_mode flips the SPA to a
  status page polling the exempt endpoint, reloading when the instance
  returns

Verified end-to-end against a live stack (fresh DB, native api + sidecar,
fake WebDAV server): configure → test → manual backup → bundle upload →
readyz/sets/status surfaces → remote restore with maintenance gate,
marker rollback and api restart; suites: shared 21, backup 9, collab 11,
api 58 files green, lint + i18n:check + typecheck clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
2026-07-12 10:39:18 +02:00

183 lines
6.5 KiB
TypeScript

import { existsSync, readFileSync } from 'node:fs';
import { join } from 'node:path';
import { Injectable, NotFoundException } from '@nestjs/common';
import { Prisma, User } from '@prisma/client';
import {
AUDIT_PAGE_SIZE,
BACKUP_FRESH_MAX_AGE_HOURS,
BACKUP_STATUS_FILE,
type AuditListQuery,
type AuditListView,
type BackupStatus,
type JobTriggerResult,
type StorageOverviewView,
type SystemBackupView,
type SystemJobView,
} from '@dorfteich/shared';
import { AuditService } from '../audit/audit.service';
import { BackupTargetService } from '../backup/backup-target.service';
import { MaintenanceStateService } from '../backup/maintenance-state.service';
import { AppConfig } from '../config/app-config.service';
import { PrismaService } from '../prisma/prisma.service';
import { SchedulerService } from '../scheduler/scheduler.service';
/**
* Data behind the Site-Admin "System" panel (issue #86): maintenance jobs
* with truthful last-run data, the backup card mirroring the sidecar's
* status.json, the persistent audit trail, and the per-pond storage top
* list. Reads only — the single write path is the manual job trigger,
* which is itself audit-logged.
*/
@Injectable()
export class SystemAdminService {
constructor(
private readonly prisma: PrismaService,
private readonly scheduler: SchedulerService,
private readonly audit: AuditService,
private readonly config: AppConfig,
private readonly backupTarget: BackupTargetService,
private readonly maintenance: MaintenanceStateService,
) {}
/**
* Registered jobs merged with their `jobs` rows: a job that never ran yet
* appears with null last-run data, and a leftover row whose job no longer
* registers in this build is flagged instead of hidden.
*/
async jobs(): Promise<SystemJobView[]> {
const definitions = this.scheduler.definitions();
const rows = await this.prisma.job.findMany();
const rowsByName = new Map(rows.map((row) => [row.name, row]));
const views: SystemJobView[] = definitions.map((definition) => {
const row = rowsByName.get(definition.name);
rowsByName.delete(definition.name);
return {
name: definition.name,
cadenceSeconds: definition.cadenceSeconds,
status: row?.status ?? 'IDLE',
lastRunAt: row?.lastRunAt?.toISOString() ?? null,
lastDurationMs: row?.lastDurationMs ?? null,
lastError: row?.lastError ?? null,
registered: true,
} satisfies SystemJobView;
});
for (const row of rowsByName.values()) {
views.push({
name: row.name,
cadenceSeconds: row.cadenceSeconds,
status: row.status,
lastRunAt: row.lastRunAt?.toISOString() ?? null,
lastDurationMs: row.lastDurationMs ?? null,
lastError: row.lastError ?? null,
registered: false,
});
}
return views.sort((a, b) => a.name.localeCompare(b.name));
}
async triggerJob(actor: User, name: string): Promise<JobTriggerResult> {
if (!this.scheduler.definitions().some((job) => job.name === name)) {
throw new NotFoundException();
}
const outcome = await this.scheduler.runNow(name);
await this.audit.record({
action: 'job.triggered',
actorId: actor.id,
targetType: 'job',
targetId: name,
details: { outcome },
});
const job = (await this.jobs()).find((view) => view.name === name)!;
return { outcome, job };
}
/** The backup card mirrors status.json including the freshness verdict,
* plus the off-host target state and restore progress (issue #103). */
async backup(): Promise<SystemBackupView> {
const path = join(this.config.env.BACKUPS_DIR, BACKUP_STATUS_FILE);
let status: BackupStatus | null = null;
if (existsSync(path)) {
try {
status = JSON.parse(readFileSync(path, 'utf8')) as BackupStatus;
} catch {
status = null;
}
}
const finishedAt = status?.lastSuccess?.finishedAt;
const ageHours = finishedAt
? (Date.now() - new Date(finishedAt).getTime()) / 3_600_000
: Number.POSITIVE_INFINITY;
return {
available: status !== null,
fresh: Number.isFinite(ageHours) && ageHours <= BACKUP_FRESH_MAX_AGE_HOURS,
status,
maxAgeHours: BACKUP_FRESH_MAX_AGE_HOURS,
remoteConfigured: (await this.backupTarget.resolveTarget()) !== null,
restore: this.maintenance.current(),
};
}
async auditLog(query: AuditListQuery): Promise<AuditListView> {
const where: Prisma.AuditEntryWhereInput = {};
if (query.actor) {
const actor = await this.prisma.user.findUnique({ where: { username: query.actor } });
// An unknown username matches nothing rather than everything.
where.actorId = actor?.id ?? '00000000-0000-0000-0000-000000000000';
}
if (query.action) where.action = { startsWith: query.action };
if (query.from || query.to) {
where.at = {
...(query.from ? { gte: query.from } : {}),
...(query.to ? { lte: query.to } : {}),
};
}
const total = await this.prisma.auditEntry.count({ where });
const pageCount = Math.max(1, Math.ceil(total / AUDIT_PAGE_SIZE));
const page = Math.min(query.page, pageCount);
const entries = await this.prisma.auditEntry.findMany({
where,
orderBy: { at: 'desc' },
skip: (page - 1) * AUDIT_PAGE_SIZE,
take: AUDIT_PAGE_SIZE,
include: { actor: { select: { id: true, username: true, displayName: true } } },
});
return {
entries: entries.map((entry) => ({
id: entry.id,
at: entry.at.toISOString(),
action: entry.action,
actor: entry.actor,
targetType: entry.targetType,
targetId: entry.targetId,
details: (entry.details as Record<string, unknown> | null) ?? null,
})),
page,
pageCount,
total,
};
}
async storage(): Promise<StorageOverviewView> {
const usages = await this.prisma.pondUsage.findMany({
where: { pond: { deletedAt: null } },
orderBy: { storageBytesUsed: 'desc' },
take: 20,
include: { pond: { select: { name: true, slug: true, type: true } } },
});
const totals = await this.prisma.pondUsage.aggregate({ _sum: { storageBytesUsed: true } });
return {
totalBytes: Number(totals._sum.storageBytesUsed ?? 0n),
ponds: usages.map((usage) => ({
pondId: usage.pondId,
name: usage.pond.name,
slug: usage.pond.slug,
type: usage.pond.type === 'PERSONAL' ? 'personal' : 'shared',
storageBytesUsed: Number(usage.storageBytesUsed),
})),
};
}
}