dorfteich/docs/architecture/deployment.md
Claude Fable 5 a9e901c449
All checks were successful
Release / Build release images and notes (push) Successful in 1m8s
CD / Build and push images (push) Successful in 1m9s
CD / Deploy to Test (push) Successful in 9s
Release / Release-candidate operations QA (push) Successful in 41s
CD / Smoke tests against Test (push) Successful in 1m10s
CD / Promote to Int (push) Successful in 10s
Prod deploy / Deploy the released images to Prod (push) Successful in 15s
CI / Lint, typecheck, test (push) Successful in 3m35s
CI / Build container images (push) Has been skipped
CI / Auth e2e pack (push) Successful in 5m32s
CI / Import/export fidelity gate (push) Successful in 48s
Run operations QA on every release candidate before the prod gate (#90)
New deploy/release-qa.sh, wired as the release workflow's second job: it
boots the PREVIOUS release with pre-seeded fixture content in a scratch
environment, swaps the api to the candidate against the same database
(migrations auto-apply, readiness green, content intact — the
seed_fixture/assert_fixture pair is the update-fixture contract future
migrations extend), asserts the degraded-readyz semantics on the
candidate (200 + warn-level checks without sidecars), and runs a full
backup/restore roundtrip with the candidate's sidecar into a second,
empty database. The wizard e2e already guards fresh installs in CI
(issue #81). Verified green on the host for v0.1.0→v0.1.1; a simulated
destructive migration made the suite fail loudly (negative test,
not committed). A human pushes the prod tag only when both release jobs
are green — the documented pre-approval gate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
2026-07-12 00:29:53 +02:00

7.2 KiB

Deployment architecture

Four stages, one Compose definition. Foundational decisions: ADR 0014 (CI/CD), ADR 0015 (backup), kickoff topology decision (Dev local on the developer's machine; Test, Int and — from M8 — Prod on the operator's dedicated Hetzner server ONE (one.101010.cloud, 168.119.32.247), which also hosts the Gitea instance and the CI runner; the architecture must keep any later move cheap). DNS status: *.dorfteich.online and *.dorfteich.cloud already point to ONE. (Until 2026-07-11 the stages ran on a shared 4-GB VPS, 188.245.116.44.)

The Compose stack

Every stage (and every self-hosted instance) runs the same services:

Service Image Notes
web dorfteich-web nginx: SPA assets, fonts; SPA fallback routing
api dorfteich-api NestJS; runs prisma migrate deploy on start
collab dorfteich-collab Hocuspocus WebSocket server
db postgres:<pinned> volume db-data
pandoc pandoc/core:<pinned> (server mode) internal only
gotenberg gotenberg/gotenberg:<pinned> internal only
backup dorfteich-backup cron sidecar: pg_dump, volume archive, prune, mirror (ADR 0015)

Volumes: db-data, uploads, plugins (installed plugin bundles), secrets (wizard-written secret store), backups (restore sets + status.json). Networks: frontend (reverse proxy ↔ web/api/collab) and internal (api/collab ↔ db/pandoc/gotenberg); db and converters are never exposed.

Ingress is a host-level reverse proxy (existing Caddy/Traefik/nginx on the host), routing:

/            → web
/api/        → api
/collab      → collab   (WebSocket upgrade required)
/media/      → api      (permission-checked file streaming)

The proxy forwards /collab without stripping the prefix, so the collab service answers its health probe at /collab/healthz externally and at /healthz for the container-internal Docker healthcheck.

Self-hosters without a proxy can enable the optional caddy Compose profile (bundled Caddy with automatic TLS).

Stages

Stage Where Domain Purpose Data
Dev contributor machine (e.g. the operator's MacBook), Docker Desktop localhost feature work; hot reload via compose.dev.yml overlay (source mounts, vite dev server) fixtures/seed script
Test ONE one.101010.cloud, /home/DOCKER/dorfteich-test/ test.dorfteich.cloud auto-deploy target of main; e2e suite runs here reset-able; seeded
Int ONE one.101010.cloud, /home/DOCKER/dorfteich-int/ int.dorfteich.cloud stable preview; manual/exploratory testing; release candidates persistent test data
Prod ONE one.101010.cloud, /home/DOCKER/dorfteich-prod/ (M8, #89) dorfteich.online public flagship instance real data; full backup + mirror

Stage layout follows the operator's Docker host convention: compose file + .env under /home/DOCKER/dorfteich-<stage>/, bulk data volumes under /home/RAID/DOCKER/dorfteich-<stage>/ (bind-mounted).

Prod relocation readiness (kickoff requirement): all state lives in the three volumes + .env; the documented move procedure is: stop stack → final backup → restore backup set on the new host → switch DNS. The backup sidecar's restore runbook doubles as the migration procedure, and Test restore drills (ADR 0015) keep it honest.

Configuration

  • One .env per stage (never in git; .env.example in the repo documents every variable): database credentials, APP_BASE_URL, collab token signing key, SMTP settings, stage name shown in the UI for non-Prod.
  • First-run setup wizard (kickoff decision, issue #80): when the API starts against an empty database it exposes only /setup (create Site Admin account, SMTP, instance name/locale, registration mode); the wizard locks itself permanently after completion (steps answer 410, also across restarts). SETUP_ADMIN_* in .env pre-seeds the whole wizard for automated deploys; instances that predate the wizard are locked by a backfill migration.
  • Secrets entered in the wizard (the SMTP password) go to the env-backed secret store — a mode-600 dotenv file on the secrets volume (SECRETS_FILE, security.md §Secrets), never into the database. Explicit container env always wins over the store, so operators can override a broken wizard entry from the stage .env.

Pipeline (ADR 0014, concrete)

flowchart LR
    PR[PR: lint + typecheck + unit + build] -->|merge| M[main]
    M --> B[build images :sha]
    B --> DT[deploy Test]
    DT --> E2E[Playwright e2e vs Test]
    E2E -->|green| DI[deploy Int - tag int]
    DI --> REL{manual: tag vX.Y.Z + approval}
    REL --> DP[deploy Prod - semver tag]
  • Deploy jobs SSH into the stage directory and run docker compose pull && docker compose up -d; migrations apply on api start. Rollback = re-deploy the previous tag (migrations must be backward-compatible one release back — contributor rule for schema stories).
  • The e2e suite is the Int-promotion gate; flaky tests are defects.
  • Release notes are generated from merged PR titles; releases with data-affecting migrations are labeled migration and called out.

Self-hosting distribution

  • Published artifacts per release: versioned images in the Gitea registry (mirrored to a public registry at first public release), a reference docker-compose.yml + .env.example, and the install/update/backup guide (docs/self-hosting/, written as part of the docs milestone).
  • Minimum requirements: Docker + Compose, 2 GB RAM, a domain (TLS via own proxy or the caddy profile). Setup = compose up + browser wizard.
  • Updates: docker compose pull && up -d on a new semver tag; migrations run automatically; the release notes flag anything manual. Downgrades are supported one release back. Every release candidate passes the automated operations QA (deploy/release-qa.sh, issue #90: update simulation from the previous release, degraded-readyz semantics, backup/restore roundtrip) before the manual prod gate; new migrations add their own fixture to the script's seed_fixture/assert_fixture pair.