dorfteich/docs/architecture/deployment.md
Claude Fable 5 f0a82bad20
Some checks failed
CD / Build and push images (push) Successful in 3m16s
CI / Lint, typecheck, test (push) Successful in 3m5s
CD / Deploy to Test (push) Successful in 13s
CD / Smoke tests against Test (push) Failing after 3m35s
CD / Promote to Int (push) Has been skipped
CI / Auth e2e pack (push) Successful in 5m6s
CI / Import/export fidelity gate (push) Successful in 43s
CI / Build container images (push) Has been skipped
Add the first-run setup wizard API with env-backed secret store (#80)
When the api runs against a database without the setup.completedAt
marker, a global SetupGuard answers every non-exempt route with 503
setup_required; only /setup/*, health probes, and the session routes
stay reachable. The wizard steps (POST /setup/admin|instance|smtp|
registration|complete) write straight to their production homes; the
Site Admin step signs its creator in, later steps require that session.
Completing sets the marker and locks every step permanently (410, also
across restarts, and not reopenable via PATCH /admin/settings).

SMTP entered in the wizard is verified with a live delivery test first
(failure blocks the step with the transport error as detail) and then
persisted to the new env-backed secret store: a mode-600 dotenv file on
the new `secrets` volume (SECRETS_FILE). Explicit container env always
wins over the store; empty compose-passed strings count as unset. The
mail transport now resolves lazily through SmtpConfigService so wizard
changes apply without a restart.

SETUP_ADMIN_* env pre-seeds the whole wizard at boot for automated
deploys; a backfill migration marks instances that already have a Site
Admin as completed, and seed/vitest global-setup do the same for
fixture databases. The setup e2e suite provisions its own fresh
database (CREATE DATABASE + migrate deploy) per run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EwZ4jR4KFAPvpjWevfUGX1
2026-07-11 15:10:28 +02:00

117 lines
6.8 KiB
Markdown

# Deployment architecture
Four stages, one Compose definition. Foundational decisions: ADR 0014
(CI/CD), ADR 0015 (backup), kickoff topology decision (Dev local on the
developer's machine; Test, Int and — from M8 — Prod on the operator's
dedicated Hetzner server ONE (`one.101010.cloud`, 168.119.32.247), which
also hosts the Gitea instance and the CI runner; the architecture must
keep any later move cheap). DNS status: `*.dorfteich.online` and
`*.dorfteich.cloud` already point to ONE. (Until 2026-07-11 the stages
ran on a shared 4-GB VPS, `188.245.116.44`.)
## The Compose stack
Every stage (and every self-hosted instance) runs the same services:
| Service | Image | Notes |
| ----------- | ------------------------------------ | --------------------------------------------------------------- |
| `web` | `dorfteich-web` | nginx: SPA assets, fonts; SPA fallback routing |
| `api` | `dorfteich-api` | NestJS; runs `prisma migrate deploy` on start |
| `collab` | `dorfteich-collab` | Hocuspocus WebSocket server |
| `db` | `postgres:<pinned>` | volume `db-data` |
| `pandoc` | `pandoc/core:<pinned>` (server mode) | internal only |
| `gotenberg` | `gotenberg/gotenberg:<pinned>` | internal only |
| `backup` | `dorfteich-backup` | cron sidecar: pg_dump, volume archive, prune, mirror (ADR 0015) |
Volumes: `db-data`, `uploads` (uploads + installed plugins), `backups`.
Networks: `frontend` (reverse proxy ↔ web/api/collab) and `internal`
(api/collab ↔ db/pandoc/gotenberg); db and converters are never exposed.
Ingress is a host-level reverse proxy (existing Caddy/Traefik/nginx on the
host), routing:
```
/ → web
/api/ → api
/collab → collab (WebSocket upgrade required)
/media/ → api (permission-checked file streaming)
```
The proxy forwards `/collab` without stripping the prefix, so the collab
service answers its health probe at `/collab/healthz` externally and at
`/healthz` for the container-internal Docker healthcheck.
Self-hosters without a proxy can enable the optional `caddy` Compose profile
(bundled Caddy with automatic TLS).
## Stages
| Stage | Where | Domain | Purpose | Data |
| -------- | ----------------------------------------------------------------- | ---------------------- | --------------------------------------------------------------------------------------- | ------------------------------- |
| **Dev** | contributor machine (e.g. the operator's MacBook), Docker Desktop | `localhost` | feature work; hot reload via `compose.dev.yml` overlay (source mounts, vite dev server) | fixtures/seed script |
| **Test** | ONE `one.101010.cloud`, `/home/DOCKER/dorfteich-test/` | `test.dorfteich.cloud` | auto-deploy target of `main`; e2e suite runs here | reset-able; seeded |
| **Int** | ONE `one.101010.cloud`, `/home/DOCKER/dorfteich-int/` | `int.dorfteich.cloud` | stable preview; manual/exploratory testing; release candidates | persistent test data |
| **Prod** | ONE `one.101010.cloud`, `/home/DOCKER/dorfteich-prod/` (M8, #89) | `dorfteich.online` | public flagship instance | real data; full backup + mirror |
Stage layout follows the operator's Docker host convention:
compose file + `.env` under `/home/DOCKER/dorfteich-<stage>/`, bulk data
volumes under `/home/RAID/DOCKER/dorfteich-<stage>/` (bind-mounted).
**Prod relocation readiness** (kickoff requirement): all state lives in the
three volumes + `.env`; the documented move procedure is: stop stack →
final backup → restore backup set on the new host → switch DNS. The backup
sidecar's restore runbook doubles as the migration procedure, and Test
restore drills (ADR 0015) keep it honest.
## Configuration
- One `.env` per stage (never in git; `.env.example` in the repo documents
every variable): database credentials, `APP_BASE_URL`, collab token
signing key, SMTP settings, stage name shown in the UI for non-Prod.
- First-run **setup wizard** (kickoff decision, issue #80): when the API
starts against an empty database it exposes only `/setup` (create Site
Admin account, SMTP, instance name/locale, registration mode); the wizard
locks itself permanently after completion (steps answer 410, also across
restarts). `SETUP_ADMIN_*` in `.env` pre-seeds the whole wizard for
automated deploys; instances that predate the wizard are locked by a
backfill migration.
- Secrets entered in the wizard (the SMTP password) go to the **env-backed
secret store** — a mode-600 dotenv file on the `secrets` volume
(`SECRETS_FILE`, security.md §Secrets), never into the database. Explicit
container env always wins over the store, so operators can override a
broken wizard entry from the stage `.env`.
## Pipeline (ADR 0014, concrete)
```mermaid
flowchart LR
PR[PR: lint + typecheck + unit + build] -->|merge| M[main]
M --> B[build images :sha]
B --> DT[deploy Test]
DT --> E2E[Playwright e2e vs Test]
E2E -->|green| DI[deploy Int - tag int]
DI --> REL{manual: tag vX.Y.Z + approval}
REL --> DP[deploy Prod - semver tag]
```
- Deploy jobs SSH into the stage directory and run
`docker compose pull && docker compose up -d`; migrations apply on api
start. Rollback = re-deploy the previous tag (migrations must be
backward-compatible one release back — contributor rule for schema
stories).
- The e2e suite is the Int-promotion gate; flaky tests are defects.
- Release notes are generated from merged PR titles; releases with
data-affecting migrations are labeled `migration` and called out.
## Self-hosting distribution
- Published artifacts per release: versioned images in the Gitea registry
(mirrored to a public registry at first public release), a reference
`docker-compose.yml` + `.env.example`, and the install/update/backup
guide (`docs/self-hosting/`, written as part of the docs milestone).
- Minimum requirements: Docker + Compose, 2 GB RAM, a domain (TLS via own
proxy or the `caddy` profile). Setup = compose up + browser wizard.
- Updates: `docker compose pull && up -d` on a new semver tag; migrations
run automatically; the release notes flag anything manual. Downgrades are
supported one release back.