dorfteich/deploy/stages.md
Claude Fable 5 db4f517e44
All checks were successful
CI / Lint, typecheck, test (pull_request) Successful in 5m32s
CI / Build container images (pull_request) Successful in 1m13s
CI / Auth e2e pack (pull_request) Successful in 8m22s
CI / Import/export fidelity gate (pull_request) Successful in 57s
CD / Build and push images (push) Successful in 16s
CD / Deploy to Test (push) Successful in 56s
CD / Smoke tests against Test (push) Successful in 1m24s
CD / Promote to Int (push) Successful in 52s
CI / Lint, typecheck, test (push) Successful in 5m36s
CI / Build container images (push) Has been skipped
CI / Auth e2e pack (push) Successful in 8m3s
CI / Import/export fidelity gate (push) Successful in 57s
#203: pin all third-party deploy images by digest
The four third-party images in the deploy compose (postgres, pandoc,
gotenberg — previously a floating MAJOR tag —, caddy) are now
name:tag@sha256 pins; the tag stays for readability, the digest decides
what runs. The pinned digests are exactly what the stages already run
(verified against the live containers' RepoDigests on ONE), so the next
recreation is byte-identical. A new early CI step fails on any
third-party compose image without a digest; compose.dev.yml is a local
convenience and deliberately exempt (its node helpers now follow the
#236 pin). Update + rollout procedure in deploy/stages.md — CD does not
sync stage composes, so the hand rollout to test/int/prod is part of
this issue's definition of done.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0168Ph5uBmHm8X28CSVpbpnJ
2026-07-31 05:13:13 +02:00

8.0 KiB

Stage provisioning on the dedicated host ONE (one.101010.cloud)

Test and Int run as Compose stacks on the operator's dedicated Hetzner server ONE (one.101010.cloud, 168.119.32.247) — the same host that runs the Gitea instance and the CI runner; DNS for *.dorfteich.cloud already points there (deployment.md §Stages). Steps marked [root] need host root access and are executed by the repo owner; everything else can be done by CI or a deploy user.

Overview

Stage Directory Domain Ports (localhost)
Test /srv/DOCKER/dorfteich-test/ test.dorfteich.cloud web 8100, api 8101, collab 8102
Int /srv/DOCKER/dorfteich-int/ int.dorfteich.cloud web 8110, api 8111, collab 8112

1. Stage directories [root]

for stage in test int; do
  mkdir -p /srv/DOCKER/dorfteich-$stage
  mkdir -p /home/RAID/DOCKER/dorfteich-$stage   # bulk data, if RAID exists on this host
done

Copy deploy/compose/docker-compose.yml and .env.example.env into each stage directory. Set per stage in .env (mode 600):

  • POSTGRES_PASSWORD: unique random value per stage
  • COLLAB_TOKEN_SECRET: long random value per stage (openssl rand -base64 32); signs/verifies the collaboration tokens (issue #34). The api and collab services read the same value from this one variable.
  • COMPOSE_PROJECT_NAME: dorfteich-test / dorfteich-int
  • WEB_PORT/API_PORT/COLLAB_PORT: 8100/8101/8102 (test), 8110/8111/8112 (int). COLLAB_PORT must be set per stage — both stacks share this host, so the compose default (8102) would make the Int collab container collide with Test's; Int needs COLLAB_PORT=8112.
  • IMAGE_PREFIX=gitea.101010.cloud/stwaidele/dorfteich
  • TAG: managed by the CD pipeline (<git-sha> on test, int on int)
  • APP_BASE_URL: https://test.dorfteich.cloud / https://int.dorfteich.cloud — e-mail links and the CSRF origin check depend on it
  • SMTP_HOST/SMTP_PORT/SMTP_SECURE/SMTP_USER/SMTP_PASS/SMTP_FROM: real relay credentials (both non-prod stages share one mailbox); without them, signup/reset mails queue up and fail

Fixture accounts on the stages are created with the regular seed, but with stage-specific passwords (never the public dev password). Always pass both override variables when re-seeding a stage — the seed re-hashes credentials on every run, so omitting them silently resets the stage accounts to the public dev password:

FIXTURE_ADMIN_PASSWORD=FIXTURE_USER_PASSWORD=\
  DATABASE_URL=postgresql://dorfteich:…@localhost:<tunnel-port>/dorfteich \
  pnpm --filter @dorfteich/api db:seed

(The stage db is not published; tunnel to the db container, e.g. ssh -L 15432:<db-container-ip>:5432 root@one.101010.cloud.)

2. Reverse proxy vhosts [root]

Both vhosts terminate TLS and route by path; WebSocket upgrade on /collab is required from milestone M3 on, configure it now. Caddy example:

test.dorfteich.cloud {
    handle /api/* {
        reverse_proxy 127.0.0.1:8101
    }
    handle /collab* {
        reverse_proxy 127.0.0.1:8102   # Hocuspocus collab (issue #33); Caddy passes the WebSocket upgrade through automatically
    }
    handle {
        reverse_proxy 127.0.0.1:8100
    }
}

(nginx equivalent: proxy_pass per location; for /collab add proxy_set_header Upgrade $http_upgrade; proxy_set_header Connection "upgrade";.)

Int: same block with int.dorfteich.cloud and ports 8110/8111/8112.

3. Gitea act_runner [root]

The CI/CD workflows (.gitea/workflows/) need one act_runner on the host with Docker access and the ubuntu-latest label:

# 1. Download act_runner (https://gitea.com/gitea/act_runner/releases)
# 2. Registration token: Gitea → Site/Repo Settings → Actions → Runners
#    (or: docker exec -u git gitea_app gitea actions generate-runner-token)
act_runner register \
  --instance https://gitea.101010.cloud \
  --token <REGISTRATION_TOKEN> \
  --name one-dorfteich \
  --labels ubuntu-latest:docker://docker.gitea.com/runner-images:ubuntu-latest
# 3. Run as a systemd service (act_runner daemon), user in the docker group.

4. Deploy user and SSH keys

The CD workflow (issue #8) deploys via SSH: ssh deploy@one.101010.cloud 'cd /srv/DOCKER/dorfteich-test && docker compose pull && docker compose up -d'.

  • [root] Create a deploy user (or reuse an existing deployment user), member of the docker group, owning the stage directories.
  • Generate one ed25519 keypair per stage; public keys into deploy's authorized_keys (optionally with a command= restriction to the compose command), private keys become the repository secrets DEPLOY_SSH_KEY_TEST / DEPLOY_SSH_KEY_INT.

5. Registry access

The pipeline pushes images to the Gitea container registry (gitea.101010.cloud/stwaidele/dorfteich-{web,api,collab}):

  • Repository secret REGISTRY_TOKEN: a Gitea access token with write:package scope (owner stwaidele or a CI account).
  • On the host, docker login gitea.101010.cloud for the deploy user with a read:package token, so compose pull works.

5a. Third-party image digests (issue #203, ADR 0024)

Every third-party image in deploy/compose/docker-compose.yml is pinned as name:tag@sha256:… — the tag stays for readability, the digest decides what runs, so the deployed artefact is exactly the reviewed one. An early CI step fails on any third-party image: reference without a digest. Our own images are pinned per release by the deploy pipeline (TAG in the stage .env); compose.dev.yml is a local convenience and deliberately not digest-pinned.

Updating a digest (e.g. to take a rebased base image or a new tag):

  1. Resolve the new digest — this prints the manifest-list digest every platform pulls:

    docker buildx imagetools inspect <name:tag>   # → Digest: sha256:…
    
  2. Update the reference in deploy/compose/docker-compose.yml to <name:tag>@sha256:… and let CI confirm.

  3. Roll out by hand: CD does NOT sync stage composes — apply the same change to /srv/DOCKER/dorfteich-{test,int,prod}/docker-compose.yml on ONE. The next compose pull && up -d (any CD run for test/int, the next release deploy for prod) recreates the containers from the pinned digest.

  4. Verify after rollout: docker inspect --format '{{.Image}}' <container> must print the pinned digest (or check RepoDigests on the image).

6. Verification checklist

  • https://test.dorfteich.cloud/healthzok
  • https://test.dorfteich.cloud/api/v1/readyz{"status":"ok",…}
  • https://test.dorfteich.cloud/collab/healthz{"status":"ok","service":"collab",…}
  • same for int
  • runner shows online under Gitea → Settings → Actions → Runners
  • a test workflow run executes on the runner
  • .env files are mode 600, owned by deploy

Provisioning log

  • 2026-07-05: Test/Int stage directories, .env files, Caddy vhosts (TLS live), deploy user, act_runner (v0.6.1, systemd) and registry login provisioned on the shared 4-GB VPS (188.245.116.44); Gitea Actions enabled instance-wide (app.ini on BASEL, backup kept). First pipeline run = this commit.
  • 2026-07-11: Everything moved to the dedicated host ONE (one.101010.cloud, 168.119.32.247) after Gitea itself relocated there: stage volumes (db-data, uploads) and .env files copied 1:1, compose files refreshed from the repo (now includes the plugins volume from #71), Caddy vhosts recreated (prod block prepared but commented out), act_runner one-dorfteich registered, old VPS runner and stacks stopped (kept as rollback reserve). DEPLOY_HOST in cd.yml and the DEPLOY_HOST_KEY secret updated accordingly.
  • 2026-07-11 (later): act_runner capacity raised from 1 to 4 in /home/deploy/act_runner/config.yaml — ONE has the headroom, so the independent CI jobs now run in parallel (the CD chain stays sequential via needs:).