← all notes

Building a publishing platform, Part 2: Zero-downtime deploys

In Part 1 I wrote about why the platform lives in a single Nx monorepo. This note covers the other end: how the whole stack ships to a single VPS with zero downtime, automatic TLS, and no manual SSH unless something goes wrong.

The moving parts

Four Docker containers on one server:

services:
  traefik:     # Reverse proxy + automatic TLS
  web:         # Angular 21 SSR frontend
  cms:         # Strapi 5 headless CMS
  postgres:    # PostgreSQL 17 database

Traefik sits in front, terminates TLS, and routes traffic to the web and cms containers based on the hostname. moen.ga hits the Angular SSR app. cms.moen.ga hits Strapi. Traefik handles Let's Encrypt certificates automatically through its own ACME challenge resolver. No certbot, no cron jobs, no manual renewals.

Why Traefik over nginx

I have done the nginx + certbot dance. It works, but it is ceremony. Traefik was built for this exact problem: it reads Docker labels, discovers services automatically, and provisions TLS certificates the moment a container starts. You declare routing in docker-compose.yml labels instead of a separate config file:

labels:
  - "traefik.enable=true"
  - "traefik.http.routers.web.rule=Host(`moen.ga`)"
  - "traefik.http.routers.web.entrypoints=websecure"
  - "traefik.http.routers.web.tls.certresolver=letsencrypt"

That is the entire reverse proxy config for the frontend. Add a new service, add labels, run docker compose up -d, and Traefik picks it up without a reload.

The deploy pipeline

A push to main triggers a GitHub Actions workflow with three stages:

Detect what changed. A dorny/paths-filter step checks whether files under apps/web/ or apps/cms/ were touched. If only the CMS changed, the frontend image is not rebuilt. If only shared libraries changed, both are rebuilt. This keeps CI fast and avoids shipping unchanged images.

Build and push. Each affected service gets its own Docker image built with BuildKit, tagged with both latest and a sha- commit hash, and pushed to GitHub Container Registry. Build cache is stored in GitHub Actions cache (type=gha), so subsequent builds only rebuild changed layers.

Deploy over SSH. The final job SSHes into the VPS and runs deploy.sh, which pulls the latest code, runs docker compose pull for the new images, and does a docker compose up -d. Docker Compose's rolling restart strategy means the old container keeps serving requests until the new one is healthy, then traffic switches over.

Health checks make zero-downtime real

Zero-downtime is not magic. It depends on Traefik knowing when a container is ready. Every service exposes a health endpoint, and Traefik polls it before routing traffic:

labels:
  - "traefik.http.services.cms.loadbalancer.healthCheck.path=/_health"
  - "traefik.http.services.cms.loadbalancer.healthCheck.interval=5s"
  - "traefik.http.services.cms.loadbalancer.healthCheck.timeout=3s"

PostgreSQL has its own health check through pg_isready, and both web and cms depend on it being healthy before they start. The chain is: Postgres is ready, then CMS starts and passes /_health, then Traefik routes to it. If the new container never becomes healthy, Traefik keeps routing to the old one. The site stays up.

Automated content pipelines

Beyond deploys, GitHub Actions also runs two scheduled scrapers: a news scraper at 01:00 UTC and an events scraper at 06:00 UTC. Both are Python scripts that fetch industry sources, summarize articles with Gemini, and push them to Strapi through its API. This means the site has fresh content every morning without anyone touching the CMS. The scrapers live in tools/agents/ and use Strapi API tokens stored as repository secrets.

Takeaways

Traefik + Docker Compose on a single VPS is enough for a platform this size. You do not need Kubernetes. Health checks on every service are what make zero-downtime actually work. And path-based CI detection saves more time than any caching strategy: if you did not change the frontend, do not build the frontend.

Next in the series: Part 3 covers designing Strapi content types for geospatial publishing: what to model as a collection type vs. a component, how taxonomy ties everything together, and why the schema decisions you make early are the ones you live with longest.