Upgrading spieli¶
Before you upgrade¶
Check GitHub Releases for the latest version and scan the release notes for breaking labels. Most releases carry none and require no manual action.
| Label | What it means | What you must do |
|---|---|---|
requires-env-update |
New or removed env var | Update .env before pulling the new image |
requires-compose-update |
compose.yml changed structurally |
Re-run install.sh or update compose.yml manually before restarting |
requires-reimport |
OSM data model changed | Trigger a full re-import after the image update lands |
requires-schema-update |
api.sql changed |
Covered automatically — no extra step needed |
To check what version is currently running:
docker inspect spieli-app-1 --format '{{index .Config.Labels "org.opencontainers.image.version"}}'
# or
docker compose images
Automatic upgrade (Watchtower)¶
If you installed with the auto-update profile active, Watchtower handles standard releases with no manual intervention:
- New image published → Watchtower pulls and restarts
appandimporter. - The daemon importer applies
api.sqlon every container startup, so schema changes take effect immediately after the restart — no separateAPI_ONLY=1step needed.
Standard releases (no breaking labels): nothing to do. Wait for Watchtower to pick up the new image (default poll interval: 5 minutes).
For breaking-label releases, follow the Breaking label procedures below in addition to the Watchtower restart.
One Watchtower covers all containers
Only the hub stack (or your single standalone stack) should run --profile auto-update. Adding a second Watchtower to any data-node stack causes both instances to fight over the same containers. The single Watchtower instance watches all containers on the Docker host automatically, including any data-node stacks that don't run auto-update themselves.
Manual upgrade¶
Use manual steps when you don't run Watchtower, or when a breaking label requires a specific action sequence.
Hub (DEPLOY_MODE=ui)¶
The hub has no database and no importer — only the app container needs updating.
Done. No API_ONLY=1, no importer, no verify step needed.
Standalone or Data-node (DEPLOY_MODE=data-node-ui or data-node)¶
Replace <mode> with your DEPLOY_MODE and <port> with your APP_PORT.
Step 1 — Pull new images
Step 2 — Stop the daemon importer
This is not optional, and it is not the same as "the daemon is idle". The daemon applies api.sql on container startup, so anything that restarts it — a reboot, Watchtower, up -d — makes it a second writer against the same schema. Stopping it removes that possibility for the duration of the apply.
Step 3 — Apply schema changes (API_ONLY)
This updates all PostgREST functions and the version number reported by get_meta. It runs as a one-shot container, and with the daemon stopped it is the only writer.
Concurrent applies queue rather than race
api.sql takes a session-level advisory lock, so if something applies the schema at the same time — a manual make db-apply, a Watchtower-triggered restart — the second one waits instead of corrupting the first.
That is a safety net, not a substitute for step 2. The last writer still wins, and the lock cannot tell an old image from a new one: if the daemon on its previous image applies last, the stack ends up serving the old schema while reporting healthy. Stopping it first is what prevents that.
If API_ONLY=1 fails mid-run
Since v0.10.0 the API stays up: playground_stats is built under a staging name and swapped in, so a crash before the swap leaves the previous view serving traffic. Re-run API_ONLY=1 after fixing the cause — the staging view is dropped and rebuilt on the next attempt. On v0.9.0 and earlier the view was dropped first, and a crash partway through left it gone, with PostgREST logging relation "public.playground_stats" does not exist; recovery there is a full re-import.
Step 4 — Recreate the daemon importer on the new image
--force-recreate matters. A plain up -d importer restarts the existing container, which leaves the daemon on the old image — and since its startup path applies api.sql, the old image then becomes the last writer and quietly restores the old schema. That is what happened during the v0.9.0 sweep (#800): the stack ran the new app against the old schema, with two filters silently wrong while get_meta and row counts looked healthy.
The API is unavailable for a few minutes after this step
The freshly recreated daemon applies api.sql again on startup. That drops and rebuilds playground_stats, which takes minutes on a large region, and terminates PostgREST's connections while it runs. During that window get_meta answers 5xx or reports a missing relation. This is expected, and it is why the verification below polls instead of being run once. Watch the importer log and wait for the apply to finish before concluding anything:
Step 5 — Restart the app container
Step 6 — Verify
curl -sf http://localhost:<port>/api/rpc/get_meta | \
python3 -c "import sys,json; d=json.load(sys.stdin); print('version:', d['version'], ' playgrounds:', d['playground_count'])"
Both version and playground_count should be non-zero. If playground_count is 0, see If playground_count is zero after upgrade below.
Retry it for a few minutes if it fails: the daemon's startup apply from step 4 has to finish first, and until it does the API is legitimately unavailable.
Verifying after step 4 is deliberate. Verifying before it checks a stack whose last writer may still be the old image, and the old schema is internally consistent enough to pass.
To confirm the schema really is the new one, compare the version get_meta reports against the release you just installed:
curl -sf http://localhost:<port>/api/rpc/get_meta | \
python3 -c "import sys,json; print(json.load(sys.stdin)['version'])"
version comes from the image that last applied api.sql, so it is the one field that distinguishes "the new schema is in place" from "an older image applied last". Do not check for the presence of a column instead: a column added in an earlier release is present in both schemas, so the check passes on the old one and gives false assurance in exactly the failure mode #800 is about.
Breaking label procedures¶
requires-env-update¶
- Read the release notes to find the new or removed variable.
- Update
.envaccordingly. - Then proceed with the normal upgrade steps for your deployment type.
requires-compose-update¶
compose.yml has structural changes (new service, renamed service, removed volume).
# Download the updated file:
curl -O https://raw.githubusercontent.com/mfuhrmann/spieli/main/compose.yml
# Recreate the stack:
docker compose --profile <mode> down
docker compose --profile <mode> up -d
Or re-run install.sh — it re-downloads and applies compose.yml automatically.
requires-reimport¶
A new OSM tag type was added. Existing data doesn't have the new columns until a fresh import.
After the image update (Watchtower or manual steps 1–2):
# Full re-import — also applies api.sql, so no separate API_ONLY step needed
docker compose --profile <mode> run --rm importer
This takes several minutes for large regions. The separate API_ONLY=1 step is not needed — a full re-import also updates all schema functions.
Single-host federation (multiple stacks on one VPS)¶
When the hub and its data-nodes share a single VPS, use scripts/upgrade-stacks.sh. It handles the correct order (data-nodes first, hub last) and runs each stack sequentially to avoid OOM from concurrent importer runs.
Copy the script to the VPS once, then edit the STACKS array to match your port layout:
STACKS=(
"$HOME/spieli-hessen:data-node-ui:8081"
"$HOME/spieli-berlin:data-node-ui:8082"
"$HOME/spieli:ui auto-update:8080" # hub last
)
Run:
Per stack, in this order — the order is the point, and it is not the one an earlier version of this page described:
- Pull images.
- Stop the daemon importer (data-nodes only), so it cannot race the schema apply. The manual procedure above calls this not optional; the script does it for you.
- Apply
api.sqlvia a one-shotAPI_ONLY=1importer, which never triggers a full reimport. If it fails, the daemon importer is brought back up before the sweep aborts, so a failed run never leaves a stack with its importer stopped. - Recreate the daemon importer with
up -d --force-recreate. The plain restart form is what left daemons running the old image during the v0.9.0 sweep. - Restart the app container — last, so it never serves against a schema the database does not have yet.
- Verify by polling
get_metafor up to 180 s, treating only200and404as settled. Verifying before step 4 is precisely what made #800 invisible: the old schema is internally consistent enough to pass.
The hub entry skips steps 2–4 automatically.
A stack that is mid-import is skipped, not forced. If get_meta reports
importing: true, the script leaves the schema alone and says so: stopping the
importer there would SIGKILL osm2pgsql after 10 s and leave the database
partially imported. The app is still upgraded, and the stack is listed
separately at the end — re-run the sweep for it once the import has finished.
The script does not touch
.envorcompose.yml. A release labelledrequires-env-updateorrequires-compose-updateneeds those applied to every stack before the sweep. It is easy to miss, because the script otherwise looks like the whole federation procedure.
Verification is not fatal to the sweep. By the time a stack is verified its images are already pulled and its containers restarted, so a failed check means the upgrade worked but the stack is not serving data yet. The script records the stack, carries on with the rest, and exits non-zero at the end listing every inconclusive one. A data-node that has never completed an import answers 404 on get_meta and is called out as such — expected until its first import finishes.
If the sweep does abort — a failed image pull or container start — it prints which stacks were upgraded, which one it stopped on, and that everything from there onward was not reached. Completed stacks are idempotent, so re-run after fixing the cause. One caveat it will warn you about: if a forced reimport was launched in the background before the abort, wait for it to finish first, or the next run's API_ONLY=1 step races it on the playground_stats rebuild.
Watchtower and the single-VPS setup
On a single-VPS federation, only the hub stack runs --profile auto-update. The single Watchtower instance restarts all containers on the host, including data-node containers — their stacks don't need their own auto-update profile.
Special cases¶
If playground_count is zero after upgrade¶
get_meta returns playground_count: 0 after an upgrade when playground_stats was left in a broken state (see the API_ONLY race warning above), or when the database volume was wiped.
Recovery: run a full re-import.
docker compose --profile <mode> run --rm \
-e REIMPORT_INTERVAL_MIN_DAYS= \
-e REIMPORT_INTERVAL_MAX_DAYS= \
importer
The empty REIMPORT_INTERVAL_* overrides force the importer to run immediately regardless of when data was last imported.
If you deleted the database volume¶
docker compose down -v or docker volume rm <stack>_pgdata wipes all imported data. Before re-importing:
-
Check whether the PBF cache volume survived:
-
If it exists, clear it — a previously interrupted import may have left a corrupt filtered PBF. The importer checks timestamps but not content; a corrupt cache causes it to process 0 objects and leave the database silently empty.
-
Run a full re-import:
-
Verify with
get_metaas shown in the manual upgrade section.
Downgrading¶
Supported only within the same minor version. Pin a specific version by editing compose.yml:
The database schema is not automatically rolled back. If the new version added columns or functions, the older app typically ignores them — but this is not tested. When in doubt, re-import from scratch.
See also¶
- Configuration reference — check for new or changed variables before upgrading
- Troubleshooting — common post-upgrade issues
- Single-host Federation — multi-stack setup and the upgrade script
- RELEASING.md — how releases are cut (maintainer reference)