Skip to main content

Operations — upgrades, backups, health & troubleshooting

This page is for the deployment administrator, once Ouroboros is running. It covers:

  • upgrading to a new version;
  • backing up and restoring;
  • checking each service's health and reading its logs;
  • fixing the failures you are most likely to meet.

How the services fit together, and how to stand them up the first time, is in Deploying Ouroboros.

The images​

Ouroboros ships as four container images:

ImageWhat it runs
ouroboros-dbThe database migrations — run once per deploy, then exits.
ouroboros-restThe REST service: the API, sign-in, scheduling, the farm gateway.
ouroboros-engineThe engine: estimation and the work REST hands it. Internal only.
ouroboros-uiThe web app people use.

Each image is published with two tags:

  • latest — the newest build of that image. Use it to track the main line.
  • The commit it was built from — a full git commit hash, such as :9c7f54cf…. It never moves, so use it to pin a deployment, to roll back, and to record exactly what you ran.

Each image is rebuilt only when its own part of Ouroboros changes, so the four images usually carry different commit tags. When you pin, pin each image to its own tag.

Upgrading​

The database schema must be migrated before the services that read it start. Upgrade in this order:

  1. Back up the database — see Backups. Migrations move forward only.
  2. Pull the new images, for example with docker compose pull.
  3. Run the migrations. Start ouroboros-db with the database's address (OURO_DB_HOST, OURO_DB_USER, OURO_DB_PASSWORD, and OURO_DB_PORT and OURO_DB_NAME if they differ from 5432 and ouroboros). It applies any new migrations and exits. Check that it exited with status 0.
  4. Restart the engine, then REST, then the UI. REST needs the migrated database and the engine, and the UI talks only to REST.
  5. Check health — see Health checks.

Never run the migrations with the development seed in production. The image's default command runs the production configuration; the seed is a separate file you would have to ask for explicitly, and it creates demonstration workspaces, people and passwords.

Rolling back. There is no "down" migration. To go back to an earlier version after its migrations have run, restore the database backup you took in step 1, then run the earlier images. Running an older ouroboros-db against a newer database changes nothing — it does not undo the newer migrations — so older services would meet a schema they were not built for.

Backups​

note

What follows is guidance from how Ouroboros stores its data. A full production runbook — TLS, backups and upgrades on a single host — has not been written yet.

A complete backup is three things, and each is needed to restore:

WhatWhereWhy it matters
The PostgreSQL database ouroborosYour database serverEvery workspace's data, sealed credentials, audit trail and build logs.
OURO_VAULT_MASTER_KEYREST's environmentUnseals every stored credential. A database restored without it cannot open a single key, token or secret.
Artifact storageThe directory or S3 bucket in OURO_ARTIFACT_STOREFiles builds uploaded. Not in the database.

Backing up the database. Use PostgreSQL's own tools, for example a nightly pg_dump in custom format:

pg_dump --format=custom --file=ouroboros-$(date -u +%F).dump \
"postgresql://<db user>:<db password>@<db host>:5432/ouroboros"

Or use continuous archiving if you need point-in-time recovery. Store dumps away from the database host.

Keep the master key separately, with your other production secrets — never in the same place as the database dumps. Anyone holding both can read every stored credential. Losing the key loses every stored credential, and there is no recovery.

Restoring.

  1. Stop REST, the engine and the UI.
  2. Restore into an empty database, for example with pg_restore --dbname=<url> ouroboros-<date>.dump.
  3. Run ouroboros-db from the same version as the backup, or a newer one, to bring the schema up to date.
  4. Start the services with the same OURO_VAULT_MASTER_KEY the backup was taken with.

A restored backup brings back everything it held — including any workspace that has since been deleted and purged. See Deleting a workspace for what a purge does and does not reach.

Health checks​

Each image runs its own container health check. You can also call these addresses yourself:

ServiceAddressAnswers
ouroboros-restGET /health/live200 while the process is running.
ouroboros-restGET /health/ready200 when the database and the engine both answer; 503 when either does not, saying which.
ouroboros-engineGET /healthz200 with {"status":"ok"} while the process is running.
ouroboros-uiGET /A redirect to the app (307), which is how its container check knows it is serving.

ouroboros-db has no health check: it runs the migrations and exits, and its exit status is the answer.

The REST health addresses are at the root of the service, not under /api/v1, and need no sign-in. A failing readiness check names the dependency and the reason:

{
"status": "error",
"info": { "database": { "status": "up" } },
"error": { "engine": { "status": "down", "message": "GET /healthz failed (ECONNREFUSED)" } }
}

The dashboard's system card is drawn from the same readiness check, so anyone signed in sees when the engine or the database is down.

What readiness does not check. It asks the engine only whether it is running. A mismatched engine secret still reports the engine up — see Troubleshooting.

Logs​

Every service writes its log to standard output, so read it with your container runtime, for example docker compose logs rest.

  • ouroboros-rest writes one line per event. At start-up it prints its configuration with every secret replaced by [redacted]. A start-up failure names the variable at fault.
  • ouroboros-engine writes one JSON object per line, at the level set by OURO_LOG_LEVEL (info by default).
  • ouroboros-ui writes the web server's log.
  • ouroboros-db writes Flyway's report of the migrations it applied, or why it refused.

The services are written never to log a credential, a token or a key.

Troubleshooting​

SymptomCauseFix
ouroboros-db exits non-zero with Validate failed and a checksum mismatchA migration already applied to this database differs from the one in the image — usually an image from a different build than the database has run, or an edited migration.Run the ouroboros-db image matching what you deployed. Do not edit applied migrations, and do not "repair" the history to make the error go away. If you cannot tell which is right, restore your backup.
ouroboros-db warns that the schema has a version … that is newer than the latest available migration, and applies nothingThe image is older than the database. It does not undo anything.Run the ouroboros-db matching the services you deploy. To go back a version, restore a backup.
ouroboros-db exits with no database to migrateOURO_DB_HOST (and the user and password) are not set.Set them, or pass -url=, -user= and -password=.
REST exits at start-up naming a variableA required variable is missing or malformed.Fix it in REST's environment; look it up in the Configuration reference.
Things that need the engine fail with The engine is not available right now, while /health/ready says the engine is up; REST's log says OURO_ENGINE_SHARED_SECRET does not match the value ouroboros-engine holds; the engine logs rejected an internal request without a valid keyREST and the engine have different values for OURO_ENGINE_SHARED_SECRET.Set the same value on both and restart them.
/health/ready answers 503 naming the engineThe engine is not running, or REST cannot reach OURO_ENGINE_URL.Start the engine; check the address from REST's network.
/health/ready answers 503 naming the databaseREST cannot reach the database at OURO_DATABASE_URL.Check the database is running and the address and password are right.
After approving on GitHub, sign-in fails with The redirect_uri is not associated with this applicationThe OAuth app's callback does not match BETTER_AUTH_URL.See Sign-in & workspace settings.
A runner's install or enrollment fails, or a runner shows offlineThe machine cannot reach OURO_FARM_PUBLIC_URL, or a proxy in front of REST drops the runner's client certificate.Check the address from the machine, and that the proxy passes the certificate — see Build farm administration and The farm gateway.
Every stored key stops working after a restartOURO_VAULT_MASTER_KEY changed.Restore the previous value. A changed key cannot open what the old one sealed.
Everyone is signed out after a restartBETTER_AUTH_SECRET changed.Expected after a rotation; otherwise restore the previous value.