Skip to content

Health checks & readiness

A probe is only useful if something acts on its answer, and a deep probe can take every replica out at once. This page answers which probes exist and who reads them, what a probe may check, and how a deploy proves the new code is the code serving traffic.

Serve /health for liveness and /ready for routing

Section titled “Serve /health for liveness and /ready for routing”

Impact: HIGH pointing the load balancer at /ready is what makes draining work

  • The paths are exactly /health and /ready: no /healthz, no /livez, no version prefix. They describe the process, not the API.
  • /health reads nothing and stays 200 while draining. /ready reads only the draining Ref and returns 503 draining while Lifecycle.isDraining. Neither answers during boot, because the handler does not exist yet.
  • The load balancer’s target check uses /ready. A check that restarts the process (Kubernetes liveness and startup) uses /health. A platform with one check for both uses /ready, which fails only while the task is already stopping.

❌ Incorrect — the load balancer probes the liveness route, so it never sees a drain:

targetGroup:
healthCheck:
path: /health

✅ Correct — routing follows readiness; restarts follow liveness:

targetGroup:
healthCheck:
path: /ready
livenessProbe:
httpGet: { path: /health }

Source: notes/09-production-concerns/health-checks-and-readiness.md · Decision 1

Mount /health in every profile, /ready inside the container’s serve

Section titled “Mount /health in every profile, /ready inside the container’s serve”

Impact: MEDIUM a route outside serve registers on no served router

  • Both are raw routes in infra/health.ts, outside the HttpApi (see HTTP API).
  • livenessLayer is part of RoutesLayer, which every profile serves. readinessLayer sits inside the container graph’s serve, reading the Lifecycle that MainLayer provides. On a Worker, /ready is a 404.
  • The revision is read once, at layer build, from SERVICE_VERSION.

❌ Incorrect — the readiness route placed beside Lifecycle.layer, outside serve:

const MainLayer = Layer.mergeAll(Lifecycle.layer, Health.readinessLayer, ServerLayer)

✅ Correct — the readiness route reads the draining flag and nothing else:

export const readinessLayer = Layer.unwrap(Effect.gen(function* () {
const headers = yield* revisionHeader
const lifecycle = yield* Lifecycle.Lifecycle
return HttpRouter.add("GET", "/ready", Effect.map(lifecycle.isDraining, (draining) =>
draining
? HttpServerResponse.text("draining", { status: 503, headers })
: HttpServerResponse.text("OK", { headers })))
}))

Source: notes/09-production-concerns/health-checks-and-readiness.md · Decision 1, amended

Never let a probe read a shared dependency

Section titled “Never let a probe read a shared dependency”

Impact: HIGH a database check in /ready empties the load balancer on the first blip

  • Neither probe runs a query, pings the pool or calls another service. Boot already proved the database: the handler exists only after SqlLayer and requireApplied built.
  • A dependency outage after boot is a request failure (the error boundary’s 5xx and a failed span), not a probe failure.
  • The test for adding something to /ready: only a fault that can differ between replicas. A shared dependency fails on all of them at once, so there is nowhere to route to.
  • Accepted cost: a green probe does not mean the database is reachable now.

❌ Incorrect — a deep readiness check with N replicas on one database:

HttpRouter.add("GET", "/ready", sql`SELECT 1`.pipe(
Effect.as(HttpServerResponse.text("OK")),
Effect.orElseSucceed(() => HttpServerResponse.text("db down", { status: 503 })),
))

✅ Correct — the probe answers for the process; requests answer for the database:

HttpRouter.add("GET", "/health", HttpServerResponse.text("OK", { headers }))

Source: notes/09-production-concerns/health-checks-and-readiness.md · Decision 2

Answer OK publicly, with the revision in a header

Section titled “Answer OK publicly, with the revision in a header”

Impact: MEDIUM the revision is the one fact a deploy cannot learn about itself

  • Both probes are unauthenticated and reachable through the public URL. The smoke check needs the public URL, and the load balancer needs no credential.
  • The body is the text OK (or draining). No JSON, dependency list, uptime, pid, hostname, environment name or error text.
  • The only thing revealed is x-app-revision, the value of SERVICE_VERSION: the deployed commit SHA, dev locally. A header keeps the body OK.
  • Probes stay out of tracing via TracerDisabledWhen (see observability).

❌ Incorrect — a JSON status page that leaks internals:

{ "status": "ok", "db": "up", "uptime": 81234, "pid": 7, "host": "ip-10-0-3-17", "env": "prd" }

✅ Correct — one header, one word:

const revisionHeader = Effect.map(
Config.String("SERVICE_VERSION").pipe(Config.withDefault("dev")),
(revision) => ({ "x-app-revision": revision }),
)

Source: notes/09-production-concerns/health-checks-and-readiness.md · Decision 3

Make the smoke check assert the deployed revision

Section titled “Make the smoke check assert the deployed revision”

Impact: HIGH a 200 from the old revision is a failed deploy

  • The check passes only on 200 and the expected revision. On failure it names the revision it saw.
  • It runs even when alchemy deploy failed, as long as the deploy step ran, because that is when a partial deploy happens.
  • On a Worker there is no probe, restart or drain, so the smoke check is the only thing that asks. /health lives inside the graph, which is built in init, so an answer proves init succeeded.
  • The smoke check is shallow too: it proves the new code is serving and booted. See deployment and environments.

❌ Incorrect — any 200 passes, including the old code:

Terminal window
curl -fsS "https://$HOST/health"

✅ Correct — status and revision, retried, per profile:

deploy job, after `alchemy deploy --stage <stg|prd>`, on always()
container: GET https://<public host>/ready → 200 and x-app-revision == $SHA
Worker: GET https://<worker host>/health → 200 and x-app-revision == $SHA
retry: up to 6 attempts, 10s apart; on failure name the revision that is serving

Source: notes/09-production-concerns/health-checks-and-readiness.md · Decision 4

  • Liveness that reads a background loop’s last-progress time — trigger: the first incident where a worker loop stopped making progress while /health kept answering 200.
  • A /ready check for a replica-local dependency (a per-replica cache or sidecar that must be warm) — trigger: the first such dependency.
  • A smoke step that reads through the database — trigger: the first deploy that passed the revision check and broke a database-touching path.