Lifecycle & shutdown
The layer graph is the lifecycle. This page answers what order things close in, when a
process is ready, what happens between SIGTERM and exit, and how a boot failure is
reported. It covers long-lived processes on the container profile (Node, and Bun, whose
runMain behaves the same). On the Cloudflare Worker profile there is no process, no signal and
no graph teardown; per-request background work uses waitUntil, and work that must survive goes
through the outbox.
Let layer order be teardown order
Section titled “Let layer order be teardown order”Impact: HIGH the graph already computes the right order
- What a layer uses is provided beneath it, and is released after its user. A shared layer is refcounted and closes after its last user. No shutdown-phase registry, no sequential scope.
- Observability is the outermost thing in
bin.ts, so its exporter closes last. - Siblings in a
Layer.mergeAllclose in parallel and must not depend on each other’s close order. If one needs the other alive, it uses it as a service, which puts it beneath. - The one deliberate exception is
drainOnShutdown, which sits above the graph so its finalizer runs first (next rules).
❌ Incorrect — a hand-written shutdown sequence beside the graph:
process.on("SIGTERM", async () => { await server.close() await worker.stop() await pool.end() await tracer.flush()})✅ Correct — dependencies beneath their users; the graph closes in reverse:
// AppLayer = mergeAll(serve(routes + /ready), Outbox…, Jobs…) over ServicesLayer over SqlLayerMainLayer.pipe( Layer.provide(NodeHttpServer.layerConfig(() => createServer(), { /* … */ })), Layer.provide(Observability.layer), // outermost: closes after the graph Layer.provide(ConfigProviderLayer), Layer.provide(NodeServices.layer), Layer.launch, NodeRuntime.runMain,)Source: notes/09-production-concerns/lifecycle-and-shutdown.md · Decision 1
Treat “the handler answers” as ready; add only draining
Section titled “Treat “the handler answers” as ready; add only draining”Impact: MEDIUM a boot flag would duplicate what the graph already guarantees
HttpRouter.serveattaches the handler only after every layer has built, migrations check included. An answering readiness probe is proof that boot finished. Nobootingflag, no command gate.- Readiness (
Health.readinessLayerininfra/health.ts) readsLifecycle.isDrainingand nothing else: 503 while draining, 200 otherwise. Liveness never reads it. See health checks and readiness. - A
LayerMap/RcMapentry still building does not affect readiness. Startup work stays a layer. - During boot the port is open but nothing answers, so the platform needs a startup probe or an initial delay.
❌ Incorrect — a boot flag flipped by a last-built layer:
const booted = yield* Ref.make(false)// … some layer later: Ref.set(booted, true)// /ready: booted ? 200 : 503✅ Correct — one Ref, flipped only on the way down:
export class Lifecycle extends Context.Service<Lifecycle, { readonly isDraining: Effect.Effect<boolean> readonly setDraining: Effect.Effect<void>}>()("@app/infra/Lifecycle") {}
export const make = Effect.gen(function* () { const draining = yield* Ref.make(false) return Lifecycle.of({ isDraining: Ref.get(draining), setDraining: Ref.set(draining, true) })})Source: notes/09-production-concerns/lifecycle-and-shutdown.md · Decision 2, amended
On SIGTERM, keep serving for SHUTDOWN_DELAY, then drain HTTP for 10s
Section titled “On SIGTERM, keep serving for SHUTDOWN_DELAY, then drain HTTP for 10s”Impact: HIGH without the delay, every deploy refuses requests
- The server refuses new connections at the first instant of its finalizer, but load balancers
stop routing seconds later.
drainOnShutdownmarks the process draining and waits, still serving, so/ready503 is actually seen. drainOnShutdownis the outermost layer ofMainLayer. Beneathserve, the delay would run after the server had closed. A wrong placement still compiles, so review checks it.SHUTDOWN_DELAYis read throughLifecycle.configand has no default:"5 seconds"in every deployed environment,"0 seconds"in.env.exampleand tests. A stage that forgets it fails at boot (see config and secrets).gracefulShutdownTimeoutis 10s, set inbin.ts. No timeout of 0 and nodisablePreemptiveShutdown; a future streaming endpoint gets its own close path.
❌ Incorrect — the delay placed beneath serve, so the server is already closed:
export const MainLayer = AppLayer.pipe( Layer.provide(Lifecycle.drainOnShutdown), Layer.provide(Lifecycle.layer),)✅ Correct — outermost, so its finalizer runs first:
export const drainOnShutdown = Layer.effectDiscard(Effect.gen(function* () { const lifecycle = yield* Lifecycle const { shutdownDelay } = yield* config yield* Effect.addFinalizer(() => lifecycle.setDraining.pipe(Effect.andThen(Effect.sleep(shutdownDelay))))}))
export const MainLayer = Lifecycle.drainOnShutdown.pipe( Layer.provideMerge(AppLayer), Layer.provide(Lifecycle.layer),)Source: notes/09-production-concerns/lifecycle-and-shutdown.md · Decision 3, amended
End a worker’s queue and drain it for up to 5s
Section titled “End a worker’s queue and drain it for up to 5s”Impact: HIGH a deploy no longer drops the backlog silently
- Every worker drains on close. It is built into
makeWorkerininfra/worker.ts, never written per worker. Queue.end, notQueue.shutdown: end stops new offers and letstakefinish the buffer; shutdown throws it away.- The default deadline is 5s. Dropped items are counted in one Warn line, and an offer refused after shutdown Warns too. Nested workers drain one after another, each adding its deadline.
- The item cut at the deadline may run again, which consumer idempotency makes safe. Work that must survive is an outbox row, not a worker item.
❌ Incorrect — the fiber is interrupted and the buffer discarded:
const queue = yield* Queue.unbounded<A>()yield* Queue.take(queue).pipe(Effect.flatMap(process), Effect.forever, Effect.forkScoped)yield* Effect.addFinalizer(() => Queue.shutdown(queue))✅ Correct — end, drain to a deadline, count what is left:
const queue = yield* Queue.unbounded<A, Cause.Done>()const fiber = yield* Queue.take(queue).pipe(Effect.flatMap(process), Effect.forever, Effect.forkScoped)yield* Effect.addFinalizer(() => Queue.end(queue).pipe( Effect.andThen(Fiber.await(fiber).pipe(Effect.timeoutOption(options?.drainTimeout ?? "5 seconds"))), Effect.flatMap(Option.match({ onSome: () => Effect.void, onNone: () => Queue.size(queue).pipe(Effect.flatMap((dropped) => Effect.logWarning(`${name}: shutdown drain deadline`, { dropped }))), }))))Source: notes/09-production-concerns/lifecycle-and-shutdown.md · Decision 4, amended
Keep every stage inside the budget; bound a release that does I/O
Section titled “Keep every stage inside the budget; bound a release that does I/O”Impact: MEDIUM a stuck release would cost exactly the telemetry that explains it
| Stage after SIGTERM | Bound |
|---|---|
1. Pre-stop delay: /ready 503, requests still served |
5s (SHUTDOWN_DELAY) |
| 2. HTTP drain | 10s (gracefulShutdownTimeout) |
| 3. Worker drain | 5s per worker level |
| 4. Releases that do I/O | 2s each |
| 5. OTLP final export | 3s |
| Total | 25s, against a platform grace of at least 30s |
- The table is the contract. A change to any stage updates it, and the sum stays at least 5s below the platform’s grace period.
- A release we write that does I/O carries its own
Effect.timeoutOption(2s by default) and logs at Warn on timeout. A release never fails, and now it never hangs either. - No process-level watchdog. If something still hangs, the platform’s SIGKILL is the backstop.
❌ Incorrect — an unbounded release, plus a watchdog that skips the final export:
Effect.acquireRelease(spawnAgent(id), (proc) => killAgent(proc))setTimeout(() => process.exit(1), 25_000)✅ Correct — the release bounds itself and leaves a Warn:
Effect.acquireRelease(spawnAgent(id), (proc) => killAgent(proc).pipe( Effect.timeoutOption("2 seconds"), Effect.flatMap(Option.match({ onNone: () => Effect.logWarning("agents: kill timed out", { id }), onSome: () => Effect.void, })), Effect.ignoreCause({ log: "Warn", message: "agents: release failed" })))Source: notes/09-production-concerns/lifecycle-and-shutdown.md · Decision 5
Report a boot failure with runMain’s default until something alerts on it
Section titled “Report a boot failure with runMain’s default until something alerts on it”Impact: LOW the better log only pays once an alert reads it
- Day 1,
bin.tskeeps every provide beforeLayer.launch. A boot failure prints asrunMain’s multi-lineCause.pretty, with notrace_id, after observability has closed. That gap is accepted until it matters. - The written plan:
Observability.reportFailurelogs the cause once, as one JSON line withtrace_id, then fails withProcessFailed, whose[Runtime.errorReported] = falsestops a second log. Observability andConfigProviderLayermove toEffect.provideafterLayer.launch. - A clean SIGTERM exits 130. That is normal, not a failure.
❌ Incorrect — building the plan before anything reads its output:
MainLayer.pipe( Layer.launch, Observability.reportFailure, // no boot-failure alert exists yet Effect.provide(Observability.layer), NodeRuntime.runMain,)✅ Correct — day 1, every provide before Layer.launch:
MainLayer.pipe( Layer.provide(NodeHttpServer.layerConfig(() => createServer(), { /* … */ })), Layer.provide(Observability.layer), // still outermost of our layers Layer.provide(ConfigProviderLayer), // beneath Observability and the server, so both read .env Layer.provide(NodeServices.layer), Layer.launch, NodeRuntime.runMain,)Source: notes/09-production-concerns/lifecycle-and-shutdown.md · Decision 6, amended
Deferred
Section titled “Deferred”- A
bootingreadiness state — trigger: the first startup work that must finish after the graph is built (a forked warm-up). ProcessFailed+reportFailure, with Observability andConfigProviderLayermoved toEffect.provideafterLayer.launch— trigger: a boot-failure alert is wired.