Skip to content

Lifecycle & shutdown

The layer graph is the lifecycle. This page answers what order things close in, when a process is ready, what happens between SIGTERM and exit, and how a boot failure is reported. It covers long-lived processes on the container profile (Node, and Bun, whose runMain behaves the same). On the Cloudflare Worker profile there is no process, no signal and no graph teardown; per-request background work uses waitUntil, and work that must survive goes through the outbox.

Impact: HIGH the graph already computes the right order

  • What a layer uses is provided beneath it, and is released after its user. A shared layer is refcounted and closes after its last user. No shutdown-phase registry, no sequential scope.
  • Observability is the outermost thing in bin.ts, so its exporter closes last.
  • Siblings in a Layer.mergeAll close in parallel and must not depend on each other’s close order. If one needs the other alive, it uses it as a service, which puts it beneath.
  • The one deliberate exception is drainOnShutdown, which sits above the graph so its finalizer runs first (next rules).

❌ Incorrect — a hand-written shutdown sequence beside the graph:

process.on("SIGTERM", async () => {
await server.close()
await worker.stop()
await pool.end()
await tracer.flush()
})

✅ Correct — dependencies beneath their users; the graph closes in reverse:

// AppLayer = mergeAll(serve(routes + /ready), Outbox…, Jobs…) over ServicesLayer over SqlLayer
MainLayer.pipe(
Layer.provide(NodeHttpServer.layerConfig(() => createServer(), { /* … */ })),
Layer.provide(Observability.layer), // outermost: closes after the graph
Layer.provide(ConfigProviderLayer),
Layer.provide(NodeServices.layer),
Layer.launch,
NodeRuntime.runMain,
)

Source: notes/09-production-concerns/lifecycle-and-shutdown.md · Decision 1

Treat “the handler answers” as ready; add only draining

Section titled “Treat “the handler answers” as ready; add only draining”

Impact: MEDIUM a boot flag would duplicate what the graph already guarantees

  • HttpRouter.serve attaches the handler only after every layer has built, migrations check included. An answering readiness probe is proof that boot finished. No booting flag, no command gate.
  • Readiness (Health.readinessLayer in infra/health.ts) reads Lifecycle.isDraining and nothing else: 503 while draining, 200 otherwise. Liveness never reads it. See health checks and readiness.
  • A LayerMap / RcMap entry still building does not affect readiness. Startup work stays a layer.
  • During boot the port is open but nothing answers, so the platform needs a startup probe or an initial delay.

❌ Incorrect — a boot flag flipped by a last-built layer:

const booted = yield* Ref.make(false)
// … some layer later: Ref.set(booted, true)
// /ready: booted ? 200 : 503

✅ Correct — one Ref, flipped only on the way down:

infra/lifecycle.ts
export class Lifecycle extends Context.Service<Lifecycle, {
readonly isDraining: Effect.Effect<boolean>
readonly setDraining: Effect.Effect<void>
}>()("@app/infra/Lifecycle") {}
export const make = Effect.gen(function* () {
const draining = yield* Ref.make(false)
return Lifecycle.of({ isDraining: Ref.get(draining), setDraining: Ref.set(draining, true) })
})

Source: notes/09-production-concerns/lifecycle-and-shutdown.md · Decision 2, amended

On SIGTERM, keep serving for SHUTDOWN_DELAY, then drain HTTP for 10s

Section titled “On SIGTERM, keep serving for SHUTDOWN_DELAY, then drain HTTP for 10s”

Impact: HIGH without the delay, every deploy refuses requests

  • The server refuses new connections at the first instant of its finalizer, but load balancers stop routing seconds later. drainOnShutdown marks the process draining and waits, still serving, so /ready 503 is actually seen.
  • drainOnShutdown is the outermost layer of MainLayer. Beneath serve, the delay would run after the server had closed. A wrong placement still compiles, so review checks it.
  • SHUTDOWN_DELAY is read through Lifecycle.config and has no default: "5 seconds" in every deployed environment, "0 seconds" in .env.example and tests. A stage that forgets it fails at boot (see config and secrets).
  • gracefulShutdownTimeout is 10s, set in bin.ts. No timeout of 0 and no disablePreemptiveShutdown; a future streaming endpoint gets its own close path.

❌ Incorrect — the delay placed beneath serve, so the server is already closed:

export const MainLayer = AppLayer.pipe(
Layer.provide(Lifecycle.drainOnShutdown),
Layer.provide(Lifecycle.layer),
)

✅ Correct — outermost, so its finalizer runs first:

export const drainOnShutdown = Layer.effectDiscard(Effect.gen(function* () {
const lifecycle = yield* Lifecycle
const { shutdownDelay } = yield* config
yield* Effect.addFinalizer(() => lifecycle.setDraining.pipe(Effect.andThen(Effect.sleep(shutdownDelay))))
}))
export const MainLayer = Lifecycle.drainOnShutdown.pipe(
Layer.provideMerge(AppLayer),
Layer.provide(Lifecycle.layer),
)

Source: notes/09-production-concerns/lifecycle-and-shutdown.md · Decision 3, amended

End a worker’s queue and drain it for up to 5s

Section titled “End a worker’s queue and drain it for up to 5s”

Impact: HIGH a deploy no longer drops the backlog silently

  • Every worker drains on close. It is built into makeWorker in infra/worker.ts, never written per worker.
  • Queue.end, not Queue.shutdown: end stops new offers and lets take finish the buffer; shutdown throws it away.
  • The default deadline is 5s. Dropped items are counted in one Warn line, and an offer refused after shutdown Warns too. Nested workers drain one after another, each adding its deadline.
  • The item cut at the deadline may run again, which consumer idempotency makes safe. Work that must survive is an outbox row, not a worker item.

❌ Incorrect — the fiber is interrupted and the buffer discarded:

const queue = yield* Queue.unbounded<A>()
yield* Queue.take(queue).pipe(Effect.flatMap(process), Effect.forever, Effect.forkScoped)
yield* Effect.addFinalizer(() => Queue.shutdown(queue))

✅ Correct — end, drain to a deadline, count what is left:

const queue = yield* Queue.unbounded<A, Cause.Done>()
const fiber = yield* Queue.take(queue).pipe(Effect.flatMap(process), Effect.forever, Effect.forkScoped)
yield* Effect.addFinalizer(() =>
Queue.end(queue).pipe(
Effect.andThen(Fiber.await(fiber).pipe(Effect.timeoutOption(options?.drainTimeout ?? "5 seconds"))),
Effect.flatMap(Option.match({
onSome: () => Effect.void,
onNone: () => Queue.size(queue).pipe(Effect.flatMap((dropped) =>
Effect.logWarning(`${name}: shutdown drain deadline`, { dropped }))),
}))))

Source: notes/09-production-concerns/lifecycle-and-shutdown.md · Decision 4, amended

Keep every stage inside the budget; bound a release that does I/O

Section titled “Keep every stage inside the budget; bound a release that does I/O”

Impact: MEDIUM a stuck release would cost exactly the telemetry that explains it

Stage after SIGTERM Bound
1. Pre-stop delay: /ready 503, requests still served 5s (SHUTDOWN_DELAY)
2. HTTP drain 10s (gracefulShutdownTimeout)
3. Worker drain 5s per worker level
4. Releases that do I/O 2s each
5. OTLP final export 3s
Total 25s, against a platform grace of at least 30s
  • The table is the contract. A change to any stage updates it, and the sum stays at least 5s below the platform’s grace period.
  • A release we write that does I/O carries its own Effect.timeoutOption (2s by default) and logs at Warn on timeout. A release never fails, and now it never hangs either.
  • No process-level watchdog. If something still hangs, the platform’s SIGKILL is the backstop.

❌ Incorrect — an unbounded release, plus a watchdog that skips the final export:

Effect.acquireRelease(spawnAgent(id), (proc) => killAgent(proc))
setTimeout(() => process.exit(1), 25_000)

✅ Correct — the release bounds itself and leaves a Warn:

Effect.acquireRelease(spawnAgent(id), (proc) =>
killAgent(proc).pipe(
Effect.timeoutOption("2 seconds"),
Effect.flatMap(Option.match({
onNone: () => Effect.logWarning("agents: kill timed out", { id }),
onSome: () => Effect.void,
})),
Effect.ignoreCause({ log: "Warn", message: "agents: release failed" })))

Source: notes/09-production-concerns/lifecycle-and-shutdown.md · Decision 5

Report a boot failure with runMain’s default until something alerts on it

Section titled “Report a boot failure with runMain’s default until something alerts on it”

Impact: LOW the better log only pays once an alert reads it

  • Day 1, bin.ts keeps every provide before Layer.launch. A boot failure prints as runMain’s multi-line Cause.pretty, with no trace_id, after observability has closed. That gap is accepted until it matters.
  • The written plan: Observability.reportFailure logs the cause once, as one JSON line with trace_id, then fails with ProcessFailed, whose [Runtime.errorReported] = false stops a second log. Observability and ConfigProviderLayer move to Effect.provide after Layer.launch.
  • A clean SIGTERM exits 130. That is normal, not a failure.

❌ Incorrect — building the plan before anything reads its output:

MainLayer.pipe(
Layer.launch,
Observability.reportFailure, // no boot-failure alert exists yet
Effect.provide(Observability.layer),
NodeRuntime.runMain,
)

✅ Correct — day 1, every provide before Layer.launch:

MainLayer.pipe(
Layer.provide(NodeHttpServer.layerConfig(() => createServer(), { /* … */ })),
Layer.provide(Observability.layer), // still outermost of our layers
Layer.provide(ConfigProviderLayer), // beneath Observability and the server, so both read .env
Layer.provide(NodeServices.layer),
Layer.launch,
NodeRuntime.runMain,
)

Source: notes/09-production-concerns/lifecycle-and-shutdown.md · Decision 6, amended

  • A booting readiness state — trigger: the first startup work that must finish after the graph is built (a forked warm-up).
  • ProcessFailed + reportFailure, with Observability and ConfigProviderLayer moved to Effect.provide after Layer.launch — trigger: a boot-failure alert is wired.