Skip to content

Error boundaries

Failures stay typed until a boundary ends them. This page answers when to die, how to catch, how wide an error union gets, and where a failure is logged: exactly once, at the place it ends.

A defect means a bug; infrastructure failures stay typed

Section titled “A defect means a bug; infrastructure failures stay typed”

Impact: HIGH an orDie in a service throws away typed retry

  • Effect.die / Effect.orDie in feature code are allowed for exactly three things: a broken invariant, boot-time configuration, and an operation that cannot fail (completing a Deferred we own). Each has a comment naming the invariant, and dies with a namespaced Schema.TaggedError, not a string. Lint rule app/no-effect-die enforces it.
  • Connection loss, deadlock, timeout and upstream 5xx are not bugs. They travel as PersistenceError or a feature’s upstream error until the surface boundary turns them into a defect (see HttpApi).
  • Client input is never a defect: a malformed request fails typed and gets a 4xx. No Cause.squash(…) as E; pass the Cause or Exit on instead. There is no orDieWith; write Effect.catch((e) => Effect.die(…)).

❌ Incorrect — a storage failure removed from the signature:

get: (id: RuleId) => Effect.gen(function* () {
const org = yield* CurrentOrg
const row = yield* repo.findById(org.orgId, id).pipe(Effect.orDie) // "can't happen"
…
})

✅ Correct — infrastructure in the signature; die only on a named invariant:

get: (id: RuleId) => Effect.Effect<LabelingRule, RuleNotFound | PersistenceError, CurrentOrg>
const stored = yield* Effect.fromOption(
yield* repo.findById(org.orgId, ruleId), // PersistenceError stays typed; orgId first
() => new StoredRuleMissing({ ruleId }),
).pipe(Effect.orDie) // invariant: the row was inserted earlier in this transaction

Source: notes/06-errors/error-boundaries.md · Decision 1, amended

Catch by tag; a catch-all only where a failure ends

Section titled “Catch by tag; a catch-all only where a failure ends”

Impact: HIGH a new service error breaks the compile of handlers that ignore it

  • To remap or recover use catchTag, catchTags or catchReason. The catch-all is allowed only at a terminal boundary: “everything else is a defect” in a surface handler, or per-item isolation in a worker.
  • Even there, name the tags you map and send the rest to the orElse argument, so a reviewer sees what goes to “internal”.
  • mapError only when wrapping every error is the intent: it also wraps the domain error you meant to keep. Isolate with Effect.catch + Effect.catchDefect, not catchCause, so interruption passes without a hand-written check.
  • No predicate catch that collapses errors: turning a database failure into Forbidden hides an outage behind a 403.

❌ Incorrect — one catch-all; a new expected error silently becomes “internal”:

createRule: ({ payload }) =>
rules.create(payload).pipe(
Effect.map(toPublicRule),
Effect.catch((e) => internalFailure(e)),
),

✅ Correct — named tags, the remainder explicit:

createRule: ({ payload }) =>
rules.create(payload).pipe(
Effect.map(toPublicRule),
Effect.catchTags({
"@app/labeling/RuleLabelTaken": (e) => new Http.Conflict({ reason: `label ${e.label} is taken` }),
"@app/labeling/StaleRevision": () => new Http.Conflict({ reason: "stale revision" }),
}, internalFailure), // orElse: the remainder, explicitly
),

Source: notes/06-errors/error-boundaries.md · Decision 2, amended

Impact: MEDIUM callers never handle errors their method cannot produce

  • Each method writes its union inline: domain errors plus PersistenceError or a feature upstream error. No infrastructure types, and no wrapper errors whose only field is a cause.
  • No service-wide Feature.Error union. One union for every method makes each caller handle errors it cannot get, and lets dead members survive.
  • When a union gets hard to read, the method is doing too much. Split the method; do not collapse the errors.

❌ Incorrect — one union for the whole service:

export type LabelingRulesError =
| RuleNotFound | RuleLabelTaken | StaleRevision | DuplicateLabelingRule
| PersistenceError | SqlError /* … 13 members */
readonly get: (id: RuleId) => Effect.Effect<LabelingRule, LabelingRulesError>

✅ Correct — each method lists what it can produce:

export type LabelingRulesShape = {
readonly get: (id: RuleId) => Effect.Effect<LabelingRule, RuleNotFound | PersistenceError, CurrentOrg>
readonly create: (input: CreateLabelingRuleInput) => Effect.Effect<
LabelingRule,
RuleLabelTaken | StaleRevision | PersistenceError,
CurrentPrincipal | CurrentOrg
>
}

Source: notes/06-errors/error-boundaries.md · Decision 3, amended

Impact: HIGH one incident, one error line, matching the 500 the client saw

  • A failure ends when it is answered (a response or a typed Unexpected), swallowed so a loop survives, or handed to a durable retry. Only that code logs it. tapError(logError) and then propagate is forbidden.
  • For HTTP the API-level catchDefect boundary writes the one log line for every defect. The handler’s internalFailure only dies.
  • Turn off HttpMiddleware.logger with HttpRouter.disableLogger; it would log the same cause again. Expected failures that become a 4xx are not error-level events.
  • Pass the Cause or the error to the logger, never String(cause), which collapses a schema error to its tag. How a cause is rendered is observability’s call.

❌ Incorrect — logged at every layer it passes through:

const internalFailure = (e: PersistenceError | UpstreamError) =>
Effect.logError("rule handler failed", e).pipe(Effect.andThen(Effect.die(e)))
// … and the boundary logs it again, and the request logger a third time

✅ Correct — the handler dies; the boundary logs:

labeling/http.ts
const internalFailure = (e: PersistenceError | UpstreamError) => Effect.die(e)
// http.ts — API-level boundary
Effect.catchDefect((defect) =>
Effect.gen(function* () {
const reference = yield* currentReference // trace id
yield* Effect.logError("unexpected failure", Cause.die(defect))
return yield* new Http.Unexpected({ reference })
}))
// server assembly
Layer.provide(HttpRouter.disableLogger)

Source: notes/06-errors/error-boundaries.md · Decision 4

Give every background loop a per-item boundary in its helper

Section titled “Give every background loop a per-item boundary in its helper”

Impact: HIGH a forked fiber that fails is never logged; the loop just stops

  • Effect logs nothing when a forked fiber fails and nobody joins it. So the boundary lives in the helper (infra/worker.ts’s makeWorker), and a caller cannot forget it. It catches per item with Effect.catch + Effect.catchDefect, logs once and continues.
  • makeWorker returns { offer }, and on close ends and drains its queue for 5s, Warning with the dropped count (see lifecycle and shutdown).
  • A scheduled job’s boundary is Jobs.run. On a Cloudflare Worker, background work goes through the infra/ helper around waitUntil, which owns the same boundary; alchemy’s waitUntil logs nothing on its own.
  • Durable-queue consumers log once and re-fail the message, so the platform owns retry. An undecodable message is acked with an Error line. The outbox relay is the named exception and owns its retry (see idempotency and outbox).
  • No bare Effect.ignoreCause / Effect.ignore. A best-effort one-shot uses Effect.ignoreCause({ log: true }).

❌ Incorrect — a loop forked bare; one bad item kills it silently:

yield* Queue.take(queue).pipe(
Effect.flatMap(process),
Effect.forever,
Effect.forkScoped,
)

✅ Correct — the helper owns the per-item boundary:

infra/worker.ts
yield* Queue.take(queue).pipe(
Effect.flatMap((item) =>
process(item).pipe(
Effect.catch((e) => Effect.logError(`${name} failed`, e)),
Effect.catchDefect((d) => Effect.logError(`${name} defect`, Cause.die(d))),
)),
Effect.forever,
Effect.forkScoped,
)
// … ends and drains the queue on close, and returns { offer }

Source: notes/06-errors/error-boundaries.md · Decision 5, amended

Impact: MEDIUM extra process handlers turn into silent no-ops

  • NodeRuntime.runMain (or BunRuntime.runMain) is the last-resort boundary. It logs any non-interrupt failure of the main program and sets the exit code (see runtime and entrypoint).
  • No process.on("unhandledRejection" | "uncaughtException") handlers.
  • A platform guard below Effect is allowed when it names the crash it prevents (for example, late EPIPE on a response socket).

❌ Incorrect — a process handler that swallows what Effect would report:

process.on("unhandledRejection", () => {})
MainLayer.pipe(Layer.launch, NodeRuntime.runMain({ disableErrorReporting: true }))

✅ Correct — the runtime’s own boundary, with no options:

MainLayer.pipe(
Layer.provide(/* platform, config, observability */),
Layer.launch,
NodeRuntime.runMain,
)

Source: notes/06-errors/error-boundaries.md · Decision 6

  • A named error union per feature, derived from a tuple of classes — trigger: the same set of errors becomes a contract used in a second place, such as an RPC’s declared errors.
  • The RPC boundary (RpcServer.layerHttp with disableFatalDefects: true, and an RpcErrorBoundary middleware added last that turns defects into a typed Unexpected and is the single log site) — trigger: an RPC server exists, together with the first server-pushed subscription in RPC contracts.