Skip to content

Observability

Traces carry the context. 4.0.3’s automatic HTTP and SQL spans are the frame, every service method adds one named span inside it, and logs are kept for what a human has to read. This page answers what gets a span, where context goes, how signals leave the process, and what counts as an error.

Make every public service method Effect.fn("Service.method")

Section titled “Make every public service method Effect.fn("Service.method")”

Impact: HIGH a slow request shows which use case issued its queries

  • The span name is <ServiceClass>.<method>: the class name, not the @app/… tag. It covers services, repositories and outbound adapters alike. Nothing else names a span.
  • Private helpers, HttpApi group builders and handlers use Effect.fnUntraced; the server span and the service span already frame the handler. Effect.withSpan is only for code that is not a function, such as the per-item body of the worker helper ("<WorkerName>.process").
  • Never add Effect.withSpan in the body or pipeables of an Effect.fn: they already run inside its span, so it doubles.
  • Accepted cost: span volume follows call count, and renaming a method does not rename its span, so dashboards keyed on the old name break silently.

❌ Incorrect — bare names, a span on a helper, and a doubled span:

const normalize = Effect.fn("normalize")(function* (input: CreateLabelingRuleInput) { … })
create: Effect.fn("create")(function* (input: CreateLabelingRuleInput) {
// …
}, Effect.withSpan("LabelingRules.create")), // a second span around the first

✅ Correct — one named span per public method; helpers untraced:

const normalize = Effect.fnUntraced(function* (input: CreateLabelingRuleInput) { … })
return {
create: Effect.fn("LabelingRules.create")(function* (input: CreateLabelingRuleInput) {
const org = yield* CurrentOrg // annotated once by OrgScope, not here
const rule = yield* normalize(input)
yield* Effect.annotateCurrentSpan({ "app.labeling.rule_id": rule.id })
return yield* repo.insert(org.orgId, rule)
}),
}

Source: notes/08-application-surfaces/observability.md · Decision 1, amended

Put context on spans; log only what a human must read

Section titled “Put context on spans; log only what a human must read”

Impact: HIGH the server span is the wide event, for free

  • The principal is annotated once, by the middleware that provides CurrentPrincipal, ids only. The org is annotated once, by OrgScope, as app.orgs.org_id. Domain ids are annotated by the service method that knows them. Id attributes carry our internal ids, never an email or an external subject.
  • Keys use the OpenTelemetry semantic convention when one exists (enduser.id, http.*, db.*), otherwise app.<feature>.<field> in dotted snake_case. Never a bare key like orgId.
  • A log line is the one error line at a boundary, a warning for degraded behaviour, or job and worker lifecycle. No “entering X” or “X done” logs, and annotateLogs only at entry points.
  • Accepted cost: without a trace backend the context is gone; stdout keeps only the human lines.

❌ Incorrect — progress logs, and context repeated on every line with bare keys:

yield* Effect.logInfo("creating rule").pipe(Effect.annotateLogs({ orgId, userId }))
const rule = yield* repo.insert(org.orgId, input)
yield* Effect.logInfo("rule created").pipe(Effect.annotateLogs({ orgId, ruleId: rule.id }))

✅ Correct — once, in the middleware that provides CurrentPrincipal:

yield* Effect.annotateCurrentSpan({
"enduser.id": principal.userId, // OTel semconv key
"app.auth.method": principal._tag, // ours
})

Source: notes/08-application-surfaces/observability.md · Decision 2, amended

Provide one infra/observability layer, outermost

Section titled “Provide one infra/observability layer, outermost”

Impact: HIGH every layer is traced, and the exporter closes last

  • infra/observability.ts holds the logger, MinimumLogLevel, the OTLP exporter and the outbound redaction lists. bin.ts provides it with Layer.provide before Layer.launch, outermost of our layers. Its 3s exporter flush is the last stage of the shutdown budget.
  • Use Otlp.layer with our own Config.all, not Otlp.layerFromConfig(), which exports nothing unless OTEL_TRACES_EXPORTER=otlp is set and reads the token as a plain string. No endpoint, no exporter: development and tests run on stdout only.
  • Production stdout is one JSON object per line, formatStructured plus trace_id and span_id; development is consolePretty. Logger.tracerLogger is always listed, because Logger.layer replaces the whole set. deployment.environment.name comes from DEPLOYMENT_ENVIRONMENT, and SERVICE_VERSION is the deployed commit SHA.
  • The exporter is container-only: a Worker isolate never closes its build scope, so nothing flushes. The logger and the redaction lists carry over to the Worker profile as they are.

❌ Incorrect — exports nothing by default, and log lines can’t be joined to traces:

export const layer = Layer.mergeAll(
Otlp.layerFromConfig(), // silent unless OTEL_TRACES_EXPORTER=otlp
Logger.layer([Logger.consoleJson]), // no trace_id, and tracerLogger dropped
)

✅ Correct — our config, opt-in exporter, trace ids in every line:

export const layer = Layer.unwrap(Effect.map(config, (c) => Layer.mergeAll(
Logger.layer([
c.env === "production" ? jsonWithTraceId : Logger.consolePretty(),
Logger.tracerLogger, // Logger.layer replaces the set; keep it
]),
Layer.succeed(References.MinimumLogLevel, c.logLevel),
Option.match(c.otlpEndpoint, {
onNone: () => Layer.empty,
onSome: (baseUrl) => Otlp.layer({ baseUrl, resource: { … }, headers: { … } })
.pipe(Layer.provide([FetchHttpClient.layer, OtlpSerialization.layerJson])),
}),
)))

Source: notes/08-application-surfaces/observability.md · Decision 3, amended

Write a metric only for what a span cannot say

Section titled “Write a metric only for what a span cannot say”

Impact: MEDIUM one fact, recorded once

  • Request rate, errors and duration come from server spans. A Metric is written only for gauges (queue depth, pool usage, backlog), business counters with no single span, and ratios such as cache hit rate. No blanket per-endpoint counter or timer.
  • Definitions live in <feature>/metrics.ts, named app.<feature>.<measure>. Labels are a closed set of low-cardinality values, never an id, a path or a user. Shared machinery keeps its own, such as the outbox gauges in infra/outbox/metrics.ts and the job metrics in infra/jobs/metrics.ts.
  • Every duration comes from a monotonic reading (Effect.timed, Effect.trackDuration, Clock.monotonicTimeNanos), never from subtracting two wall-clock readings.
  • Accepted cost: RED from spans needs the backend to compute span metrics, so any sampling must stay downstream of that computation.

❌ Incorrect — a per-request timer that repeats the server span, labelled by an open value:

yield* Metric.update(Metric.withAttributes(requestDuration, { path: request.url, user: userId }), ms)

✅ Correct — a business counter, closed labels, defined beside the feature:

labeling/metrics.ts
export const rulesApplied = Metric.counter("app.labeling.rules_applied", {
description: "Labeling rules evaluated against an event, by outcome",
})
// labeling/apply-rules.ts
yield* Metric.update(Metric.withAttributes(rulesApplied, { outcome: "labeled" }), 1)

Source: notes/08-application-surfaces/observability.md · Decision 4, amended

Let a 5xx mark the server span Error; keep noise out

Section titled “Let a 5xx mark the server span Error; keep noise out”

Impact: MEDIUM alerts fire on the 500 we return, with no code of ours

  • Server span status follows OTel HTTP server semantics: 5xx is Error, 4xx is Ok. 4.0.3’s tracer already does this, including the successful 500 encoded from our Unexpected. A boundary test (see testing) pins it.
  • Inner spans keep their real Exit, so an expected RuleNotFound is red on its own span. Error rates and alerts read server spans only (span.kind = server).
  • ErrorReporter stays empty. The error signal is the boundary’s one log line with trace_id, plus the 500’s reference.
  • One HttpMiddleware.TracerDisabledWhen predicate, in apps/server/src/http.ts, drops /health, /ready, OPTIONS and the OAuth callback, whose ?code= must not reach telemetry.

❌ Incorrect — a hand-written middleware that duplicates the framework:

// ordered outside every other app middleware, and one more class for a future tracker to ignore
const serverErrorSpan = HttpMiddleware.make((app) => …)

✅ Correct — no span middleware; one noise list:

apps/server/src/http.ts
const OAUTH_CALLBACK = /^\/oauth\/[^/?]+\/callback/
export const NoTraceLayer = Layer.succeed(HttpMiddleware.TracerDisabledWhen, (r) =>
r.method === "OPTIONS" || r.url === "/health" || r.url === "/ready" || OAUTH_CALLBACK.test(r.url))

Source: notes/08-application-surfaces/observability.md · Decisions 1, 5, amended

  • Telemetry export on the Worker profile — trigger: the first Worker deployed to staging.
  • fnUntraced or a batched call for one hot method — trigger: more than 100 spans with that method’s name in one trace.
  • Tail or ratio sampling in a collector — trigger: the monthly telemetry bill exceeds 10% of the monthly hosting bill.
  • @effect/opentelemetry NodeSdk for traces — trigger: a library that emits through the global OTel API, such as an AI SDK, is added.
  • An error tracker as a sink, through ErrorReporter — trigger: a tracker is picked.
  • A local trace viewer in compose.yaml — trigger: the first bug whose local diagnosis needed span attributes that stdout does not show.