Skip to content

Testing

The service under test is always real. Its collaborators are partial fakes, except the database, which is real, embedded and fresh for every test. This page answers which harness to use, what to fake and what to run for real, how time is controlled, and what each test level must cover.

Impact: HIGH deterministic time and a closed scope become the default

  • it.effect is the default. It gives every test its own scope plus TestClock and TestConsole. it.live is only for a real clock or live OS behaviour: subprocesses, filesystem mtimes, real sockets.
  • Provide the test layer per test with Effect.provide(testLayer), so nothing leaks between tests. Reserve it.layer for something immutable and expensive, never for state or the database.
  • No Effect.runPromise, runSync, runFork or ManagedRuntime.make in a test file. A lint rule in the shape of t3code’s no-manual-effect-runtime-in-tests enforces it.
  • Flaky tests are never retried (test: { retry: 0 }). A flake not fixed the same day is quarantined with it.effect.skip and flaky, #<issue> in its name; app/no-unlinked-flaky enforces both. See CI pipelines.

❌ Incorrect — a hand-run runtime inside a vitest test:

import { it } from "@effect/vitest"
it("publishes a draft", async () => {
const version = await Effect.runPromise(publish(slug).pipe(Effect.provide(testLayer)))
expect(version.revision).toBe(5)
})

✅ Correct — it.effect, layer provided per test:

import { assert, it } from "@effect/vitest"
it.effect("publishes a draft", () =>
Effect.gen(function* () {
const policies = yield* Policies.Policies
const version = yield* policies.publish(slug)
assert.strictEqual(version.revision, 5)
}).pipe(Effect.provide(testLayer)))

Source: notes/08-application-surfaces/testing.md · Decision 1, amended

Fake collaborators with a partial Layer.mock

Section titled “Fake collaborators with a partial Layer.mock”

Impact: HIGH a test says what it depends on and nothing else

  • X.layer leaves its dependencies unprovided (see service definition), so a test provides fakes beneath it. Write only the methods this path uses; any other method dies with UnimplementedError, naming it.
  • Record calls in a Ref and assert on the Ref. No spy library. External HTTP is faked at the HttpClient port, never by patching fetch or globalThis.
  • A canned fake that three test files would copy moves into the owning feature’s testing.ts (see test placement). A fake for a dual method supports both call styles.
  • Never vi.mock a module that exports an Effect service, and never cast a fake into shape.

❌ Incorrect — a module mock and a cast:

vi.mock("./policies-repo.ts")
const repo = { findBySlug: vi.fn() } as unknown as PoliciesRepo.PoliciesRepoShape

✅ Correct — the real service over partial fakes:

const testLayer = Policies.layer.pipe(
Layer.provide([
Layer.mock(PoliciesRepo.PoliciesRepo)({
findBySlug: () => Effect.succeed(Option.some(policy)),
}),
Layer.mock(GitHubClient.GitHubClient)({
addLabels: (input) => Ref.update(calls, Arr.append(input)),
}),
]),
)

Source: notes/08-application-surfaces/testing.md · Decision 2, amended

Test repositories against a real embedded database, fresh per test

Section titled “Test repositories against a real embedded database, fresh per test”

Impact: HIGH the repository is where untyped rows become typed data

  • Every repository has tests against a real database. Those tests execute the Row schema, fromRow/toRow and the SqlSchema codecs. A repository is never tested by faking SqlClient; services above it fake the repository instead.
  • The database is embedded, in the production dialect: PGlite for Postgres, :memory: SQLite through the production driver family for SQLite, and a local D1 for D1 (a SQLite test would accept code that dies on D1). The D1 test graph provides no Transactions layer.
  • TestDatabase.layer (in packages/db/src/testing.ts) is provided per test and runs the migrations on build, so every repository test also proves the migrations apply. A migration that transforms data gets its own test: migrate to N-1, insert, migrate to N, assert.
  • Every tenant repository method has a test named “<method> never sees another org”, using the shared seedTwoOrgs fixture (see multi-tenancy).

❌ Incorrect — one database shared by a whole block:

layer(TestDatabase.layer)("PoliciesRepo", (it) => {
it.effect("inserts", () => /* leaves rows behind for the next test */ Effect.void)
})

✅ Correct — a fresh database per test, orgId first:

const testLayer = PoliciesRepo.layer.pipe(Layer.provide(TestDatabase.layer))
it.effect("round-trips a policy", () =>
Effect.gen(function* () {
const repo = yield* PoliciesRepo.PoliciesRepo
yield* repo.insert(orgId, policy)
assert.deepStrictEqual(yield* repo.findBySlug(orgId, slug), Option.some(policy))
}).pipe(Effect.provide(testLayer)))

Source: notes/08-application-surfaces/testing.md · Decision 3, amended

Drive time with TestClock, never with the wall clock

Section titled “Drive time with TestClock, never with the wall clock”

Impact: HIGH retry, expiry and timeout bugs only show under a controlled clock

  • Production reads time only through Clock, DateTime.now, Effect.sleep and Schedule (see time and clocks). A test that depends on “now” calls TestClock.setTime before anything reads it; TestClock starts at epoch 0.
  • Timeouts, retries and schedules use fork → adjust → assert → join, stepping to just before the boundary and then across it. Never wait on wall time; wait on a Deferred, Latch or Queue.
  • Every lease gets a skew test, every calendar computation a DST test, and a backwards jump is tested with setTime to an earlier instant. The suite runs under TZ=America/New_York, set in the test script. Randomness is pinned with Random.withSeed(seed).
  • The infra/ tests all run under TestClock: makeWorker’s 5s drain; infra/outbox (no row runs twice, expired lease reclaimed, Retry.durableDelay, unknown tag retried); infra/jobs (one winner per slot, no re-run, skip under a live lease, reclaim, holder-guarded release); infra/retry (isContention per SqlError reason, every policy terminates).

❌ Incorrect — a real sleep, or Vitest’s fake timers:

it.live("retries", () =>
client.send(req).pipe(Effect.forkChild, Effect.andThen(Effect.sleep("2 seconds"))))
// or: vi.useFakeTimers() — does not drive Effect's clock

✅ Correct — fork, adjust to the boundary, assert, cross it, join:

const fiber = yield* client.send(req).pipe(Effect.forkChild)
yield* TestClock.adjust("1999 millis")
assert.strictEqual(yield* Ref.get(attempts), 1)
yield* TestClock.adjust("1 milli")
yield* Fiber.join(fiber)

Source: notes/08-application-surfaces/testing.md · Decision 4, amended

Test handlers through HttpApiTest.groups with the service mocked

Section titled “Test handlers through HttpApiTest.groups with the service mocked”

Impact: MEDIUM handler tests stay about translation, typed end to end

  • HttpApiTest.groups(Api, [group]) runs the real pipeline without a socket: decode, middleware, handler, boundary, encode and client decode, through the generated typed client. Unselected groups are stubbed. HttpServer.layerServices supplies the platform services.
  • A handler is an adapter, so its test mocks the feature service and proves only the mapping. Business rules, including authorization deny cases, are tested once at the service.
  • Auth is replaced at the middleware: Identity.authenticatedAs(principal) provides CurrentPrincipal without minting tokens. The real Authentication middleware is tested once, directly, with a Layer.mock Authenticator.
  • Webhooks, OAuth callbacks, SSE and health stay raw HttpRouter routes, tested with a built Request.

❌ Incorrect — the whole stack behind every handler test, with hand-written auth headers:

const handler = HttpRouter.toWebHandler(AllRoutes.pipe(Layer.provide(RealServicesLayer)))
await handler(new Request(url, { headers: { authorization: `Bearer ${token}` } }))

✅ Correct — the adapter alone, through the typed client:

const makeClient = HttpApiTest.groups(Api, ["labeling"])
const testLayer = (rules: Partial<LabelingRules.LabelingRulesShape>, principal = alice) =>
Labeling.http.layer.pipe(
Layer.provide(Layer.mock(LabelingRules.LabelingRules)(rules)),
Layer.provideMerge(Layer.mergeAll(
Identity.authenticatedAs(principal), ErrorBoundary.layer, HttpServer.layerServices,
)),
)

Source: notes/08-application-surfaces/testing.md · Decision 5, amended

Pin the API contract once: boundary, structure and compatibility

Section titled “Pin the API contract once: boundary, structure and compatibility”

Impact: HIGH catches the middleware and contract mistakes the compiler misses

  • Per endpoint: the happy path, plus one test per error mapping that is not 1:1 (a tenant mismatch surfacing as NotFound, say). A 1:1 catchTag line gets no test of its own.
  • api.boundary.test.ts: a malformed payload is @app/http/InvalidRequest 400 with the field path; a defect and a response-encode failure are @app/http/Unexpected 500 with no internal message; that 500 ends its server span as Error.
  • api.contract.test.ts: every endpoint outside a *Public group has Authentication and the wire Forbidden; OrgScope groups carry orgId and are not outermost; /orgs/:orgId groups have OrgScope. When RPC exists, it also iterates AppRpcs.
  • api.compat.test.ts: the previous release’s responses, kept in fixtures/compat/previous.json, decode with the current schema and vice versa. See API evolution.

❌ Incorrect — one test per declared error, restating each catchTag line:

it.effect("maps RuleNotFound to NotFound", () => /* … */ Effect.void)
it.effect("maps Forbidden to Forbidden", () => /* … */ Effect.void)
// …and no test that every endpoint is protected

✅ Correct — one structural test covers every endpoint:

it("protects every non-public endpoint", () => {
for (const [groupName, group] of Object.entries(Api.groups)) {
if (groupName.endsWith("Public")) continue
for (const [name, endpoint] of Object.entries(group.endpoints)) {
assert.isTrue(endpoint.middlewares.has(Authentication), `${groupName}.${name} is unprotected`)
assert.isTrue(endpoint.error.has(ContractErrors.Forbidden), `${groupName}.${name} cannot return 403`)
}
}
})

Source: notes/08-application-surfaces/testing.md · Decision 5, amended

  • A shared fake module for one port — trigger: a third test file would copy the same canned fake.
  • A cached post-migration database snapshot — trigger: migrations take more than half of the test wall time.
  • The choice of local-D1 harness — trigger: the first D1 repository test.
  • An integration suite against the production engine (testcontainers Postgres, or a real D1) — trigger: a bug that passed the embedded engine fails on the production engine.
  • API-key cases in handler tests — trigger: the ApiKey principal variant is built.