Skip to content

CI pipelines

Every PR and every main commit runs the same CI: two parallel jobs folded into one gate. This page answers what runs on every push, which checks a merge requires, what happens when main goes red, how deploys hang off CI, and what to do with a flaky test.

Run two parallel jobs behind one gate, with no path filters

Section titled “Run two parallel jobs behind one gate, with no path filters”

Impact: HIGH one round trip shows every failure

  • check runs linting and formatting’s one command (pnpm run check: graph scripts, oxfmt, oxlint, tsc --noEmit), then ci:migrations-immutable and ci:gate-complete. test runs the whole suite, migration contiguity and contract fixtures included.
  • ci-passed needs both and treats skipped as a failure. ci:gate-complete fails if any job is missing from ci-passed.needs.
  • No path filters, ever: a modular monolith’s dependency graph is dense, and a filter is a hand-kept copy of it. Every job has timeout-minutes. Default permissions are contents: read.
  • One composite .github/actions/setup installs the toolchain and dependencies, cached on the lockfile. Accepted cost: a docs-only PR pays for the whole suite.

❌ Incorrect — a path filter guessing what a change can reach:

on:
pull_request:
paths: ["apps/server/**"] # a change in packages/db never runs the server tests

✅ Correct — everything runs, and one gate reports it:

.github/workflows/ci.yml
jobs:
check:
timeout-minutes: 10
steps:
- uses: actions/checkout@<sha>
with: { fetch-depth: 0 } # the immutability check diffs against the base
- uses: ./.github/actions/setup
- run: pnpm run check
- run: pnpm run ci:migrations-immutable # PR: vs merge base; main: vs HEAD^
- run: pnpm run ci:gate-complete
test:
timeout-minutes: 15
steps: [{ uses: actions/checkout@<sha> }, { uses: ./.github/actions/setup }, { run: pnpm run test }]
ci-passed:
needs: [check, test]
if: always()
steps:
- if: contains(needs.*.result, 'failure') || contains(needs.*.result, 'cancelled') || contains(needs.*.result, 'skipped')
run: exit 1

Source: notes/11-repo-operations/ci-pipelines.md · Decision 1, amended

Impact: HIGH renaming or sharding a job never touches the ruleset

  • Two stable names are the whole contract between the ruleset and the workflows. Renaming, splitting or sharding a job changes ci-passed.needs, never the ruleset.
  • pr-title stays its own workflow, because its edited trigger would re-run the whole suite on every title edit (see git workflow).
  • deploy.yml, ai-review.yml, labeler.yml and anything on request are never required. A required check that does not report leaves the PR pending forever.
  • Accepted cost: ci:gate-complete is one more script to keep working.

❌ Incorrect — every job by name, or a gate nobody requires:

required_status_checks [check, test, test (shard 1/2), ai-review]

✅ Correct — the gate and the title, nothing else:

ruleset on main
required_status_checks [ci-passed, pr-title]
strict false

Source: notes/11-repo-operations/ci-pipelines.md · Decision 2

Let main’s CI be the backstop, and never cancel it

Section titled “Let main’s CI be the backstop, and never cancel it”

Impact: HIGH every main SHA gets a real verdict

  • No “require branches to be up to date”, no merge queue. When two green PRs collide (both add migration 0048), main goes red, nothing deploys, and the author of the second merge opens a fix PR at once. No direct push, even for a red main.
  • While main is red, nobody merges except the fix.
  • main runs the same ci.yml as PRs. A PR’s stale runs are cancelled; a main run never is, because deploy.yml starts from each run’s conclusion and a cancelled run would read as a failed deploy.
  • Accepted cost: main is red for one fix PR (about 15–20 minutes), and every main SHA pays a full run.

❌ Incorrect — cancelling by ref, main included:

concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true # the next merge cancels this main run; the deploy skips

✅ Correct — cancel only a PR’s own stale runs:

concurrency:
group: ci-${{ github.event.pull_request.number || github.sha }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}

Source: notes/11-repo-operations/ci-pipelines.md · Decisions 3, 4

Impact: HIGH production only ever takes a SHA that passed CI and staging

  • workflow_run on CI completed on main, deploying its head_sha. Never push. deploy-prd needs deploy-stg in the same run, so it deploys the SHA staging just took.
  • Inside a deploy job, migrate is the first step after setup, with the owner credential in that step’s env only. Smoke (readiness 200 through the public URL) is the last.
  • After a successful deploy-prd, release tags the SHA and creates the GitHub Release. It is the only job with contents: write. The rollback dispatch skips migrate and release.
  • deployment-gate: a skipped or failed deploy is red; a superseded one is cancelled, grey. Reviewers approve the newest waiting production run, never reject an older one. See deployment and environments and release and versioning.

❌ Incorrect — a deploy that does not wait for CI:

on: { push: { branches: [main] } }

✅ Correct — chained off CI’s conclusion:

.github/workflows/deploy.yml
on:
workflow_run: { workflows: [CI], types: [completed], branches: [main] }
workflow_dispatch:
inputs: { sha: { required: true, type: string } } # rollback
jobs:
deploy-stg:
if: github.event_name == 'workflow_run' && github.event.workflow_run.conclusion == 'success'
environment: staging
concurrency: { group: deploy-stg, cancel-in-progress: false }
deploy-prd:
needs: deploy-stg
environment: production # required reviewer

Source: notes/11-repo-operations/ci-pipelines.md · Decision 5, amended

Keep slow and scheduled work off the per-push path

Section titled “Keep slow and scheduled work off the per-push path”

Impact: MEDIUM a failing job with no owner teaches people to ignore red

  • Load tests and benchmarks run on workflow_dispatch only. There is no per-PR benchmark comparison and no nightly suite.
  • There is no cron job on day 1. The plan’s only one, the preview sweep, is deferred with previews. A scheduled job that can fail needs an owner, and a cron run has no PR author.
  • An informational job (AI review, labels, a measurement) never fails and is never inside ci-passed.
  • Production deploy waits for a reviewer, rollback is a dispatch with a SHA, and the restore drill is done by hand.

❌ Incorrect — a nightly suite nobody owns:

on: { schedule: [{ cron: "0 3 * * *" }] }
jobs: { full-suite: { steps: [{ run: pnpm run test:slow }] } }

✅ Correct — expensive suites on request:

.github/workflows/load-test.yml
on: { workflow_dispatch: {} }

Source: notes/11-repo-operations/ci-pipelines.md · Decision 6, amended

Test on production’s runtime only, from one version file

Section titled “Test on production’s runtime only, from one version file”

Impact: MEDIUM one runtime shipped, one runtime tested

  • No runtime matrix and no OS matrix; runs-on: ubuntu-latest everywhere. A server ships one Linux container.
  • .node-version is the single source for CI, the production image and local tools (see local development setup). A Node upgrade is one PR, and its CI is the compatibility test.
  • On the Bun profile, Bun is pinned the same way and pnpm still installs. A Cloudflare Worker project runs its tests on Node; workerd is not the test runtime.
  • Accepted cost: breakage from the next Node version is found in the upgrade PR, not before.

❌ Incorrect — a matrix for users a server does not have:

strategy:
matrix: { node: [22, 24], os: [ubuntu-latest, macos-latest] }

✅ Correct — production’s version, read from the file:

.github/actions/setup/action.yml
- uses: pnpm/action-setup@<sha>
- uses: actions/setup-node@<sha>
with: { node-version-file: .node-version, cache: pnpm }
- run: pnpm install --frozen-lockfile
shell: bash

Source: notes/11-repo-operations/ci-pipelines.md · Decision 7, amended

Never retry a test; fix or quarantine it the same day

Section titled “Never retry a test; fix or quarantine it the same day”

Impact: HIGH a retry hides exactly the real race

  • retry: 0 everywhere, CI included. it.flakyTest is not used.
  • A flake is fixed or quarantined on the day it is seen. The fix usually moves the test onto TestClock or off shared state. Quarantine is it.effect.skip with the issue number in the name. The lint rule app/no-unlinked-flaky fails any it.flakyTest and any .skip( without #<issue>.
  • Re-running a failed job to unblock a merge is allowed. Whoever presses re-run opens the issue.
  • Accepted cost: a flake blocks others until someone deals with it, and a quarantined test covers nothing until it is fixed.

❌ Incorrect — retries that let a race ship:

vitest.config.ts
test: { retry: process.env.CI ? 2 : 0 }
it.flakyTest(…)

✅ Correct — no retries, and visible debt:

vitest.config.ts
test: { retry: 0 }
it.effect.skip("applies the rule kind — flaky, #123", () => …)

Source: notes/11-repo-operations/ci-pipelines.md · Decision 8

  • Sharding the test job — trigger: the test job’s p50 exceeds 10 minutes.
  • A merge queue, or “require branches to be up to date” — trigger: two or more main freezes in four weeks.
  • preview.yml and preview-sweep.yml — trigger: previews are built, when a second human reviewer does not run branches locally.
  • A build cache — trigger: check p50 over 5 minutes.
  • Moving an informational job behind ci-passed — trigger: the first time it is made to gate.