CI pipelines
Every PR and every main commit runs the same CI: two parallel jobs folded into one gate. This
page answers what runs on every push, which checks a merge requires, what happens when
main goes red, how deploys hang off CI, and what to do with a flaky test.
Run two parallel jobs behind one gate, with no path filters
Section titled “Run two parallel jobs behind one gate, with no path filters”Impact: HIGH one round trip shows every failure
checkruns linting and formatting’s one command (pnpm run check: graph scripts, oxfmt, oxlint,tsc --noEmit), thenci:migrations-immutableandci:gate-complete.testruns the whole suite, migration contiguity and contract fixtures included.ci-passedneeds both and treatsskippedas a failure.ci:gate-completefails if any job is missing fromci-passed.needs.- No path filters, ever: a modular monolith’s dependency graph is dense, and a filter is a
hand-kept copy of it. Every job has
timeout-minutes. Default permissions arecontents: read. - One composite
.github/actions/setupinstalls the toolchain and dependencies, cached on the lockfile. Accepted cost: a docs-only PR pays for the whole suite.
❌ Incorrect — a path filter guessing what a change can reach:
on: pull_request: paths: ["apps/server/**"] # a change in packages/db never runs the server tests✅ Correct — everything runs, and one gate reports it:
jobs: check: timeout-minutes: 10 steps: - uses: actions/checkout@<sha> with: { fetch-depth: 0 } # the immutability check diffs against the base - uses: ./.github/actions/setup - run: pnpm run check - run: pnpm run ci:migrations-immutable # PR: vs merge base; main: vs HEAD^ - run: pnpm run ci:gate-complete test: timeout-minutes: 15 steps: [{ uses: actions/checkout@<sha> }, { uses: ./.github/actions/setup }, { run: pnpm run test }] ci-passed: needs: [check, test] if: always() steps: - if: contains(needs.*.result, 'failure') || contains(needs.*.result, 'cancelled') || contains(needs.*.result, 'skipped') run: exit 1Source: notes/11-repo-operations/ci-pipelines.md · Decision 1, amended
Require exactly ci-passed and pr-title
Section titled “Require exactly ci-passed and pr-title”Impact: HIGH renaming or sharding a job never touches the ruleset
- Two stable names are the whole contract between the ruleset and the workflows. Renaming,
splitting or sharding a job changes
ci-passed.needs, never the ruleset. pr-titlestays its own workflow, because itseditedtrigger would re-run the whole suite on every title edit (see git workflow).deploy.yml,ai-review.yml,labeler.ymland anything on request are never required. A required check that does not report leaves the PR pending forever.- Accepted cost:
ci:gate-completeis one more script to keep working.
❌ Incorrect — every job by name, or a gate nobody requires:
required_status_checks [check, test, test (shard 1/2), ai-review]✅ Correct — the gate and the title, nothing else:
ruleset on main required_status_checks [ci-passed, pr-title] strict falseSource: notes/11-repo-operations/ci-pipelines.md · Decision 2
Let main’s CI be the backstop, and never cancel it
Section titled “Let main’s CI be the backstop, and never cancel it”Impact: HIGH every main SHA gets a real verdict
- No “require branches to be up to date”, no merge queue. When two green PRs collide (both add
migration
0048),maingoes red, nothing deploys, and the author of the second merge opens a fix PR at once. No direct push, even for a redmain. - While
mainis red, nobody merges except the fix. mainruns the sameci.ymlas PRs. A PR’s stale runs are cancelled; amainrun never is, becausedeploy.ymlstarts from each run’s conclusion and a cancelled run would read as a failed deploy.- Accepted cost:
mainis red for one fix PR (about 15–20 minutes), and everymainSHA pays a full run.
❌ Incorrect — cancelling by ref, main included:
concurrency: group: ci-${{ github.ref }} cancel-in-progress: true # the next merge cancels this main run; the deploy skips✅ Correct — cancel only a PR’s own stale runs:
concurrency: group: ci-${{ github.event.pull_request.number || github.sha }} cancel-in-progress: ${{ github.event_name == 'pull_request' }}Source: notes/11-repo-operations/ci-pipelines.md · Decisions 3, 4
Deploy from green main runs only
Section titled “Deploy from green main runs only”Impact: HIGH production only ever takes a SHA that passed CI and staging
workflow_runon CI completed onmain, deploying itshead_sha. Neverpush.deploy-prdneedsdeploy-stgin the same run, so it deploys the SHA staging just took.- Inside a deploy job,
migrateis the first step after setup, with the owner credential in that step’senvonly. Smoke (readiness 200 through the public URL) is the last. - After a successful
deploy-prd,releasetags the SHA and creates the GitHub Release. It is the only job withcontents: write. The rollback dispatch skipsmigrateandrelease. deployment-gate: a skipped or failed deploy is red; a superseded one iscancelled, grey. Reviewers approve the newest waiting production run, never reject an older one. See deployment and environments and release and versioning.
❌ Incorrect — a deploy that does not wait for CI:
on: { push: { branches: [main] } }✅ Correct — chained off CI’s conclusion:
on: workflow_run: { workflows: [CI], types: [completed], branches: [main] } workflow_dispatch: inputs: { sha: { required: true, type: string } } # rollbackjobs: deploy-stg: if: github.event_name == 'workflow_run' && github.event.workflow_run.conclusion == 'success' environment: staging concurrency: { group: deploy-stg, cancel-in-progress: false } deploy-prd: needs: deploy-stg environment: production # required reviewerSource: notes/11-repo-operations/ci-pipelines.md · Decision 5, amended
Keep slow and scheduled work off the per-push path
Section titled “Keep slow and scheduled work off the per-push path”Impact: MEDIUM a failing job with no owner teaches people to ignore red
- Load tests and benchmarks run on
workflow_dispatchonly. There is no per-PR benchmark comparison and no nightly suite. - There is no cron job on day 1. The plan’s only one, the preview sweep, is deferred with previews. A scheduled job that can fail needs an owner, and a cron run has no PR author.
- An informational job (AI review, labels, a measurement) never fails and is never inside
ci-passed. - Production deploy waits for a reviewer, rollback is a dispatch with a SHA, and the restore drill is done by hand.
❌ Incorrect — a nightly suite nobody owns:
on: { schedule: [{ cron: "0 3 * * *" }] }jobs: { full-suite: { steps: [{ run: pnpm run test:slow }] } }✅ Correct — expensive suites on request:
on: { workflow_dispatch: {} }Source: notes/11-repo-operations/ci-pipelines.md · Decision 6, amended
Test on production’s runtime only, from one version file
Section titled “Test on production’s runtime only, from one version file”Impact: MEDIUM one runtime shipped, one runtime tested
- No runtime matrix and no OS matrix;
runs-on: ubuntu-latesteverywhere. A server ships one Linux container. .node-versionis the single source for CI, the production image and local tools (see local development setup). A Node upgrade is one PR, and its CI is the compatibility test.- On the Bun profile, Bun is pinned the same way and pnpm still installs. A Cloudflare Worker project runs its tests on Node; workerd is not the test runtime.
- Accepted cost: breakage from the next Node version is found in the upgrade PR, not before.
❌ Incorrect — a matrix for users a server does not have:
strategy: matrix: { node: [22, 24], os: [ubuntu-latest, macos-latest] }✅ Correct — production’s version, read from the file:
- uses: pnpm/action-setup@<sha>- uses: actions/setup-node@<sha> with: { node-version-file: .node-version, cache: pnpm }- run: pnpm install --frozen-lockfile shell: bashSource: notes/11-repo-operations/ci-pipelines.md · Decision 7, amended
Never retry a test; fix or quarantine it the same day
Section titled “Never retry a test; fix or quarantine it the same day”Impact: HIGH a retry hides exactly the real race
retry: 0everywhere, CI included.it.flakyTestis not used.- A flake is fixed or quarantined on the day it is seen. The fix usually moves the test onto
TestClockor off shared state. Quarantine isit.effect.skipwith the issue number in the name. The lint ruleapp/no-unlinked-flakyfails anyit.flakyTestand any.skip(without#<issue>. - Re-running a failed job to unblock a merge is allowed. Whoever presses re-run opens the issue.
- Accepted cost: a flake blocks others until someone deals with it, and a quarantined test covers nothing until it is fixed.
❌ Incorrect — retries that let a race ship:
test: { retry: process.env.CI ? 2 : 0 }
it.flakyTest(…)✅ Correct — no retries, and visible debt:
test: { retry: 0 }
it.effect.skip("applies the rule kind — flaky, #123", () => …)Source: notes/11-repo-operations/ci-pipelines.md · Decision 8
Deferred
Section titled “Deferred”- Sharding the
testjob — trigger: thetestjob’s p50 exceeds 10 minutes. - A merge queue, or “require branches to be up to date” — trigger: two or more
mainfreezes in four weeks. preview.ymlandpreview-sweep.yml— trigger: previews are built, when a second human reviewer does not run branches locally.- A build cache — trigger:
checkp50 over 5 minutes. - Moving an informational job behind
ci-passed— trigger: the first time it is made to gate.