Deployment and environments
Two environments on day 1, one alchemy program, and one path to production. This page answers which environments exist, how a commit gets promoted, where migrations and secrets run, and how to undo code and data.
Run production and staging; keep previews as a written plan
Section titled “Run production and staging; keep previews as a written plan”Impact: HIGH staging is where an old database meets new code
- Production runs N (≥ 2) replicas behind the load balancer, with its own database and PITR. Staging runs 1 replica and its own long-lived database with synthetic data.
- Staging catches what tests cannot: a migration meeting a database that lived through earlier releases, and a key rotation rehearsal.
- Label-gated per-PR previews named
pr-<n>are designed but not built. With one developer and no required approvals, nobody needs them yet. On the Cloudflare Worker profile, “replicas” does not apply; the environments do.
❌ Incorrect — production only, so the first old-database migration runs in production:
main → production✅ Correct — two environments, previews waiting on a trigger:
production main, after approval N ≥ 2 replicas own database, PITRstaging main, after green CI 1 replica own database, synthetic datapr-<n> deferred — —Source: notes/11-repo-operations/deployment-and-environments.md · Decision 1, amended
Describe every environment in one alchemy program with a parsed stage
Section titled “Describe every environment in one alchemy program with a parsed stage”Impact: HIGH reviewed, type-checked, and destroy is exact
- The stage is parsed into
{ kind: "prd" } | { kind: "stg" } | { kind: "dev" }. Any other name fails the deploy.devis accepted only underalchemy dev(ALCHEMY_DEV === "true"), uses local state, and exists on the Worker profile only. - Every environment difference is a
switch (stage.kind)inside the program. Nothing is set in a dashboard, and no workflow passes per-environment infrastructure flags. - Container platform settings live here: 30s stop timeout, load balancer health check on
/ready(a restarting probe uses/health), 60s grace period, nopreStop, exit 130 after SIGTERM is a normal stop,TZ=UTC. alchemy is pinned exactly, and every deploy job hastimeout-minutes.
❌ Incorrect — the stage is a free string, and differences live elsewhere:
const stage = process.env.STAGE ?? "dev" // a typo deploys a new environmentconst replicas = stage === "production" ? 2 : 1 // plus settings edited in a dashboard✅ Correct — one program, a parsed union:
type Stage = { kind: "prd" } | { kind: "stg" } | { kind: "dev" }
export default Alchemy.Stack("app", { providers, state: isDevServer ? Alchemy.localState() : remoteState }, Effect.gen(function* () { const stage = parseStage(yield* Alchemy.Stage) // replicas, domains, database wiring, env: one `switch (stage.kind)` each }))Source: notes/11-repo-operations/deployment-and-environments.md · Decision 2, amended
Promote one SHA: CI, then staging, then an approval
Section titled “Promote one SHA: CI, then staging, then an approval”Impact: HIGH a red main never ships; staging is a real gate
- Deploys start from CI’s completion (
workflow_run, deployinghead_sha), never frompush.deploy-prdruns only afterdeploy-stgsucceeded for the same SHA, in onedeploy.yml. - The
productionGitHub environment has a required reviewer and acceptsmainonly. Approve the newest waiting run; never reject an older one (a rejection turns the run red), and the concurrency group cancels the older ones. Concurrency is one group per environment,cancel-in-progress: false. - The smoke check runs on
always()after the deploy step and asserts 200 andx-app-revisionequal to the deployed SHA, retried up to 6 × 10s. Thedeployment-gateis red when staging failed or skipped because CI failed, or production failed; a superseded run is grey. A successfuldeploy-prdis a release.
❌ Incorrect — deploy on push, before CI has said anything:
on: push: branches: [main]jobs: deploy: # a red main ships to production✅ Correct — the tested SHA, through staging, behind a reviewer:
push to main (a1b2c3d) CI ✅ └─ deploy-stg [staging] migrate → alchemy deploy --stage stg → smoke → pending notes └─ deploy-prd [production, required reviewer] migrate → alchemy deploy --stage prd → smoke └─ release [release] tag → GitHub Release → compat snapshotSource: notes/11-repo-operations/deployment-and-environments.md · Decision 3, amended
Run migrate as a step in the deploy job, before the deploy
Section titled “Run migrate as a step in the deploy job, before the deploy”Impact: HIGH guarantees order; the owner credential never reaches the runtime
pnpm run migrateruns beforealchemy deployin each deploy job, overDATABASE_OWNER_URL(a direct connection, never a pooler). A failed migrate stops the job and the old release keeps serving.- The owner credential lives only in the environment’s secrets and only in the
migratestep’senv. The runtime gets the DML-only role. Default privileges for that role are set by the first migration. - On the Worker profile, alchemy’s D1 resource applies migrations during
alchemy deploy, before the Worker that binds it. See database migrations.
❌ Incorrect — the owner credential handed to the whole job, deploy unguarded:
deploy-stg: env: { DATABASE_OWNER_URL: ${{ secrets.DATABASE_OWNER_URL }} } # visible to every step steps: - run: alchemy deploy --stage stg - run: pnpm run migrate # after the new code is already serving✅ Correct — migrate first, credential scoped to its step:
deploy-stg: environment: staging steps: - run: pnpm run migrate # DATABASE_OWNER_URL in this step's env only - run: alchemy deploy --stage stg # never runs if migrate failedSource: notes/11-repo-operations/deployment-and-environments.md · Decision 4, amended
Keep secrets in one GitHub Environment per stage
Section titled “Keep secrets in one GitHub Environment per stage”Impact: HIGH the environment is the secret scope and the approval gate at once
productionhas a required reviewer and acceptsmainonly.stagingholds the same names with staging values.release(no reviewer,mainonly) holds only the snapshot GitHub App’s key.- The deploy job declares its
environment:. alchemy writes runtime secrets into the platform’s secret store, and the container receives them as environment variables, read by config and secrets. - Every environment has its own
SECRET_KEYS. Production keys are never copied to staging, so a value sealed in one environment cannot be opened in another.
❌ Incorrect — one set of keys shared across environments:
staging.SECRET_KEYS = <copied from production> # staging can open production's sealed values✅ Correct — one environment per stage, each with its own keys:
environment reviewer branches holdsproduction required main prd owner URL, runtime secrets, SECRET_KEYS, deploy credsstaging — main the same names, staging valuesrelease — main the snapshot GitHub App's key onlySource: notes/11-repo-operations/deployment-and-environments.md · Decision 7, amended
Run production mode everywhere and tell each environment its name
Section titled “Run production mode everywhere and tell each environment its name”Impact: MEDIUM staging rehearses production’s code path, telemetry still tells them apart
APP_ENVis the mode:productionin every deployed environment,developmentlocally and in tests.DEPLOYMENT_ENVIRONMENTis the name stamped asdeployment.environment.name. It has no default.- The stack sets
SHUTDOWN_DELAY: "5 seconds"in every deployed stage (no default in code) andSERVICE_VERSIONto the deployed SHA. - No code branches on
DEPLOYMENT_ENVIRONMENT. Behaviour differs only through configuration values.
❌ Incorrect — a widened mode that code branches on:
if (env.APP_ENV === "staging") { /* a code path production never runs */ }✅ Correct — one mode, a name for telemetry, set by the stack:
env: { APP_ENV: "production", // stg and prd alike DEPLOYMENT_ENVIRONMENT: formatStage(stage), // "production" | "staging" SHUTDOWN_DELAY: "5 seconds", SERVICE_VERSION: sha,}Source: notes/11-repo-operations/deployment-and-environments.md · Decision 8, amended
Undo code by redeploying an earlier SHA; undo data with PITR into a new database
Section titled “Undo code by redeploying an earlier SHA; undo data with PITR into a new database”Impact: HIGH minutes to recover, and backup is the only undo for data
- Rollback is a
workflow_dispatchofdeploy-prdwith ashathat the deployments API shows was in production before. It needs the reviewer, skipsmigrate, and creates no tag. A revert PR is mandatory, and nothing else merges until it lands. - Production’s Postgres has provider-managed PITR with at least 7 days’ retention. A restore goes
into a new database; lost rows are copied back by a reviewed script, or
DATABASE_URLis pointed at it through a normal deploy. On D1, the provider’s own restore replaces PITR. - A restore drill runs once before launch: restore into a temporary database, migrate, boot the current release, time it, delete it.
❌ Incorrect — roll back by migrating down, or restore over production:
pnpm run migrate:down # migrations are forward-onlypg_restore --clean -d prod … # overwrites the live database✅ Correct — redeploy the old SHA on the current schema, then revert:
workflow_dispatch deploy-prd sha=<a SHA previously deployed to production> → [production, reviewer] → skip migrate → alchemy deploy --stage prd → smokethen: PR "revert: …" through the normal pipelineSource: notes/11-repo-operations/deployment-and-environments.md · Decisions 9–10, amended
Deferred
Section titled “Deferred”- Label-gated per-PR previews (
pr-<n>, a shared preview database server, event teardown plus a 6-hourly sweep) — trigger: a second human reviewer who does not run branches locally. migrateas a one-off task inside the network — trigger: the chosen platform keeps the database on a private network the CI runner cannot reach.- A longer health-check grace period — trigger: twice the slowest boot ever observed exceeds 60s.
- The quarterly restore drill — trigger: production holds its first customer data.