feat(deploy): drop the IaC layer, and make both apps runnable as containers

Moving the issuer's state into Postgres removed the last thing that tied
either app to one hosting provider. What was left was a deployment tool
describing resources that no longer existed — so this replaces it with
`wrangler`, which is what actually deploys a Worker, and adds a second way
to run each app that involves no provider at all.

Each app now has a `wrangler.jsonc` with an environment per stage, and a
`Dockerfile` beside it. The handler is the same one in both cases; what
differs is only where its settings come from. Two of them gained a second
spelling so that nothing has to branch on the runtime: Postgres arrives as
a pooled binding or as `DATABASE_URL`, and the route to the issuer is a
service binding or `AUTH_INTERNAL_URL`.

That last one is new, and it is a split the binding was already making
without saying so. `AUTH_ISSUER_URL` has to be the issuer's public name,
because it is compared literally against every token's `iss` claim — but
the public name is often not routable from inside a deployment. So the
name and the route are two settings now rather than one that cannot be
both.

DNS moves out of code and into `docs/dns.md`, which lists every hostname
and what it is for. Six records that change roughly never did not need a
tool, and the table outlives whatever is answering the names — which is
the point, since some of them will stop being Workers. The sandbox
hostnames are hyphenated rather than nested for the same reason: a
certificate covering `*.nestri.io` covers one label and not two, so
`api-sandbox.nestri.io` can become an ordinary origin later without a
certificate having to be ordered for it first.

Also drops `EMAIL_DEV_LOG` from committed configuration into `.dev.vars`,
which `wrangler deploy` cannot upload. Printing a live sign-in code to a
log should not be one forgotten override away from production.
This commit is contained in:
Wanjohi
2026-09-05 15:27:56 +03:00
parent 3d0dcf3e46
commit 51ababc900
28 changed files with 1001 additions and 1322 deletions

View File

@@ -1,345 +0,0 @@
# Alchemy — infrastructure as code
This project uses [Alchemy](https://alchemy.run) (v0.93.12) for infrastructure-as-code — the equivalent of SST, but targeting Cloudflare Workers instead of AWS Lambda.
## Project structure
```
web/
alchemy.run.ts # Entry point — creates scope, imports infra
infra/
stage.ts # Stage detection (Scope.getCurrentScope().stage)
secret.ts # Encrypted secrets via alchemy.secret()
auth.ts # Auth Worker resource
api.ts # API Worker resource
```
## `alchemy.run.ts` — entry point
```ts
import alchemy from 'alchemy';
const app = await alchemy('nestri', {
password: process.env.ALCHEMY_PASSWORD // required for secrets
});
// Import infra modules in dependency order (SST-style)
await import('./infra/stage.ts');
await import('./infra/secret.ts');
await import('./infra/auth.ts');
await import('./infra/api.ts');
await app.finalize();
```
Key rules:
- `alchemy(appName, opts)` creates a **scope** — resources register into this scope automatically
- `app.finalize()` must be called at the end to persist state
- Import order matters — resources that depend on others must be imported after
- `--dev` flag runs locally via Miniflare; omit it to deploy to Cloudflare
## Infra resources
Each resource is imported from `alchemy/cloudflare` and called with an ID + props:
```ts
import { Worker, KVNamespace, D1Database } from 'alchemy/cloudflare';
export const kv = await KVNamespace('my-kv');
export const db = await D1Database('my-db');
export const worker = await Worker('my-worker', {
entrypoint: 'apps/some-app/src/index.ts',
compatibility: 'node', // enables nodejs_compat flag
url: true, // assign workers.dev URL
bindings: {
KV: kv, // resource binding → KVNamespace at runtime
DB: db, // → D1Database
PLAIN_VAR: 'hello' // → plain_text binding
}
});
```
### Supported resources (subset)
| Resource | Import | Purpose |
| ------------- | -------------------- | ----------------------------------------------- |
| `Worker` | `alchemy/cloudflare` | Cloudflare Worker (entrypoint or inline script) |
| `KVNamespace` | `alchemy/cloudflare` | KV storage |
| `D1Database` | `alchemy/cloudflare` | D1 SQL database |
| `R2Bucket` | `alchemy/cloudflare` | R2 object storage |
| `Queue` | `alchemy/cloudflare` | Queue/pub-sub |
### Compatibility flag
Always add `compatibility: 'node'` to Workers that use Node.js built-ins (`node:async_hooks`, `crypto`, `node:stream`, etc.):
```ts
Worker('api', {
entrypoint: 'apps/api/app/index.ts',
compatibility: 'node' // enables nodejs_compat
});
```
## Stage detection
```ts
// infra/stage.ts
import { Scope } from 'alchemy';
const scope = Scope.getCurrentScope();
export const stage = scope?.stage ?? 'dev';
export const isPermanent = ['production', 'dev'].includes(stage);
```
Use stage for conditional infrastructure:
```ts
const api = await Worker('api', {
...(isPermanent && {
observability: { enabled: true },
logpush: true
})
});
```
Pass `--stage` flag at runtime: `bun alchemy.run.ts --stage production`
## Secrets and environment variables
Three levels of env management, from most-secure to least:
### 1. `alchemy.secret.env.X` (preferred)
```ts
// infra/secret.ts
import alchemy from 'alchemy';
export const secret = {
steamApiKey: alchemy.secret.env.STEAM_API_KEY // reads process.env at deploy time
// Equivalent to:
// steamApiKey: alchemy.secret(process.env.STEAM_API_KEY),
};
```
- Reads from `process.env` at deploy time
- Throws a descriptive error if the env var is missing
- Encrypted in Alchemy state files (`.alchemy/`)
- Deployed as `secret_text` binding (hidden from Cloudflare API)
### 2. `alchemy.env()` (non-secret config)
```ts
export const frontendUrl = alchemy.env('FRONTEND_URL', 'http://localhost:5173');
```
- Optional default value
- Plain text — not encrypted
- Deployed as `plain_text` binding
### 3. Plain strings in `bindings` (inline)
```ts
bindings: {
MY_VAR: 'hello';
}
```
- Hard-coded, visible in state files
- Deployed as `plain_text` binding
### How bindings map to runtime types
| Alchemy binding type | Deployed as | Runtime type |
| -------------------- | -------------- | -------------------------- |
| `Worker` | `service` | `Service` (has `.fetch()`) |
| `KVNamespace` | `kv_namespace` | `KVNamespace` |
| `D1Database` | `d1` | `D1Database` |
| `alchemy.secret()` | `secret_text` | `string` |
| plain `string` | `plain_text` | `string` |
| `Json(...)` | `json` | `typeof json` |
## Service bindings (Worker → Worker)
Pass one Worker as a binding to another:
```ts
// infra/auth.ts
export const auth = await Worker('auth', {
entrypoint: 'apps/auth/src/index.ts',
compatibility: 'node',
bindings: { ... },
});
// infra/api.ts
import { auth } from './auth.ts';
export const api = await Worker('api', {
entrypoint: 'apps/api/app/index.ts',
bindings: { AUTH: auth },
});
```
At runtime, `env.AUTH` is a `Service` — call it directly:
```ts
const response = await env.AUTH.fetch(request);
```
### OpenAuth client + service binding
The `@openauthjs/openauth/client` only accepts a URL string for `issuer`, so use a custom `fetch` to route through the service binding:
```ts
function getClient(env: Record<string, unknown>) {
return createClient({
issuer: 'https://auth.internal', // dummy — used for path construction
clientID: 'api',
fetch: (input, init) => {
const url = new URL(typeof input === 'string' ? input : input.url);
const request = new Request(url.pathname + url.search, init);
return (env.AUTH as { fetch: typeof fetch }).fetch(request);
}
});
}
```
## Env propagation to Workers
CF Workers receive env vars as the second argument to the `fetch` handler (`env`), NOT via `process.env`. Bridge the gap with a lazy + overridable schema:
```ts
// packages/core/src/env.ts
import { memo } from '../utils/memo.ts';
let _overrides: Record<string, unknown> = {};
export namespace Env {
export const Info = z.object({
FRONTEND_URL: z.string().optional(),
STEAM_API_KEY: z.string().optional(),
AUTH_ISSUER_URL: z.string().optional()
});
export type Info = z.infer<typeof Info>;
const _get = memo(() => Info.parse({ ...process.env, ..._overrides }));
export function get(): Info {
return _get();
}
export function init(bindings: Record<string, unknown>) {
_overrides = bindings;
_get.reset();
}
}
```
Wire in the Hono entrypoint:
```ts
export default {
fetch(request, env, ctx) {
Env.init(env); // merge CF bindings into Env
return app.fetch(request, env, ctx);
}
};
```
Now any module that imports `Env.get()` gets the correct values — on Bun dev `process.env` provides them, on CF Workers the bindings override.
## CLI usage
```sh
# Local dev (Miniflare)
bun alchemy.run.ts --dev
# Deploy to Cloudflare
bun alchemy.run.ts --stage production
# Destroy all resources
bun alchemy.run.ts --destroy
# With custom stage
bun alchemy.run.ts --stage wanjohiryan
# Password (for encrypting secrets)
export ALCHEMY_PASSWORD="some-passphrase"
```
When deploying, set `CLOUDFLARE_API_TOKEN` or configure `alchemy login`.
## Common patterns
### Conditional infra per-stage
```ts
Worker('api', {
...(isPermanent && { logpush: true }),
...(stage === 'production' && { scaling: { min: 3, max: 10 } })
});
```
### Across-app resource references
Alchemy uses top-level await in infra files — resources resolve at import time within the active scope. The scope propagates via `AsyncLocalStorage`, so any `await import()` after `alchemy(appName)` picks it up.
### .alchemy/ directory
Created automatically — contains Miniflare state, build output, and encrypted state files. Add to `.gitignore`.
```gitignore
.alchemy/
```
---
### Index Rule: Null-Safe Exclusions (`IS DISTINCT FROM`)
When writing indices to track data drift, synchronization deltas, or pending background worker states where values might be nullable, **always build a partial index utilizing Postgres-native `IS DISTINCT FROM`**.
Standard inequality operators (`!=` or `<>`) evaluate to `NULL` if either column is `NULL`, causing them to bypass standard `WHERE` index filters. Using `is distinct from` allows Postgres to treat `NULL` as a real value for state comparison:
- Excludes perfectly synchronized records completely from the index footprint.
- Optimizes heavy background worker poll queries directly into small, lightning-fast index scans.
2. Add to the Pattern: index.ts (Domain Namespace) section
Replace the existing create block and add the upsert block inside SomeModule:
```ts
// ── create ───────────────────────────────────────────────────────────
// Use Info.pick({…}) for the schema — keeps fields in sync with Info.
// Always use .returning() to get the updated row context in one database trip.
export const create = fn(Info.pick({ id: true, name: true, email: true }), async (input) => {
return Database.use(async (tx) => {
const [row] = await tx
.insert(SomeTable)
.values({
id: input.id,
name: input.name,
email: input.email ?? null
})
.returning();
return row;
});
});
// ── upsert ───────────────────────────────────────────────────────────
// Simple copies use the input values directly. For coalesce-style set
// expressions, reference the excluded pseudo-table with unqualified
// identifiers: sql`excluded.${sql.identifier(SomeTable.name.name)}` —
// interpolating a column object (or its .name string) is invalid.
export const upsert = fn(Info.pick({ id: true, name: true }), async (input) => {
return Database.use(async (tx) => {
const [row] = await tx
.insert(SomeTable)
.values({ id: input.id, name: input.name })
.onConflictDoUpdate({
target: SomeTable.id,
set: { name: input.name }
})
.returning();
return row;
});
});
```

147
docs/deploy.md Normal file
View File

@@ -0,0 +1,147 @@
# Deploying the control plane
Two apps — [`apps/api`](../apps/api) and [`apps/auth`](../apps/auth) — and two
ways to run each of them. There is one handler per app and it is the same
handler both ways: a function from a request to a response, holding no opinion
about what is calling it.
| | |
| --- | --- |
| **Cloudflare Workers**, via `wrangler` | what production and sandbox are today |
| **A container**, via the `Dockerfile` in each app | what a self-hoster runs, and where this is going |
Settings arrive as bindings in the first case and as environment variables in
the second, and `@nestri/core`'s `Env` resolves the two into one shape — so
`HYPERDRIVE` and `DATABASE_URL` are two spellings of the database, and an
`AUTH` service binding and `AUTH_INTERNAL_URL` are two spellings of the route
to the issuer. Nothing in either app branches on which it got.
Hostnames, and why they are shaped the way they are: [`dns.md`](dns.md).
## Locally
Three ways, in increasing order of how much they resemble a deployment.
```sh
docker compose up postgres # the database, for either of the next two
bun dev # both apps under the Workers runtime
bun run dev:server # both apps as plain processes
docker compose up --build # both apps as containers, plus the database
```
`bun dev` runs two `wrangler dev` sessions, on ports 1337 and 3000. They find
each other through wrangler's local registry, so the API reaches the issuer
over the same service binding it uses in production rather than over the
network — which is the point of running it this way. Neither needs a
Cloudflare account: the Hyperdrive binding falls back to
`localConnectionString`, which is the compose database.
Settings that exist only locally live in `apps/auth/.dev.vars` rather than in
`vars`. `wrangler dev` reads that file and `wrangler deploy` cannot upload it,
which is the guarantee wanted for the one setting in it — the one that prints
sign-in codes to the log.
`docker compose` here is Docker's plugin or `podman-compose`; both read the
file unchanged.
Migrations are never run for you, in any of the three:
```sh
bun run db:migrate # against DATABASE_URL
```
## Cloudflare Workers
Configuration is [`apps/auth/wrangler.jsonc`](../apps/auth/wrangler.jsonc) and
[`apps/api/wrangler.jsonc`](../apps/api/wrangler.jsonc). Each has two named
environments, `sandbox` and `production`, plus an unnamed default that is the
local one.
Wrangler's named environments do **not** inherit bindings from the top level —
`vars`, `services` and `hyperdrive` are repeated in each on purpose, and a
setting added to one environment and not the other is a silent hole rather than
an error.
### One-time setup
```sh
bunx wrangler login
# Once per database. Prints an id; paste it into both wrangler.jsonc files,
# replacing the placeholder for that environment.
bunx wrangler hyperdrive create nestri-production --connection-string "postgres://…"
bunx wrangler hyperdrive create nestri-sandbox --connection-string "postgres://…"
```
Hyperdrive is a connection pool in front of Postgres, and it is there because
each Worker isolate would otherwise open a connection of its own — which
Postgres answers, at some point in a busy hour, with *"sorry, too many clients
already"*. A container has one pool per process and needs none of this.
### Secrets
Set per app and per environment, and held by Cloudflare rather than by this
repository:
```sh
cd apps/auth
bunx wrangler secret put EMAIL_SEND_URL --env production
bunx wrangler secret put EMAIL_API_KEY --env production
bunx wrangler secret put EMAIL_FROM --env production
cd ../api
bunx wrangler secret put STEAM_API_KEY --env production
bunx wrangler secret put ADMIN_SHARED_SECRET --env production
```
The issuer refuses to send a sign-in code with its mail settings half
configured or absent, rather than falling back to printing codes to the log —
so a deployment that forgets these fails at the first sign-in attempt with a
message naming what is missing, instead of quietly logging usable codes.
### Deploying
```sh
bun run deploy:sandbox
bun run deploy:production
```
Both deploy the issuer first and the API second, because the API's `AUTH`
binding names a script that has to exist. The custom domains in the config are
what create the DNS records — there is no separate step, and no separate tool
holding the other half of that fact.
## Containers
```sh
docker build -f apps/api/Dockerfile -t nestri-api .
docker build -f apps/auth/Dockerfile -t nestri-auth .
```
The context is the repository root in both cases: the lockfile and the two
shared packages are there, and a context rooted at the app directory could not
reach them. Both use the repository-wide `.dockerignore`; only the guest rootfs
build has one of its own, as `build/Dockerfile.dockerignore` — a
`<Dockerfile>.dockerignore` **replaces** the repository-wide file rather than
adding to it, which is worth knowing before writing a third.
Both images are stateless and hold no configuration. What they need:
| | `auth` | `api` |
| --- | --- | --- |
| `DATABASE_URL` | required | required |
| `AUTH_ISSUER_URL` | — | required, the issuer's **public** URL |
| `AUTH_INTERNAL_URL` | — | only if that URL is unroutable from here |
| `EMAIL_SEND_URL` `EMAIL_API_KEY` `EMAIL_FROM` | all three, or none | — |
| `EMAIL_DEV_LOG` | `true` prints codes instead of sending | — |
| `STEAM_API_KEY` | — | to link a Steam account |
| `ADMIN_SHARED_SECRET` | — | operator access |
| `PORT` | default `1337` | default `3000` |
[`docker-compose.yml`](../docker-compose.yml) at the root wires all of it
together with a Postgres, and is the smallest complete answer to *"how do I run
this myself"*.
Neither image terminates TLS or serves a certificate. Put a reverse proxy in
front of them, point the hostnames at it, and keep the origin unreachable
except through it.

69
docs/dns.md Normal file
View File

@@ -0,0 +1,69 @@
# DNS
Cloudflare holds the zones. Once the control plane moves off Workers that is
the only thing it holds, so this file is deliberately written to survive the
move: it says what each name **is for**, and treats what currently answers it
as a detail that changes.
There is no infrastructure-as-code here, on purpose. There are six records.
They change roughly never, they outlive several generations of whatever serves
them, and the failure mode of getting one wrong is that sign-in stops working
for everybody — which is a thing to do slowly, by hand, having read this table,
rather than as a side effect of a deploy. What *is* automated is only the part
that must stay in step with a deploy: while the control plane is a set of
Workers, `wrangler` creates and owns the four control-plane records itself,
because a route and its hostname are one fact and splitting them across two
tools is how they drift.
## The rule
**One label deep on `nestri.io`.** A certificate for `*.nestri.io` covers
`api-sandbox.nestri.io` and does not cover `api.sandbox.nestri.io`, and that is
the whole reason the sandbox names are hyphenated rather than nested. It costs
nothing while these are Workers — a custom domain gets its own certificate for
the exact hostname either way — and it is what lets any of these names become
an ordinary proxied origin later without also needing a certificate ordered for
it. A name should not have to change because the thing behind it did.
## `nestri.io`
| Name | What it is | Answered today by |
| ------------------------ | --------------------------------- | ----------------------- |
| `api.nestri.io` | The API, production | Worker custom domain |
| `auth.nestri.io` | The issuer, production | Worker custom domain |
| `api-sandbox.nestri.io` | The API, sandbox | Worker custom domain |
| `auth-sandbox.nestri.io` | The issuer, sandbox | Worker custom domain |
| `doctor.nestri.io` | Where `nesdoctor` is downloaded | Static site |
| `nestri.io` | The website, and `ssh nestri.io` | Website |
`auth.nestri.io` is the one name that cannot be changed casually. A token
carries the address it was minted through in its `iss` claim, and every API
request verifies that claim literally — so renaming the issuer invalidates
every token in circulation at once, including the refresh tokens that would
otherwise have recovered from it.
## After the move off Workers
Each of the first four becomes a proxied `A` record pointing at the host
running the containers, and nothing else about them changes: same names, same
certificates, same `iss` claim. Cloudflare keeps terminating public TLS, so
there is no certificate on our own host to renew, and the origin is not
addressable except through the proxy.
The order that matters, on the day: create the `A` records with the proxy on,
confirm the containers answer through them, *then* remove the Worker routes.
Doing it the other way leaves a window where the name resolves to nothing.
## `nestri.link`
A second zone, reserved and not yet serving anything. It exists so that a
per-box hostname — one name, one box, the address a person opens to set their
box up — never has to live under `nestri.io` beside the control plane. Two
reasons, both of which get worse to fix later than to decide now: a box serves
content we do not write, and cookie scope is a property of the registrable
domain, so a name under `nestri.io` would put that content inside the same
cookie boundary as sign-in.
`*.nestri.link` will be proxied for the same reason the control plane is: the
public certificate stays Cloudflare's, and the only key on our own host is an
origin certificate that is useless anywhere else.