fix(deploy): require every credential, and give sandbox its own domain

Three things review caught, and one shape correction.

**No credential has a default any more.** The compose file shipped
`ADMIN_SHARED_SECRET` falling back to a value written in this repository —
and that header bypasses token verification entirely, so anyone reading
the file could act as an operator against any deployment that had not
overridden it. A default is worth less than it looks here: the deployment
that never set the variable is exactly the one where the default is public.
Every credential now comes from `.env`, and compose refuses to start naming
the variable it wanted. That also takes the last literal password out of a
tracked file.

**The origin ports are on loopback.** Both services speak plain HTTP and
mark no cookie `Secure`, because both expect to sit behind something that
terminates TLS. Published on every interface they were a way to reach the
issuer around that proxy, with sign-in codes and tokens in clear text.

**Mail settings are passed through rather than fixed.** The issuer was
pinned to printing sign-in codes to its log, and the three delivery
settings never reached it — so the documented way to configure mail could
not work, and every code and recipient went to the container log instead.
Printing codes is now asked for in `.env` like everything else, and with
nothing configured the issuer refuses to send rather than logging.

**Sandbox becomes a domain rather than a prefix.** `api.sandbox.nestri.io`
and `auth.sandbox.nestri.io`, because sandbox holds whatever is not
production and that set grows. One certificate for `*.sandbox.nestri.io`
then covers all of it, including unpredictable per-pull-request names,
and cannot be presented for production's own domain — which the zone-wide
wildcard the previous shape leaned on could.

Also drops `STEAM_API_KEY`. It was declared in two type definitions and
read by nothing: linking an account makes no outbound call that needs it.
This commit is contained in:
Wanjohi
2026-09-05 15:58:21 +03:00
parent 51ababc900
commit f30a1432f8
11 changed files with 133 additions and 63 deletions

View File

@@ -1,4 +1,14 @@
# Local development database (docker compose up postgres)
# Copy to `.env` before `docker compose up`. Compose reads every credential
# from here and has no defaults of its own — it refuses to start naming the
# variable it wanted rather than falling back to a value that would be public.
# The local database. Throwaway values are fine; these three are what compose
# creates the container with and what it builds DATABASE_URL from.
POSTGRES_USER=postgres
POSTGRES_PASSWORD=postgres
POSTGRES_DB=nestri
# For anything run outside a container — `bun dev`, `bun run db:migrate`.
DATABASE_URL=postgres://postgres:postgres@localhost:5432/nestri
# Isolated database for tests. Required by DB-backed tests; unset tests fail.
@@ -10,15 +20,14 @@ AUTH_ISSUER_URL=http://localhost:1337
# public name is unroutable from where the API runs; docker compose sets it.
AUTH_INTERNAL_URL=
# Linking a Steam account
STEAM_API_KEY=
# Operator access to the API
# Turns any request carrying it into an operator, so generate one rather than
# typing something: `openssl rand -hex 32`.
ADMIN_SHARED_SECRET=
# Mail delivery. All three together, or none of them plus EMAIL_DEV_LOG=true,
# which prints sign-in codes to the log instead of sending them.
# which prints sign-in codes to the log instead of sending them. Printing them
# is a local-development convenience and nothing else.
EMAIL_SEND_URL=
EMAIL_API_KEY=
EMAIL_FROM=
EMAIL_DEV_LOG=
EMAIL_DEV_LOG=true

View File

@@ -114,6 +114,7 @@ sandboxes, one GPU" possible instead of one tenant per card.
```sh
bun install
cp .env.example .env # compose reads every credential from here
docker compose up postgres # the database
bun run db:migrate # schema
bun dev # control plane, local Cloudflare runtime

View File

@@ -39,7 +39,12 @@ test/ # Route tests
## Key details
- Auth: `Authorization: Bearer <JWT>` verified against `@nestri/auth`; or the `x-nestri-admin-token` header (see `ADMIN_SHARED_SECRET`).
- Auth: `Authorization: Bearer <JWT>` verified against `@nestri/auth`; or the `x-nestri-admin-token` header
carrying `ADMIN_SHARED_SECRET`, which bypasses JWT verification entirely and is required — it has no
default anywhere. It is what authenticates the callers that have no user identity to present:
`POST /pairing-code/claim` (a device being paired has no identity yet, which is the whole point),
`POST /games`, `POST /games/sync`, `POST /library/sync`, `GET /waitlist`, `POST /steam/link` on behalf
of another user, and `POST /games/download-state` when an operator is repairing state a box reported.
- Errors: centralized `VisibleError` → typed JSON responses.
- Settings arrive as bindings or as environment variables, and two of them have one spelling of
each: Postgres is `HYPERDRIVE` or `DATABASE_URL`, and the route to the issuer is an `AUTH`

View File

@@ -123,7 +123,6 @@ export type ApiEnv = {
AUTH_INTERNAL_URL?: string;
HYPERDRIVE?: Hyperdrive;
DATABASE_URL?: string;
STEAM_API_KEY?: string;
ADMIN_SHARED_SECRET?: string;
};

View File

@@ -44,11 +44,11 @@
"sandbox": {
"name": "nestri-api-sandbox",
"workers_dev": false,
"routes": [{ "pattern": "api-sandbox.nestri.io", "custom_domain": true }],
"routes": [{ "pattern": "api.sandbox.nestri.io", "custom_domain": true }],
"observability": { "enabled": true },
"services": [{ "binding": "AUTH", "service": "nestri-auth-sandbox" }],
"hyperdrive": [{ "binding": "HYPERDRIVE", "id": "<sandbox-hyperdrive-id>" }],
"vars": { "AUTH_ISSUER_URL": "https://auth-sandbox.nestri.io" }
"vars": { "AUTH_ISSUER_URL": "https://auth.sandbox.nestri.io" }
},
"production": {
"name": "nestri-api",

View File

@@ -44,7 +44,7 @@
"sandbox": {
"name": "nestri-auth-sandbox",
"workers_dev": false,
"routes": [{ "pattern": "auth-sandbox.nestri.io", "custom_domain": true }],
"routes": [{ "pattern": "auth.sandbox.nestri.io", "custom_domain": true }],
"observability": { "enabled": true },
"hyperdrive": [{ "binding": "HYPERDRIVE", "id": "<sandbox-hyperdrive-id>" }]
},

View File

@@ -5,27 +5,43 @@
# stateless processes and a database, with a reverse proxy in front of them
# terminating TLS. Nothing here knows about a hosting provider.
#
# cp .env.example .env # then fill it in
# docker compose up --build everything, built from source
# docker compose up postgres just the database, for `bun dev`
#
# **There are no credentials in this file, and none of them have defaults.**
# Every one is read from `.env`, and compose refuses to start naming the
# variable it wanted rather than falling back to something. A default is worth
# less than it looks: the deployment that never set the variable is exactly the
# one where the default is a publicly known value, and `ADMIN_SHARED_SECRET`
# below bypasses authentication entirely.
#
# Migrations are not run for you — `bun run db:migrate` against DATABASE_URL,
# because a container that migrates on boot races with the second copy of
# itself and there is eventually a second copy.
x-postgres-url: &postgres-url
DATABASE_URL: postgres://${POSTGRES_USER:?set POSTGRES_USER in .env}:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}@postgres:5432/${POSTGRES_DB:?set POSTGRES_DB in .env}
services:
postgres:
image: docker.io/postgres:18-alpine
container_name: nestri_postgres
environment:
POSTGRES_USER: postgres
POSTGRES_PASSWORD: postgres
POSTGRES_DB: nestri
POSTGRES_USER: ${POSTGRES_USER:?set POSTGRES_USER in .env}
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD in .env}
POSTGRES_DB: ${POSTGRES_DB:?set POSTGRES_DB in .env}
# Loopback, not every interface. `5432:5432` would publish the database to
# anything that can reach this host. The three services below talk to each
# other over the compose network and do not use this mapping at all; it is
# here only so `bun dev` and `bun run db:migrate` can reach the database
# from outside a container.
ports:
- '5432:5432'
- '127.0.0.1:5432:5432'
volumes:
- nestri_data:/var/lib/postgresql
healthcheck:
test: ['CMD-SHELL', 'pg_isready -U postgres -d nestri']
test: ['CMD-SHELL', 'pg_isready -U "$$POSTGRES_USER" -d "$$POSTGRES_DB"']
interval: 5s
timeout: 5s
retries: 10
@@ -33,7 +49,7 @@ services:
auth:
build:
# The repository root, because the lockfile and the shared packages are
# there. Same reason for both images below.
# there. Same reason for the API below.
context: .
dockerfile: apps/auth/Dockerfile
container_name: nestri_auth
@@ -41,13 +57,24 @@ services:
postgres:
condition: service_healthy
environment:
DATABASE_URL: postgres://postgres:postgres@postgres:5432/nestri
<<: *postgres-url
# Passed through rather than fixed here, so that setting them in `.env`
# is enough to make this deployment deliver mail. All three together or
# none of them: the issuer refuses to send when they are half configured.
EMAIL_SEND_URL: ${EMAIL_SEND_URL:-}
EMAIL_API_KEY: ${EMAIL_API_KEY:-}
EMAIL_FROM: ${EMAIL_FROM:-}
# Printing a live sign-in code to the log is a thing you ask for by name,
# and this file is the local machine. Set the three EMAIL_* settings
# instead and codes are delivered rather than printed.
EMAIL_DEV_LOG: 'true'
# and it is asked for in `.env` — not defaulted to here. With mail
# unconfigured and this unset, the issuer refuses to send rather than
# logging codes, which is the failure a self-hoster should get.
EMAIL_DEV_LOG: ${EMAIL_DEV_LOG:-}
# Loopback. This listener speaks plain HTTP and sets no `Secure` on the
# cookies it issues, because it expects to be behind something that
# terminates TLS. Published on every interface it would be a way to reach
# the issuer *around* that proxy, with codes and tokens in clear text.
ports:
- '1337:1337'
- '127.0.0.1:1337:1337'
api:
build:
@@ -60,17 +87,19 @@ services:
auth:
condition: service_started
environment:
DATABASE_URL: postgres://postgres:postgres@postgres:5432/nestri
<<: *postgres-url
# The issuer's public URL, and not `http://auth:1337`. A token carries
# the address it was minted through, and verification compares the two
# literally — so the name a browser used is the only one that can appear
# here. `AUTH_INTERNAL_URL` is how this container actually gets there.
AUTH_ISSUER_URL: http://localhost:1337
AUTH_ISSUER_URL: ${AUTH_ISSUER_URL:?set AUTH_ISSUER_URL in .env}
AUTH_INTERNAL_URL: http://auth:1337
STEAM_API_KEY: ${STEAM_API_KEY:-}
ADMIN_SHARED_SECRET: ${ADMIN_SHARED_SECRET:-dev-admin-shared-secret-change-in-prod}
# A shared secret that turns any request carrying it into an operator.
# Required, with no default, for that reason.
ADMIN_SHARED_SECRET: ${ADMIN_SHARED_SECRET:?set ADMIN_SHARED_SECRET in .env to a value you generated}
# Loopback, for the same reason as the issuer above.
ports:
- '3000:3000'
- '127.0.0.1:3000:3000'
volumes:
nestri_data:

View File

@@ -23,6 +23,7 @@ Hostnames, and why they are shaped the way they are: [`dns.md`](dns.md).
Three ways, in increasing order of how much they resemble a deployment.
```sh
cp .env.example .env # once; compose has no credentials of its own
docker compose up postgres # the database, for either of the next two
bun dev # both apps under the Workers runtime
bun run dev:server # both apps as plain processes
@@ -90,10 +91,14 @@ bunx wrangler secret put EMAIL_API_KEY --env production
bunx wrangler secret put EMAIL_FROM --env production
cd ../api
bunx wrangler secret put STEAM_API_KEY --env production
bunx wrangler secret put ADMIN_SHARED_SECRET --env production
```
`ADMIN_SHARED_SECRET` turns any request carrying it into an operator, so
generate it rather than choosing it — `openssl rand -hex 32` — and never give
it a default anywhere. What it is for is listed in
[`apps/api/README.md`](../apps/api/README.md).
The issuer refuses to send a sign-in code with its mail settings half
configured or absent, rather than falling back to printing codes to the log —
so a deployment that forgets these fails at the first sign-in attempt with a
@@ -134,14 +139,16 @@ Both images are stateless and hold no configuration. What they need:
| `AUTH_INTERNAL_URL` | — | only if that URL is unroutable from here |
| `EMAIL_SEND_URL` `EMAIL_API_KEY` `EMAIL_FROM` | all three, or none | — |
| `EMAIL_DEV_LOG` | `true` prints codes instead of sending | — |
| `STEAM_API_KEY` | — | to link a Steam account |
| `ADMIN_SHARED_SECRET` | — | operator access |
| `ADMIN_SHARED_SECRET` | — | required; operator access |
| `PORT` | default `1337` | default `3000` |
[`docker-compose.yml`](../docker-compose.yml) at the root wires all of it
together with a Postgres, and is the smallest complete answer to *"how do I run
this myself"*.
Neither image terminates TLS or serves a certificate. Put a reverse proxy in
front of them, point the hostnames at it, and keep the origin unreachable
except through it.
Neither image terminates TLS or serves a certificate, and neither marks the
cookies it sets `Secure`, because both expect to sit behind something that does
terminate TLS. So put a reverse proxy in front of them and keep the origin
unreachable except through it — `docker-compose.yml` publishes their ports on
loopback only for exactly this reason, and changing that to `0.0.0.0` is a way
to reach the issuer *around* the proxy with codes and tokens in clear text.

View File

@@ -17,38 +17,61 @@ tools is how they drift.
## The rule
**One label deep on `nestri.io`.** A certificate for `*.nestri.io` covers
`api-sandbox.nestri.io` and does not cover `api.sandbox.nestri.io`, and that is
the whole reason the sandbox names are hyphenated rather than nested. It costs
nothing while these are Workers — a custom domain gets its own certificate for
the exact hostname either way — and it is what lets any of these names become
an ordinary proxied origin later without also needing a certificate ordered for
it. A name should not have to change because the thing behind it did.
**A domain gets one certificate, obtained once, and it covers that domain and
nothing else.** Names are then grouped so that the grouping is the same shape
as the certificate: production sits directly under `nestri.io`, and everything
that is not production sits under `sandbox.nestri.io`.
That is why the sandbox names are nested rather than hyphenated. `sandbox` is a
domain, not a prefix — it holds whatever is not production, which today is the
API and the issuer and later is more. Once the shape is a domain, a single
certificate for `*.sandbox.nestri.io` covers all of it, including per-pull-
request deployments at `pr-<id>.sandbox.nestri.io` if those ever arrive; those
would be unbounded and unpredictable names, which is precisely the case that a
name-by-name certificate cannot serve and a domain-wide one can.
It also means a certificate that can be presented for a sandbox name cannot be
presented for `api.nestri.io`. Leaning on the zone-wide `*.nestri.io` instead
would have given every scratch deployment a certificate for production's own
domain, which is the opposite of what a sandbox is for.
Nothing extra is needed while these are Workers — a custom domain is issued its
own certificate for the exact hostname, at any depth. The rule binds on the day
they become origins, and it is written down now because that is the day it is
expensive to have got wrong.
## `nestri.io`
| Name | What it is | Answered today by |
| ------------------------ | --------------------------------- | ----------------------- |
| ------------------------ | -------------------------------- | -------------------- |
| `nestri.io` | The website, and `ssh nestri.io` | Website |
| `api.nestri.io` | The API, production | Worker custom domain |
| `auth.nestri.io` | The issuer, production | Worker custom domain |
| `api-sandbox.nestri.io` | The API, sandbox | Worker custom domain |
| `auth-sandbox.nestri.io` | The issuer, sandbox | Worker custom domain |
| `doctor.nestri.io` | Where `nesdoctor` is downloaded | Static site |
| `nestri.io` | The website, and `ssh nestri.io` | Website |
`auth.nestri.io` is the one name that cannot be changed casually. A token
carries the address it was minted through in its `iss` claim, and every API
request verifies that claim literally — so renaming the issuer invalidates
every token in circulation at once, including the refresh tokens that would
otherwise have recovered from it.
## `sandbox.nestri.io`
Everything that is not production, under one domain and one certificate.
| Name | What it is | Answered today by |
| ------------------------- | ------------------ | -------------------- |
| `api.sandbox.nestri.io` | The API, sandbox | Worker custom domain |
| `auth.sandbox.nestri.io` | The issuer, sandbox| Worker custom domain |
`auth.nestri.io` is the one name in either table that cannot be changed
casually. A token carries the address it was minted through in its `iss` claim,
and every API request verifies that claim literally — so renaming the issuer
invalidates every token in circulation at once, including the refresh tokens
that would otherwise have recovered from it. The sandbox issuer has the same
property and none of the consequences, which is the point of having one.
## After the move off Workers
Each of the first four becomes a proxied `A` record pointing at the host
running the containers, and nothing else about them changes: same names, same
certificates, same `iss` claim. Cloudflare keeps terminating public TLS, so
there is no certificate on our own host to renew, and the origin is not
addressable except through the proxy.
Each of the four control-plane names becomes a proxied `A` record pointing at
the host running the containers, and nothing else about them changes: same
names, same `iss` claim. Cloudflare keeps terminating public TLS, so there is
no certificate on our own host to renew, and the origin is not addressable
except through the proxy.
The order that matters, on the day: create the `A` records with the proxy on,
confirm the containers answer through them, *then* remove the Worker routes.

View File

@@ -401,7 +401,6 @@ export namespace Env {
export const Info = z.object({
NODE_ENV: z.enum(['development', 'production', 'test']).default('development'),
FRONTEND_URL: z.string().optional(),
STEAM_API_KEY: z.string().optional(),
AUTH_ISSUER_URL: z.string().optional() // used by API auth middleware to verify tokens
});
export type Info = z.infer<typeof Info>;
@@ -577,7 +576,7 @@ A Cloudflare Worker using `@nestri/auth` (OpenAuth). Entry point is the `success
Server-to-server calls can authenticate as an `admin` actor by setting the `x-nestri-admin-token` header to the value of `ADMIN_SHARED_SECRET`. This bypasses JWT auth entirely and grants a system-level actor with no user scope — useful for operations like adding games to the DB, syncing data, or other admin tasks.
Configure `ADMIN_SHARED_SECRET` in `.env`; defaults to the same value as `SSH_AUTH_KEY` in dev (`dev-ssh-auth-key-change-in-prod`).
Configure `ADMIN_SHARED_SECRET` in `.env`. It has no default anywhere — a known value here is an authentication bypass, so nothing falls back to one.
### Auth Middleware (`apps/api/app/middleware/auth.ts`)

View File

@@ -8,8 +8,6 @@ export namespace Env {
export const Info = z.object({
NODE_ENV: z.enum(['development', 'production', 'test']).default('development'),
STEAM_API_KEY: z.string().optional(),
AUTH_ISSUER_URL: z.string().optional(),
/**