Files
netris-nestri/build/README.md
Kristian Ollikainen 8246aa5538 feat: resident guest init (#333)
Get this thing going..







<!-- greptile_comment -->

<!-- greptile_summary -->

<h2><a
href="https://app.greptile.com/api/retrigger?id=63134761"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/RetriggerDark.svg?v=1"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=1"><img
alt="Retrigger"
src="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=1"
align="right"></picture></a>Confidence Score: 5/5</h2>

The PR appears safe to merge; all previous findings are resolved and the
latest readiness change introduces no established actionable regression.

<h3>Summary</h3>

- Establishes required guest filesystems, runtime directories, device
permissions, and service processes.
- Reports initialization and service deaths over the lifecycle channel.
- Supports launch, restart, and shutdown commands for a resident guest.
- Separates service and workload identities and configures per-launch
runtime environments.
- Removes the currently inactive nescope screenshot option and makes
capture-chain verification fail explicitly when compositor readback is
unavailable.
- Reworks the guest image around `nesinit` as PID 1 without a
distribution service manager.

<h3>Diagram</h3>

```mermaid
sequenceDiagram
    participant Host
    participant Init as nesinit
    participant FS as Guest filesystems
    participant Services as Service stack
    participant Workload

    Init->>Host: Ready(protocol version)
    Host->>Init: Boot(mount descriptors)
    Init->>FS: Establish and mount shares
    Init->>Services: Spawn services in order
    Services-->>Init: Required sockets ready
    Init->>Host: Initialized(service names)
    Host->>Init: Launch(id, exec, on_exit)
    Init->>Workload: Spawn with isolated UID/runtime
    Init->>Host: Started(id)
    Workload-->>Init: Exit status
    Init->>Host: WorkloadExited(id, status)
    Host->>Init: Launch / Restart / Shutdown
```

<sub>Reviews (4) · Last reviewed commit: ["fix(nesinit): readiness is a
socket
that..."](731d34df9d)</sub>

<!-- /greptile_comment -->

---------

Co-authored-by: DatCaptainHorse <DatCaptainHorse@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 14:45:13 +03:00

248 lines
13 KiB
Markdown

# build/ — the guest rootfs
Builds a bootable Arch image for the box's virtio-blk root: Mesa (virtio-gpu
native context) plus the five open guest components —
[`nesinit`](../apps/nesinit), [`nescope`](../apps/nescope),
[`neshub`](../apps/neshub), [`neswire`](../apps/neswire),
[`nescapture`](../apps/nescapture) — laid out
the way [borealis](https://chromium.googlesource.com/chromiumos/overlays/board-overlays/+/main/project-borealis)
lays out its `build/`: one big multi-stage `Containerfile`, `--target` picks the
flavor, `etc/` holds the files that get overlaid onto the image verbatim.
```
build/
├── Containerfile everything, in stages: mesa-build, nestri-build,
│ os-base, runtime, runtime_prod, runtime_debug
├── etc/ overlaid onto the image's /etc as-is
├── scripts/
│ └── mkimage.sh docker export → raw ext4, for nesbox's virtio-blk
├── Makefile
└── output/ `make image` writes here (gitignored)
```
```sh
make build # docker build --target runtime_prod → ghcr.io/nestrilabs/nestri/base:latest
make build-debug # docker build --target runtime_debug → ghcr.io/nestrilabs/nestri/base:debug
make image # + pack into output/rootfs.ext4
make image-debug # + pack into output/rootfs-debug.ext4
make proton-image # build Proton from source — hours
make proton-push # build it and publish it
```
## Design notes
Three things worth knowing about how this is put together:
1. **No privileged host chroot.** A bare `chroot` into a hand-extracted
rootfs needs `/proc`, `/sys`, `/dev` bind-mounted in first — they don't
exist inside a chroot target until something puts them there. `os-base`
here is `FROM archlinux:base` directly, with `pacman -S` as plain `RUN`
steps — a Docker build step already runs inside a real container with its
own `/proc`, `/sys`, `/dev`, so that whole bind-mount mechanism has nothing
to do.
2. **No host-side ownership bug to guard against.** `COPY --from=` runs as
root inside the build with no host user in the loop, so there's no
invoking-user uid getting stamped onto `/`, `/usr/bin`, or anywhere else
a build step touches — a failure mode some overlay approaches need an
explicit sanity check for doesn't exist here to check for.
3. **One `cargo build --release --workspace`, not one stage per binary.**
`nescope`, `neshub`, `neswire` and `nescapture` share one Cargo workspace
and one `Cargo.lock` — a BuildKit cache mount on `target/` gives cargo's
own incremental compiler per-crate isolation without needing a separate
Docker stage (and a separate full rebuild of `nesprotocol`) per binary.
**What is deliberately not here: Valve's `steamclient.so`.**
`nestri/CLAUDE.md` is explicit — *"Nothing closed may enter this repo. Not
source, not a dependency, not a directory that 'looked convenient'."* That one
is closed, so whatever layers it on top is a build outside this repo — not
something this repo names, links to, or depends on.
**Proton is here, and this paragraph used to say it was not.** The old wording
put Proton and `steamclient.so` together and called both closed, which is
wrong about Proton: it is compiled from source, which is not a thing you can
do with closed software. Keeping it out cost a box the only way it has to run
a Windows title, for a rule that did not apply to it.
What it is: **proton-cachyos built with `--enable-wow64`**, pulled by tag as a
published image rather than rebuilt here, because it takes hours and moves
only when its own tag does. `PROTON_IMAGE` overrides the tag, and it has to be
declared before the first `FROM` — an `ARG` a `FROM` expands is global or it
is nothing, and getting that wrong fails with `no FROM statement found`, which
says nothing about the actual mistake. wow64 is the whole reason it is a build of ours
and not the distribution's package — it runs 32-bit Windows code inside a
64-bit unix process, so a box needs no lib32 glibc, no second Mesa for i686,
and no second capture layer for 32-bit titles to be captured. The
distribution's package is built without the flag, which is exactly why it
depends on `lib32-*`.
It costs about 1.4 GB of image, and it is the one thing in here that is
payload-shaped: a compatibility layer for Windows games in an image that is
otherwise indifferent to what it runs. The guest components stay indifferent
regardless — none of them branches on it, and the init does not know it
exists. What names it is the command a caller sends.
`runtime_prod` from this Containerfile — tagged
`ghcr.io/nestrilabs/nestri/base:latest` — is a complete, bootable guest image,
and also the shared foundation other builds start from: nesbox's jail image
(see `nesbox/build/`) extracts **Mesa** from it so the guest and host sides of
the virtio-gpu native-context protocol never drift apart. Only Mesa —
`virglrenderer` is the host half of that protocol and nesbox builds its own,
patched, from `nesbox/patches/`; nothing in this image carries it.
## Two packages that look droppable and are not
`llvm-libs` is 164 MB, the largest single thing in the image after Proton, and
`lm_sensors` is only there because something links `libsensors`. Both look like
leftovers of a Mesa configuration that has since been trimmed, and both have
been checked rather than reasoned about: **`libgbm` links them**, and the
compositor needs GBM. Trimming the Mesa build does not reach them.
`lm_sensors` in particular was found the hard way. It used to arrive as a
dependency of the distribution's Mesa package, and dropping that package took
it away — leaving our own Mesa unable to resolve `libsensors.so.5`. Nothing in
a package list says that; the check below is what said it.
## Proton has its own cadence, and its own Containerfile
`make build` **pulls** Proton by tag; it does not build it. Building it takes
hours and it changes only when its tag moves, so it is one image published
once and copied into every guest image after that. `Containerfile.proton` is
that build, and it lives here so the published tag stays reproducible from
this tree rather than from somebody's laptop.
```sh
make proton-image # the current tag
make PROTON_TAG=cachyos-11.1-20261115-native proton-image
```
**`PROTON_TAG` is the only thing to change.** The published version is derived
from it in the `Makefile` rather than written a second time, because the two
are the same number in two spellings — and an image whose name does not say
which Proton is inside it is worse than no image. The `Containerfile`'s own
`PROTON_IMAGE` default is a fallback for a bare container build; going through
`make` is what keeps them in step.
Its **context is `build/`**, not the repository root the guest build uses. All
it needs is the two scripts beside it, and `Containerfile.proton.containerignore`
keeps `output/` out of that context — a build context is copied before the
first instruction runs, so without it every Proton build would begin by moving
the last rootfs image it produced.
Two things in the recipe are worth knowing before changing it:
- **Fetch and build are separate layers on purpose.** The submodule checkout
runs well past ten minutes, and a build that fails on a flag or a missing
tool must not pay for that again. Keep anything that can fail *fast* in
`proton-build.sh`.
- **`widl` is built by hand from the mingw-w64 release.** Without it autoconf
quietly sets `HAVE_WIDL` to false, vkd3d's public headers are never
generated, and the build dies an hour later on a missing header. Arch ships
`widl` only inside `wine`, which wants multilib — which is the thing
`--enable-wow64` exists to avoid.
## There is no init system in here, and that is the design
`nesinit` is PID 1. The image carries **no service manager, no init scripts,
no `udev` and no systemd** — `systemd-libs` stays, because `dbus-daemon` and
`wireplumber` link `libsystemd.so.0`, but nothing in the image can be PID 1
except `nesinit`, and the build fails if anything that could be turns up.
That is why this is plain Arch. The image used to be Artix, chosen for OpenRC,
and every cost of that choice — no `eudev`, no `agetty-openrc`, `udev` being
systemd's anyway, a runlevel edit not stopping a service another one still
needs — was paid for an init system that is no longer here.
**What replaced fourteen `rc-update` lines and nine init scripts:**
| was | now |
|---|---|
| `devfs`, `dmesg`, `udev`, `udev-trigger` | `devtmpfs` makes the nodes; init sets the two modes that matter. The compositor takes input through Wayland and opens nothing `udev` provides |
| `guest-net`, `hostname`, `xdg-runtime`, `cgroups` | init, before it dials out |
| `dbus`, `dbus-session`, `pipewire`, `wireplumber`, `neshub`, `neswire` | a table compiled into `nesinit` |
| `nescope` in the `default` runlevel | **not a service.** It wraps the workload and is started by a launch, with that launch's geometry, and dies with it |
| `agetty` on `hvc0` | nothing. See below |
| `/etc/fstab` | init's own mounts, and shares named in the boot descriptor |
**A box is launched into, not booted into something.** Init mounts what the
descriptor names, brings the table up, says it is ready, and then takes
commands — so an image on its own runs nothing at all, which is the point:
this image is payload-independent and there is no payload in it.
### Getting into a guest that will not boot
`init=/bin/bash` on the kernel command line. Nothing in the image offers a
login prompt — there is no getty in either flavour — and that is cheaper than
carrying one: `nesinit` is an ordinary program, so from that shell you can run
it by hand and watch it fail. `make build-debug` adds `vulkaninfo` and friends
and gives root a password for `su`; it does not add a console.
**A shell is not a booted box, and the difference bites immediately.** Nothing
the init does has happened: the root is read-only, `/run` and `/tmp` are still
directories on it rather than tmpfs, `/run/user/1000` is unwritable, and the
hostname is `(none)` rather than `nesbox` — which is the quickest way to tell
the two states apart. A compositor started in that shell fails on its own
socket, and the error names the runtime directory rather than the cause.
So run `nesinit` first. It mounts, prepares the directories, brings the
services up, then fails to reach a control channel that is not there and
exits — **leaving everything it prepared behind**, which is exactly what makes
the hand-run useful. Then start what you came to debug.
It is safe to run outside a box, and that took fixing: the shutdown path signals
every process it may signal and then powers the machine off, which is right for
PID 1 of a box and catastrophic anywhere else. Both steps are refused when it is
not PID 1, and it says so rather than doing it quietly.
If you would rather not run it at all, the two mounts it does that a compositor
needs are:
```sh
mount -t tmpfs -o mode=1777,size=64m tmpfs /tmp
mount -t tmpfs -o mode=755,size=32m tmpfs /run
mkdir -p /run/user/1000 && chown 1000:1000 /run/user/1000 && chmod 0700 /run/user/1000
```
Making `/run/user/1000` writable in the *image* does not help, and is worth
saying because it is the obvious first thing to try: before the init runs, the
directory is on a read-only root, so its ownership is not what stops a write;
after the init runs, a fresh tmpfs is mounted over `/run` and the image's copy
of the directory is hidden underneath it.
The one thing this does not reach is a failure *before* the shell. If that
happens the evidence is on `console=hvc0` and nowhere else.
### Two build-time checks worth knowing about
Both exist because the failure they catch is invisible at runtime rather than
loud, which is the same reason the old build checked its `conf.d` files:
- **No hook may point at a program that is not in the image.** Removing
systemd removes the script `dbus-reload.hook` calls, and a leftover hook
produces `error: command failed to execute correctly` on every future pacman
transaction — indistinguishable, in a log, from something that matters.
- **Nothing `nesinit` will look for may be missing.** Its service table is
compiled in, so an absent `dbus-daemon` is not a build error by itself; it
is a service that does not come up in a box somebody is waiting on.
## Network defaults
`nesinit` reads `nestri.ip=`/`nestri.gw=` off the kernel command line,
falling back to `172.30.0.2/24` via `172.30.0.1` if neither is set — the
host's own default tap addressing. Keep these in step if that changes on the
host side. A box started with no network device at all is a valid box and
boots without one.
The address is a per-boot parameter rather than an image setting because the
alternative makes every box built from this image the same host on the
network, and two of them collide the moment they run together. Same reasoning
for `/etc/machine-id`, which is a symlink into a tmpfs that init fills at
boot — the previous image baked one in, so every box built from it was the
same machine to anything that asked.
This is also the one thing in the image that keeps `iproute2` installed:
init runs `ip` rather than talking netlink, which is a hundred lines of
`unsafe` saved for an interface configured once.