mirror of
https://github.com/nestriness/nestri.git
synced 2026-09-19 17:25:19 +03:00
fix(nesinit): do not mount over the share tree, and check who serves an address
Four findings from review, all of them real. The relay's directory was mounted on the tree a session's shares live in. A fresh tmpfs there hides every directory the image prepared underneath it: the install, the user state, the work directory, and the mount point the log share is attached to from fstab. A box would have come up with a socket and without any of the places its workload looks for its files, and the exact-path check could not notice, because what fstab mounts is a directory inside that tree rather than the tree itself. It moves to /run, which is where a runtime socket belongs, is a tmpfs already, and has nothing else mounted inside it. It was also owned by this process and closed to everyone else, which stopped the workload traversing it to reach the relay at all. The directory is now readable and searchable, and still writable by nothing but this process, which is what makes the socket in it unreplaceable; the socket itself is what the workload is allowed to connect to. The permission belongs on the socket rather than on the path. The address served to a reader was built once at startup and served forever, so a reader that polls for a better one could only ever get the first. An endpoint does not know all of its own addresses when it binds: the first is the one that works on the same network and fails from anywhere else. It is now rebuilt per read, which is what makes polling for it worth doing. And the address was taken from whoever held a path in a directory the workload can write. Workload code could unlink the socket a service was listening on, bind its own, and every read afterwards would hand the client an address of its choosing -- a session given to somebody else rather than a session that fails. The peer's credentials are now checked before a byte is read, from the kernel rather than from anything the peer says about itself, and an address served by the workload's own user is refused and said loudly. That check is only worth something while the workload has a user of its own, so the image grows one. Two users, and they must stay two: one runs the services that ship in the image, the other is who a workload runs as. Sharing one does not weaken the check, it makes every session fail it. A workload running as root is every user at once and cannot be told apart from anything; the check stands down there and says so at boot instead, because refusing root would refuse whatever legitimately serves the address as well. Also bumps tinyvec by a patch release. It does not build on this toolchain -- `vec` resolves to the module and not the macro -- which made every crate that depends on an endpoint, including this one, unbuildable. Pre-existing and nothing to do with this change; the lockfile said the same version before it.
This commit is contained in:
@@ -9,6 +9,7 @@
|
||||
// anything above it.
|
||||
|
||||
use std::io;
|
||||
use std::os::unix::fs::PermissionsExt;
|
||||
use std::path::Path;
|
||||
|
||||
use nesprotocol::lifecycle::{Payload, from_line, to_line};
|
||||
@@ -22,10 +23,18 @@ use tokio::sync::mpsc::{Receiver, Sender};
|
||||
/// is read-only, so this is a tmpfs that `filesystems` puts there, and a
|
||||
/// rename here that did not reach the mount table would take the relay down
|
||||
/// with an `EROFS` that looks like nothing to do with a path.
|
||||
pub const DIRECTORY: &str = "/nestri";
|
||||
///
|
||||
/// **Under `/run` rather than in the tree the session's shares live in**, and
|
||||
/// that was a correction. Putting a fresh tmpfs on the share tree hid every
|
||||
/// directory the image had prepared underneath it — the install, the user
|
||||
/// state, the work directory and the log share's own mount point — so a box
|
||||
/// gained a socket and lost the places its workload was supposed to find its
|
||||
/// files. `/run` is where a runtime socket belongs, it is a tmpfs already, and
|
||||
/// nothing else is mounted inside it.
|
||||
pub const DIRECTORY: &str = "/run/nestri";
|
||||
|
||||
/// Where the workload finds the relay.
|
||||
pub const SOCKET: &str = "/nestri/payload.sock";
|
||||
pub const SOCKET: &str = "/run/nestri/payload.sock";
|
||||
|
||||
/// The longest envelope this will assemble before giving up on the connection.
|
||||
///
|
||||
@@ -63,6 +72,17 @@ pub async fn serve(
|
||||
let _ = std::fs::remove_file(path);
|
||||
let listener = UnixListener::bind(path)?;
|
||||
|
||||
// The workload does not run as this process does, and it has to be able to
|
||||
// connect. It reaches the socket through a directory this process owns and
|
||||
// nothing else may write to, so the permission that matters is on the
|
||||
// socket rather than on the path: the directory is what stops anybody
|
||||
// replacing this listener, and this is what lets the workload talk to it.
|
||||
//
|
||||
// Without it the workload gets `EACCES` on connect and the payload layer
|
||||
// is dead in both directions, silently, because nothing in the guest is
|
||||
// waiting to be told about a relay it cannot reach.
|
||||
std::fs::set_permissions(path, std::fs::Permissions::from_mode(0o666))?;
|
||||
|
||||
loop {
|
||||
let stream = tokio::select! {
|
||||
accepted = listener.accept() => accepted?.0,
|
||||
|
||||
Reference in New Issue
Block a user