# Root Cause Analysis: gVisor directfs `openHandle()` Host-FD Identity TOCTOU

## Summary

gVisor's directfs implementation in release-20260928.0 reopens a cached dentry by name in `directfsInode.openHandle()` without checking that the newly opened host file descriptor still identifies the file and type that the sentry previously revalidated. A host-side attacker who can modify the backing directory can preserve a cached regular-file dentry, then atomically race that name to a device node between revalidation and the later `openat()`. The vulnerable sentry uses the resulting device descriptor as a regular-file backing handle. This run proved the issue through the real `runsc`/directfs/systrap product path by racing a regular file with an otherwise sandbox-inaccessible host loop block device: the guest read a host-only secret and wrote unique guest-controlled markers into host-only backing images in two attempts. release-20261005.0 rejected the same swaps and wrote no marker.

## Impact

- **Affected component:** `google/gvisor`, `pkg/sentry/fsimpl/gofer/directfs_inode.go`, function `(*directfsInode).openHandle()`.
- **Affected version tested:** `release-20260928.0`, commit `6485c3fbdfe28172ced5d6aed49554cf494967db`.
- **Fixed version tested:** `release-20261005.0`, commit `97e8e896904cb364319cd9758241d3f0bd27a6f4`.
- **Fix commit:** `cd968a36e5557215b5dcb8989011ce0e558b5dd0`.
- **Required attacker/environment conditions:** the attacker controls or can replace entries in a directfs-backed host directory after a guest has cached a regular-file dentry; the runsc host context can open the swapped device; the race wins between revalidation and open-by-name. This proof requires privileged host setup for an unattached loop device and uses the systrap platform.
- **Risk:** high. A sandboxed guest obtained read/write access to a host block device that was neither present in the OCI rootfs nor mounted into the sandbox. On a real host, equivalent access to a security-relevant host block device can disclose or modify host filesystems and privileged state.

## Impact Parity

- **Disclosed/claimed maximum impact:** guest-to-host sandbox escape, described as host-root code execution through an unsafe host device FD.
- **Reproduced impact:** repeatable guest-controlled read and write to an otherwise inaccessible host block device through production `runsc`. In two vulnerable attempts, the guest read `HOST_ONLY_SECRET_4ec7309a` from offset 4096 and wrote a unique 4096-byte marker at offset 8192. The host verified both markers directly in backing images outside the guest. Two fixed attempts created no markers and rejected the raced device opens.
- **Parity:** `partial`.
- **Not demonstrated:** no host process command was executed and no host-root shell or host executable modification was attempted. The proof establishes a concrete sandbox-boundary-crossing host storage read/write primitive, not a complete command-execution chain. The earlier CUSE path was not used as the final proof because gofer's regular-file I/O path forwards positional `preadv2`/`pwritev2`, which does not provide the normal CUSE handshake used by deterministic character-device passthrough.

## Root Cause

The directfs dentry lifecycle separates identity validation from the host descriptor later used for data operations:

1. The guest resolves `/race/target` while it is a stable regular file, so gVisor creates and caches a regular-file dentry/inode.
2. Directfs revalidation examines a child reached from a parent control descriptor and validates metadata associated with that lookup.
3. `directfsInode.openHandle()` subsequently obtains a usable handle by resolving the dentry name again:

   ```go
   flags |= hostOpenFlags
   openFD, err := unix.Openat(parent.inode.impl.(*directfsInode).controlFD, d.name, int(flags), 0)
   if err != nil {
       return noHandle, err
   }
   return handle{fd: int32(openFD)}, nil
   ```

4. In release-20260928.0, the returned `openFD` is accepted without `fstat()` or comparison against the cached inode type/device identity. A host `renameat2(RENAME_EXCHANGE)` can therefore let revalidation observe the original regular inode while the later `openat()` observes a device node at the same name.
5. Because the sentry still treats the file description as a regular file, reads and writes are forwarded through the gofer host handle. With the raced loop block descriptor, guest `pread()` and `pwrite()` directly affected host block storage.

The fix commit [cd968a36e5557215b5dcb8989011ce0e558b5dd0](https://github.com/google/gvisor/commit/cd968a36e5557215b5dcb8989011ce0e558b5dd0) performs `Fstat(openFD)`, checks that the file type is supported, compares the reopened type with the cached inode type, and for character devices compares major/minor identity. A mismatch is closed and rejected as stale. For this block-device replacement, the fixed build rejects the block-device replacement with `EPERM`; the separate same-input `/dev/cuse` identity control returns `ESTALE`, exactly exercising the fix’s cached S_IFREG versus reopened S_IFCHR comparison. In both cases, the guest never receives a usable host device handle.

## Reproduction Steps

1. Run `bundle/repro/reproduction_steps.sh` from any directory. It accepts `PRUVA_ROOT`; the default resolves the bundle directory portably.
2. The script reads `bundle/project_cache_context.json`, reuses the prepared gVisor repository and build cache when available, and otherwise creates `bundle/artifacts/gvisor-cache`.
3. It anchors source/build identity to vulnerable commit `6485c3f...` and fixed commit `97e8e896...`, verifies that fix commit `cd968a36...` is absent/present respectively, and checks the actual `Fstat(openFD)` patch hunk.
4. It builds or reuses genuine optimized `runsc` binaries, compiles the static guest and host racers, and starts a privileged Docker host fixture.
5. For each of two vulnerable and two fixed attempts, the fixture creates a fresh 16 MiB host-only backing image and unattached loop device, writes a secret at offset 4096, starts a real `runsc --directfs=true --platform=systrap` sandbox, primes the target as a regular dentry, then begins atomic regular-file/block-device swaps.
6. The guest races `open(O_RDWR|O_TRUNC)`, reads the host secret, and attempts to write a unique marker at offset 8192. After `runsc` exits, the host reads that offset directly from the backing image.
7. Exit code 0 requires both vulnerable attempts to read/write and create their exact markers, both fixed attempts to create no marker, and both fixed attempts to record device-open rejection.

Expected summary:

```text
vulnerable_host_read_write_marker_attempts=2/2
fixed_host_write_marker_attempts=0/2
fixed_rejection_attempts=2/2
vulnerable_cuse_fd_control=1/1
fixed_cuse_estale_control=1/1
```

## Evidence

Primary current-run evidence is under `bundle/repro/results/` and is digest-bound by `bundle/repro/runtime_manifest.json`.

- `results/source-identity.log`: exact vulnerable/fixed commits and patch presence.
- `results/runsc-version-{vuln,fixed}.txt`: product versions exercised.
- `results/guest-vuln-1.log`:

  ```text
  GUEST_HOST_BLOCK_READ: secret=HOST_ONLY_SECRET_4ec7309a opens=2 offset=4096 bytes=4096
  GUEST_HOST_BLOCK_WRITE: marker=GUEST_HOST_MARKER_vuln_1_91d7c3ee offset=8192 bytes=4096 errno=0
  ```

- `results/guest-vuln-2.log`: an independent vulnerable process repeats the read/write after 11 opens.
- `results/host-observation-vuln-1.log` and `host-observation-vuln-2.log`: direct host inspection reports `MATCH=true` and includes marker bytes in hexadecimal.
- `results/host-marker-vuln-{1,2}.bin`: exact marker bytes read from host-only backing images.
- `results/guest-fixed-1.log` and `guest-fixed-2.log`: hundreds of thousands of block-device attempts produce no win; device-node windows return `EPERM` while regular-file windows remain usable.
- `results/cuse-guest-vuln.log`: the vulnerable build accepts the raced `/dev/cuse` descriptor (`GUEST_CUSE_FD_OPEN`).
- `results/cuse-guest-fixed.log`: the fixed build has no CUSE win and reports large nonzero `estale` counts for the identical synchronized swap, directly demonstrating the fix’s required `ESTALE` identity failure.
- `results/host-observation-fixed-{1,2}.log` and `host-marker-fixed-{1,2}.bin`: host-side negative controls are all zero and report `MATCH=false`.
- `results/summary.txt`: aggregate 2/2 vulnerable effects versus 0/2 fixed markers and 2/2 fixed rejection.
- `bundle/logs/reproduction_steps.log`: complete final product-run diagnostics.
- `bundle/logs/final_run1_console.log` and `final_run2_console.log`: two consecutive successful top-level executions.
- `bundle/repro/runtime_manifest.json`: source identity, runtime stack, artifact list, and SHA-256 map.

The final run used Linux `6.8.0-142-generic`, x86-64, Docker `29.1.3`, gVisor systrap, directfs enabled, and no sanitizer.

The fdwatch files are retained as diagnostics only. They intentionally scope observations to runsc-reported process trees and exact loop major/minor, but sampled no rows because the vulnerable guest won and completed I/O faster than the `/proc` watcher acquired the process tree. Unlike the previous broad watcher, no unrelated host character-device rows are treated as proof. Exact guest output plus direct host backing-image marker comparison is the primary descriptor/effect correlation.

## Recommendations / Next Steps

- Upgrade to `release-20261005.0` or later.
- Preserve the fix's post-`openat()` identity checks. At minimum, compare file type against the cached inode; for device files, compare major/minor identity; close the descriptor and return `ESTALE`/an error on mismatch.
- Add a regression test that synchronizes a cached regular dentry with an atomic replacement before `openHandle()`, covering character and block device substitutions.
- Test all open modes that force a fresh handle, especially `O_TRUNC`, and test repeated revalidation/open races.
- Treat any host device descriptor exposed through a cached regular dentry as a sandbox-boundary violation even if one specific device protocol is not usable through positional I/O.
- For terminal exploit research, a host filesystem block device can be used to study controlled modification of a non-production fixture filesystem, but that escalation was deliberately not claimed here without current-run host command-execution evidence.

## Additional Notes

- The final script passed twice consecutively in the current runtime. Each invocation itself performs two vulnerable and two fixed process attempts.
- The script is self-contained apart from declared source helpers and standard network/build prerequisites. It recompiles helpers and rebuilds `runsc` if compatible cached binaries are unavailable.
- It bounds every `runsc` invocation with `timeout`, isolates each attempt with a fresh bundle/runsc root/tmpfs/backing image/loop mapping, and restores caller ownership of bind-mounted result paths.
- The fixture is intentionally an unattached loop device rather than a live system disk. This safely proves host block read/write without risking corruption of the worker host.
- `runtime_manifest.json` is rewritten on success and on premature failure; proof artifacts are finalized before they are hashed.
