16 Commits
Author SHA1 Message Date
Daniel Samson 4037e746aa docs: storage remount-on-replug is bench-pending, not bench-verified
The physical unplug/replug remount was never actually bench-run — the claim was
carried over from S3. QEMU cannot re-present a usb-storage device_add, so this
path is only reachable on real hardware. Correct both docs to say bench-pending,
keeping the honest distinction: the re-adopt+remount CODE PATH is QEMU-proven by
the driver-crash rebuild; only the physical-replug end-to-end awaits a bench pass.
2026-08-10 19:35:04 +01:00
Daniel Samson 4dfb5012c0 docs: the removal lifecycle closes — three triggers, one path (S5)
Storage removal is now robust to all three ways a volume can leave, and the docs
say so. storage-architecture.md and storage-design-rationale.md move medium_changed
from "planned to be consumed" to consumed, and record the third trigger:

  - the DEVICE leaving the tree (a pulled stick)          — presence polling
  - the MEDIUM leaving while its device stays (a reader)  — the volume manager
    now consumes the pushed medium_changed event
  - the storage DRIVER crashing while its device stays    — a channel-liveness
    geometry() probe reaps the volume and rebuilds it on the restarted driver's
    fresh channel; presence polling alone cannot see this (the V4 open edge)

The re-adopt-and-remount path is QEMU-proven by the driver-crash rebuild; a
physical unplug/replug is bench-verified (QEMU cannot re-present a usb-storage
device_add). The transport-native eject signal (SCSI UNIT ATTENTION, AHCI
PxSSTS, NVMe namespace-change AER) in place of the TEST UNIT READY poll stays
the documented future refinement.
2026-08-10 06:08:56 +01:00
Daniel Samson 061eb7c004 volume-manager: drop the medium-event dedup — it only ever misfired (S5 review)
The adversarial S5 review found a real, unrecoverable defect: onMediumEvent
deduped medium_changed events on a module-global last_medium_change compared by
equality against the event's change_count. But change_count is a PER-DRIVER
counter that restarts at 0 in every usb-storage instance — it is bumped +%=1 and
published only on a real medium transition, so within one instance every count
is unique and monotonic and an equality dedup can never legitimately fire.

The global was carried across a driver restart — S5's OWN crash-rebuild path —
so a fresh instance's first eject (count=1) collided with a stale last==1 and was
dropped. removeDevice never ran; the filesystem kept serving I/O against absent
media forever, and nothing else recovered it: onGeometry answers from a cached
block_count so channelAlive stays true, and isDevicePresent stays true (the
device never left the tree). medium_changed is the sole eject oracle there.

The dedup guarded a re-delivery the driver already makes impossible, and its
only observable effect was this bug. Remove it: react to each present-edge
directly. Both branches are idempotent (a freed device stops matching dev.used)
and the poll reconciles, so acting on every genuine edge is safe — and a fresh
driver instance's counter can no longer be mistaken for the previous one's.

Verified: build + bounds green; volume-removal, volume-medium-change and
volume-driver-restart all pass, so both removal paths survive the change. The
driver-restart-then-eject intersection that triggered the collision cannot be
staged in QEMU — the internal driver-kill cannot be ordered against a QMP eject,
and a usb-storage device_add is not re-presented — so the guarantee rests on the
driver's one-publish-per-transition-with-unique-count contract.
2026-08-10 05:56:06 +01:00
Daniel Samson 5dc966838a volume-manager: rebuild a volume when its storage driver dies (S5)
The V4 review's open edge: a storage driver that crashes while its device
stays in the tree left fat wedged on a dead channel — device-presence
polling (a device-manager enumerate) still reported the device present,
so nothing reaped it. pollTick now also probes channelAlive(dev), a
geometry() on the block channel that fails fast on the dead endpoint; a
present device with a dead channel is reaped like a pull, and the adopt
loop re-adopts it on the restarted driver's fresh channel — the rebuild.
The manager's own liveness probe makes fat self-detection unnecessary:
it rebuilds regardless of the wedged filesystem's state.

The drill: the device manager gains a test-storage-restart mode that
kills usb-storage once, ~2s after its hello (post-mount); a new
volume-driver-restart kernel case boots a manual tree with it, and the
QEMU case asserts a SECOND mount of the same id-path after the reap —
the rebuild. A pre-S5 manager, checking only device presence, never
reaps, so the second mount never appears. The manually-spawned volume
manager needed kernel-supervisor protocol grants (bind its name, open
the device manager), as the other manual-tree services already have.

Full suite 133/133.
2026-08-10 05:21:48 +01:00
Daniel Samson 700452dc4e volume-manager: consume medium_changed — the second removal trigger (S5)
The block client gains subscribeMedium / unsubscribeMedium /
decodeMediumChanged (the reserved subscribe/unsubscribe verbs plus the
MediumChanged decode), so a consumer never hand-rolls the wire format.
The volume manager subscribes to each device it adopts and consumes the
event through the new on_buffered_message seam — never the protocol
dispatch, whose op numbers collide with the manager's own hello.

A medium leaving while its device stays in the tree (a card reader, an
eject) now runs the SAME kill-retire-remount path as a pulled stick:
absent retires the volume, present re-probes it. That closes the "two
triggers, one lifecycle" the architecture specifies — device-presence
polling alone could never see a medium leave under a present device.
dropDevice unsubscribes before closing so the driver's bounded
subscriber table frees the slot; on a dead channel (a real pull) the
call fails fast, proven by volume-removal still passing.

New volume-medium-change case: eject the medium (not the device) ->
"medium absent" -> the manager unmounts. Fails against a pre-S5 manager
that never subscribed. Suite 132/132 (device-authority is the known
child-cleanup flake, green on rerun).
2026-08-10 04:58:54 +01:00
Daniel Samson 0faa0fd21b service: on_buffered_message for pushed events (S5)
An additive, behavior-neutral callback. A buffered async message
(Received.isMessage — a pushed event from a provider this service
subscribed to) carries a payload in the receive buffer; run() now hands
it to on_buffered_message before falling through to on_notification with
the badge, so a coalesced timer/exit riding the same wake is not lost. A
service that does not set the callback (all of them today) is unchanged:
the isMessage branch is a no-op and on_notification still runs, exactly
as before.

This is the seam the volume manager needs to consume block
medium_changed: the event's reserved op number collides with the volume
manager's own hello, so it must be decoded by hand here, never through
the protocol dispatch. Full suite neutral (the lone device-authority
miss is a known child-cleanup race that passes on rerun).
2026-08-10 04:43:27 +01:00
Daniel Samson c47215821c docs: the second engine is built — exFAT + proven harness reuse (S4)
Flip the storage docs to record exFAT as a built second engine reusing
the shared filesystem harness wholesale (the reuse the architecture
promised): full read + write, directories, rename, on-disk up-case
folding, routed by VBR to fat or exfat at an exfat-<serial> id-path.
Note the two surface limits shared by BOTH engines as vfs-layer concerns,
not exFAT shortcuts: u32 file offsets (a 4 GiB cap) and ASCII-only names.
2026-08-10 04:15:03 +01:00
Daniel Samson 6ddb08091d exfat: adversarial-review fixes — overflow safety, sparse gaps, dir size, big-image bitmap (S4 step 9)
A 7-dimension adversarial review of the engine, tool, and routing found
nine real defects (host tests + the in-VM drill missed them). Fixed:

- geometryOf now rejects a crafted VBR whose cluster shift exceeds the
  exFAT ceiling (bytes+sectors shift > 25) or whose cluster_count exceeds
  the spec max (0xFFFFFFF5) — either would overflow the engine's u32
  cluster-byte / cluster-bounds arithmetic and panic under ReleaseSafe on
  untrusted removable media. validCluster/allocateCluster widened to u64,
  and writeFile's clusters_needed widened, for a >4 GiB file near the u32
  offset boundary.
- writeFile no longer claims valid_data_length = size unconditionally: a
  sparse write past a foreign file's old valid boundary now zero-fills the
  skipped gap on disk, so a read there returns zero, not stale bytes.
- ensureDirCapacity rewrites a grown subdirectory's own DataLength, so a
  spec-compliant reader that bounds a directory by DataLength sees the new
  entries (danos itself bounds by the end marker, but chkdsk / other OSes
  do not).
- make-exfat-image lays the allocation bitmap across as many clusters as
  it needs; a >128 MiB image (whose bitmap exceeds one cluster) was
  self-inconsistent. Verified: the engine mounts+reads both the 48 MiB
  fixture and a 256 MiB image.

Documented (not fixed here — a shared vfs-layer limit, like the u32
offset cap): non-ASCII names fold to '?', the same as the FAT engine.

New host tests pin each fix (crafted-VBR rejection, sparse-gap zero,
subdir-grows-and-records-size). Full suite 131/131, bounds green.
2026-08-10 04:14:54 +01:00
Daniel Samson 77b64229c2 exfat: the in-VM mount + mutation drill (S4 step 8)
exfat-volume attaches a bare exFAT data device (serial e0fa0001, from
make-exfat-image.py) beside the FAT boot volume. The volume manager
content-routes it to the exFAT service — not fat — which mounts it at
its id-path /volumes/exfat-e0fa0001; the exfat-test client then reads the
seeded HELLO.TXT and mutates through the mount (mkdir/write/rename/read/
remove). The second engine reusing the shared harness is now proven end
to end, on a real device, in QEMU.

The drill caught what host tests could not: the exfat service was denied
openEndpoint("volume-manager") — it had no protocol.csv grant — so its
hello never reached the manager and its volume never mounted (a silent
spin, no fault). Added the grant mirroring fat's. A new exfatVolumeTest
kernel case boots the tree and spawns the fixture; the run harness grows
an exfat data-volume flavor. Fails against pre-S4 (no exfat binary, csv
row, VBR recognizer, or grant).
2026-08-10 03:44:29 +01:00
Daniel Samson 90906bcefe exfat: cross-engine discrimination + the exfat-test fixture (S4 step 7)
Each engine now refuses the other's volume at mount(): the exFAT engine
rejects a FAT boot sector (a non-zero byte where exFAT keeps MustBeZero),
and the FAT engine rejects an exFAT one (MustBeZero reads as a zero
bytes-per-sector), each mounting its own as a control. The two engines
can never claim the same medium.

exfat-test is the mount round-trip client, cloned from fat-test: it waits
for /volumes/exfat-e0fa0001, reads the seeded HELLO.TXT, then exercises
mkdir + write + rename + read-back + remove through the VFS and reports
"exfat-test: ok". It ships as a lazy fixture in test builds (the
exfat-volume QEMU case, step 8, drives it). build + zig build test +
bounds green.
2026-08-10 03:32:55 +01:00
Daniel Samson 2c2745e9e5 volume-manager: recognize exFAT and route it by content (S4 step 6)
partition.zig gains an exFAT VBR recognizer: the "EXFAT   " name (where a
FAT BPB keeps its OEM string, so the two never collide) plus the 0x55AA
signature, with VolumeSerialNumber (offset 100) as a new exfat_serial
identity rung. A `recognize` helper tries exFAT, then FAT, then the MBR
disk-signature fallback, and sets each volume's FilesystemKind — so
allVolumes/firstVolume and the GPT path all tag a volume with the engine
its content needs. volume-map renders exfat-<serial> as the id-path, and
filesystems.csv adds the exfat -> /system/services/exfat row: an exFAT
stick now spawns the exFAT service, at its own content id-path.

Host tests: a bare exFAT volume recognized with its serial; a FAT VBR
still recognized as fat (the discrimination); the exfat-<serial> id
render. build + zig build test + bounds green.
2026-08-10 03:28:08 +01:00
Daniel Samson 2a6d604577 tools: make-exfat-image.py — a real exFAT image builder (S4 step 5)
Pure stdlib, no mkfs.exfat: writes a Main Boot Sector + its boot-region
checksum + a backup region, the 32-bit FAT, an allocation bitmap, an
up-case table (with its checksum), and a root directory carrying the
bitmap/up-case/label entries plus a seeded HELLO.TXT file set. --serial
sets the volume serial (the content id-path); --verify re-reads the VBR,
recomputes the boot checksum, and validates each file set's checksum.

Cross-checked against the engine: the Zig exFAT engine mounts a
Python-produced image and reads HELLO.TXT through a case-insensitive
resolve ("exfat hello danos"), proving the tool and engine agree on the
layout — VBR, geometry, bitmap, up-case, and the entry-set checksum. No
image is committed; the drill (step 8) generates one per run.
2026-08-10 03:22:19 +01:00
Daniel Samson e81a4e6f1d exfat: the service + build wiring (S4 step 4)
exfat.zig is a near-clone of fat.zig — the reuse the architecture
promised, now real: same IpcBlock DMA-bounce wrapper, same
acquire-volume-by-hello to the volume manager, same content-conditional
boot rewrites (resolve /system/configuration to decide the system
volume), all differences confined to the engine it wraps. That the
harness's Server(engine.FileSystem) compiles is the proof the exFAT
engine meets the same pub-fn contract as FAT — a comptime check, not a
hope.

The exfat package (build.zig + build.zig.zon) mirrors fat's: it ships in
production_ship, its on-disk + engine host tests run through the root
test aggregate (replacing the temporary standalone entry), and the root
zon declares it. build + zig build test + bounds all green.
2026-08-10 03:18:00 +01:00
Daniel Samson e240341bfb exfat: the engine write path (S4 step 3)
create / write / truncate / remove / mkdir / rename, completing the
engine's pub-fn contract with the shared harness. The allocation BITMAP
is the authority: setAllocated IS the allocation (a set bit), and the
32-bit FAT only records a fragmented chain's order — forgetting the bit
would hand a live cluster out twice, so the write path never touches the
FAT without also owning the bit.

This engine writes FAT-linked (no_fat_chain=0) files: createFile makes
an empty set; writeFile allocates+links+zeroes clusters (so a sparse gap
reads zero and valid_data_length can honestly equal data_length) and
rewrites the Stream entry; a contiguous file opened for growth is first
threaded through the FAT. truncate frees the chain; removeFile clears
each set entry's InUse bit and frees the chain (a non-empty directory is
refused); rename re-homes the same clusters under a new name set.
createDirectory allocates one zeroed cluster — exFAT directories carry no
"." / ".." entries. Every entry-set mutation recomputes the set
checksum. Offsets clamp to the vfs u32 surface.

Ten engine host tests now (read + write across clusters, truncate+reuse,
remove, subdir+inner file, rename); 17 total with on-disk. bounds green.
2026-08-10 03:11:05 +01:00
Daniel Samson 62eb2a748a exfat: the engine read path (S4 step 2)
mount + resolve + list + read over a BlockDevice, host-tested against a
RAM-backed image the tests build with formatExfat. mount reads the VBR,
loads the geometry, and scans the root for the Allocation Bitmap (0x81)
and Up-case Table (0x82); the up-case prefix is decompressed (0xFFFF
identity runs) into a bounded table so names fold correctly.

A directory is read as consecutive 32-byte entries via readChain, so a
File/Stream/Name SET that straddles a sector or cluster boundary
assembles cleanly; each set's checksum is validated before it counts as
a file. A stream's no_fat_chain flag picks contiguous-arithmetic vs
FAT-follow cluster walking. reads honor valid_data_length (allocated-
but-unwritten tail reads as zero) and clamp data_length to the vfs u32
offset surface.

Tests cover mount + up-case fold, listing (skipping the metadata
entries), case-insensitive resolve, and reads across a cluster boundary
on both a contiguous and a FAT-fragmented file. Wired via engine.zig
(which imports on-disk.zig) into the host-test aggregate. 12/12, bounds
green. Write path is step 3.
2026-08-10 03:02:58 +01:00
Daniel Samson 56bd2e7678 exfat: the on-disk layout (S4 step 1)
The pure, host-testable byte layer of the second engine: the Main Boot
Sector (VBR) and the six 32-byte directory-entry types — Allocation
Bitmap, Up-case Table, Volume Label, File, Stream Extension, File Name —
as align(1) extern structs, plus the three exFAT checksums (boot region,
up-case table, directory-entry set), the name hash, and the packed
timestamp <-> Unix-epoch conversion.

geometryOf accepts only "EXFAT   " + 0xAA55 + an all-zero MustBeZero
region; that last guard is the mutual exclusion with FAT — a FAT prober
reads a zero bytes-per-sector there and rejects the volume, and this one
rejects a FAT boot sector for want of the exFAT name. Wire-format widths
are named consts (spec facts, no bare literals) so the bounds gate stays
green; the layout is pinned by @offsetOf/@sizeOf tests.

Wired into the host-test aggregate directly for now; it moves into the
exfat package's own test step when that lands (step 4). 7/7 host tests,
bounds green.
2026-08-10 02:46:23 +01:00
24 changed files with 2966 additions and 52 deletions
+3
View File
@@ -122,6 +122,7 @@ fn driverArtifact(comptime package: []const u8, comptime artifact: []const u8) S
/// (docs/build-packages-plan.md). /// (docs/build-packages-plan.md).
const production_ship = [_]ShipRow{ const production_ship = [_]ShipRow{
service("fat"), service("fat"),
service("exfat"),
service("display"), service("display"),
service("display-demo"), service("display-demo"),
service("device-manager"), service("device-manager"),
@@ -332,6 +333,7 @@ pub fn build(b: *std.Build) void {
if (test_case != null) for ([_][]const u8{ if (test_case != null) for ([_][]const u8{
"vfs-test", // the user-space VFS round-trip client "vfs-test", // the user-space VFS round-trip client
"fat-test", "fat-test",
"exfat-test", // the exFAT mount round-trip client (S4)
"badge-scope-test", // the guessable-id probe: a second process names the first's node and layer "badge-scope-test", // the guessable-id probe: a second process names the first's node and layer
"shared-memory-server", "shared-memory-server",
"shared-memory-client", "shared-memory-client",
@@ -447,6 +449,7 @@ pub fn build(b: *std.Build) void {
csv_library, csv_library,
xkeyboard_config_library, xkeyboard_config_library,
b.dependency("fat", .{}), b.dependency("fat", .{}),
b.dependency("exfat", .{}),
b.dependency("volume-manager", .{}), b.dependency("volume-manager", .{}),
b.dependency("display", .{}), b.dependency("display", .{}),
b.dependency("ps2-bus", .{}), b.dependency("ps2-bus", .{}),
+2
View File
@@ -46,6 +46,7 @@
.@"pci-bus" = .{ .path = "system/drivers/pci-bus" }, .@"pci-bus" = .{ .path = "system/drivers/pci-bus" },
.init = .{ .path = "system/services/init" }, .init = .{ .path = "system/services/init" },
.fat = .{ .path = "system/services/fat" }, .fat = .{ .path = "system/services/fat" },
.exfat = .{ .path = "system/services/exfat" },
.display = .{ .path = "system/services/display" }, .display = .{ .path = "system/services/display" },
.@"display-demo" = .{ .path = "system/services/display-demo" }, .@"display-demo" = .{ .path = "system/services/display-demo" },
.@"device-manager" = .{ .path = "system/services/device-manager" }, .@"device-manager" = .{ .path = "system/services/device-manager" },
@@ -64,6 +65,7 @@
.@"virtio-gpu" = .{ .path = "system/drivers/virtio-gpu" }, .@"virtio-gpu" = .{ .path = "system/drivers/virtio-gpu" },
.@"vfs-test" = .{ .path = "test/system/services/vfs-test", .lazy = true }, .@"vfs-test" = .{ .path = "test/system/services/vfs-test", .lazy = true },
.@"fat-test" = .{ .path = "test/system/services/fat-test", .lazy = true }, .@"fat-test" = .{ .path = "test/system/services/fat-test", .lazy = true },
.@"exfat-test" = .{ .path = "test/system/services/exfat-test", .lazy = true },
.@"badge-scope-test" = .{ .path = "test/system/services/badge-scope-test", .lazy = true }, .@"badge-scope-test" = .{ .path = "test/system/services/badge-scope-test", .lazy = true },
.@"shared-memory-server" = .{ .path = "test/system/services/shared-memory-server", .lazy = true }, .@"shared-memory-server" = .{ .path = "test/system/services/shared-memory-server", .lazy = true },
.@"shared-memory-client" = .{ .path = "test/system/services/shared-memory-client", .lazy = true }, .@"shared-memory-client" = .{ .path = "test/system/services/shared-memory-client", .lazy = true },
@@ -18,13 +18,21 @@
> its own `/volumes/<id>` path with its own supervision. The boot volume is > its own `/volumes/<id>` path with its own supervision. The boot volume is
> identified by **content** (a volume backs `/system/configuration` + `/system/logs` > identified by **content** (a volume backs `/system/configuration` + `/system/logs`
> only when it resolves `/system/configuration` on its own media), so it works as > only when it resolves `/system/configuration` on its own media), so it works as
> any partition of any device. **Still pending**: the `filesystem UUID` rung and > any partition of any device. exFAT is **built** as a second engine
> a second engine (exFAT, S4); the volume manager *consuming* `medium_changed` > (`system/services/exfat`): full read + write, directories, rename, and on-disk
> (removal is detected by device-presence polling; the event is published but only > up-case folding, reusing `library/kernel/file-system-harness` wholesale — the
> a card-reader medium change needs the subscription); the remount-on-replug > reuse claim, proven — and a volume routes to fat or exfat by its VBR, at an
> end-to-end (the logic is in place; QEMU can't re-present the boot-controller > `exfat-<serial>` id-path. Removal is robust to all three triggers now: a
> device, so it is bench-verified); and arbitration when two volumes both resolve > pulled device (presence polling), a medium that leaves while its device stays
> the boot markers (S3 mounts both and logs each claim; picking one is S4). A few > (the volume manager CONSUMES `medium_changed`), and a storage driver that
> crashes while its device stays present (a channel-liveness `geometry()` probe
> reaps the volume and rebuilds it on the restarted driver's fresh channel). The
> re-adopt-and-remount path is QEMU-proven by the driver-crash rebuild; a physical
> unplug/replug exercises the same path but is bench-pending (QEMU cannot
> re-present a usb-storage `device_add`). **Still pending**: the `filesystem UUID`
> rung (ext-family superblocks, which need such an engine); and arbitration when
> two volumes both resolve the boot markers (S3 mounts both and logs each claim;
> picking one is deferred). A few
> markers below are left where a duty is still pending. > markers below are left where a duty is still pending.
## The model ## The model
@@ -182,25 +190,26 @@ surprise-removal path — kill the filesystem process, retire its mounts,
respawn on return. No half-alive states, no `remount-ro`, no mounts that respawn on return. No half-alive states, no `remount-ro`, no mounts that
error forever (Plan 9's dead-server wart). error forever (Plan 9's dead-server wart).
The path has **two triggers, one lifecycle**: the *device* leaving (the The path folds **three triggers into one lifecycle**: the *device* leaving (a
storage driver dies — channel death, the table below), and the *medium* pulled stick — presence polling); the *medium* leaving while the device stays
leaving while the device stays (an SD card pulled from its reader, an ATAPI (an SD card pulled from its reader, an ATAPI tray opened, a USB card reader);
tray opened — including USB card readers today). The second trigger is the and a storage *driver crashing* while its device stays in the tree. The second
pushed `medium_changed` event on the block protocol — published today from a trigger is the pushed `medium_changed` event on the block protocol, published
TEST UNIT READY poll; still *planned* is the volume manager *consuming* it from a TEST UNIT READY poll — the volume manager now **consumes** it (subscribed
(today removal is driven only by device-presence polling) and translating the per device), running the same kill-retire path and re-probing on medium return,
transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS, NVMe so a swapped card is never served with the previous card's filesystem state. The
namespace-change AER) in place of the poll. On the event the volume manager third is caught by a channel-liveness `geometry()` probe: presence polling alone
runs the same kill-retire path, then re-probes on medium return exactly as on sees the device still present, but the channel is dead, so the manager reaps the
device return. Without it, a swapped card would be served with the previous volume and rebuilds it on the restarted driver's fresh channel. Still *planned*
card's filesystem state. is translating the transport's native signal (SCSI UNIT ATTENTION, AHCI PxSSTS,
NVMe namespace-change AER) in place of the presence poll.
| Layer | Observes | Must do | Guarantees | | Layer | Observes | Must do | Guarantees |
|---|---|---|---| |---|---|---|---|
| Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick | | Bus driver | port/hub status change | tear down the device's slots (children first, recursively — built, hot-plug matrix), report `child_removed` per interface | the device tree is honest within one reconcile tick |
| Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds | | Device manager | `child_removed` / reporter death | prune the child; **reap the bound driver** (built) — the storage driver for that stick dies now, not never | no zombie storage processes; re-report rebinds |
| Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang | | Storage driver | its own death (it IS the removed device's driver) | nothing — dying is its removal handling; DMA/IOMMU/claims release mechanically at death | in-flight transfers fail visibly to callers, never hang |
| Volume manager *(removal built; remount bench-pending)* | the storage device leaving the device-manager tree (poll) | kill the filesystem service of that device's volume; its kernel mounts retire | one removal path; mounts never dangle; log persistence stops *cleanly* | | Volume manager *(built)* | a device leaving the tree (poll), a `medium_changed` event, or a dead channel under a still-present device (a crashed driver — `geometry()` liveness probe) | kill that volume's filesystem service (its mounts retire), then re-adopt + remount on return or on the restarted driver's fresh channel | one removal path for all three triggers; mounts never dangle; the manager never serves from behind a dead channel |
| Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel | | Filesystem service | its block channel dies (`EPEER`) mid-operation, or it is killed by the volume manager | if it observes the death first: flush nothing (the medium is gone), answer in-flight requests with errors, exit; dirty write-back data is **lost and said to be lost** | the unflushed write-back window is dropped on a surprise yank — danos writes no on-disk dirty/clean-shutdown marker today; the process never serves from behind a dead channel |
| Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang | | Kernel | backend endpoint death | lazy mount-slot sweep on next resolve (built); ownership-gated `fs_unmount` (built, V0) | resolution under a dead mount is `not_found`, not a hang |
| Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media | | Application | `not_found` / error on paths under the vanished mount | its own error handling — the contract is honest absence, identical to the path never existing | no operation blocks forever on removed media |
@@ -104,12 +104,19 @@ and /system/logs), closing the two-sticks question honestly.
**Filesystems (per volume, one process).** The proven unit everywhere from **Filesystems (per volume, one process).** The proven unit everywhere from
Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol Plan 9's `dossrv` to Minix to Fuchsia: block-client + engine + file-protocol
provider in one binary, one process per volume (9front practice; per-volume provider in one binary, one process per volume (9front practice; per-volume
fault isolation is what our supervision makes cheap). fat's shell becomes a fault isolation is what our supervision makes cheap). fat's shell became a
shared *filesystem harness* library before a second engine is written; a shared *filesystem harness* library (`library/kernel/file-system-harness`), and
partition walk is added in the volume manager (`partition.zig`) — the engine's the second engine — **exFAT**, `system/services/exfat` — now reuses it wholesale:
own MBR walk currently remains alongside it; write caching stays the reuse this design promised, proven. exfat is nothing but the exFAT engine +
inside the process (the anti-fsyncgate rule). Each mounts its prefixes into the a near-clone of fat's thin service, full read + write + directories + rename +
kernel mount table itself, exactly as today. on-disk up-case folding, differing only in the format it wraps. A partition walk
lives in the volume manager (`partition.zig`), which recognizes fat vs exFAT by
VBR and routes each to its engine; write caching stays inside the process (the
anti-fsyncgate rule). Each mounts its prefixes into the kernel mount table
itself. Two surface limits are shared across both engines and are the vfs
layer's, not an exFAT shortcut: file offsets are u32 (a 4 GiB addressable cap),
and file names are ASCII bytes (a non-ASCII unit becomes `?`) — teaching the vfs
name layer UTF-8 is a separate cross-cutting change.
**Kernel: two small changes only.** `fs_unmount` gains ownership (only the **Kernel: two small changes only.** `fs_unmount` gains ownership (only the
mounting endpoint's holder may unmount — possession-is-capability, consistent mounting endpoint's holder may unmount — possession-is-capability, consistent
@@ -202,18 +209,23 @@ matrix-proven shape; genuinely open.
names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the names it as the 256-byte ceiling's unlock — Fuchsia's FIFO+VMO is the
precedent); format-level crash honesty (a Power-Safe-style journaling or COW precedent); format-level crash honesty (a Power-Safe-style journaling or COW
filesystem) once danos outgrows FAT; per-process namespaces. filesystem) once danos outgrows FAT; per-process namespaces.
7. **The media-presence event** (settled in principle; lands with the volume 7. **The media-presence event** (the consuming half is BUILT; the
manager): the block protocol gains a pushed event — `medium_changed`, with transport-native signal stays future): the block protocol carries a pushed
present/absent and a change counter — produced by the storage driver from event — `medium_changed`, with present/absent and a change counter —
its transport's native signal (SCSI UNIT ATTENTION / TEST UNIT READY for produced today by the storage driver from a TEST UNIT READY poll (the
USB and ATAPI, PxSSTS for AHCI, namespace-change AER for NVMe) and transport's native signal — SCSI UNIT ATTENTION, PxSSTS for AHCI,
consumed by the volume manager, which runs the SAME kill-retire-remount namespace-change AER for NVMe — is the future refinement in place of the
path it runs on channel death — one lifecycle, two triggers. The driver poll) and now **consumed** by the volume manager, which subscribes per
reports presence, never content; a pushed event carries no capability, device and runs the SAME kill-retire-remount path it runs on channel death.
which the kernel already guarantees. The device staying while its medium The driver reports presence, never content; a pushed event carries no
leaves is the one removable-media case the channel-death trigger cannot capability, which the kernel already guarantees. The device staying while
see; without this event a swapped SD card would be served with the old its medium leaves is the one removable-media case the channel-death trigger
card's filesystem state. cannot see; without this event a swapped SD card would be served with the
old card's filesystem state. A THIRD trigger closes the last gap — a
storage driver that *crashes* while its device stays present: channel death
there is invisible to presence polling, so the volume manager probes channel
liveness (`geometry()`) each tick and reaps-then-rebuilds the volume on the
restarted driver's fresh channel. One lifecycle, three triggers.
8. **Volume identity, and the mount map as danos's fstab** (settled). The 8. **Volume identity, and the mount map as danos's fstab** (settled). The
lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and lesson is Linux's own history: fstab keyed on `/dev/sda1` for years and
broke whenever a drive changed ports or enumeration order; `UUID=` entries broke whenever a drive changed ports or enumeration order; `UUID=` entries
@@ -239,8 +251,9 @@ matrix-proven shape; genuinely open.
whether USB port, hub depth, SATA port, or a stick that left as USB and whether USB port, hub depth, SATA port, or a stick that left as USB and
returned in a SATA dock; **replug remounts at the same path** (the id-path is returned in a SATA dock; **replug remounts at the same path** (the id-path is
content-derived, so a volume returns to `/volumes/<id>` wherever it reappears; content-derived, so a volume returns to `/volumes/<id>` wherever it reappears;
remount-on-replug end-to-end is bench-verified, not QEMU-tested, because QEMU the re-adopt+remount code path is QEMU-proven by the driver-crash rebuild, but
can't re-present the boot-controller device); **the boot volume** is the volume remount on a *physical* replug end-to-end is bench-pending, not QEMU-testable,
because QEMU can't re-present the boot-controller device); **the boot volume** is the volume
that resolves `/system/configuration` on its own media, findable on any port or that resolves `/system/configuration` on its own media, findable on any port or
partition; and **duplicate identity is a known S4 gap** — two cloned sticks partition; and **duplicate identity is a known S4 gap** — two cloned sticks
share one content id, so today they collide on `/volumes/<id>` (the kernel share one content id, so today they collide on `/volumes/<id>` (the kernel
+37
View File
@@ -7,12 +7,25 @@
//! `runtime.dma.alloc`), so whole sectors move without crossing the IPC size //! `runtime.dma.alloc`), so whole sectors move without crossing the IPC size
//! limit — the same handoff usb-storage uses toward the controller. //! limit — the same handoff usb-storage uses toward the controller.
const std = @import("std");
const envelope = @import("envelope"); const envelope = @import("envelope");
const ipc = @import("ipc"); const ipc = @import("ipc");
const block_protocol = @import("block-protocol"); const block_protocol = @import("block-protocol");
const Protocol = block_protocol.Protocol; const Protocol = block_protocol.Protocol;
/// The medium_changed event payload, re-exported so a consumer decodes it without
/// reaching into the wire-format module.
pub const MediumChanged = block_protocol.MediumChanged;
/// Decode a medium_changed event from a buffered-message payload a subscriber
/// received (a `Received.isMessage` wake). Null if the bytes are too short to be
/// one — a caller ignores anything that is not a well-formed event.
pub fn decodeMediumChanged(payload: []const u8) ?MediumChanged {
if (payload.len < envelope.prefix_size + @sizeOf(MediumChanged)) return null;
return std.mem.bytesToValue(MediumChanged, payload[envelope.prefix_size..][0..@sizeOf(MediumChanged)]);
}
pub const Geometry = struct { block_size: u32, block_count: u64 }; pub const Geometry = struct { block_size: u32, block_count: u64 };
pub const Device = struct { pub const Device = struct {
@@ -88,6 +101,30 @@ pub const Device = struct {
if (status.status != 0) return null; if (status.status != 0) return null;
return reply[0..answer.len]; return reply[0..answer.len];
} }
/// Subscribe `subscriber` (an endpoint) to this device's medium_changed
/// events: the reserved `subscribe` verb carries the subscriber's endpoint as
/// the capability, and the driver then ipc.sends each medium transition to it.
pub fn subscribeMedium(self: Device, subscriber: ipc.Handle) bool {
var packet: [block_protocol.message_maximum]u8 = undefined;
const framed = envelope.encodeSubscribe(0, &packet) orelse return false; // interest 0: every event (block has one)
var reply: [block_protocol.message_maximum]u8 = undefined;
const answer = ipc.callCap(self.endpoint, framed, &reply, subscriber) catch return false;
const status = envelope.statusOf(reply[0..answer.len]) orelse return false;
return status.status == 0;
}
/// Unsubscribe from this device's medium_changed events. Call before closing
/// the channel so the driver's bounded subscriber table frees the slot rather
/// than holding a dead endpoint until an exit sweep notices.
pub fn unsubscribeMedium(self: Device) bool {
var packet: [block_protocol.message_maximum]u8 = undefined;
const framed = envelope.encodeUnsubscribe(&packet) orelse return false;
var reply: [block_protocol.message_maximum]u8 = undefined;
const answer = ipc.callCap(self.endpoint, framed, &reply, null) catch return false;
const status = envelope.statusOf(reply[0..answer.len]) orelse return false;
return status.status == 0;
}
}; };
// There is deliberately no open-by-name here: `block` is not a registry name. // There is deliberately no open-by-name here: `block` is not a registry name.
+16
View File
@@ -67,6 +67,15 @@ pub const Callbacks = struct {
/// A notification that is not a signal — a subscribed exit event, a bound /// A notification that is not a signal — a subscribed exit event, a bound
/// IRQ, a timer landing. The raw badge; decode with the ipc helpers. /// IRQ, a timer landing. The raw badge; decode with the ipc helpers.
on_notification: ?*const fn (badge: u64) void = null, on_notification: ?*const fn (badge: u64) void = null,
/// A buffered async message (`Received.isMessage`): a pushed event from a
/// provider this service subscribed to, its payload in the receive buffer.
/// Unlike `on_message`, it never goes through the protocol dispatch — so an
/// event whose reserved op number collides with one of this service's own
/// verbs (a `block` `medium_changed` reaching the volume manager, whose own
/// protocol numbers `hello` the same) is decoded by hand here, not
/// mis-dispatched. Default null: the badge alone still reaches
/// `on_notification`, exactly as before this callback existed.
on_buffered_message: ?*const fn (message: []const u8) void = null,
/// The reload signal. Default: ignored. /// The reload signal. Default: ignored.
on_reload: ?*const fn () void = null, on_reload: ?*const fn () void = null,
/// The terminate signal, called before the loop returns. The clean exit is /// The terminate signal, called before the loop returns. The clean exit is
@@ -372,6 +381,13 @@ pub fn run(comptime maximum_message: usize, callbacks: Callbacks) void {
if (got.isChildExit()) { if (got.isChildExit()) {
if (callbacks.subscribers) |subscribers| subscribers.forget(got.childProcessId()); if (callbacks.subscribers) |subscribers| subscribers.forget(got.childProcessId());
} }
// A buffered async message (a pushed event) carries a payload; hand it
// to the service that asked for it. The badge still reaches
// on_notification below, so a coalesced timer/exit riding the same wake
// is not lost — and a service without this callback is unchanged.
if (got.isMessage()) {
if (callbacks.on_buffered_message) |onBuffered| onBuffered(receive[0..got.len]);
}
if (callbacks.on_notification) |onNotification| onNotification(got.badge); if (callbacks.on_notification) |onNotification| onNotification(got.badge);
continue; continue;
} }
+1
View File
@@ -5,3 +5,4 @@
# #
# signature, binary # signature, binary
fat, /system/services/fat fat, /system/services/fat
exfat, /system/services/exfat
1 # The filesystem map: a probed volume's content signature -> the service binary
5 #
6 # signature, binary
7 fat, /system/services/fat
8 exfat, /system/services/exfat
+6
View File
@@ -79,6 +79,7 @@
# 'kernel' as the supervisor. Nothing else changes: the binary must still match. # 'kernel' as the supervisor. Nothing else changes: the binary must still match.
/system/services/input, kernel, bind, input /system/services/input, kernel, bind, input
/system/services/device-manager, kernel, bind, device-manager /system/services/device-manager, kernel, bind, device-manager
/system/services/volume-manager, kernel, bind, volume-manager
/system/services/fat, kernel, bind, vfs /system/services/fat, kernel, bind, vfs
/system/services/display, kernel, bind, display /system/services/display, kernel, bind, display
/system/services/discovery, kernel, bind, power /system/services/discovery, kernel, bind, power
@@ -102,9 +103,14 @@
# own endpoint (the mouse-listener thread opens /protocol/display like any other # own endpoint (the mouse-listener thread opens /protocol/display like any other
# client — threads share no handles), and the input stream that moves the cursor. # client — threads share no handles), and the input stream that moves the cursor.
/system/services/fat, /system/services/volume-manager, open, volume-manager /system/services/fat, /system/services/volume-manager, open, volume-manager
# exfat reaches the volume manager the same way — the second engine, same lineage.
/system/services/exfat, /system/services/volume-manager, open, volume-manager
# The volume manager reaches the device manager to be routed to each storage # The volume manager reaches the device manager to be routed to each storage
# provider's block channel, then confines a filesystem to each volume. # provider's block channel, then confines a filesystem to each volume.
/system/services/volume-manager, /system/services/init, open, device-manager /system/services/volume-manager, /system/services/init, open, device-manager
# ...and again under the kernel supervisor for the manual-tree drills (S5's
# volume-driver-restart spawns the volume manager directly, not via init).
/system/services/volume-manager, kernel, open, device-manager
/system/services/display, /system/services/init, open, scanout /system/services/display, /system/services/init, open, scanout
/system/services/display, /system/services/init, open, display /system/services/display, /system/services/init, open, display
/system/services/display, /system/services/init, open, input /system/services/display, /system/services/init, open, input
Can't render this file because it contains an unexpected character in line 12 and column 15.
+68
View File
@@ -227,6 +227,10 @@ pub fn run(case: []const u8, boot_information: *const BootInformation) void {
usbStorageTest(boot_information); usbStorageTest(boot_information);
} else if (eql(case, "fat-mount")) { } else if (eql(case, "fat-mount")) {
fatMountTest(boot_information); fatMountTest(boot_information);
} else if (eql(case, "exfat-volume")) {
exfatVolumeTest(boot_information);
} else if (eql(case, "volume-driver-restart")) {
volumeDriverRestartTest(boot_information);
} else if (eql(case, "device-list")) { } else if (eql(case, "device-list")) {
deviceListTest(boot_information); deviceListTest(boot_information);
} else if (eql(case, "pci-scan")) { } else if (eql(case, "pci-scan")) {
@@ -3036,6 +3040,32 @@ fn fatMountTest(boot_information: *const BootInformation) void {
result(); result();
} }
/// The exFAT mount chain (S4): boot the full tree, which brings up the USB storage
/// chain. The harness attaches a SECOND device — a data-only exFAT volume — beside
/// the FAT boot volume, so the volume manager spawns the exfat service for it
/// (content-routed, its id-path /volumes/exfat-<serial>). Then spawn exfat-test,
/// which reads the seeded file and mutates through the mount. The reuse of the
/// shared harness by a second engine is proven end to end here.
fn exfatVolumeTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: exfat-volume\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over the initial_ramdisk", false);
result();
return;
}
const ramdisk = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(ramdisk) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(ramdisk);
const init_ok = if (process.spawnBundled("/system/services/init")) true else |_| false;
check("init spawned (boots the tree, incl. the volume manager)", init_ok);
check("exfat-test client spawned", spawnNamed(rd, "exfat-test"));
result();
}
/// Per-sender range confinement (V2a, docs/volume-manager-plan.md): the fixture /// Per-sender range confinement (V2a, docs/volume-manager-plan.md): the fixture
/// acquires the block channel, confines ITSELF to a sub-range, and asserts it /// acquires the block channel, confines ITSELF to a sub-range, and asserts it
/// cannot read past that range or widen it. Boots init in REGISTRY-ONLY mode /// cannot read past that range or widen it. Boots init in REGISTRY-ONLY mode
@@ -3657,6 +3687,44 @@ fn displayReattachTest(boot_information: *const BootInformation) void {
while (true) scheduler.yield(); while (true) scheduler.yield();
} }
/// Storage-driver-crash rebuild (S5): the device manager runs in
/// "test-storage-restart" mode and kills the usb-storage driver once, a moment
/// after its volume has mounted. The driver's device stays in the tree, so the
/// volume manager's presence poll alone would miss the death and leave fat wedged
/// on a dead channel; its channel-liveness probe must notice, reap the volume, and
/// rebuild on the restarted driver's fresh channel — a SECOND mount of the same
/// id-path is the proof. (A pre-S5 manager, checking only device presence, never
/// reaps, so the second mount never appears.)
fn volumeDriverRestartTest(boot_information: *const BootInformation) void {
log("DANOS-TEST-BEGIN: volume-driver-restart\n", .{});
if (boot_information.initial_ramdisk_len == 0) {
check("bootloader handed over an initial_ramdisk", false);
result();
return;
}
const image = @as([*]const u8, @ptrFromInt(boot_handoff.physicalToVirtual(boot_information.initial_ramdisk_base)))[0..boot_information.initial_ramdisk_len];
const rd = initial_ramdisk.Reader.init(image) orelse {
check("initial_ramdisk image is valid", false);
result();
return;
};
process.setInitialRamdisk(image);
_ = spawnRegistry(rd);
var manager: u32 = 0;
var i: u32 = 0;
while (i < rd.count) : (i += 1) {
const item = rd.entry(i) orelse continue;
if (!eql(initial_ramdisk.basename(item.name), "device-manager")) continue;
manager = process.spawnProcessSupervised(item.blob, 4, &.{ item.name, "test-storage-restart" }, scheduler.currentId(), null) catch 0;
break;
}
check("device-manager spawned (test-storage-restart mode)", manager != 0);
check("volume-manager spawned", spawnNamed(rd, "volume-manager"));
check("fat-test client spawned", spawnNamed(rd, "fat-test"));
scheduler.setPriority(1); // below the tree, so it runs
while (true) scheduler.yield();
}
/// Process arguments, end to end: spawn args-echo bare (its argv[0] is the /// Process arguments, end to end: spawn args-echo bare (its argv[0] is the
/// initial-ramdisk name). Instance 1 sees argc == 1 and respawns itself through /// initial-ramdisk name). Instance 1 sees argc == 1 and respawns itself through
/// `system_spawn` with the extra arguments "alpha beta-42" — the syscall argument /// `system_spawn` with the extra arguments "alpha beta-42" — the syscall argument
@@ -170,6 +170,8 @@ var test_usb_killed = false;
var test_pci_restart_mode = false; var test_pci_restart_mode = false;
var test_scanout_restart_mode = false; var test_scanout_restart_mode = false;
var test_scanout_killed = false; var test_scanout_killed = false;
var test_storage_restart_mode = false;
var test_storage_killed = false;
var test_kill_pid: u32 = 0; var test_kill_pid: u32 = 0;
var test_kill_due_ns: u64 = 0; var test_kill_due_ns: u64 = 0;
@@ -600,6 +602,16 @@ fn onHello(_: void, invocation: Invocation(device_manager_protocol.Hello), _: An
test_kill_due_ns = time.clock() + 1_500_000_000; test_kill_due_ns = time.clock() + 1_500_000_000;
_ = time.timerOnce(manager_endpoint, 1600); _ = time.timerOnce(manager_endpoint, 1600);
} }
// Storage-driver-crash drill (S5): once, a moment after usb-storage hellos —
// long enough that its volume has mounted — kill it. The manager re-delegates
// the still-present device to a restarted driver on a fresh channel; the volume
// manager's channel-liveness probe must notice the dead channel and rebuild.
if (test_storage_restart_mode and !test_storage_killed and std.mem.eql(u8, driver.name(), "/system/drivers/usb-storage")) {
test_storage_killed = true;
test_kill_pid = invocation.sender;
test_kill_due_ns = time.clock() + 2_000_000_000; // after the ~0.6s mount
_ = time.timerOnce(manager_endpoint, 2100);
}
return 0; return 0;
} }
@@ -763,6 +775,7 @@ pub fn main(init: process.Init) void {
test_usb_restart_mode = std.mem.eql(u8, mode, "test-usb-restart"); test_usb_restart_mode = std.mem.eql(u8, mode, "test-usb-restart");
test_pci_restart_mode = std.mem.eql(u8, mode, "test-pci-restart"); test_pci_restart_mode = std.mem.eql(u8, mode, "test-pci-restart");
test_scanout_restart_mode = std.mem.eql(u8, mode, "test-scanout-restart"); test_scanout_restart_mode = std.mem.eql(u8, mode, "test-scanout-restart");
test_storage_restart_mode = std.mem.eql(u8, mode, "test-storage-restart");
} }
service.run(device_manager_protocol.message_maximum, .{ service.run(device_manager_protocol.message_maximum, .{
.service = "device-manager", .service = "device-manager",
+34
View File
@@ -0,0 +1,34 @@
//! The exfat service as a binary package (docs/build-packages-plan.md):
//! this file names the binary and EXACTLY the modules its source imports —
//! build-support resolves each name from the domains this zon declares.
const std = @import("std");
const build_support = @import("build-support");
pub fn build(b: *std.Build) void {
const exe = build_support.userBinary(b, .{
.name = "exfat",
.root_source_file = b.path("exfat.zig"),
.imports = &.{
"block", "channel", "envelope", "file-system-harness",
"ipc", "logging", "memory", "process",
"time", "volume-manager-protocol",
},
});
b.installArtifact(exe);
// Standalone `zig build test`; the root aggregate depends on this step.
const test_step = b.step("test", "Run the exfat unit tests");
for ([_][]const u8{
"on-disk.zig", // exFAT on-disk struct sizes + geometry + checksums
"engine.zig", // exFAT read/write over a RAM-backed image
}) |test_root| {
const unit_tests = b.addTest(.{
.root_module = b.createModule(.{
.root_source_file = b.path(test_root),
.target = b.resolveTargetQuery(.{}),
}),
});
test_step.dependOn(&b.addRunArtifact(unit_tests).step);
}
}
+16
View File
@@ -0,0 +1,16 @@
.{
.name = .exfat,
.version = "0.0.0",
.fingerprint = 0x5eafdf02d20f93dd, // Changing this has security and trust implications.
.minimum_zig_version = "0.16.0",
.dependencies = .{
// build-support supplies the shared recipe; kernel is implicit in
// every binary (the root shim + link script live there). The rest
// are exactly the homes of this binary's declared imports.
.@"build-support" = .{ .path = "../../../build-support" },
.kernel = .{ .path = "../../../library/kernel" },
.device = .{ .path = "../../../library/device" },
.protocol = .{ .path = "../../../library/protocol" },
},
.paths = .{""},
}
File diff suppressed because it is too large Load Diff
+200
View File
@@ -0,0 +1,200 @@
//! system/services/exfat — the exFAT filesystem service. Like fat.zig, this is
//! only the format-specific half: it finds its block device, sets up the DMA
//! bounce buffer, mounts the exFAT engine on it, and hands the mounted volume to
//! the shared filesystem harness (library/kernel/file-system-harness), which owns
//! everything else — vfs serving, the open-node table, mount registration, the
//! exit sweep, durable-on-close. The engine (engine.zig) is the pure,
//! host-testable format code; on-disk.zig its byte layout.
//!
//! This service is a near-clone of fat.zig: the second engine reuses the harness
//! wholesale, which is the reuse the storage architecture promised
//! (docs/file-system-development/storage-architecture.md). The block data path
//! never crosses IPC: a DMA bounce buffer is handed to the block driver by
//! physical address, and the engine copies sectors in and out.
const std = @import("std");
const channel = @import("channel");
const volume_manager_protocol = @import("volume-manager-protocol");
const ipc = @import("ipc");
const process = @import("process");
const block = @import("block");
const memory = @import("memory");
const logging = @import("logging");
const time = @import("time");
const engine = @import("engine.zig");
const envelope = @import("envelope");
const harness = @import("file-system-harness");
/// The serving harness, specialized for the exFAT engine. One volume per process.
const Harness = harness.Server(engine.FileSystem);
// The engine's BlockDevice, backed by the `.block` driver plus a DMA bounce
// buffer the driver reads/writes by physical address.
const IpcBlock = struct {
device: block.Device,
bounce: memory.DmaRegion, // engine.max_transfer_sectors * 512 bytes
fn readBlocks(context: *anyopaque, lba: u64, count: u32, buffer: []u8) bool {
const self: *IpcBlock = @ptrCast(@alignCast(context));
if (count == 0 or count > engine.max_transfer_sectors) return false;
const len = count * 512;
if (!self.device.read(lba, count, self.bounce.physical)) return false;
const source: [*]const u8 = @ptrFromInt(self.bounce.virtual);
@memcpy(buffer[0..len], source[0..len]);
return true;
}
fn writeBlocks(context: *anyopaque, lba: u64, count: u32, buffer: []const u8) bool {
const self: *IpcBlock = @ptrCast(@alignCast(context));
if (count == 0 or count > engine.max_transfer_sectors) return false;
const len = count * 512;
const destination: [*]u8 = @ptrFromInt(self.bounce.virtual);
@memcpy(destination[0..len], buffer[0..len]);
if (!self.device.write(lba, count, self.bounce.physical)) return false;
device_dirty = true; // a block reached the device; a close will flush it
return true;
}
};
var ipc_block: IpcBlock = undefined;
// Set whenever a block is written, cleared when the device cache is flushed on a
// file close — so writes are committed to stable media before a power-off.
var device_dirty: bool = false;
var filesystem: engine.FileSystem = undefined;
/// The volume this exFAT process serves, its id given as argv[1] by the volume
/// manager that spawned it. The startup hello names it so the manager returns the
/// right volume's channel.
var my_volume_id: u64 = 0;
/// The volume's own mount path, handed in as argv[2] by the volume manager: the
/// volume's content id-path (e.g. /volumes/exfat-12345678). Defaults to
/// /volumes/exfat only for a bare launch with no argument; the manager always
/// passes it. The slice points into the entry block, valid for the process life.
var volume_mount_prefix: []const u8 = "/volumes/exfat";
/// The mounts this volume installs: its own root, plus — only if it is the boot
/// volume (it resolves /system/configuration) — the two FHS rewrites, so the
/// logger's /system/logs stays decoupled from which volume backs it. Boot-volume
/// detection is by content, so it works no matter which volume carries /system.
/// bound: mounts one volume installs (its root + the two boot rewrites)
/// decided-by: ours
/// protects: the mount_specs array
/// at-limit: truncate - unreachable today (fixed at 3); more configured mounts
/// would need this raised, a deliberate change
/// observed-by: a mount silently missing from the harness's mount log
const maximum_mounts_per_volume = 4;
var mount_specs: [maximum_mounts_per_volume]harness.MountSpec = undefined;
/// Get this volume's block channel from the volume manager (establishment by
/// lineage — `block` is not a registry name). The manager spawned this process,
/// confined it to its partition, and answers the hello with the channel; the
/// channel is range-confined to this process's badge. Null until the manager has
/// the volume ready — this retries.
fn acquireVolume() ?block.Device {
var attempts: u32 = 0;
const vm = while (attempts < 500) : (attempts += 1) {
if (channel.openEndpoint("volume-manager")) |handle| break handle;
time.sleepMillis(20);
} else return null;
attempts = 0;
while (attempts < 500) : (attempts += 1) {
var packet: [volume_manager_protocol.message_maximum]u8 = undefined;
const framed = volume_manager_protocol.Protocol.encodeRequest(.hello, my_volume_id, .{}, &.{}, &packet) orelse return null;
var reply: [volume_manager_protocol.message_maximum]u8 = undefined;
const answered = ipc.callCap(vm, framed, &reply, null) catch return null;
const status = envelope.statusOf(reply[0..answered.len]) orelse return null;
if (status.status != 0) {
if (answered.cap) |stray| _ = ipc.close(stray);
_ = logging.write("/system/services/exfat: volume manager refused the hello\n");
return null;
}
if (answered.cap) |bus| return .{ .endpoint = bus };
// Acked with no channel: the volume is not ready yet — retry.
time.sleepMillis(20);
}
return null;
}
/// Durable-on-close: commit the device write cache if any block reached it since
/// the last flush. The harness calls this on every close; the dirty check keeps
/// it cheap.
fn flushIfDirty() void {
if (device_dirty) {
_ = ipc_block.device.flush();
device_dirty = false;
}
}
/// exFAT bring-up: find the block device, set up DMA, mount the engine, and hand
/// the volume to the harness — or null to retry on the harness's timer.
fn exfatBringUp(endpoint: ipc.Handle) ?Harness.Volume {
_ = endpoint;
const device = acquireVolume() orelse return null;
const geometry = device.geometry() orelse {
_ = logging.write("/system/services/exfat: block geometry unavailable\n");
return null;
};
// Shareable so the buffer's capability can be attached down the chain, making
// its physical addresses reachable under an enforcing IOMMU. No-op otherwise.
const bounce = memory.dmaAlloc(engine.max_transfer_sectors * 512, memory.dma_coherent | memory.dma_shareable) orelse return null;
if (bounce.handle) |handle| {
// Attach, detach, and attach again: the round trip exercises BOTH verbs of
// the DMA-window lifecycle through the whole chain on every boot.
if (!device.attach(handle)) {
_ = logging.write("/system/services/exfat: could not attach the DMA bounce buffer\n");
return null;
}
if (!device.detach(handle)) {
_ = logging.write("/system/services/exfat: could not detach the DMA bounce buffer\n");
return null;
}
if (!device.attach(handle)) {
_ = logging.write("/system/services/exfat: could not re-attach the DMA bounce buffer\n");
return null;
}
_ = ipc.close(handle); // the binding holds its own reference now
}
ipc_block = .{ .device = device, .bounce = bounce };
const block_device = engine.BlockDevice{
.context = &ipc_block,
.block_size = geometry.block_size,
.block_count = geometry.block_count,
.readBlocksFn = IpcBlock.readBlocks,
.writeBlocksFn = IpcBlock.writeBlocks,
};
filesystem = engine.FileSystem.mount(block_device) orelse {
_ = logging.write("/system/services/exfat: not an exFAT filesystem\n");
return null;
};
std.log.info("mounted exFAT ({d} clusters, {d} sectors/cluster, serial 0x{x})", .{ filesystem.geometry.cluster_count, filesystem.geometry.sectors_per_cluster, filesystem.geometry.volume_serial_number });
// The volume mounts at its id-path (argv[2]). The boot/system volume — the one
// carrying the /system tree — additionally installs the two FHS rewrites, by
// CONTENT: it resolves /system/configuration on its own media. A data volume
// mounts only at its id-path and never shadows the running system.
mount_specs[0] = .{ .prefix = volume_mount_prefix };
var mount_count: usize = 1;
if (filesystem.resolve("/system/configuration") != null) {
std.log.info("volume {d} carries the system tree; backing /system/configuration and /system/logs", .{my_volume_id});
mount_specs[1] = .{ .prefix = "/system/configuration", .rewrite = "/system/configuration" };
mount_specs[2] = .{ .prefix = "/system/logs", .rewrite = "/system/logs" };
mount_count = 3;
} else {
std.log.info("volume {d} is a data volume; mounted at {s}", .{ my_volume_id, volume_mount_prefix });
}
return .{ .engine = &filesystem, .mounts = mount_specs[0..mount_count], .flush = flushIfDirty };
}
pub fn main(init: process.Init) void {
// The volume manager spawns this process with its volume id as argv[1] and the
// volume's mount path (its id-path) as argv[2].
if (init.arguments.get(1)) |id| {
my_volume_id = std.fmt.parseInt(u64, id, 10) catch 0;
}
if (init.arguments.get(2)) |prefix| {
volume_mount_prefix = prefix;
}
_ = logging.write("/system/services/exfat: starting, waiting for a block device\n");
Harness.run(.{ .bringUp = exfatBringUp });
}
+475
View File
@@ -0,0 +1,475 @@
//! The on-disk layout of an exFAT filesystem — the Main Boot Sector (VBR) and the
//! six 32-byte directory-entry types — as `align(1)` extern structs that bit-cast
//! straight out of a sector (multi-byte fields little-endian). Pure data, plus the
//! three exFAT checksums (boot region, up-case table, directory-entry set), the
//! name hash, and the packed timestamp <-> Unix-epoch conversion. Host-testable.
//!
//! exFAT departs from FAT in three ways this file encodes: geometry lives in a
//! MustBeZero-guarded VBR (byte 11 is zero, which is exactly why the FAT prober
//! rejects an exFAT volume — it reads a zero bytes-per-sector); a file is a SET of
//! entries (a File entry, a Stream Extension, and one or more File Name entries)
//! validated by a rotate-right checksum; and names are compared case-folded through
//! the volume's own on-disk up-case table (the folding itself lives in the engine,
//! which holds the loaded table; the hash it feeds is here).
const std = @import("std");
// exFAT's fixed on-disk widths — byte and UTF-16-unit counts the FORMAT defines,
// not ceilings danos chooses. Named so the wire-format structs carry no bare
// literal lengths; the values are facts of the spec.
pub const entry_bytes: usize = 32; // every directory entry
const jump_boot_bytes = 3;
const filesystem_name_bytes = 8; // "EXFAT "
const must_be_zero_bytes = 53; // the FAT-BPB overlap the format holds zero
const boot_code_bytes = 390;
const volume_label_units = 11;
// --- the Main Boot Sector (VBR, sector 0) ------------------------------------
/// The exFAT Main Boot Sector. `must_be_zero` (offset 11..64) overlaps where a
/// FAT BPB keeps bytes-per-sector/sectors-per-cluster/etc.; exFAT holds it zero,
/// so a FAT prober reading a zero bytes-per-sector rejects the volume — the
/// mutual-exclusion the two engines rely on.
pub const MainBootSector = extern struct {
jump_boot: [jump_boot_bytes]u8, // 0
filesystem_name: [filesystem_name_bytes]u8, // 3 "EXFAT "
must_be_zero: [must_be_zero_bytes]u8, // 11
partition_offset: u64 align(1), // 64 sectors, informational
volume_length: u64 align(1), // 72 sectors
fat_offset: u32 align(1), // 80 sectors from volume start
fat_length: u32 align(1), // 84 sectors, per FAT
cluster_heap_offset: u32 align(1), // 88 sectors from volume start
cluster_count: u32 align(1), // 92
first_cluster_of_root: u32 align(1), // 96
volume_serial_number: u32 align(1), // 100
filesystem_revision: u16 align(1), // 104
volume_flags: u16 align(1), // 106 (skipped by the boot checksum)
bytes_per_sector_shift: u8, // 108 9..12
sectors_per_cluster_shift: u8, // 109
number_of_fats: u8, // 110 1 (2 for TexFAT)
drive_select: u8, // 111
percent_in_use: u8, // 112 (skipped by the boot checksum)
reserved: [7]u8, // 113
boot_code: [boot_code_bytes]u8, // 120
boot_signature: u16 align(1), // 510 0xAA55
};
// --- directory entries (32 bytes each) ---------------------------------------
/// Entry-type bytes. The high bit (0x80) is InUse: a type with it clear is not in
/// use, and 0x00 ends the directory. Deleting an entry clears bit 7 (0x85 -> 0x05).
pub const entry_type_allocation_bitmap: u8 = 0x81;
pub const entry_type_upcase_table: u8 = 0x82;
pub const entry_type_volume_label: u8 = 0x83;
pub const entry_type_file: u8 = 0x85;
pub const entry_type_stream_extension: u8 = 0xC0;
pub const entry_type_file_name: u8 = 0xC1;
pub const entry_type_in_use_bit: u8 = 0x80;
pub const entry_type_end_of_directory: u8 = 0x00;
/// A raw 32-byte entry, for type dispatch before it is reinterpreted as a
/// specific entry.
pub const RawEntry = extern struct {
entry_type: u8,
data: [entry_bytes - 1]u8,
pub fn inUse(self: RawEntry) bool {
return self.entry_type & entry_type_in_use_bit != 0;
}
pub fn isEnd(self: RawEntry) bool {
return self.entry_type == entry_type_end_of_directory;
}
};
/// 0x81 — the Allocation Bitmap: one bit per cluster (cluster 2 = bit 0), the
/// authority for which clusters are free. The deepest departure from FAT, where
/// the chain itself was the authority.
pub const AllocationBitmapEntry = extern struct {
entry_type: u8, // 0 0x81
bitmap_flags: u8, // 1
reserved: [18]u8, // 2
first_cluster: u32 align(1), // 20
data_length: u64 align(1), // 24
};
/// 0x82 — the Up-case Table: the on-disk case-fold map (code unit -> uppercase),
/// referenced by cluster and validated by `table_checksum`.
pub const UpcaseTableEntry = extern struct {
entry_type: u8, // 0 0x82
reserved1: [3]u8, // 1
table_checksum: u32 align(1), // 4
reserved2: [12]u8, // 8
first_cluster: u32 align(1), // 20
data_length: u64 align(1), // 24
};
/// 0x83 — the Volume Label (up to 11 UTF-16 units).
pub const VolumeLabelEntry = extern struct {
entry_type: u8, // 0 0x83
character_count: u8, // 1
volume_label: [volume_label_units]u16 align(1), // 2
reserved: [8]u8, // 24
};
/// 0x85 — the File entry: the head of a set, carrying the attributes,
/// timestamps, the secondary-entry count, and the set checksum.
pub const FileEntry = extern struct {
entry_type: u8, // 0 0x85
secondary_count: u8, // 1 stream (1) + name entries
set_checksum: u16 align(1), // 2 over the whole set, skipping these two bytes
file_attributes: u16 align(1), // 4
reserved1: u16 align(1), // 6
create_timestamp: u32 align(1), // 8
last_modified_timestamp: u32 align(1), // 12
last_accessed_timestamp: u32 align(1), // 16
create_10ms: u8, // 20
last_modified_10ms: u8, // 21
create_utc_offset: u8, // 22
last_modified_utc_offset: u8, // 23
last_accessed_utc_offset: u8, // 24
reserved2: [7]u8, // 25
};
/// 0xC0 — the Stream Extension: the second entry of every file set, carrying the
/// name length + hash and the data location (first cluster, sizes, the
/// no-FAT-chain flag).
pub const StreamExtensionEntry = extern struct {
entry_type: u8, // 0 0xC0
general_secondary_flags: u8, // 1
reserved1: u8, // 2
name_length: u8, // 3 UTF-16 units
name_hash: u16 align(1), // 4
reserved2: u16 align(1), // 6
valid_data_length: u64 align(1), // 8
reserved3: u32 align(1), // 16
first_cluster: u32 align(1), // 20
data_length: u64 align(1), // 24
};
/// 0xC1 — a File Name entry: 15 UTF-16 units of the name; a set carries
/// ceil(name_length / 15) of them.
pub const FileNameEntry = extern struct {
entry_type: u8, // 0 0xC1
general_secondary_flags: u8, // 1
file_name: [name_units_per_entry]u16 align(1), // 2
};
pub const name_units_per_entry: usize = 15;
// General secondary flags (Stream Extension + File Name entries).
pub const secondary_flag_allocation_possible: u8 = 0x01;
pub const secondary_flag_no_fat_chain: u8 = 0x02;
// File attributes (same bit assignments as FAT).
pub const attribute_read_only: u16 = 0x0001;
pub const attribute_hidden: u16 = 0x0002;
pub const attribute_system: u16 = 0x0004;
pub const attribute_directory: u16 = 0x0010;
pub const attribute_archive: u16 = 0x0020;
// FAT special cluster values (exFAT's FAT is 32-bit; used only for a fragmented
// chain, i.e. when no_fat_chain is clear).
pub const first_data_cluster: u32 = 2;
pub const end_of_chain: u32 = 0xFFFFFFFF;
pub const bad_cluster: u32 = 0xFFFFFFF7;
pub const boot_signature_offset: usize = 510; // 0x55 0xAA
// --- geometry ----------------------------------------------------------------
pub const Geometry = struct {
bytes_per_sector: u32,
sectors_per_cluster: u32,
cluster_count: u32,
fat_offset_sectors: u32, // from volume start
fat_length_sectors: u32,
cluster_heap_offset_sectors: u32, // from volume start
first_cluster_of_root: u32,
volume_serial_number: u32,
volume_length_sectors: u64,
number_of_fats: u32,
};
/// Derive the geometry from a Main Boot Sector. Returns null unless it is a
/// plausible exFAT VBR: the "EXFAT " name, an all-zero MustBeZero region, the
/// 0xAA55 signature, and sane shifts. Accepting ONLY these is what keeps exFAT and
/// FAT from ever claiming each other's volumes.
pub fn geometryOf(sector: []const u8) ?Geometry {
if (sector.len < 512) return null;
if (sector[boot_signature_offset] != 0x55 or sector[boot_signature_offset + 1] != 0xAA) return null;
const vbr = std.mem.bytesToValue(MainBootSector, sector[0..@sizeOf(MainBootSector)]);
if (!std.mem.eql(u8, &vbr.filesystem_name, "EXFAT ")) return null;
for (vbr.must_be_zero) |byte| if (byte != 0) return null;
if (vbr.bytes_per_sector_shift < 9 or vbr.bytes_per_sector_shift > 12) return null;
// The exFAT spec caps a cluster at 2^25 bytes (32 MiB): bytes-per-sector-shift
// plus sectors-per-cluster-shift must not exceed 25. Enforcing it here is also
// what keeps the engine's u32 cluster-byte arithmetic (sectors_per_cluster *
// 512) from overflowing on a crafted VBR off untrusted removable media.
if (@as(u16, vbr.bytes_per_sector_shift) + vbr.sectors_per_cluster_shift > 25) return null;
// cluster_count is capped at 0xFFFFFFF5 (the spec's ClusterCount maximum), so
// cluster_count + first_data_cluster cannot overflow u32 in the bounds checks.
if (vbr.number_of_fats == 0 or vbr.cluster_count == 0 or vbr.cluster_count > 0xFFFFFFF5) return null;
if (vbr.first_cluster_of_root < first_data_cluster) return null;
return .{
.bytes_per_sector = @as(u32, 1) << @intCast(vbr.bytes_per_sector_shift),
.sectors_per_cluster = @as(u32, 1) << @intCast(vbr.sectors_per_cluster_shift),
.cluster_count = vbr.cluster_count,
.fat_offset_sectors = vbr.fat_offset,
.fat_length_sectors = vbr.fat_length,
.cluster_heap_offset_sectors = vbr.cluster_heap_offset,
.first_cluster_of_root = vbr.first_cluster_of_root,
.volume_serial_number = vbr.volume_serial_number,
.volume_length_sectors = vbr.volume_length,
.number_of_fats = vbr.number_of_fats,
};
}
// --- checksums and the name hash ---------------------------------------------
/// The directory-entry-SET checksum (a File entry's `set_checksum`), a 16-bit
/// rotate-right sum over every byte of the set, skipping the two checksum bytes
/// themselves (offset 2..3 of the first entry). `entries` is the whole set:
/// (secondary_count + 1) * 32 bytes.
pub fn setChecksum(entries: []const u8) u16 {
var checksum: u16 = 0;
for (entries, 0..) |byte, i| {
if (i == 2 or i == 3) continue;
checksum = std.math.rotr(u16, checksum, 1) +% byte;
}
return checksum;
}
/// The up-case-table checksum (an Up-case entry's `table_checksum`), a 32-bit
/// rotate-right sum over the table's on-disk bytes.
pub fn upcaseChecksum(table_bytes: []const u8) u32 {
var checksum: u32 = 0;
for (table_bytes) |byte| checksum = std.math.rotr(u32, checksum, 1) +% byte;
return checksum;
}
/// The boot-region checksum — the u32 the checksum sector repeats — a 32-bit
/// rotate-right sum over the first eleven sectors, skipping VolumeFlags (offset
/// 106..107) and PercentInUse (offset 112) of the first sector.
pub fn bootChecksum(region: []const u8) u32 {
var checksum: u32 = 0;
for (region, 0..) |byte, i| {
if (i == 106 or i == 107 or i == 112) continue;
checksum = std.math.rotr(u32, checksum, 1) +% byte;
}
return checksum;
}
/// The name hash a Stream entry carries: a 16-bit rotate-right sum over the
/// UP-CASED name's bytes (low byte then high byte of each UTF-16 unit). The caller
/// up-cases through the volume's table first; a mismatch lets a lookup reject a
/// name without reading its File Name entries.
pub fn nameHash(upcased: []const u16) u16 {
var hash: u16 = 0;
for (upcased) |unit| {
hash = std.math.rotr(u16, hash, 1) +% @as(u8, @truncate(unit));
hash = std.math.rotr(u16, hash, 1) +% @as(u8, @truncate(unit >> 8));
}
return hash;
}
// --- timestamps --------------------------------------------------------------
//
// exFAT packs a timestamp into one u32: the high 16 bits are a DOS date
// (year-1980 | month | day), the low 16 a DOS time (hour | minute | second/2).
// Same field layout as FAT, so the epoch math matches; there is no timezone in
// the packed value (a separate UTC-offset byte carries that, which danos leaves
// zero = UTC).
fn isLeapYear(year: u32) bool {
return (year % 4 == 0 and year % 100 != 0) or (year % 400 == 0);
}
const days_in_month = [_]u8{ 31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31 };
/// Convert a packed exFAT timestamp to Unix epoch seconds (UTC). 0 for unset.
pub fn timestampToEpoch(timestamp: u32) u64 {
if (timestamp == 0) return 0;
const date: u32 = timestamp >> 16;
const time: u32 = timestamp & 0xFFFF;
const day: u32 = date & 0x1F;
const month: u32 = (date >> 5) & 0x0F;
const year: u32 = 1980 + (date >> 9);
if (month < 1 or month > 12 or day < 1) return 0;
const second: u32 = (time & 0x1F) * 2;
const minute: u32 = (time >> 5) & 0x3F;
const hour: u32 = (time >> 11) & 0x1F;
var days: u64 = 0;
var y: u32 = 1970;
while (y < year) : (y += 1) days += if (isLeapYear(y)) 366 else 365;
var m: u32 = 1;
while (m < month) : (m += 1) {
days += days_in_month[m - 1];
if (m == 2 and isLeapYear(year)) days += 1;
}
days += day - 1;
return ((days * 24 + hour) * 60 + minute) * 60 + second;
}
/// Convert Unix epoch seconds (UTC) to a packed exFAT timestamp. 0 for epoch 0 or
/// any time before 1980 (unrepresentable).
pub fn epochToTimestamp(epoch: u64) u32 {
if (epoch == 0) return 0;
var remaining = epoch;
const second: u32 = @intCast(remaining % 60);
remaining /= 60;
const minute: u32 = @intCast(remaining % 60);
remaining /= 60;
const hour: u32 = @intCast(remaining % 24);
remaining /= 24;
var days: u32 = @intCast(remaining);
var year: u32 = 1970;
while (true) {
const year_days: u32 = if (isLeapYear(year)) 366 else 365;
if (days < year_days) break;
days -= year_days;
year += 1;
}
if (year < 1980) return 0;
var month: u32 = 1;
while (true) {
var month_days: u32 = days_in_month[month - 1];
if (month == 2 and isLeapYear(year)) month_days += 1;
if (days < month_days) break;
days -= month_days;
month += 1;
}
const day = days + 1;
const date: u32 = ((year - 1980) << 9) | (month << 5) | day;
const time: u32 = (hour << 11) | (minute << 5) | (second / 2);
return (date << 16) | time;
}
// --- tests -------------------------------------------------------------------
test "on-disk struct sizes match the specification" {
try std.testing.expectEqual(@as(usize, 512), @sizeOf(MainBootSector));
try std.testing.expectEqual(@as(usize, 32), @sizeOf(RawEntry));
try std.testing.expectEqual(@as(usize, 32), @sizeOf(AllocationBitmapEntry));
try std.testing.expectEqual(@as(usize, 32), @sizeOf(UpcaseTableEntry));
try std.testing.expectEqual(@as(usize, 32), @sizeOf(VolumeLabelEntry));
try std.testing.expectEqual(@as(usize, 32), @sizeOf(FileEntry));
try std.testing.expectEqual(@as(usize, 32), @sizeOf(StreamExtensionEntry));
try std.testing.expectEqual(@as(usize, 32), @sizeOf(FileNameEntry));
}
test "MainBootSector field offsets" {
try std.testing.expectEqual(@as(usize, 3), @offsetOf(MainBootSector, "filesystem_name"));
try std.testing.expectEqual(@as(usize, 11), @offsetOf(MainBootSector, "must_be_zero"));
try std.testing.expectEqual(@as(usize, 80), @offsetOf(MainBootSector, "fat_offset"));
try std.testing.expectEqual(@as(usize, 88), @offsetOf(MainBootSector, "cluster_heap_offset"));
try std.testing.expectEqual(@as(usize, 96), @offsetOf(MainBootSector, "first_cluster_of_root"));
try std.testing.expectEqual(@as(usize, 106), @offsetOf(MainBootSector, "volume_flags"));
try std.testing.expectEqual(@as(usize, 112), @offsetOf(MainBootSector, "percent_in_use"));
try std.testing.expectEqual(@as(usize, 510), @offsetOf(MainBootSector, "boot_signature"));
// The Stream Extension's data location must sit where the spec places it.
try std.testing.expectEqual(@as(usize, 20), @offsetOf(StreamExtensionEntry, "first_cluster"));
try std.testing.expectEqual(@as(usize, 24), @offsetOf(StreamExtensionEntry, "data_length"));
}
test "geometryOf accepts exFAT and the MustBeZero guard rejects a FAT-shaped sector" {
var sector = [_]u8{0} ** 512;
@memcpy(sector[3..11], "EXFAT ");
sector[510] = 0x55;
sector[511] = 0xAA;
// fat_offset=128, fat_length=64, cluster_heap_offset=256, cluster_count=1000,
// root cluster=5, bytes/sector=512 (shift 9), sectors/cluster=8 (shift 3), 1 FAT.
std.mem.writeInt(u32, sector[80..84], 128, .little);
std.mem.writeInt(u32, sector[84..88], 64, .little);
std.mem.writeInt(u32, sector[88..92], 256, .little);
std.mem.writeInt(u32, sector[92..96], 1000, .little);
std.mem.writeInt(u32, sector[96..100], 5, .little);
sector[108] = 9; // bytes_per_sector_shift
sector[109] = 3; // sectors_per_cluster_shift
sector[110] = 1; // number_of_fats
const geo = geometryOf(&sector) orelse return error.ShouldParse;
try std.testing.expectEqual(@as(u32, 512), geo.bytes_per_sector);
try std.testing.expectEqual(@as(u32, 8), geo.sectors_per_cluster);
try std.testing.expectEqual(@as(u32, 1000), geo.cluster_count);
try std.testing.expectEqual(@as(u32, 5), geo.first_cluster_of_root);
// A non-zero byte in MustBeZero (where a FAT BPB keeps bytes-per-sector) is
// rejected — the mutual exclusion between the engines.
sector[11] = 0x02;
try std.testing.expect(geometryOf(&sector) == null);
sector[11] = 0;
// Wrong name is rejected too.
sector[3] = 'F';
try std.testing.expect(geometryOf(&sector) == null);
sector[3] = 'E';
}
test "geometryOf rejects crafted VBRs that would overflow u32 cluster arithmetic" {
var sector = [_]u8{0} ** 512;
@memcpy(sector[3..11], "EXFAT ");
sector[510] = 0x55;
sector[511] = 0xAA;
std.mem.writeInt(u32, sector[92..96], 1000, .little); // cluster_count
std.mem.writeInt(u32, sector[96..100], 5, .little); // root cluster
sector[108] = 9; // bytes_per_sector_shift
sector[110] = 1; // number_of_fats
// A cluster shift past the exFAT ceiling (9 + 17 = 26 > 25) would make
// sectors_per_cluster * 512 overflow u32 — rejected.
sector[109] = 17;
try std.testing.expect(geometryOf(&sector) == null);
sector[109] = 3; // sane again
try std.testing.expect(geometryOf(&sector) != null);
// cluster_count above the spec maximum (0xFFFFFFF5) would overflow
// cluster_count + first_data_cluster in the bounds checks — rejected.
std.mem.writeInt(u32, sector[92..96], 0xFFFFFFFF, .little);
try std.testing.expect(geometryOf(&sector) == null);
}
test "set checksum skips its own two bytes and depends on the rest" {
var set = [_]u8{0} ** 64; // a File entry + one secondary
set[0] = entry_type_file;
set[1] = 1;
set[4] = 0x20; // an attribute byte
set[40] = 0xAB; // a byte in the secondary entry
const base = setChecksum(&set);
// Changing the checksum field itself must NOT change the computed checksum.
set[2] = 0xFF;
set[3] = 0xEE;
try std.testing.expectEqual(base, setChecksum(&set));
// Changing any other byte MUST change it.
set[4] = 0x21;
try std.testing.expect(setChecksum(&set) != base);
}
test "name hash is deterministic and order-sensitive" {
const readme = [_]u16{ 'R', 'E', 'A', 'D', 'M', 'E' };
const different = [_]u16{ 'E', 'R', 'A', 'D', 'M', 'E' };
try std.testing.expectEqual(nameHash(&readme), nameHash(&readme));
try std.testing.expect(nameHash(&readme) != nameHash(&different));
}
test "boot checksum skips VolumeFlags and PercentInUse" {
var region = [_]u8{0} ** 1536; // three 512-byte sectors is enough to exercise the skips
region[64] = 0x11;
const base = bootChecksum(&region);
for ([_]usize{ 106, 107, 112 }) |skipped| {
var copy = region;
copy[skipped] = 0xFF;
try std.testing.expectEqual(base, bootChecksum(&copy));
}
var copy = region;
copy[108] = 0xFF; // a non-skipped byte
try std.testing.expect(bootChecksum(&copy) != base);
}
test "exFAT timestamp <-> Unix epoch round trip" {
for ([_]u64{ 1_577_836_800, 1_700_000_000, 1_262_304_000, 1_783_971_244 }) |epoch| {
try std.testing.expectEqual(epoch, timestampToEpoch(epochToTimestamp(epoch)));
}
// 1577836800 is 2020-01-01 00:00:00 UTC.
const stamp = epochToTimestamp(1_577_836_800);
try std.testing.expectEqual(@as(u32, 2020), 1980 + (stamp >> 16 >> 9));
try std.testing.expectEqual(@as(u64, 0), timestampToEpoch(0));
try std.testing.expectEqual(@as(u32, 0), epochToTimestamp(0));
}
+18
View File
@@ -1657,3 +1657,21 @@ test "short-name checksum matches the reference vector" {
const c = FileSystem.shortChecksum("REDAME TXT".*); const c = FileSystem.shortChecksum("REDAME TXT".*);
try std.testing.expect(a != c); try std.testing.expect(a != c);
} }
test "the FAT engine rejects an exFAT volume (mutual exclusion at mount)" {
const allocator = std.testing.allocator;
const bytes = try allocator.alloc(u8, 5000 * sector_size);
defer allocator.free(bytes);
@memset(bytes, 0);
// An exFAT boot sector: the "EXFAT " name and 0x55AA, but MustBeZero (offset
// 11, where a FAT BPB keeps bytes-per-sector) stays zero — so this engine's
// geometryOf reads a zero bytes-per-sector and rejects it.
@memcpy(bytes[3..11], "EXFAT ");
bytes[on_disk.boot_signature_offset] = 0x55;
bytes[on_disk.boot_signature_offset + 1] = 0xAA;
var disk = RamDisk{ .bytes = bytes };
try std.testing.expect(FileSystem.mount(disk.device()) == null);
// Control: a real FAT16 mounts.
formatFat16(bytes);
try std.testing.expect(FileSystem.mount(disk.device()) != null);
}
+67 -8
View File
@@ -31,10 +31,11 @@ pub const label_maximum = 36;
/// stay distinct, and it drives how the mount path is rendered from the id. /// stay distinct, and it drives how the mount path is rendered from the id.
pub const Rung = enum(u8) { pub const Rung = enum(u8) {
gpt_guid = 1, gpt_guid = 1,
filesystem_uuid = 2, // reserved: no non-FAT engine reads a superblock UUID yet filesystem_uuid = 2, // reserved: no engine reads a superblock UUID yet
fat_serial = 3, fat_serial = 3,
mbr_index = 4, mbr_index = 4,
anonymous = 5, anonymous = 5,
exfat_serial = 6, // exFAT's VolumeSerialNumber — content-strong like fat_serial
}; };
/// A volume's content identity. `key` is the ID — the stable, unique handle the /// A volume's content identity. `key` is the ID — the stable, unique handle the
@@ -61,15 +62,16 @@ pub const Identity = struct {
}; };
/// Which filesystem a volume's content is — the key `filesystems.csv` maps to a /// Which filesystem a volume's content is — the key `filesystems.csv` maps to a
/// service binary. Today only FAT is recognized (S4 adds exFAT with a real VBR /// service binary. FAT and exFAT are recognized by their VBRs; content that is
/// recognizer); until then every probed volume is `.fat`, matching the volume /// neither falls back to `.fat`, the volume manager's historical hand-off.
/// manager's historical hand-off of everything to the FAT service.
pub const FilesystemKind = enum { pub const FilesystemKind = enum {
fat, fat,
exfat,
unknown, unknown,
pub fn fromToken(token: []const u8) FilesystemKind { pub fn fromToken(token: []const u8) FilesystemKind {
if (std.mem.eql(u8, token, "fat")) return .fat; if (std.mem.eql(u8, token, "fat")) return .fat;
if (std.mem.eql(u8, token, "exfat")) return .exfat;
return .unknown; return .unknown;
} }
}; };
@@ -226,7 +228,7 @@ fn gptAllVolumes(reader: SectorReader, device_blocks: u64, out: []Volume) usize
if (start == 0 or end < start or end >= device_blocks) continue; if (start == 0 or end < start or end >= device_blocks) continue;
var id = Identity{ .rung = .gpt_guid, .key = std.mem.readInt(u128, entry[16..32], .little) }; var id = Identity{ .rung = .gpt_guid, .key = std.mem.readInt(u128, entry[16..32], .little) };
setLabelFromUtf16(&id, entry[56..128]); setLabelFromUtf16(&id, entry[56..128]);
out[count] = .{ .base_lba = start, .block_count = end - start + 1, .identity = id }; out[count] = .{ .base_lba = start, .block_count = end - start + 1, .identity = id, .signature = signatureAt(reader, start) };
count += 1; count += 1;
} }
return count; return count;
@@ -262,6 +264,36 @@ fn fatIdentity(reader: SectorReader, start_lba: u64) ?Identity {
return id; return id;
} }
/// The exFAT VolumeSerialNumber (offset 100) read from the Main Boot Sector at
/// `start_lba` — its content identity, rung `exfat_serial`. Null unless the sector
/// is an exFAT VBR (the "EXFAT " name at offset 3 + the 0x55AA signature; the
/// name is where a FAT BPB keeps its OEM string, so the two never collide). The
/// label lives in a root-directory entry, not the VBR, so it is left empty here.
fn exfatIdentity(reader: SectorReader, start_lba: u64) ?Identity {
var vbr: [sector_bytes]u8 = undefined;
if (!reader.read(start_lba, &vbr)) return null;
if (vbr[510] != 0x55 or vbr[511] != 0xAA) return null;
if (!std.mem.eql(u8, vbr[3..11], "EXFAT ")) return null;
return .{ .rung = .exfat_serial, .key = std.mem.readInt(u32, vbr[100..104], .little) };
}
const Recognized = struct { identity: Identity, signature: FilesystemKind };
/// Recognize the filesystem at `start_lba` by its VBR: exFAT first (its serial and
/// the `.exfat` signature), else FAT (its serial), else unknown content that keeps
/// the MBR disk-signature identity and the historical `.fat` hand-off.
fn recognize(reader: SectorReader, start_lba: u64, block0: *const [sector_bytes]u8, index: u8) Recognized {
if (exfatIdentity(reader, start_lba)) |id| return .{ .identity = id, .signature = .exfat };
if (fatIdentity(reader, start_lba)) |id| return .{ .identity = id, .signature = .fat };
return .{ .identity = mbrIdentity(block0, index), .signature = .fat };
}
/// The filesystem signature at `start_lba` when the identity is decided elsewhere
/// (a GPT partition keeps its GUID identity but still needs its content's kind).
fn signatureAt(reader: SectorReader, start_lba: u64) FilesystemKind {
return if (exfatIdentity(reader, start_lba) != null) .exfat else .fat;
}
/// Append every volume on the device `reader` addresses, whose whole-device size /// Append every volume on the device `reader` addresses, whose whole-device size
/// is `device_blocks`, to `out` (up to `out.len`), returning the count. A GPT /// is `device_blocks`, to `out` (up to `out.len`), returning the count. A GPT
/// disk (protective MBR) is enumerated by GPT, authoritatively — a zero count is /// disk (protective MBR) is enumerated by GPT, authoritatively — a zero count is
@@ -290,12 +322,14 @@ pub fn allVolumes(reader: SectorReader, device_blocks: u64, out: []Volume) usize
// device (usb-storage.zig resolveTransfer), which only holds because the // device (usb-storage.zig resolveTransfer), which only holds because the
// range handed down is validated here. The subtraction cannot overflow. // range handed down is validated here. The subtraction cannot overflow.
if (start > device_blocks or device_blocks - start < size) continue; if (start > device_blocks or device_blocks - start < size) continue;
out[count] = .{ .base_lba = start, .block_count = size, .identity = fatIdentity(reader, start) orelse mbrIdentity(&block0, index) }; const found = recognize(reader, start, &block0, index);
out[count] = .{ .base_lba = start, .block_count = size, .identity = found.identity, .signature = found.signature };
count += 1; count += 1;
} }
if (count == 0 and out.len > 0) { if (count == 0 and out.len > 0) {
// No partition entries: a bare FAT spanning the device. // No partition entries: a bare FAT or exFAT spanning the device.
out[0] = .{ .base_lba = 0, .block_count = device_blocks, .identity = fatIdentity(reader, 0) orelse mbrIdentity(&block0, 0) }; const found = recognize(reader, 0, &block0, 0);
out[0] = .{ .base_lba = 0, .block_count = device_blocks, .identity = found.identity, .signature = found.signature };
return 1; return 1;
} }
return count; return count;
@@ -360,6 +394,31 @@ test "no boot signature is no volume" {
const disk = RamDisk{ .sectors = &block0 }; const disk = RamDisk{ .sectors = &block0 };
try std.testing.expect(firstVolume(disk.reader(), 65536) == null); try std.testing.expect(firstVolume(disk.reader(), 65536) == null);
} }
test "a bare exFAT volume is recognized by its VBR, with its serial as the id" {
var block0 = [_]u8{0} ** 512;
@memcpy(block0[3..11], "EXFAT ");
block0[510] = 0x55;
block0[511] = 0xAA;
std.mem.writeInt(u32, block0[100..104], 0xDA7A0001, .little); // VolumeSerialNumber
const disk = RamDisk{ .sectors = &block0 };
const v = firstVolume(disk.reader(), 65536).?;
try std.testing.expectEqual(FilesystemKind.exfat, v.signature);
try std.testing.expectEqual(Rung.exfat_serial, v.identity.rung);
try std.testing.expectEqual(@as(u128, 0xDA7A0001), v.identity.key);
}
test "a FAT VBR is recognized as fat, not exfat — the signatures never collide" {
var block0 = [_]u8{0} ** 512;
@memcpy(block0[3..11], "MSWIN4.1"); // a FAT OEM name, not "EXFAT "
block0[510] = 0x55;
block0[511] = 0xAA;
std.mem.writeInt(u16, block0[22..24], 16, .little); // fat_size_16 != 0 -> FAT16 shape
block0[38] = 0x29; // extended boot signature
std.mem.writeInt(u32, block0[39..43], 0x12345678, .little); // volume id
const disk = RamDisk{ .sectors = &block0 };
const v = firstVolume(disk.reader(), 65536).?;
try std.testing.expectEqual(FilesystemKind.fat, v.signature);
try std.testing.expectEqual(Rung.fat_serial, v.identity.rung);
}
test "a partition that runs past the device is skipped, not trusted" { test "a partition that runs past the device is skipped, not trusted" {
var block0 = [_]u8{0} ** 512; var block0 = [_]u8{0} ** 512;
@@ -306,6 +306,12 @@ fn bringUpVolume() bool {
return false; return false;
}; };
dev.* = .{ .used = true, .device_id = opened.device_id, .channel = opened.device }; dev.* = .{ .used = true, .device_id = opened.device_id, .channel = opened.device };
// Consume this device's medium_changed events (the second of the removal
// lifecycle's two triggers: the device stays in the tree while its medium
// leaves — a card reader, an eject). Best effort: a provider that never
// publishes the event simply never wakes us, and device-pull is still caught
// by the presence poll.
_ = opened.device.subscribeMedium(service_endpoint);
const device = opened.device; const device = opened.device;
// Attach the read buffer to THIS device (a no-op without an enforcing IOMMU). // Attach the read buffer to THIS device (a no-op without an enforcing IOMMU).
// The handle is kept, not closed, so it can be re-attached after a replug. A // The handle is kept, not closed, so it can be re-attached after a replug. A
@@ -370,6 +376,10 @@ fn bringUpVolume() bool {
/// Close a device's channel and free its slot. No volumes are touched (the caller /// Close a device's channel and free its slot. No volumes are touched (the caller
/// ensures none remain, or there never were any). /// ensures none remain, or there never were any).
fn dropDevice(dev: *StorageDevice) void { fn dropDevice(dev: *StorageDevice) void {
// Free the driver's subscriber slot before the channel closes. On a still-live
// channel (a medium eject) this frees the slot; on a dead one (a device pull)
// the call fails fast and the exit sweep frees it anyway.
_ = dev.channel.unsubscribeMedium();
_ = ipc.close(dev.channel.endpoint); _ = ipc.close(dev.channel.endpoint);
dev.* = .{}; dev.* = .{};
} }
@@ -395,13 +405,25 @@ fn removeDevice(dev: *StorageDevice) void {
dropDevice(dev); dropDevice(dev);
} }
/// Whether a device's block channel still answers — a geometry() probe. A storage
/// driver that DIED while its device stays in the tree (it crashed; the device
/// manager will re-delegate the device to a restarted driver on a FRESH channel)
/// leaves a dead channel here, even though isDevicePresent still reports the device
/// present. geometry() on the dead endpoint fails fast, so this catches the crash
/// that presence-polling alone cannot — the V4 review's open edge.
fn channelAlive(dev: *StorageDevice) bool {
return dev.channel.geometry() != null;
}
/// One poll tick. Device removal is reconciled FIRST and supersedes a pending /// One poll tick. Device removal is reconciled FIRST and supersedes a pending
/// restart: a volume whose device left is retired before its restart could fire, /// restart: a volume whose device left (a pull) OR whose driver died on a channel
/// so nothing respawns against a dead channel. Then due restarts fire for present /// that no longer answers is retired before its restart could fire, so nothing
/// volumes; then, if no device is adopted, a present device is brought up. /// respawns against a dead channel. Dropping the device frees its slot, so the
/// adopt loop below re-adopts the still-present device on the restarted driver's
/// fresh channel — the rebuild. Then due restarts fire for present volumes.
fn pollTick() void { fn pollTick() void {
for (&devices) |*dev| { for (&devices) |*dev| {
if (dev.used and !isDevicePresent(dev.device_id)) removeDevice(dev); if (dev.used and (!isDevicePresent(dev.device_id) or !channelAlive(dev))) removeDevice(dev);
} }
for (&volumes) |*v| { for (&volumes) |*v| {
if (v.used and v.restart_pending and time.clock() >= v.restart_due_ns) { if (v.used and v.restart_pending and time.clock() >= v.restart_due_ns) {
@@ -533,6 +555,43 @@ fn onNotification(badge: u64) void {
} }
} }
fn anyVolumeOn(device_id: u64) bool {
for (&volumes) |*v| if (v.used and v.device_id == device_id) return true;
return false;
}
/// A storage device published `medium_changed` — the second removal trigger: the
/// device stays in the tree while its medium leaves or returns (a card reader, an
/// eject). This arrives as a buffered async message, NOT a protocol request, so it
/// never reaches `Serve.dispatch` (its event op number collides with the manager's
/// own `hello`); it is decoded here by hand. Single-volume scope: the event names
/// no device, so `absent` retires every adopted device (its volumes unmount and
/// the poll re-adopts the still-present device with its now-empty medium), and
/// `present` frees any empty adopted device so the poll re-probes and remounts it.
///
/// We act on every edge and do NOT dedup on `change_count`. The driver publishes
/// exactly once per transition, each with a unique monotonic count, so a count is
/// never legitimately re-sent within one subscription — an equality dedup could
/// only ever fire spuriously, and it did: `change_count` restarts at 0 in each
/// driver instance (usb-storage.zig), so a global "last count" carried across a
/// driver restart (S5's own crash-rebuild) mistook the fresh instance's first
/// edge for a re-delivery and dropped a real eject, wedging a mount over absent
/// media. Both branches are idempotent (a freed device stops matching `dev.used`)
/// and the poll reconciles, so reacting to each genuine edge is safe.
fn onMediumEvent(payload: []const u8) void {
const event = block.decodeMediumChanged(payload) orelse return;
if (event.present == 0) {
std.log.info("medium left a storage device; unmounting its volume(s)", .{});
for (&devices) |*dev| {
if (dev.used) removeDevice(dev);
}
} else {
for (&devices) |*dev| {
if (dev.used and !anyVolumeOn(dev.device_id)) removeDevice(dev);
}
}
}
pub fn main(init: process.Init) void { pub fn main(init: process.Init) void {
_ = init; _ = init;
service.run(volume_manager_protocol.message_maximum, .{ service.run(volume_manager_protocol.message_maximum, .{
@@ -540,5 +599,6 @@ pub fn main(init: process.Init) void {
.init = initialise, .init = initialise,
.on_message = onMessage, .on_message = onMessage,
.on_notification = onNotification, .on_notification = onNotification,
.on_buffered_message = onMediumEvent,
}); });
} }
@@ -42,6 +42,7 @@ pub fn idString(identity: partition.Identity, buf: []u8) []const u8 {
.gpt_guid => std.fmt.bufPrint(buf, "gpt-{x:0>32}", .{identity.key}) catch "", .gpt_guid => std.fmt.bufPrint(buf, "gpt-{x:0>32}", .{identity.key}) catch "",
.filesystem_uuid => std.fmt.bufPrint(buf, "uuid-{x:0>32}", .{identity.key}) catch "", .filesystem_uuid => std.fmt.bufPrint(buf, "uuid-{x:0>32}", .{identity.key}) catch "",
.fat_serial => std.fmt.bufPrint(buf, "fat-{x:0>8}", .{@as(u32, @truncate(identity.key))}) catch "", .fat_serial => std.fmt.bufPrint(buf, "fat-{x:0>8}", .{@as(u32, @truncate(identity.key))}) catch "",
.exfat_serial => std.fmt.bufPrint(buf, "exfat-{x:0>8}", .{@as(u32, @truncate(identity.key))}) catch "",
.mbr_index => std.fmt.bufPrint(buf, "mbr-{x}-{d}", .{ .mbr_index => std.fmt.bufPrint(buf, "mbr-{x}-{d}", .{
@as(u32, @truncate(identity.key >> 8)), @as(u32, @truncate(identity.key >> 8)),
@as(u8, @truncate(identity.key & 0xff)), @as(u8, @truncate(identity.key & 0xff)),
@@ -105,6 +106,7 @@ const test_override_slots = 4;
test "idString renders each rung's id token" { test "idString renders each rung's id token" {
var buf: [id_maximum]u8 = undefined; var buf: [id_maximum]u8 = undefined;
try testing.expectEqualStrings("fat-12345678", idString(.{ .rung = .fat_serial, .key = 0x12345678 }, &buf)); try testing.expectEqualStrings("fat-12345678", idString(.{ .rung = .fat_serial, .key = 0x12345678 }, &buf));
try testing.expectEqualStrings("exfat-da7a0001", idString(.{ .rung = .exfat_serial, .key = 0xDA7A0001 }, &buf));
try testing.expectEqualStrings("mbr-deadbeef-1", idString(.{ .rung = .mbr_index, .key = (@as(u128, 0xDEADBEEF) << 8) | 1 }, &buf)); try testing.expectEqualStrings("mbr-deadbeef-1", idString(.{ .rung = .mbr_index, .key = (@as(u128, 0xDEADBEEF) << 8) | 1 }, &buf));
const guid: u128 = 0x00112233445566778899AABBCCDDEEFF; const guid: u128 = 0x00112233445566778899AABBCCDDEEFF;
try testing.expectEqualStrings("gpt-00112233445566778899aabbccddeeff", idString(.{ .rung = .gpt_guid, .key = guid }, &buf)); try testing.expectEqualStrings("gpt-00112233445566778899aabbccddeeff", idString(.{ .rung = .gpt_guid, .key = guid }, &buf));
+57
View File
@@ -796,6 +796,42 @@ CASES = [
"expect": r"(?s)fat: mounted /volumes/fat-12345678" "expect": r"(?s)fat: mounted /volumes/fat-12345678"
r"[\s\S]*volume-manager: storage for volume \d+ removed; unmounting", r"[\s\S]*volume-manager: storage for volume \d+ removed; unmounting",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# S5 medium_changed: the SECOND removal trigger. QMP-eject the MEDIUM (the
# block backend, not the device) — the usb-storage device stays in the tree,
# but its TEST UNIT READY poll reports not-ready and publishes medium_changed
# (absent). The volume manager, now a subscriber, runs the same unmount path as
# a device pull. Discrimination: before S5 the manager never subscribed, so the
# event reached no one and the mount persisted (device-presence polling cannot
# see a medium leave while the device stays). One lifecycle, two triggers.
{"name": "volume-medium-change",
"build_case": "fat-mount",
"smp": 4,
"timeout": 150,
"qmp_sequence": [
{"delay": 8, "command": "eject", "arguments": {"device": "bootusb", "force": True}},
],
"expect": r"(?s)fat: mounted /volumes/fat-12345678"
r"[\s\S]*usb-storage: medium absent"
r"[\s\S]*volume-manager: medium left a storage device; unmounting"
r"[\s\S]*volume-manager: storage for volume \d+ removed; unmounting",
"fail": r"DANOS-TEST-RESULT: FAIL"},
# S5 storage-driver-crash rebuild. The device manager (test-storage-restart
# mode) kills usb-storage once, ~2s in — after its volume mounted. The device
# stays in the tree, so device-presence polling alone would leave fat wedged on
# the dead channel; the volume manager's channel-liveness probe (a geometry()
# that fails on the dead endpoint) must notice, reap the volume, and rebuild on
# the restarted driver's fresh channel — a SECOND mount of the same id-path.
# Discrimination: a pre-S5 manager checks only isDevicePresent (still true), so
# it never reaps and the second mount never appears (it would restart fat on
# the stale channel and crash-loop).
{"name": "volume-driver-restart",
"build_case": "volume-driver-restart",
"smp": 4,
"timeout": 150,
"expect": r"(?s)fat: mounted /volumes/fat-12345678"
r"[\s\S]*volume-manager: storage for volume \d+ removed; unmounting"
r"[\s\S]*fat: mounted /volumes/fat-12345678",
"fail": r"failing repeatedly; giving up|\[FAIL\]|DANOS-TEST-RESULT: FAIL"},
# Volume-manager discovery + probe (V3a, docs/volume-manager-plan.md). Reuses # Volume-manager discovery + probe (V3a, docs/volume-manager-plan.md). Reuses
# the fat-mount kernel build (the default boot now spawns the volume manager # the fat-mount kernel build (the default boot now spawns the volume manager
# from init.csv). It acquires the mass-storage block channel through the # from init.csv). It acquires the mass-storage block channel through the
@@ -857,6 +893,23 @@ CASES = [
r"(?=.*fat: mounted /volumes/fat-da7a0001)" r"(?=.*fat: mounted /volumes/fat-da7a0001)"
r"(?=.*fat: mounted /volumes/fat-da7a0002)", r"(?=.*fat: mounted /volumes/fat-da7a0002)",
"fail": r"DANOS-TEST-RESULT: FAIL"}, "fail": r"DANOS-TEST-RESULT: FAIL"},
# S4 second engine: a bare exFAT data device (serial e0fa0001) attached beside
# the FAT boot volume. The volume manager content-routes it to the exFAT
# service (not fat), which mounts it at its id-path /volumes/exfat-e0fa0001;
# the exfat-test client then reads the seeded HELLO.TXT and mutates through the
# mount (mkdir/write/rename/read/remove). Proves the second engine reuses the
# shared harness end to end. Fails against pre-S4 (no exfat binary, csv row, or
# VBR recognizer — the device would go to fat, which rejects the exFAT VBR).
{"name": "exfat-volume",
"build_case": "exfat-volume",
"smp": 4,
"timeout": 150,
"data_volume": {"exfat": True, "serial": "E0FA0001", "size_mib": 48},
"expect": r"(?s)(?=.*volume 0x0*e0fa0001 -> /system/services/exfat )"
r"(?=.*exfat: mounted /volumes/exfat-e0fa0001)"
r"(?=.*exfat-test: read HELLO.TXT ok)"
r"(?=.*exfat-test: ok)",
"fail": r"exfat-test: FAILED|DANOS-TEST-RESULT: FAIL"},
# Phase 2b: mkdir/unlink through the mount. Reuses the fat-mount build — the # Phase 2b: mkdir/unlink through the mount. Reuses the fat-mount build — the
# fat-test client, after listing, makes a directory, writes+reads a file inside # fat-test client, after listing, makes a directory, writes+reads a file inside
# it, then removes the file, exercising the whole VFS -> fat mutation path. # it, then removes the file, exercising the whole VFS -> fat mutation path.
@@ -1477,6 +1530,10 @@ def run_case(arch, case):
gen = [sys.executable, os.path.join(REPO, "tools", "make-partitioned-image.py"), data_img] gen = [sys.executable, os.path.join(REPO, "tools", "make-partitioned-image.py"), data_img]
for part in dv["partitions"]: for part in dv["partitions"]:
gen += [part["serial"], str(part.get("size_mib", 40))] gen += [part["serial"], str(part.get("size_mib", 40))]
elif dv.get("exfat"):
# One device, a bare exFAT volume — the second engine's medium.
gen = [sys.executable, os.path.join(REPO, "tools", "make-exfat-image.py"),
"--serial", dv["serial"], data_img, str(dv.get("size_mib", 48))]
else: else:
# One device, one bare FAT volume. # One device, one bare FAT volume.
gen = [sys.executable, os.path.join(REPO, "tools", "make-fat-image.py"), gen = [sys.executable, os.path.join(REPO, "tools", "make-fat-image.py"),
+15
View File
@@ -0,0 +1,15 @@
//! The exfat-test test fixture as a binary package (docs/build-packages-plan.md):
//! this file names the binary and EXACTLY the modules its source imports —
//! build-support resolves each name from the domains this zon declares.
const std = @import("std");
const build_support = @import("build-support");
pub fn build(b: *std.Build) void {
const exe = build_support.userBinary(b, .{
.name = "exfat-test",
.root_source_file = b.path("exfat-test.zig"),
.imports = &.{ "file-system", "logging", "process", "time" },
});
b.installArtifact(exe);
}
@@ -0,0 +1,14 @@
.{
.name = .exfat_test,
.version = "0.0.0",
.fingerprint = 0x77b19e3f7ee43ce3, // Changing this has security and trust implications.
.minimum_zig_version = "0.16.0",
.dependencies = .{
// build-support supplies the shared recipe; kernel is implicit in
// every binary (the root shim + link script live there). The rest
// are exactly the homes of this binary's declared imports.
.@"build-support" = .{ .path = "../../../../build-support" },
.kernel = .{ .path = "../../../../library/kernel" },
},
.paths = .{""},
}
@@ -0,0 +1,88 @@
//! test/system/services/exfat-test — a client that proves the exFAT mount end to
//! end: it waits for the exfat server to mount the volume at /volumes/exfat-
//! e0fa0001, reads the seeded HELLO.TXT off it, and exercises mkdir / write /
//! read / rename / remove through the VFS (which routes the id-path to the exfat
//! backend). Shipped in the initial_ramdisk; the `exfat-volume` QEMU case spawns
//! it alongside init with an exFAT data device attached beside the FAT boot volume.
const std = @import("std");
const fs = @import("file-system");
const process = @import("process");
const time = @import("time");
const logging = @import("logging");
// The exFAT data device's content id-path — its VolumeSerialNumber is 0xE0FA0001
// (the `exfat-volume` case passes --serial E0FA0001 to make-exfat-image.py).
const mount = "/volumes/exfat-e0fa0001";
fn writeLine(comptime fmt: []const u8, arguments: anytype) void {
var line: [128]u8 = undefined;
_ = logging.write(std.fmt.bufPrint(&line, fmt, arguments) catch return);
}
pub fn main(init: process.Init) void {
_ = init;
// Wait for the exfat server to bring up the USB storage chain and mount.
var opened: ?fs.Directory = null;
var tries: u32 = 0;
while (opened == null and tries < 1400) : (tries += 1) {
opened = fs.openDirectory(mount);
if (opened == null) time.sleepMillis(50);
}
var dir = opened orelse {
_ = logging.write("exfat-test: " ++ mount ++ " never became available\n");
return;
};
var count: u32 = 0;
var entry: fs.Entry = .{};
while (dir.next(&entry)) {
writeLine("exfat-test: entry '{s}' size={d}\n", .{ entry.name(), entry.size });
count += 1;
if (count > 32) break;
}
dir.close();
// Read the seeded HELLO.TXT (make-exfat-image.py writes "exfat hello danos\n").
var read_ok = false;
if (fs.open(mount ++ "/HELLO.TXT", .{})) |opened_file| {
var file = opened_file;
var buf: [32]u8 = undefined;
const n = file.read(&buf) orelse 0;
file.close();
read_ok = std.mem.startsWith(u8, buf[0..n], "exfat hello danos");
}
if (read_ok) _ = logging.write("exfat-test: read HELLO.TXT ok\n");
// Mutation through the mount: mkdir, create + write, rename, read back, remove
// — proof the write path reaches the engine over a real device.
var mut_ok = false;
if (fs.makeDirectory(mount ++ "/TESTDIR")) {
var wrote = false;
if (fs.open(mount ++ "/TESTDIR/W.TXT", .{ .create = true, .truncate = true })) |created| {
var f = created;
wrote = (f.writeAll("exfat-mutation-ok") orelse 0) == "exfat-mutation-ok".len;
f.close();
}
const renamed = fs.rename(mount ++ "/TESTDIR/W.TXT", mount ++ "/TESTDIR/R.TXT");
var readback = false;
if (fs.open(mount ++ "/TESTDIR/R.TXT", .{})) |reopened| {
var f = reopened;
var buf: [32]u8 = undefined;
const got = f.read(&buf) orelse 0;
f.close();
readback = std.mem.eql(u8, buf[0..got], "exfat-mutation-ok");
}
const removed = fs.remove(mount ++ "/TESTDIR/R.TXT");
mut_ok = wrote and renamed and readback and removed;
}
if (mut_ok) _ = logging.write("exfat-test: mutations ok\n");
if (read_ok and mut_ok) {
while (true) {
_ = logging.write("exfat-test: ok\n");
time.sleepMillis(1000);
}
}
writeLine("exfat-test: FAILED (read={} mutations={})\n", .{ read_ok, mut_ok });
}
+307
View File
@@ -0,0 +1,307 @@
#!/usr/bin/env python3
"""Format a real exFAT image from scratch — the danos exFAT test volume.
Pure Python 3 standard library (no mkfs.exfat / mtools). It writes a valid exFAT
filesystem — a Main Boot Sector + its boot-region checksum + a backup region, the
32-bit FAT, an allocation bitmap, an up-case table (with its checksum), and a root
directory whose entry sets a real exFAT reader (and the danos exfat engine) mount
and walk. Mirrors tools/make-fat-image.py in spirit.
make-exfat-image.py [--serial <hex>] [--label <name>] <out.img> <size-MiB>
make-exfat-image.py --verify <out.img>
The image seeds one file, HELLO.TXT, so a mount can be proven by reading it.
"""
import struct
import sys
SECTOR = 512
UPCASE_UNITS = 256 # a-z -> A-Z, the rest identity; covers ASCII names
def align_up(value, to):
return (value + to - 1) // to * to
def rotr16(v):
return ((v >> 1) | (v << 15)) & 0xFFFF
def rotr32(v):
return ((v >> 1) | (v << 31)) & 0xFFFFFFFF
def boot_checksum(region):
"""32-bit rotate-right sum over the boot region, skipping VolumeFlags
(106,107) and PercentInUse (112) of the first sector."""
checksum = 0
for i, byte in enumerate(region):
if i in (106, 107, 112):
continue
checksum = (rotr32(checksum) + byte) & 0xFFFFFFFF
return checksum
def upcase_checksum(table_bytes):
checksum = 0
for byte in table_bytes:
checksum = (rotr32(checksum) + byte) & 0xFFFFFFFF
return checksum
def set_checksum(entries):
"""16-bit rotate-right sum over a directory-entry set, skipping its own two
checksum bytes (offset 2..3 of the first entry)."""
checksum = 0
for i, byte in enumerate(entries):
if i in (2, 3):
continue
checksum = (rotr16(checksum) + byte) & 0xFFFF
return checksum
def name_hash(upcased_units):
h = 0
for unit in upcased_units:
h = (rotr16(h) + (unit & 0xFF)) & 0xFFFF
h = (rotr16(h) + (unit >> 8)) & 0xFFFF
return h
def ascii_upper(unit):
return unit - ord("a") + ord("A") if ord("a") <= unit <= ord("z") else unit
def solve_geometry(total_sectors, spc):
"""Solve for cluster count / FAT length / heap offset that fit. The FAT sits
after the main + backup boot regions (24 sectors)."""
fat_offset = 24
fat_length = 1
while True:
heap_offset = align_up(fat_offset + fat_length, spc)
cluster_count = (total_sectors - heap_offset) // spc
needed = ((cluster_count + 2) * 4 + SECTOR - 1) // SECTOR
if needed <= fat_length:
return cluster_count, fat_offset, fat_length, heap_offset
fat_length = needed
class ExfatImage:
def __init__(self, size_mib, volume_id=0x1234ABCD, label="DANOS"):
self.total_sectors = size_mib * 1024 * 1024 // SECTOR
self.spc = 8 # 4 KiB clusters
self.volume_id = volume_id & 0xFFFFFFFF
self.label = label
self.cluster_count, self.fat_offset, self.fat_length, self.heap_offset = solve_geometry(self.total_sectors, self.spc)
if self.cluster_count < 16:
sys.exit(f"error: image too small for exFAT ({self.cluster_count} clusters)")
self.cluster_bytes = self.spc * SECTOR
# Layout: the allocation bitmap (as many clusters as it needs — one per
# 8*cluster_bytes clusters of the volume), then the up-case table, the root
# directory, and the seeded file. A single-cluster bitmap (small volumes,
# e.g. the 48 MiB fixture) puts root at cluster 4, as before.
self.bitmap_bytes = (self.cluster_count + 7) // 8
self.bitmap_clusters = (self.bitmap_bytes + self.cluster_bytes - 1) // self.cluster_bytes
self.bitmap_cluster = 2
self.upcase_cluster = self.bitmap_cluster + self.bitmap_clusters
self.root_cluster = self.upcase_cluster + 1
self.hello_cluster = self.root_cluster + 1
self.image = bytearray(self.total_sectors * SECTOR)
def cluster_offset(self, cluster):
return (self.heap_offset + (cluster - 2) * self.spc) * SECTOR
def set_fat(self, cluster, value):
struct.pack_into("<I", self.image, self.fat_offset * SECTOR + cluster * 4, value)
def mark_allocated(self, cluster):
bit = cluster - 2
pos = self.cluster_offset(2) + bit // 8
self.image[pos] |= 1 << (bit % 8)
def main_boot_sector(self):
sector = bytearray(SECTOR)
sector[0:3] = b"\xEB\x76\x90" # jump boot
sector[3:11] = b"EXFAT " # filesystem name
# 11..64 MustBeZero (already zero)
struct.pack_into("<Q", sector, 72, self.total_sectors) # volume length
struct.pack_into("<I", sector, 80, self.fat_offset) # fat offset
struct.pack_into("<I", sector, 84, self.fat_length) # fat length
struct.pack_into("<I", sector, 88, self.heap_offset) # cluster heap offset
struct.pack_into("<I", sector, 92, self.cluster_count) # cluster count
struct.pack_into("<I", sector, 96, self.root_cluster) # first cluster of root
struct.pack_into("<I", sector, 100, self.volume_id) # volume serial number
struct.pack_into("<H", sector, 104, 0x0100) # filesystem revision 1.0
sector[108] = 9 # bytes per sector shift (512)
sector[109] = self.spc.bit_length() - 1 # sectors per cluster shift
sector[110] = 1 # number of FATs
sector[111] = 0x80 # drive select
sector[112] = 0xFF # percent in use (unknown)
sector[510] = 0x55
sector[511] = 0xAA
return sector
def build(self):
# Main boot region (sectors 0..11): VBR, eight extended boot sectors, OEM
# parameters, reserved, then the checksum sector.
vbr = self.main_boot_sector()
self.image[0:SECTOR] = vbr
for s in range(1, 9): # extended boot sectors carry the 0xAA550000 signature
struct.pack_into("<I", self.image, s * SECTOR + 508, 0xAA550000)
# sectors 9 (OEM) and 10 (reserved) stay zero
region = bytes(self.image[0 : 11 * SECTOR])
checksum = boot_checksum(region)
for i in range(SECTOR // 4):
struct.pack_into("<I", self.image, 11 * SECTOR + i * 4, checksum)
# Backup boot region (sectors 12..23) is a copy of 0..11.
self.image[12 * SECTOR : 24 * SECTOR] = self.image[0 : 12 * SECTOR]
# FAT: reserved entries, then a single-cluster chain per metadata object,
# except the bitmap which spans self.bitmap_clusters (a real FAT chain).
self.set_fat(0, 0xFFFFFFF8)
self.set_fat(1, 0xFFFFFFFF)
used = []
for i in range(self.bitmap_clusters):
cluster = self.bitmap_cluster + i
self.set_fat(cluster, 0xFFFFFFFF if i == self.bitmap_clusters - 1 else cluster + 1)
used.append(cluster)
for cluster in (self.upcase_cluster, self.root_cluster, self.hello_cluster):
self.set_fat(cluster, 0xFFFFFFFF)
used.append(cluster)
# Allocation bitmap: every metadata/file cluster in use.
for cluster in used:
self.mark_allocated(cluster)
# Up-case table: 256 explicit units, a-z -> A-Z.
upcase = bytearray(UPCASE_UNITS * 2)
for i in range(UPCASE_UNITS):
struct.pack_into("<H", upcase, i * 2, ascii_upper(i))
off = self.cluster_offset(self.upcase_cluster)
self.image[off : off + len(upcase)] = upcase
table_checksum = upcase_checksum(upcase)
# Seed file HELLO.TXT (contiguous, one cluster).
content = b"exfat hello danos\n"
off = self.cluster_offset(self.hello_cluster)
self.image[off : off + len(content)] = content
# Root directory: bitmap, up-case, volume label, HELLO set.
root = self.cluster_offset(self.root_cluster)
# 0x81 Allocation Bitmap
struct.pack_into("<BBB", self.image, root, 0x81, 0, 0)
struct.pack_into("<I", self.image, root + 20, self.bitmap_cluster)
struct.pack_into("<Q", self.image, root + 24, self.bitmap_bytes)
# 0x82 Up-case Table
struct.pack_into("<B", self.image, root + 32, 0x82)
struct.pack_into("<I", self.image, root + 32 + 4, table_checksum)
struct.pack_into("<I", self.image, root + 32 + 20, self.upcase_cluster)
struct.pack_into("<Q", self.image, root + 32 + 24, UPCASE_UNITS * 2)
# 0x83 Volume Label
label_units = [ord(c) for c in self.label[:11]]
struct.pack_into("<BB", self.image, root + 64, 0x83, len(label_units))
for i, u in enumerate(label_units):
struct.pack_into("<H", self.image, root + 64 + 2 + i * 2, u)
# HELLO.TXT set: File (0x85) + Stream (0xC0) + Name (0xC1)
name = "HELLO.TXT"
self.write_file_set(root + 96, name, first_cluster=self.hello_cluster, length=len(content))
def write_file_set(self, offset, name, first_cluster, length):
entries = bytearray(32 * 3)
# File entry
entries[0] = 0x85
entries[1] = 2 # stream + one name entry
struct.pack_into("<H", entries, 4, 0x20) # attributes: archive
# Stream entry
entries[32 + 0] = 0xC0
entries[32 + 1] = 0x01 | 0x02 # allocation possible + no FAT chain (contiguous)
entries[32 + 3] = len(name)
upname = [ascii_upper(ord(c)) for c in name]
struct.pack_into("<H", entries, 32 + 4, name_hash(upname))
struct.pack_into("<Q", entries, 32 + 8, length) # valid data length
struct.pack_into("<I", entries, 32 + 20, first_cluster)
struct.pack_into("<Q", entries, 32 + 24, length) # data length
# File Name entry
entries[64 + 0] = 0xC1
for i, c in enumerate(name):
struct.pack_into("<H", entries, 64 + 2 + i * 2, ord(c))
struct.pack_into("<H", entries, 2, set_checksum(entries))
self.image[offset : offset + len(entries)] = entries
def serialize(self):
self.build()
return bytes(self.image)
def verify(path):
with open(path, "rb") as handle:
data = handle.read()
if len(data) < 512 or data[510] != 0x55 or data[511] != 0xAA:
sys.exit("verify: missing 0x55AA boot signature")
if data[3:11] != b"EXFAT ":
sys.exit("verify: not an exFAT boot sector")
if any(data[11:64]):
sys.exit("verify: MustBeZero region is not zero")
fat_offset = struct.unpack_from("<I", data, 80)[0]
heap_offset = struct.unpack_from("<I", data, 88)[0]
cluster_count = struct.unpack_from("<I", data, 92)[0]
root_cluster = struct.unpack_from("<I", data, 96)[0]
spc = 1 << data[109]
# Boot checksum sector 11 must match a fresh checksum over sectors 0..10.
expected = boot_checksum(data[0 : 11 * SECTOR])
got = struct.unpack_from("<I", data, 11 * SECTOR)[0]
if expected != got:
sys.exit(f"verify: boot checksum mismatch (0x{got:08X} != 0x{expected:08X})")
# Walk the root directory for the HELLO.TXT set and check its checksum.
root = (heap_offset + (root_cluster - 2) * spc) * SECTOR
found = False
for i in range(spc * SECTOR // 32):
entry = root + i * 32
if data[entry] == 0x00:
break
if data[entry] == 0x85:
secondary = data[entry + 1]
total = (secondary + 1) * 32
stored = struct.unpack_from("<H", data, entry + 2)[0]
if set_checksum(data[entry : entry + total]) != stored:
sys.exit("verify: a file set checksum is wrong")
found = True
if not found:
sys.exit("verify: no file set in the root directory")
print(f"make-exfat-image: {path} OK "
f"({cluster_count} clusters of {spc * SECTOR} bytes, fat@{fat_offset}, heap@{heap_offset})")
def main(argv):
if len(argv) == 3 and argv[1] == "--verify":
verify(argv[2])
return 0
argv = list(argv)
volume_id = 0x1234ABCD
label = "DANOS"
i = 1
while i < len(argv):
if argv[i] == "--serial" and i + 1 < len(argv):
volume_id = int(argv[i + 1], 16)
del argv[i : i + 2]
elif argv[i] == "--label" and i + 1 < len(argv):
label = argv[i + 1]
del argv[i : i + 2]
else:
i += 1
if len(argv) != 3:
sys.exit("usage: make-exfat-image.py [--serial <hex>] [--label <name>] <out.img> <size-MiB>\n"
" make-exfat-image.py --verify <out.img>")
out_path = argv[1]
size_mib = int(argv[2])
image = ExfatImage(size_mib, volume_id, label)
with open(out_path, "wb") as handle:
handle.write(image.serialize())
print(f"make-exfat-image: wrote {out_path} "
f"({size_mib} MiB exFAT, {image.cluster_count} clusters, serial 0x{image.volume_id:08X})")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv))