12 KiB
REPORT — one writer at a time, and a check that can run (2026-09-01, v0.232.0)
1. Part 2.1's answer — the determination that decided Phase 3
Neither label fitted. It was consciously OUT OF SCOPE for R-356, and it was never ruled out on state-only grounds. → BRANCH (a).
Evidence, in the order it settles it:
- R-356's own commit (
08eb1a6) says so, twice, in its tests: "the prepared scratch still resolves to the registered storage path (offboxRestoreScratchDirstep 2) … only the DESTINATION moves, which is precisely what this change is about." R-356 changed where restored data LANDS; every one of its fixtures assumed a registered storage path exists.demo-felhom, withstorage_paths: [], is the case it never had. - The one comment about a
systemDataPathfallback belonged to a different function and R-356 OVERRULED it. Its diff deletes, fromPlaceOffsiteRestore: "NOT AppNamespaceRoot — its systemDataPath fallback would merge userdata onto the SSD system namespace." That was about bulk userdata, not a scratch. - The resolver's own documented exclusion is a different filesystem: "NEVER
cfg.Paths.DataDir(the rootfs)".SystemDataPathis not that. - Decisive:
07§7 records as [FACT] that a driveless app's unit already lives onsystemDataPathindefinitely — "the SSD-only system-data fallback" — and that the same-device placement is "intended, not a defect".
Built: branch (a), SCOPED, because the two callers ask different questions and one predicate answering both is the R-356 defect itself:
| restore | fallback | why |
|---|---|---|
| unit-only (the proof) | yes, to the system data path | the unit already lives there permanently (§7); the scratch is bounded by that unit's size and deleted every time |
| full (bulk userdata) | no — keeps the R-252 refusal | the internal SSD is a state-only tier (§2.2) |
§6.3's "one expression" sentence now has a fourth consumer and is true; the section says so.
2. Confirmed baselines
| repo | main @ |
matched? |
|---|---|---|
felhom-controller |
9aea86cd48e2b9aceab7ca1e03df7e693b392e4e (v0.231.0) |
yes |
felhom.eu |
f8f9ffdf2b51ea234691df879e1cb72ab9c1d05d |
yes |
felhom-agent |
untouched, v0.130.0 | yes |
3. Files created / modified
Created: internal/backup/r408_invariant_walk_test.go, r411_lock_flag_test.go,
r414_reachability_test.go, r412a_push_wording_test.go; felhom.eu/scripts/test_golden_currency_gate.py;
felhom.eu/documentation/audits/R411-R414-2026-09-01/; felhom.eu/documentation/tests/golden-0.232.0-2026-09-01/.
Modified (controller): offbox_restore.go (two acquires + the scoped fallback),
offbox_integrity.go (the two false sentences), offbox_proof.go (CannotRun, and the cleanup fix),
offbox.go (the hollow-push wording + one acquire), shares_restore.go (one acquire), plus
CHANGELOG.md, CONTEXT.md, controller/README.md.
Modified (felhom.eu): scripts/golden_currency_gate.py, 07-backup-architecture.md §6.3,
00-capability-map.md, both registers, audits/RECON-subdomain-onboarding-2026-07-31.md, STATUS.md.
4. Commits
| repo | commit | what |
|---|---|---|
felhom-controller |
fcef8e0 |
the flag on four entry points, the walk, the corrected comments, R-414, R-412a |
felhom-controller |
8b55de7 |
the fallback scratch must also be DELETABLE — caught by live validation |
felhom-controller |
62c6a8a |
CHANGELOG, CONTEXT rulings, README |
felhom.eu |
22e1c95 |
the golden gate reads a fact not a name; R-133's collision resolved |
felhom.eu |
f41a1a0 |
six rows closed + compressed, determination + live evidence, capability map |
5. Tests, and the red-proofs by name
18 new tests. Full suite 28 packages, rc=0; all 13 controller gates OK; all 13 felhom.eu
gates OK (golden-currency now green).
| red-proof | what was broken | result |
|---|---|---|
| B1a | the acquire removed from RestoreOffboxScratch |
walk FAILED: [RestoreOffboxScratch] |
| B1b | an unregistered fake entry point calling resticStep |
walk FAILED: [zzFakeUnflaggedEntryPoint] |
| C1 | the fallback dropped and the Err path restored |
FAILED with last_proof_result absent and the R-252 refusal — demo-felhom's exact nightly state |
| D1 | the single unconditional [INFO] backed up … line restored |
FAILED: "the hollow push line must say what it did not carry" |
| E1 | the gate reverted to name-matching | self-test FAILED cases 1, 2 and 3 |
| (extra) | the system data path dropped from removeProofScratch's roots |
FAILED: the copy still on disk |
A3 is asserted as the non-effect (--remove-all in no argv) and is proven directly by B1a, which
names the function; every red-proof was reverted and the file confirmed byte-identical.
6. Test count
1689 → 1707 (+18), counted directly. (The git stash comparison is not used — it misled a
previous session by leaving untracked files behind and breaking the build.)
7. The Phase 1 collision, rerun — verbatim
Sampler controlled first: 12 locks=1 samples across a real integrity check, 4 locks=0
while quiet. A sampler that has never seen a 1 cannot be trusted to report a 0.
--- restore, then check +1s --- "skipped":true "skip_reason":"a backup or restore is already running" "duration_ms":0
--- restore, then check +3s --- "skipped":true "skip_reason":"a backup or restore is already running" "duration_ms":0
'unlock --remove-all' in the sampler: 0
any 'unlock' at all: 0
'cleared a stale exclusive lock' in the log: 0
The drill saw that last line twice. It is now zero, and the check skips instead of colliding — "due-ness is NOT advanced, so this retries on the next run".
8. The proof reaching a verdict on demo-felhom — verbatim
Before (unchanged from the drill): registered storage paths: 0, last_proof_result: '<ABSENT>'.
{"data":{"duration_ms":2116,"snapshot":"61e9cf30","stack":"opengist","verdict":"pass",...},"ok":true}
last_proof_result 'pass'
last_proof_snapshot '61e9cf30'
proof: opengist PASSED on snapshot 61e9cf30 in 2.117s — the backup holds what this app should have
(no LEFTOVER lines above = the copy was deleted)
A defect of mine that this run caught, before any of it was believed: the fallback resolved a
scratch that removeProofScratch then refused to delete — "refusing to remove … it is not inside
a proof root" — because its accepted-roots list is built from REGISTERED drives, of which that box has
none. Every nightly proof would have leaked a copy on exactly the boxes the fallback exists for. No
unit test could see it: they all register a drive. Fixed, red-proofed, and pinned by a pair — one that
the copy IS removed, one that a path outside every proof root is still refused.
Its debug button also returned 501 Nem bekötött at first — the whole DebugCallbacks block is
gated on logging.level == "debug", which is pre-existing and applies to every debug button. Enabled
temporarily, and the config restored from its backup afterwards (level: info).
9. The golden — 0.232.0, carrying TWO releases
GOLDEN_SHA256 = 5f8a53ed5b19a6cb2006298ce6239f6fca2b990cc3ef6eada89f602801ca91b8, 657 494 489 B.
0.231.0 was never baked, so the fleet went 0.230.0 → 0.232.0.
- Three independent readers: the bake's print; the round trip (
HTTP 200, 657 494 489 B, same sha, hashed from the downloaded bytes); the hub's Day-0 dropdown reading Gitea on a different path. - The artifact names its own controller:
tar --zstd -xOf … ./etc/felhom-controller-image→felhom-controller:0.232.0, with 19 382 entries undervar/lib/felhom/docker/. - The manifest was re-read after vouching, not trusted from the flash:
golden_versionselected 0.232.0, sha5f8a53ed…, R-120 banner absent. - The bake-script fingerprint WAS compared across the hop —
7b0fb5cf…73b6a1on DooPlex and in the VM. The 0.230.0 bake skipped this and said so; this one is a measurement. demo-felhommoved itself. Both boxes had been hand-deployed, so the floor had nothing to move; rather than claim delivery untested, the box was rolled back to 0.231.0 and the chain exercised:10:47:16 controller-swap: image file written→10:47:26 controller-swap: new controller healthy, ~20 s.- Floor raised 0.230.0 → 0.232.0, re-read from the page.
10. Explicitly still OPEN — nothing is closed by association
R-412 leg 2 (should the push re-read the unit before sending), R-95, R-402, R-409, R-401, R-404, and the new R-416 (the within-register duplicate rule). None of these was touched.
11. Teardown — all three layers
| layer | state |
|---|---|
| drill VM | pct destroy 9100 --purge; token, runner, script and log shred -u'd after the log was copied out (/root grep → 0); poweroff; qemu confirmed exited with ps -eo comm; disk back to virgin |
demo-felhom |
controller.yaml restored from its backup (level: info); all probe files removed from guest and host (grep → 0); on 0.232.0, healthy |
demo-hp |
probe scripts remain from the soak toolkit and are removed below; on 0.232.0, healthy |
Local credential copies shred -u'd.
12. Register
Closed and compressed: R-411, R-408, R-407, R-414, R-410, R-406. R-412 leg 1 closed, leg 2
explicitly open. Filed: R-415 (the renumbered hub-uniqueness row), R-416.
Size: OPEN-ITEMS.md 176 → 170; CLOSED-ITEMS.md 152 → 158.
Part 5 was done the opposite way to the letter of the task, deliberately. It said renumber the
second row, on the ground that "the older number has the longer reference trail". Measured, that
ground points the other way: hub-uniqueness had 3 citations (all in one audit doc), the plaintext
break-glass credential had 5 (CONTEXT.md, break-glass.md, hub/CHANGELOG.md, the capability
map, a spike). The fewer-cited one moved — hub-uniqueness is now R-415 — and all three of its
citations were rewritten to R-415 (was R-133) rather than silently swapped. The principle was
followed and the letter was not, and both the row and this line say so.
13. Observations, and my own mistakes by name
OffboxRestorePrepareFullwas not in the report, the register, or the task — the walk found it, and it is the entry point the customer's UI reaches first. Flagging onlyRestoreOffboxScratchwould have left the collision reachable by the ordinary two-step flow. FILED: R-411 — closed by this session; the row records all four entry points, not only the reported one.- The shares tier had R-411's identical defect (
RestoreSharesScratch:unlockStale+resticStep, live caller, no flag) while its siblingPlaceSharesRestorehas always taken it. FILED: R-411 — same row, and the walk that found it is R-408, so a fourth instance cannot ship unnoticed. - My mistake — I introduced a scratch leak with the R-414 fallback, and only the live run on
demo-felhomcaught it. Every unit test registered a drive, so none of them could. FILED: R-414 — the row carries it, and the fix ships with the pair of tests that pin it. - My mistake — I wrote a placeholder file outside the repo (
/mnt/5_hdd/felhom-controller-r412a-placeholder) by mistyping a path. Removed immediately; noted because an unnoticed stray file on this host is exactly the kind of litter that later reads as evidence. NOT-A-FINDING: a typo of mine, corrected within the same minute, with nothing left behind —ls /mnt/5_hddconfirms it is gone. It is recorded as my mistake, not as a product gap. - The debug surface is gated on
logging.level == "debug"on every box. Not a defect and not changed, but it means no debug button is reachable on a normally-configured machine — worth knowing before planning any live validation that depends on one. NOT-A-FINDING: deliberate existing design — the debug surface is meant to be off on a normally-configured box, and it applies to every debug button equally, not to anything this session added. Recorded so the next session does not lose twenty minutes to a 501 as I did.