2d64ee7241a319e32c2e74af48559bc82b104689
436 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2d64ee7241 |
Campaign 10: OPEN observation — backup_target_* pair went silent under rapid cycling
Three I1/I1-pair violations in ~5 minutes, all "expected event absent". Recorded as an OPEN observation, NOT a finding: the system was mid-abuse when it was seen, and a verdict taken on a system being hammered is worth little. Established: it is not hub-side suppression and not a truncated log. The hub pod has 0 restarts over 43h and the controller's own log matches it line for line, so the events were never emitted. It is specific to the backup_target_* pair - the generic storage_disconnected/reconnected pair for the other drive kept firing normally throughout the same window. Also sampled, and the more serious half if it survives quiescence: mentes reads bound_under_parent=False while the backup-target state simultaneously reports degraded=false. Those cannot both be right - a false healthy on the backup target is I5/I6's failure mode. NOT established: whether the pair recovers once cycling stops (the harness detaches every ~2 min; a customer does not), whether the 02:25:37 controller restart is implicated, and whether the degraded=false sample was transient. Disposition written into the doc: after the run ends, quiesce with both drives attached, then do ONE slow detach/reattach and see whether the pair fires. That distinguishes "does not survive rapid cycling" from "the target alarm has silently stopped working", which would be severe. |
||
|
|
3d4c5365c1 |
Campaign 10: correct R-157 — the failure is INTERMITTENT (3 of 6), not deterministic
The first write-up said R-157 reproduced "at the same cycle in both runs - deterministic, not a coincidence". Wrong. The cycle numbers matched only because the runner's RNG is seeded so both runs drew the same permutation. The failure itself is a coin flip: run 2b's four hard resets went PASS(c2), FAIL(c10), PASS(c18), FAIL(c26); run 2a went PASS(c2), FAIL(c10). Three failures in six. The correction matters because it changes what kind of bug this is, and it strengthens rather than weakens the root cause: intermittency is exactly what a race against container-state settling predicts, whereas a wrong predicate would fail every time. Signature is identical on all three occurrences: rallly Exited 255 with rallly-postgres healthy, bootrecon reporting "no boot-orphaned apps" about 5s after controller start, and the container count still churning after the sweep (third occurrence 01:05: refresh 8, bootrecon 01:05:13, then 8 -> 7 -> 8). |
||
|
|
7f6b00375b |
Campaign 10: R-157 — bootrecon's start-once sweep misses the boot orphan it exists to recover
Reproduced twice, two independent runs, same cycle (the runner's RNG is seeded so both drew the same permutation - deterministic, not coincidence). A hard reset mid-backup brought everything back except the app half of the DB-backed stack: rallly left Exited 255, oom=false, restarts=0, its own log ending "Ready" - it died healthy - while rallly-postgres returned healthy. 20:28:13 Status refresh: 8 containers across 55 stacks <-- docker ps -a shows NINE 20:28:18 [bootrecon] Boot reconciliation: no boot-orphaned apps 20:28:25 Status refresh: 7 ... 8 containers <-- still churning AFTER the sweep 20:39:14 [deadapp] 20 scans, 5 deployed evaluated, 1 currently down The predicate is sound: once settled the controller reports rallly state=degraded containers=2, and IsDownState includes StateDegraded, so len>0 && IsDownState holds. The SNAPSHOT was wrong. bootrecon fires as a goroutine ~5s after start while docker is still restoring containers, and is start-once by design, so it never re-checks. Consequence: the app stays down indefinitely. Detection is perfect and recovery never happens - R-52's original shape, an alarm with no recovery. Not fixed. Distinguished from this campaign's two earlier HARNESS defects: both drives bound, every other app returned incl. the drive-backed one, only the app half of a two-container stack missing while its DB is healthy, and it surfaced through the fixed check written for exactly this. |
||
|
|
80db2c103a |
Campaign 10: full write-up of the run-2a harness defects
The previous commit message was truncated by an unescaped paren in the shell, so the fix detail and the product observations were lost from the record. This adds them as evidence, where they belong. Covers: the cc_proof table showing no C010-A row at all (the seed never landed); both harness defects; why an ambiguous I7 justified stopping a 10-cycle run; the red-proofed controls; and two transient product observations recorded but NOT filed as findings - the health probe naming the DB container on the app's port for about 70s during recovery, and a ValidateDump WARN on a dump taken while the app was down. |
||
|
|
9ca57e591b |
Campaign 10: two run-2a violations were HARNESS defects, not product defects — fixed
Run 2a hit its first two violations at cycle 10 and BOTH trace to my harness, not the product. Recorded in full because a check that fails for the wrong reason is as corrosive as one that passes for the wrong reason. HARD-RESET VM returned=True canaries_intact=False I7 want=C10-C010-A-194530 got=C10-C009-A-192929 restore_ok=True Root cause, evidenced: the cc_proof table's highest row is C10-C009-A — there is NO C010-A row at all, so the seed never landed. The hard-reset atom ran earlier in the same cycle and left rallly Exited(255); atom_restore_verify called seed() and never checked its return value, so an unwritten generation became a fake stale |
||
|
|
ac6c05bd7b |
Campaign 10: add monotonic-growth sampling — the half the invariants cannot see
I1-I11 are CORRECTNESS invariants: they answer 'is the system telling the truth this cycle'. All 586 of them passed in run 1 while nothing at all watched whether disk usage, snapshot count, log volume, fd count or RSS climbs. Accumulation is exactly what depth was for, and it was missing from the invariant list. Adds c10growth.py (Campaign 2's controller_rss.tsv precedent, widened to 19 metrics) sampling every 90s as a SEPARATE process, so the in-flight run 2 did not have to be restarted. Attributes every sample to a cycle by reading the runner's status.txt, and records NA rather than dying when the box is down during a hard-reset or reboot atom. c10growth_report.py turns it into Campaign 2's table shape (start/end/min/max/ slope-per-cycle) and splits verdicts by class: growth in RSS/fd/volumes/images/ restarts is a LEAK; growth in backup storage or the qcow2 is expected accumulation, reported with a projection to cycle 45. Caught a bug in the sampler itself on the first analysis: MENTES_USED_MB appeared to jump 623 -> 5667 MB, which is exactly ROOT_USED_MB — when a drive is detached, /mnt/<name> reverts to a plain directory on root and df silently reports the ROOT filesystem. The same class of error as the agent's exactMount check, in the measurement code. Gated on mountpoint and red-proofed both ways: a real mount returns a number, a non-mount returns NA. |
||
|
|
816c59c43a |
Campaign 10: R-117 Q7 (fs aborted in place, device present) proven PASS; extended atom set
The case R-117's spike called the worse half — a drive dying with no detach/return cycle, which before agent v0.117.0 emitted nothing on any channel indefinitely. Box runs 0.119.0. Aborted ext4 in place (abort,emergency_ro; device still present): bound_under_parent went false, storage_disconnected fired, the storage page named the stopped app, and calibre-web (whose library binds that drive) was STOPPED rather than restarted onto the dead namespace. Recovery needed a full device close, not a remount — exactly as the fix intends (BindAborted => quiet no-op). Runner extended with the 7 atom families run 1 skipped: abort-fs-in-place, kill-agent-mid-backup, hard-reset-VM-mid-write, reboot-VM, concurrent backup+restore, concurrent backup+detach, fill-drive-near-full. Also fixes a run-1 flaw recorded in the audit: reboot was appended AFTER the shuffle so it never interleaved with a detach; heavy atoms are now permuted in with the rest. Run-1 evidence preserved as *-run1.* (cycle numbering restarts per run). |
||
|
|
69f896d3cd |
Campaign 10 Phase B: 27 cycles, 586 invariant checks, 0 violations
Ran the soak on the Phase A rig. Ended on its own deadline — no watchdog halt, no atom exception, no I11 breach. I1 28+28 pairs, I2 28+28 pairs, I3 56, I4 56, I5/I6 28 each, I7 28, I10 135, I11 28. Zero violations. The row counts are themselves the no-silent-skip check: I3/I4 twice per cycle (both drives), I10 = 5 secret-class fields x 27, REBOOT on cycles 7/14/21 only. I7 is the headline: 28 restores, 28 correct discriminators — never stale, never empty. RTO (Tier 1, rallly, 66 MB): min 38.8s, median 42.0s, p90 42.5s, max 44.3s. That is the S band's lower end ONLY; the 5.5s spread over 28 runs says fixed work dominates, so nothing extrapolates to M or L. RPO not measured. Every atom and invariant was proven BY HAND before automation — the runner asserts nothing that was not first observed live. Caught a Phase A gap before starting: no app had HDD_PATH, so all data sat on the system disk and I3 could never have fired. Deployed calibre-web onto adatok first; otherwise the run would have produced 27 green cycles that tested nothing cross-drive. Investigated and DISPROVED a suspected defect (audit 5.2): /api/disks reports state=attached for a physically absent drive, and intermediary.go:230 really does compute presence from State=="attached". It is inert — planDriveGates only gates paths under /mnt/felhom-drives/ and uses BoundUnderParent there, which was correctly false. The gate fired; the storage page showed "Meghajtó leválasztva". No R-n minted. Honest gaps: 6 of ~12 atom families ran. Not run — Tier 3 (structurally un-isolatable), abort-fs-in-place, kill-agent-mid-backup, hard-reset-mid-write, reboot-VM, both concurrency atoms, fill-drive-near-full. I8 not checked, I9 not automated (cited from the tester-gate run, not re-claimed). kill_controller is NOT mid-backup and reboot_guest never interleaved with a detach. 27 cycles does not answer the brief's question about drift at the thirty-eighth. Teardown still OWED, including hub customer c10-soak (disposition: DELETE). |
||
|
|
4691aa1a35 |
Campaign 10: Phase A complete + gated; Phase B not run; R-156 filed
Phase A passed every gate on a fresh box built from the PUBLISHED ISO 1.26.1: install, claim, two drives enrolled through the real endpoints with the backup target healthy, four apps spanning both sides of D5's secret split, and a working discriminator across all four. Isolation gate: both denials captured, each with a positive control. The PBS control FAILED first — four clean-looking 403s were worthless because the token was denied on its own datastore too (PBS token privilege separation). Fixed and re-run; the denials stand. R-156 (new, register grepped): papra's data is neither persisted nor backed up, and it reports healthy. The template mounts papra_data:/app/data; the app writes /app/app-data/db/db.sqlite. Volume empty and root-owned against a -rootless image, real DB in the container writable layer, healthcheck only probes the HTTP port. Its Tier-1/2 backup is real, verifiable and contains nothing. Not fixed. Tier 3 could not be isolated so it was not run: offsite hard-requires the DR tier (configs.go:1300) and the DR tier only provisions on ep0 (per-endpoint allocation deferred, hub/README.md:260). Both are recorded deliberate positions, so no R-n minted. The campaign touched neither ep0 nor the Storage Box. Phase B did not start. Phase A was budgeted at ~1h and took ~5.5h (1.26.1 is a public release image with no auto-install path, so the install was a blind screendump+sendkey walk). That left the runner — which judges eleven invariants and fires destructive atoms unattended — to be written at 04:00 with ~3h of night left. Stopped on the brief's own fence: a rig producing false negatives is worse than no rig. The rig is built and idle; teardown is OWED and itemised, including hub customer c10-soak (disposition: DELETE). |
||
|
|
e9a74a0019 |
docs: remove a gate criterion that could never pass, and close three register rows
PART 1 — the release gate.
G7 required the packaged .deb to sha256-match the one built from committed source. That is
unsatisfiable BY CONSTRUCTION: dpkg-deb stamps the build time into every archive, so two builds of
byte-identical source differ. It was already failing when the 1.26.1 release ran it. A criterion
nobody can satisfy gets waived once and read as advisory ever after — which is how R-29's shelf of
never-run gates was built. Sub-clause dropped, reason recorded in G7's own note the way G6's
amendment was, so a future reader can restore it if SOURCE_DATE_EPOCH ever makes it meaningful.
RULING ASKED FOR — is payload integrity covered by G9 alone? NO, and G9 is widened rather than a new
criterion invented. The package ships TWO payload files (build-deb.sh:54-55); G9 checked only the
script. The systemd UNIT was covered by nothing: G7 covered the container, G8 covers the postinst
behaviourally, G13 covers directory presence. The unit is not incidental — its After=, its
ConditionPathExists= and its Restart= decide WHEN AND WHETHER day-0 runs at all, so a drifted unit
would have shipped silently. Same shape as the /etc/felhom miss that G13 exists to prevent: a check
that proved the thing present and said nothing about what it depended on. The check passes today.
G13 moved to sit after G12 — it was minted late and left between G10 and G11.
PART 2 — register dispositions. BASELINE DISCREPANCY, reported rather than worked around: only R-128
had a row. R-154 and R-155 had NO row in either file — minted in a spike document and never carried
across, which is R-123's class, not the drift the task described. Rows created, closed, with the
reasoning, because in all three cases the reasoning is the durable part:
R-128 closed by CORRECTING a false claim, not by making the assertion real — the coupling does not
exist and asserting it would invent a constraint. Flagged so nobody 'restores' it.
R-154 closed with the measurement and where it now lives in pushed source.
R-155 NARROWED, not deleted — unchanged for FELHOM_MENU=single, inapplicable to release. Flagged so
the guard is not later removed wholesale on the strength of 'R-155 closed it'.
Documentation only: no code, no build, no ISO, no upload, no box touched.
|
||
|
|
f2fc76ec4b |
ISO v1.26.1 PUBLISHED — both entries proven, round trip verified
Live: https://iso.felhom.eu/felhom-installer-1.26.1-pve9.2-1.iso sha256 f3cc86d5f0ec68bba4155c994b4fa84e208d50209bb6e815636c99e5441059a6, 1705322496 bytes. PART 5 PASSED ON BOTH MENU ENTRIES, four observables each: Graphical spikegfx.felhom.eu pairing code J7N-2DA TerminalUI spikesix.felhom.eu pairing code ZY5-YY4 Both: manual install, own disk, own password, real completion signal, and the journal's 'not bound yet — polling every 30s ... normal waiting state, not an error'. Spike 4 had REASONED the graphical path follows from shared Install.pm; it is now measured. PART 6: G1-G10 + G13 all PASS against the uploaded file. G4's single hit is felhom-bootstrap.sh:480's substring TEST ('$envtext' != *FELHOM_RETRIEVAL_PASSPHRASE=*), not a value — my own regex matched the glob's asterisk. PART 7: uploaded via rclone in a container configured ENTIRELY by environment variables, so no credential file was ever written. Round trip verified from the public URL — not the local file. Bucket stays private: unauthenticated GET to the S3 endpoint 400, custom domain has no index (404). CORRECTED BEFORE UPLOAD: the generated manifest described a single automated entry with a 5s timeout and listed Graphical/Terminal UI as 'menu-removed'. Generator fixed, sidecar regenerated, and the ISO verified byte-identical before and after — the published file IS the file Part 5 validated. Hub-side cleared: appliances 16, 17, 18 discarded (303 each); zero rows remain. The endpoint is /appliances/<id>/discard, POST only (server.go:345) — not /delete. Teardown: VMs purged, spike5 storage removed, demo-hp back to 6.6G, drill-r50 and 9201 untouched. Still open and named: OPEN-ITEMS/ROADMAP dispositions for R-128/R-154/R-155 are not written; the .deb is not byte-reproducible (G7 sub-clause); before-network stub unreached; Secure Boot and real hardware not exercised. |
||
|
|
a967da7d2c |
iso 1.26.1: ship /etc/felhom/ — the directory the bootstrap writes its state into
FIX for the Part-5 failure. felhom-bootstrap.sh writes the appliance token (:431), the pairing code
(:435) and .bootstrap-done into /etc/felhom/. The old stub-first-boot.sh created it explicitly
('install -d -m 0755 /etc/felhom /usr/local/sbin'); packaging dropped the env FILE correctly and the
DIRECTORY with it. Measured consequence on a real interactive install: the box registered at the hub,
could not persist its token, and polled 'HTTP 401 — still retrying' forever with no claim code.
- build-deb.sh now ships ./etc/felhom/ (0755, empty) and ASSERTS it, plus ./usr/local/sbin/ and
./lib/systemd/system/, as G13. RED-PROOFED: removing the install -d makes the build exit 3 with
'is not in the package (G13)', and restoring it goes green.
- The gate gains G13 with the reasoning: G7/G8/G9 all passed on the broken package. G9 proves the
payload is the right payload and says NOTHING about what the payload depends on.
ISO_VERSION -> 1.26.1.
|
||
|
|
01a8155c5a |
iso v1.26.0: the PUBLIC release image — no answer file, interactive install, day-0 by .deb
Design inputs: SPIKE-universal-iso-{1,2,3,4}-2026-07-31.md. Every choice below is a measurement.
NEW: scripts/iso/pkg/ — the felhom-bootstrap .deb, built from committed source.
Two files only (script + unit), NOT three: felhom-bootstrap.sh:91 reads /etc/felhom/bootstrap.env
only 'if [[ -r ]]', and its defaults at :95-96 are EXACTLY what the pairing env set
(build-felhom-iso.sh:257-258) — so shipping it would add a 0600 file to a public package to express
values the script already defaults to. NO dependencies: the binaries it calls run at FIRST BOOT,
not at postinst time, so SPIKE 4's open 'dpkg --configure -a' ordering question does not arise.
The postinst is structurally incapable of failing (no 'set -e', every statement guarded, ends
'exit 0'); build-deb.sh self-asserts G8/G9 and REFUSES to emit a package that violates them.
iso-repack.sh — two changes, both narrowing rather than deleting:
- R-155 guard: now applies to FELHOM_MENU=single ONLY. It protected the single-entry mode's promise
(one button labelled 'install' must not drop into a disk-picker); a release image carries no
auto-installer-mode.toml BY DESIGN (gate G1), so refusing it would be the guard firing on the
shape it describes rather than the one it prevents.
- the menu collapse now has a release mode: two INTERACTIVE entries, Graphical default, timeout 15.
Entry-count and banned-token gates are per-mode; the six-token list is UNCHANGED for single mode.
- .deb injection into /proxmox/packages/, with a skip-list collision check (a colliding name would
be dropped silently — the inert-payload class) and a post-remaster assertion that it landed in
final.iso, not merely in the extract tree.
build-felhom-iso.sh — --release: no profile, no root hash, no answer.toml, no prepare-iso at all.
Skipping prepare-iso is what removes the Automated entry by construction, since the stock grub.cfg
emits it only inside 'if [ -f auto-installer-mode.toml ]'.
R-128 RULING — FIXED, by correcting the claim rather than inventing an assertion for it. The comment
said ISO_VERSION 'aligns with SCRIPT_VERSION'; nothing evaluated it and the two had drifted. The
coupling does not exist: the ISO is frozen, felhom-host-install.sh is fetched at run time from main
(R-94/R-110), so an assertion would invent a constraint. Comment corrected, ISO_VERSION -> 1.26.0.
Release gate G6 AMENDED before the build, with its reasoning recorded in the runbook: the six-token
ban existed to keep users away from the manual installer, which the ruling makes the product.
'proxtui' (the TUI installer we ship) and 'nomodeset' (its graphics fallback) are dropped for
release images; proxdebug/Rescue Boot/memtest/fwsetup stay banned in both modes.
|
||
|
|
e787391c0a |
docs: the public ISO release gate, written BEFORE the first release image
A standard defined in advance cannot be rationalised afterwards, and this is the artifact that most needs one: once a file is on iso.felhom.eu and someone has downloaded it, it cannot be recalled. Twelve criteria, each checkable against the UPLOADED FILE rather than the build inputs, and each carrying the spike measurement that justifies it: - G1 no answer.toml / auto-installer-mode.toml — deletes the whole Spike 1-2 problem space and removes the Automated menu entry by construction rather than by a guard - G2/G3/G4 no root hash, no SSH key, no customer identity — the shared-credential classes - G5 credential scan by ENUMERATION against the stock ISO, not a pattern sweep (Spike 1 found /answer.toml precisely because the earlier recon grepped the wrong file) - G6 menu present, interactive default, timeout >= 10 (Spike 2 lost a probe to a 1-second menu), underscore timeout_style, and the banned-token safety gate kept unchanged - G7/G8 the felhom .deb present, and a postinst that cannot fail: no systemctl start/daemon-reload (no systemd runs in the installer chroot), no network use (the cable may be out), no 'set -e', ends 'exit 0' - G9 felhom-bootstrap.sh byte-identical to repo HEAD — the one frozen, drift-capable payload - G10 every build input committed (Spike 1: demo-felhom came from an uncommitted profile) - G11 published checksum AND a verified download round trip - G12 bucket Public Access stays Disabled Committed on its own, before any build. |
||
|
|
61e9b55737 |
SPIKE 4: a .deb in the ISO DOES deliver on an interactive install
Findings only — no script, profile or build file changed; no release ISO built, nothing published. documentation/audits/SPIKE-universal-iso-4-2026-07-31.md MEASURED, with a control, and the negative control is in the SAME box. One ISO (15 GRUB entries), a trivial probe .deb injected into /proxmox/packages/, two qm-created VMs on demo-hp (400 interactive / 401 automated control) on a scratch dir storage at the /mnt/nvme-1tb mount ROOT. Interactive (Terminal UI) install: - package installed (ii felhom-spike4-probe 0.0.1) - postinst RAN (marker + content intact) - it enabled a systemd unit, and that unit FIRED ON FIRST BOOT (uptime 7.98s, pid1=systemd) - while on the same machine proxmox-first-boot is NOT installed and /var/lib/proxmox-first-boot does not exist — Spike 3's negative reproduced, not assumed. Postinst environment (identical both paths): pid1=unconfigured.sh, NO running systemd, but 'systemctl enable' SUCCEEDS; /proc+/sys mounted; network+DNS happened to be up (inherited from the installer's DHCP — must NOT be relied on). Constraints: never systemctl start/daemon-reload, never require network, never fail, do the real work in the unit at first boot. Repack preserves it, but a naive 'xorriso -boot_image any replay' fails with 'Overlapping MBR partition entries' — iso-repack.sh:270-292 already documents that exact failure and its fix. R-153 RETRACTED into R-94 leg (b): OPEN-ITEMS.md:15 carries it verbatim at READY (XS), and R-29 says explicitly 'do not mint a new ID for a new instance'. Spike 3's further claim that the drift leaves the generator 'three minor versions stale' was FALSE and is corrected — R-94 retracts that exact reading; the served script is always main, so 1.22.0 is what every install already gets. No new R-rows opened. |
||
|
|
bb29186d62 |
SPIKE 3: [first-boot] does NOT fire on an interactive install
Findings only — no script, profile or build file changed; no release ISO built, nothing published. documentation/audits/SPIKE-universal-iso-3-2026-07-31.md MEASURED with a control from the SAME image (one ISO, 15 GRUB entries): - Automated entry -> hook fires: ttyS0 marker, marker file, /var/lib/proxmox-first-boot/proxmox-first-boot (0700), activation symlink, unit active. - Terminal-UI entry, normal manual install -> ALL absent, and the proxmox-first-boot PACKAGE is not installed at all. A whole-filesystem grep for the marker returns nothing. Mechanism cited: Config.pm:118 defaults first_boot.enabled=0 and set_first_boot_opt is never called in the Perl tree; Install.pm:746 returns early without it; Install.pm:1360 skips the package. proxinstall (graphical) has ZERO occurrences of first-boot. [first-boot] is an automated-installer feature, unavailable on every interactive path by construction. R-154. A delivery mechanism DOES exist and is UNTESTED: Install.pm:1343-1372 unpacks every .deb in the ISO's /proxmox/packages/ into the target on every path (fixed skip-list), then dpkg --configure -a runs postinsts (:1378) — how PVE ships first-boot itself. Read from source, not measured. Q5: the public image should carry NO answer.toml at all — that removes the baked root hash, the disk profile and the whole Spike 1-2 problem space, and makes it a one-line release gate. But iso-repack.sh:100-106 refuses an ISO without auto-installer-mode.toml. R-155. Incidental R-153: hub hostInstallVersion=1.19.0 vs SCRIPT_VERSION=1.22.0; hostinstall_gates.py detects it and exits 1 — the gate works, nothing runs it. Q3 (real stub at before-network) was NOT reached and is recorded as not reached. |
||
|
|
19c932a693 |
SPIKE 2 complete: locked root closes the PVE web UI; before-network gives a measured zero window
Findings only — no script, profile or build file changed; no ISO built, nothing published.
documentation/audits/SPIKE-universal-iso-2-2026-07-31.md
Both Tier 0 boxes went offline mid-session (provider cable fault; four routes tried, no Tier 2
fallback used) and returned. All three scenarios then ran to completion on real PVE, each signalled
by reboot-mode='power-off' rather than a disk hash.
- A LOCKED ROOT CLOSES THE PVE WEB INTERFACE. Measured at the exact endpoint the UI uses
(POST /api2/json/access/ticket, root@pam) WITH A WORKING CONTROL: known-password install returns
HTTP 200 + ticket; locked install returns 401 for every password and none can exist.
passwd -S root = L, shadow = literal-asterisk, PVE uses the stock PAM stack.
- GRUB recovery mode is also closed ('the root account is locked') — but init=/bin/bash still gives
an unauthenticated root@(none):/#. A locked box is recoverable, operator-only, at the console.
The installed GRUB has NO password, so locking root is not a physical-security measure. R-152.
- before-network MEASURED (A/B, same image): the hook RUNS (marker, uptime 6.58s) with entropy 256,
writable /etc, all binaries and openssl_rand_len=32, while ip_global is EMPTY and
listen_22_8006 = 0. fully-up is the converse: sshd+pveproxy active, 3 listening. Zero window.
- R-148: answer.toml.tmpl:27 justifies fully-up with a pvesh/pct dependency the stub does not have
(grep rc=1) — it blocked the ordering now measured as the fix.
- R-149 three ordering values; R-150 Condition-guarded hooks skip silently; R-151 demo-felhom built
from an uncommitted profile.
Three probes failed and are recorded as failed: a container probe that ran as uid 0, a GRUB probe
that missed the 1-second menu timeout, and a kernel-line edit one line off (caught by a pre-typing
verification screendump). The interim 'Layer 1 teardown INCOMPLETE' is corrected — the fixture had
never landed, because the staging mkdir was in the SSH call that timed out.
|
||
|
|
5bdd8372f8 |
SPIKE 2: before-network gives a zero window by construction; locked root closes sulogin
Findings only — no script, profile or build file changed; no ISO built, nothing published. documentation/audits/SPIKE-universal-iso-2-2026-07-31.md BOTH Tier 0 boxes went offline mid-session (remote site, 12:28 CEST; four routes tried, our tailscale pod healthy). Q1/Q2/Q3 each keep a part needing a nested VM: those are BLOCKED, not answered. DooPlex was NOT used as a fallback — Tier 2, and this task did not authorise it. Established without them: - STRUCTURAL: ordering='before-network' maps to proxmox-first-boot-network-pre.service (Before=network-pre.target, Type=oneshot) — it completes before ANY interface is configured, so a rotation there has a zero-length window BY CONSTRUCTION, not by being fast. - R-148: the stub does not need 'fully-up'. stub-first-boot.sh has no pvesh/pct/pveum/qm call (grep rc=1); that usage is in felhom-bootstrap.sh under its own After=network-online unit. answer.toml.tmpl:27 justifies the current ordering with a dependency that does not exist. - R-149: the ordering enum has THREE values (before-network, network-online, fully-up), not two. - MECHANISM (container, not PVE): locked root closes sulogin — 'the root account is locked' for both '*' and '!', with a working control. So 'discard' and 'lock' are the SAME outcome for recovery, making the escrow decision binary. - R-150: all four proxmox-first-boot-* units are Condition-guarded; a failed condition is a SKIP, so a hook that never ran looks identical to one that succeeded. - R-151: demo-felhom was installed from an UNCOMMITTED profile — a Tier 0 reference box is not reproducible from main. - Q4: four gates in iso-repack.sh enforce the single-entry menu; default/timeout already settable. The first mechanism probe was invalid (uid 0 bypassed pam_unix; sulogin had no tty) and a teardown error (shredding the control plaintext) are both recorded as failures, not massaged. demo-hp teardown is INCOMPLETE and named as such; the command is recorded, not claimed done. |
||
|
|
ea00976403 |
SPIKE: a universal ISO needs a different disk strategy and a locked root
Findings only — no script, profile or build file changed; no ISO built, nothing published. documentation/audits/SPIKE-universal-iso-2026-07-31.md - R-139 (HIGH): a disk filter matching >1 device does NOT fail safe. Observed in a nested VM — the installer silently picked one of two matching disks and wiped it; validate-answer accepts such an answer. The 'filter did not match any devices' guard covers the ZERO-match case only. - No udev property distinguishes an internal system disk from external media. Measured on demo-felhom with its 1TB external attached: ID_BUS='ata' for BOTH, lsblk RM=0 for both, and device-info exposes no removability property. demo-hp's NVMe carries no ID_BUS/ID_TYPE at all. - R-141 (HIGH): the answer schema makes a root credential mandatory, but root-password-hashed='*' validates AND installs to completion. [first-boot].ordering accepts 'before-network', the only ordering that closes the exposure window structurally. - Q3: prepare-iso leaves grub.cfg byte-identical to stock (15 entries, automated AND interactive) — a two-entry menu is purely a Felhom grub.cfg.tmpl change. - R-129 resolved: demo-hp's key is the operator's own, added post-install; demo-felhom's IS baked by an uncommitted profile. The reachable-before-rotation measurement FAILED twice and is recorded as failed, not inferred. Opens R-139..R-147; restates R-128. |
||
|
|
5825ceeabf | docs: v0.86.0 copy-without-reveal + the break-glass credential leg is now proven (PVE ticket minted) | ||
|
|
9e079c7883 |
RECON: a Felhom-issued subdomain works in the product — the blocker is Cloudflare edge-cert depth
Question A: YES, no code change. customer.domain is a trimmed string with no UNIQUE, no CHECK, no format rule (store.go:114, configs.go:673), copied verbatim into controller.yaml (configgen.go:48), and every one of its 30 consumers on the box interpolates it without parsing. Zero hits for registrable/eTLD/publicsuffix across both repos. Nothing creates DNS records (zero hits for dns_records) — the two Cloudflare clients are WAF-only. And the zone-ownership assumption is a SWITCH, not a requirement: traefik.yml.tmpl selects DNS-01 when cf_api_token is set and HTTP-01 when it is empty. The real blocker is Cloudflare, proven live: the edge certificate covers exactly one wildcard level (SAN = demo-felhom.eu, *.demo-felhom.eu), so a two-label hostname — which a per-tester subdomain forces — gets "tls alert handshake failure" and no peer certificate at all. That makes Advanced Certificate Manager a prerequisite of the separate-domain plan, not an optional extra. Whether ACM is available on the account could not be established read-only: the only Cloudflare tokens in reach are the Zone:DNS:Edit tokens on the demo boxes, which the fence forbids using. Question C, measured rather than reasoned: r.Cookie returns the FIRST match and never tries the others (BOGUS+real = 302, real+BOGUS = 200), so a tossed cookie wins outright — DoS and confusion, not takeover, since it fails closed on mutations. CSRF is a single choke point (server.go:256) and the token carries the whole load against a same-registrable-domain attacker. But it is SKIPPED entirely when no session cookie is present, which with browser-cached Basic auth is cross-origin CSRF on every mutating route (R-135). Agreeing with the separate-domain recommendation, with the caveat the brief asked for: it is necessary but not sufficient. It does not solve Question D, because that is a shared-zone problem and the new domain is a shared zone. Filed R-133..R-138: duplicate domains accepted; hub/controller zone-resolvers disagree on depth; CSRF skipped on the no-cookie path; __Host- rename (one line, preconditions verified met); geo-WAF rules zone-scoped and non-namespaced (four cross-tenant faults, blocks shared-zone onboarding); shared-zone cf_api_token is a zone-wide DNS-write capability on a customer's box. Nothing created: no customer, DNS record, tunnel, route or code change. |
||
|
|
eb5d05f496 | docs: host-addresses audit + capability-map row + REPORT (agent 0.119.0 / hub 0.85.0) | ||
|
|
b4edc087fa |
Tester gate: golden re-baked to 0.188.0, fresh-install proof PASSED — a fresh box is safe to hand to a tester
§7.2 answer: YES. A real day-0 from the existing v1.25.0 ISO reached a claimable,
app-serving box in ~10 minutes unattended, and an app's data came back from the
drive with the guest's app.yaml gone — proven readable by the application over
its own TCP path, with a discriminator (PRE-BACKUP row = 1, POST-BACKUP row = 0).
Part 0: NO ISO rebuild needed, verified against the ISO on disk rather than from
source. It bakes only felhom-bootstrap.sh, its unit and the secret-free pairing
env (full-base64 match, 1 hit each) and 0 hits for any installer, controller or
golden marker. The installer is fetched at run time; the live URL is byte-identical
to repo HEAD (v1.22.0, six days newer than the ISO) and the fresh box ran it.
Part 1: baked 0.188.0 rather than the brief's 0.187.0 — 0.187.0 lacks D5, which
is the very claim Part 2 step 6 tests. Published (404 pre-gate with a 200 control;
anonymous download, 649310288 bytes, sha match), vouched, and consumed by a real
box. R-120's gate exercised BOTH ways: 0.185.1 refused with no write, 0.188.0
allowed — evaluated, not silently skipped.
Part 3: RUNBOOK-manual-build.md cited a "RECORDED" qemu line that is itself
labelled reconstructed and whose source says it was never saved. The real
invocation is now captured from this bake as §4.0, with the bake/publish/teardown
steps; the old entry is marked SUPERSEDED.
Teardown all three layers, hub disposition stated: VM destroyed, scratch storage
removed with space returned exactly, customer sess-g DELETED via full cascade.
sess-f deliberately left (R-131) with its command recorded.
Filed, none fixed: R-128 (false ISO_VERSION invariant comment), R-129 (demo-hp's
"no baked SSH key" is stale — key auth works), R-130 (HARD_MIN_LVM_GIB warns and
proceeds), R-131 (fourth orphaned scratch customer), R-132 (curl's %{redirect_url}
printed the hub operator password into a transcript — HUB_PW needs rotating).
|
||
|
|
1956e5d390 |
hub v0.84.0 — break-glass console credential on the host page
The credential existed and was not reachable when it was wanted. Every box has
had a strong random root@pam password since TASK G1, vaulted in the hub at day 0
and used for real during the sshd incident — but the only way to read it back was
a hand-written curl carrying the global operator key, a secret kept out-of-band.
In practice the PVE web console on a demo box felt locked.
The host page grows a Console access card: presence + username + set_at by
default, Reveal fetches the plaintext on demand for 60 s with a Copy button.
Masking clears the JS variable, and also fires on a second click and on
visibilitychange. A host with nothing vaulted says so, and says why.
The secret is NEVER rendered into the page, and that constraint shapes the
change. The render path uses a new store.GetHostRecoveryMeta whose struct and
SELECT both omit the secret column, so it is structurally incapable of carrying
one. The plaintext crosses the wire only in the response to POST
/hosts/{id}/reveal-recovery-credential (Cache-Control: no-store, CSRF-gated at
the ServeHTTP level; POST precisely so that gate applies and so no secret is
retrievable by URL alone). Deliberately NOT the customer page's data-secret
widget, which embeds the plaintext on every load.
A delivered reveal writes one recovery_credential_revealed event on the host's
customer timeline (info, source hub, Hungarian) via SaveEvent alone — no
dispatcher, nobody emailed, the log_tail_requested shape. Two reveals write two
events: the register records accesses, not states. A 404 is not an access. An
unbound host reveals fine and writes no event; the [INFO] hub line, carrying the
username and a length only, is then the record.
The global-key API path is untouched by design — it is the route for when the
hub UI itself is broken, and coupling it to the session layer would delete the
independence that makes it a fallback.
Recorded as a real trade: the hub session password alone now unlocks console root
fleet-wide, where retrieval previously also needed the global key. Accepted for a
single-operator, HU-geo-fenced hub that already stores these passwords in
plaintext at rest (CONTEXT.md ruling S-4). The plaintext-at-rest half is filed as
R-133 — every hub DB backup is a fleet-wide console-credential dump.
Tests 550 -> 559; four red-proofs (page leak, audit event, CSRF gate, route
order) each run, observed failing, and reverted. The route-order proof is a seam
test driving ServeHTTP: a handler-level test cannot see that defect, because the
handler is correct and simply never runs.
|
||
|
|
0a9bd3829d |
D5 SHIPPED: Tier-1/2 restore no longer depends on the whole-guest tier
Records controller v0.188.0 across the four coupled artifacts. 07-backup-architecture.md is the owning doc: - new 7.4 = the recovery chain AFTER D5 (7.1 leg 1 superseded; leg 2, the living-app dependency, explicitly unchanged so this is not read as more than it is) - 7.3 collapsed to history, with the correction that the target as written (data_key-only) was tested in Part 0 and rejected - 3 records that the two-lane split is now real, not just intended - matrix rows 3 / 3c (new) / 13; 10.1 D5 itself shipped Also: new capability-map row, D5 collapsed in ROADMAP + OPEN-ITEMS, and R-127 filed in both (data_key flag unreliable; O4 can regenerate a DB password that no longer matches the restored data directory). The audit is named D5-drive-alone-restore rather than "...secrets..." because .gitignore blocks *secret* -- a guard worth respecting, not forcing past. |
||
|
|
d42d90fed7 |
R-108 CLOSED — D5's precondition is met (controller v0.187.0)
Four-artifact update per the coupling rule, plus the audit. 07-backup-architecture.md: §10.1 retitled CLOSED with the ruling and the D5 sentence; the FileBrowser network-share row flipped YES->NO, closed at the PLACEMENT rather than at the bind; the exposure chain annotated with the fifth surface (decommission-with-migrate guarded only its source) and the correction that the boundary is the deploy POST, not the dropdown; §7.3 retitled UNBLOCKED; register row collapsed; open question F answered. 00-capability-map.md: new §D row PROVEN-LIVE, with the un-exercised legs named — the deploy-POST and decommission refusals are unit-tested, not live-fired. OPEN-ITEMS.md: R-108 dispositioned; D5 given its OWN row as READY/UNBLOCKED (it had existed only inside other rows' prose — the R-123 thread-loss pattern); R-126 registered. ROADMAP.md: R-108 collapsed to a shipped one-liner; R-126 added. R-126 filed not fixed: a .fab bundle (plaintext secrets, optional password) can be exported ONTO a NAS. Split out of R-108 rather than folded in — it is an explicit customer-chosen export destination, not a browsing surface reaching a backup tree, so it was never part of D5's precondition. Live evidence: same-box before/after on demo-felhom through the real authenticated endpoint, the network-specific refusal on demo-hp, non-effect verified in the registry, and R-67's share-root bind diffed byte-identical across the deploy. |
||
|
|
70f84941d4 |
R-106/R-109 audit + registers: shipped at agent 0.118.1, plus R-125
Adds the full audit: Part 0's three answers, the pre/post recipe for both boxes, the on-disk proof that `local` froze at the 2026-07-28 target move while felhom-backup kept running, all seven red-proofs, and the three publish observables. R-125 filed: v0.118.0's R-106 half shipped INERT. Two tests ran the real Collector.Collect() but both injected a fakeObserver, and the break was one layer below in mergeConfig, which dropped the pbs namespace. The recipe still said "root" — now with namespace_state "resolved" beside it, confident and wrong. Caught by live validation, not by the green suite. Fixed in 0.118.1; filed for the doctrine point that a production-path claim must name the seam it injects at. |
||
|
|
acfc2b7e95 |
R-109 + R-122: the recipe assembly stops dropping sections (hub v0.83.0)
AssembleDRRecipe's hostHalfShape/appHalfShape are ALLOW-LISTS, not the forward-compat their comment advertised: a section an emitter adds is silently discarded until it is named in both the shape struct and AssembledRecipe. No error, no log, no failing test. R-122 (found this session): that already happened and shipped. The controller has emitted offsite_restic since fork-4 — the offsite recovery LOCATION — the hub stored it for all three real customers, and appHalfShape never listed the key, so no delivered recipe has ever contained it. It stayed green because the fixture drAppHalf is hand-written and omits the field. R-109: the agent's new backup_target is a new top-level host-half section and would have been dropped identically, making the fix read as shipped while changing nothing an operator can see. 3 tests built on halves read verbatim out of the live dr_recipe table, plus 2 red-proofs (each mutation asserted to have landed). vet rc=0, suite rc=0, 17 ok. Registers: R-106 + R-109 dispositioned; R-105/R-106 were READY in ROADMAP with no OPEN-ITEMS row (→ R-123, registered); R-124 filed on the "root" spelling. |
||
|
|
3d504d58c8 |
docs(R-117): CLOSED — proven live on demo-hp; R-121 filed for agent-on-box drift
R-117 row → SHIPPED + PROVEN-LIVE (agent v0.117.0), with the full validation in audits/R117-v0117-2026-07-30.md. Both dead states detected on real hardware through the shipped predicate: RETURN raw 8:32 /dev/sdc | bind 8:16 shutdown → stale-device, usable false IN-PLACE both 252:11 emergency_ro, raw unit active → filesystem-aborted, usable false healthy → live 340-497us per call. No block I/O proven by strace (only /proc/self/mountinfo, 0 statfs) — the Part 1 CLAUDE.md fence applied to its own first consumer. No regression through the real pipeline: the live backup-target drive reads bound_under_parent=True via GET /disks with the controller's own credential, with 32 gate lines in 3 min as the positive observable and zero spurious transitions. The ruling asked for in §2.2 is recorded in full and flagged for overrule: Aborted must NOT self-heal. A re-bind lands on the same dead superblock and the call site runs every 20s, so repairing would be an infinite silent retry that masks the state. It surfaces instead. No operator decision was taken quietly — the reasoning is that it routes an already-broken state into the existing gate, event types and Hungarian copy, so no new concept reaches the customer. R-121 filed: a box's installed agent can sit releases behind the vouched one and nothing notices. demo-hp ran 0.113.0 against a vouched 0.116.0 through the whole R-116/R-117 arc. Confirmed at source that R-120's gate cannot catch it — it compares goldenVer against NewestReportedControllerVersion(), i.e. golden-artifact vs fleet-CONTROLLER. MinAgent is protective, not an alarm, and 0.113.0 equalled the floor. Fourth instance of the drift family. Also filed: R-117g (an aborted filesystem is never cleared automatically by design, so it alarms until a human acts, with no guided recovery) and R-117h (StablePathForRaw hardcodes the parent, so the repair path cannot be exercised on hardware without writing into a live customer guest's namespace). |
||
|
|
37515cda7c |
docs(R-117 Part 1): a health check issues no block I/O — and narrow one R-116 claim
Two record items, banked before any Go file is opened. 1. CLAUDE.md gains a standing rule beside the seam-wiring rule: a health check issues no block I/O. A probe that touches a wedged device enters uninterruptible sleep, survives SIGKILL, and cannot be recovered until the device returns or the host reboots — so `systemctl restart` hangs too. A timeout protects the caller's control flow and nothing else. Liveness is decided from /proc and kernel state. Measured in the R-117 spike §6.3: D state 3m50s after kill -9; a buffered write with no fsync blocked too (O_CREAT needs journal access); statfs and getdents returned HEALTHY on a namespace that EIOs every byte. Repeated as a one-line pointer in felhom-agent/CLAUDE.md, because health checks are written in that repo and felhom.eu/CLAUDE.md does not load in an agent-only session — a standing rule that does not load where it binds is the inert-seam shape applied to a rule. 2. The R-116 row gains the clause the spike recommended but did not apply. Its verdict stands and every input to the pairing fix is configuration-derived. But the over-correction window's degraded:false was read off a drive whose bind was dead, so it evidences "the gate did not over-fire", not "the drive was healthy". The two RETURNED lines remain a genuine positive observable, so rule 3 is still satisfied. Nothing else about the row changed. |
||
|
|
e70b5feebe |
docs(R-117): the hang case measured — an I/O probe turns a wedged drive into an unkillable agent
Completes the spike once the venue came back. Q4's hang case and teardown are now measurements, not plans. Against a dmsetup-suspended device (I/O queues instead of returning EIO): - P1 (devno compare) and P2 (ext4 abort flags) completed in 364us / 206us. They read /proc, so no block device is involved. - statfs and getdents completed and reported HEALTHY — on a wedged device they do not even hang. R-117b confirmed in a second failure mode. - EVERY probe that touches the device blocked, including a buffered write with no fsync: the O_CREAT metadata path needs journal access (wchan=do_get_write_access). There is no cheap-and-safe write probe. - The blocked process survived SIGTERM AND SIGKILL (stat=D, wchan=folio_wait_bit_common, still alive 3m50s after kill -9) and died only when the device was resumed. So `systemctl restart felhom-agent` would hang, leaving the agent unrecoverable until the device returns or the host reboots. The thread count does not reveal the leak (5->5, 5->6). Filed as R-117f. A timeout protects the caller's control flow and nothing else, so "the fix must issue no block I/O" is now a fence rather than a preference — the thread-leak hypothesis the probes were built to test turned out to be the weaker half of the result. Teardown done, all three layers: guest 9301 destroyed, r117scratch removed, both dm and both loop devices gone, scsi_debug unloaded, local back to 37.02% against a 37.00% session start. Fences re-verified AFTER teardown: 9201 running, drill-r50 stopped, local-lvm 38.84% byte-identical, felhom-backup content unchanged, live /mnt/felhom-drives intact with both submounts, agent active. Layer 3 genuinely empty — 9301 had no NIC and ran no controller. Trap recorded: a suspended dm device must be resumed BEFORE any umount, or the teardown blocks on the same uninterruptible sleep. |
||
|
|
c949389c95 |
docs(R-117): spike — the mechanism, a recipe, and a steady-state half nobody had looked for
Both halves of the R-113 conjunction are path-presence tests: GuestSeesMount (intermediary.go:276) and isHostMountpoint (:394) compare field 5 of a mountinfo line and never read field 3, so neither can see that the bind and the raw mount name different devices. Measured BoundUnderParent=TRUE over a namespace that EIOs on every read and write. Reproduced 3/3 on a purpose-built scratch LXC on demo-hp; predicates evaluated by a throwaway probe calling the real localapi code from d4eb259. Three results that change the shape of the fix: - Q7: a bind can die in STEADY STATE with no detach/return cycle. The gate produces no action and nothing is emitted on any channel. A Return-branch fix cannot reach this half, and a devno comparison does not detect it. - Q6/R-117d: AttachDrive's normalize leg already performs the repair, and three call sites already invoke it - including the controller's Return branch before it restarts apps. All defeated by one early return at :235. Unblock the existing path; do not add a new one. - Q1: the device-node change is a CONSEQUENCE, not a precondition. The stale bind pins the dead superblock, forcing the returning device onto a new number. Control test: released, the letter is reused. Not established: the hang case. Venue and probes built, run lost to a site internet outage; the thread-leak hypothesis is not claimed as a result. Teardown of the spike venue is owed - commands in the findings doc; nothing fenced was touched and no hub-side record was created. |
||
|
|
29bcfeb214 |
docs(R-120): CLOSED on both halves — golden current, and the class has a gate that refuses
Half 1, the artifact: golden 0.186.0 baked, published, vouched, and proven on a REAL day-0 on demo-hp (not the fixture, per the rule committed in Part 1). With the target detached, the fresh box's endpoint returned the TargetAbsent copy -- "A rendszermentés meghajtója nem érhető el — amíg vissza nem csatlakoztatod..." -- with offer_path absent entirely. The day-old read on the 0.185.1 golden had returned the false system-disk message plus an offer of the other drive. That is the customer-visible defect closed. Half 2, the mechanism: operator ruled REFUSE, shipped as hub v0.82.0 and DEPLOYED. Proven live by re-attempting the original mistake -- vouching the stale 0.185.1 golden now yields HTTP 303 flash=golden_behind_fleet plus [WARN] artifact vouch REFUSED, and the manifest reads back unchanged at 0.186.0. Refused AND unwritten, against the real fleet signal rather than a unit fixture. Recorded on R-29's audit list as the first ENFORCED gate beside its three orphans, so the contrast is kept rather than lost. The orphans are unchanged -- this proves the pattern is available, not that the backlog moved. Teardown all three layers: VM 9402 purged, r120-images removed with the space measured back, hub layer gate-blocked on ONLINE with the command recorded. Last session's sess-e was deleted this run, discharging its recorded layer 3. |
||
|
|
1a68b53b06 |
hub v0.82.0 (R-120): the vouch path REFUSES a golden the fleet has already outrun
The golden's version IS the controller it bakes (build-golden.sh:345 defaults GOLDEN_VERSION to the controller tag), so a golden behind the newest deployed controller means every FRESH install lands on stale application code. On the R-120 occurrence that stale code shipped a customer-facing falsehood: a box from the 0.185.1 golden told a customer whose backup drive had fallen out that the backup was on the same disk as the system -- false, the drive was gone -- and offered a different drive as the remedy. WHY A GATE, NOT A REMINDER. The gap has opened three times: R-111 (golden's agent 17 releases behind), R-115 (agent built and deployed, never published), R-120 (this). The first two were closed by re-baking and remembering; remembering then failed again. R-29 is the standing proof that a check nobody runs is worse than none because it reads as coverage -- hostinstall_gates.py sat RED and uninvoked across three version bumps and hub_confirm_gate.py has never run at all. So the property that matters is not whether a check exists but whether it BLOCKS. - Wired into handleSetArtifacts (internal/web/configs.go), immediately before the only write, on the sole UI path to SetArtifactManifest -- it runs on every vouch without anyone choosing to. A script in scripts/ would have been a fourth orphan. - It REFUSES (operator ruling, 2026-07-30), with a flash naming the remedy. - Signal: store.NewestReportedControllerVersion() over reports.controller_version, SEMVER-compared in Go -- MAX() in SQL ranks 0.99.0 above 0.186.0, a pair this fleet has shipped. No outbound call, no new credential. - Fail-open in exactly two deliberate cases: an empty golden field (clearing the manifest is legitimate) and an unknown fleet version (a new hub must vouch its first golden). NEAR-MISS RECORDED: the first draft read guests.controller_version, a column that exists in the schema and that NOTHING writes -- it would always have seen "" and failed open, i.e. inert, this gate's own failure shape. Caught by grepping for a writer before trusting the column. Blind spot stated rather than papered over: a controller no box has ever run is invisible to this signal. Not the failure that has bitten -- all three instances were deployed-newer-than-baked. 4 tests through the PRODUCTION handler over httptest, never an injected seam. The refusal asserts both the flash and that the manifest was NOT written, because a gate that redirects and saves anyway reads as enforcement while providing none. Red-proof: deleting the block makes the stale golden vouchable and both assertions fail. ROADMAP R-29's audit list now records this as the FIRST enforced gate, so the contrast with its three orphans is kept rather than lost. The orphans are unchanged. Suite rc=0 read separately from this commit. |
||
|
|
49b627684c |
docs(R-120): golden rebaked to 0.186.0, published, vouched, proven on a real day-0
The golden baked controller 0.185.1 -- confirmed from the golden's OWN record (drill/bake-0.185.1.log:1 and :330) and from build-golden.sh:345, which derives GOLDEN_VERSION from the controller tag. 0.185.1 predates R-114 + R-112, so every freshly installed box told a customer whose backup drive had fallen out that the backup was on the same disk as the system (false) and offered a different drive as the remedy. Baked golden 0.186.0 from main's controller in the DooPlex bake fixture: overlay2 OK, 3 mounts included, FATAL 0, exclusions 0, 618 MB, upload HTTP 201, GOLDEN_SHA256 b760ac6a33e70700..., token-leak grep 0, GL-1 teardown with drill.qcow2 back to virgin. Three observables, quoted as returned: PUBLISHED (anonymous GET -- what the installer does -- 200 / 648930639 bytes / sha identical to the bake); VOUCHED (manifest read BACK, not the 303); RESOLVED BY A CONSUMER (Artifact manifest served for customer sess-f, golden=0.186.0). Floor NOT touched per publish-train rule 2 -- it is a separate form and min_controller_version still reads 0.156.0. MinAgent left 0.113.0 because 0.186.0 declares it unchanged. Proven on a REAL day-0 on demo-hp, not the fixture, per the rule committed in Part 1: VM 9402 from the v1.25.0 ISO -> Controller elindult (0.186.0), box confirms felhom-controller:0.186.0 + agent 0.116.0. A fresh box now runs 0.186.0 where it ran 0.185.1. The procedure was NOT unwritten: RUNBOOK-manual-build.md:101-115 documents it and build-golden.sh carries its own usage and publishes to Gitea itself. One documentation-integrity finding: that runbook says to use the RECORDED qemu line and not reconstruct, while the line it cites is itself labelled reconstructed, the canonical one never having been saved. NOT done and not claimed: the TargetAbsent/empty-offer_path endpoint capture (the claim gate runs before auth with no Bearer escape -- R-119's fourth instance), and the Part 3 mechanism, which awaits the operator ruling. Recommendation and exact wiring recorded in the audit rather than built. VM 9402 + r120-images + customer sess-f retained pending that read, with teardown commands recorded. Previous session's sess-e layer-3 is now DISCHARGED -- it aged to STALE and the cascade completed, full residue purge logged. |
||
|
|
376365bb12 |
docs(target-selection): a fixture may prove a mechanism; only a fresh box may prove a path
The page said which machine is safe to break but not when reusing a test box is legitimate. That distinction is exactly what surfaced R-120: R-116's closing run deliberately did a real day-0 from the ISO instead of reusing the standing fixture, and the fresh box installed the golden's controller -- a release behind -- and showed the customer the wrong absent-target message. A fixture would have shown a controller nobody installs. Adds to the Tier 1 section: a reusable snapshot-reset fixture is the right default for MECHANISM work (payload capture, fix cycles, claims about code behaviour), while a fresh day-0 from the ISO is REQUIRED for any claim about the install path, the golden image, agent publish/vouch or first-boot state -- naming the drift family it exists to catch (R-111, R-115, R-120). Also: a fixture must record its provenance (which golden, agent and controller, and when), because a fixture whose versions drift silently is R-120's mechanism turned into a permanent installation -- worse than no fixture, since it produces confident wrong results quickly. Part 1 of the R-120 task, committed alone and before the bake. Docs only. |
||
|
|
772956d214 |
docs(R-116): CLOSED — proven live; capability row F to PROVEN-LIVE; R-120 filed
The events leg the previous commit reported as not-reached is now done. The operator relayed the claim code (the only route: bcrypt-hashed hub-side, emailed only), the two storage paths were registered through the real POST /api/storage/register, and the cycle ran on the fresh box: 07:20:04 backup_target_absent (error) Cel meghajto <- TARGET, specific 07:22:34 backup_target_restored (info) Cel meghajto <- its matching pair 07:24:04 storage_disconnected (error) Adat meghajto <- NON-target, generic 07:25:34 storage_reconnected (info) Adat meghajto All four at the hub; gate fired in 3 s. Two matched pairs, correctly discriminated -- and discrimination is proven NON-trivially for the first time, since both prior runs had the target itself emit the generic event. Over-correction passes on a positive observable, with two RETURNED lines proving the gate was ticking. 00-capability-map row F: PARTIAL -> PROVEN-LIVE with the evidence and the caveat. R-120 filed: the golden bakes controller 0.185.1, which PREDATES R-114 + R-112, so a freshly installed box shows the customer the WRONG absent-target message -- observed live on the drill box: the generic "the backup is on the same disk as the system" copy (false; the target is a drive that vanished) plus an offer of the other drive as the remedy. That is E2D 5.3's exact payload, still reachable on any new install. R-115's class one layer up -- R-111 closed by re-baking the golden, 0.186.0 then shipped, the golden did not move, and the gap reopened silently; this time the stale artifact carries a customer-facing falsehood in exactly the state R-116 now alarms about correctly. Teardown recorded for all three layers, hub layer gate-blocked with the command. |
||
|
|
315c469fc8 |
docs(R-116): v0.116.0 proven live at the payload layer; events leg blocked on an emailed claim code
audits/R116-v0116-2026-07-30.md + the R-116 register row. WHAT PASSED, on real hardware. Agent 0.116.0 published (independent registry GET verified the bytes), vouched, and installed UNAIDED by a fresh box -- "Artifact manifest served for customer sess-e (agent=0.116.0 golden=0.185.1)", host sess-e-5d4427 ... 0.116.0 ONLINE. Real day-0 on a nested PVE on demo-hp (per runbooks/target-selection.md, which sent this run there rather than to the DooPlex fixture the previous run used), both drives enrolled through the real endpoints, device loss a real hot-detach. Captured live, absent state: the target is now ONE row carrying backup_target:true AND guest_path:/mnt/felhom-drives/cel with mount_path:"", so isTarget[/mnt/felhom-drives/cel] = TRUE -- it was false through v0.115.0. RETURNED gives true as well, so the pair matches. All three guards pass from the same payload: R-114 preserved (no row combines the flag with a non-empty mount_path), no over-correction (bound_under_parent:false), and discrimination at the payload layer (the non-target carries the flag on no row) -- the thing neither prior run could show. WHAT DID NOT HAPPEN, and is not claimed. No backup_target_absent or backup_target_restored event was observed on the wire. planDriveGates iterates registered StoragePaths and the drill controller has none ([WARN] Storage paths: no storage paths registered); every storage route answers 401 "dashboard not yet claimed". The claim code is bcrypt-hashed and emailed-only, and handleSelfBindLinkSend (selfbind_mint.go:139-161) renders a flash and never the token, so no operator-side route exists. A gen-2 code was re-sent; the drill VM, its storage and customer sess-e are DELIBERATELY RETAINED with teardown commands recorded, so the leg finishes without a rebuild. Reported as not-reached rather than as a third trivial pass. R-119 filed: the claim gate makes drive-gate legs unreachable to CC by design, and has now stopped three sessions at the same wall -- needs a ruling (operator-scoped test affordance, or a documented prerequisite step), not a fix. R-117 reproduced on real hardware with a read/write probe (EIO both directions while /disks reports attached + bound_under_parent:true) and §5 records how it colours the reattach leg. R-118's symptom vanishes incidentally on this one row; R-118 is NOT fixed. sess-c and sess-d verified GONE (404, absent from both tables) -- cleared by the operator using the previously recorded commands, not by this session. |
||
|
|
1aa1bd17c2 |
docs(template): §13 gains a teardown step, §15 gains its evidence line
Three drills, three orphaned hub customers -- drill-r50, sess-c, sess-d -- because §13 covered the clean-tree gate, build/deploy, live validation and the STOP point and said nothing about teardown at all. Layers 1 and 2 (the VM and its volumes; the host's reclaimed space) get remembered because they are visible on the box. Layer 3, the hub-side customer or appliance record, is invisible from there and has been missed every time -- sess-c was not even recorded by its own report, so the record claimed a clean teardown that had not happened. §13: a Teardown subsection at the end, before §14. All three layers, with the hub layer requiring an EXPLICIT disposition -- deleted, retained as a fixture with the reason, or gate-blocked with the command recorded -- because silence is how drill-r50 became simultaneously a blocked customer and the only drift fixture. Cites runbooks/target-selection.md for which machine to provision on rather than restating it. §15: deliverable 8 demands the evidence for all three layers and names the failure it prevents; the former 8 (Observations) becomes 9. No section renumbered, §13/§15 not restructured, author checklist untouched. Part 1 of the R-116 join task, committed alone and before any Go file is opened -- the code half ends in a live run and live runs have stalled twice, while the record work is unconditional. |
||
|
|
e6b5fa1e63 |
docs: retract the expired NVMe fence, drop component versions from the inventory, fence acts in the template
Follow-up acting on the observations filed with runbooks/target-selection.md. operations/nodes.md - The demo-hp NVMe was documented "PRESENT AND UNENROLLED -- do not touch" and listed under "What is NOT enrolled here (deliberately)". Both are FALSE and had been for eight days: it was enrolled 2026-07-22 through the normal Tarhely flow and is now /mnt/nvme-1tb -- the enrolled user-data drive AND the felhom-backup target (verified live 2026-07-30: nvme0n1 -> /mnt/nvme-1tb, and dir: felhom-backup / path /mnt/nvme-1tb / is_mountpoint 1). The fence's own condition (join via Tarhely, not the installer, not by hand) was SATISFIED, so the prohibition expired with it -- while still contradicting the task specs that correctly sent drill-VM disks there. Retracted with its reason recorded, and the caution that IS still live kept (dir storage at the mountpoint ROOT, else exactMount fails and the storage reads disconnected forever). - Component versions REMOVED and a note explains why: agent/controller/hub versions change several times a day, so a number written in an inventory is wrong within hours and then read as fact -- and the fleet is not uniform (on 2026-07-30 the two boxes ran different agent AND different controller versions). Points at the authorities instead: hub /hosts + /configs, felhom-agent --version, docker ps. - Site addresses now say re-check rather than asserting one (the N100 read .162, not the recorded .147); records that LAN literals are unreachable from DooPlex while the boxes are away. Adds the target-selection pointer: this page is what the hardware IS, that page is what may be done to it. PROMPT-TEMPLATE.md -- the upstream generator of the defect - Section 12's "Do NOT touch [the untouchable]" asked the spec author to name a THING. Now asks for the forbidden ACT plus its REASON, with the demo-hp case as the worked example of how a bare object-fence over-reads. - Section 13 gains the positive counterpart, which was the actual gap: if a task needs a machine to break, NAME IT. Listing only what is off-limits leaves the most valuable unfenced machine as the residual choice. runbooks/workspace-CLAUDE.md (+ the untracked root copy re-synced, verified identical) - Host table gains a Blast radius column and the missing demo-hp row, notes felhotest as Connection refused, and points at target-selection.md. This is the file that loads FIRST every session, so leaving it with the old table would have undercut the whole fix. No code, no build, no deploy, no host reconfigured or renamed. |
||
|
|
699790b12d |
docs: write down which boxes are disposable (target selection by blast radius)
Nothing in the repo said which machines are safe to break. The host table gave access and role and stopped there, so a session needing a victim had to guess -- and the guessing inverted: the two boxes that exist to be broken were treated as sacred, and DooPlex (the recovery chain) got used because it was the only box no spec had fenced. New documentation/runbooks/target-selection.md -- one page, three tiers, and per machine what is freely permitted / needs care / forbidden, each carrying its REASON so a rule can be correctly narrowed later instead of ossifying. States the selection rule positively (start at Tier 0; a Tier 2 box only when a task says so explicitly; an absent fence is not permission) and that fences name ACTS, not machines -- demo-hp's over-subscribed local-lvm is one dangerous storage, not a dangerous box. CLAUDE.md: host table gains a Blast radius column, gains the missing demo-hp row (it was where the drill VMs ran and it was not in the table at all), and a pointer line to the new runbook. CORRECTION to the spec's problem statement: the designation was not missing. The 2026-07-25 operator ruling naming the t740 as drill+build VM host -- explicitly "moved off DooPlex" -- already existed in operations/nodes.md. It sat where no session reads at start, while the prohibitions were repeated in every task spec. The defect is reachability of the ruling, not its absence, and the R-116 drill on DooPlex contradicted a written ruling rather than filling a vacuum. CORRECTION to the R-116 record, same commit: the baseline claimed controller 0.186.0 on both demo boxes. Only felhom-pve was sampled and generalised; demo-hp re-checked directly runs 0.185.1, so the fleet is split and R-114's TargetAbsent branch is absent from demo-hp. Fixed in the audit table and REPORT-r116-diag. Docs only -- no code, no build, no deploy, no host reconfigured, no host renamed. |
||
|
|
d56e395a2a |
docs(R-116): isolate the mechanism from the real /disks payload; file R-117 + R-118
The absent-state /disks payload was captured on a genuine device loss, after a present-drive control run proved the query works (Part 5's three attempts failed on token extraction, and its control returned 0 rows). The answer is theory #1 -- "the registry-union row writes false" -- which was raised, declared wrong and retracted. The retraction was the error. Absent state returns 4 rows, not 3. The drive appears twice and the two facts the controller needs sit on different rows: the Observe row has backup_target:true but mount_path:"" and guest_path:"", so it contributes no key to driveTargetByPath; the registry-union row owns /mnt/felhom-drives/<name> and omits BackupTarget from its struct literal (disks.go:301-306) => false. The union row is not deduped because seen is keyed on MountPath (:290-295), the one field the absent state empties, and its own MountPath comes from the systemd .mount unit FILE (registry_known.go:40-75), which never reads the mount table. Theory #2 (the basis of the shipped v0.115.0) is false on both halves; #3 is false too. v0.115.0 is provably inert: StablePathForRaw("") returns "". Also files the read path verbatim -- the token plaintext lives only in bootstrap.json on the Proxmox host; the agent's store keeps hashes only. New: R-117 (READY M, outranks R-116) -- a returned drive's guest bind is a DEAD mount (EIO both ways) while /disks reports attached + bound_under_parent:true, so the gate restarts the customer's apps onto it and reports healthy with no alarm. R-118 (READY XS) -- an absent drive's union row advertises the root filesystem's capacity as its own. Docs only. No code written, nothing built or published; v0.115.0 untouched. Both demo boxes read-only; drill fixture restored to virgin. |
||
|
|
c3ce4c7b20 |
R-116 Part 5 FAILED: the fix shipped, C5 still fails, mechanism NOT isolated
A fresh box running the fully shipped stack -- agent 0.115.0 from the Day-0 manifest plus controller 0.185.1 from the vouched golden, no hand-deploy -- still fired the GENERIC storage_disconnected on detach and the SPECIFIC backup_target_restored on return. backup_target_absent count 0. Identical to Session C. The v0.115.0 fix changed nothing observable. Part 4's three positive observables were all obtained before the run (registry newest 0.115.0, hub vouches 0.115.0, felhom-pve running 0.115.0 clean), so the publish step forgotten twice was not forgotten a third time, and the box demonstrably installed the fix under test. Discrimination FAILS: the target itself produced the generic event, so the two cannot be told apart regardless of the non-target leg -- which was therefore not staged. Reported as a fail, not as Session C's trivial pass. Over-correction guard PASSES: 0 ABSENT lines with the drive present, target degraded:false. THE HONEST PART. The fix targets a shape that does not occur live, and which shape does occur is NOT ISOLATED. With the drive detached PVE reports the storage inactive with zeroed fields -- a shape the unit fixture did not model. Three attempts to read the real /disks payload failed on token extraction across the ssh -> guest -> container layers, and a present-drive CONTROL query also returned 0 rows, proving the query was broken rather than the payload. Without that control this run would have recorded a third false mechanism, after "the union row writes false" (wrong, corrected yesterday) and "no row carries the guest path" (unverified). The leading hypothesis -- an inactive storage reaching Observe with an empty MountPath, so StablePathForRaw returns "" -- is consistent with the pvesm output but is NOT evidence and is recorded as such. Next session's first job is a working /disks read, with a present-drive control run FIRST, before any further code. agent v0.115.0 is published, vouched and INERT. Not reverted: reverting is itself a change, the runbook forbids fixing mid-run, and the code is tested and harmless. Capability-map row F stays PARTIAL, now citing the re-test. Teardown clean: pvesm status after == before (local-lvm 38.83%), guest 9201 and drill-r50 untouched. Customer sess-d pending the usual ONLINE-ages-to-DOWN gate. |
||
|
|
e87d6b26bb |
Correct the Session C audit: the union row is DEDUPED AWAY, not written false
The audit said the union row "writes false" for the guest-path key. That is wrong, and the next reader would have inherited the error. Isolated during R-116's Phase 0: RoleForStorage returns RoleSystem whenever backingDevice == "" (felhom-agent internal/storage/role.go:180-181). When the device vanishes the target row's role flips to system and it loses its guest path, but KEEPS its MountPath -- and the union loop skips any drive whose MountPath is already seen, so the registry row is never emitted at all. /disks therefore carries NO row with that guest path: isTarget[guestPath] is a MISSING KEY, not a false value. The practical difference is decisive -- the obvious fix (set BackupTarget on the union row) could not have worked, because that row does not exist in the state where the alarm is needed. The section's own "not isolated" caveat is replaced by the isolated answer. |
||
|
|
952ebf4862 |
Record work, banked first: shrink the E-2d row, create the missing capability-map rows
Unconditional and three sessions overdue, so it commits before any code is
touched — E-2d itself stopped at Phase 0 and banked nothing.
E-2d row: 822 words -> 121, and the contradiction resolved. Its State read
CLOSED — PARTIALLY PROVEN while the cell's final sentence read "This row stays
OPEN only for the residue"; a reader could not tell which. It is CLOSED, with
R-116 the single named open leg.
Nothing unique was binned. Three facts existed ONLY in that cell and are moved
into audits/E2D-fresh-vm-2026-07-29.md as a new §1a: the local-lvm fence figures
with the 888 GB nvme alternative, the exactMount subdirectory caveat and why the
subdirectory is nonetheless the safe placement (no durable_id collision), and
the ISO/PAIRING -> DIRECT fall-through derived at source with its line
citations. drill-r50's blocked status was already in both audits.
Capability map: it had ZERO rows for the backup-target work — grep gives 0 hits
for backup_target and one for "E-2" that is a campaign date string. Three
scenario rows added, at today's honest status, not the value hoped for later:
C. Protection & recovery — installer Case A/B, DEGRADED recorded not hidden
PROVEN-LIVE, cites E2D-fresh-vm C1+C2
D. Storage & devices — the offer, and that registration confers no role
PROVEN-LIVE, cites SESSION-C C4 + the decline path
F. Notifications & monitoring — the absent-target alarm and its pairing
PARTIAL, cites SESSION-C C5, leg named, -> R-116
Row F is PARTIAL today per the doc's own strict enum (a leg not exercised live
is PARTIAL with the leg named, never PROVEN-LIVE). A later session may flip it;
this commit must not.
|
||
|
|
06d7788392 |
Session C: R-113/R-114/R-112 PROVEN LIVE; C5 fails on a new defect (R-116)
Full ISO/PAIRING run on a fresh nested box. Agent 0.114.0 came from the Day-0 manifest -- the SHIPPED binary -- so C5 tested the real artifact. Controller 0.186.0 hand-deployed after install per the §3.1 ruling; the vouched golden bakes 0.185.1, so C3/C4 prove the code not the shipped golden, and that lag is filed against R-115 rather than a new ID. R-113 PROVEN: detach 18:43:50, gate fired 18:43:54 -- four seconds, where E-2d measured zero over 4.5 minutes -- and SetDisconnected was reached. It fired on exactly the shape that defeated it: raw /mnt/mentes NOT mounted while the bind /mnt/felhom-drives/mentes still read /dev/sdb[/felhom-data]. R-114 PROVEN: with the target absent the page rendered the absent copy, the system-disk copy 0 and the offer block 0. Both of E-2d's falsehoods are gone. R-112 PROVEN: the banner reached a customer's page for the first time. Healthy renders nothing, proven POSITIVELY -- idle delta 0 /backup/tiers calls, page load delta +1, single caller, so the seam ran and chose silence. C5 FAILED on a fourth, separate defect. The alarm fires but as the GENERIC storage_disconnected, while the recovery is the SPECIFIC backup_target_restored -- a pair an operator cannot match, which is what notifyDriveReturned's own comment forbids. backup_target_absent count 0 across the run. Root cause: the drive is TWO /disks rows and BackupTarget and GuestPath sit on different ones; absent they separate, on return they rejoin. v0.184.1 fixed the keying, not this. Only reachable because R-113 made the gate fire at all. Filed as R-116. Mirror + over-correction guard PASS: non-target drive -> storage_disconnected, backup_target_absent 0; both drives present -> 0 ABSENT lines and the target stayed healthy. Caveat recorded: the mirror passes trivially because the target also produced the generic event. E-2 and E-2d CLOSED as partially proven with R-116 the one named open leg, per the runbook's §9 rule decided in advance rather than mid-run. Capability map NOT touched: it has no E-2 rows at all, so nothing could move to PROVEN-LIVE. Creating them is a design act, not a validation act. Teardown clean: pvesm status after == before (local-lvm 38.78%), guest 9201 and drill-r50 untouched. Customer delete attempted and correctly refused while the host still reads ONLINE; command recorded for once it ages to DOWN. |
||
|
|
af518ba151 |
R-114 + R-112 code shipped (controller v0.186.0) — seam proven live, copy not
R-114: new BackupTargetState.TargetAbsent separates configured-and-gone from never-configured. Degraded keeps its meaning so the wire contract is unchanged; TargetAbsent answers which problem, because the remedies are opposite. Copy is verbatim the hub's backup_target_absent email. The offer is suppressed on the branch itself, not left to firstOfferableDrive's Disconnected skip -- that flag comes from R-113 in another repo and this state must be right without it. R-112: the state finally has a consumer. Server-rendered on /backups via backupsHandler -> backupTargetView -> backups.html, not a 19th JS fetch. The view is nil for healthy and unknown so those render nothing at all. SEAM PROVEN LIVE by a DIFFERENTIAL positive observable rather than by an absent banner: idle 8s produced 0 new /backup/tiers agent calls; each /backups load produced exactly +1, and that call has a single caller. The demo box is healthy and correctly rendered nothing, which matches its real state but is a negative and so proves nothing about wiring on its own. MinAgent unchanged at 0.113.0 -- R-114 reads BackupTarget/MountPath/GuestPath/ Role, none of which R-113 altered. demo-hp is not held. Session C scope unchanged: neither fix touches the agent, so the leg awaiting proof is still device loss -> gate Stop -> SetDisconnected -> backup_target_absent on the wire. One rebuild validates all three. |
||
|
|
338b2ccf86 |
agent 0.114.0 published + vouched; R-115 files the recurring publish gap
PART 1 — Session C unblocked. Agent 0.114.0 (the R-113 fix) was built, pushed and deployed but never published, so a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix. Published from the clean tree at b58d7bc via scripts/publish-agent.sh; sha 5e4c15ebee2d7583d57301d1f7c9cc7d4276262966bf738b05e34653bfd18c31, verified by an INDEPENDENT round-trip GET (http=200, sha match, binary self-reports 0.114.0), and the hub manifest read back after the write. Deliberately NOT done, each with a reason: - No golden bake. The golden bakes the CONTROLLER, not the agent, and host-install fetches them as separate generic packages (:1945 / :2573). Golden 0.185.1 is current, so there is no new-agent-against-old-golden risk. - min_agent NOT raised, stays 0.113.0. It expresses what the CONTROLLER requires of the agent, and controller v0.185.0 declares MinAgent 0.113.0 — which 0.114.0 already satisfies. Raising it to 0.114.0 would have been a false claim AND would have held demo-hp and drill-r50. No box is held; no §3 STOP fired. - Global controller floor NOT raised (v0.156.0), per R-111's reasoning. - wrapper_sha256 preserved verbatim; re-checked against configs/felhom-pbs-apply before and after — no drift both times. demo-hp RULING: left on 0.113.0. The R-113 fix is not live-validated, so putting it on a second box widens exposure for no proof, and Session C's nested box takes its agent from the manifest, not from demo-hp's host agent. Move the fleet once, after Session C. PART 2 — R-115 opened (WAITING-ON-OPERATOR). The finding is the RECURRENCE, not either instance: publishing is a remembered step, and it was forgotten within eight hours of R-111 documenting it as forgettable. Filed as a new ID with a back-pointer rather than reopening R-111, because R-111's finding (the channel WAS stale) is closed and verified end-to-end, while the process defect that caused it is a distinct problem with a distinct fix and owner. Class cross-linked to R-29 (a control that exists and is never walked) WITHOUT minting a second ID for it. Options are stated as the operator's decision, with mechanisms (build-step, deploy gate) separated from reminders (checklist, manual) — R-29's whole finding being that reminders do not hold. No code written, by design. R-111 gains a deferred-leg-recurred line; its shipped evidence is untouched and it is NOT reopened. R-113 records that Session C is now unblocked. |
||
|
|
ca4c8b3afc |
R-113 code shipped (agent v0.114.0) — NOT live-validated, awaiting Session C
BoundUnderParent is now a CONJUNCTION: bound under the parent AND the drive's raw host mount still mounted. The raw mount is the device-bound systemd unit that dies with the device; the agent's own bind is not, which is why the bind outlived the device and the gate could never fire. Conjunction deliberately, not replacement: the device half alone would regress boot ordering (raw mounts early, bind lands ~18s later — that window must keep reading absent), so existing behaviour is byte-identical and only the unreachable case is closed. Unknown is never absent. Controller UNCHANGED, no MinAgent bump — BoundUnderParent has exactly one functional consumer (planDriveGates:226). A new DevicePresent bool was rejected: absent-from-JSON decodes to false, so every drive on an older agent would have read ABSENT and stopped its apps. +6 tests (208->214), 4 red-proofs run and reverted. Deployed to demo-felhom and the over-correction guard verified in production: raw mount present, drive still reads present, 10/10 apps untouched, no gate action, no false alarm. demo-hp deliberately left on 0.113.0 (the spec scoped deploy to felhom-pve). SESSION C BLOCKER recorded on the row: the hub Day-0 manifest vouches agent 0.113.0, so a fresh drill box would install WITHOUT this fix and validate nothing. Publish + vouch 0.114.0 first — R-111's trap in the same shape. |
||
|
|
d839ddcb60 |
E-2d teardown complete: drill customer + host removed from the hub
The delete was correctly refused at four successive gates while the host still read ONLINE (acknowledgements -> typed confirm_id -> expect_hosts stale-preview -> "host is ONLINE"). Rather than force it, the run waited for the destroyed host to age to DOWN; delete-impact then reported deletable:true and the documented cascade ran: host deleted (escrow demoted to retained custody), tenantsync deprovisioned, PBS tenancy deprovisioned, claim reset to unclaimed, residue purged (reports=5 app_telemetry=5 notif_prefs=1 appliance_registrations=1) Verified after: 0 occurrences of "e2d" anywhere on the hosts page; demo-felhom and demo-hp ONLINE on agent 0.113.0; drill-r50 and peti-felhom unchanged; demo-hp carries only guest 9201 and VM 300. Scoping checked rather than assumed: the single purged appliance_registration was this run's own appliance (810d10c5, bound to e2d-fresh). The unrelated stale 2026-07-25 appliance (206c8838 / QWA-WJE) was NOT touched by the cascade — the operator removed it separately. - OPEN-ITEMS.md: the drill-cleanup WATCHING row is removed (done, not open). - audits/E2D-fresh-vm-2026-07-29.md §8 + REPORT-e2d.md: teardown recorded as complete, with the cascade output and the appliance-scoping note. |