Commit Graph

163 Commits

Author SHA1 Message Date
admin e34b614e5b docs: R-182 closed, R-90 closed on measurement, R-86 unblocked, ep0 record corrected
gates / gates (push) Successful in 7s
R-182 CLOSED (controller v0.194.0 + hub v0.90.0/.1), proven live on demo-hp. The
hub's notification_log for the run reads: two per-app failures RECORDED, one
digest SENT naming both, and the customer channel SKIPPED with operator_only.
Against the measured previous behaviour — two failures, one email naming one
app, one leaving no trace anywhere.

Scenario D proved itself on an event I had not planned: disk_critical alarmed on
two filesystems, the second was collapsed by the cooldown, and that collapse is
now visible WITH ITS KEY. Yesterday it would have left nothing at all.

A gap the spec did not anticipate is recorded with its fix: the per-app event
also fires from the periodic sweep, outside any run, so making it record-only
would have created a NEW silence. The sweep emits a digest too, with no run_id,
so it stays under the ordinary hourly cooldown.

ep0: MEASURED on the box — 7757 MB (8 GB), 4 vCPU, and the 4 GiB swapfile
SURVIVED the resize and is active (checked, because a resize is a stop/start).
The 40 GB local disk is UNCHANGED, so no disk figure was touched anywhere.

Five documents corrected — three of which the task's list did not name, found by
searching. Two audit/evidence documents ANNOTATED, body untouched: they record
what was true when written and that is their value.

R-90 CLOSED. R-86 unblocked and re-ranked, stated honestly: 8 GB is comfortable,
not unbounded — the original OOM was a 14.46 GB restore — so the restore-test
cadence should still be paced, just not by fear of the endpoint.

target-selection.md's "D-d did not name ep0 either way" is deliberately left
standing. It is the operator's question, not CC's.

STATUS.md 127 -> 83 lines, items rather than sentences.
2026-08-03 14:00:54 +02:00
admin b0b269b28d docs: R-110 + R-115 closed, R-182 re-scoped by measurement, R-183/R-184 filed
gates / gates (push) Successful in 7s
R-182's direction REVERSED by Part 0's measurement. Filed yesterday as "the
reserve re-alerts on every status refresh" — too many alerts, seen at the
sending end. Measured at the receiving end: 9 events received today, 2 operator
emails sent. When two apps are refused in the same second the operator is told
about ONE; the other is dropped before LogNotification, so it leaves no row on
any channel and cannot be audited. The operator cooldown key is
customerID:eventType(+tier) and the capture-failed event carries `app` but no
`tier`, so the key has no app identifier. Same failure mode as R-97a, in a
second event type that never opted into the narrow fix. Nothing changed —
Part 0 was investigation only.

Correction owed: yesterday's report said "one recovery_unit_capture_failed per
app, HTTP 200". True of what the CONTROLLER pushed; a reader would take it as
"the operator was told about each app", which is false.

R-110 CLOSED (installer v1.23.0). Both channels moved. The spec's mechanism for
channel 2 rested on a factual error — the run-time fetches are sixteen, not
nine, and come from felhom-agent, not this repo — so no tag here could cover
them; pinned to the agent version being installed instead, on the operator's
ruling. Channel 3 needed no change: the URL never carried a ref, so no hub
change and no hub bump.

R-115 CLOSED. release-agent.sh builds, tags, publishes and verifies by an
independent download; check-published-versions.py refuses a tag with no package;
CI now runs the full gate set so it actually runs.

R-183 NEW+CLOSED: a fresh install fetched the vouched agent binary and its
sixteen config files from two different refs, and nothing compared them.

R-184 NEW: nothing stops the hub vouching a version that was never released.
The R-115 gate cannot see it — measured, the hub manifest and Gitea's package
listing are both 401 anonymously.

capability map: new PROVEN-LIVE row for the published installer channel.
STATUS.md 138 -> 127 lines.
2026-08-03 12:44:08 +02:00
admin fb652024ea docs: R-181 closed, R-156 closed, R-110 + R-115 rulings recorded, R-182 filed
gates / gates (push) Successful in 7s
R-181 CLOSED (controller v0.193.0 + v0.193.1) and proven live on demo-hp for
BOTH reserve terms. The reserve is now a per-app, per-run ADMISSION decision
taken before the app's first write and covering all three write legs, and it
gained a size term. The refusal's wording was not weakened; the behaviour moved
so it became true, verified by sha256 tree fingerprint.

R-156 CLOSED — papra's template mounts the app's own data root. Precondition
re-measured rather than inherited (both boxes were wiped today).

Part 4, documentation only, nothing built:
- R-110 WAITING-ON-OPERATOR -> READY. Ruling: option (b), the installer's
  publish channel moves to a TAG. Recorded with the condition that decides
  whether it works at all — it must cover BOTH the /scripts/ git-sync and the
  nine files the installer fetches from raw/branch/main.
- R-115 WAITING-ON-OPERATOR -> READY. Ruling: mechanism (b), a build-side gate
  refusing to deploy or vouch an unpublished version. The third instance (agent
  v0.120.0) would have silently downgraded both demo boxes while succeeding.

R-182 NEW: the periodic status refresh has no admission scope, so a refused app
re-alerts on every poll (measured: a second alert pair 13s after the run's).
Pre-existing in v0.192.0; deliberately not fixed in the R-181 task.

capability map: the local-backup row moves to PROVEN-LIVE in BOTH halves.
ROADMAP: R-165 collapses to CLOSED; R-181 collapsed into it.
07-backup-architecture.md: the reserve's contract stated as what the code
provides (S-1 — an architectural contract changed in the same session).
STATUS.md trimmed 150 -> 111 lines, "What's broken" no longer holds shipped
work, and the stale "After:" line (pointing at work that shipped on 2 August)
is fixed.
2026-08-03 11:36:16 +02:00
admin aa62449694 R-178 CLOSED: both demo boxes reinstalled from the merged golden and proven
gates / gates (push) Successful in 8s
Two boxes, two DIFFERENT supply paths, so the session proved the disk shape and
the delivery route rather than one of them twice.

demo-hp (layout proof, --golden <local volid>): mp0 at /var/lib/felhom,
backup=1, 70G, no mp1; /var/lib/docker and /mnt/sys_drive both real mounts of
its subdirectories via fstab; one df figure and one device id (64519) on all
three paths; reboots 3/3 with the binds surviving each.

demo-felhom (pipeline proof, --force-gitea-golden): fetch_verify succeeding
against the vouched manifest for BOTH artifacts -- 'verified sha256
54e2a4c431daf580... matches the hub manifest' for the golden, a7763d31... for
the agent. 250G single volume, grep -c '^mp1:' = 0, reboots 3/3.

Journey proven on both, endpoint-level: claim -> deploy -> back up -> restore,
with a planted marker returning byte-identical on each box. Ceiling measured
gone: 65 GiB and 233 GiB available to a recovery unit, against 19 and 45.

R-165 -> IMPLEMENTED, not PROVEN-LIVE, on the operator's ruling. B2, which that
row records as the bulkhead's replacement, fired live for the first time and
does refuse per app, delete nothing and alert -- but it is checked only in
captureAllRecoveryUnits while runVolumeDumps writes the bulk unguarded, and its
'the previous unit is untouched' claim was measured false (182,272 B dump
replaced by 2,147,666,432 B under a manifest still dated 06:34:26). -> R-181.

New: R-179 (uninstall leaves NAS network-storage units), R-180 (--archive-storage
not cross-checked against the ACL grant; 403 at step 8/8 after root@pam is
rotated), R-181. Third instance of R-115 recorded (agent 0.120.0 unpublished).

No code written, no version bumps -- this was a runbook.
2026-08-03 09:34:15 +02:00
admin 14d8c00781 docs: R-165 merge built and proven at the bake; R-163 + R-175 closed, R-178 filed
gates / gates (push) Successful in 8s
07-backup-architecture.md gains §7.5.1 (S-1: the contract changed in the same
session): the ceiling §7.5 describes no longer exists for a box built from
golden >= 0.192.0, the bulkhead's replacement is recorded, and R-175 is FIXED
here rather than left standing — the bound is restated as a function of mp1
and scoped to split-layout boxes, naming all three real shapes.

Capability map: new row as IMPLEMENTED, deliberately NOT proven-live, with
the missing leg named — no box has been reinstalled from the golden, and
"the golden baked" is not "a box built from it works".

R-163 CLOSED: the ceiling it recorded stops existing. R-176(a) answered by
P1; (b) WITHDRAWN, since every node is reinstalled rather than migrated.
R-178 filed for the reinstalls, which were not done this session.

CONTEXT S-13 (the variant chosen on measurement; pruning rejected with its
reason) and S-14 (prove first, then vouch — the golden is published but
deliberately unvouched, because vouching is what makes a fresh install pick
up a layout no box has been proven from).

STATUS: plain-language section; both operator questions now answered, so the
waiting-on-you item is cleared. Two older entries trimmed so the page did
not grow.
2026-08-03 07:16:15 +02:00
admin 41dbecb264 docs: R-167 + R-158 CLOSED, R-165 SPIKED, R-174..R-177 filed
gates / gates (push) Successful in 8s
R-167/R-158 shipped and proven live (controller v0.191.x, hub v0.89.0):
two new capability-map rows PROVEN-LIVE with live citations, and
07-backup-architecture.md §7.5's closing claim "nothing warns when an app
crosses the line" is now false and rewritten (S-1: an architectural
contract changed in the same session). §7.5 also gains the caveat that its
size bound is ONE BOX'S, not the fleet's.

Part 3 SPIKE (audits/SPIKE-r165-mp1-merge-2026-08-02.md): M1-M5 measured,
NO layout touched. Three findings the merge session must not re-derive:
"the layout" is not one thing (200G/50G vs 50G/20G vs 16G/8G); mp1 is a
BULKHEAD and not only a ceiling, so after the merge an overflow reaches
/var/lib/docker; the golden fails closed on the split in four places.
D-a's condition (1) is currently SATISFIED — no external box is in the
hub's register, and both demo boxes are Tier 0 and reinstallable.
Recommendation given, choice NOT made — it ends at the operator's ruling.

CONTEXT.md S-11 (D-c's routing, and why R-158's own backup_failed proposal
was overruled) and S-12 (the monitoring landed BEFORE the merge).
STATUS.md gains the plain-language section and the merge decision, with two
older entries trimmed so the page did not grow.

New rows R-174 (closed same session), R-175, R-176, R-177; each ID grepped
free before minting.
2026-08-02 23:56:16 +02:00
admin 8ef92a3fa7 docs: R-172 CLOSED (hub v0.88.0), R-173 filed, session report
gates / gates (push) Successful in 7s
R-172's root cause was not tuning — the WAL/busy_timeout pragmas had never been
applied, because the DSN used mattn/go-sqlite3 syntax against modernc.org/sqlite,
which ignores unknown parameters without an error. Recorded that way so nobody
re-reads it as "SQLite was slow".

R-173 NEW: while establishing who copies hub.db for the WAL change, found
pvc/hub-data labelled recurring-job-group.longhorn.io/default: disabled, with
backup-daily and backup-weekly the only recurring jobs and both on the default
group — so the hub database has no volume-level backup, and it holds every box's
break-glass root password plus the escrow custody records. Filed, not fixed:
whether the exclusion is deliberate is an operator question.

The session report is REPORT-r172-hub-wal.md, not REPORT.md, per the
parallel-session rule — REPORT.md belongs to the controller session that ran
immediately before this one.

It also records, plainly, that a 60-concurrent load test I ran OOM-killed the hub
pod three times against a 256Mi limit. Not the WAL change, and not a test I
should have run against a Tier-2 box; the unit tests already proved the property.
2026-08-02 21:22:14 +02:00
admin 2c35c4204a OPEN-ITEMS: R-172 — false host_stale when SQLite refuses two consecutive host reports
gates / gates (push) Successful in 6s
The hub's /data/hub.db (128 MB) is in rollback-journal mode, not WAL, so a UI
render can block a report write; the hub returns 500 on SQLITE_BUSY without
retrying, and the agent waits its full 15-minute interval rather than retrying.
Staleness fires at 30 minutes, so two consecutive collisions produce a false
host_stale and an operator email for a healthy host. Observed twice on
2026-08-02 while the agent was up 2 days and reconciling throughout.

Pre-existing: 13 collisions in one pod lifetime, first ~3h before that day's
controller work, though a burst of restarts amplifies it.
2026-08-02 20:48:04 +02:00
admin ad28699761 docs: R-157 A / R-170 / R-171 closed — boot recovery finished
gates / gates (push) Successful in 7s
Controller v0.190.0. Docs only here; no hub change, no hub version bump.

- audits/DIAG-bootrecon-drive-absent-2026-08-02.md — NEW. The Part 0 diagnosis,
  including the run that produced a FALSE NEGATIVE and the mechanism behind it
  (the agent re-binds an unmounted drive within ~60s, so the drive gate's startup
  reconcile restarted the apps one second before the sweep looked). Records that
  the write hazard was blocked only by an ACCIDENTAL filesystem permission that no
  code owns and no test pins.
- architecture/02 §0a — the boot-recovery contract (S-1): both gates read desired
  state; the sweep observes a SETTLED fleet and each sample must refresh first;
  nothing is started without asking, fail-safe. Plus the durable warning:
  Manager.StartStack has no gate of its own.
- 00-capability-map — the boot-recovery row, with the repeat count cited per N.5
  (6 of 6 hard resets) rather than a bare PROVEN-LIVE.
- OPEN-ITEMS / ROADMAP — R-157 CLOSED (both mechanisms), R-170 CLOSED, R-171 NEW
  and closed the same session, marked a regression from v0.189.0.
- STATUS.md — the power-cut line moved from "What's broken" to "What works right
  now" with its repeat count; one dated bullet in the change log.
- CONTEXT.md S-13 — the lessons worth carrying: "it didn't happen this time" is
  not a disproof; widening a window makes previously-unreachable overlaps
  reachable; and a settle detector is only as good as the freshness of what it
  samples — the fix's own defect, found live rather than by review.
2026-08-02 20:38:21 +02:00
admin 5c97fbc397 docs: R-166 SHIPPED — the desired/in-flight/observed split (D-b)
gates / gates (push) Successful in 8s
Controller v0.189.0 implements operator decision D-b. Docs only here; no hub
change and no hub version bump.

- architecture/02-controller-module-map.md §0a — NEW, and it is the S-1 contract:
  desired (app.yaml) / in-flight (own marker file) / observed (not persisted),
  with the rule that ties them — never derive one from another. Absent desired
  state means UNKNOWN, never "running". One file, one writer. D-b's binding
  safety rule quoted verbatim.
- 00-capability-map.md — the boot-recovery row now rests on a recorded signal,
  with the three live flows from 9201. The interrupted-operation half is marked
  IMPLEMENTED, not PROVEN-LIVE: nobody killed the controller mid-backup on metal.
- OPEN-ITEMS/ROADMAP — R-166 SHIPPED with both blocking facts and their answers;
  R-157 mechanism B CLOSED and A restated as the whole item; R-170 NEW (the
  drive-backed boot gate still infers a Stop from a container count).
- STATUS.md — the "an app can stay switched off and nothing says so" line
  rewritten to what is actually left: timing.
- CLAUDE.md — end-of-session checklist gains: confirm your own last push's CI run
  went green, BY RUN ID. The failure email is a push signal; this is the pull check.
- CONTEXT.md S-12 — the rulings, and the two lessons worth carrying: a test that
  constructs the thing it should prove the caller constructs is hollow (its
  red-proof will say so), and a field-by-field struct rebuild in a save path is a
  defect on sight.
2026-08-02 18:58:27 +02:00
admin c718aad1bc docs: R-168 SHIPPED, R-29 CLOSED on the demonstrated alarm, R-169 minted
gates / gates (push) Successful in 7s
SPIKE-ci-runner-2026-08-02.md: all six probes with method, measurement and ruling; none
STOPped. P2 (stock image has git but no python3) and P6 (a runner that loses its state
re-registers and orphans the old record) changed the design; P5 (a failed run signals
NOTHING) is why the alarm exists at all.

R-168 SHIPPED with its evidence. R-29 CLOSED — on the demonstrated alarm and not on a green
run, as required: the class it opened is answered at both ends, the hook refusing locally and
CI catching a --no-verify bypass and emailing. R-161 noted: its automatic half now exists for
the STATIC gate, while its original scope, the runtime gate, is deliberately still not
automatic and should stay that way.

NEW R-169 (grep established R-168 was the highest in use): CI can only report, because there
is no gate in the road. Making it blocking needs branch protection plus a PR workflow, both
of which change how the operator works — so it is theirs to decide, and the row states the
cost honestly rather than recommending it.

CONTEXT gains S-8 (CI detects, does not block, and why that is structural), S-9 (a detector
that tells no one is not finished, plus the curl and Cloudflare-1010 traps), S-10 (the runner
is unprivileged because DooPlex is Tier 2), S-11 (CI reproduces the sibling layout).

CLAUDE.md gains the rule earned by red-proofing: a go test -run pattern that matches no test
prints ok and exits 0, and an instrument that can silently drop results is not a measurement.
2026-08-02 16:35:34 +02:00
admin 4707be755c docs: R-94 closed, R-29 leg (a) closed + leg (b) half, R-168 minted
hub/CHANGELOG v0.87.0 + scripts/CHANGELOG gate-enforcement entry. CONTEXT gains S-6 (the
hub renders no host-install version and the gate pins its absence) and S-7 (gates run from
one entry point per repo; reuse_refs_check was fixed rather than the REUSE.md convention,
with both rejected alternatives recorded).

OPEN-ITEMS: R-94 CLOSED all three legs, leg (a) by DELETION with its reason; R-29 leg (a)
CLOSED and leg (b) HALF-SHIPPED with the census result written into the row (13 gates; every
gate a CLAUDE.md names was green, two of the four unnamed were red); R-161 gains its
successor pointer. NEW R-168 (grep established R-167 was the highest in use): Gitea Actions
runner — measured 2026-08-02 as Gitea 1.26.2, Actions enabled on all four repos, 0 runners,
0 workflow runs, 0 branch protections, and the consequence that trunk-based direct-to-main
pushes leave no merge for a status check to gate, so CI here can detect but not block.
BLOCKED on a spike over host-mode vs privileged DinD on DooPlex and whether the workflow can
avoid JavaScript actions.

ROADMAP: R-94 collapsed to its one-liner, R-29 updated, R-168 added.
2026-08-02 15:28:31 +02:00
admin e994bf35d2 STATUS.md: a plain-language operator page, and today's four decisions recorded
Documentation only — no code, no box, no build.

STATUS.md (repo root, 652 words / 67 lines): what works · what's broken ·
what we're working on · waiting on you · changed since. A VIEW of
OPEN-ITEMS.md, holding nothing of its own; not CONTEXT.md, and both files
now say why they stay separate. No R-n is the subject of a sentence —
identifiers are bracketed pointers only.

CONTEXT.md S-5 records the four operator decisions taken 2026-08-02
(D-a … D-d), none of them implemented:
  D-a merge mp1 into mp0 rather than resize it — before any external
      install, and D-c ships in the same step        → R-165
  D-b desired/observed app state in its own store, with the state-store
      safety rule verbatim                           → R-166 (BLOCKED)
  D-c customer fill warning + operator backup-failure alert → R-167
  D-d only DooPlex and Peti's box are protected      → target-selection.md

R-163 RE-FRAMED, not closed: the sizing question is withdrawn rather than
answered; the row survives as the record of the constraint until R-165
lands. R-156's papra referral RESOLVED — deployed nowhere, so the template
fix strands nothing; the docker ps evidence is recorded with its
provenance and its scope limit.

target-selection.md: two protected machines, everything else disposable.
ep0 is no longer Tier 2 but is not scratch (it holds the only off-premises
copy of real customer data) — flagged for explicit operator confirmation.
The demo-box backup-target fence drops from prohibition to stated cost,
because D-d spends that reference anyway.

CLAUDE.md gains an End-of-session checklist carrying the STATUS.md
maintenance rule and "a finding goes in OPEN-ITEMS.md first".
2026-08-02 14:20:29 +02:00
admin 260a8f6e58 register: R-161 ruled and shipped at reduced scope; re-ranked
The operator ruled on R-161 and the runner shipped in app-catalog-felhom.eu
(fd7747d), so the row moves from BLOCKED-needs-a-ruling to REDUCED SCOPE - open.

Both obvious enforcement points were rejected for measured reasons, and the row
now records them rather than leaving the rejection implicit. Controller-side at
template load: rejected because such a check can only read the file, and a static
audit of all 53 templates reports the catalog clean INCLUDING papra - it would
pass on the exact defect it exists to catch, the property being decidable only at
runtime. CI: rejected for now, neither repo has any and there are no users yet.

Shipped instead: scripts/catalog_gates.py, one entry point over all three gates,
non-zero exit on any failure, mandated in the catalog's CLAUDE.md the way
site_gates.py is. The rationale is recorded because it is the transferable part -
of this project's gates, the only ones that ever get run are those with a single
entry point named in a CLAUDE.md; site_gates.py is run and R-29's three orphans
are named nowhere and have stopped nothing.

What stays open is only the automatic half, which is sufficient while ONE person
touches templates - revisit when a second does.

Re-ranked accordingly: R-161 drops from 2nd to 7th, and R-156 is promoted to 2nd,
since R-161 was ranked high precisely because nothing ran the gate and that is no
longer true. The de-ranking is recorded inline with its reason, matching how R-94's
de-ranking is recorded, so a later reader sees a decision rather than drift.
2026-08-02 14:05:13 +02:00
admin b06ea9c877 register: file R-156..R-164 in one pass, ranked; and record what mp1 is actually for
Nine rows into OPEN-ITEMS.md and ROADMAP.md, matching each file's column shape.
R-156 and R-157 had lived only in audit documents - the identical "minted in a
spike doc and never carried across" failure the register already records for
R-153/R-154/R-155, caught by the catalog sweep's own section 8.0 while it was
happening. R-158 was minted by a second session the same day for an unrelated
finding, which is why the sweep's proposals were renumbered R-159..R-162 at filing
time. All nine IDs verified free in BOTH backlog files before use.

Part 0 settled the question the sizing item depended on, by reading:

mp1 is RETENTION, not staging, and neither of the two framings was right. A unit
is the KEPT copy on the app's OWN drive (backup.go:245-255); for an app with no
HDD_PATH the namespace falls back to the system SSD - "the SSD-only system-data
fallback" (appbackup/paths.go:26-27). There is no post-copy deletion: the only
prune is F5 residue-on-old-drives when an app MOVES (backup.go:1053-1112). So mp1
retains the units of driveless apps only - not every app, but not transient
either. Confirmed against the spike: sys_drive held exactly the four driveless
apps and not calibre-web, which had a drive and was still backed up.

A unit is volume tars + DB dumps only, never mp8 userdata
(recovery_unit.go:20-25), so a 1 TB photo library can never overflow one. And mp1
gates the WHOLE chain, not just Tier 1: Tier-2 mirrors the unit "(always)" from
RecoveryUnitPath (tier2.go:302,368) and Tier-3 carries it, so a unit that cannot
be written leaves both with nothing to copy.

Part 2 fired on both triggers - retention, and the fallback undocumented - so
07-backup-architecture.md gains section 7.5. Section 6.1 said a unit lives "on the
app's own drive", which is true and was the whole story only for drive-resident
apps; the no-drive case was undocumented, as was the sizing constraint. 7.5
records the mp0-50G-vs-mp1-20G mismatch, the measured ratios (DB app up to ~2x,
21.1GB -> 40.2GB; file-only 1.00x), and the bound this puts on D5's Lane-1
independence: restorable from the drive alone only while the unit still fits -
about 19 GB file-only, about 10 GB DB-backed. No number proposed; the ratio is the
operator's ruling (R-163).

R-159/R-160 marked SHIPPED only after verifying the template changes are in
app-catalog origin/main, and R-156's gate likewise (check-volume-persistence.py
present). papra is NOT fixed - referred - so R-156 stays open on that one app.

Ranked, with one line of reasoning each: R-157 first (an app can stay down
indefinitely with mechanism B silent on every channel), then R-161 (the gate
exists and nothing runs it, which is why R-156's class recurs - R-29's record is
three orphaned gates and one enforced), R-156, R-163, R-158, R-164, R-162.
2026-08-02 12:33:12 +02:00
admin e9a74a0019 docs: remove a gate criterion that could never pass, and close three register rows
PART 1 — the release gate.

G7 required the packaged .deb to sha256-match the one built from committed source. That is
unsatisfiable BY CONSTRUCTION: dpkg-deb stamps the build time into every archive, so two builds of
byte-identical source differ. It was already failing when the 1.26.1 release ran it. A criterion
nobody can satisfy gets waived once and read as advisory ever after — which is how R-29's shelf of
never-run gates was built. Sub-clause dropped, reason recorded in G7's own note the way G6's
amendment was, so a future reader can restore it if SOURCE_DATE_EPOCH ever makes it meaningful.

RULING ASKED FOR — is payload integrity covered by G9 alone? NO, and G9 is widened rather than a new
criterion invented. The package ships TWO payload files (build-deb.sh:54-55); G9 checked only the
script. The systemd UNIT was covered by nothing: G7 covered the container, G8 covers the postinst
behaviourally, G13 covers directory presence. The unit is not incidental — its After=, its
ConditionPathExists= and its Restart= decide WHEN AND WHETHER day-0 runs at all, so a drifted unit
would have shipped silently. Same shape as the /etc/felhom miss that G13 exists to prevent: a check
that proved the thing present and said nothing about what it depended on. The check passes today.

G13 moved to sit after G12 — it was minted late and left between G10 and G11.

PART 2 — register dispositions. BASELINE DISCREPANCY, reported rather than worked around: only R-128
had a row. R-154 and R-155 had NO row in either file — minted in a spike document and never carried
across, which is R-123's class, not the drift the task described. Rows created, closed, with the
reasoning, because in all three cases the reasoning is the durable part:

  R-128 closed by CORRECTING a false claim, not by making the assertion real — the coupling does not
        exist and asserting it would invent a constraint. Flagged so nobody 'restores' it.
  R-154 closed with the measurement and where it now lives in pushed source.
  R-155 NARROWED, not deleted — unchanged for FELHOM_MENU=single, inapplicable to release. Flagged so
        the guard is not later removed wholesale on the strength of 'R-155 closed it'.

Documentation only: no code, no build, no ISO, no upload, no box touched.
2026-07-31 21:31:54 +02:00
admin 9e079c7883 RECON: a Felhom-issued subdomain works in the product — the blocker is Cloudflare edge-cert depth
Question A: YES, no code change. customer.domain is a trimmed string with no
UNIQUE, no CHECK, no format rule (store.go:114, configs.go:673), copied verbatim
into controller.yaml (configgen.go:48), and every one of its 30 consumers on the
box interpolates it without parsing. Zero hits for registrable/eTLD/publicsuffix
across both repos. Nothing creates DNS records (zero hits for dns_records) — the
two Cloudflare clients are WAF-only. And the zone-ownership assumption is a
SWITCH, not a requirement: traefik.yml.tmpl selects DNS-01 when cf_api_token is
set and HTTP-01 when it is empty.

The real blocker is Cloudflare, proven live: the edge certificate covers exactly
one wildcard level (SAN = demo-felhom.eu, *.demo-felhom.eu), so a two-label
hostname — which a per-tester subdomain forces — gets "tls alert handshake
failure" and no peer certificate at all. That makes Advanced Certificate Manager
a prerequisite of the separate-domain plan, not an optional extra. Whether ACM is
available on the account could not be established read-only: the only Cloudflare
tokens in reach are the Zone:DNS:Edit tokens on the demo boxes, which the fence
forbids using.

Question C, measured rather than reasoned: r.Cookie returns the FIRST match and
never tries the others (BOGUS+real = 302, real+BOGUS = 200), so a tossed cookie
wins outright — DoS and confusion, not takeover, since it fails closed on
mutations. CSRF is a single choke point (server.go:256) and the token carries the
whole load against a same-registrable-domain attacker. But it is SKIPPED entirely
when no session cookie is present, which with browser-cached Basic auth is
cross-origin CSRF on every mutating route (R-135).

Agreeing with the separate-domain recommendation, with the caveat the brief asked
for: it is necessary but not sufficient. It does not solve Question D, because
that is a shared-zone problem and the new domain is a shared zone.

Filed R-133..R-138: duplicate domains accepted; hub/controller zone-resolvers
disagree on depth; CSRF skipped on the no-cookie path; __Host- rename (one line,
preconditions verified met); geo-WAF rules zone-scoped and non-namespaced (four
cross-tenant faults, blocks shared-zone onboarding); shared-zone cf_api_token is
a zone-wide DNS-write capability on a customer's box.

Nothing created: no customer, DNS record, tunnel, route or code change.
2026-07-31 09:00:11 +02:00
admin b4edc087fa Tester gate: golden re-baked to 0.188.0, fresh-install proof PASSED — a fresh box is safe to hand to a tester
§7.2 answer: YES. A real day-0 from the existing v1.25.0 ISO reached a claimable,
app-serving box in ~10 minutes unattended, and an app's data came back from the
drive with the guest's app.yaml gone — proven readable by the application over
its own TCP path, with a discriminator (PRE-BACKUP row = 1, POST-BACKUP row = 0).

Part 0: NO ISO rebuild needed, verified against the ISO on disk rather than from
source. It bakes only felhom-bootstrap.sh, its unit and the secret-free pairing
env (full-base64 match, 1 hit each) and 0 hits for any installer, controller or
golden marker. The installer is fetched at run time; the live URL is byte-identical
to repo HEAD (v1.22.0, six days newer than the ISO) and the fresh box ran it.

Part 1: baked 0.188.0 rather than the brief's 0.187.0 — 0.187.0 lacks D5, which
is the very claim Part 2 step 6 tests. Published (404 pre-gate with a 200 control;
anonymous download, 649310288 bytes, sha match), vouched, and consumed by a real
box. R-120's gate exercised BOTH ways: 0.185.1 refused with no write, 0.188.0
allowed — evaluated, not silently skipped.

Part 3: RUNBOOK-manual-build.md cited a "RECORDED" qemu line that is itself
labelled reconstructed and whose source says it was never saved. The real
invocation is now captured from this bake as §4.0, with the bake/publish/teardown
steps; the old entry is marked SUPERSEDED.

Teardown all three layers, hub disposition stated: VM destroyed, scratch storage
removed with space returned exactly, customer sess-g DELETED via full cascade.
sess-f deliberately left (R-131) with its command recorded.

Filed, none fixed: R-128 (false ISO_VERSION invariant comment), R-129 (demo-hp's
"no baked SSH key" is stale — key auth works), R-130 (HARD_MIN_LVM_GIB warns and
proceeds), R-131 (fourth orphaned scratch customer), R-132 (curl's %{redirect_url}
printed the hub operator password into a transcript — HUB_PW needs rotating).
2026-07-31 08:27:36 +02:00
admin 1956e5d390 hub v0.84.0 — break-glass console credential on the host page
The credential existed and was not reachable when it was wanted. Every box has
had a strong random root@pam password since TASK G1, vaulted in the hub at day 0
and used for real during the sshd incident — but the only way to read it back was
a hand-written curl carrying the global operator key, a secret kept out-of-band.
In practice the PVE web console on a demo box felt locked.

The host page grows a Console access card: presence + username + set_at by
default, Reveal fetches the plaintext on demand for 60 s with a Copy button.
Masking clears the JS variable, and also fires on a second click and on
visibilitychange. A host with nothing vaulted says so, and says why.

The secret is NEVER rendered into the page, and that constraint shapes the
change. The render path uses a new store.GetHostRecoveryMeta whose struct and
SELECT both omit the secret column, so it is structurally incapable of carrying
one. The plaintext crosses the wire only in the response to POST
/hosts/{id}/reveal-recovery-credential (Cache-Control: no-store, CSRF-gated at
the ServeHTTP level; POST precisely so that gate applies and so no secret is
retrievable by URL alone). Deliberately NOT the customer page's data-secret
widget, which embeds the plaintext on every load.

A delivered reveal writes one recovery_credential_revealed event on the host's
customer timeline (info, source hub, Hungarian) via SaveEvent alone — no
dispatcher, nobody emailed, the log_tail_requested shape. Two reveals write two
events: the register records accesses, not states. A 404 is not an access. An
unbound host reveals fine and writes no event; the [INFO] hub line, carrying the
username and a length only, is then the record.

The global-key API path is untouched by design — it is the route for when the
hub UI itself is broken, and coupling it to the session layer would delete the
independence that makes it a fallback.

Recorded as a real trade: the hub session password alone now unlocks console root
fleet-wide, where retrieval previously also needed the global key. Accepted for a
single-operator, HU-geo-fenced hub that already stores these passwords in
plaintext at rest (CONTEXT.md ruling S-4). The plaintext-at-rest half is filed as
R-133 — every hub DB backup is a fleet-wide console-credential dump.

Tests 550 -> 559; four red-proofs (page leak, audit event, CSRF gate, route
order) each run, observed failing, and reverted. The route-order proof is a seam
test driving ServeHTTP: a handler-level test cannot see that defect, because the
handler is correct and simply never runs.
2026-07-31 08:19:36 +02:00
admin 0a9bd3829d D5 SHIPPED: Tier-1/2 restore no longer depends on the whole-guest tier
Records controller v0.188.0 across the four coupled artifacts.

07-backup-architecture.md is the owning doc:
- new 7.4 = the recovery chain AFTER D5 (7.1 leg 1 superseded; leg 2,
  the living-app dependency, explicitly unchanged so this is not read
  as more than it is)
- 7.3 collapsed to history, with the correction that the target as
  written (data_key-only) was tested in Part 0 and rejected
- 3 records that the two-lane split is now real, not just intended
- matrix rows 3 / 3c (new) / 13; 10.1 D5 itself shipped

Also: new capability-map row, D5 collapsed in ROADMAP + OPEN-ITEMS,
and R-127 filed in both (data_key flag unreliable; O4 can regenerate a
DB password that no longer matches the restored data directory).

The audit is named D5-drive-alone-restore rather than "...secrets..."
because .gitignore blocks *secret* -- a guard worth respecting, not
forcing past.
2026-07-30 16:58:06 +02:00
admin d42d90fed7 R-108 CLOSED — D5's precondition is met (controller v0.187.0)
Four-artifact update per the coupling rule, plus the audit.

07-backup-architecture.md: §10.1 retitled CLOSED with the ruling and the D5
sentence; the FileBrowser network-share row flipped YES->NO, closed at the
PLACEMENT rather than at the bind; the exposure chain annotated with the fifth
surface (decommission-with-migrate guarded only its source) and the correction
that the boundary is the deploy POST, not the dropdown; §7.3 retitled UNBLOCKED;
register row collapsed; open question F answered.

00-capability-map.md: new §D row PROVEN-LIVE, with the un-exercised legs named —
the deploy-POST and decommission refusals are unit-tested, not live-fired.

OPEN-ITEMS.md: R-108 dispositioned; D5 given its OWN row as READY/UNBLOCKED (it
had existed only inside other rows' prose — the R-123 thread-loss pattern);
R-126 registered.

ROADMAP.md: R-108 collapsed to a shipped one-liner; R-126 added.

R-126 filed not fixed: a .fab bundle (plaintext secrets, optional password) can
be exported ONTO a NAS. Split out of R-108 rather than folded in — it is an
explicit customer-chosen export destination, not a browsing surface reaching a
backup tree, so it was never part of D5's precondition.

Live evidence: same-box before/after on demo-felhom through the real authenticated
endpoint, the network-specific refusal on demo-hp, non-effect verified in the
registry, and R-67's share-root bind diffed byte-identical across the deploy.
2026-07-30 14:21:44 +02:00
admin 70f84941d4 R-106/R-109 audit + registers: shipped at agent 0.118.1, plus R-125
Adds the full audit: Part 0's three answers, the pre/post recipe for both boxes,
the on-disk proof that `local` froze at the 2026-07-28 target move while
felhom-backup kept running, all seven red-proofs, and the three publish
observables.

R-125 filed: v0.118.0's R-106 half shipped INERT. Two tests ran the real
Collector.Collect() but both injected a fakeObserver, and the break was one layer
below in mergeConfig, which dropped the pbs namespace. The recipe still said
"root" — now with namespace_state "resolved" beside it, confident and wrong.
Caught by live validation, not by the green suite. Fixed in 0.118.1; filed for
the doctrine point that a production-path claim must name the seam it injects at.
2026-07-30 13:28:10 +02:00
admin acfc2b7e95 R-109 + R-122: the recipe assembly stops dropping sections (hub v0.83.0)
AssembleDRRecipe's hostHalfShape/appHalfShape are ALLOW-LISTS, not the
forward-compat their comment advertised: a section an emitter adds is silently
discarded until it is named in both the shape struct and AssembledRecipe. No
error, no log, no failing test.

R-122 (found this session): that already happened and shipped. The controller
has emitted offsite_restic since fork-4 — the offsite recovery LOCATION — the
hub stored it for all three real customers, and appHalfShape never listed the
key, so no delivered recipe has ever contained it. It stayed green because the
fixture drAppHalf is hand-written and omits the field.

R-109: the agent's new backup_target is a new top-level host-half section and
would have been dropped identically, making the fix read as shipped while
changing nothing an operator can see.

3 tests built on halves read verbatim out of the live dr_recipe table, plus
2 red-proofs (each mutation asserted to have landed). vet rc=0, suite rc=0, 17 ok.

Registers: R-106 + R-109 dispositioned; R-105/R-106 were READY in ROADMAP with
no OPEN-ITEMS row (→ R-123, registered); R-124 filed on the "root" spelling.
2026-07-30 13:13:56 +02:00
admin 3d504d58c8 docs(R-117): CLOSED — proven live on demo-hp; R-121 filed for agent-on-box drift
R-117 row → SHIPPED + PROVEN-LIVE (agent v0.117.0), with the full validation in
audits/R117-v0117-2026-07-30.md.

Both dead states detected on real hardware through the shipped predicate:
  RETURN   raw 8:32 /dev/sdc  | bind 8:16 shutdown      → stale-device,        usable false
  IN-PLACE both 252:11 emergency_ro, raw unit active    → filesystem-aborted,  usable false
  healthy                                               → live
340-497us per call. No block I/O proven by strace (only /proc/self/mountinfo,
0 statfs) — the Part 1 CLAUDE.md fence applied to its own first consumer. No
regression through the real pipeline: the live backup-target drive reads
bound_under_parent=True via GET /disks with the controller's own credential, with
32 gate lines in 3 min as the positive observable and zero spurious transitions.

The ruling asked for in §2.2 is recorded in full and flagged for overrule:
Aborted must NOT self-heal. A re-bind lands on the same dead superblock and the
call site runs every 20s, so repairing would be an infinite silent retry that
masks the state. It surfaces instead. No operator decision was taken quietly —
the reasoning is that it routes an already-broken state into the existing gate,
event types and Hungarian copy, so no new concept reaches the customer.

R-121 filed: a box's installed agent can sit releases behind the vouched one and
nothing notices. demo-hp ran 0.113.0 against a vouched 0.116.0 through the whole
R-116/R-117 arc. Confirmed at source that R-120's gate cannot catch it — it
compares goldenVer against NewestReportedControllerVersion(), i.e.
golden-artifact vs fleet-CONTROLLER. MinAgent is protective, not an alarm, and
0.113.0 equalled the floor. Fourth instance of the drift family.

Also filed: R-117g (an aborted filesystem is never cleared automatically by
design, so it alarms until a human acts, with no guided recovery) and R-117h
(StablePathForRaw hardcodes the parent, so the repair path cannot be exercised on
hardware without writing into a live customer guest's namespace).
2026-07-30 12:42:26 +02:00
admin 37515cda7c docs(R-117 Part 1): a health check issues no block I/O — and narrow one R-116 claim
Two record items, banked before any Go file is opened.

1. CLAUDE.md gains a standing rule beside the seam-wiring rule: a health check
   issues no block I/O. A probe that touches a wedged device enters
   uninterruptible sleep, survives SIGKILL, and cannot be recovered until the
   device returns or the host reboots — so `systemctl restart` hangs too. A
   timeout protects the caller's control flow and nothing else. Liveness is
   decided from /proc and kernel state.

   Measured in the R-117 spike §6.3: D state 3m50s after kill -9; a buffered
   write with no fsync blocked too (O_CREAT needs journal access); statfs and
   getdents returned HEALTHY on a namespace that EIOs every byte.

   Repeated as a one-line pointer in felhom-agent/CLAUDE.md, because health
   checks are written in that repo and felhom.eu/CLAUDE.md does not load in an
   agent-only session — a standing rule that does not load where it binds is the
   inert-seam shape applied to a rule.

2. The R-116 row gains the clause the spike recommended but did not apply. Its
   verdict stands and every input to the pairing fix is configuration-derived.
   But the over-correction window's degraded:false was read off a drive whose
   bind was dead, so it evidences "the gate did not over-fire", not "the drive
   was healthy". The two RETURNED lines remain a genuine positive observable, so
   rule 3 is still satisfied. Nothing else about the row changed.
2026-07-30 12:10:10 +02:00
admin e70b5feebe docs(R-117): the hang case measured — an I/O probe turns a wedged drive into an unkillable agent
Completes the spike once the venue came back. Q4's hang case and teardown are
now measurements, not plans.

Against a dmsetup-suspended device (I/O queues instead of returning EIO):

- P1 (devno compare) and P2 (ext4 abort flags) completed in 364us / 206us.
  They read /proc, so no block device is involved.
- statfs and getdents completed and reported HEALTHY — on a wedged device they
  do not even hang. R-117b confirmed in a second failure mode.
- EVERY probe that touches the device blocked, including a buffered write with
  no fsync: the O_CREAT metadata path needs journal access
  (wchan=do_get_write_access). There is no cheap-and-safe write probe.
- The blocked process survived SIGTERM AND SIGKILL (stat=D,
  wchan=folio_wait_bit_common, still alive 3m50s after kill -9) and died only
  when the device was resumed. So `systemctl restart felhom-agent` would hang,
  leaving the agent unrecoverable until the device returns or the host reboots.
  The thread count does not reveal the leak (5->5, 5->6).

Filed as R-117f. A timeout protects the caller's control flow and nothing else,
so "the fix must issue no block I/O" is now a fence rather than a preference —
the thread-leak hypothesis the probes were built to test turned out to be the
weaker half of the result.

Teardown done, all three layers: guest 9301 destroyed, r117scratch removed, both
dm and both loop devices gone, scsi_debug unloaded, local back to 37.02% against
a 37.00% session start. Fences re-verified AFTER teardown: 9201 running,
drill-r50 stopped, local-lvm 38.84% byte-identical, felhom-backup content
unchanged, live /mnt/felhom-drives intact with both submounts, agent active.
Layer 3 genuinely empty — 9301 had no NIC and ran no controller.

Trap recorded: a suspended dm device must be resumed BEFORE any umount, or the
teardown blocks on the same uninterruptible sleep.
2026-07-30 11:49:59 +02:00
admin c949389c95 docs(R-117): spike — the mechanism, a recipe, and a steady-state half nobody had looked for
Both halves of the R-113 conjunction are path-presence tests: GuestSeesMount
(intermediary.go:276) and isHostMountpoint (:394) compare field 5 of a mountinfo
line and never read field 3, so neither can see that the bind and the raw mount
name different devices. Measured BoundUnderParent=TRUE over a namespace that
EIOs on every read and write.

Reproduced 3/3 on a purpose-built scratch LXC on demo-hp; predicates evaluated
by a throwaway probe calling the real localapi code from d4eb259.

Three results that change the shape of the fix:

- Q7: a bind can die in STEADY STATE with no detach/return cycle. The gate
  produces no action and nothing is emitted on any channel. A Return-branch fix
  cannot reach this half, and a devno comparison does not detect it.
- Q6/R-117d: AttachDrive's normalize leg already performs the repair, and three
  call sites already invoke it - including the controller's Return branch before
  it restarts apps. All defeated by one early return at :235. Unblock the
  existing path; do not add a new one.
- Q1: the device-node change is a CONSEQUENCE, not a precondition. The stale
  bind pins the dead superblock, forcing the returning device onto a new number.
  Control test: released, the letter is reused.

Not established: the hang case. Venue and probes built, run lost to a site
internet outage; the thread-leak hypothesis is not claimed as a result.
Teardown of the spike venue is owed - commands in the findings doc; nothing
fenced was touched and no hub-side record was created.
2026-07-30 11:38:22 +02:00
admin 29bcfeb214 docs(R-120): CLOSED on both halves — golden current, and the class has a gate that refuses
Half 1, the artifact: golden 0.186.0 baked, published, vouched, and proven on a REAL
day-0 on demo-hp (not the fixture, per the rule committed in Part 1). With the target
detached, the fresh box's endpoint returned the TargetAbsent copy -- "A rendszermentés
meghajtója nem érhető el — amíg vissza nem csatlakoztatod..." -- with offer_path
absent entirely. The day-old read on the 0.185.1 golden had returned the false
system-disk message plus an offer of the other drive. That is the customer-visible
defect closed.

Half 2, the mechanism: operator ruled REFUSE, shipped as hub v0.82.0 and DEPLOYED.
Proven live by re-attempting the original mistake -- vouching the stale 0.185.1 golden
now yields HTTP 303 flash=golden_behind_fleet plus [WARN] artifact vouch REFUSED, and
the manifest reads back unchanged at 0.186.0. Refused AND unwritten, against the real
fleet signal rather than a unit fixture.

Recorded on R-29's audit list as the first ENFORCED gate beside its three orphans, so
the contrast is kept rather than lost. The orphans are unchanged -- this proves the
pattern is available, not that the backlog moved.

Teardown all three layers: VM 9402 purged, r120-images removed with the space measured
back, hub layer gate-blocked on ONLINE with the command recorded. Last session's sess-e
was deleted this run, discharging its recorded layer 3.
2026-07-30 10:48:16 +02:00
admin 1a68b53b06 hub v0.82.0 (R-120): the vouch path REFUSES a golden the fleet has already outrun
The golden's version IS the controller it bakes (build-golden.sh:345 defaults
GOLDEN_VERSION to the controller tag), so a golden behind the newest deployed
controller means every FRESH install lands on stale application code. On the R-120
occurrence that stale code shipped a customer-facing falsehood: a box from the
0.185.1 golden told a customer whose backup drive had fallen out that the backup was
on the same disk as the system -- false, the drive was gone -- and offered a
different drive as the remedy.

WHY A GATE, NOT A REMINDER. The gap has opened three times: R-111 (golden's agent 17
releases behind), R-115 (agent built and deployed, never published), R-120 (this).
The first two were closed by re-baking and remembering; remembering then failed
again. R-29 is the standing proof that a check nobody runs is worse than none because
it reads as coverage -- hostinstall_gates.py sat RED and uninvoked across three
version bumps and hub_confirm_gate.py has never run at all. So the property that
matters is not whether a check exists but whether it BLOCKS.

- Wired into handleSetArtifacts (internal/web/configs.go), immediately before the
  only write, on the sole UI path to SetArtifactManifest -- it runs on every vouch
  without anyone choosing to. A script in scripts/ would have been a fourth orphan.
- It REFUSES (operator ruling, 2026-07-30), with a flash naming the remedy.
- Signal: store.NewestReportedControllerVersion() over reports.controller_version,
  SEMVER-compared in Go -- MAX() in SQL ranks 0.99.0 above 0.186.0, a pair this
  fleet has shipped. No outbound call, no new credential.
- Fail-open in exactly two deliberate cases: an empty golden field (clearing the
  manifest is legitimate) and an unknown fleet version (a new hub must vouch its
  first golden).

NEAR-MISS RECORDED: the first draft read guests.controller_version, a column that
exists in the schema and that NOTHING writes -- it would always have seen "" and
failed open, i.e. inert, this gate's own failure shape. Caught by grepping for a
writer before trusting the column.

Blind spot stated rather than papered over: a controller no box has ever run is
invisible to this signal. Not the failure that has bitten -- all three instances were
deployed-newer-than-baked.

4 tests through the PRODUCTION handler over httptest, never an injected seam. The
refusal asserts both the flash and that the manifest was NOT written, because a gate
that redirects and saves anyway reads as enforcement while providing none. Red-proof:
deleting the block makes the stale golden vouchable and both assertions fail.

ROADMAP R-29's audit list now records this as the FIRST enforced gate, so the
contrast with its three orphans is kept rather than lost. The orphans are unchanged.

Suite rc=0 read separately from this commit.
2026-07-30 10:42:43 +02:00
admin 49b627684c docs(R-120): golden rebaked to 0.186.0, published, vouched, proven on a real day-0
The golden baked controller 0.185.1 -- confirmed from the golden's OWN record
(drill/bake-0.185.1.log:1 and :330) and from build-golden.sh:345, which derives
GOLDEN_VERSION from the controller tag. 0.185.1 predates R-114 + R-112, so every
freshly installed box told a customer whose backup drive had fallen out that the
backup was on the same disk as the system (false) and offered a different drive as
the remedy.

Baked golden 0.186.0 from main's controller in the DooPlex bake fixture: overlay2 OK,
3 mounts included, FATAL 0, exclusions 0, 618 MB, upload HTTP 201, GOLDEN_SHA256
b760ac6a33e70700..., token-leak grep 0, GL-1 teardown with drill.qcow2 back to
virgin.

Three observables, quoted as returned: PUBLISHED (anonymous GET -- what the installer
does -- 200 / 648930639 bytes / sha identical to the bake); VOUCHED (manifest read
BACK, not the 303); RESOLVED BY A CONSUMER (Artifact manifest served for customer
sess-f, golden=0.186.0). Floor NOT touched per publish-train rule 2 -- it is a
separate form and min_controller_version still reads 0.156.0. MinAgent left 0.113.0
because 0.186.0 declares it unchanged.

Proven on a REAL day-0 on demo-hp, not the fixture, per the rule committed in Part 1:
VM 9402 from the v1.25.0 ISO -> Controller elindult (0.186.0), box confirms
felhom-controller:0.186.0 + agent 0.116.0. A fresh box now runs 0.186.0 where it ran
0.185.1.

The procedure was NOT unwritten: RUNBOOK-manual-build.md:101-115 documents it and
build-golden.sh carries its own usage and publishes to Gitea itself. One
documentation-integrity finding: that runbook says to use the RECORDED qemu line and
not reconstruct, while the line it cites is itself labelled reconstructed, the
canonical one never having been saved.

NOT done and not claimed: the TargetAbsent/empty-offer_path endpoint capture (the
claim gate runs before auth with no Bearer escape -- R-119's fourth instance), and
the Part 3 mechanism, which awaits the operator ruling. Recommendation and exact
wiring recorded in the audit rather than built.

VM 9402 + r120-images + customer sess-f retained pending that read, with teardown
commands recorded. Previous session's sess-e layer-3 is now DISCHARGED -- it aged to
STALE and the cascade completed, full residue purge logged.
2026-07-30 10:33:03 +02:00
admin 772956d214 docs(R-116): CLOSED — proven live; capability row F to PROVEN-LIVE; R-120 filed
The events leg the previous commit reported as not-reached is now done. The operator
relayed the claim code (the only route: bcrypt-hashed hub-side, emailed only), the
two storage paths were registered through the real POST /api/storage/register, and
the cycle ran on the fresh box:

  07:20:04  backup_target_absent   (error)  Cel meghajto   <- TARGET, specific
  07:22:34  backup_target_restored (info)   Cel meghajto   <- its matching pair
  07:24:04  storage_disconnected   (error)  Adat meghajto  <- NON-target, generic
  07:25:34  storage_reconnected    (info)   Adat meghajto

All four at the hub; gate fired in 3 s. Two matched pairs, correctly discriminated
-- and discrimination is proven NON-trivially for the first time, since both prior
runs had the target itself emit the generic event. Over-correction passes on a
positive observable, with two RETURNED lines proving the gate was ticking.

00-capability-map row F: PARTIAL -> PROVEN-LIVE with the evidence and the caveat.

R-120 filed: the golden bakes controller 0.185.1, which PREDATES R-114 + R-112, so
a freshly installed box shows the customer the WRONG absent-target message --
observed live on the drill box: the generic "the backup is on the same disk as the
system" copy (false; the target is a drive that vanished) plus an offer of the other
drive as the remedy. That is E2D 5.3's exact payload, still reachable on any new
install. R-115's class one layer up -- R-111 closed by re-baking the golden, 0.186.0
then shipped, the golden did not move, and the gap reopened silently; this time the
stale artifact carries a customer-facing falsehood in exactly the state R-116 now
alarms about correctly.

Teardown recorded for all three layers, hub layer gate-blocked with the command.
2026-07-30 09:34:08 +02:00
admin 315c469fc8 docs(R-116): v0.116.0 proven live at the payload layer; events leg blocked on an emailed claim code
audits/R116-v0116-2026-07-30.md + the R-116 register row.

WHAT PASSED, on real hardware. Agent 0.116.0 published (independent registry GET
verified the bytes), vouched, and installed UNAIDED by a fresh box -- "Artifact
manifest served for customer sess-e (agent=0.116.0 golden=0.185.1)", host
sess-e-5d4427 ... 0.116.0 ONLINE. Real day-0 on a nested PVE on demo-hp (per
runbooks/target-selection.md, which sent this run there rather than to the DooPlex
fixture the previous run used), both drives enrolled through the real endpoints,
device loss a real hot-detach.

Captured live, absent state: the target is now ONE row carrying backup_target:true
AND guest_path:/mnt/felhom-drives/cel with mount_path:"", so
isTarget[/mnt/felhom-drives/cel] = TRUE -- it was false through v0.115.0. RETURNED
gives true as well, so the pair matches. All three guards pass from the same
payload: R-114 preserved (no row combines the flag with a non-empty mount_path),
no over-correction (bound_under_parent:false), and discrimination at the payload
layer (the non-target carries the flag on no row) -- the thing neither prior run
could show.

WHAT DID NOT HAPPEN, and is not claimed. No backup_target_absent or
backup_target_restored event was observed on the wire. planDriveGates iterates
registered StoragePaths and the drill controller has none ([WARN] Storage paths:
no storage paths registered); every storage route answers 401 "dashboard not yet
claimed". The claim code is bcrypt-hashed and emailed-only, and
handleSelfBindLinkSend (selfbind_mint.go:139-161) renders a flash and never the
token, so no operator-side route exists. A gen-2 code was re-sent; the drill VM,
its storage and customer sess-e are DELIBERATELY RETAINED with teardown commands
recorded, so the leg finishes without a rebuild. Reported as not-reached rather
than as a third trivial pass.

R-119 filed: the claim gate makes drive-gate legs unreachable to CC by design, and
has now stopped three sessions at the same wall -- needs a ruling (operator-scoped
test affordance, or a documented prerequisite step), not a fix.

R-117 reproduced on real hardware with a read/write probe (EIO both directions
while /disks reports attached + bound_under_parent:true) and §5 records how it
colours the reattach leg. R-118's symptom vanishes incidentally on this one row;
R-118 is NOT fixed.

sess-c and sess-d verified GONE (404, absent from both tables) -- cleared by the
operator using the previously recorded commands, not by this session.
2026-07-30 09:13:00 +02:00
admin d56e395a2a docs(R-116): isolate the mechanism from the real /disks payload; file R-117 + R-118
The absent-state /disks payload was captured on a genuine device loss, after a
present-drive control run proved the query works (Part 5's three attempts failed
on token extraction, and its control returned 0 rows).

The answer is theory #1 -- "the registry-union row writes false" -- which was
raised, declared wrong and retracted. The retraction was the error.

Absent state returns 4 rows, not 3. The drive appears twice and the two facts the
controller needs sit on different rows: the Observe row has backup_target:true but
mount_path:"" and guest_path:"", so it contributes no key to driveTargetByPath;
the registry-union row owns /mnt/felhom-drives/<name> and omits BackupTarget from
its struct literal (disks.go:301-306) => false. The union row is not deduped
because seen is keyed on MountPath (:290-295), the one field the absent state
empties, and its own MountPath comes from the systemd .mount unit FILE
(registry_known.go:40-75), which never reads the mount table.

Theory #2 (the basis of the shipped v0.115.0) is false on both halves; #3 is false
too. v0.115.0 is provably inert: StablePathForRaw("") returns "".

Also files the read path verbatim -- the token plaintext lives only in
bootstrap.json on the Proxmox host; the agent's store keeps hashes only.

New: R-117 (READY M, outranks R-116) -- a returned drive's guest bind is a DEAD
mount (EIO both ways) while /disks reports attached + bound_under_parent:true, so
the gate restarts the customer's apps onto it and reports healthy with no alarm.
R-118 (READY XS) -- an absent drive's union row advertises the root filesystem's
capacity as its own.

Docs only. No code written, nothing built or published; v0.115.0 untouched.
Both demo boxes read-only; drill fixture restored to virgin.
2026-07-30 08:06:21 +02:00
admin c3ce4c7b20 R-116 Part 5 FAILED: the fix shipped, C5 still fails, mechanism NOT isolated
A fresh box running the fully shipped stack -- agent 0.115.0 from the Day-0
manifest plus controller 0.185.1 from the vouched golden, no hand-deploy -- still
fired the GENERIC storage_disconnected on detach and the SPECIFIC
backup_target_restored on return. backup_target_absent count 0. Identical to
Session C. The v0.115.0 fix changed nothing observable.

Part 4's three positive observables were all obtained before the run (registry
newest 0.115.0, hub vouches 0.115.0, felhom-pve running 0.115.0 clean), so the
publish step forgotten twice was not forgotten a third time, and the box
demonstrably installed the fix under test.

Discrimination FAILS: the target itself produced the generic event, so the two
cannot be told apart regardless of the non-target leg -- which was therefore not
staged. Reported as a fail, not as Session C's trivial pass.

Over-correction guard PASSES: 0 ABSENT lines with the drive present, target
degraded:false.

THE HONEST PART. The fix targets a shape that does not occur live, and which
shape does occur is NOT ISOLATED. With the drive detached PVE reports the
storage inactive with zeroed fields -- a shape the unit fixture did not model.
Three attempts to read the real /disks payload failed on token extraction across
the ssh -> guest -> container layers, and a present-drive CONTROL query also
returned 0 rows, proving the query was broken rather than the payload. Without
that control this run would have recorded a third false mechanism, after "the
union row writes false" (wrong, corrected yesterday) and "no row carries the
guest path" (unverified). The leading hypothesis -- an inactive storage reaching
Observe with an empty MountPath, so StablePathForRaw returns "" -- is consistent
with the pvesm output but is NOT evidence and is recorded as such.

Next session's first job is a working /disks read, with a present-drive control
run FIRST, before any further code.

agent v0.115.0 is published, vouched and INERT. Not reverted: reverting is
itself a change, the runbook forbids fixing mid-run, and the code is tested and
harmless.

Capability-map row F stays PARTIAL, now citing the re-test.
Teardown clean: pvesm status after == before (local-lvm 38.83%), guest 9201 and
drill-r50 untouched. Customer sess-d pending the usual ONLINE-ages-to-DOWN gate.
2026-07-30 07:15:31 +02:00
admin 952ebf4862 Record work, banked first: shrink the E-2d row, create the missing capability-map rows
Unconditional and three sessions overdue, so it commits before any code is
touched — E-2d itself stopped at Phase 0 and banked nothing.

E-2d row: 822 words -> 121, and the contradiction resolved. Its State read
CLOSED — PARTIALLY PROVEN while the cell's final sentence read "This row stays
OPEN only for the residue"; a reader could not tell which. It is CLOSED, with
R-116 the single named open leg.

Nothing unique was binned. Three facts existed ONLY in that cell and are moved
into audits/E2D-fresh-vm-2026-07-29.md as a new §1a: the local-lvm fence figures
with the 888 GB nvme alternative, the exactMount subdirectory caveat and why the
subdirectory is nonetheless the safe placement (no durable_id collision), and
the ISO/PAIRING -> DIRECT fall-through derived at source with its line
citations. drill-r50's blocked status was already in both audits.

Capability map: it had ZERO rows for the backup-target work — grep gives 0 hits
for backup_target and one for "E-2" that is a campaign date string. Three
scenario rows added, at today's honest status, not the value hoped for later:

  C. Protection & recovery — installer Case A/B, DEGRADED recorded not hidden
     PROVEN-LIVE, cites E2D-fresh-vm C1+C2
  D. Storage & devices — the offer, and that registration confers no role
     PROVEN-LIVE, cites SESSION-C C4 + the decline path
  F. Notifications & monitoring — the absent-target alarm and its pairing
     PARTIAL, cites SESSION-C C5, leg named, -> R-116

Row F is PARTIAL today per the doc's own strict enum (a leg not exercised live
is PARTIAL with the leg named, never PROVEN-LIVE). A later session may flip it;
this commit must not.
2026-07-29 23:34:06 +02:00
admin 06d7788392 Session C: R-113/R-114/R-112 PROVEN LIVE; C5 fails on a new defect (R-116)
Full ISO/PAIRING run on a fresh nested box. Agent 0.114.0 came from the Day-0
manifest -- the SHIPPED binary -- so C5 tested the real artifact. Controller
0.186.0 hand-deployed after install per the §3.1 ruling; the vouched golden
bakes 0.185.1, so C3/C4 prove the code not the shipped golden, and that lag is
filed against R-115 rather than a new ID.

R-113 PROVEN: detach 18:43:50, gate fired 18:43:54 -- four seconds, where E-2d
measured zero over 4.5 minutes -- and SetDisconnected was reached. It fired on
exactly the shape that defeated it: raw /mnt/mentes NOT mounted while the bind
/mnt/felhom-drives/mentes still read /dev/sdb[/felhom-data].

R-114 PROVEN: with the target absent the page rendered the absent copy, the
system-disk copy 0 and the offer block 0. Both of E-2d's falsehoods are gone.

R-112 PROVEN: the banner reached a customer's page for the first time. Healthy
renders nothing, proven POSITIVELY -- idle delta 0 /backup/tiers calls, page
load delta +1, single caller, so the seam ran and chose silence.

C5 FAILED on a fourth, separate defect. The alarm fires but as the GENERIC
storage_disconnected, while the recovery is the SPECIFIC backup_target_restored
-- a pair an operator cannot match, which is what notifyDriveReturned's own
comment forbids. backup_target_absent count 0 across the run. Root cause: the
drive is TWO /disks rows and BackupTarget and GuestPath sit on different ones;
absent they separate, on return they rejoin. v0.184.1 fixed the keying, not
this. Only reachable because R-113 made the gate fire at all. Filed as R-116.

Mirror + over-correction guard PASS: non-target drive -> storage_disconnected,
backup_target_absent 0; both drives present -> 0 ABSENT lines and the target
stayed healthy. Caveat recorded: the mirror passes trivially because the target
also produced the generic event.

E-2 and E-2d CLOSED as partially proven with R-116 the one named open leg, per
the runbook's §9 rule decided in advance rather than mid-run.

Capability map NOT touched: it has no E-2 rows at all, so nothing could move to
PROVEN-LIVE. Creating them is a design act, not a validation act.

Teardown clean: pvesm status after == before (local-lvm 38.78%), guest 9201 and
drill-r50 untouched. Customer delete attempted and correctly refused while the
host still reads ONLINE; command recorded for once it ages to DOWN.
2026-07-29 20:55:48 +02:00
admin af518ba151 R-114 + R-112 code shipped (controller v0.186.0) — seam proven live, copy not
R-114: new BackupTargetState.TargetAbsent separates configured-and-gone from
never-configured. Degraded keeps its meaning so the wire contract is unchanged;
TargetAbsent answers which problem, because the remedies are opposite. Copy is
verbatim the hub's backup_target_absent email. The offer is suppressed on the
branch itself, not left to firstOfferableDrive's Disconnected skip -- that flag
comes from R-113 in another repo and this state must be right without it.

R-112: the state finally has a consumer. Server-rendered on /backups via
backupsHandler -> backupTargetView -> backups.html, not a 19th JS fetch. The
view is nil for healthy and unknown so those render nothing at all.

SEAM PROVEN LIVE by a DIFFERENTIAL positive observable rather than by an absent
banner: idle 8s produced 0 new /backup/tiers agent calls; each /backups load
produced exactly +1, and that call has a single caller. The demo box is healthy
and correctly rendered nothing, which matches its real state but is a negative
and so proves nothing about wiring on its own.

MinAgent unchanged at 0.113.0 -- R-114 reads BackupTarget/MountPath/GuestPath/
Role, none of which R-113 altered. demo-hp is not held.

Session C scope unchanged: neither fix touches the agent, so the leg awaiting
proof is still device loss -> gate Stop -> SetDisconnected ->
backup_target_absent on the wire. One rebuild validates all three.
2026-07-29 19:26:09 +02:00
admin 338b2ccf86 agent 0.114.0 published + vouched; R-115 files the recurring publish gap
PART 1 — Session C unblocked.
Agent 0.114.0 (the R-113 fix) was built, pushed and deployed but never
published, so a fresh drill box would have installed 0.113.0 and proven the bug
rather than the fix. Published from the clean tree at b58d7bc via
scripts/publish-agent.sh; sha 5e4c15ebee2d7583d57301d1f7c9cc7d4276262966bf738b05e34653bfd18c31,
verified by an INDEPENDENT round-trip GET (http=200, sha match, binary
self-reports 0.114.0), and the hub manifest read back after the write.

Deliberately NOT done, each with a reason:
- No golden bake. The golden bakes the CONTROLLER, not the agent, and
  host-install fetches them as separate generic packages (:1945 / :2573). Golden
  0.185.1 is current, so there is no new-agent-against-old-golden risk.
- min_agent NOT raised, stays 0.113.0. It expresses what the CONTROLLER requires
  of the agent, and controller v0.185.0 declares MinAgent 0.113.0 — which
  0.114.0 already satisfies. Raising it to 0.114.0 would have been a false claim
  AND would have held demo-hp and drill-r50. No box is held; no §3 STOP fired.
- Global controller floor NOT raised (v0.156.0), per R-111's reasoning.
- wrapper_sha256 preserved verbatim; re-checked against configs/felhom-pbs-apply
  before and after — no drift both times.

demo-hp RULING: left on 0.113.0. The R-113 fix is not live-validated, so putting
it on a second box widens exposure for no proof, and Session C's nested box
takes its agent from the manifest, not from demo-hp's host agent. Move the fleet
once, after Session C.

PART 2 — R-115 opened (WAITING-ON-OPERATOR).
The finding is the RECURRENCE, not either instance: publishing is a remembered
step, and it was forgotten within eight hours of R-111 documenting it as
forgettable. Filed as a new ID with a back-pointer rather than reopening R-111,
because R-111's finding (the channel WAS stale) is closed and verified
end-to-end, while the process defect that caused it is a distinct problem with a
distinct fix and owner. Class cross-linked to R-29 (a control that exists and is
never walked) WITHOUT minting a second ID for it. Options are stated as the
operator's decision, with mechanisms (build-step, deploy gate) separated from
reminders (checklist, manual) — R-29's whole finding being that reminders do not
hold. No code written, by design.

R-111 gains a deferred-leg-recurred line; its shipped evidence is untouched and
it is NOT reopened. R-113 records that Session C is now unblocked.
2026-07-29 19:03:59 +02:00
admin ca4c8b3afc R-113 code shipped (agent v0.114.0) — NOT live-validated, awaiting Session C
BoundUnderParent is now a CONJUNCTION: bound under the parent AND the drive's
raw host mount still mounted. The raw mount is the device-bound systemd unit
that dies with the device; the agent's own bind is not, which is why the bind
outlived the device and the gate could never fire.

Conjunction deliberately, not replacement: the device half alone would regress
boot ordering (raw mounts early, bind lands ~18s later — that window must keep
reading absent), so existing behaviour is byte-identical and only the
unreachable case is closed. Unknown is never absent.

Controller UNCHANGED, no MinAgent bump — BoundUnderParent has exactly one
functional consumer (planDriveGates:226). A new DevicePresent bool was rejected:
absent-from-JSON decodes to false, so every drive on an older agent would have
read ABSENT and stopped its apps.

+6 tests (208->214), 4 red-proofs run and reverted. Deployed to demo-felhom and
the over-correction guard verified in production: raw mount present, drive still
reads present, 10/10 apps untouched, no gate action, no false alarm. demo-hp
deliberately left on 0.113.0 (the spec scoped deploy to felhom-pve).

SESSION C BLOCKER recorded on the row: the hub Day-0 manifest vouches agent
0.113.0, so a fresh drill box would install WITHOUT this fix and validate
nothing. Publish + vouch 0.114.0 first — R-111's trap in the same shape.
2026-07-29 17:23:09 +02:00
admin d839ddcb60 E-2d teardown complete: drill customer + host removed from the hub
The delete was correctly refused at four successive gates while the host still
read ONLINE (acknowledgements -> typed confirm_id -> expect_hosts stale-preview
-> "host is ONLINE"). Rather than force it, the run waited for the destroyed
host to age to DOWN; delete-impact then reported deletable:true and the
documented cascade ran:

  host deleted (escrow demoted to retained custody), tenantsync deprovisioned,
  PBS tenancy deprovisioned, claim reset to unclaimed, residue purged
  (reports=5 app_telemetry=5 notif_prefs=1 appliance_registrations=1)

Verified after: 0 occurrences of "e2d" anywhere on the hosts page; demo-felhom
and demo-hp ONLINE on agent 0.113.0; drill-r50 and peti-felhom unchanged;
demo-hp carries only guest 9201 and VM 300.

Scoping checked rather than assumed: the single purged appliance_registration
was this run's own appliance (810d10c5, bound to e2d-fresh). The unrelated stale
2026-07-25 appliance (206c8838 / QWA-WJE) was NOT touched by the cascade — the
operator removed it separately.

- OPEN-ITEMS.md: the drill-cleanup WATCHING row is removed (done, not open).
- audits/E2D-fresh-vm-2026-07-29.md §8 + REPORT-e2d.md: teardown recorded as
  complete, with the cascade output and the appliance-scoping note.
2026-07-29 15:56:49 +02:00
admin f3975cf5bc E-2d executed on a fresh box: C1/C2 proven, C3/C4 partial, C5 FAILS — R-112/113/114
Full ISO/PAIRING route on a nested PVE VM on demo-hp, after R-111 was fixed
earlier in the session. Bind -> running controller in 3m35s. The install fetched
the artifacts published an hour before and restored the golden baked 20 minutes
before, so the publish train is proven end to end on a real install.

C1 PROVEN: "felhom-host-install v1.22.0", "Day-0 provision SUCCESS", guest 9201
running, bootstrap unit wrote its done-flag and self-disabled. This retires E-2's
"installer-logic-tested, not install-tested".

C2 PROVEN: both DEGRADED warning lines verbatim, backup.local_backup_target=local,
no felhom-backup storage created, and the install did not abort.

C3/C4 PARTIAL and C5 FAILED — three findings, none fixed:

R-112 (P1): E-2's degraded banner and offer have NO UI CONSUMER. The endpoint
returns byte-exact copy; grep 'backup-target' across every html/js/css is 0 hits
and no page handler injects the state. Templates fetch 18 distinct /api/storage/*
endpoints; these two are the only ones with zero references. v0.185.1 fixed the
router mount and stopped one layer short of the render. Fifth instance of
seam-built-but-never-wired.

R-113 (P1): the drive-absent gate CANNOT FIRE on device loss. planDriveGates
reads presence from BoundUnderParent = "is this path in the guest's mountinfo".
The raw mount is a device-bound systemd unit and dies with the device; the
agent's own bind is not device-bound and outlives it, so the gate sees "present"
forever. Live: agent reported the drive absent every 20s for 4.5 minutes, the
controller logged 0 [gate] lines, the hub received zero events -- neither
backup_target_absent nor the generic storage_disconnected. Sixth instance of the
class: E-2b wired the seam to a condition that cannot occur.

R-114: on target-drive loss the message claims the backup is on the system disk
(false) and offers the drive that just vanished. Invisible only because of R-112,
so it must be fixed BEFORE R-112 is wired.

Also filed as a second instance under R-110 rather than a new ID: host-install
fetches nine files from raw/branch/main and the hub vouches a sha for one;
E-2a's wrapper is installed 0755 to /usr/local/sbin, root-fenced in sudoers,
validated only by bash -n.

C4 is fully proven at API level: decline path (registration confers no role),
restart_required:true, agent did NOT self-restart (in-flight check performed and
recorded first), E-2a wrapper created the storage at the drive's own mountpoint,
and healthy renders nothing.

Teardown: VM destroyed, scratch storage removed, pvesm status after == before
(local-lvm 38.77%), guest 9201 and drill-r50 untouched. Hub records for e2d-fresh
remain -- delete correctly refused at four gates, finally "host is ONLINE";
deletable once it ages to DOWN. Command recorded in OPEN-ITEMS.md.

capability-map NOT touched: the customer-facing legs are broken rather than
proven, and the map has no E-2 rows at all.
2026-07-29 13:13:56 +02:00
admin 3dff3573f7 R-111 SHIPPED: the Day-0 artifact channel now serves agent 0.113.0 + golden 0.185.1
Found and fixed the same day. The channel was 17 agent releases stale — a box
installed today would have received agent 0.96.0 and controller 0.161.0.

- agent 0.113.0 built from the clean tree @ 58b598b and published via
  scripts/publish-agent.sh; sha 5f3247f756cb658e…, round-trip GET verified.
- golden 0.185.1 baked on the nested drill VM embedding controller 0.185.1;
  sha dba00f3e845c415e…. Bake clean: Result=success, overlay2, all 3 mounts
  included (rootfs+mp0+mp1), 0 FATAL/exclusions, HTTP 201, token-leak grep 0.
  GL-1 teardown: guest 9100 purged, secrets shredded, drill disk restored to
  the virgin snapshot. Log saved to drill/bake-0.185.1.log.
- Hub Day-0 manifest: agent and golden moved TOGETHER in one POST so the
  manifest never vouched a new agent against an old golden. min_agent
  0.93.0 -> 0.113.0, which is what controller v0.185.0 declares. Zero fleet
  impact, verified: all three enrolled hosts already run agent 0.113.0.
  wrapper_sha256 preserved verbatim (re-checked, no drift).
- The global controller floor was deliberately NOT raised: the golden now
  bakes 0.185.1, so a fresh box needs no self-update.

This unblocks E-2d C3/C4/C5, which the Phase 0 gate had blocked.
2026-07-29 12:10:35 +02:00
admin f3f0d58844 E-2d: Phase 0 STOP — the Day-0 artifact channel cannot deliver the code under test
No VM created, no install run, no box touched. The run stopped at the Phase 0
gate per runbook §3, before provisioning.

felhom-host-install.sh does not install what is on main. resolve_artifacts()
(:423-436) reads the hub-vouched manifest and fetches Gitea GENERIC PACKAGES
(agent :1945, golden :2573). Gitea's newest are agent 0.96.0 and golden 0.161.0;
the hub manifest selects exactly those; the global floor v0.156.0 is below the
golden's 0.161.0 so nothing self-updates. A fresh box therefore lands on agent
0.96.0 + controller 0.161.0 against main's 0.113.0 / 0.185.1. Agent 0.113.0
reached both demo boxes by direct deploy and is not in the channel at all.

Claim impact, each pinned to its introducing commit:
- C1 (real rc=0 1.22.0 install) and C2 (Case B natural) — ACHIEVABLE, not run;
  both are installer-side and host-install is served at 1.22.0.
- C3 — BLOCKED: banner + GET /api/storage/backup-target are controller v0.185.1
  (cdaeb36), copy v0.185.0 (3f7cf2a). Unblocks cheaply by raising the hub floor
  to >=0.185.0; measured fleet impact nil (both demo boxes already 0.185.1).
- C4 — BLOCKED: needs controller v0.185.1 + agent v0.113.0 (58b598b).
- C5 — BLOCKED: needs controller v0.184.0 (c1a63de) + agent v0.112.0.

Filed R-111 (P1): 17 unpublished agent releases (v0.97.0-v0.113.0) strand the
entire R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT, so a new customer's
box installs without them. Mirror of R-110, not a duplicate.

- audits/E2D-fresh-vm-2026-07-29.md — all four Phase 0 answers recorded so a
  resumed run does not re-derive them (cadence 30s; hot-detach available; ISO
  present; local-lvm fence re-measured at 38.77%, unchanged).
- OPEN-ITEMS.md — R-111 opened; E-2d re-stated, NOT closed.
- ROADMAP.md — R-111 under P1.
- capability map NOT touched: nothing was proven live.

The §5.1a operator STOP is retired — HUB_PW is in ~/.config/credentials and hub
auth was verified, so CC can bind on a resumed run.
2026-07-29 11:43:00 +02:00
admin de5a3e5765 docs: retire the last two false gate-enforcement claims; scope the ranking heading
Closes the record-hygiene rider. Part 3 of the spec (documenting a
ROADMAP/OPEN-ITEMS state convention) is deliberately NOT done — its stated
evidence is false; see REPORT-record-correction-2026-07-29.md.

- CONTEXT.md:540 — "scripts/hub_confirm_gate.py enforces" was present tense
  about a gate invoked by nothing. Now says it asserts but is not enforced
  (R-29). Third instance of the class after :564 and configs.go:27.
- REUSE.md:62 — same claim, "enforces zero". The RULE stays (never native
  confirm()/prompt() is correct guidance and this is a reuse-reference row);
  only the enforcement claim changes, and it now says the rule holds only as
  long as you keep it.
- OPEN-ITEMS.md:4 — root REPORT.md is the overwritten per-session file;
  REPORT-<topic>.md is the non-clobbering sibling form (CLAUDE.md:82-87), of
  which 14 exist. The prohibition on durable content living only there stays.
- OPEN-ITEMS.md:55 — "Why the READY rows rank this way" promised a complete
  ordering and listed 5 of ~15 open rows. Scoped to TOP, with a half-sentence
  saying it is deliberately not a full ordering. No row added to the list.

hub/internal/web/configs.go:27 left alone (R-94 leg (b), needs a hub build).
No gate wired, run or fixed. Documentation only, no version bump, no CHANGELOG.
2026-07-29 11:22:07 +02:00
admin 7383400a23 docs: file R-29 to the register; attach the gate-orphan instance to its class
d4c07873 filed "hostinstall_gates.py is invoked by nothing" as a novel
observation. It is not novel — R-29 already names the class (green gates are
enforced nowhere; one sat RED for 16 releases while every REPORT said green),
and R-29 was missing from OPEN-ITEMS.md entirely, having never been carried
across the 2026-07-27 register rebuild. An open item about work not getting
done was absent from the page that decides what gets done.

Ruling on whether R-29 is the right home for a non-design-v2 gate: YES. Its
title says design-v2, but its own audit list already spans mount-safety,
secrets and dedup gates across four repos, and its part (b) — "the systemic
half is the real item" — is about the enforcement mechanism, which is
gate-agnostic. hub_confirm_gate.py is already on its list and sits in the same
scripts/ directory. No new ID minted; R-29's own text forbids it, and this is
the third re-raise it has absorbed.

- OPEN-ITEMS.md: open R-29 (READY, S(a)/M(b)), with the orphan evidence and
  the two separable parts R-29 already defines.
- OPEN-ITEMS.md: R-94 leg (b) now points at R-29 as its class.
- ROADMAP.md:158: audit list extended with hostinstall_gates.py (RED today,
  1.19.0 != 1.22.0) + hub_confirm_gate.py verified orphan. Entry not rewritten.
- ROADMAP.md:147: cited a non-existent R-164 — it means controller v0.164.0.
- CONTEXT.md:564: asserted in the present tense that the version cross-check is
  "gated by scripts/hostinstall_gates.py". It exists, is red, and runs nowhere.
- OPEN-ITEMS.md: READY #1/#3/#4 markers dropped — they duplicated ranked-list
  positions and the gap was left by the row merged in d4c07873.
- OPEN-ITEMS.md: E-2d citation :322-341 widened to :322-343; the invocation it
  describes is at :343, two lines outside the old range.
- backlog/README.md: two-line lead naming OPEN-ITEMS.md and ROADMAP.md.
- REPORT-record-correction-2026-07-29.md: the report CLAUDE.md:82-87 requires
  for both commits. Root REPORT.md (E-2 increment 1) untouched.

No gate wired, fixed, run or deleted — that is R-29 part (b), its own task.
Documentation only. No version bump, no CHANGELOG entry, no box touched.
2026-07-29 11:13:23 +02:00
admin d4c07873ca docs: correct the installer-channel record — R-94 retracted and re-scoped, R-110 opened
The 2026-07-29 R-94/E-2d finding was written from an unverified claim and was
false. `felhom-bootstrap.sh:96` fetches the installer from the WEBSITE, not the
hub; the website git-syncs /scripts/ from main on a 30s period; every install
since 1.22.0 hit main this morning already runs 1.22.0. Confirmed by live fetch.

- OPEN-ITEMS.md: merge the two duplicate R-94 rows into one, retract the false
  framing, re-scope to what it actually is (a drifting hand-synced constant plus
  two pieces of dead safety equipment), unblock it from E-2d.
- OPEN-ITEMS.md: de-rank R-94 in the ranked list — the "high-consequence" reason
  was the false claim in its most load-bearing form.
- OPEN-ITEMS.md: E-2d — the ISO is the STRONGER proof route, not an obstacle.
  Phase 0 question answered at source: PAIRING falls through to run_direct in
  the same invocation (:495-499), so it reaches the identical installer call.
- ROADMAP.md:149: same retraction; the original diagnosis (a hand-synced
  constant in a second repo drifts every time the first ships) survives.
- ROADMAP.md + OPEN-ITEMS.md: new R-110 — main is the installer's publish
  channel and there is no staging, tag, pinned path or rollback, for the one
  artifact that runs as root on a virgin box. Operator ruling, not a defect.
- day0-install.md C.1: one sentence recording the same about the fetch URL.

Documentation only. No version bump, no CHANGELOG entry, no code, no box touched.
2026-07-29 10:54:59 +02:00
admin 36d635a4cd E-2d: file the fresh-VM proof plan; R-94 blocked on it, with the ISO finding
Space checked on the t740 -- NOT a blocker, with one constraint: the VM disk must
not go on local-lvm. That thin pool is over-subscribed (144G allocated against a
54G pool) on a box running a live customer guest, and a full thin pool corrupts
every guest on it. local has 23.7G on pve-root. Use /mnt/nvme-1tb (888G free).

Confirmed the ISO does NOT bake felhom-host-install.sh -- it ships
felhom-bootstrap.sh, which fetches the installer FROM THE HUB. Since the hub
serves 1.19.0, a fresh ISO install today would run the pre-E-2 installer and
exercise neither Case A nor Case B. So R-94 must be bumped only AFTER a real
1.22.0 run, not before -- which is the ordering already decided.

drill-r50 stays blocked and was restored to its r50pre state: the agent upgrade,
the added disk and the moved backup target from this session are all reverted.
2026-07-29 09:56:14 +02:00
admin bcbe2707d6 E-2 complete: wrapper, installer Case A/B, offer flow, degraded banner
Live: hub 0.81.0, agent 0.113.0, controller 0.185.1 on both demo boxes;
host-install 1.22.0 (script; no reinstall performed).

E-2a wrapper proven live as root on demo-hp: F-1 subdirectory refused, F-2
unmounted path refused, root device refused, idempotent re-apply is a no-op,
repointing refused -- 0 stray storages. The agent PVE role was NOT widened.

Scenario E proven live on BOTH boxes: healthy renders nothing, no message key.

Records three defects I introduced and caught: unreachable routes (mounted
outside /api/storage/, caught by the first live call), a hollow test exposed by
its own red-proof, and another gofmt-realignment no-op.

Not live-proven: the degraded banner and offer acceptance (both boxes healthy),
backup_target_absent end-to-end, Case A/B on a real install, drive-loss recovery.
2026-07-29 09:16:59 +02:00
admin 3696188636 E-2 increment 1: report + close E-2b/E-2c as shipped and proven live
hub 0.81.0, agent 0.112.0, controller 0.184.1 live on BOTH demo boxes.

E-2c: eject/decommission of the backup-target drive refused 409 on both boxes,
drives unmoved. E-2b: the never-called disconnect seam is wired, with the target
case raising the specific backup_target_absent.

Records the keying bug caught before deploy (a.Path is the GUEST path, so the
target branch was unreachable -- 0.184.0 superseded, never deployed) and states
plainly that backup_target_absent is NOT proven end-to-end live: proving it needs
a live enrolled drive to go absent.

Parts 2/3/4 and E-2a remain open; Peti risk stays parked.
2026-07-29 08:34:41 +02:00
admin 2508788d38 E-2: file the remaining work, three Phase 0 findings, and the parked Peti risk
E-2 is partially shipped (hub v0.81.0 + controller Part 1). Filing the rest so a
foundation with no UI cannot quietly become a sixth seam-built-but-never-wired.

  E-2   remaining: installer Case A/B, the offer + agent-side move, the degraded
        banner, the controller half of the signal, red-proofs E/F, live validation.
        Phase 0 INVERTED the emphasis: the installer has no drive-enrollment step,
        so the common case at install is system-drive-only and Part 3 (drive added
        later) is the PRIMARY path, not Case A.
  E-2a  the move needs a root-fenced wrapper -- the agent holds neither
        Datastore.Allocate at /storage nor Permissions.Modify, and its sudoers has
        no pvesm and no pveum. Use the guarded-wrapper pattern; do NOT widen the
        agent's PVE role.
  E-2b  NotifyStorageDisconnected/Reconnected are defined and called NOWHERE, so a
        drive going absent emits no event at all. Hub side is already plumbed, so
        wiring needs no hub change.
  E-2c  E-1 put the whole-guest backups on a drive POST /disks/eject will eject
        (RoleForStorage returns user-data for a local-dir on a non-system device).
        Guard the eject specifically -- reclassifying the drive RoleBackup would
        block legitimate ejects, since it is also the enrolled user-data drive.
  PETI  peti-felhom deliberately NOT migrated; drive failure there is offsite-only
        recovery. Accepted until the operator's reinstall; re-evaluate if that
        slips past ~2026-09-01.
2026-07-29 08:01:42 +02:00