Merge branch 'main' of https://gitea.dooplex.hu/admin/felhom.eu
gates / gates (push) Successful in 7s

This commit is contained in:
2026-08-03 13:14:17 +02:00
15 changed files with 988 additions and 235 deletions
+16
View File
@@ -213,6 +213,22 @@ local and skippable, and only CI is neither.
- **Website** auto-deploys via git-sync; just push to `main` (live in 1–2 min). Website changes go
through `repo_gates.py` above (it runs `site_gates.py`); new pages go into that gate's `PAGES`
list. Emergency edits: https://files.felhom.eu. All `website/` HTML is **UTF-8 with BOM** — preserve it.
- **THE INSTALLER DOES NOT (R-110, 2026-08-03).** `manifests/webpage.yaml` runs **two** git-syncs:
the website from `main` as above, and `/scripts/` from the tag **`installer-v<SCRIPT_VERSION>`**.
Pushing `scripts/felhom-host-install.sh` therefore changes nothing that any machine downloads —
which it used to, within thirty seconds, for the one artifact that runs as **root on a virgin box**.
- **To publish:** cut `installer-v<new SCRIPT_VERSION>`, bump the `--ref` in `webpage.yaml`
(both the sidecar and the init container), commit, and sync. `hostinstall_gates.py` gate 6
fails if the manifest stops naming an `installer-v…` tag or if the website stops tracking `main`.
- **To roll back:** move the tag back to the previous commit and wait ~30 s. **No ArgoCD sync and
no deploy** — git-sync picks up a moved tag on its next period, measured live on 2026-08-03 in
both directions. That is the emergency lever; fix forward with a new version afterwards.
- **Do NOT pin the website to the tag.** The sparse-checkout used to cover `/website/` and
`/scripts/` in one sync, and pinning that would turn every copy edit into a release.
- The **URL never carries a ref** (`https://felhom.eu/scripts/felhom-host-install.sh`), so
`felhom-bootstrap.sh` and the hub's day-0 command follow the tag with no edit — do not add one.
- The installer's own sixteen run-time fetches are pinned separately, to `raw/tag/v$ART_AGENT_VER`
in the **agent** repo (R-183) — they are the agent's configs, not this repo's.
- **Manifests** are GitOps via the `felhom` app — commit to `main`, then deliberate sync.
## Key patterns
+106
View File
@@ -17,6 +17,112 @@
## Standing rulings
**S-13 — the `mp1` merge landed, and the variant was chosen on measurement (2026-08-03, R-165 / D-a).**
The appliance's two data volumes are one. **Variant V-c**: the volume mounts at the NEUTRAL path
`/var/lib/felhom`, and both `/var/lib/docker` and `/mnt/sys_drive` are binds of subdirectories of it.
**Three shapes were built and rebooted before choosing** (`audits/SPIKE-r165-phase0-2026-08-03.md`) —
all three boot, reboot 3/3, give ONE `df` figure and keep a container's `statfs("/")` on the merged
volume, so **the ordering risk that motivated the probe was not what mattered.** They differ only in
which documented guarantee they break: volume-at-`/var/lib/docker` puts customer backups inside
Docker's data-root, so the ordinary "clear `/var/lib/docker`" reflex destroys every local unit;
volume-at-`/mnt/sys_drive` puts Docker's entire data-root under `/mnt`, which the controller container
mounts wholesale — **measured: it then sees `/mnt/sys_drive/docker`**, falsifying the bootstrap's own
comment that `/mnt` holds only Felhom's namespace mounts. **V-c breaks neither**, for one extra path.
**B2 is the bulkhead replacement, and "or prune the oldest" is REJECTED with its reason**, because the
question will be asked again: nothing on that filesystem is generational — a unit is ONE fixed path per
app (`backups/primary/<app>`) refreshed in place, and a DB dump is `<stack>-<dbtype>.sql`, also fixed —
so pruning could only mean deleting a **different** app's only local recovery unit.
`pruneStalePrimaryDirs` is an ORPHAN sweep with no notion of age and must never be repurposed.
**No migration exists, and that is a ruling not an omission:** every node is REINSTALLED. Both demo
boxes are Tier 0; the colleague's box carries none of our customer data and is clean-installed shortly.
So R-176's in-place migration rehearsal is **withdrawn**, not deferred.
**S-14 — prove first, then vouch (2026-08-03) — SPENT, and the ordering did not survive contact.** The
rule was: golden **0.192.0** stays UNVOUCHED until a box has been proven from it, because vouching is
what makes a fresh install pick a golden up. **In the event the golden was vouched at 07:23:26 CEST on
2026-08-03, before any box was reinstalled** (hub log `Artifact manifest set: agent=0.119.0
golden=0.192.0`), so the ordering was already spent when R-178's session opened; the operator elected
to accept it rather than revert the manifest. **Both boxes were then reinstalled and proven** (R-178,
`REPORT.md`), so the end state is the intended one and no unproven layout was ever in front of a real
install — but the rule protected nothing, because nothing enforced it. **The lesson is R-115's, one
layer up:** an ordering that lives only in a `CONTEXT.md` sentence and a runbook's §7 is a reminder,
and reminders do not hold. If prove-then-vouch is to be a rule it needs the shape R-120's gate has —
a refusal at `handleSetArtifacts`, the sole path to `SetArtifactManifest`, which runs without anyone
choosing to run it.
**S-15 — the merged layout is proven live, by two different supply paths (2026-08-03, R-178).** Both
demo boxes were wiped and reinstalled from golden 0.192.0 and taken through claim → deploy → back up →
**restore**. **demo-hp** was installed with `--golden <local volid>` (the layout proof) and
**demo-felhom** by the normal manifest route with `--force-gitea-golden` (the pipeline proof —
`verified sha256 54e2a4c431daf580… matches the hub manifest`), deliberately different so the session
proved the disk shape *and* the delivery route rather than one of them twice. Live shape on both:
`mp0` at `/var/lib/felhom`, `backup=1`, **no `mp1`**; `/var/lib/docker` and `/mnt/sys_drive` both real
mounts of its subdirectories via `/etc/fstab`; ONE `df` figure and one device id on all three paths;
3/3 reboots each with the binds surviving every time. B2 was **not** proven on that pass → **R-181**:
the floor guarded `captureAllRecoveryUnits` and not `runVolumeDumps`, the leg that fills the volume,
and its refusal's "the previous unit is untouched" was measured false. **R-181 CLOSED the same day
(controller v0.193.0 + v0.193.1), so R-165 is now PROVEN-LIVE in both halves** — see S-14.
**S-14 — the reserve is a per-app, per-run ADMISSION decision, not a capture check (2026-08-03, R-181;
controller v0.193.0 + v0.193.1).** B2 as first shipped was consulted in exactly one place —
`captureAllRecoveryUnits`, a few KB — while `RunDBDumps`' database leg and `runVolumeDumps` wrote the
bulk into the same `backups/primary/<app>` tree, first and unguarded. The reserve was therefore
consumed by the very write it exists to bound, and the refusal then claimed *"the previous unit is
untouched"* about a tree the earlier leg had already rewritten (182,272 B → 2,147,666,432 B under a
manifest that had not moved). **Sixth entry in `CLAUDE.md`'s table of shipped guarantees the code did
not provide, and the fourth of those found on live hardware rather than by review.**
- **`internal/backup/admission.go` — `admitApp` is now THE gate**, and every per-app write leg calls
it. One verdict per app per run covers all three; they share one per-app root, which is what makes
that honest.
- **Decided lazily at the app's first write, never once at run start** (app A's dump can put app B
under the reserve), **never re-decided between an app's own legs** (that is the split it closes),
and **reset per run**.
- **Ahead of `DumpAppVolumesSafe`**, which stops the stack as its first act — a refusal decided
inside it has already bounced the app. **After** the volume-less check, which has no write to gate.
- **Size term added:** *would THIS app's write cross the reserve?*, estimated from the app's previous
`.sql` + `.tar`. **No history → headroom-only**, deliberately — otherwise the first backup is the
one that can never happen.
- **A container-based `du` was MEASURED and rejected**, not waved away: median **~355 ms/volume** over
66 runs on demo-hp, on volumes holding tens of KB (container start-up, not the walk). Decisive on
top: `docker run` needs the writable layer, so the instrument can fail under exactly the pressure
the reserve handles.
- **The wording was NOT weakened; the behaviour moved so it became true**, and it is checked by
sha256 tree fingerprint, never by reading the log line — the log line is what lied.
- **v0.193.1**, found by the proof run itself: a 178 KB estimate printed as `0.00 GiB`, which reads as
*no estimate available*. Rendering moved to `humanizeBytes`; arithmetic still in GiB.
- **New finding, deliberately not fixed here → R-182**: `GetFullStatus`'s periodic capture sweep has
no run scope, so a refused app re-alerts on every status refresh (measured: a second identical alert
pair 13 s after the run's). Pre-existing in v0.192.0; R-181 changed neither caller.
**S-15 — publishing is an act, not a side-effect of pushing (2026-08-03, R-110 + R-115 + R-183).**
Two rulings, one shape: something became live because someone pushed, not because anyone decided.
- **The installer.** `/scripts/` now git-syncs the tag `installer-v<SCRIPT_VERSION>`; the **website
keeps tracking `main`** in a second sync, because pinning both would make every copy edit a
release. Publish = cut the next tag + bump the manifest `--ref` + sync. **Roll back = move the tag
back**, which takes ~30 s and needs no ArgoCD sync at all — git-sync v4.4.0 follows a moved tag,
and that half was measured before the manifest was touched because the whole model rests on it.
- **The sixteen run-time fetches were NOT what the spec described** — sixteen, not nine, and from
`felhom-agent`, not this repo — so no tag here could cover them. They are pinned to
`raw/tag/v$ART_AGENT_VER` instead, which is strictly better: the agent's configs now come from the
same ref as the agent binary being installed. That closed a real skew (**R-183**), not just a
channel.
- **The URL needed no change**, and that is worth knowing rather than re-deriving: it never carried
a ref, so both producers follow the tag automatically — and no hub change means no hub bump.
- **The agent.** `scripts/release-agent.sh` is THE release path: build → tag → publish → **verify by
an independent download**. It does not vouch. `check-published-versions.py` refuses a `v<semver>`
tag with no downloadable package, and **CI now runs the full gate set** rather than `--fast`,
without which that gate would have been registered and never run.
- **The gate's invariant is not the one specified, and P-C is why:** the hub manifest and Gitea's
package listing are both **401** anonymously; the package download and the tags api are not. So CI
can ask *is this installable* but not *what is vouched*. The residue is **R-184**.
- **Neither gate asserts "the newest version is published."** That would go red on the very push
that bumps a version, before publishing — and a gate that fails on the normal path is one people
learn to ignore.
**S-11 — D-c's routing, and why R-158's own proposal was overruled (2026-08-02, R-167 SHIPPED).**
Decision D-c splits two signals by AUDIENCE, and the split is the ruling: **a fill warning is the
CUSTOMER's** (they can free space, delete files, add a drive) and **a per-app backup capture failure
+179 -126
View File
@@ -1,147 +1,200 @@
# REPORT — hub v0.89.0: the two halves of decision D-c, plus the R-165 merge spike (2026-08-02)
# REPORT — publishing becomes an act, not a side-effect (R-110, R-115) + R-182 measured, R-183/R-184 filed
**Overwritten** per the standing rule. The prior contents (R-168, the CI runner, same day) have their
durable record in `scripts/CHANGELOG.md` and `CONTEXT.md` S-8/S-9/S-10.
**Date:** 2026-08-03 · **Repos:** `felhom.eu` (installer **v1.22.0 → v1.23.0**), `felhom-agent` (**no bump**)
**Nothing was built** — no image, no binary, no golden. **Hub stays v0.89.0.**
**Companion report:** `felhom-controller/REPORT.md` holds the controller side (v0.191.0/.1/.2), the
full red-proof table, the Hungarian copy, and the live evidence for all three flows. This file covers
the hub change, the documentation coupling, and **Part 3's spike**.
## 1. Baselines — re-read on arrival, both matched §1
---
## 1. Baseline drift — recorded, because the task's §1 was wrong
The task targeted hub **v0.87.0 → v0.88.0**. On arrival `main` was at `8ef92a3f` with hub **v0.88.0
already shipped** (R-172, the WAL fix), not `d5774d318941`/v0.87.0. Target corrected to **v0.89.0**.
Highest register ID in use was **R-173**, not R-171.
## 2. Hub change (v0.89.0)
**One new event type, not two.** The task called for a new customer-facing type *and* a new
operator-only one. Reconnaissance found `disk_warning`/`disk_critical` already allowlisted here, with
Hungarian `customerMessages`, in the controller's `DefaultEnabledEvents` and behind a UI checkbox —
**and with no producer in any repo.** The operator chose to wire that inert pair rather than mint a
near-duplicate, so only the operator type is new.
| Change | File | Why |
|---|---|---|
| `+ "recovery_unit_capture_failed"` | `internal/api/handler.go` (`allowedEventTypes`) | without it the controller's POST 400s and the event vanishes |
| `+ "recovery_unit_capture_failed"` | `internal/notify/dispatcher.go` (`operatorOnlyEvents`) | **this** is what makes it operator-only; the allowlist does not, and v0.78.0 claimed otherwise and shipped the defect |
| `- customerMessages["disk_warning"]`, `- ["disk_critical"]` | `internal/notify/templates.go` | `FormatCustomerEmail` PREFERS the entry over the message, so a static template would discard the drive label and the free-space figures the controller now sends. Same reason `offbox_enlarge_blocked` and `disk_health_degraded` have none |
| `+ func IsOperatorOnly` | `internal/notify/dispatcher.go` | lets the `api` package pin BOTH registers in ONE test; checked separately, allowlisted-but-not-operator-only is invisible. Read-only — the register stays unexported so nothing can widen it at runtime |
| `REUSE.md` §5 "new event type" rewritten | `REUSE.md` | it told readers to always add a `customerMessages` entry, which is **wrong** for operator-only types and **harmful** for dynamic-message ones |
**Tests 574 → 579**, full suite green (`go build ./... && go vet ./... && go test ./...`), all five
`repo_gates.py` gates OK.
**Red-proof (Scenario G), demonstrated not argued:** removing `recovery_unit_capture_failed` from
`operatorOnlyEvents` fails two tests, one reading *"a customer was emailed the OPERATOR-ONLY
recovery_unit_capture_failed (customer@example.com)"*. The dispatch test runs under the **breaking**
configuration — the customer has the event enabled and an email set — because that is the only
configuration in which the missing entry is visible.
**Live (guest 9201 → hub):** both event types accepted and stored; `operator | sent`; and the positive
observable `customer | recovery_unit_capture_failed | skipped | operator_only` read from
`notification_log`. The customer half: `customer | disk_warning | sent` and `customer | disk_critical
| sent` with the dynamic Hungarian intact.
**Deploy:** GitOps only — `manifests/hub.yaml` bumped 0.88.0 → 0.89.0 (`6d359a5`), pushed, then a
deliberate ArgoCD hard-refresh + sync. Never `kubectl set image`. App `felhom` **Synced / Healthy**,
`deploy/hub` rolled out, running `gitea.dooplex.hu/admin/felhom-hub:0.89.0`, startup log clean.
## 3. Part 3 — the R-165 spike. **M1-M5 each answered; nothing was changed.**
Full document: `documentation/audits/SPIKE-r165-mp1-merge-2026-08-02.md`. No partition was created,
resized, moved or deleted; no golden rebuilt; no guest config edited. `ep0` and Peti's box were not
contacted (D-d, `runbooks/target-selection.md`).
**M1 — what is actually there. ANSWERED, and it contradicts the architecture doc.**
| | demo-felhom | demo-hp | golden default |
| Repo | @ arrival | Version | Result |
|---|---|---|---|
| `mp0` `/var/lib/docker` | **200 G** (13 G used) | **50 G** (5.4 G used) | 16 G |
| `mp1` `/mnt/sys_drive` | **50 G** (2.0 G used, 5%) | **20 G** (92 M used, 1%) | 8 G |
| `felhom.eu` | `8360f940bfb2` | hub v0.89.0, `SCRIPT_VERSION="1.22.0"`, **0 tags** (confirmed) | installer **v1.23.0**, first tag `installer-v1.23.0` |
| `felhom-agent` | `9dfd89cb947e` | v0.120.0 | **unchanged** — scripts and gates only |
§7.5 documents the appliance as `mp0 50G / mp1 20G` — that is demo-hp exactly and **not** demo-felhom.
Any merge plan expressed as a fixed pair is already wrong for one of the two boxes that exist. §7.5's
headline bound (*"≈ 19 GB … ≈ 10 GB"*) is derived from `mp1 = 20 G` and is therefore one box's, not
the fleet's → **R-175**, filed and §7.5 annotated in this session.
## 2. Part 0 — the R-182 measurement, and it REVERSED the row
**M2 — what lives on `mp1`. ANSWERED, and it is not only backups.** Four things would move:
Tier-1 units of driveless apps (269 M, ~30 apps on demo-felhom), **Tier-2 mirrors (1.7 G — i.e. the
MAJORITY is Tier 2, not Tier 1)**, the `userdata/import` drop zone which lives on the system drive by
**contract** (R-75), and the system-data userdata namespace. Observed fill is 5% / 1%: the constraint
is a **ceiling** problem, not a current-fill one.
Filed yesterday as *"the reserve re-alerts on every status refresh"* — **too many** alerts, observed
at the sending end. Measured at the **receiving end**, it is the opposite.
**M3 — which merge shapes exist. ANSWERED for three shapes, with ONE item explicitly unmeasured.**
The golden **fails closed on the split in four places**, not one (`build-golden.sh:126,130`
separate-mount asserts + `:315,319` vzdump-exclusion guards). The archive scope `rootfs+mp0+mp1` stays
complete after a merge (the data moves onto `mp0`). `mountParity` holds for new archives. **Unmeasured
and reported as such:** whether a *pre-merge* archive restore-tests into a *merged* guest — reading
`mountParity` says it should, but that is reasoning from source about an unvalidated mechanism, which
this project has got wrong four times → **R-176**.
Method: the hub's SQLite copied **with its `-wal`** (4 MB and newer than the db — copying `hub.db`
alone would have read stale data, the exact trap this project recorded before), freshness confirmed by
the newest `notification_log` row post-dating the session.
**M4 — the bulkhead. ANSWERED, and it is the important one.** `mp1` is not only a ceiling: today an
overflow is refused per app with the last good unit byte-identical **and cannot reach
`/var/lib/docker`**. After the merge it can, and a full Docker data-root is a stopped box, not a slow
one. Four replacements costed — a reserved block percentage, **a refusal threshold in the capture
path**, a project quota, or deeming R-167's warnings sufficient — with the trade-off of each.
**Deliberately not chosen: this is the operator's ruling.**
**9 `recovery_unit_capture_failed` events received today → 2 operator emails sent.**
**M5 — existing boxes. ANSWERED for the measurable population; one part honestly UNMEASURED.** The
hub's `/hosts` register holds four hosts, **two ONLINE**, both demo boxes — and both are **Tier 0,
therefore reinstallable rather than migratable (D-d)**, so migration cost for the measurable population
is **zero**. **D-a's condition (1) — "before any external install" — is currently SATISFIED**, which
makes this the cheapest this decision will ever be. **`peti-felhom` exists as a customer with NO host
in the register**, so its layout is not knowable from the hub and the box was not contacted; whether it
needs converting or reinstalling is the operator's information. The in-place migration procedure has
**never been rehearsed**, so "is the box restorable at every point of it?" is currently unknown → also
**R-176**.
| time | apps refused (events in) | operator emails out |
|---|---|---|
| 06:40:03 | privatebin, opengist | **opengist only** |
| 08:59:46/47 | opengist, privatebin | **privatebin only** |
| 08:59:59 | privatebin, opengist | **none** |
| 09:03:00 | opengist | **none** |
| 09:07:06 | privatebin, opengist | **none** |
**Ranked options and recommendation:** (1) **S1 — one volume with the two paths as directories — plus
B2, a refusal threshold in the capture path**, shipped as a fresh-install shape with the demo boxes
reinstalled; (2) S1 + warnings only; (3) S3, grow `mp1` and keep the split (D-a's rejected baseline,
measured for comparison); (4) S2, two mounts on one pool — **not recommended at all**, it satisfies
every assertion while delivering none of the benefit and converts a clean per-app refusal into a
shared-pool exhaustion neither `df` can see coming.
**Cause, confirmed at source:** the operator cooldown key is
`customerID + ":" + eventType + cooldownTierSuffix(details)` (`dispatcher.go:268`, 1 hour hardcoded).
`RecoveryUnitFailureDetails` carries **`app`** and **no `tier`**, so the suffix is empty and the key
holds **no app identifier**. The first refused app takes the slot; every other app's refusal for the
next hour is dropped — and dropped **before `LogNotification`**, so it leaves **no row on any
channel** and cannot be audited afterwards.
**STOPPED at the operator's question**, per the task. The merge is next session's supervised work.
This is **R-97a's failure mode in a second event type**; that row's own comment states it
(*"`felhom-pbs` failing at 09:00 would swallow `local` failing at 09:20"*). `cooldownTierSuffix` was
written narrow on purpose; `recovery_unit_capture_failed` simply never opted in.
## 4. Documentation coupling
**A correction I owe on yesterday's report.** It said *"one `recovery_unit_capture_failed` per app,
HTTP 200"*. That was true of what the **controller pushed**, and a reader would take it as *the
operator was told about each app* — which is false. The gap between an accepted event and a sent
email is the whole of this row.
| File | Change |
**Nothing was changed** (§8.5). R-182 is re-scoped with the evidence and the fix shape.
## 3. Probes
| | Question | Method | Verdict |
|---|---|---|---|
| **P-A** | does git-sync v4.4.0 follow a tag, and notice a **moved** one? | throwaway `docker run` git-sync against this repo, tag moved under it | **PASS both halves** — `update required … local:fb65202 remote:8360f94` → `updated successfully`, one period (~20 s) |
| **P-B** | does Gitea serve `raw/tag/<tag>/<path>`? | one fetch on a throwaway tag | **PASS** — HTTP 200, byte-identical to `raw/branch/main` |
| **P-C** | can CI read the package registry? | anonymous fetches | **PARTIAL, and it changed the gate's design** — package **download** 200 (and **404** for a fake version, so it discriminates), **tags** api 200; package **listing** api **401**, hub artifact manifest **401** |
**Publish model P-A implies:** publishing is **moving the tag**; rollback is **moving it back**, in
~30 s with no ArgoCD sync and no deploy. Probe teardown: container, sync tree and probe tag all gone
(`git ls-remote --tags` → 0 at the time).
## 4. §8.2's three channels — enumerated
| Channel | Before | After | |
|---|---|---|---|
| 1. the served script | `main`, 30 s | **`installer-v1.23.0`** | **MOVED** — `webpage.yaml` split into two syncs |
| 2. the run-time fetches | `raw/branch/main` | **`raw/tag/v$ART_AGENT_VER`** | **MOVED** — but see below |
| 3. the URL producers | `main` | unchanged | **NO CHANGE NEEDED** — and that is a finding, not an omission |
**Channel 2 was not what the spec described, and the spec's mechanism for it was unimplementable.**
There are **sixteen** fetches, not nine, and they come from **`felhom-agent`**, not `felhom.eu` — so
no tag on this repo could ever have covered them, and §8.1's *"derive the tag from `SCRIPT_VERSION`"*
was impossible for them. Raised before building; operator ruled to pin them to **the agent version
being installed**, which the installer already resolves from the hub manifest and already sha-verifies.
That is strictly better than any installer-derived tag: binary and configs now come from one ref.
**Channel 3 needed no change because the URL never carried a ref** —
`https://felhom.eu/scripts/felhom-host-install.sh` is path-based; the ref lives in the manifest. So
`felhom-bootstrap.sh` and the hub's day-0 command follow the tag automatically. **No hub template
change ⇒ no hub bump**, so §1's rule was never in tension and the STOP it anticipated never arose.
## 5. The tag convention
- **Shape:** `installer-v<SCRIPT_VERSION>` in `felhom.eu` (prefixed so it cannot be read as a hub,
agent, controller or golden version); `v<semver>` in `felhom-agent` (that repo versions one thing).
**No new constant in the installer** — channel 2 derives its ref from `$ART_AGENT_VER` at run time,
and channel 1's ref lives only in the manifest.
- **Publish:** cut `installer-v<new SCRIPT_VERSION>`, bump the `--ref` in `webpage.yaml` (sidecar *and*
init container), commit, sync.
- **Roll back:** move the tag back to the previous commit — takes ~30 s, **no ArgoCD sync, no deploy**.
## 6. Scenario A — proven by HTTP
A real commit was pushed to `main` (a marker comment in the installer) **without moving the tag**, and
three sync periods were allowed to pass so "unchanged" means "had every chance to change":
```
website tree (main): .worktrees/6a82719… <- ADVANCED to the new commit
scripts tree (tag): .worktrees/bee6848… <- STAYED
sha256 before push: 2f859555382c4c69c18c48dccd8d8b132ffd49b4dbe4e03e5dbb192e8d883555
sha256 after push: 2f859555382c4c69c18c48dccd8d8b132ffd49b4dbe4e03e5dbb192e8d883555
marker present at the served URL? 0
https://felhom.eu/ -> HTTP 200
```
Both halves of the split in one observation: the site still tracks `main`, the installer does not.
## 7. Scenario B — publish and rollback, both directions
| act | result |
|---|---|
| `documentation/backlog/OPEN-ITEMS.md` | **R-158** closed (by R-167 — *no second row for the same wire*); **R-167** closed; **R-165** updated with M1-M5 + the operator question, stays open; **4 new rows** R-174/175/176/177 |
| `documentation/backlog/ROADMAP.md` | R-158 collapsed to a shipped one-liner; R-167 added as shipped; R-165 added as spiked/waiting-on-operator |
| `documentation/architecture/00-capability-map.md` | **two new rows**, both **PROVEN-LIVE** with live citations |
| `documentation/architecture/07-backup-architecture.md` §7.5 | **S-1: the contract changed in the same session.** The section's closing claim *"nothing warns when an app crosses the line"* is now false; the alerting is written in, and the one-box-vs-fleet caveat added |
| `CONTEXT.md` | **S-11** (D-c's routing, and why R-158's own `backup_failed` proposal was overruled) and **S-12** (the monitoring landed *before* the merge, not with it) |
| `STATUS.md` | new plain-language section; the merge decision added to *Waiting on you*; **two older entries trimmed** so the page did not grow — one screen, per its own rule |
| `REUSE.md` | the "new event type" extension point rewritten (see §2) |
| tag moved `bee6848 → 6a82719` | scripts tree moved in **~40 s**; served `sha256 ea2b4aa9…`; **marker present** |
| tag moved back `→ bee6848` | scripts tree back in **~40 s**; served `sha256 2f859555…` — **exactly** the pre-publish sha; **marker gone** |
## 5. Register IDs
`https://felhom.eu/` returned 200 throughout. The marker commit was then reverted, and the tag moved
to `main`'s head — a **byte no-op**, verified by the served sha not changing.
**Opened:** R-174, R-175, R-176, R-177. Each established free by
`grep -ro "R-17n\b" documentation/ *.md` → **0 hits**, run before minting.
**Closed:** R-158, R-167, R-174. **Updated, still open:** R-165, R-163 (unchanged — it stays the
record of the constraint until the merge lands).
## 8. Files, commits, tags
## 6. CI — run ids and conclusions
**`felhom.eu`** — `bee6848` (installer + gate + manifest), `6a82719` (Scenario A marker), `e79a20b`
(marker removed), plus the docs commit below.
`scripts/felhom-host-install.sh` · `scripts/hostinstall_gates.py` · `scripts/CHANGELOG.md` ·
`manifests/webpage.yaml` · `CLAUDE.md` · `CONTEXT.md` · `STATUS.md` · `REPORT.md` ·
`documentation/backlog/{OPEN-ITEMS,ROADMAP}.md` · `documentation/architecture/00-capability-map.md`
Checked by PULL from `…/actions/tasks`, matching `head_sha` to each commit — CI emails only on
failure, so a green that was never looked at is an assumption, not an observation. **Every commit this
session, both repos, is green.**
**`felhom-agent`** — `dd2d1fe` (release path + gate + CI), `0db7766` (REPORT).
`scripts/release-agent.sh` **(new)** · `scripts/check-published-versions.py` **(new)** ·
`scripts/agent_gates.py` · `.gitea/workflows/gates.yml` · `CLAUDE.md` · `CHANGELOG.md` · `REPORT.md`
| Repo | Commit | Task id | Run # | Conclusion |
|---|---|---|---|---|
| `felhom-controller` | `cf48214` (v0.191.0) | 31 | 11 | **success** |
| `felhom-controller` | `5adae4d` (v0.191.1) | 34 | 12 | **success** |
| `felhom-controller` | `9a3c485` (v0.191.2) | 35 | 13 | **success** |
| `felhom.eu` | `179dd79` (hub v0.89.0) | 32 | 17 | **success** |
| `felhom.eu` | `6d359a5` (manifest 0.89.0) | 33 | 18 | **success** |
| `felhom.eu` | `41dbecb` (docs) | 36 | 19 | **success** |
**Tags created:** `felhom.eu/installer-v1.23.0` (the first tag this repo has ever had) and
`felhom-agent/v0.120.0` (retroactive, at `cd6e267` — the commit the published binary was built from;
`configs/` is byte-identical there and at `main`, so nothing depended on the choice).
## 7. `--no-verify`
## 9. Tests and red-proofs
**Not used anywhere.** Every push in this session ran `.githooks/pre-push` (`repo_gates.py --fast` /
`controller_gates.py --fast`) and passed.
| Check | Result |
|---|---|
| `felhom.eu` `repo_gates.py --fast` | all 5 gates OK |
| `felhom-agent` `go build ./... && go vet ./...` | OK |
| `felhom-agent` `go test ./...` | **29 packages ok, rc=0** (read separately from any commit) |
| `agent_gates.py --fast` | `published` correctly **SKIPPED** (hook must not fail on a network blip) |
| `agent_gates.py` (full) | both OK |
**Red-proofs, each demonstrated failing then restored:**
| # | Mutation | Result |
|---|---|---|
| C | one of the sixteen fetches reverted to `raw/branch/main` | **RED** — gate 6a *and* 6b both fired |
| D | assertions 6a **and** 6b removed (every guard the test covers), same bad installer | **zero** mentions of the regression — the guards are what catch it |
| 6c | the manifest before the split | **RED** on its own, before I fixed it — the gate was demonstrated red by the real pre-change state |
| F | `v9.9.9` tagged and not published | **RED**, `binary NOT downloadable (HTTP 404 …)`, rc=1 |
| F′ | the gate **deregistered** from `agent_gates.py`, same bad state | **rc=0, "all agent gates OK"** — restored → `CONVICTED: published`, rc=1 |
**Scenario F measured on real CI, not inferred.** Runs **69** and **70** are on the *same commit*
`0db7766`: **success** before `v9.9.9` existed, **failure** after pushing it. One variable. This also
retrospectively explains runs 67/68. **One deliberate CI failure email reached the operator — that was
this proof, not an incident.** I could not read CI's own step log: the jobs endpoint needs a Gitea API
token, and the only credential available (`~/.docker/config.json`) is a registry password that the API
rejects — so the controlled before/after replaced the log rather than an assumption standing in for it.
## 10. No version bumps, nothing built
`felhom-agent` **v0.120.0** unchanged (no Go code changed). Hub **v0.89.0** unchanged (no hub file
touched). The installer's `SCRIPT_VERSION` **did** go 1.22.0 → 1.23.0 — the installer is not in §12's
no-bump list, its behaviour changed materially, and the tag derives from it. No image, binary or
golden was built.
## 11. Register
| ID | Outcome |
|---|---|
| **R-110** | **CLOSED — SHIPPED** (installer v1.23.0), both-channels condition honoured, though not in the shape the ruling assumed |
| **R-115** | **CLOSED — SHIPPED** (`release-agent.sh` + `check-published-versions.py`, no bump) |
| **R-182** | **RE-SCOPED — the direction reversed** by Part 0's measurement; still open, now correctly described |
| **R-183** | **NEW, and CLOSED the same session** — binary and configs came from two different refs |
| **R-184** | **NEW, open** — nothing stops the hub vouching a version that was never released |
**IDs established free:** `^| \*\*R-183\*\*` / `^| \*\*R-184\*\*` in `OPEN-ITEMS.md` → **0 rows** each;
all other hits are this session's own code and changelogs (forward references I wrote). `R-185` → 0
hits anywhere and remains free.
## 12. Observations — noticed, documented, NOT acted on
1. **The gate cannot see what is vouched** — filed as R-184 rather than papered over. Closing it needs
either a hub credential in CI (operator's call) or a check at vouch time in the hub (better: fails
closed where the mistake is made, needs no new credential).
2. **A suppressed operator alert leaves no row at all.** The cooldown returns before `LogNotification`,
so the hub's own records cannot distinguish "never happened" from "held back". Recorded inside
R-182 because it is what made that row take a day to get the right way round.
3. **`on: [push]` fires CI for tag pushes too.** Useful (it is how Scenario F was measured), but it
means a tag push runs the full gate set — worth knowing before anyone adds an expensive gate.
4. **`felhom.eu` CI still runs `--fast`.** Correct today, since all its gates are network-free; if a
network gate is ever added there, that workflow needs the same change the agent's just got.
## 13. Teardown
Probe container, probe sync tree and probe tag (`probe-r110-delete-me`) removed; the red-proof tag
`v9.9.9` deleted (`git ls-remote --tags` → only `v0.120.0`); the Scenario A marker reverted from
`main` and the installer confirmed byte-identical to the published tag; the throwaway in-cluster curl
pod removed; the hub DB copy is scratch-only and holds no secret material in any committed file.
+80 -73
View File
@@ -1,6 +1,6 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-08-02.**
**Updated 2026-08-03.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority on open work; this
> page restates part of it in plain words, and **nothing may exist only here**. **Not `CONTEXT.md`**,
@@ -27,94 +27,101 @@ time, and an app switched off deliberately stayed off every time.
also delete it. A daily snapshot is armed as a stopgap, and we have never restored from that copy.
*(R-95, R-87)*
**Three apps out of fifty-three kept their data where backups never looked.** They reported healthy;
the data would vanish on the next update. Two are fixed, the third is now clear to fix because it is
installed nowhere. *(R-156)*
**A full disk tells you about ONE app and silently swallows the rest.** Yesterday this was written
down the wrong way round — as *too many* emails. Measuring the receiving end reversed it: of nine
refusals the machine reported today, **two emails were sent**. When two apps are refused in the same
second you are told about one of them, and the other leaves no trace anywhere — not an email, not
even a line in the log saying it was held back. So a second app can be going unbacked-up while you
have already been told the problem is handled. It is the same fault we fixed once before for
whole-machine backups, in a second place that never opted into the fix. *(R-182)*
**Local backups get 20 GB while apps get 50 GB.** An app that outgrows the smaller space stops being
backed up locally — and the off-site copy is made from the local one, so that stops too. Nothing is
lost: the last good copy is kept intact. *(R-163)*
## What shipped recently
**When that happens, only one page says so** — no email, no alert. The page that answers "is this app
backed up?" is the one that stays silent. *(R-158)*
**Pushing the installer no longer publishes it.** The script that runs as root on a brand-new
machine was copied from the main branch and served within thirty seconds, so pushing it *was*
publishing it, with no staging and no way back but another push. It now comes from a **labelled**
version: publishing is moving the label, and undoing it is moving the label back — about half a
minute, no deploy. The website is untouched by this and still updates in thirty seconds, because a
typo fix must never need a release. Proven by actually doing it: a real push changed nothing that
anyone downloads, moving the label published it, moving it back restored the previous bytes exactly.
**The checks now have two nets, and the second one emails you.** Every repository has one command
that runs all of its checks; it runs by itself before every push and refuses a push that fails. That
one lives on the workstation and can be skipped. So the build server now runs the same checks again,
on a machine that does not care who pushed or what they typed — and **when they fail it sends you an
email**, because a red mark on a page nobody watches is not a warning. Proven with a real broken
change, not assumed. The one thing it still cannot do is *stop* the change: every change here goes
straight to the main copy with no review step, so there is no point in the road for it to stand at.
It notices, quickly, and tells you. *(R-29, R-161, R-168, R-169)*
**The catch that would have made it cosmetic was found and covered.** While it runs, the installer
fetches sixteen more files — not nine, and from the *agent's* repository, not the website's. They now
come from the same version of the agent the machine is installing. That closed a real fault nobody
had noticed: a new machine was getting the agent's tested program and its untested settings files, in
one install, from two different places. *(R-110, R-183)*
**A filling disk now warns the customer before anything breaks, and a failed backup now reaches you.**
Until today the first sign that a disk was filling up was a backup that did not happen — nothing said
anything beforehand. Two things changed. The customer is now warned while there is still room to act,
naming the drive and how much space is left, in plain Hungarian that says what to do about it. And
when one app's backup fails for any reason, **you** are told which app and why, with the disk figures
attached — the page that answers "is this app backed up?" was, until now, the one page that never
said. The customer is deliberately *not* told about that second one: they can free up space, but they
can do nothing about a backup that failed, so telling them would only alarm them.
**Releasing the agent now publishes it, in one command.** Putting a built agent where a new machine
can download it was a step someone had to remember, and it was forgotten three times in five days —
the last time leaving both demo machines running a version nobody could download, so a rebuild would
have quietly installed the *older* one and reported success. There is now one command that builds,
labels, publishes and then **downloads it back to check** — and a check that refuses to stay quiet if
a released version cannot actually be fetched. Proven by making CI fail on purpose and then go green
again on the same code. *(R-115)*
Both were proven on the demo machine by actually filling a disk. One detail is worth knowing because
it is why there are two rules and not one: the serious warning fired when free space dropped below a
fixed amount while the disk was only 91% full — a percentage on its own would have missed it.
**The backup partition is gone and both demo machines run on the new shape** — wiped, rebuilt and
taken through the whole customer journey on 3 August, by two deliberately different routes so the disk
shape and the delivery route are both proven. The space a backup can use went from 19 GB to 65 GB on
the small machine and 45 GB to 233 GB on the big one. Their previous demo apps and data are gone; that
was the point of a wipe, and you approved it. *(R-165, R-178)*
**These went in *before* the partition change deliberately.** The partition being removed is also a
barrier against a runaway backup filling the space the machine needs to run; putting the warnings in
first means that when it comes down, the thing watching is already working and already tested.
*(R-167, R-158)*
**What replaced the wall now watches the right moment.** The wall was quietly keeping a runaway
backup from eating the space the machine needs to run. As first built, that replacement was checked
too late — the big write happened first, unchecked — while still promising your last good copy was
untouched. Fixed and proven on 3 August: the machine decides once, per app, **before it writes
anything**, and that one answer covers all three steps, so a refused app writes nothing, is not
restarted, and the promise is now literally true. It also stopped being blind to size. Nothing is ever
deleted to make room. *(R-181)*
**The last of the three apps that never saved their data is fixed.** Installed nowhere, so nothing was
stranded — checked on both demo machines and in the fleet list rather than assumed. Proven by the check
that caught it, run in both directions: it clears the fixed version and still convicts the old one.
*(R-156)*
**A filling disk warns the customer before anything breaks, and a failed backup reaches you** — the
customer while there is still room to act, naming the drive and the space left; you when one app's
backup fails, with the disk figures. The customer is deliberately not told about the second: they can
free space, but they can do nothing about a failed backup. Both proven by filling a real disk. There
are two rules and not one because the serious warning fired on free space while the disk was only 91%
full — a percentage alone would have missed it. *(R-167, R-158)*
**The checks have two nets and the second emails you.** Every repository has one command that runs all
its checks, before every push. That one can be skipped, so the build server runs them again and emails
you on failure. It cannot *stop* a change — everything goes straight to the main copy with no review
step — but it notices quickly and tells you. *(R-29, R-161, R-168, R-169)*
## What we're working on
- **Now:** the last app whose data was never saved; today's decisions written down.
- **Next:** merging the small backup partition into the large one — **the warnings for it are already
done and working**, so this step is now only the partition change. It needs one decision from you
first (below).
- **After:** rebuilding how the machine records whether an app is meant to be running.
- **Now:** nothing outstanding from today — the reserve, the last unsaved app, and both of your
decisions are all built and proven.
- **Next:** the alert that tells you about one app and swallows the second *(R-182)*.
- **After:** the off-site copy that the machine making it can still erase *(R-95, R-87)*.
## Waiting on you
- **How a new version reaches a machine.** Pushing the installer publishes it — half a minute later
every new machine downloads it, with no staging and no way back but another push. And publishing is
a step we remember rather than one the release performs, forgotten twice: a fix can be live here
and still not reach a new machine. Nothing is installing today, so this is the cheapest moment to
settle both. *(R-110, R-115)*
- **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a
session log; nothing suggests anyone else saw it. *(R-132)*
- **The partition merge: one decision, now measured.** Removing the backup partition also removes a
barrier — today a runaway backup is refused on its own and cannot touch the space the machine needs
to run; afterwards it can, and a machine out of that space is stopped, not slow. So: do we add a
hard stop that refuses a backup before it eats the last of the room, or do we rely on the new
warnings? The recommendation is the hard stop, because it keeps exactly what the barrier gave us.
**Second question, which only you can answer:** does the tester's box need converting in place, or
can it be reinstalled? It does not report to the hub, so nothing here can tell. *(R-165, R-176)*
- **Nothing else.** Both decisions you took on 3 August are now built and proven. One small question
will come back later: the automatic check cannot see which version you have told machines to
install, only which ones exist — closing that either needs a password given to the build server or
a check inside the hub itself. Filed, not urgent. *(R-184)*
## Changed since last update
- **2026-08-02** — The false "host offline" warning is fixed, and the cause was not what it looked
like. The hub's database was supposed to be in a mode where reading a page cannot block a machine's
status update — the code said so, but a one-word syntax difference meant the setting had **never
taken effect**, for the hub's whole life. So opening an operator page could make a machine's report
fail; two failures in a row crossed the half-hour threshold and sent you an alert about a machine
that was up and healthy. It had already done that twice that day. Now genuinely fixed and verified
live. **Also found while checking it: the hub's own database is not in any automatic backup** — it
holds every machine's emergency password and the escrow records. Filed, not yet fixed.
- **2026-08-03** — Publishing became something you do rather than something that happens: the
installer and the agent both moved onto labelled versions with a way back, and a check now refuses
a release nobody can download. Earlier the same day: the reserve now guards the step that fills the
disk and its promise is true, and the last app whose data was never saved is fixed. All proven on
real machines, not just in tests.
Earlier the same day: both demo machines wiped and rebuilt from the new base image and taken through
set-up → install an app → back it up → restore it, with the backup space ceiling gone and measured.
- **2026-08-02** — Boot recovery finished: the machine records what the customer asked for, and waits
for the system to finish starting before deciding what is missing. Six hard resets, everything back
every time. A hole the previous day's change had opened — starting an app whose external drive was
missing — was found by reading the code, reproduced on the demo box first, and fixed the same day.
**A second instance of the same hole was found and fixed today**, on the path that restarts an app
after an interrupted backup.
- **2026-08-02** — The false "host offline" warning is fixed. The hub's database was supposed to be in
a mode where reading a page cannot block a machine's status update; a one-word difference meant that
setting had **never taken effect**, for the hub's whole life. Fixed and verified live. **Also found:
the hub's own database is in no automatic backup** — it holds every machine's emergency password.
Filed, not yet fixed.
- **2026-08-02** — Thirteen mechanical checks had built up across the four repositories and nothing
ran most of them; two were failing quietly, one since 14 July. Both fixed, and the arrangement that
replaced them is described above.
- **2026-08-02** — Decided: the 20 GB backup partition goes away and shares space with app data. That
changes the disk layout, so it happens before any machine is installed outside the house. **Measured
since:** no machine outside the house is registered yet, so this is as cheap now as it will ever be;
and the two demo machines can simply be reinstalled rather than converted.
- **2026-08-02** — Decided: only this machine and the tester's box are protected; every other box,
demo boxes included, may be broken or reinstalled freely. Two of the three apps that never saved
their data are fixed; this page created.
- **2026-08-02** — Thirteen mechanical checks had built up and nothing ran most of them; two were
failing quietly. Fixed. Decided the same day: the 20 GB backup partition goes away; and only this
machine and the tester's box are protected, every other box may be broken or reinstalled freely.
@@ -31,6 +31,7 @@
|---|---|---|---|---|
| Appliance day-0 install: golden image → first boot → auto-confirm (zero clicks) → claimable box | installer, agent, hub, golden | **PROVEN-LIVE** (nested VM) | `DRILL-day0-vm-2026-07-12`, `DRILL-day0-take2-2026-07-12` | First firing on real customer hardware pending → R-1 |
| BYO install: `--mode byo`, mandatory caps, host-mutation disclosure, coexistence guards | installer v1.15+, agent | **PARTIAL** | `DRILL-GL6-2026-07-08` (demo box); GL-8 coexistence fixes | Peti clean-slate reinstall on proxmox2 is the first real BYO run of the current path → R-1 |
| **The installer is PUBLISHED, not pushed — the artifact that runs as root on a virgin box is served from a version-controlled ref, and rolling back is one act** | scripts **v1.23.0** + `manifests/webpage.yaml` (R-110, operator ruling option (b)) | **PROVEN-LIVE (2026-08-03)** | `scripts/CHANGELOG.md` v1.23.0 + `REPORT.md`. **Proven by HTTP against the real URL, not from a pod's filesystem.** *Scenario A:* a real push to `main` without moving the tag left the served script **byte-identical** (`sha256 2f859555…`), and a marker comment planted in that very commit was **absent** from the served bytes, while the website tree advanced to the new commit in the same observation — both halves of the split in one measurement. *Scenario B:* moving the tag published in **~40 s** (`sha → ea2b4aa9…`, marker present) and moving it back restored **exactly** the pre-publish sha. `https://felhom.eu/` returned 200 throughout. *P-A, measured BEFORE the manifest was touched because the model rests on it:* git-sync v4.4.0 follows a tag **and notices a moved one** (`update required … local:<old> remote:<new>` → `updated successfully`) | **Two syncs, deliberately: the WEBSITE still tracks `main`.** Pinning both would turn every copy edit into a release, which makes the release meaningless and the site slow to fix. **Publish** = cut `installer-v<SCRIPT_VERSION>` + bump the manifest `--ref` + sync; **roll back** = move the tag back, which needs **no ArgoCD sync and no deploy**. Continuity is structural rather than lucky: both trees are seeded by init containers so a fresh pod is not Ready until the tag is checked out, and `maxUnavailable` rounds to 0 on one replica, so a failed scripts-init leaves the OLD pod serving — the failure direction is *no update*, never *no `/scripts/`*. **The URL never carried a ref**, so the bootstrap script and the hub's day-0 command follow the tag with no edit and **no hub change**. The installer's own sixteen run-time fetches are a separate channel pinned to the AGENT's version (**R-183**), because they are the agent's configs and not this repo's — leaving them on `main` would have made the whole change cosmetic |
| Bare-metal Felhom ISO (blank hardware → zero-touch auto-install → first-boot `host-install`); selectable UEFI loader; **universal secret-free / operator-bind** mode | scripts v1.19.0 (`scripts/iso/`) + hub v0.62.0 + assistant container | **PROVEN-LIVE on TWO different boards** (N100 2026-07-18; HP t740 2026-07-21) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — the full chain on real metal in a single pass:** the generic reusable pairing ISO (v1.20.0, `--loader mkimage`, SB off) booted the cheap AMI board that F1 had blocked, installed unattended, and the box **self-registered as an unclaimed appliance at 16:17:14 — the same second it first booted** (`appliance_registrations` id=3), then bound → credential-delivered → day-0 SUCCESS 16:32:32 → floor-lifted to current. **F1 is closed on physical hardware.** Prior nested legs: slice A `SPIKE-baremetal-iso-2026-07-16` (build gate, disk-filter fail-safe, stub→host-install fetch); slice B RUNBOOK-B (shim boots+installs OVMF SB-enforcing + SeaBIOS; `--loader mkimage` boots+installs SB-off; mkimage SB-enforcing **FAILS** `Access Denied`; surgery byte-identical); **slice C (2026-07-17): the GENERIC secret-free ISO** — box self-registers as an unclaimed appliance (`POST /api/v1/appliance/register`, one-shot poll delivery, 404-no-oracle — all live-verified through the public ingress), operator binds on the Hosts page, hub delivers credentials once; bootstrap harness proves direct(zero-appliance-calls)/pairing/delivery; artifact proven secret-free (baked env = hub URL only) | **F1 loader caveat:** `--loader mkimage` fixes cheap AMI firmware that can't USB-boot the stock GRUB — UNSIGNED → **Secure Boot must be OFF**; default `shim` keeps SB. **Slice C bind is operator-password-gated** (CC stages, Viktor binds) → the live boot→register→bind→day-0 composition + physical N100 boot fold into the supervised rehearsal (R-1). Customer-facing **self-bind page = R-27 slice 1 SHIPPED (hub v0.66.0, 2026-07-17)** — see the dedicated self-bind row | **Second board, 2026-07-21 (demo-hp, HP t740 / Ryzen V1756B / AMI M42):** the whole chain ran on virgin hardware in one pass — armed install → self-registration as an unclaimed appliance → operator bind → day-0 → running guest 9201 + agent 0.92.1 as `demo-hp-bb76ea`. **The shim loader booted with Secure Boot ENABLED**, which retires the assumption that Felhom installs need SB off — that was an N100-firmware workaround. The exact-serial disk filter took the system SSD and left the box's 1TB NVMe untouched/unenrolled on hardware it had never seen. Two failures filed rather than smoothed over: **R-59** (no DHCP → the installer baked a static fallback instead of aborting) and **R-61** (baked root password unknowable → no console access).
| Box survives a wrong-NIC install: hub-unreachable first boot → legible Hungarian console screen (NIC table + remedy) + NIC sweep self-heal (bounded DHCP + hub probe per NIC, success-only persist), and the baked root password is operator-knowable (`<iso>.rootpw.txt`) | scripts v1.24.0 (`scripts/iso/felhom-bootstrap.sh` `network_gate`/`sweep_nics`, `build-felhom-iso.sh` rootpw emission) | **PROVEN-LIVE (nested drill — nested ≠ metal: metal proof rides the next real multi-NIC install)** | `audits/SPIKE-firstboot-nic-sweep-2026-07-22.md` — dead-NIC install from the virgin v1.24.0 ISO baked the 192.168.100.2 fallback (WITH a dead default gateway), the R-59 screen painted on the console (screendump captured), and after the cable move the box swept to the working NIC, re-leased and **self-registered at the hub unaided in under a minute**; the drill also caught + fixed the stale-fallback-route trap (flush before the bounded dhclient) and verified the emitted rootpw against the installed box's shadow hash | R-59 ships as a first-boot gate, not an install-time abort (recorded deviation — the fallback is the auto-installer's own, initrd hook out of scope); sweep is structurally first-boot-only (`state.json` gate + unit done-flag condition); a box past install-start gets the screen but its interfaces are never touched |
| Customer claim: one-time emailed code → customer sets own password (bcrypt, operator never sees it) | controller v0.122, hub v0.50 | **PROVEN-LIVE** (drill VM) | `DRILL-day0-vm-2026-07-12` §10/F-4 (gate ON via real edge; claimed, code consumed) | Never executed by a non-Viktor human → R-3. **Deliverability (R-4), gmail half DONE 2026-07-18:** the rehearsal's claim email was the first sent under the tightened DMARC `p=quarantine` and **landed in the gmail Inbox, not spam** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`). **freemail.hu remains Viktor's open half.** (Dropped mis-cited `CAMPAIGN-4` F-C — that is the escrow-claim 502, not password claim) |
@@ -86,6 +87,7 @@
| Box survives a **site/network change** (relocation, different subnet, DHCP re-lease) with the control plane intact | agent v0.96.0 (island NIC), host-install v1.19.0, controller (unchanged), bootstrap | **PROVEN-LIVE (2026-07-25)** | **R-50 SHIPPED and deployed to the whole fleet.** The control plane now rides a host-internal, portless island bridge (`vmbr9`, `169.254.253.1/30`↔`.2/30`) with a fixed private address that no LAN/DHCP/site move can invalidate. Proven end-to-end: the spike's F1 replay (renumber the LAN → agent stays bound on the island, control plane HTTP 200; the LAN-literal contrast reproduces the original `bind: cannot assign requested address` daemon-death) + cold-reboot survival (`SPIKE-island-bridge-2026-07-25.md`), the migration runbook run verbatim (`RUNBOOK-island-migration.md`), a fresh provision auto-attaching the island `net1` (A4), and the live migration of **both demo boxes** (demo-hp + demo-felhom, 2026-07-25) — island `/storage` HTTP 200, LAN DNS pinned to the LAN IP (Finding-1), **apps served throughout (0 container restarts)**, hub reporting 0.96.0. **Origin:** `audits/AUDIT-vacation-remote-ops-2026-07-20.md` — the real relocation where the agent's LAN-literal bind took storage/PBS/quiesce/restore-test/DR down silently; that is now structurally impossible on a migrated box | **Fleet: DONE.** Remaining: **R-74** — bring the island to Peti's 2-node cluster (SDN vnet / bridge parity), its own supervised runbook. Related historical: R-51 (dead-primary alerting), R-52 (boot desired-state reconciliation), both shipped |
| **The customer is warned BEFORE a filesystem fills** — per filesystem, in Hungarian, naming the drive and the free space, edge-triggered | controller **v0.191.0/.1/.2**, hub **v0.89.0** (R-167, decision D-c) | **PROVEN-LIVE (2026-08-02)** | `audits/SPIKE-r165-mp1-merge-2026-08-02.md` (context) + `felhom-controller/REPORT.md`. Exercised on guest 9201 against a REAL filesystem (`/mnt/sys_drive` filled with `fallocate`): **`disk_warning` at 90% used / 4.7 GB free** → hub `notification_log` `customer | disk_warning | sent` with the dynamic Hungarian rendered; grown to 1.7 GB free → **`disk_critical`** → `customer | sent`; file removed → `critical → ok … cleared silently, re-armed` and the persisted state emptied. **Exactly two events across three boots** — the boot in between produced none, which is the edge trigger holding | **Nothing warned before this.** The only prior signal was the healthcheck's generic `health_degraded` at 90%, for REGISTERED STORAGE PATHS ONLY — it never looked at the docker area or the system-data area, never gave a free-byte figure and never named a drive. **The two event types already existed with NO PRODUCER** (`disk_warning`/`disk_critical`: allowlisted, copy'd, in `DefaultEnabledEvents`, checkbox'd) — the **sixth** *built-but-never-wired* instance here; this ships their producer rather than a seventh near-duplicate type. **Two threshold terms, whichever trips first, and the live proof vindicated the design:** the critical crossing fired on the FREE-BYTE term (1.7 GB) at only **91%** used — a percentage-only rule would have missed it. The hub's generic `customerMessages` entries were REMOVED, because `FormatCustomerEmail` prefers the entry over the message and would discard the label and figures. **Known gap → R-177:** there is no operator-triggerable run-now path; the check is daily 03:30 + once at startup, so confirming a cleared warning on a support call needs a controller restart or a wait |
| **A failed per-app Tier-1 backup reaches the OPERATOR** (app, error, and the target filesystem's used/free bytes at the moment of failure) | controller **v0.191.0**, hub **v0.89.0** (R-158, closed by R-167) | **PROVEN-LIVE (2026-08-02)** | `felhom-controller/REPORT.md`. Two real capture failures on guest 9201 (`mkdir …/backups: permission denied`) → both accepted and stored by the hub, `operator | recovery_unit_capture_failed | sent`, and the positive observable **`customer | recovery_unit_capture_failed | skipped | operator_only`** read from the hub's `notification_log`. One event per app, loop continuing | **Before this the failure was a `[WARN]` line and nothing else** — the manager carried three notify seams and none for the unit capture, so `/backups/apps`, the page you open to ask whether ONE app is backed up, was the one page that never said. **Deliberately NOT `backup_failed`:** that type is customer-enabled by default and carries Hungarian copy, so reusing it — which R-158's own proposal said — would email the customer about a failure they cannot act on. **D-c routes it to the operator and overrides the proposal.** Operator-only is enforced by `notify.operatorOnlyEvents`, NOT by the absence of a `customerMessages` entry (the v0.78.0 defect); a red-proof removing the register entry shows the customer receiving it |
| **A local backup is bounded by the box's FREE SPACE, not by a partition set at build time** — the appliance ships ONE data volume, and a capture that would exhaust it is refused per app rather than allowed to stop the container runtime | golden `build-golden.sh` **v3.0.0**, agent **v0.120.0**, controller **v0.193.1** (R-165 / D-a / B2, completed by R-181) | **PROVEN-LIVE (2026-08-03) — BOTH halves** | `REPORT.md` (R-178 reinstalls) + `audits/SPIKE-r165-phase0-2026-08-03.md` (P1/P2/P3) + the bake transcript. **The golden bake is real evidence and is cited as such:** `build-golden.sh v3.0.0` produced `including mount point mp0 ('/var/lib/felhom')` with **no `mp1` line at all**, and its own guards printed `/var/lib/docker is a real mount`, `/mnt/sys_drive is a real mount` and `both paths are ONE filesystem`. Archive published (registry HTTP 200, sha `54e2a4c4…`). The B2 floor is unit-proven with 3 red-proofs and live on 9201 | **The row's FIRST clause is now PROVEN-LIVE; its SECOND is not, and they are separated deliberately.** **Proven (R-178, 2026-08-03):** *"a local backup is bounded by the box's FREE SPACE, not by a partition set at build time"* — both demo boxes reinstalled from this golden, by two different supply paths (demo-hp `--golden <local volid>`; demo-felhom the normal manifest route with **`verified sha256 54e2a4c431daf580… matches the hub manifest`**), each showing `mp0` at `/var/lib/felhom` with **no `mp1`**, both consumer paths real mounts on ONE filesystem (`stat -c %d` = `64519` on all three), 3/3 reboots each, and claim → deploy → backup → **restore** with a planted marker returning byte-identical. Space available to a recovery unit measured at **65 GiB / 233 GiB**, against the **19 GiB / 45 GiB** those boxes' `mp1` slices offered. **NOT proven — and measured FALSE in part:** *"a capture that would exhaust it is refused per app rather than allowed to stop the container runtime"*. The floor fired live for the first time (demo-hp 06:40:03) and does refuse per app, delete nothing, and alert — **but it is checked only in `captureAllRecoveryUnits`, while `runVolumeDumps` writes the bulk with no floor check at all**, so the leg that exhausts the volume is the unguarded one; and the refusal's claim that the previous unit is untouched was measured false (a 182,272 B dump replaced by 2,147,666,432 B under a manifest still dated 06:34:26). → **R-181, CLOSED THE SAME DAY (controller v0.193.0 + v0.193.1) and the second half is now PROVEN-LIVE TOO.** The reserve became a **per-app, per-run ADMISSION decision** taken before the app's FIRST write and covering all three legs (DB dump, volume dump, capture) — they write under one per-app root, which is what lets one verdict cover them honestly — and it gained a **size term**, so an app is no longer admitted at 96% and then allowed to write 2 GB. **Re-proven by filling demo-hp deliberately, once for EACH term, using the method that found the defect.** *Headroom @ 08:59:46* (906 MB free / 99%): both apps refused, **the whole `backups/primary` tree byte-identical — `TREE_SHA` 111d1760c18d3440f700634ab325f8b8 before and after**, opengist's tar still at its original 182,272 B; **no `Stopping <app> for safe volume dump` line at all**, which is the positive-by-absence observable that matters because that line IS present in the 08:58 baseline run; 0 volume dumps; one alert per app, HTTP 200. Space freed, re-run @ 09:01:33 → both captured normally. *Size @ 09:03:00*, reproducing the original sequence with a real 2 GiB file in opengist's volume (previous tar **2,147,666,432 B**, the exact figure the defect was measured at) and the filesystem at **91% used / 2.9 GB free — both headroom terms deliberately clear**: opengist refused `(size)` while **privatebin was ADMITTED and dumped normally**, proving the term is per-app rather than a global halt. **The refusal's wording was NOT weakened to fit** — the behaviour moved so the wording became true, and it is verified by tree fingerprint rather than by reading the log line, which is what lied. The `fallocate` instrument was re-proven on the rebuilt box before use (5 GiB step moved guest `df` while thin-pool `data_percent` held **36.83 → 36.83**), and teardown returned the pool to **29.43%**, below its own baseline. The golden **is now VOUCHED** (2026-08-03, hub `Artifact manifest set: … golden=0.192.0`), so fresh installs pick up the merged layout. Every box in the field that has not been reinstalled is still on the SPLIT layout and is unaffected: nothing assumes the merged shape at runtime, the controller's system_data_path is a path rather than a volume, and agent v0.120.0 FOLDS the retired `-sysdata-grow` into the single grow so an older `felhom-host-install.sh` still provisions the same total capacity |
| Soft-quota: usage bar, pre-push enlargement block, customer notification | controller v0.109/134, hub v0.41/55 | **PROVEN-LIVE** | 6D/6E; hub OffsiteChecker | |
| **A customer (not the operator) performs a restore via UI alone** | all | **MISSING** (as evidence) | — | Alpha will produce this; script it into R-3. **2026-07-19:** the C6 evidence attempt ran and found a **product gap instead of evidence** — `audits/DIAG-immich-restore-2026-07-19.md`. A customer-driven UI restore of a DB-indexed app cannot currently succeed (R-43 file-only restore, R-44 stale dump), so this row cannot flip until those close. Row stays MISSING **by finding, not by absence of attempt** — the rehearsal system working, not failing. **2026-07-19: the blocking product gaps are CLOSED in controller v0.148.0** (R-43 + R-44 shipped), so this row is now blocked only on the evidence run itself, not on missing capability. It flips the moment the §9 acceptance produces screenshots + the outcome flash + a snapshot ID. **2026-07-19 round 2 — PARTIAL EVIDENCE ONLY, row NOT flipped** (`audits/DIAG-immich-restore-round2-2026-07-19.md`): a deliberate run from snapshot `49e7cb46` did recover all 11 assets (`status=active`, files resolve), but the operation **reported failure** and left immich reporting schema drift, because the replay aborted against the running app (H4). Photos back ≠ clean acceptance. **2026-07-20: H4 closed in controller v0.153.0 (R-47) on BOTH paths, AND THE EVIDENCE RUN HAPPENED.** *(The "closing in v0.149" wording above was wrong — v0.149.0 was the F3 dashboard fix; R-47 shipped in v0.153.0.)* The C6 drill ran end-to-end **through the UI**: photos deleted, **trash emptied**, the full files+database restore pressed on `/backups/restore`, 40 files placed + 1 DB dump replayed rc-0, 11 assets back, no drift, timeline visually confirmed. The method note below is now DEMONSTRATED, not merely written down. Evidence: `felhom-controller/REPORT.md` 4e. **Residual: the run was performed by the OPERATOR, not by a customer** — for this row literal wording the alpha still owes one genuinely customer-driven pass, but no product gap blocks it. Method note for R-3's script: deleting in an app's own UI usually means *trash*, not deletion, so a drill written that way merges 0 files, flashes success and proves nothing — a real drill must empty the trash **and** verify the app's *content*, not the file count **Lane split → `07-backup-architecture.md` §3**: this row is Lane 1 (customer, unassisted). §8 rows 1–5 are the routes it would exercise |
@@ -545,11 +545,61 @@ decision D-c; controller v0.191.x + hub v0.89.0).** The last sentence of this se
**operator-tier** (`notify.operatorOnlyEvents`) and deliberately not `backup_failed`: a customer can
take no action on a capture failure.
**A caveat this section must carry, because the bound below depends on it.** The `mp0 50G / mp1 20G`
table above is **demo-hp's** shape, not the fleet's — demo-felhom ships `mp0 200G / mp1 50G`, where
the same bound is ≈ 49 GB / ≈ 24 GB, and the golden's own defaults are `16 G / 8 G` before provision
grows them. **The bound below is a FUNCTION of `mp1`, not a constant.** Measured 2026-08-02,
`audits/SPIKE-r165-mp1-merge-2026-08-02.md` M1; correcting the numbers throughout is **R-175**.
### 7.5.1 — THE CEILING THIS SECTION DESCRIBES HAS BEEN REMOVED (2026-08-03, R-165 / decision D-a)
**Everything above describes the SPLIT layout, which is now the legacy shape.** A golden built by
`build-golden.sh` **v3.0.0** ships **one** data volume; `mp1` does not exist. Both consumer paths are
binds of subdirectories of it (variant **V-c**):
```
mp0 -> /var/lib/felhom ├─ docker/ --bind--> /var/lib/docker
└─ sys_drive/ --bind--> /mnt/sys_drive
```
**So the size bound below no longer applies to a box built from that golden.** A driveless app's
recovery unit is limited by the box's actual free space, not by a partition set at build time. The
mismatch table above (`mp0` 50 G vs `mp1` 20 G) describes what a merged box no longer has.
**R-175, fixed here rather than left standing.** The bound below was stated as the fleet's and was
**one box's**: it is derived from `mp1 = 20 G`, which is demo-hp exactly and never was demo-felhom
(`mp0 200G / mp1 50G`, where the same arithmetic gives ≈ 49 GB / ≈ 24 GB), nor the golden (`16 G / 8 G`
before provision grew them). **Read it as a function of `mp1`, and only for a box still on the split
layout.** Measured: `audits/SPIKE-r165-mp1-merge-2026-08-02.md` M1.
**What replaced the partition's second job — the reserve.** `mp1` was also a BULKHEAD: an overflow was
refused per app with the last good unit byte-identical, and it **could not reach `/var/lib/docker`**,
because that was a different filesystem. On a merged box it can. Decision **B2**, shipped in controller
**v0.192.0**, is that bulkhead made deliberate — a two-term reserve (97% used or 1 GiB free) in
`internal/fillwatch`'s shape, sitting beyond its critical band so the customer is always warned first.
It **refuses per app and never deletes**: nothing on this filesystem is generational, so pruning could
only destroy a different app's only local copy.
**THE CONTRACT, stated as what the code provides (controller v0.193.0, R-181).** The reserve is a
**per-app, per-run ADMISSION decision, not a capture check.** It is taken once for an app, immediately
before that app's FIRST write of the run, and it covers **all three write legs — the database dump, the
volume dump and the recovery-unit capture**. Those three write under one per-app root
(`backups/primary/<app>`), which is what makes one verdict able to cover them honestly.
- **What it guarantees.** A refused app has **nothing written for it in that run**, its previous unit
is **byte-identical**, it is **not stopped**, nothing anywhere is deleted, and the operator gets
**exactly one** alert naming the app, the term that bound and the disk figures.
- **Two terms, two questions.** *Headroom*: is the filesystem already below the reserve? *Size*: would
THIS app's write take it below? The size estimate is the app's previous `.sql` + `.tar` on disk;
with no history the decision degrades to headroom alone, deliberately — otherwise the first backup
is the one that can never happen.
- **Why it is decided lazily and not once per run.** Space changes during a run: app A's dump can put
app B under the reserve, so a verdict taken at run start reads a disk that no longer exists.
- **Why it is never re-decided between an app's own legs.** That is precisely the shape v0.192.0 had —
the two dump legs unguarded and only the capture refused — under which the reserve was consumed by
the very write it exists to bound, and the refusal's *"the previous unit is untouched"* was measured
false. Proven live on demo-hp 2026-08-03 (R-181), fixed the same day, and re-proven by filling the
box for each of the two terms.
- **It sits ahead of `DumpAppVolumesSafe`**, which stops the stack as its first act — a refusal
decided inside it would already have bounced the app it is refusing to back up.
**Status caveat, deliberately explicit:** every box in the field that has not been reinstalled is still
on the split layout and everything above still describes them exactly. This subsection describes what a
box built from golden ≥ 0.192.0 gets. Both demo boxes were reinstalled from it on 2026-08-03 (R-178).
Two things are deliberately **not** recorded here. **The sizing ratio is the operator's ruling**
(**R-163**) — this section states the constraint, not a number. And **the same-device placement is
@@ -0,0 +1,183 @@
# SPIKE R-165 Phase 0 — the two probes the merge spike left unmeasured
**Date:** 2026-08-03 · **Author:** Claude Code · **Status:** MEASURED — **no layout changed**
`SPIKE-r165-mp1-merge-2026-08-02.md` named two things as unmeasured and both are load-bearing. This
document measures them. **It changes nothing**: no golden rebuilt, no box reinstalled, no guest config
edited outside the throwaway probe guests, which are destroyed at the end.
**Host: `felhom-pve` (N100) — Tier 0.** demo-hp would have been the default per the 2026-07-25 ruling,
but `target-selection.md` records its `local-lvm` as an **over-subscribed thin pool backing live guest
9201** (~144 GiB allocated over ~54 GiB) where filling it corrupts every guest, and it holds **no
container template**. felhom-pve has the template and 258 GiB free at 29% pool usage. Both are Tier 0;
this picked the one where a probe cannot damage a live guest. **Neither DooPlex, `ep0`, nor the
colleague's box was contacted at any point.**
---
## P1 — does a PRE-merge archive restore cleanly into the merged world?
**Verdict: PASS.**
### Method
The spike reasoned from `mountParity` that it should pass and said plainly that this had never been
executed. It is now executed, on real hardware, with the real restore-test path — not a hand-assembled
restore.
- **Archive:** `local:backup/vzdump-lxc-9201-2026_07_28-17_43_05.tar.zst` from **demo-hp**, 1.68 GB.
Confirmed pre-merge by its own vzdump log rather than by assumption:
```
including mount point rootfs ('/') in backup
including mount point mp0 ('/var/lib/docker') in backup
including mount point mp1 ('/mnt/sys_drive') in backup
```
- **Command:** `felhom-agent --selftest=restore-test -archive <volid>` on demo-hp.
### Measured result
```
"pass": true,
"verified": "boot+running",
"mount_parity": "ok",
"duration_seconds": 84.2,
"mount_inventory": [
"mp0=/var/lib/docker (50G)",
"mp1=/mnt/sys_drive (20G)",
"mp8=/mnt/felhom-drives (throwaway for the archived bind)",
"mp9=/etc/felhom-bootstrap (throwaway for the archived bind)"
]
=== selftest=restore-test OK (scratch 990000 restored+booted+verified+torn-down in 1m24s) ===
```
`mountParity` was **not** relaxed, weakened or touched in any way.
### The limit of this result, stated rather than glossed
**It was run with the CURRENT agent (v0.119.0), because the merged agent does not exist yet** — Part 2
sits after this session's STOP. What it proves is that `mountParity` compares the **archive** against
**its own restore**, so a pre-merge archive recreates its own `mp0 + mp1` in the scratch guest and the
two agree. That comparison never consults the *host's* golden layout, which is why the merge cannot
invalidate it — **provided Part 2 honours its own constraint not to touch `mountParity` or the restore
path** (§5, §12 of the task). It is an 84-second command and **should be re-run once the merged agent
exists**, which is cheap and turns a sound inference into an observation.
---
## P2 — which S1 variant actually works on this platform?
**Verdict: all three probed variants are mechanically clean. They are separated by SCOPING, not by
mechanics — and the deciding fact was not in the task's table.**
### Method
A throwaway unprivileged LXC per variant (`nesting=1,keyctl=1`, rootfs 8 G + one 10 G volume,
`backup=1`), Docker installed from the same repo with the **same `daemon.json` the golden bakes**
(`containerd-snapshotter: false`, overlay2, log caps). Per variant, measured at first boot and after
**each of three reboots**:
1. both `/var/lib/docker` and `/mnt/sys_drive` present and **writable** (write → read back → delete,
a positive observable rather than an `ls`);
2. **one** filesystem — same source device **and** the same free-space figure for both paths;
3. `dockerd` active and `docker run hello-world` succeeding;
4. then once: the **real bootstrap propagation sequence** (`mount --rbind /mnt /mnt`,
`mount --make-rshared /mnt`, `docker run -v /mnt:/mnt:rslave`);
5. and: what a container mounting `/mnt:rslave` **actually sees** — the scoping check.
No pipe hides a non-zero; every check reports its own rc and the script aborts with `MEASURED-FAIL`.
### The variants
| | volume mounted at | then |
|---|---|---|
| **V-a** | `/var/lib/docker` | `/mnt/sys_drive` = bind of `/var/lib/docker/sys_drive` |
| **V-b** | `/mnt/sys_drive` | `/var/lib/docker` = bind of `/mnt/sys_drive/docker` |
| **V-c** | `/var/lib/felhom` *(neutral)* | **both** consumer paths are binds of subdirectories |
**V-c was not in the task's table.** It was probed *because* the measurements below showed V-a and V-b
each violate a different documented invariant, and V-c is the shape that violates neither. It is
offered as a **measured option for the operator at the STOP**, not adopted — the task is explicit that
a variant is not chosen mid-session.
### Measured results
| check | V-a | V-b | V-c |
|---|---|---|---|
| both paths present + writable | **yes** | **yes** | **yes** |
| ONE filesystem, ONE free-space figure | **yes** | **yes** | **yes** |
| dockerd active + `docker run` — initial | **yes** | **yes** | **yes** |
| dockerd active + `docker run` — reboots **1/2/3** | **3/3** | **3/3** | **3/3** |
| `/mnt` propagation `shared`; container sees `/mnt` | **yes** | **yes** | **yes** |
| both paths still real mountpoints (`findmnt` non-empty — the form the golden's assertions use) | **yes** | **yes** | **yes** |
| container `statfs("/")` reports the merged volume | **yes** (10218772 KiB) | **yes** | **yes** |
| **what a container mounting `/mnt:rslave` SEES** | `sys_drive` only — **8.0K** | `sys_drive` **+ `sys_drive/docker`** — **17.9M** | `sys_drive` only — **8.0K** |
| customer data inside Docker's data-root | **YES** | no | no |
**The mechanical worry was misplaced.** The task flagged V-b's ordering risk — `/var/lib/docker` must
be bound before dockerd starts. An `/etc/fstab` bind is ordered by `local-fs.target`, which precedes
`basic.target` and therefore `docker.service`, and it held **3 reboots out of 3**. Ordering is not what
separates these variants.
### What actually separates them
**V-b breaks a documented scoping invariant, and the measurement is the proof.** The controller
container is started with `-v /mnt:/mnt:rslave`, and the bootstrap script's own comment states the
scope it relies on:
> *"scoped to /mnt, which (Model A) holds only Felhom's felhom-data-namespace mounts, never the
> customer's other on-drive data"*
Under V-b the container sees `/mnt/sys_drive/docker` — **Docker's entire data-root**, 17.9 MB on an
empty probe box and growing with every image and every app volume. That sentence becomes false. The
controller already holds the Docker socket, so this is **not a capability escalation** — but it puts
Docker's internal tree inside the one path the controller's own scanners, the FileBrowser surface and
the data-migration engine (which works *"in-process over the controller's `/mnt:/mnt:rslave` RW
mount"*) treat as Felhom-only.
**V-a breaks the other one:** customer backups live inside Docker's data-root, so `du` on the data-root
stops meaning what it says, and the ordinary operator reflex for a sick Docker — clear `/var/lib/docker`
— destroys every local recovery unit on the box.
**V-c breaks neither**, at the cost of one new mount path and two fstab lines instead of one.
---
## P3 — the golden's four assertions
**Not yet run — it belongs to Part 1, which is after the STOP.** Recorded here so it is not lost:
`build-golden.sh` fails closed on the split in **four** places (`:126`, `:130` separate-mount asserts;
`:315`, `:319` vzdump-exclusion guards), and each retargeted guard must be shown to **abort** against a
deliberately wrong shape before the golden is trusted.
**One measurement already de-risks it:** under all three variants both `/var/lib/docker` and
`/mnt/sys_drive` remain **real mountpoints**, so `findmnt -no SOURCE,FSTYPE <path> | grep -q .` — the
exact form the two existing assertions use — still returns non-empty. The assertions can be
**retargeted with a changed message and an added guard for the single volume**, rather than rewritten
from scratch. They must not be deleted.
---
## Teardown
| layer | action | evidence |
|---|---|---|
| **the machine** | probe guests `9401`, `9402`, `9403` destroyed; P1's scratch `990000` was torn down by the restore-test itself | `pct list` shows only `9201` |
| **the host** | `local-lvm` **112398205 → 107204406 KiB** used (30.73% → 29.31%) — **5.19 GB actually returned**, not merely deallocated | `pvesm status` before/after |
| **the hub** | **none created, and verified rather than assumed.** No probe claimed a box, minted a customer or registered an appliance — none ever ran a controller. The registers hold the same **5 customers** (`david`, `demo-felhom`, `demo-hp`, `peti-felhom`, `sess-f`) and same **4 hosts** (`demo-felhom-8363b5`, `demo-hp-bb76ea`, `drill-r50-0a4f9a`, `sess-f-2670b5`) as before Phase 0 | hub `/` + `/hosts` |
Scratch scripts and logs removed from `felhom-pve` (`/root/p2-probe.sh`, `/tmp/p2-*.log`,
`/tmp/p1-restoretest.log`). `pct list` on both demo hosts shows only their own `9201`.
---
## What this changes about the merge
1. **P1 removes the restore risk from the decision.** A pre-merge archive restores clean with parity
ok; it should be re-confirmed against the merged agent, which is one command.
2. **The variant question is not "will it boot" — it is "which invariant do we break".** Both named
variants work perfectly and each violates one documented guarantee. That is the operator's call and
is the subject of this session's STOP.
3. **V-c exists and is measured.** It costs one extra mount path and one extra fstab line, and is the
only probed shape that keeps both guarantees.
File diff suppressed because one or more lines are too long
+4 -4
View File
@@ -20,7 +20,7 @@
| ID | Item | Size | Status | Notes / map rows flipped |
|----|------|------|--------|--------------------------|
| R-115 | **Publishing is a remembered step — forgotten within eight hours of being documented as forgettable** | M | idea — **WAITING-ON-OPERATOR**, 2026-07-29 | A box installs the agent from a Gitea generic package the hub explicitly vouches, never from git. Nothing in the build, deploy or session-end path publishes or checks that a version was published, so **"deployed" and "installable" are independent states that drift silently**. **Instance 1 — R-111** (morning): 17 agent releases v0.97.0–v0.113.0 stranded; a new customer would have installed without the whole R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT. Found only because the E-2d Phase 0 gate happened to look. **Instance 2 — agent 0.114.0** (same afternoon): the R-113 fix, built and pushed at `b58d7bc`, deployed to felhom-pve, never published — which blocked Session C, since a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix. **The finding is the RECURRENCE, not either instance** — both are fixed. R-111 named this leg in its own text (*"decide whether publishing joins the release train rather than staying a remembered step"*) and closed SHIPPED without it; it recurred the same day, which is the evidence that **a note is not a mechanism**. **Class: → R-29, one layer up** (a control that exists and is never walked) — deliberately NOT given a second ID. **Filed as its own item rather than reopening R-111** because R-111's finding (the channel *was* stale) is closed and verified end-to-end by the E-2d install, while the process defect that caused it is a distinct problem with a distinct fix and a distinct owner. **Operator's decision, mechanisms first:** (a) publish as a step in the build/release path so deployed and installable cannot diverge; (b) a gate that refuses to deploy an unpublished+unvouched version — strongest, fails closed; (c) a session-end checklist entry; (d) accept manual + a pre-Session-C verification. **(a)/(b) are mechanisms, (c)/(d) are reminders — and R-29's whole finding is that reminders do not hold.** No code written when filed, by design |
| R-115 | ~~**Publishing is a remembered step — forgotten within eight hours of being documented as forgettable**~~ | M | **CLOSED — SHIPPED 2026-08-03** (`release-agent.sh` + `check-published-versions.py`, no version bump). Releasing now builds, tags, publishes and **verifies by an independent download** in one act; a `v<semver>` tag with no package fails the gate, and **CI runs the full gate set** so it actually runs. Red-proof measured on real CI: same commit, green before a tagged-unpublished version existed, red after. Residue → **R-184** | A box installs the agent from a Gitea generic package the hub explicitly vouches, never from git. Nothing in the build, deploy or session-end path publishes or checks that a version was published, so **"deployed" and "installable" are independent states that drift silently**. **Instance 1 — R-111** (morning): 17 agent releases v0.97.0–v0.113.0 stranded; a new customer would have installed without the whole R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT. Found only because the E-2d Phase 0 gate happened to look. **Instance 2 — agent 0.114.0** (same afternoon): the R-113 fix, built and pushed at `b58d7bc`, deployed to felhom-pve, never published — which blocked Session C, since a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix. **The finding is the RECURRENCE, not either instance** — both are fixed. R-111 named this leg in its own text (*"decide whether publishing joins the release train rather than staying a remembered step"*) and closed SHIPPED without it; it recurred the same day, which is the evidence that **a note is not a mechanism**. **Class: → R-29, one layer up** (a control that exists and is never walked) — deliberately NOT given a second ID. **Filed as its own item rather than reopening R-111** because R-111's finding (the channel *was* stale) is closed and verified end-to-end by the E-2d install, while the process defect that caused it is a distinct problem with a distinct fix and a distinct owner. **Operator's decision, mechanisms first:** (a) publish as a step in the build/release path so deployed and installable cannot diverge; (b) a gate that refuses to deploy an unpublished+unvouched version — strongest, fails closed; (c) a session-end checklist entry; (d) accept manual + a pre-Session-C verification. **(a)/(b) are mechanisms, (c)/(d) are reminders — and R-29's whole finding is that reminders do not hold.** No code written when filed, by design |
| R-116 | **The drive-absent alarm and its recovery are a mismatched pair — generic on the way out, specific on the way back** | S | idea — **PROVEN LIVE 2026-07-29** | Absent fires `storage_disconnected`; return fires `backup_target_restored`. `backup_target_absent` never fires at all (count 0 across a full Session-C run), so an operator gets an alarm they cannot match to its recovery — exactly what `notifyDriveReturned`'s own comment forbids. Root cause: `notifyDriveAbsent` (`intermediary.go:635-646`) branches on `isTarget[a.Path]` with `a.Path` the GUEST path, and `driveTargetByPath` (`:602-616`) builds it as `out[GuestPath] = d.BackupTarget` — but **the drive is TWO `/disks` rows and the flag and the guest path sit on different ones**: the `felhom-backup` storage row has `BackupTarget: true` (`felhom-agent/internal/localapi/disks.go:211`) and gets a guest path only while classified user-data, while the registry union row has the guest path and **never assigns `BackupTarget`** (`disks.go:265-267`). Absent ⇒ the flagged row loses its guest path ⇒ the union row writes `false` ⇒ generic. On return the rows rejoin ⇒ specific. v0.184.1 fixed the KEYING, not this. **Only reachable because R-113 made the gate fire at all.** Fix likely agent-side; decide the repo first. Blocks E-2's C5. Evidence: `audits/SESSION-C-2026-07-29.md` §5 |
| R-113 | **The drive-absent gate cannot fire on device loss — E-2b's alarm is wired to an unreachable condition** | M | idea — **PROVEN LIVE 2026-07-29** | `planDriveGates` (`felhom-controller/internal/web/intermediary.go:216-262`) treats a path as present by OR-ing in `d.BoundUnderParent`, which the agent derives from `GuestSeesMount()` — *"is this path a mount target in the guest's `/proc/<pid>/mountinfo`"* (`internal/localapi/disks.go:210`). The raw drive mount is a **device-bound systemd unit** and dies with the device; **the agent's own bind under the shared parent is not device-bound and its mountinfo entry outlives the device**, so the gate reads it as present and `notifyDriveAbsent` is never called. Live on a fresh box: target drive hot-detached, agent said `enrolled drive absent by UUID` every 20 s for 4½ min, controller logged **0** `[gate]` lines, hub received **zero** events — neither `backup_target_absent` nor the generic `storage_disconnected`. Not a virtualisation artefact (device-bound-mount vs manual-bind is the same on metal); caveat: SCSI hot-detach, physical unplug not staged. **Sixth instance of seam-built-but-never-wired — E-2b wired the seam to a condition that cannot occur.** Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.2 |
| R-112 | **E-2's degraded banner and offer have no UI consumer — correct endpoint, invisible to the customer** | S | idea — **PROVEN LIVE 2026-07-29** | `GET /api/storage/backup-target` returns byte-exact Hungarian copy (verified on a live box), and nothing in the product asks for it: `grep 'backup-target'` across every `*.html`/`*.js`/`*.css` → **0 hits**; no template references `OfferPath`/`Degraded`/the copy; `resolveBackupTargetState` and `degradedMessageFor` are consumed **only** by the JSON handler, with **no page handler injecting the state**. Decisive contrast: the templates fetch **18 distinct `/api/storage/*` endpoints** — `backup-target` and `backup-target/assign` are the only two with zero references. The handler's own comment calls itself *"the dashboard's source for the degraded banner and the offer"*. v0.185.1 shipped as *"the offer endpoints were mounted where nothing routed to them"* and fixed the **mount**, stopping one layer short of the **render**; its test pins dispatch, not reachability. **Fifth instance of the class. Fix R-114 first** — wiring this alone starts showing customers a wrong message. Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.1 |
@@ -157,18 +157,18 @@ Self-resolves the moment the target answers (the storage read succeeds, sees the
| R-97 | **The whole-guest backup tier has NO failure signal to the hub — `internal/quiesce` never notifies** | S | **SHIPPED (controller v0.177.0 + hub v0.78.0, 2026-07-27)** — **R-97a:** `quiesce.TierNotifier`, a seam (not an import) wired by an init-only setter, edge-triggered on the R-88 breaker ARMING so a failing tier is reported once per run rather than once per retry; recovery rides `recordSuccess`'s existing bool. **NEW operator-only event types** `whole_guest_backup_failed`/`_recovered` — deliberately NOT `backup_failed`, which carries a customer Hungarian template AND sits in demo-felhom's live `enabled_events`, so reusing it would have emailed the CUSTOMER about a backup they cannot act on while it was still retrying. The recovery joins `recoveredPairedDownTypes` because its `info` severity would otherwise be dropped by `severityNotifies` — the operator would hear it break and never hear it heal. **The hub's operator cooldown was keyed `customerID:eventType` alone**, so one tier would have masked the other for an hour; now narrowly extended with a `tier` suffix taken from the event details, leaving every other event type unchanged. **R-97b:** a suppression window keyed to the quiesce CYCLE (not a state test — v0.164.0's `!= StateStopped` filter cannot see an app caught MID-RESTART, which is exactly how BookStack alarmed), consumed at the same single derivation point `classifyRunStates`. Grace = **180 s**, derived from the deploy flow's 120 s health timeout and Mealie's 60 s `start_period`; it **expires**, so an app that genuinely fails to come back still alarms. **PROVEN LIVE end-to-end with a control:** the new type POSTs 200 from inside guest 9201 while a bogus type 400s, and `notification_log` shows **1 operator row, 0 customer rows**. The quiesce→notify link itself is unit-proven only. | On 2026-07-27 three whole-guest backups failed and three quiesce cycles stopped and restarted every customer app stack, and **not one `backup_failed` event reached the hub.** It is not the allowlist — the hub already carries `backup_failed` and `backup_completed` (they are emitted by the controller's *app-data* backup path). The cause is that **`internal/quiesce` does not import `internal/notify` at all**: the tier R-82 built has no route to the hub, so a whole-guest backup can fail indefinitely in silence. The loop's only trace was `app_start_failed` — **info** severity, **Hungarian**, on the **customer** channel — telling the customer BookStack was down (it had been caught mid-restart by the third cycle) without saying why, during an outage the system itself caused. So the one signal that did fire was both the wrong tier and the wrong story. **Shape:** emit `backup_failed`/`backup_completed` from `quiesceAndPollTiers` naming the TIER, operator-tier; and decide whether a quiesce-induced restart should suppress `app_start_failed` the way controller **v0.164.0**'s deliberate-stop filter does — an app the backup stopped on purpose is not a fault. R-88's breaker bounds the repetition but changes nothing about the silence |
| R-95 | **The restic offsite tier's credential CAN DELETE — R-89's "parallel question", now ANSWERED** | M | idea — established read-only 2026-07-27 | **The exposure closed on the weekly PBS tier is fully open on the daily restic tier**, which holds the customer's actual documents and photos and is the only tier that survives losing the box. Established without mutating anything: **(1) Identity** — a per-customer *subaccount* on `storage-box-pool-1` (box 611714, bx11, `u629488`): `u629488-sub1` home `felhom-demo-felhom`, `sub2` peti-felhom, `sub3` demo-hp, each labelled `felhom-customer`. Auth is an **SSH key stored ON THE BOX** (`…/felhom-controller-data/_data/data/offbox/ssh_key`, 0600, beside `repo_password` + a pinned `known_hosts`) — customer-side, not hub-side, so a compromised guest holds it. **(2) Read-write: YES** — the API reports **`readonly=False` on all three subaccounts**, and it is not merely latent: the controller runs `restic forget --group-by host,tags --keep-daily 7 --keep-weekly … --prune` **from the box** (`backup/offbox.go:984`, also `:1070`). Delete rights are exercised on every run. **(3) Append-only: NO, and not expressible** — the repo is built as `sftp:` (`offbox.go:482`); restic's append-only mode requires the **REST server** backend, which plain SFTP cannot provide. **(4) A zero-code mitigation exists and is unused:** the box type carries `snapshot_limit=10` and the API reports `snapshot_plan=null` with **0 snapshots** and `size_snapshots=0`. Hetzner Storage Box snapshots are taken **server-side, outside the SFTP namespace** — an SFTP subaccount cannot delete them — so they are a genuine immutability layer at no extra cost and with no code change. **Rule once for both tiers, per R-89.** Options, cheapest first: enable a snapshot plan (operator click, immediate); split backup-write from prune so pruning runs somewhere the box cannot reach; or move the repo to restic's REST server with `--append-only`. Flips the capability-map row for offsite immutability |
| R-94 | ~~A hand-synced version constant drifts, and the gate that would catch it is never run~~ | XS | **CLOSED — SHIPPED hub v0.87.0, 2026-08-02** | Closed by **deleting** the label rather than deriving it: the Setup command fetches the installer at run time from a 30 s-git-synced website (R-110), so no build-time value in the hub can be true. `hostinstall_gates.py` gate 1 inverted to pin the ABSENCE of a version literal; the tautological `render_test.go` assertion deleted (demonstrated passing at `9.9.9`). Detail: `OPEN-ITEMS.md` R-94 |
| R-110 | **`main` is the installer's publish channel — there is no staging** | S | idea — found 2026-07-29, **WAITING-ON-OPERATOR (a ruling, not a defect)** | `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a `--period=30s`, and nginx serves that working tree directly (`location /scripts/`, `root /usr/share/nginx/html/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`, `:1262`; `runbooks/day0-install.md` C.1) receives. **There is no tag, no pinned-version path, no staging copy and no rollback other than another push** — for the one artifact that runs as **root on a virgin box**, the most privileged thing Felhom ships. **Two consequences worth stating plainly:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29, so the proof run confirms what customers already receive rather than clearing it for release; and the precaution the old R-94 row recorded ("do not point every new box at an installer that has never run") **was never available to take**, because nothing points boxes at a version. **Open question for the operator, not a defect to fix blind:** should `/scripts/` serve a pinned release — a tag-tracked git-sync ref, or a versioned directory (`/scripts/1.22.0/…`) with the hub's generated command naming a version — or is `main`-tracking the accepted shape for a one-operator product where the alternative is a release ritual nobody performs? **Exposure today is zero** (no boxes are installing), which is exactly why it is cheap to decide now. Whichever way it goes, it decides whether R-94 leg (a) makes the label a *fact* (derived from the served script) or keeps it a *claim*. Flips no capability-map row — the map states what the platform does, and this changes nothing about that |
| R-110 | ~~**`main` is the installer's publish channel — there is no staging**~~ | S | **CLOSED — SHIPPED 2026-08-03** (installer v1.23.0). `/scripts/` syncs `installer-v1.23.0`; the website still tracks `main`. Proven by HTTP: a push to `main` left the served bytes byte-identical, moving the tag published in ~40 s, moving it back restored the exact prior sha. The run-time fetches turned out to be **sixteen from the agent repo**, not nine from here — pinned to the agent version instead → **R-183** | `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a `--period=30s`, and nginx serves that working tree directly (`location /scripts/`, `root /usr/share/nginx/html/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`, `:1262`; `runbooks/day0-install.md` C.1) receives. **There is no tag, no pinned-version path, no staging copy and no rollback other than another push** — for the one artifact that runs as **root on a virgin box**, the most privileged thing Felhom ships. **Two consequences worth stating plainly:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29, so the proof run confirms what customers already receive rather than clearing it for release; and the precaution the old R-94 row recorded ("do not point every new box at an installer that has never run") **was never available to take**, because nothing points boxes at a version. **Open question for the operator, not a defect to fix blind:** should `/scripts/` serve a pinned release — a tag-tracked git-sync ref, or a versioned directory (`/scripts/1.22.0/…`) with the hub's generated command naming a version — or is `main`-tracking the accepted shape for a one-operator product where the alternative is a release ritual nobody performs? **Exposure today is zero** (no boxes are installing), which is exactly why it is cheap to decide now. Whichever way it goes, it decides whether R-94 leg (a) makes the label a *fact* (derived from the served script) or keeps it a *claim*. Flips no capability-map row — the map states what the platform does, and this changes nothing about that |
| R-128 | **`ISO_VERSION` "aligns with SCRIPT_VERSION" was a comment nothing evaluated** | XS | **CLOSED — iso v1.26.0, 2026-07-31** | Closed by **correcting the claim, not asserting it**: the ISO is frozen while `felhom-host-install.sh` is fetched at run time from `main` (R-94/R-110), so an assertion would invent a constraint. `build-felhom-iso.sh:45-52`. Full reasoning in `OPEN-ITEMS.md` |
| R-154 | **`[first-boot]` is automated-install-only, and nothing in the tree said so** | XS | **CLOSED — iso v1.26.0, 2026-07-31** | A PVE property, measured with a same-image control (`audits/SPIKE-universal-iso-3-2026-07-31.md` §2); recorded at `scripts/iso/pkg/build-deb.sh:6-11`. Superseded in practice by the `.deb` delivery route |
| R-155 | **`iso-repack.sh` refused any ISO without `auto-installer-mode.toml`** | XS | **CLOSED — iso v1.26.0, 2026-07-31** | **Narrowed, not deleted** — unchanged for `FELHOM_MENU=single` (`iso-repack.sh:121-128`), does not apply to `release` where the file's absence *is* gate G1. Do not remove it wholesale |
| R-156 | **An app's data is neither persisted nor backed up, and it reports healthy** — a template mounts a volume at a path the app never writes, so data sits in the container's writable layer: lost on redeploy, tarred nightly as an empty dir, healthcheck green | S | **DETECTED + GATE SHIPPED; papra REFERRED** (catalog sweep, 2026-08-02) | Found by Campaign 10 on **papra**; the 53-template sweep convicted **gramps-web** and **wishlist** too and **fixed both**. The gate is `app-catalog-felhom.eu/scripts/check-volume-persistence.py` — a runtime probe (`docker diff` + mount occupancy + writability, canary self-test, fails closed); the detector and the gate are one program. **papra is NOT fixed**: the fix needs the app to use `/app/data` or the template to mount `/app/app-data`, so it is referred. Sweep evidence: `app-catalog-felhom.eu/audits/persistence-sweep-2026-08-02/` (53 probe.json). Ranking in `OPEN-ITEMS.md` |
| R-156 | ~~**An app's data is neither persisted nor backed up, and it reports healthy**~~ | S | **CLOSED — all three apps fixed, 2026-08-03.** papra's template now mounts `papra_data:/app/app-data` (the app's own data ROOT, chosen over reconfiguring three env vars so a future upstream path cannot escape). Deployed nowhere — re-checked on both demo guests and the hub fleet view, not inherited. Gate run in BOTH directions: CLEAN fixed, BROKEN reverted | Found by Campaign 10 on **papra**; the 53-template sweep convicted **gramps-web** and **wishlist** too and **fixed both**. The gate is `app-catalog-felhom.eu/scripts/check-volume-persistence.py` — a runtime probe (`docker diff` + mount occupancy + writability, canary self-test, fails closed); the detector and the gate are one program. **papra is NOT fixed**: the fix needs the app to use `/app/data` or the template to mount `/app/app-data`, so it is referred. Sweep evidence: `app-catalog-felhom.eu/audits/persistence-sweep-2026-08-02/` (53 probe.json). Ranking in `OPEN-ITEMS.md` |
| R-157 | **`bootrecon`'s start-ONCE sweep misses the boot orphan it exists to recover** | S | **CLOSED — SHIPPED + PROVEN-LIVE (v0.189.0 + v0.190.0, 2026-08-02)** | Both mechanisms closed. **B:** the container count → recorded intent (R-166). **A:** the single T+5 s observation → a settle-then-sweep window (sample every 5 s, settled after 3 identical samples, one sweep at the end, ends on settled OR a 50 s budget). Budget is 50 s because settle+budget+one retry must fit the 90 s `deadAppBootGrace`; a test rejected 60 s at 95 s. Late recoveries are REPORTED, not hidden by widening the grace. **The fix had its own defect, found live:** sampling the Manager's 10 s-refreshed cache let "settled" mean "the cache did not update" — `sampleBootFleet` now refreshes first. **6/6 hard resets clean on the shipped build**, plus a same-app before/after (missed 18:08:35, recovered 18:18:50) |
| R-170 | **The drive-backed boot gate infers a Stop from a container count** | S | **CLOSED — SHIPPED + PROVEN-LIVE (v0.190.0, 2026-08-02)** | `shouldRecreateOnBoot` reads `desired_state` with the same three-way table as `isBootOrphan`; absent keeps the old `hasContainers` behaviour exactly; `presentStable` untouched. Agreement pinned from both sides against one fixture table. Live: calibre-web (`running`, zero containers) recreated and immich (`stopped`) left alone in the same reboot |
| R-171 | **The boot sweep started apps whose data drive was ABSENT** — a regression from v0.189.0 | S | **CLOSED — SHIPPED + PROVEN-LIVE (v0.190.0, 2026-08-02)** | Reasoned from the diff, CONFIRMED on hardware first (`audits/DIAG-bootrecon-drive-absent-2026-08-02.md`). The write hazard was blocked only by an ACCIDENTAL filesystem permission (host-root-owned mountpoint + unprivileged guest); the false dead-app alarm was real on every box. New fail-safe `bootrecon.StartGate` seam, also covering quiesce and in-flight app-data operations (§8.2). The rule already existed on the API path (`startGatedByMissingDrive`); the sweep bypassed it |
| R-166 | **App state gets a desired/observed model with its own store** (operator decision D-b) | M | **SHIPPED + PROVEN-LIVE — controller v0.189.0, 2026-08-02** | Tri-state `desired_state` in `app.yaml`, written ONLY by the customer's action; `bootrecon` reads intent instead of `len(Containers) > 0`; **absent means UNKNOWN** so legacy boxes keep byte-identical behaviour; running-only backfill; `backup.AppStopGuard` covers every stop→work→start window in its own marker file. **The durable fix for R-157 mechanism B.** Both blocking facts answered at source first: the crash-safe journal pattern DID already exist (quiesce, migrate) but covered **none** of the app-data path, and the SQLite store was rejected as a home because `metrics.db` is optional by design. `07`/`02` architecture docs carry the desired/in-flight/observed split. **Left for R-170:** the drive-backed boot gate still uses the container-count inference |
| R-158 | ~~A local Tier-1 backup failure reaches no hub channel~~ | S | **SHIPPED — controller v0.191.0 + hub v0.89.0, 2026-08-02** | Collapsed per the lifecycle rule. Closed by **R-167** (D-c's operator half) — no second row was filed for the same wire. New `unitNotify` seam fired per app from `captureAllRecoveryUnits` with the loop continuing, carrying the target filesystem's used/free bytes. **Routed to a new OPERATOR-ONLY type `recovery_unit_capture_failed`, NOT the `backup_failed` this row proposed** — that type is customer-enabled by default and would email the customer about a failure they cannot act on; decision D-c overrides the proposal. **Flips:** the new capability-map row *"A failed per-app Tier-1 backup reaches the OPERATOR"* → **PROVEN-LIVE** (`customer | … | skipped | operator_only` observed in the hub's `notification_log`, 9201, 2026-08-02) |
| R-167 | ~~Storage monitoring and backup alerts (decision D-c)~~ | M | **SHIPPED — controller v0.191.0/.1/.2 + hub v0.89.0, 2026-08-02** | Collapsed per the lifecycle rule. **Landed BEFORE the R-165 merge rather than with it** — D-a's condition (2) requires the same step and never after, and first is strictly better: the warnings were proven on hardware while the wall is still standing. Customer half = `internal/fillwatch`, per filesystem, two threshold terms (85% / 5 GiB; critical 95% / 2 GiB), edge-triggered with persisted state and a 75% / 7 GiB hysteresis return. **It emits the pre-existing `disk_warning`/`disk_critical` pair, which had NO producer in any repo — the sixth *built-but-never-wired* instance**; the hub's two generic `customerMessages` entries were deleted so the dynamic Hungarian survives. Operator half = R-158. **Flips:** two new capability-map rows → **PROVEN-LIVE**. **Follow-ups:** R-177 (no run-now path), R-175 (§7.5's bound is one box's) |
| R-165 | Merge `mp1` into `mp0` — the dedicated backup partition stops existing (decision D-a) | M | **SPIKED 2026-08-02 — WAITING-ON-OPERATOR** | `audits/SPIKE-r165-mp1-merge-2026-08-02.md`, M1-M5, **no layout touched**. Prerequisite R-167 is now SHIPPED. **Three findings the merge session must not re-derive:** (1) *"the layout" is not one thing* — demo-felhom `200G/50G`, demo-hp `50G/20G`, golden `16G/8G`, so any fixed pair is already wrong for one box (→ R-175); (2) **`mp1` is a BULKHEAD, not only a ceiling** — an overflow today cannot reach `/var/lib/docker`, and after the merge it can, which is the one place "the merge is cheap" stops being true; (3) the golden fails closed on the split in **four** places, not one. **M5: D-a's condition (1) is currently SATISFIED** — no external box is in the hub's host register, and both demo boxes are Tier 0 and reinstallable, so migration cost for the measurable population is zero. **Recommendation: S1 (one volume, two directories) + B2 (a refusal threshold in the capture path), fresh-install shape.** **Two prerequisites unmeasured → R-176.** Does NOT close R-163 until it lands |
| R-165 | ~~Merge `mp1` into `mp0` — the dedicated backup partition stops existing (decision D-a)~~ | M | **CLOSED — PROVEN-LIVE 2026-08-03.** Variant V-c shipped (golden v3.0.0 + agent v0.120.0 + controller v0.192.0); both demo boxes reinstalled and proven end to end (R-178). **Its stated replacement for the bulkhead — B2 — was completed by R-181 the same day** (controller v0.193.0/.1): the reserve became a per-app, per-run admission covering all three write legs and gained a size term, and both terms were re-proven by filling demo-hp. R-181 is collapsed into this row rather than carried separately | `audits/SPIKE-r165-mp1-merge-2026-08-02.md`, M1-M5, **no layout touched**. Prerequisite R-167 is now SHIPPED. **Three findings the merge session must not re-derive:** (1) *"the layout" is not one thing* — demo-felhom `200G/50G`, demo-hp `50G/20G`, golden `16G/8G`, so any fixed pair is already wrong for one box (→ R-175); (2) **`mp1` is a BULKHEAD, not only a ceiling** — an overflow today cannot reach `/var/lib/docker`, and after the merge it can, which is the one place "the merge is cheap" stops being true; (3) the golden fails closed on the split in **four** places, not one. **M5: D-a's condition (1) is currently SATISFIED** — no external box is in the hub's host register, and both demo boxes are Tier 0 and reinstallable, so migration cost for the measurable population is zero. **Recommendation: S1 (one volume, two directories) + B2 (a refusal threshold in the capture path), fresh-install shape.** **Two prerequisites unmeasured → R-176.** Does NOT close R-163 until it lands |
| R-159 | **wishlist's data landed in an ANONYMOUS volume — never backed up, orphaned by a redeploy** | XS | **SHIPPED** (`templates/wishlist/`, 2026-08-02) — filed for the CLASS | Image declares `VOLUME /usr/src/app/data`; template mounted `wishlist_data:/data`, a path the app never writes. `ResolveDockerVolumeNames` returns names only for compose-declared volumes, so `DumpAppVolumes` never sees an anonymous one. **The class is open:** any image `VOLUME` at an unmounted path is silent unbacked-up storage — **`immich-server` has one today** at `/data`, empty when measured. Proposed `REUSE.md` rule: a template mounts every path in `Config.Volumes`, or says why not |
| R-160 | **gramps-web persisted three paths and wrote to none of them** | XS | **SHIPPED** (`templates/gramps-web/`, 2026-08-02) | `/app/data` appears nowhere in the image's env; the accounts DB and **the family tree** (`GRAMPS_DATABASE_PATH=/root/.gramps/grampsdb`) both landed in the writable layer. Upstream persists **eight** paths, the template three, one a phantom. **Severity above papra's:** papra loses documents the customer may hold elsewhere; gramps-web loses the artefact built inside the app, of which no other copy exists by construction |
| R-161 | **The volume-persistence gate is enforced by convention, not automatically** | M | **RULED + SHIPPED at reduced scope** (operator, 2026-08-02; `app-catalog` `fd7747d`) | Enforcement is **convention**: the catalog repo has no CI at all (`.gitea/workflows`, `.github`, drone/woodpecker — searched, none). **R-29's record: three orphaned gates, one enforced, and only the enforced one ever stopped anything — so the ruling copied the shape that works.** **Controller-side REJECTED with a measured reason:** a check at template load can only read the file, and a static audit of all 53 templates reports the catalog clean **including papra** — it would pass on the very defect it exists to catch (the property is runtime-only; see `check-volume-persistence.py`'s header). **CI REJECTED for now** — neither repo has any, no users yet. **Shipped:** `scripts/catalog_gates.py`, one entry point over all three gates, non-zero on any failure, mandated in the catalog's `CLAUDE.md` as `site_gates.py` is. **Open residue is only the automatic half**, sufficient while one person touches templates |
+8 -4
View File
@@ -178,7 +178,7 @@ turns a true alarm into one the operator dismisses.
### A comment asserting an invariant needs a test pinning it, or it is a wish
**Six instances in this project have shipped guarantees the code did not provide** — each survived
**Seven instances in this project have shipped guarantees the code did not provide** — each survived
review because the comment read as settled:
| # | Comment | What it claimed | What the code did |
@@ -189,10 +189,14 @@ review because the comment read as settled:
| 4 | `classifyRunStates` I1 | *"StateStopped means deliberately stopped by the user"* | quiesce stops stacks the same way — a failed restart was silent (F-CRIT-1) |
| 5 | `inflight.go` | *"a caller that cannot acquire DEFERS"* | the backup caller recorded a failure and paged the operator (F-A1) |
| 6 | `quiesce.go` | the agent's 409 *prevents* "a spurious failure" | on the start path it produced one (F-A1) |
| 7 | `recovery_unit.go` B2 refusal (R-181) | *"the previous unit is untouched and NOTHING was deleted"* | *nothing deleted* held; **untouched was measured false** — the floor was checked ONLY in `captureAllRecoveryUnits`, while the two dump legs wrote the bulk into the same tree first and unguarded, so a 182,272 B tar became 2,147,666,432 B under a manifest that had not moved |
Two of these (4 and 5/6) were found by Campaign 8 **on live hardware**, not by review or unit tests
— #4 had a green, red-proofed test suite over a production path that was broken two independent
ways. So:
Three of these (4, 5/6 and 7) were found **on live hardware**, not by review or unit tests — #4 had a
green, red-proofed test suite over a production path that was broken two independent ways, and #7
survived a full green suite plus three of its own red-proofs, because every one of them asserted the
mechanism inside `captureAllRecoveryUnits` and none asserted the **consequence** across the whole
backup run. The test that would have caught it is the one #7's fix ships: fingerprint the tree before
and after, and compare. So:
- If a comment states an invariant, **name the test that pins it**, or write one.
- If an invariant has a stated dependency (*"if either invariant changes, revisit this"*), that is
+83 -3
View File
@@ -71,8 +71,12 @@ data:
# Host-install script. It lives at the repo's /scripts (outside the website doc-root),
# synced into .../current/scripts by git-sync (see the sparse-checkout ConfigMap). Served
# as text/plain so operators can inspect it in a browser before download-then-run.
# R-110: served from the INSTALLER TAG's tree, not the website's. The URL is unchanged
# (https://felhom.eu/scripts/felhom-host-install.sh) — it never carried a ref, so every
# producer of it (the bootstrap script, the hub's day-0 command) follows the tag with no
# edit. What changed is which tree this root points at.
location /scripts/ {
root /usr/share/nginx/html/current;
root /usr/share/nginx/scripts/current;
default_type text/plain;
}
@@ -213,8 +217,12 @@ metadata:
name: git-sync-sparse-checkout
namespace: felhom-system
data:
# R-110: TWO sparse-checkouts, because there are now two syncs with two different refs.
# The website tracks `main` (a copy edit must never need a release); /scripts/ tracks the
# installer TAG (pushing the installer must never publish it).
sparse-checkout: |
/website/
sparse-checkout-scripts: |
/scripts/
---
# ===================
@@ -246,6 +254,9 @@ spec:
- name: git-data
mountPath: /usr/share/nginx/html
readOnly: true
- name: git-data-scripts
mountPath: /usr/share/nginx/scripts
readOnly: true
- name: nginx-config
mountPath: /etc/nginx/conf.d/default.conf
subPath: default.conf
@@ -269,11 +280,14 @@ spec:
initialDelaySeconds: 3
periodSeconds: 10
# ── The WEBSITE sync — tracks `main`, unchanged cadence ──────────────────────────────
# Deliberately still a branch: the site is content, and a typo fix must reach felhom.eu in
# thirty seconds without cutting a release. Only /scripts/ moved to a tag (R-110).
- name: git-sync
image: registry.k8s.io/git-sync/git-sync:v4.4.0
args:
- --repo=https://gitea.dooplex.hu/admin/felhom.eu.git
- --branch=main
- --ref=main
- --root=/git
- --link=current
- --period=30s
@@ -294,13 +308,52 @@ spec:
securityContext:
runAsUser: 65534 # nobody
# ── The INSTALLER sync — tracks a TAG (R-110, operator ruling 2026-08-03) ─────────────
# felhom-host-install.sh runs as root on a virgin machine. Before this it was served
# straight from `main`, so pushing it WAS publishing it: within thirty seconds it was what
# every new machine downloaded and ran, with no staging and no rollback but another push.
#
# Publishing is now moving this tag; rolling back is moving it back. PROVEN, not assumed:
# git-sync v4.4.0 follows a tag AND notices a moved one — measured 2026-08-03 on a
# throwaway sync against this very repo (`update required … local:<old> remote:<new>` →
# `updated successfully`, one period, ~20 s).
#
# Bump this ref when the installer's published version changes. `hostinstall_gates.py`
# gate 6 fails if this sync stops naming an `installer-v…` tag.
- name: git-sync-scripts
image: registry.k8s.io/git-sync/git-sync:v4.4.0
args:
- --repo=https://gitea.dooplex.hu/admin/felhom.eu.git
- --ref=installer-v1.23.0
- --root=/git-scripts
- --link=current
- --period=30s
- --sparse-checkout-file=/etc/git-sync-scripts/sparse-checkout
volumeMounts:
- name: git-data-scripts
mountPath: /git-scripts
- name: sparse-checkout-scripts
mountPath: /etc/git-sync-scripts
resources:
requests:
memory: "32Mi"
cpu: "10m"
limits:
memory: "128Mi"
cpu: "100m"
securityContext:
runAsUser: 65534 # nobody
# Init container: wait for first sync before nginx starts
initContainers:
# BOTH trees are seeded before nginx accepts traffic. The second one is why /scripts/ has
# no 404 window across this change: a fresh pod does not become ready until the installer
# tag has been checked out, exactly as the website already worked.
- name: git-sync-init
image: registry.k8s.io/git-sync/git-sync:v4.4.0
args:
- --repo=https://gitea.dooplex.hu/admin/felhom.eu.git
- --branch=main
- --ref=main
- --root=/git
- --link=current
- --one-time
@@ -312,16 +365,43 @@ spec:
mountPath: /etc/git-sync
securityContext:
runAsUser: 65534
- name: git-sync-scripts-init
image: registry.k8s.io/git-sync/git-sync:v4.4.0
args:
- --repo=https://gitea.dooplex.hu/admin/felhom.eu.git
- --ref=installer-v1.23.0
- --root=/git-scripts
- --link=current
- --one-time
- --sparse-checkout-file=/etc/git-sync-scripts/sparse-checkout
volumeMounts:
- name: git-data-scripts
mountPath: /git-scripts
- name: sparse-checkout-scripts
mountPath: /etc/git-sync-scripts
securityContext:
runAsUser: 65534
volumes:
- name: git-data
emptyDir: {}
- name: git-data-scripts
emptyDir: {}
- name: nginx-config
configMap:
name: nginx-config
- name: sparse-checkout
configMap:
name: git-sync-sparse-checkout
items:
- key: sparse-checkout
path: sparse-checkout
- name: sparse-checkout-scripts
configMap:
name: git-sync-sparse-checkout
items:
- key: sparse-checkout-scripts
path: sparse-checkout
---
apiVersion: v1
kind: Service
+93
View File
@@ -1,3 +1,96 @@
## v1.23.0 — the installer is published, not pushed (2026-08-03, R-110 + R-183)
**Two channels moved off `main` in the same change, because either one left behind makes the other
cosmetic.**
**Channel 1 — the served script.** `manifests/webpage.yaml` git-synced `/scripts/` from
`--branch=main` on a 30 s period and nginx served that working tree, so **pushing this file WAS
publishing it**: within half a minute it was what every new machine downloaded and ran as root, with
no staging and no rollback but another push. The sync is now **split in two**: the website keeps
tracking `main` at the same cadence (a copy edit must never need a release), and `/scripts/` tracks
the tag **`installer-v<SCRIPT_VERSION>`**. Publishing is moving that tag; rolling back is moving it
back.
**PROVEN, not assumed:** git-sync v4.4.0 follows a tag *and* notices a **moved** one — measured on a
throwaway sync against this repo, `update required … local:<old> remote:<new>` → `updated
successfully`, within one period (~20 s). The moved-tag half is what the whole publish model rests
on, so it was measured before the manifest was touched.
**Channel 2 — the sixteen files the installer fetches while it runs.** `fetch_raw` pulled from
`$AGENT_REPO/raw/branch/main`. It now pulls from **`raw/tag/v$ART_AGENT_VER`** — the agent version the
hub has vouched and whose binary sha this script already verifies.
**That is a correctness fix, not only a publish-channel one (→ R-183).** These are the AGENT's
configs — its systemd unit, its sudoers, its guarded wrappers — and a fresh install was fetching the
**vouched binary** while taking its configs from **whatever `main` held**. Two refs, one install, and
nothing compared them. The right ref for them was never this script's `SCRIPT_VERSION`: they do not
live in this repo and have no relationship to its version line.
**No fallback to a branch.** A vouched version whose tag is missing fails loudly rather than quietly
serving `main` — a silent fallback is the appearance of control with none of it. `felhom-agent`
carries `v<version>` tags from now on, `release-agent.sh` creates them, and `agent_gates.py` fails if
the vouched version is not downloadable.
**Channel 3 — the URL — needed no change, and that is worth recording rather than leaving as a
silence.** `https://felhom.eu/scripts/felhom-host-install.sh` never carried a ref: the ref lives in
the manifest. So both producers of that URL (`scripts/iso/felhom-bootstrap.sh`, the hub's day-0
command) follow the tag with no edit — **and no hub change, so no hub version bump.**
**Gate 6 in `hostinstall_gates.py`** pins all three structurally, with no network so it stays in
`--fast` and runs in CI on every push: no `raw/branch/` ref anywhere in the installer; `fetch_raw`
still pins to `$ART_AGENT_VER`; the manifest still syncs `/scripts/` from an `installer-v…` tag and
the website still from `main`.
**It deliberately does NOT assert "a tag exists for the current SCRIPT_VERSION".** That gate would go
red on the very push that bumps the version, before publishing — and publishing being a separate
deliberate act is the entire ruling. A gate that fails on the normal path is one people learn to
ignore.
## docs — v1.22.0 exercised end to end on two real reinstalls (2026-08-03, R-178) — **no script change**
**Nothing shipped.** `felhom-host-install.sh` stayed at **v1.22.0**; the published copy at
`https://felhom.eu/scripts/felhom-host-install.sh` was confirmed byte-identical to the repo copy
(`sha256 ed02acb2da46c8d2b5c486ce99d5b9a2747e8786c6eb03652cf755ed1abdd9f4`) before use. Both demo
boxes were uninstalled and reinstalled with it, by **two deliberately different supply paths**:
demo-hp with `--golden <local volid>` (the `:2584` alternative), demo-felhom with
`--force-gitea-golden` (the canonical C.3 customer command). The merge-aware `step_grows` produced
`data +46G (->70G, ONE volume)` and `+226G (->250G)` respectively, and `fetch_verify` was observed
succeeding against the vouched manifest for **both** artifacts on demo-felhom
(`verified sha256 a7763d31b55b5ce7…` agent, `verified sha256 54e2a4c431daf580…` golden).
**Two script-side findings, filed not fixed** (the session was a runbook; §7 forbade code):
- **R-180** — `--archive-storage` is validated for existence (`:1583`) and for golden resolution
(`:1661`), but never against the ACL storage set it is about to grant (the fixed default
`local local-lvm felhom-pbs`). Staging the golden on `felhom-backup` therefore passed every
pre-flight gate and died at **step 8/8**: `HTTP 403: permission denied at /storage/felhom-backup
(missing privilege Datastore.AllocateSpace)` — *after* step 2 minted the token, step 4b **rotated
and vaulted root@pam**, and step 5 installed the agent. `ARCHIVE_STORAGE ∈ PVE_STORAGES` is a
one-line assertion over two variables both known at `:1583`.
- **R-179** — `--uninstall` leaves the NAS network-storage systemd units behind
(`mnt-felhom\x2ddrives-<share>.{mount,automount}`; automount left `failed`, parent bind left
mounted). The Part E residue-diff provenance is from **v1.9.1**, which predates the feature — and
demo-felhom, which never had a share configured, left nothing, which is exactly why a diff on such
a box reported clean.
Full evidence: root `REPORT.md`.
## host-install: one data volume, derived from the disk (2026-08-03, R-165)
**Forced by a census, not planned.** `felhom-agent` v0.120.0 merges the appliance's two data volumes
into one (decision D-a). `step_grows` computed **two** numbers and the install call passed both, so
this script had to change with the agent or every install would have provisioned a half-sized box.
- **`step_grows` computes ONE total.** The old 80/20 docker-vs-sysdata split is summed: `226` where it
was `184 + 42`, `106` where it was `84 + 22`, `46` where it was `34 + 12`. **A standard appliance
keeps exactly the capacity it had — 250 G — it is simply no longer split by a wall.**
- **The size still comes from the physical disk.** `step_grows` already read the thin pool's real free
space (`lvs /dev/pve/data`); the merge only collapsed its two outputs into one. This is what makes
the merge safe to ship: an unflagged install does **not** get the golden's 24 G base.
- **`--sysdata-grow` is DEPRECATED but still honoured.** It is no longer auto-computed (set to 0), and
a hand-passed value still counts because the agent **folds** it into the single volume's grow rather
than dropping it — so an operator reproducing an old command line gets the same total.
## CI — a Gitea Actions runner, and a red run that reaches a person (2026-08-02, R-168)
**No version bump anywhere: nothing in the product repos is compiled, built or deployed by this.**
+42 -12
View File
@@ -115,8 +115,8 @@
# (default: appliance → island 169.254.253.1:8443; byo → vmbr0 IP:8443)
# --no-island appliance only: keep the historical LAN bind instead of the R-50 island
# --rootfs-grow N grow OS rootfs by N GiB (default: auto-compute)
# --datavol-grow N grow Docker-data vol by N GiB (default: auto-compute)
# --sysdata-grow N grow user-data vol by N GiB (default: auto-compute)
# --datavol-grow N grow the single data volume by N GiB (default: auto-compute from the pool)
# --sysdata-grow N DEPRECATED (R-165): added to --datavol-grow; there is one volume now
#
# Guest cap (appliance: optional — protect a SHARED host's other guests; byo: BOTH REQUIRED —
# the only noisy-neighbor protection on a host you do not own; needs agent >= v0.52.0):
@@ -184,7 +184,7 @@
set -euo pipefail
SCRIPT_VERSION="1.22.0" # the SINGLE version source (F-1): -h and the run banners follow it.
SCRIPT_VERSION="1.23.0" # the SINGLE version source (F-1): -h and the run banners follow it.
# The hub used to carry a copy for its Setup tab; R-94 DELETED it
# (2026-08-02) because the hub cannot know which version a box runs —
# the Setup command fetches this script at run time. scripts/
@@ -492,12 +492,31 @@ fetch_verify() {
# exists, else anonymous). These are non-executable text (not the integrity-checked binary); the
# sudoers is `visudo -cf`-validated before install, which catches corruption/tampering that would
# matter. $1=repo-path $2=dest
#
# R-110 / R-183: PINNED TO THE AGENT VERSION BEING INSTALLED, never to a branch.
#
# These sixteen files are the AGENT's configs — its systemd unit, its sudoers, its guarded wrappers —
# so the ref that is correct for them is the agent version this run is installing, which the hub has
# vouched and whose binary sha this script verifies. It is NOT the installer's own SCRIPT_VERSION:
# these files do not live in the installer's repo and have no relationship to its version line.
#
# Before this they came from `raw/branch/main`, which is a REAL SKEW and not only a publish-channel
# defect (R-183): a fresh install fetched the vouched agent BINARY while taking its unit file and
# sudoers from whatever `main` happened to hold — two refs, one install, and nothing compared them.
#
# NO FALLBACK TO A BRANCH. A vouched version whose tag is missing must fail loudly here rather than
# quietly serving `main`, because a silent fallback is exactly the "appearance of control with none of
# it" this change exists to remove. `agent_gates.py`'s published-version gate keeps the tag and the
# vouched version in step, so this die is a backstop and not the primary control.
fetch_raw() {
local path="$1" dest="$2"
# Late steps (mgmt-watchdog, OOB) can run without step 5 having resolved the manifest.
[[ -n "$ART_AGENT_VER" ]] || resolve_artifacts
[[ -n "$ART_AGENT_VER" ]] || die "cannot pin $path: no agent version resolved from the hub manifest"
local -a _auth; _git_auth_args _auth
curl -fsS "${_auth[@]}" -o "$dest" \
"$GITEA_BASE/$GITEA_OWNER/$AGENT_REPO/raw/branch/main/$path" \
|| die "raw fetch failed: $path"
"$GITEA_BASE/$GITEA_OWNER/$AGENT_REPO/raw/tag/v$ART_AGENT_VER/$path" \
|| die "raw fetch failed: $path (agent tag v$ART_AGENT_VER — is that version tagged in $AGENT_REPO?)"
[[ -s "$dest" ]] || die "raw fetch empty: $path"
}
@@ -1772,23 +1791,34 @@ step_token() {
#-------------------------------------------------------------------------------
step_grows() {
log_step "3/8 compute volume grows"
# Golden base: rootfs 32G + Docker-data 16G + user-data 8G (build-golden.sh).
# Golden base since build-golden.sh v3.0.0 (R-165): rootfs 32G + ONE data volume 24G. The separate
# 8G user-data volume was MERGED AWAY — one volume, one free-space figure, no ceiling — so there is
# one number to compute here instead of two.
#
# THE SIZE IS DERIVED FROM THE PHYSICAL DISK, which is what makes the merge safe to ship: an
# unflagged install does NOT get the golden's 24G, it gets a share of the thin pool's real free
# space. (Before R-165 this same block already did the deriving; the merge only collapsed its
# 80/20 docker-vs-sysdata split into a single total.)
if [[ -z "$ROOTFS_GROW$DATAVOL_GROW$SYSDATA_GROW" ]]; then
local free_gib
free_gib=$(lvs --noheadings --units g -o lv_size,data_percent /dev/pve/data 2>/dev/null | awk '{gsub(/[^0-9.]/,"",$1); used=$2; print int($1*(100-used)/100)}' 2>/dev/null || echo 0)
# Reserve headroom; split the rest ~ docker 80% / sysdata 20%; rootfs stays golden.
# Reserve headroom; the totals below are the pre-merge pair SUMMED, so an appliance gets the
# same capacity it did before — it is simply no longer split by a wall.
ROOTFS_GROW=0
if [[ "${free_gib:-0}" -ge 300 ]]; then
DATAVOL_GROW=184; SYSDATA_GROW=42 # reproduces the standard 200G/50G appliance
DATAVOL_GROW=226 # 184+42 -> the standard 250G appliance (was 200G+50G)
elif [[ "${free_gib:-0}" -ge 150 ]]; then
DATAVOL_GROW=84; SYSDATA_GROW=22
DATAVOL_GROW=106 # 84+22
else
DATAVOL_GROW=34; SYSDATA_GROW=12 # minimal floors
DATAVOL_GROW=46 # 34+12 — minimal floor
fi
log_info " auto-computed from ~${free_gib} GiB free"
SYSDATA_GROW=0
log_info " auto-computed from ~${free_gib} GiB free (ONE volume since R-165)"
fi
ROOTFS_GROW="${ROOTFS_GROW:-0}"; DATAVOL_GROW="${DATAVOL_GROW:-0}"; SYSDATA_GROW="${SYSDATA_GROW:-0}"
log_info " grows: rootfs +${ROOTFS_GROW}G (->$((32+ROOTFS_GROW))G), docker +${DATAVOL_GROW}G (->$((16+DATAVOL_GROW))G), sys_drive +${SYSDATA_GROW}G (->$((8+SYSDATA_GROW))G)"
# A hand-passed --sysdata-grow is still ACCEPTED and still counts: the agent folds it into the one
# volume (bringup.go 4b), so an operator reproducing an old command line gets the same total.
log_info " grows: rootfs +${ROOTFS_GROW}G (->$((32+ROOTFS_GROW))G), data +$((DATAVOL_GROW+SYSDATA_GROW))G (->$((24+DATAVOL_GROW+SYSDATA_GROW))G, ONE volume)"
_state_mark grows
}
+56
View File
@@ -146,6 +146,62 @@ if re.search(r'^PVE_STORAGES=\([^)]*felhom-pbs[^)]*\)', src, re.M):
else:
fail("felhom-pbs missing from the default PVE_STORAGES — narrowing it 403s the PBS-DR apply-bridge")
# ── 6. the publish channel is pinned, not floating (R-110 / R-183) ──────────────
#
# WHAT THIS ASSERTS, AND WHAT IT DELIBERATELY DOES NOT.
#
# It does NOT assert "a tag exists for the current SCRIPT_VERSION". That gate would fail the very
# push that bumps SCRIPT_VERSION, before publishing has happened — and publishing being a SEPARATE
# deliberate act is the whole point of R-110's ruling. A gate that goes red on the normal path is a
# gate people learn to ignore, which is the reasoning the task's own §8.4 applies to the agent-side
# gate; it applies here identically. "Is the vouched version actually downloadable" is a real
# invariant and it lives where a missing artifact genuinely breaks day-0 — `felhom-agent`'s
# `agent_gates.py`, which has the network access to answer it.
#
# What it asserts instead are the two STRUCTURAL regressions that would silently return the
# installer to a floating channel, both answerable by reading files (no network, so this stays in
# `--fast` and therefore runs in CI on every push):
#
# 6a. no `raw/branch/` ref anywhere in the installer — one of the sixteen agent-config fetches
# slipping back to `main` is exactly how a channel stays floating unnoticed, and it is
# invisible in a diff that touches one line.
# 6b. `fetch_raw` still pins to the resolved agent version — the positive form, so the mechanism
# cannot be quietly deleted rather than regressed.
# 6c. the website manifest still syncs `/scripts/` from a TAG ref and the website from `main` —
# the split is the deploy-side half of the same channel, and reverting it is one word.
branch_refs = [l for l in lines if "raw/branch/" in l and not l.lstrip().startswith("#")]
if branch_refs:
fail("installer still fetches from a BRANCH ref — the run-time channel is floating again "
"(R-110/R-183). Offending line(s): %s" % "; ".join(l.strip()[:90] for l in branch_refs))
else:
ok("no raw/branch/ ref in the installer — every run-time fetch is pinned")
if re.search(r'raw/tag/v\$ART_AGENT_VER/', src):
ok("fetch_raw pins the agent configs to the vouched agent version")
else:
fail("fetch_raw no longer pins to $ART_AGENT_VER — the agent's configs and its binary can "
"again come from different refs in one install (R-183)")
WEBPAGE = os.path.join(ROOT, "manifests", "webpage.yaml")
try:
with io.open(WEBPAGE, "r", encoding="utf-8") as f:
wp = f.read()
except IOError as e:
fail("cannot read manifests/webpage.yaml to check the publish channel: %s" % e)
wp = None
if wp is not None:
# The scripts sync must name a tag ref; the website sync must still track main.
if re.search(r'--ref=installer-v', wp):
ok("manifest: /scripts/ syncs from an installer tag")
else:
fail("manifests/webpage.yaml has no `--ref=installer-v…` sync — /scripts/ is not served "
"from a tag, so pushing the installer publishes it again (R-110)")
if re.search(r'--(branch|ref)=main', wp):
ok("manifest: the website still tracks main (a copy edit must not need a release)")
else:
fail("manifests/webpage.yaml no longer tracks main for the website — pinning the SITE to "
"the installer tag turns every copy edit into a release")
print()
if fails:
print("hostinstall gates: %d FAILURE(S)" % len(fails))
+67 -1
View File
@@ -1,6 +1,6 @@
---
name: felhom-build-deploy
description: Build, deploy, publish, or verify ANY Felhom artifact — felhom-controller image (guest 9201 bootstrap deploy), felhom-agent binary (felhom-pve), felhom-hub (GitOps/ArgoCD), the felhom.eu website (git-sync), or the app catalog. Use whenever the task says build, deploy, ship, release, publish, bump version, restart the controller/agent/hub, or verify what version is live. Contains the exact verified commands and the gotchas that silently break deploys.
description: Build, deploy, publish, or verify ANY Felhom artifact — felhom-controller image (guest 9201 bootstrap deploy), felhom-agent binary (felhom-pve), felhom-hub (GitOps/ArgoCD), the felhom.eu website (git-sync), the app catalog, or the PUBLIC installer ISO (iso.felhom.eu). Use whenever the task says build, deploy, ship, release, publish, bump version, restart the controller/agent/hub, or verify what version is live. Contains the exact verified commands and the gotchas that silently break deploys.
---
# Felhom build & deploy runbooks
@@ -75,6 +75,72 @@ Publish to Gitea (so Day-0 self-install can fetch it): `scripts/publish-agent.sh
`REGISTRY_*` creds. The hub's Day-0 artifact manifest must then vouch the new version — that UI is
operator-password-gated (CC cannot); flag it as an operator follow-up.
## Installer ISO (felhom.eu/scripts/iso → iso.felhom.eu) — PUBLIC, irreversible
**Two modes, and picking the wrong one ships the wrong product.**
| Mode | What it is | Menu |
|---|---|---|
| `--release` | the **public** image. No `answer.toml`, no root password, no SSH key, no disk profile. Day-0 rides a `.deb`. | TWO interactive entries, graphical default, timeout 15s |
| `--pairing` / `--bootstrap-env` | operator-built **appliance** image: baked answer file, baked root hash, a disk profile pinned to one machine | ONE automated entry |
```bash
# build the public image (clean-tree gate first — an unpushed change does not exist)
export FELHOM_ISO_OUT=$FELHOM_ROOT/felhom-iso/out
bash scripts/iso/build-felhom-iso.sh \
--pve-iso $FELHOM_ROOT/drill/proxmox-ve_9.2-1.iso \
--iso-sha256 4e88fe416df9b527624a175f24c9aa07c714d3332afb1ee3dbf3879573ef2c6c --release
# -> $FELHOM_ISO_OUT/felhom-installer-<VER>-pve<PVE>.iso + .sha256 + .manifest.txt (NO .rootpw.txt)
```
**Before publishing, two hard gates — neither is optional and neither is a script yet:**
1. **`documentation/runbooks/iso-release-gate.md`** — 13 criteria, run against **the exact file you
will upload**, not the build inputs. It is a manual checklist; `repo_gates.py` does **not** cover
it, so nothing will remind you (R-29's shape — say so if you skip it).
2. **A proof install from the built image on BOTH menu entries** (graphical *and* Terminal UI), each
showing: package installed, unit `enabled`, unit fired on first boot, and a pairing code in
`/etc/felhom/appliance-pairing-code`. Spike 4 *reasoned* the graphical path follows from shared
`Install.pm`; the 1.26.0 run proved that reasoning insufficient in a different place — do both.
```bash
# publish — rclone in a container, env-only config, so NO credential file is ever written
source ~/.config/credentials # ISO_S3_CLIENT_AK / _SK / ISO_S3_URL — never echo, never log
docker run --rm -v $FELHOM_ISO_OUT:/data:ro \
-e RCLONE_CONFIG_R2_TYPE=s3 -e RCLONE_CONFIG_R2_PROVIDER=Cloudflare \
-e RCLONE_CONFIG_R2_ACCESS_KEY_ID="$ISO_S3_CLIENT_AK" \
-e RCLONE_CONFIG_R2_SECRET_ACCESS_KEY="$ISO_S3_CLIENT_SK" \
-e RCLONE_CONFIG_R2_ENDPOINT="$ISO_S3_URL" \
-e RCLONE_CONFIG_R2_REGION=auto -e RCLONE_CONFIG_R2_NO_CHECK_BUCKET=true \
rclone/rclone:latest copy /data R2:felhom-iso --include "felhom-installer-<VER>*" --s3-chunk-size 64M
# verify by ROUND TRIP — the downloaded bytes, not the local file
curl -fsSL -o /tmp/rt.iso https://iso.felhom.eu/felhom-installer-<VER>-pve<PVE>.iso && sha256sum /tmp/rt.iso
```
`ListBuckets` 403s — the token is object-scoped; list with `lsf R2:felhom-iso`, not `lsd R2:`.
### Proof-VM traps — every one of these cost a wrong diagnosis
- **Set `--boot` in a SEPARATE `qm set`, after the disk exists.** `qm set <id> --scsi0 … --boot order="scsi0;ide2"`
in one call silently yields `boot: order=net0;ide2`; the VM netboots, fails, falls through to the CD.
- **After the install, detach the CD** (`qm set <id> --delete ide2; qm set <id> --boot order="scsi0"`)
or the machine re-enters the installer on reboot — **a completed install looks exactly like a stuck one.**
Judge completion from `qm config` + disk usage, never from the screen.
- **Verify focus by screendump before every `Enter`.** TUI: red-highlighted button, tab order. GTK:
dashed focus ring, and `Enter` lands in text *fields*, not `Next`. Not checking once aborted an install.
- **Proof installs register unclaimed appliances at the hub — discard them** or R-131 grows:
`curl -u ":$HUB_PW" -X POST http://<hub-clusterIP>:8080/appliances/<id>/discard` → 303.
The verb is **`/discard`**, POST only (`hub/internal/web/server.go:345`); `/delete` 404s.
- Venue: `demo-hp`, scratch `dir` storage at **`/mnt/nvme-1tb` root** (a subdirectory reads
`disconnected` forever — the agent's `exactMount` check). Never `local-lvm`. Remove the storage at teardown.
**Why the shape is what it is** (do not re-derive; four spikes measured it):
`documentation/audits/SPIKE-universal-iso-{1,2,3,4}-2026-07-31.md`. In short — no udev property
distinguishes an internal disk from a customer's backup drive and a two-disk filter match silently
wipes one, so **there is no safe automated disk selection for unseen hardware**; and `[first-boot]`
is never placed on the system by an interactive install, so day-0 rides a `.deb` in
`/proxmox/packages/` instead (`Install.pm:1343-1372`).
## Hub (felhom.eu/hub → k3s, GitOps via ArgoCD app `felhom`)
**The manifest is the truth.** A code push + image build deploys NOTHING until `manifests/hub.yaml`'s