9c1d05d3600a078a578410c6eb3d298fbf025c4c
1001 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4cc123809c |
revert the Scenario B breakage — main is green again
gates / gates (push) Successful in 7s
The deliberate hostInstallVersion const is removed. It existed only to produce a real red run (#3-#6) and the demonstrated alarm; R-94's deletion stands. |
||
|
|
9530de722d |
ci: the alarm needs a User-Agent — Cloudflare 403s python-urllib
gates / gates (push) Failing after 8s
Resend sits behind Cloudflare, which blocks the default 'Python-urllib/3.x' agent with its own 403 (error 1010). That failure looks exactly like an auth failure and is not one, so the reason is recorded next to the header. Verified from the runner image with a deliberately invalid payload: with the agent set, Resend answers 422 missing-field, i.e. the request now reaches the API. |
||
|
|
f7dbc335ac |
ci: send the failure alarm with python3/urllib, not curl
gates / gates (push) Failing after 8s
The first version died on 'curl: command not found' — the runner image carries python3 and git and nothing else on purpose. Reaching for a bigger image to send one HTTP request would have been the wrong trade, so the step uses urllib. Verified from the image itself that HTTPS to api.resend.com resolves and the certificate verifies. The step also fails LOUDLY on an empty key or a non-2xx from Resend: a silent alarm is worse than no alarm, because it reads as coverage. |
||
|
|
dd13f632c8 |
ci: a failed run sends its own alarm (R-168, probe P5)
gates / gates (push) Failing after 7s
P5 measured: a failed run produces NO mail, NO notification row and NO log line from Gitea. A red tick in a web UI nobody watches is exactly the defect R-29 filed, rebuilt one layer up, so the run alarms itself on the project's existing transactional path (Resend, the same one the hub uses) and prints the provider's accepted id, making 'it was sent' an observable rather than an assumption. The key is a user-level Gitea Actions secret created out-of-band; it is in no committed file. The recipient is the operator address the hub already uses and is not a secret. This push is deliberately made while main is still carrying the Scenario B breakage, so the resulting run fails and demonstrates the alarm end to end. |
||
|
|
3252d51104 |
SCENARIO B: deliberately break the hostinstall gate (reverted immediately)
gates / gates (push) Failing after 7s
Pushed with --no-verify ON PURPOSE: this simulates exactly the bypass that CI exists to catch. The local pre-push hook would have refused this commit. |
||
|
|
666a34da88 |
ci: run the gate entry point on every push (R-168)
gates / gates (push) Successful in 7s
Replaces the Part 0 probe workflow, whose four measurements are recorded in documentation/audits/SPIKE-ci-runner-2026-08-02.md. Reports, does not refuse: pushes go straight to main with no pull request, so there is no merge for a status check to stand at. The refusing half is .githooks/pre-push, which is per-clone and --no-verify-able; this half notices when that was skipped. No uses: step anywhere — JavaScript actions need a node runtime the host-mode runner does not have. Probe P3 measured that a shallow git fetch of the exact pushed SHA from the in-cluster Gitea service is sufficient, and that it equals the pushed commit. The alarm step is deliberately absent until probe P5 measures whether Gitea already mails on a failed run. |
||
|
|
bbd6231909 |
probe: temporary CI workflow for R-168 Part 0 (P1/P2/P3 + Scenario E)
probe / probe (push) Successful in 48s
TEMPORARY. Deleted before the session ends. Measures whether a registered runner picks up a job at all, whether python3 and git are visible to the JOB (not merely present in the image), whether the source can be obtained with no JavaScript action, and that docker is NOT reachable from a job. |
||
|
|
af2d103880 |
REPORT + STATUS: gate enforcement session, hub v0.87.0 live
REPORT overwritten per the standing rule; every red-proof, the core.hooksPath probe's four measured outcomes, Scenario C's refusal-and-bypass, the hub deployment and the live Setup-tab read are recorded there, plus three observations and two deliberate deviations from the spec (a comment-only edit to felhom-host-install.sh, and __pycache__ in .gitignore). STATUS: the 'check that needs a person to remember it' line is rewritten — the checks now run themselves before every push, with both honest limits stated in plain words; and one entry records the thirteen-check cleanup and the deleted installer version number. |
||
|
|
8d9b78c153 | manifests: hub 0.86.0 -> 0.87.0 (R-94, the Setup tab renders no version) | ||
|
|
4707be755c |
docs: R-94 closed, R-29 leg (a) closed + leg (b) half, R-168 minted
hub/CHANGELOG v0.87.0 + scripts/CHANGELOG gate-enforcement entry. CONTEXT gains S-6 (the hub renders no host-install version and the gate pins its absence) and S-7 (gates run from one entry point per repo; reuse_refs_check was fixed rather than the REUSE.md convention, with both rejected alternatives recorded). OPEN-ITEMS: R-94 CLOSED all three legs, leg (a) by DELETION with its reason; R-29 leg (a) CLOSED and leg (b) HALF-SHIPPED with the census result written into the row (13 gates; every gate a CLAUDE.md names was green, two of the four unnamed were red); R-161 gains its successor pointer. NEW R-168 (grep established R-167 was the highest in use): Gitea Actions runner — measured 2026-08-02 as Gitea 1.26.2, Actions enabled on all four repos, 0 runners, 0 workflow runs, 0 branch protections, and the consequence that trunk-based direct-to-main pushes leave no merge for a status check to gate, so CI here can detect but not block. BLOCKED on a spike over host-mode vs privileged DinD on DooPlex and whether the workflow can avoid JavaScript actions. ROADMAP: R-94 collapsed to its one-liner, R-29 updated, R-168 added. |
||
|
|
9bd1a54d71 |
gates: one entry point (scripts/repo_gates.py) + pre-push hook
A census of all thirteen gate scripts across the four felhom repos on 2026-08-02 found one clean correlation: every check a CLAUDE.md tells a person to run was passing, and two of the four nobody is told to run were failing — one since 14 July. Neither failure was harmful in effect (checked line by line); nothing would have said so if they had been. The fix is not more gates, it is one place to run them from. repo_gates.py runs site + hostinstall + hub-confirm + manifest-bearer + reuse-refs, streams each gate's own output, and exits worst-wins non-zero. A missing gate script is a FAILURE and prints the path tried — fail-closed, because a runner that quietly skips a gate is the inert-seam failure this project has shipped four times. It copies catalog_gates.py (R-161), NOT site_gates.py, which is a gate and not a runner. .githooks/pre-push runs it with --fast and refuses the push. Honest limits are written into the hook itself: per-clone (core.hooksPath is local config), and --no-verify bypasses it on purpose. Any manual run WARNS when the clone is unarmed. Measured on git 2.47.3: a relative core.hooksPath resolves correctly and the hook's cwd is the repo root from any subdirectory. test_repo_gates.py is a SEAM test — it asserts each member gate's own distinctive stdout, not the runner's summary line, which an inert runner prints while calling nothing. Red-proofed: replacing run_gate's body with 'return 0' still prints 'all felhom.eu gates OK' and exits 0, and turns the seam test red. |
||
|
|
2137094799 |
scripts: reuse_refs_check resolves package shorthand and sibling repos
RED on all four repos with 13 findings, and a hand audit of all 13 on 2026-08-02 found ZERO genuine drift: twelve were package shorthand whose file sits a couple of directories deeper, and one (wgsync/reconciler.go, cited by the controller) lives in the hub. REUSE.md cites by package shorthand and across repos on purpose; the tool was what was wrong. Resolution order, first hit wins, every non-exact hit PRINTED so a weakening is visible: exact / suffix / ambiguous (real citation, imprecise shorthand — not a failure) / sibling repo (as-is or with the sibling's own name stripped from the token) / FAIL. A failure lists every resolution attempted, so a 'not found' claim names what was tried. Per-root tallies are the positive observable: '0 failures' alone cannot tell a working checker from a blind one. Evidence trees (audits/, documentation/tests/) are excluded from the suffix index — a copy of a file is not the file. An absent sibling is never a failure; an unreadable parent says so and continues. Result: 13/13 resolve, all four roots exit 0. felhom.eu 60 exact + 1 suffix; controller 126 exact + 6 suffix + 1 cross-repo; agent 88 + 1 + 1; catalog 17 exact + 3 cross-repo. New scripts/test_reuse_refs_check.py: 13 fixture tests, one per resolution row plus the kill condition. Red-proof: making resolve() return 'exact' for an unresolvable token turns 4 of them red. |
||
|
|
d319ae573e |
hub: delete the host-install version label (R-94) + invert hostinstall gate 1
The Setup tab said 'host-install 1.19.0' while the served script was 1.22.0, and had been wrong since 2026-07-14. Deriving the number honestly is not possible: the Option-1 command downloads felhom-host-install.sh from the website at RUN TIME and the website git-syncs main every 30s (R-110), so no build-time value in the hub can be true. R-94(a) offered derive-or-delete; deleted, which removes the drift class instead of automating it. - configs.go: hostInstallVersion const, pageData.ScriptVersion field and its assignment all removed; a NOTE in their place records why there is no constant here. - customer_unified.html: the sentence now says the command always fetches the current installer, and renders no version. - hostinstall_gates.py gate 1: the third assertion INVERTS — it used to require the hub const to equal SCRIPT_VERSION, it now asserts the hub carries no host-install version literal at all, matched in six code shapes across every .go/.html under hub/ (comments are deliberately not stripped: a // inside a URL literal would blind the scan). - render_test.go: the assertion 'html contains hostInstallVersion' compared the constant to itself and passed at ANY value — demonstrated green with the const at 9.9.9 while the script was 1.22.0. Deleted, not replaced: there is no longer a version to assert. - felhom-host-install.sh: COMMENT ONLY (SCRIPT_VERSION untouched) — it claimed the gate keeps the hub copy equal, an invariant that no longer exists. Red-proofs: restoring the const fails the rewritten gate 1 (3 shapes hit); the old render_test assertion passes at 9.9.9. |
||
|
|
e994bf35d2 |
STATUS.md: a plain-language operator page, and today's four decisions recorded
Documentation only — no code, no box, no build.
STATUS.md (repo root, 652 words / 67 lines): what works · what's broken ·
what we're working on · waiting on you · changed since. A VIEW of
OPEN-ITEMS.md, holding nothing of its own; not CONTEXT.md, and both files
now say why they stay separate. No R-n is the subject of a sentence —
identifiers are bracketed pointers only.
CONTEXT.md S-5 records the four operator decisions taken 2026-08-02
(D-a … D-d), none of them implemented:
D-a merge mp1 into mp0 rather than resize it — before any external
install, and D-c ships in the same step → R-165
D-b desired/observed app state in its own store, with the state-store
safety rule verbatim → R-166 (BLOCKED)
D-c customer fill warning + operator backup-failure alert → R-167
D-d only DooPlex and Peti's box are protected → target-selection.md
R-163 RE-FRAMED, not closed: the sizing question is withdrawn rather than
answered; the row survives as the record of the constraint until R-165
lands. R-156's papra referral RESOLVED — deployed nowhere, so the template
fix strands nothing; the docker ps evidence is recorded with its
provenance and its scope limit.
target-selection.md: two protected machines, everything else disposable.
ep0 is no longer Tier 2 but is not scratch (it holds the only off-premises
copy of real customer data) — flagged for explicit operator confirmation.
The demo-box backup-target fence drops from prohibition to stated cost,
because D-d spends that reference anyway.
CLAUDE.md gains an End-of-session checklist carrying the STATUS.md
maintenance rule and "a finding goes in OPEN-ITEMS.md first".
|
||
|
|
260a8f6e58 |
register: R-161 ruled and shipped at reduced scope; re-ranked
The operator ruled on R-161 and the runner shipped in app-catalog-felhom.eu (fd7747d), so the row moves from BLOCKED-needs-a-ruling to REDUCED SCOPE - open. Both obvious enforcement points were rejected for measured reasons, and the row now records them rather than leaving the rejection implicit. Controller-side at template load: rejected because such a check can only read the file, and a static audit of all 53 templates reports the catalog clean INCLUDING papra - it would pass on the exact defect it exists to catch, the property being decidable only at runtime. CI: rejected for now, neither repo has any and there are no users yet. Shipped instead: scripts/catalog_gates.py, one entry point over all three gates, non-zero exit on any failure, mandated in the catalog's CLAUDE.md the way site_gates.py is. The rationale is recorded because it is the transferable part - of this project's gates, the only ones that ever get run are those with a single entry point named in a CLAUDE.md; site_gates.py is run and R-29's three orphans are named nowhere and have stopped nothing. What stays open is only the automatic half, which is sufficient while ONE person touches templates - revisit when a second does. Re-ranked accordingly: R-161 drops from 2nd to 7th, and R-156 is promoted to 2nd, since R-161 was ranked high precisely because nothing ran the gate and that is no longer true. The de-ranking is recorded inline with its reason, matching how R-94's de-ranking is recorded, so a later reader sees a decision rather than drift. |
||
|
|
b06ea9c877 |
register: file R-156..R-164 in one pass, ranked; and record what mp1 is actually for
Nine rows into OPEN-ITEMS.md and ROADMAP.md, matching each file's column shape. R-156 and R-157 had lived only in audit documents - the identical "minted in a spike doc and never carried across" failure the register already records for R-153/R-154/R-155, caught by the catalog sweep's own section 8.0 while it was happening. R-158 was minted by a second session the same day for an unrelated finding, which is why the sweep's proposals were renumbered R-159..R-162 at filing time. All nine IDs verified free in BOTH backlog files before use. Part 0 settled the question the sizing item depended on, by reading: mp1 is RETENTION, not staging, and neither of the two framings was right. A unit is the KEPT copy on the app's OWN drive (backup.go:245-255); for an app with no HDD_PATH the namespace falls back to the system SSD - "the SSD-only system-data fallback" (appbackup/paths.go:26-27). There is no post-copy deletion: the only prune is F5 residue-on-old-drives when an app MOVES (backup.go:1053-1112). So mp1 retains the units of driveless apps only - not every app, but not transient either. Confirmed against the spike: sys_drive held exactly the four driveless apps and not calibre-web, which had a drive and was still backed up. A unit is volume tars + DB dumps only, never mp8 userdata (recovery_unit.go:20-25), so a 1 TB photo library can never overflow one. And mp1 gates the WHOLE chain, not just Tier 1: Tier-2 mirrors the unit "(always)" from RecoveryUnitPath (tier2.go:302,368) and Tier-3 carries it, so a unit that cannot be written leaves both with nothing to copy. Part 2 fired on both triggers - retention, and the fallback undocumented - so 07-backup-architecture.md gains section 7.5. Section 6.1 said a unit lives "on the app's own drive", which is true and was the whole story only for drive-resident apps; the no-drive case was undocumented, as was the sizing constraint. 7.5 records the mp0-50G-vs-mp1-20G mismatch, the measured ratios (DB app up to ~2x, 21.1GB -> 40.2GB; file-only 1.00x), and the bound this puts on D5's Lane-1 independence: restorable from the drive alone only while the unit still fits - about 19 GB file-only, about 10 GB DB-backed. No number proposed; the ratio is the operator's ruling (R-163). R-159/R-160 marked SHIPPED only after verifying the template changes are in app-catalog origin/main, and R-156's gate likewise (check-volume-persistence.py present). papra is NOT fixed - referred - so R-156 stays open on that one app. Ranked, with one line of reasoning each: R-157 first (an app can stay down indefinitely with mechanism B silent on every channel), then R-161 (the gate exists and nothing runs it, which is why R-156's class recurs - R-29's record is three orphaned gates and one enforced), R-156, R-163, R-158, R-164, R-162. |
||
|
|
482af37b7d |
Campaign 10 closeout Part 2: teardown complete — five layers, each verified gone
Evidence-survival check FIRST: HEAD == origin/main ==
|
||
|
|
7efb7a53d3 |
Campaign 10 closeout Part 1: Q1 lowers R-158's rank; Q2 clears ValidateDump and kills C2's gate
Q1 - what the customer sees when a backup refuses for lack of space. The failure
IS customer-visible: /backups renders "Adatmentés sikertelen" with a cross mark.
It is absent from the dashboard, the launcher, the app detail page, and - the one
worth fixing - from /backups/apps, the per-app page where you would naturally ask
whether a given app is backed up.
Point 5 measured across three runs: it retries, stays failed while constrained
(marker persists, unit mtime unchanged at 07:30:50), and clears on recovery with a
fresh unit at 07:37:38. /backups/apps reading "Utolsó: 3 perce" tracks the unit's
REAL mtime, not the failed run, so it is honest about the age of the last good
unit rather than claiming a fresh one. Explicitly NOT the R-156 family.
So R-158 is a NOTIFICATION GAP, not a silent-failure defect, and ranks BELOW
R-157 - whose mechanism B leaves a deployed app not running while deadapp reports
"0 currently down", silent on every channel.
Q2 - ValidateDump was right and no bad dumps are shipping. The live DB genuinely
had zero accounts (only _prisma_migrations 129, cc_proof 82, instance_settings 1).
An empty table proves nothing, so an account was SEEDED as the task required: the
warning then stopped entirely and the dump provably contained the rows (c10acct 1,
c10user 2; 102766 -> 103029 bytes).
But that kills C2's proposed ordering. A fresh appliance legitimately has zero
accounts, so gating on "accounts has rows" would block the backups of every new
customer until someone registers. The validator's fact is right; its inference
("may predate the customer's data") is wrong - there was no data to predate. The
chain is therefore longer: a sound predicate first (compare the dump against the
LIVE db, per-table counts, not an absolute expectation), then warn->gate, then the
tar-drop. Until then the DB volume tar stays load-bearing - not because dumps are
bad, but because nothing can yet prove one is good.
No new R-n; register grepped. Nothing fixed. Part 2 (teardown) follows.
|
||
|
|
0afadbdeff |
SPIKE: recovery-unit space — the ceiling is real on mp1, overflow is clean, but silent (R-158)
Three headline answers.
1. The ceiling is REAL and on mp1 (/mnt/sys_drive), but its shape is a MISMATCH
rather than a single number. A1: docker's data-root is a SEPARATE 50G volume
(mp0) and every app volume resolves there, so app DBs are NOT on sys_drive -
build-golden.sh:68's "like the Docker-data" reading is correct. A2: the
recovery units ARE on sys_drive, which the golden ships at 20G. So a box
permits 50 GB of live app data while capping local backup at 20 GB, and
crossing that line is invisible until a backup fails.
A3 rules out the lab-default explanation: --sysdata-grow defaults to 0
(main.go:178) and is not computed from the drive. demo-hp's REAL guest 9201
runs a bare --config ExecStart and shows mp0 50G / mp1 20G; agent.json has no
sizing keys at all. A4, measured not read: restore extracts IN PLACE on the
docker volume - sys_drive avail was 799.2M before and after a restore run under
constraint - so the constrained mount is written only during backup.
2. Overflow behaves WELL. With sys_drive ballasted to 799 MB, backup refused
per-app ("No space left on device"), other apps continued, status reported
success=false, and the "last good dump preserved" claim VERIFIED byte-for-byte:
size and md5 unchanged, tar valid end-to-end, no .tmp residue. Restoring that
preserved unit under the same constraint returned correct data and claimed
success honestly. Explicitly NOT the R-156 family.
3. But it is SILENT - R-158, filed. Zero events reached the hub.
NotifyBackupFailed exists and the hub allowlists backup_failed, but the only
production caller is the off-box/NAS leg (main.go:659); the backup manager has
tier2/offbox/offbox-enlarge notify seams and none for the local recovery-unit
capture. This is R-97's shipped defect exactly one tier over, and the fifth
instance of "seam built but never wired" - a pattern the codebase names in its
own R-97 wiring test.
Sizing rule corrected: unit ~= volume-tar bytes + logical dump bytes, not a
constant 1.90x. Measured C1: file-only apps are 1.00x (homebox 2305->2305 MB, no
db-dumps dir at all), and the SAME DB app with an empty DB is also 1.00x. So a 20G
sys_drive holds ~19 GB file-only or ~10 GB DB-backed. That is the bound on D5's
Lane-1 independence.
C2: both representations are used for a reason stated in code (F17 - the dump is
authoritative and WINS over the tar; R-47 - replayed with only the DB service up).
The dump is single-database pg_dump --no-owner, so a fresh initdb plus the dump is
logically sufficient and the tar is a PHYSICAL FALLBACK. Dropping it would halve
DB-app units and also close the D5/R-127(b) password trap (restored PGDATA makes
postgres skip initdb and ignore POSTGRES_PASSWORD) - but only after ValidateDump
is promoted from a warning to a gate, since it currently WARNS on a dump whose
accounts table has no rows. In its present form the tar is load-bearing.
No production code, no template change. Teardown still owed and itemised.
|
||
|
|
5f35aa0346 |
Campaign 10: M-band RTO measured — RTO ~= 40s + 26.9s/GB, and a capacity ceiling that matters more
The S figures (66 MB -> 42.0s, two passes agreeing to 0.6s) had a spread tight enough to prove fixed work dominates, which is exactly why they said nothing about M. Second point taken 327x larger, same app, same method: clock from restore request to the app serving the correct discriminator. rallly's postgres volume grown 66 MB -> 21.1 GB (200k rows, STORAGE EXTERNAL so TOAST cannot compress it into a fake number). Two reps: rep 1 backup 406.4s unit 41149 MB RTO 624.5s discriminator correct rep 2 backup 387.2s unit 41133 MB RTO 591.8s discriminator correct 327x the data cost 14.5x the time - strongly sub-linear: RTO ~= 40s + 26.9 s/GB backup ~= 29s + 17.4 s/GB 10 GB -> 5.2 min 20 GB -> 9.6 min (measured 10.1) 100 GB -> 46 min The fixed ~40s dominates below ~1.5 GB, which IS the S band and explains its tight clustering. The more consequential result is capacity. A DB-backed app's recovery unit is 1.90x its data (volume tar PLUS SQL dump): 21.1 GB produced a 40.2 GB unit. The default appliance ships /mnt/sys_drive at 20 GB, so the largest app that can hold a local Tier-1/2 recovery unit on a default box is about 10 GB - and that fills the volume. The M band does not fit on a default box at all; this test only reached 21 GB because sys_drive was first grown 20G -> 70G with the same operation the product performs via SysDataGrowGB. A tier-sizing decision, not a defect, but it is invisible until an app crosses it. Caveats stated in the doc: two points define a line but do not test linearity; the 1.90x is DB-app-specific and a file-only app should be nearer 1.0x (inferred, not measured); synthetic incompressible data; one app, one box. |
||
|
|
7ba7c2a271 |
Campaign 10: final results — 39 cycles, full atom set, R-156 + R-157, no leaks
Phase B completed in three passes: run 1 (27 cycles, 6 atom families, 0 violations), run 2a (10 cycles, stopped deliberately - two violations were harness defects), run 2b (39 cycles, 12 of the brief's ~13 atom families). 1461 invariant checks. Depth reached 39 consecutive cycles, past the brief's "drift at the thirty-eighth", with c34-c39 clean on every invariant. I7 headline: 66 restores across both passes, 66 correct discriminators - never stale, never empty. I2/I3/I4/I5/I6/I10/I11 zero violations in either pass. I1-under-load 5/5: the target pulled WHILE a backup ran still produced backup_target_absent and a clean recovery. R-117's Q7 case holds - a filesystem aborted in place surfaces and the gate stops the app on the dead namespace. RTO Tier-1 rallly 66MB: run 1 median 42.0s, run 2b median 41.4s over 38 restores - two independent passes agreeing to 0.6s. S band's lower end only; nothing extrapolates to M or L. RPO not measured. Monotonic growth, 9457 samples of 19 metrics over 13.5h: NO leak. Controller and agent RSS flat, fds flat, no orphaned volumes/images/containers despite dozens of redeploys, kills, reboots and hard resets. Only curve with real slope is the agent journal at ~20MB/h, bounded by journald. Findings: R-156 (papra's data neither persisted nor backed up, reports healthy) and R-157 (bootrecon's start-once sweep, two mechanisms - the zero-container one is silent on every channel). Four suspicions investigated and DISPROVED, each recorded with what settled it. |
||
|
|
405a795e32 |
Campaign 10: RESOLVED — the backup_target_* silence was transient and self-recovered; SQLITE_BUSY drops are absorbed by retry
Both halves of the disposition were run and neither survived as a finding. The I1/I1-pair violations cluster at cycles 31-33 and nowhere else across 39 cycles; c34-c39 are clean, so it recovered with no intervention. Final tally I1 37 PASS / 2 VIOLATION, I1-pair 36 PASS / 3 VIOLATION. On the quiesced box one slow detach with 4 minutes either side produced a perfect pair. And the alarming false-healthy (mentes bound=False while degraded=false) does not survive quiescence - I had been reading the two halves at different instants of a detach. No R-n. Separately cleared: the hub's SQLITE_BUSY event drops. 7 in 24h including one for the real customer demo-felhom, and the hub does return 500 with notification dispatch only after a successful save - so a lost event would be a lost alarm. But the controller retries 3 times and ZERO events exhausted their attempts; the 07:04:39 drop landed at 07:04:42. Nothing lost. Only cosmetic note: the ERROR line reads like data loss and is not. |
||
|
|
1931dfcb0c |
Campaign 10: R-157 second mechanism — the zero-container case, which is SILENT
The 4th hard-reset failure had a different signature, verified not assumed: all of rallly healthy, papra missing entirely with state=stopped deployed=True containers=0. Zero containers is exactly what bootrecon deliberately never touches, because the UI's Stop is compose down which removes containers - but a hard reset landing during a compose operation produces the identical state. The signature the safety rule depends on cannot distinguish the two. Worse: in that state the deadapp check reported 0 currently down while a deployed app was not running. No app_start_failed, no banner. That is the workspace's own false-invariant #4 (F-CRIT-1, StateStopped assumed deliberate) recurring through a hard reset rather than quiesce. NOT filed as new - CLAUDE.md already records it - but confirmed live on 0.188.0 via a new path. papra returned after ~15 min, later than the harness's 10-min window, so this instance was slow rather than permanent and the doc says so. What restarted it is not established. Mechanism A (Exited, missed by the unsettled snapshot) alarms; mechanism B (zero containers) is invisible on every channel. A settle-condition fix closes A only. |
||
|
|
2d64ee7241 |
Campaign 10: OPEN observation — backup_target_* pair went silent under rapid cycling
Three I1/I1-pair violations in ~5 minutes, all "expected event absent". Recorded as an OPEN observation, NOT a finding: the system was mid-abuse when it was seen, and a verdict taken on a system being hammered is worth little. Established: it is not hub-side suppression and not a truncated log. The hub pod has 0 restarts over 43h and the controller's own log matches it line for line, so the events were never emitted. It is specific to the backup_target_* pair - the generic storage_disconnected/reconnected pair for the other drive kept firing normally throughout the same window. Also sampled, and the more serious half if it survives quiescence: mentes reads bound_under_parent=False while the backup-target state simultaneously reports degraded=false. Those cannot both be right - a false healthy on the backup target is I5/I6's failure mode. NOT established: whether the pair recovers once cycling stops (the harness detaches every ~2 min; a customer does not), whether the 02:25:37 controller restart is implicated, and whether the degraded=false sample was transient. Disposition written into the doc: after the run ends, quiesce with both drives attached, then do ONE slow detach/reattach and see whether the pair fires. That distinguishes "does not survive rapid cycling" from "the target alarm has silently stopped working", which would be severe. |
||
|
|
3d4c5365c1 |
Campaign 10: correct R-157 — the failure is INTERMITTENT (3 of 6), not deterministic
The first write-up said R-157 reproduced "at the same cycle in both runs - deterministic, not a coincidence". Wrong. The cycle numbers matched only because the runner's RNG is seeded so both runs drew the same permutation. The failure itself is a coin flip: run 2b's four hard resets went PASS(c2), FAIL(c10), PASS(c18), FAIL(c26); run 2a went PASS(c2), FAIL(c10). Three failures in six. The correction matters because it changes what kind of bug this is, and it strengthens rather than weakens the root cause: intermittency is exactly what a race against container-state settling predicts, whereas a wrong predicate would fail every time. Signature is identical on all three occurrences: rallly Exited 255 with rallly-postgres healthy, bootrecon reporting "no boot-orphaned apps" about 5s after controller start, and the container count still churning after the sweep (third occurrence 01:05: refresh 8, bootrecon 01:05:13, then 8 -> 7 -> 8). |
||
|
|
7f6b00375b |
Campaign 10: R-157 — bootrecon's start-once sweep misses the boot orphan it exists to recover
Reproduced twice, two independent runs, same cycle (the runner's RNG is seeded so both drew the same permutation - deterministic, not coincidence). A hard reset mid-backup brought everything back except the app half of the DB-backed stack: rallly left Exited 255, oom=false, restarts=0, its own log ending "Ready" - it died healthy - while rallly-postgres returned healthy. 20:28:13 Status refresh: 8 containers across 55 stacks <-- docker ps -a shows NINE 20:28:18 [bootrecon] Boot reconciliation: no boot-orphaned apps 20:28:25 Status refresh: 7 ... 8 containers <-- still churning AFTER the sweep 20:39:14 [deadapp] 20 scans, 5 deployed evaluated, 1 currently down The predicate is sound: once settled the controller reports rallly state=degraded containers=2, and IsDownState includes StateDegraded, so len>0 && IsDownState holds. The SNAPSHOT was wrong. bootrecon fires as a goroutine ~5s after start while docker is still restoring containers, and is start-once by design, so it never re-checks. Consequence: the app stays down indefinitely. Detection is perfect and recovery never happens - R-52's original shape, an alarm with no recovery. Not fixed. Distinguished from this campaign's two earlier HARNESS defects: both drives bound, every other app returned incl. the drive-backed one, only the app half of a two-container stack missing while its DB is healthy, and it surfaced through the fixed check written for exactly this. |
||
|
|
80db2c103a |
Campaign 10: full write-up of the run-2a harness defects
The previous commit message was truncated by an unescaped paren in the shell, so the fix detail and the product observations were lost from the record. This adds them as evidence, where they belong. Covers: the cc_proof table showing no C010-A row at all (the seed never landed); both harness defects; why an ambiguous I7 justified stopping a 10-cycle run; the red-proofed controls; and two transient product observations recorded but NOT filed as findings - the health probe naming the DB container on the app's port for about 70s during recovery, and a ValidateDump WARN on a dump taken while the app was down. |
||
|
|
9ca57e591b |
Campaign 10: two run-2a violations were HARNESS defects, not product defects — fixed
Run 2a hit its first two violations at cycle 10 and BOTH trace to my harness, not the product. Recorded in full because a check that fails for the wrong reason is as corrosive as one that passes for the wrong reason. HARD-RESET VM returned=True canaries_intact=False I7 want=C10-C010-A-194530 got=C10-C009-A-192929 restore_ok=True Root cause, evidenced: the cc_proof table's highest row is C10-C009-A — there is NO C010-A row at all, so the seed never landed. The hard-reset atom ran earlier in the same cycle and left rallly Exited(255); atom_restore_verify called seed() and never checked its return value, so an unwritten generation became a fake stale |
||
|
|
ac6c05bd7b |
Campaign 10: add monotonic-growth sampling — the half the invariants cannot see
I1-I11 are CORRECTNESS invariants: they answer 'is the system telling the truth this cycle'. All 586 of them passed in run 1 while nothing at all watched whether disk usage, snapshot count, log volume, fd count or RSS climbs. Accumulation is exactly what depth was for, and it was missing from the invariant list. Adds c10growth.py (Campaign 2's controller_rss.tsv precedent, widened to 19 metrics) sampling every 90s as a SEPARATE process, so the in-flight run 2 did not have to be restarted. Attributes every sample to a cycle by reading the runner's status.txt, and records NA rather than dying when the box is down during a hard-reset or reboot atom. c10growth_report.py turns it into Campaign 2's table shape (start/end/min/max/ slope-per-cycle) and splits verdicts by class: growth in RSS/fd/volumes/images/ restarts is a LEAK; growth in backup storage or the qcow2 is expected accumulation, reported with a projection to cycle 45. Caught a bug in the sampler itself on the first analysis: MENTES_USED_MB appeared to jump 623 -> 5667 MB, which is exactly ROOT_USED_MB — when a drive is detached, /mnt/<name> reverts to a plain directory on root and df silently reports the ROOT filesystem. The same class of error as the agent's exactMount check, in the measurement code. Gated on mountpoint and red-proofed both ways: a real mount returns a number, a non-mount returns NA. |
||
|
|
816c59c43a |
Campaign 10: R-117 Q7 (fs aborted in place, device present) proven PASS; extended atom set
The case R-117's spike called the worse half — a drive dying with no detach/return cycle, which before agent v0.117.0 emitted nothing on any channel indefinitely. Box runs 0.119.0. Aborted ext4 in place (abort,emergency_ro; device still present): bound_under_parent went false, storage_disconnected fired, the storage page named the stopped app, and calibre-web (whose library binds that drive) was STOPPED rather than restarted onto the dead namespace. Recovery needed a full device close, not a remount — exactly as the fix intends (BindAborted => quiet no-op). Runner extended with the 7 atom families run 1 skipped: abort-fs-in-place, kill-agent-mid-backup, hard-reset-VM-mid-write, reboot-VM, concurrent backup+restore, concurrent backup+detach, fill-drive-near-full. Also fixes a run-1 flaw recorded in the audit: reboot was appended AFTER the shuffle so it never interleaved with a detach; heavy atoms are now permuted in with the rest. Run-1 evidence preserved as *-run1.* (cycle numbering restarts per run). |
||
|
|
69f896d3cd |
Campaign 10 Phase B: 27 cycles, 586 invariant checks, 0 violations
Ran the soak on the Phase A rig. Ended on its own deadline — no watchdog halt, no atom exception, no I11 breach. I1 28+28 pairs, I2 28+28 pairs, I3 56, I4 56, I5/I6 28 each, I7 28, I10 135, I11 28. Zero violations. The row counts are themselves the no-silent-skip check: I3/I4 twice per cycle (both drives), I10 = 5 secret-class fields x 27, REBOOT on cycles 7/14/21 only. I7 is the headline: 28 restores, 28 correct discriminators — never stale, never empty. RTO (Tier 1, rallly, 66 MB): min 38.8s, median 42.0s, p90 42.5s, max 44.3s. That is the S band's lower end ONLY; the 5.5s spread over 28 runs says fixed work dominates, so nothing extrapolates to M or L. RPO not measured. Every atom and invariant was proven BY HAND before automation — the runner asserts nothing that was not first observed live. Caught a Phase A gap before starting: no app had HDD_PATH, so all data sat on the system disk and I3 could never have fired. Deployed calibre-web onto adatok first; otherwise the run would have produced 27 green cycles that tested nothing cross-drive. Investigated and DISPROVED a suspected defect (audit 5.2): /api/disks reports state=attached for a physically absent drive, and intermediary.go:230 really does compute presence from State=="attached". It is inert — planDriveGates only gates paths under /mnt/felhom-drives/ and uses BoundUnderParent there, which was correctly false. The gate fired; the storage page showed "Meghajtó leválasztva". No R-n minted. Honest gaps: 6 of ~12 atom families ran. Not run — Tier 3 (structurally un-isolatable), abort-fs-in-place, kill-agent-mid-backup, hard-reset-mid-write, reboot-VM, both concurrency atoms, fill-drive-near-full. I8 not checked, I9 not automated (cited from the tester-gate run, not re-claimed). kill_controller is NOT mid-backup and reboot_guest never interleaved with a detach. 27 cycles does not answer the brief's question about drift at the thirty-eighth. Teardown still OWED, including hub customer c10-soak (disposition: DELETE). |
||
|
|
4691aa1a35 |
Campaign 10: Phase A complete + gated; Phase B not run; R-156 filed
Phase A passed every gate on a fresh box built from the PUBLISHED ISO 1.26.1: install, claim, two drives enrolled through the real endpoints with the backup target healthy, four apps spanning both sides of D5's secret split, and a working discriminator across all four. Isolation gate: both denials captured, each with a positive control. The PBS control FAILED first — four clean-looking 403s were worthless because the token was denied on its own datastore too (PBS token privilege separation). Fixed and re-run; the denials stand. R-156 (new, register grepped): papra's data is neither persisted nor backed up, and it reports healthy. The template mounts papra_data:/app/data; the app writes /app/app-data/db/db.sqlite. Volume empty and root-owned against a -rootless image, real DB in the container writable layer, healthcheck only probes the HTTP port. Its Tier-1/2 backup is real, verifiable and contains nothing. Not fixed. Tier 3 could not be isolated so it was not run: offsite hard-requires the DR tier (configs.go:1300) and the DR tier only provisions on ep0 (per-endpoint allocation deferred, hub/README.md:260). Both are recorded deliberate positions, so no R-n minted. The campaign touched neither ep0 nor the Storage Box. Phase B did not start. Phase A was budgeted at ~1h and took ~5.5h (1.26.1 is a public release image with no auto-install path, so the install was a blind screendump+sendkey walk). That left the runner — which judges eleven invariants and fires destructive atoms unattended — to be written at 04:00 with ~3h of night left. Stopped on the brief's own fence: a rig producing false negatives is worse than no rig. The rig is built and idle; teardown is OWED and itemised, including hub customer c10-soak (disposition: DELETE). |
||
|
|
e9a74a0019 |
docs: remove a gate criterion that could never pass, and close three register rows
PART 1 — the release gate.
G7 required the packaged .deb to sha256-match the one built from committed source. That is
unsatisfiable BY CONSTRUCTION: dpkg-deb stamps the build time into every archive, so two builds of
byte-identical source differ. It was already failing when the 1.26.1 release ran it. A criterion
nobody can satisfy gets waived once and read as advisory ever after — which is how R-29's shelf of
never-run gates was built. Sub-clause dropped, reason recorded in G7's own note the way G6's
amendment was, so a future reader can restore it if SOURCE_DATE_EPOCH ever makes it meaningful.
RULING ASKED FOR — is payload integrity covered by G9 alone? NO, and G9 is widened rather than a new
criterion invented. The package ships TWO payload files (build-deb.sh:54-55); G9 checked only the
script. The systemd UNIT was covered by nothing: G7 covered the container, G8 covers the postinst
behaviourally, G13 covers directory presence. The unit is not incidental — its After=, its
ConditionPathExists= and its Restart= decide WHEN AND WHETHER day-0 runs at all, so a drifted unit
would have shipped silently. Same shape as the /etc/felhom miss that G13 exists to prevent: a check
that proved the thing present and said nothing about what it depended on. The check passes today.
G13 moved to sit after G12 — it was minted late and left between G10 and G11.
PART 2 — register dispositions. BASELINE DISCREPANCY, reported rather than worked around: only R-128
had a row. R-154 and R-155 had NO row in either file — minted in a spike document and never carried
across, which is R-123's class, not the drift the task described. Rows created, closed, with the
reasoning, because in all three cases the reasoning is the durable part:
R-128 closed by CORRECTING a false claim, not by making the assertion real — the coupling does not
exist and asserting it would invent a constraint. Flagged so nobody 'restores' it.
R-154 closed with the measurement and where it now lives in pushed source.
R-155 NARROWED, not deleted — unchanged for FELHOM_MENU=single, inapplicable to release. Flagged so
the guard is not later removed wholesale on the strength of 'R-155 closed it'.
Documentation only: no code, no build, no ISO, no upload, no box touched.
|
||
|
|
f2fc76ec4b |
ISO v1.26.1 PUBLISHED — both entries proven, round trip verified
Live: https://iso.felhom.eu/felhom-installer-1.26.1-pve9.2-1.iso sha256 f3cc86d5f0ec68bba4155c994b4fa84e208d50209bb6e815636c99e5441059a6, 1705322496 bytes. PART 5 PASSED ON BOTH MENU ENTRIES, four observables each: Graphical spikegfx.felhom.eu pairing code J7N-2DA TerminalUI spikesix.felhom.eu pairing code ZY5-YY4 Both: manual install, own disk, own password, real completion signal, and the journal's 'not bound yet — polling every 30s ... normal waiting state, not an error'. Spike 4 had REASONED the graphical path follows from shared Install.pm; it is now measured. PART 6: G1-G10 + G13 all PASS against the uploaded file. G4's single hit is felhom-bootstrap.sh:480's substring TEST ('$envtext' != *FELHOM_RETRIEVAL_PASSPHRASE=*), not a value — my own regex matched the glob's asterisk. PART 7: uploaded via rclone in a container configured ENTIRELY by environment variables, so no credential file was ever written. Round trip verified from the public URL — not the local file. Bucket stays private: unauthenticated GET to the S3 endpoint 400, custom domain has no index (404). CORRECTED BEFORE UPLOAD: the generated manifest described a single automated entry with a 5s timeout and listed Graphical/Terminal UI as 'menu-removed'. Generator fixed, sidecar regenerated, and the ISO verified byte-identical before and after — the published file IS the file Part 5 validated. Hub-side cleared: appliances 16, 17, 18 discarded (303 each); zero rows remain. The endpoint is /appliances/<id>/discard, POST only (server.go:345) — not /delete. Teardown: VMs purged, spike5 storage removed, demo-hp back to 6.6G, drill-r50 and 9201 untouched. Still open and named: OPEN-ITEMS/ROADMAP dispositions for R-128/R-154/R-155 are not written; the .deb is not byte-reproducible (G7 sub-clause); before-network stub unreached; Secure Boot and real hardware not exercised. |
||
|
|
70d034a3f3 |
iso: the release manifest described a different image than it shipped
The 1.26.1 manifest — the file a tester reads to know what they have, and which is published alongside the ISO — carried four statements that were false for a release build: boot-menu 'single entry Felhom telepítés, default, 5s' -> it has TWO, timeout 15 menu-entries '1 (... timeout 5s)' -> 2 menu-removed 'Graphical, Terminal UI, ...' -> those are exactly what it SHIPS kernel-line '... proxmox-start-auto-installer' -> the release menu deliberately has none secret-bearing 'no (embeds the customer retrieval passphrase...)' -> self-contradictory All four came from branding/pairing notes that predate --release and were emitted unconditionally. A public artifact whose own manifest misdescribes it is the false-claim class this arc exists to correct, so it is fixed before publication rather than after. |
||
|
|
52e5cdb86a |
REPORT: /etc/felhom fix verified — TUI entry PASSES all four. Graphical untested, so still NOT PUBLISHED
iso 1.26.1, sha256 f3cc86d5f0ec68bba4155c994b4fa84e208d50209bb6e815636c99e5441059a6.
Terminal-UI interactive install, host spikesix.felhom.eu — all four observables PASS:
1 package installed ii felhom-bootstrap 1.26.1
2 unit enabled enabled
3 unit fired first boot activating; 'registering unclaimed appliance at the hub'
4 box wants a claim code /etc/felhom/appliance-pairing-code = ZY5-YY4, token 64B mode 600,
'not bound yet — polling every 30s ... normal waiting state, not an error'
That is the product working end-to-end from a public image on a manual install.
G13 added and RED-PROOFED (removing install -d -> exit 3 with the G13 message; restoring -> green).
The first red-proof attempt was INVALID — a copied script failed on a missing control file, i.e.
non-zero for the wrong reason — and was redone in place.
NOT PUBLISHED: Part 5 requires BOTH entries. The Graphical entry reached the Target-Harddisk screen
but was not driven to completion (Enter lands in the Country field; monitor mouse_move does not move
the guest cursor), so Part 5 is not fully passed and Part 7 did not run.
TWO HUB ITEMS OUTSTANDING and NOT disposed of: unclaimed appliances 16 and 17. GET/POST on
/appliances, /appliances/17 and /appliances/17/delete all 404; the rows appear only inside /hosts,
which offers 'Bind & deliver' and no delete affordance. Stated at the top of the report too, because
R-131 is four orphaned objects that recorded commands never cleared.
Also recorded: 'qm set --scsi0 ... --boot' silently yields net0;ide2, and ide2-first sends a finished
install back into the installer — both made a COMPLETED install look like a stuck one.
|
||
|
|
a967da7d2c |
iso 1.26.1: ship /etc/felhom/ — the directory the bootstrap writes its state into
FIX for the Part-5 failure. felhom-bootstrap.sh writes the appliance token (:431), the pairing code
(:435) and .bootstrap-done into /etc/felhom/. The old stub-first-boot.sh created it explicitly
('install -d -m 0755 /etc/felhom /usr/local/sbin'); packaging dropped the env FILE correctly and the
DIRECTORY with it. Measured consequence on a real interactive install: the box registered at the hub,
could not persist its token, and polled 'HTTP 401 — still retrying' forever with no claim code.
- build-deb.sh now ships ./etc/felhom/ (0755, empty) and ASSERTS it, plus ./usr/local/sbin/ and
./lib/systemd/system/, as G13. RED-PROOFED: removing the install -d makes the build exit 3 with
'is not in the package (G13)', and restoring it goes green.
- The gate gains G13 with the reasoning: G7/G8/G9 all passed on the broken package. G9 proves the
payload is the right payload and says NOTHING about what the payload depends on.
ISO_VERSION -> 1.26.1.
|
||
|
|
4ea211f67f |
REPORT: Part 5 FAILED — the package never creates /etc/felhom. NOT PUBLISHED
The Terminal-UI interactive install ran to completion from the release image and gave 3 of 4 required observables: 1 package installed PASS ii felhom-bootstrap 1.26.0 2 unit enabled PASS wants-symlink present; postinst enabled it from the chroot 3 unit FIRED first boot PASS journal shows PAIRING mode, registering unclaimed appliance 4 box wants a claim code FAIL /etc/felhom/ does not exist on the installed system, so felhom-bootstrap.sh cannot write the appliance token (:431) or the pairing code (:435), and the hub poll then 401s forever. The box can never finish pairing and the customer never sees a claim code. ROOT CAUSE, mine: stub-first-boot.sh opened with 'install -d -m 0755 /etc/felhom /usr/local/sbin'. This task correctly dropped the env FILE from the package and dropped the DIRECTORY with it. felhom-bootstrap.sh uses /etc/felhom for its runtime state (token, pairing code, .bootstrap-done). WHY THE GATE MISSED IT: G9 proves the packaged script is byte-identical to HEAD, and it is. I verified the payload files and never the directory the payload writes into — a check that proves the thing present and not the thing it depends on. Added as G13. The fix is one line and is deliberately NOT applied: a failing Part 5 stops the task, and proving a fix needs both installs re-run. Also recorded: 'qm set --scsi0 ... --boot order=scsi0;ide2' silently yields boot: order=net0;ide2, so a COMPLETED install looked like a machine sitting in the installer. Set --boot separately. Nothing uploaded; R2 credentials never read. Teardown complete: VMs purged, spike5 storage removed, demo-hp back to 6.6G, drill-r50 and 9201 untouched. Hub-side: no appliance object was created (searched /, /hosts, /configs for the hostname — zero hits), so R-131 gains no row. |
||
|
|
246036605a |
REPORT: universal ISO built and gated — NOT PUBLISHED (Part 5 incomplete)
Publication is gated on Part 5 (two interactive installs proving delivery end-to-end). Neither was carried to completion, so nothing was uploaded. Per the task: stopping is the good outcome. BUILT: felhom-installer-1.26.0-pve9.2-1.iso sha256 24977bafd24d73262745fc1b3040939469c9b23b87ead927a8af86de73044a90, 1705322496 bytes. GATE (Part 6) against that exact file: G1-G10 PASS, G11/G12 not run (nothing published). G5 by ENUMERATION vs the stock PVE ISO: exactly four added paths — three felhomtheme files and /proxmox/packages/felhom-bootstrap_1.26.0_all.deb. G9: the packaged felhom-bootstrap.sh is byte-identical to repo HEAD. No .rootpw.txt is emitted at all, which is G2's own evidence. PROVEN in Part 5 before stopping: the image boots to the branded TWO-ENTRY release menu and the Terminal-UI entry reaches the stock PVE installer. NOT proven: package installed, unit enabled, unit fired, box asking for a claim code — on either entry. The Graphical entry was never driven. Reporting partial observables would be the LastRun-class error this arc has corrected three times. R2 credentials were never read, never used, never echoed; no rclone/aws config was created. Teardown complete: VMs purged, scratch storage spike5 removed, demo-hp back to 6.6G, drill-r50 and 9201 untouched, hub-side nothing created (verified by fetching the customer list). |
||
|
|
8505954df7 |
iso: ship the package under its real filename, not felhom.deb
The release build copied the .deb to $WORK/felhom.deb before handing it to the repack, so the ISO carried '/proxmox/packages/felhom.deb' — the version invisible from the image, and not matching the release gate's 'exactly one felhom-*.deb' check (G7). Caught by running G5's enumeration against the built artifact rather than trusting the build log. |
||
|
|
0f22447d15 |
iso: two build-log lines stated things that were not true
Neither changes an artifact, but both are read by an operator deciding whether a build is sound: - the closing banner printed 'root-pw : <iso>.rootpw.txt ... the console credential for this build' unconditionally. In --release mode no password is minted and no such file is written (verified: the release build emits only .iso, .sha256 and .manifest.txt). It now says so. - the repack's menu-surgery line hardcoded '1 entry, 0 submenus' and printed it after a gate that had just accepted TWO. It now reports the counts it actually asserted. |
||
|
|
57b87ec9be |
iso: actually pass FELHOM_MENU/FELHOM_DEB to the repack
The repack read both correctly; the caller never set them, so a --release build reached the narrowed R-155 guard still in 'single' mode and was refused (rc=10). Caught by the build's true exit code. Also copies grub-release.cfg.tmpl into the brand dir and fixes the branding log line, which claimed 'single-entry menu' unconditionally. |
||
|
|
a4d7dfd336 |
iso: --release must satisfy the mode check it is a third case of
The mode validation still required one of --bootstrap-env / --pairing, so --release died at 'one of --bootstrap-env (direct) or --pairing (generic) is required' before reaching its own validated branch. Caught by the build's true exit code (rc=1), not by a pipe. |
||
|
|
01a8155c5a |
iso v1.26.0: the PUBLIC release image — no answer file, interactive install, day-0 by .deb
Design inputs: SPIKE-universal-iso-{1,2,3,4}-2026-07-31.md. Every choice below is a measurement.
NEW: scripts/iso/pkg/ — the felhom-bootstrap .deb, built from committed source.
Two files only (script + unit), NOT three: felhom-bootstrap.sh:91 reads /etc/felhom/bootstrap.env
only 'if [[ -r ]]', and its defaults at :95-96 are EXACTLY what the pairing env set
(build-felhom-iso.sh:257-258) — so shipping it would add a 0600 file to a public package to express
values the script already defaults to. NO dependencies: the binaries it calls run at FIRST BOOT,
not at postinst time, so SPIKE 4's open 'dpkg --configure -a' ordering question does not arise.
The postinst is structurally incapable of failing (no 'set -e', every statement guarded, ends
'exit 0'); build-deb.sh self-asserts G8/G9 and REFUSES to emit a package that violates them.
iso-repack.sh — two changes, both narrowing rather than deleting:
- R-155 guard: now applies to FELHOM_MENU=single ONLY. It protected the single-entry mode's promise
(one button labelled 'install' must not drop into a disk-picker); a release image carries no
auto-installer-mode.toml BY DESIGN (gate G1), so refusing it would be the guard firing on the
shape it describes rather than the one it prevents.
- the menu collapse now has a release mode: two INTERACTIVE entries, Graphical default, timeout 15.
Entry-count and banned-token gates are per-mode; the six-token list is UNCHANGED for single mode.
- .deb injection into /proxmox/packages/, with a skip-list collision check (a colliding name would
be dropped silently — the inert-payload class) and a post-remaster assertion that it landed in
final.iso, not merely in the extract tree.
build-felhom-iso.sh — --release: no profile, no root hash, no answer.toml, no prepare-iso at all.
Skipping prepare-iso is what removes the Automated entry by construction, since the stock grub.cfg
emits it only inside 'if [ -f auto-installer-mode.toml ]'.
R-128 RULING — FIXED, by correcting the claim rather than inventing an assertion for it. The comment
said ISO_VERSION 'aligns with SCRIPT_VERSION'; nothing evaluated it and the two had drifted. The
coupling does not exist: the ISO is frozen, felhom-host-install.sh is fetched at run time from main
(R-94/R-110), so an assertion would invent a constraint. Comment corrected, ISO_VERSION -> 1.26.0.
Release gate G6 AMENDED before the build, with its reasoning recorded in the runbook: the six-token
ban existed to keep users away from the manual installer, which the ruling makes the product.
'proxtui' (the TUI installer we ship) and 'nomodeset' (its graphics fallback) are dropped for
release images; proxdebug/Rescue Boot/memtest/fwsetup stay banned in both modes.
|
||
|
|
e787391c0a |
docs: the public ISO release gate, written BEFORE the first release image
A standard defined in advance cannot be rationalised afterwards, and this is the artifact that most needs one: once a file is on iso.felhom.eu and someone has downloaded it, it cannot be recalled. Twelve criteria, each checkable against the UPLOADED FILE rather than the build inputs, and each carrying the spike measurement that justifies it: - G1 no answer.toml / auto-installer-mode.toml — deletes the whole Spike 1-2 problem space and removes the Automated menu entry by construction rather than by a guard - G2/G3/G4 no root hash, no SSH key, no customer identity — the shared-credential classes - G5 credential scan by ENUMERATION against the stock ISO, not a pattern sweep (Spike 1 found /answer.toml precisely because the earlier recon grepped the wrong file) - G6 menu present, interactive default, timeout >= 10 (Spike 2 lost a probe to a 1-second menu), underscore timeout_style, and the banned-token safety gate kept unchanged - G7/G8 the felhom .deb present, and a postinst that cannot fail: no systemctl start/daemon-reload (no systemd runs in the installer chroot), no network use (the cable may be out), no 'set -e', ends 'exit 0' - G9 felhom-bootstrap.sh byte-identical to repo HEAD — the one frozen, drift-capable payload - G10 every build input committed (Spike 1: demo-felhom came from an uncommitted profile) - G11 published checksum AND a verified download round trip - G12 bucket Public Access stays Disabled Committed on its own, before any build. |
||
|
|
61e9b55737 |
SPIKE 4: a .deb in the ISO DOES deliver on an interactive install
Findings only — no script, profile or build file changed; no release ISO built, nothing published. documentation/audits/SPIKE-universal-iso-4-2026-07-31.md MEASURED, with a control, and the negative control is in the SAME box. One ISO (15 GRUB entries), a trivial probe .deb injected into /proxmox/packages/, two qm-created VMs on demo-hp (400 interactive / 401 automated control) on a scratch dir storage at the /mnt/nvme-1tb mount ROOT. Interactive (Terminal UI) install: - package installed (ii felhom-spike4-probe 0.0.1) - postinst RAN (marker + content intact) - it enabled a systemd unit, and that unit FIRED ON FIRST BOOT (uptime 7.98s, pid1=systemd) - while on the same machine proxmox-first-boot is NOT installed and /var/lib/proxmox-first-boot does not exist — Spike 3's negative reproduced, not assumed. Postinst environment (identical both paths): pid1=unconfigured.sh, NO running systemd, but 'systemctl enable' SUCCEEDS; /proc+/sys mounted; network+DNS happened to be up (inherited from the installer's DHCP — must NOT be relied on). Constraints: never systemctl start/daemon-reload, never require network, never fail, do the real work in the unit at first boot. Repack preserves it, but a naive 'xorriso -boot_image any replay' fails with 'Overlapping MBR partition entries' — iso-repack.sh:270-292 already documents that exact failure and its fix. R-153 RETRACTED into R-94 leg (b): OPEN-ITEMS.md:15 carries it verbatim at READY (XS), and R-29 says explicitly 'do not mint a new ID for a new instance'. Spike 3's further claim that the drift leaves the generator 'three minor versions stale' was FALSE and is corrected — R-94 retracts that exact reading; the served script is always main, so 1.22.0 is what every install already gets. No new R-rows opened. |
||
|
|
bb29186d62 |
SPIKE 3: [first-boot] does NOT fire on an interactive install
Findings only — no script, profile or build file changed; no release ISO built, nothing published. documentation/audits/SPIKE-universal-iso-3-2026-07-31.md MEASURED with a control from the SAME image (one ISO, 15 GRUB entries): - Automated entry -> hook fires: ttyS0 marker, marker file, /var/lib/proxmox-first-boot/proxmox-first-boot (0700), activation symlink, unit active. - Terminal-UI entry, normal manual install -> ALL absent, and the proxmox-first-boot PACKAGE is not installed at all. A whole-filesystem grep for the marker returns nothing. Mechanism cited: Config.pm:118 defaults first_boot.enabled=0 and set_first_boot_opt is never called in the Perl tree; Install.pm:746 returns early without it; Install.pm:1360 skips the package. proxinstall (graphical) has ZERO occurrences of first-boot. [first-boot] is an automated-installer feature, unavailable on every interactive path by construction. R-154. A delivery mechanism DOES exist and is UNTESTED: Install.pm:1343-1372 unpacks every .deb in the ISO's /proxmox/packages/ into the target on every path (fixed skip-list), then dpkg --configure -a runs postinsts (:1378) — how PVE ships first-boot itself. Read from source, not measured. Q5: the public image should carry NO answer.toml at all — that removes the baked root hash, the disk profile and the whole Spike 1-2 problem space, and makes it a one-line release gate. But iso-repack.sh:100-106 refuses an ISO without auto-installer-mode.toml. R-155. Incidental R-153: hub hostInstallVersion=1.19.0 vs SCRIPT_VERSION=1.22.0; hostinstall_gates.py detects it and exits 1 — the gate works, nothing runs it. Q3 (real stub at before-network) was NOT reached and is recorded as not reached. |
||
|
|
19c932a693 |
SPIKE 2 complete: locked root closes the PVE web UI; before-network gives a measured zero window
Findings only — no script, profile or build file changed; no ISO built, nothing published.
documentation/audits/SPIKE-universal-iso-2-2026-07-31.md
Both Tier 0 boxes went offline mid-session (provider cable fault; four routes tried, no Tier 2
fallback used) and returned. All three scenarios then ran to completion on real PVE, each signalled
by reboot-mode='power-off' rather than a disk hash.
- A LOCKED ROOT CLOSES THE PVE WEB INTERFACE. Measured at the exact endpoint the UI uses
(POST /api2/json/access/ticket, root@pam) WITH A WORKING CONTROL: known-password install returns
HTTP 200 + ticket; locked install returns 401 for every password and none can exist.
passwd -S root = L, shadow = literal-asterisk, PVE uses the stock PAM stack.
- GRUB recovery mode is also closed ('the root account is locked') — but init=/bin/bash still gives
an unauthenticated root@(none):/#. A locked box is recoverable, operator-only, at the console.
The installed GRUB has NO password, so locking root is not a physical-security measure. R-152.
- before-network MEASURED (A/B, same image): the hook RUNS (marker, uptime 6.58s) with entropy 256,
writable /etc, all binaries and openssl_rand_len=32, while ip_global is EMPTY and
listen_22_8006 = 0. fully-up is the converse: sshd+pveproxy active, 3 listening. Zero window.
- R-148: answer.toml.tmpl:27 justifies fully-up with a pvesh/pct dependency the stub does not have
(grep rc=1) — it blocked the ordering now measured as the fix.
- R-149 three ordering values; R-150 Condition-guarded hooks skip silently; R-151 demo-felhom built
from an uncommitted profile.
Three probes failed and are recorded as failed: a container probe that ran as uid 0, a GRUB probe
that missed the 1-second menu timeout, and a kernel-line edit one line off (caught by a pre-typing
verification screendump). The interim 'Layer 1 teardown INCOMPLETE' is corrected — the fixture had
never landed, because the staging mkdir was in the SSH call that timed out.
|
||
|
|
5bdd8372f8 |
SPIKE 2: before-network gives a zero window by construction; locked root closes sulogin
Findings only — no script, profile or build file changed; no ISO built, nothing published. documentation/audits/SPIKE-universal-iso-2-2026-07-31.md BOTH Tier 0 boxes went offline mid-session (remote site, 12:28 CEST; four routes tried, our tailscale pod healthy). Q1/Q2/Q3 each keep a part needing a nested VM: those are BLOCKED, not answered. DooPlex was NOT used as a fallback — Tier 2, and this task did not authorise it. Established without them: - STRUCTURAL: ordering='before-network' maps to proxmox-first-boot-network-pre.service (Before=network-pre.target, Type=oneshot) — it completes before ANY interface is configured, so a rotation there has a zero-length window BY CONSTRUCTION, not by being fast. - R-148: the stub does not need 'fully-up'. stub-first-boot.sh has no pvesh/pct/pveum/qm call (grep rc=1); that usage is in felhom-bootstrap.sh under its own After=network-online unit. answer.toml.tmpl:27 justifies the current ordering with a dependency that does not exist. - R-149: the ordering enum has THREE values (before-network, network-online, fully-up), not two. - MECHANISM (container, not PVE): locked root closes sulogin — 'the root account is locked' for both '*' and '!', with a working control. So 'discard' and 'lock' are the SAME outcome for recovery, making the escrow decision binary. - R-150: all four proxmox-first-boot-* units are Condition-guarded; a failed condition is a SKIP, so a hook that never ran looks identical to one that succeeded. - R-151: demo-felhom was installed from an UNCOMMITTED profile — a Tier 0 reference box is not reproducible from main. - Q4: four gates in iso-repack.sh enforce the single-entry menu; default/timeout already settable. The first mechanism probe was invalid (uid 0 bypassed pam_unix; sulogin had no tty) and a teardown error (shredding the control plaintext) are both recorded as failures, not massaged. demo-hp teardown is INCOMPLETE and named as such; the command is recorded, not claimed done. |
||
|
|
ea00976403 |
SPIKE: a universal ISO needs a different disk strategy and a locked root
Findings only — no script, profile or build file changed; no ISO built, nothing published. documentation/audits/SPIKE-universal-iso-2026-07-31.md - R-139 (HIGH): a disk filter matching >1 device does NOT fail safe. Observed in a nested VM — the installer silently picked one of two matching disks and wiped it; validate-answer accepts such an answer. The 'filter did not match any devices' guard covers the ZERO-match case only. - No udev property distinguishes an internal system disk from external media. Measured on demo-felhom with its 1TB external attached: ID_BUS='ata' for BOTH, lsblk RM=0 for both, and device-info exposes no removability property. demo-hp's NVMe carries no ID_BUS/ID_TYPE at all. - R-141 (HIGH): the answer schema makes a root credential mandatory, but root-password-hashed='*' validates AND installs to completion. [first-boot].ordering accepts 'before-network', the only ordering that closes the exposure window structurally. - Q3: prepare-iso leaves grub.cfg byte-identical to stock (15 entries, automated AND interactive) — a two-entry menu is purely a Felhom grub.cfg.tmpl change. - R-129 resolved: demo-hp's key is the operator's own, added post-install; demo-felhom's IS baked by an uncommitted profile. The reachable-before-rotation measurement FAILED twice and is recorded as failed, not inferred. Opens R-139..R-147; restates R-128. |