e78726314ea43aecaea37a2fc081edef922ea607
96 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
eaae94a4b8 |
felhom-host-install.sh 1.29.0: installs the OS-update wrapper (felhom-os-apply) from the pinned agent tag; skipped with a warning on an agent older than 0.140.0; uninstall removes it
gates / gates (push) Successful in 38s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
33bbf8e94a |
Backlog triage Part D: ROADMAP cleaned (161 -> 123 lines, 81 -> 42 KB): intentions re-sorted P2/P3/P4, each names its 00 row; 33 shipped/killed/closed/moved items + the pre-invite checklist -> ROADMAP-HISTORY; UPDATE-ARC collapsed; one_register_gate reads suffix ids (R-50b found, moved to OPEN-ITEMS; decoy seen red)
gates / gates (push) Successful in 31s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
9f77865c2a |
Backlog triage Part C: every open row has a Category (11) and a Sev (P1-P4); OPEN-ITEMS grouped by category, severity order; register_shape_gate RULES 5-8 (columns by header, category, sev, defined state) with five decoys seen red; CLAUDE.md + PROMPT-TEMPLATE: new rows filed into their category
gates / gates (push) Successful in 29s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
71b8c8c62b |
Backlog triage Part B: 125 finished rows + 20 id-less rows moved to CLOSED-ITEMS (full text at 9e2786c); open rows normalised to one 6-column shape; narratives archived verbatim; closed_register_gate RULE 3 refuses a finished row in OPEN-ITEMS (decoys, seen red); rules rehomed to CONTEXT + 07 §11; loose notes triaged; R-814..R-819 filed; register 444 -> 325
gates / gates (push) Successful in 30s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
2866f6a318 |
immich's first start: cause measured, fixed in the catalog (R-732 closed); ISO clean-tree gate (R-730 closed)
gates / gates (push) Successful in 27s
- R-732: the first-start geodata import runs up to 9 concurrent 5000-row INSERTs; the database needs ~400 MB anon + ~170 MB touched shared_buffers (the image's FIXED 512MB, not host-RAM sizing). 512M fits only with swap (bench swap 0: 61-104 kills; 9202 swap 512 MiB: survived by swapping). Controls: swap alone, limit alone flip it; shared_buffers 128MB alone does not. Catalog 56c4888: v3.2.4 + 768M, proven with swap off on both venues. audits/immich-first-start-2026-09-30/A-cause.md. - R-730: scripts/iso/build-felhom-iso.sh refuses an uncommitted/untracked/unpushed tree (no bypass), records repo-commit from the gate and iso-v<version>; test iso/test/clean-tree.sh, red-proof run (status check removed -> 2 of 4 cases fail -> restored). - R-731 narrowed (gitea 28.0.0 GA; mariadb 13.0 a short-term Rolling line). R-676 note. - New rows R-733 (bench has no swap, boxes 512 MiB), R-734 (immich .immich markers -> files_may_change). - STATUS: the golden line corrected (no bake is due; 0.283.1 is the newest release). Register 364 -> 366. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
31eeb36e88 |
ISO 1.29.0 PUBLISHED — bilingual console, proven on both menu entries (R-559)
gates / gates (push) Successful in 22s
Live at iso.felhom.eu, sha256 dceacae5da247d76cad065bf6c0d3bbefac8d8a5f8e571 db2fd8449a97e94829, and both download pages now name it. Every gate criterion is recorded with its OBSERVED value in documentation/tests/iso-release-1.29.0-2026-09-18/ — including two proof installs from the published bytes, one per boot-menu entry, each with a first boot AND one reboot: /etc/issue bilingual with zero hits for 8006, pvebanner masked, package 1.29.0 installed, unit enabled and fired, pairing code present, and the installed script byte-identical to repo HEAD. G11: the downloaded bytes hash to the published checksum. A defect was caught BETWEEN builds by looking at the screen rather than at the config: the second menu entry read "Felhom telepítés (szöveges mód) / Install Felhom (text mode)" — 58 characters — and the GRUB menu box cut it at "Instal". The English half was unreadable on the boot screen. Shortened to "… / text" and rebuilt; the published image is the rebuilt one. The Hungarian half is the part that may not change, so the English half is the part that gave. Teardown: VMs 323/324/325 destroyed, the two unclaimed appliance registrations discarded (zero left in `registered`), guest 9201 untouched — 23 containers before and after. The two stale *.rootpw.txt files were shredded from the publish source directory before the upload ran from it (R-587, files gone; the guard that would stop it recurring is still open). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
eb1ae37095 |
ISO 1.29.0 source + an English download page (R-559 slice 4, source only)
gates / gates (push) Successful in 23s
THE IMAGE IS NOT BUILT AND NOT PUBLISHED BY THIS COMMIT. Publishing to iso.felhom.eu is public and irreversible and its runbook requires a proof install on BOTH menu entries plus the 16-criterion gate run against the exact uploaded bytes. That is the operator's step. The download pages therefore still name 1.28.0 - the image that is actually published - and a new site gate refuses the two pages naming different files or hashes. Three texts a person meets before any dashboard become bilingual: Hungarian block first, byte for byte as before, then English, inside the same frame. The pairing banner, the bound banner, /etc/issue (and the postinst's byte-coupled copy), plus an English half on the GRUB entries. The Hungarian is a GOLDEN, not a grep: test/golden/*.hu.txt were captured from the script at |
||
|
|
5b7b2b22f1 |
i18n starter: inventory (script + audit), rows R-553..R-562, language allowlisted in wire-contract gate
gates / gates (push) Successful in 23s
Phase 0 of the localisation starter: i18n_inventory.py counts every customer-visible Hungarian string; the audit names six further claims in the prompt that live source disproved. Rows for the compare-not-show sites, wizard deletion, the wire-contract comment blind spot, and localisation slices 1-6. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
832218dca4 |
ISO 1.28.0 PUBLISHED on the operator's yes; download page and R-535 updated
gates / gates (push) Successful in 20s
Uploaded with env-only credentials and verified by ROUND TRIP: the downloaded bytes checksum to a4cd9b6d…, identical to the built file, and the checksum file is served. 1.27.1 stays in the bucket; nothing was overwritten. The download page now names 1.28.0 with the published checksum (BOM preserved, site gates green). R-535 closes with an honest caveat: the new banner ships byte-identical to repo HEAD and the string is in the published payload, but it was never seen on a screen — the box bound itself while the walk was headless. Also corrected: the 1.27.1 heading still said NOT PUBLISHED although it went out on the big night. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
c033b3b617 |
ISO 1.28.0 source: the console stops showing the pairing code once bound (R-535)
gates / gates (push) Successful in 20s
Measured 2026-09-16: 25 minutes after a successful bind AND claim the console still showed the pairing code under a line promising the screen refreshes itself. print_bound_banner is printed the moment the bind delivery lands. It does NOT name the dashboard URL: the one-shot delivery carries the customer id, passphrase and mode, not the domain, so naming an address would mean inventing one. The residue — the console still does not reflect the later CLAIM, because this unit has exited by then — is recorded in the changelog rather than implied away. Not published: the built image needs the release gate and the operator's yes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
65790672d5 |
doorstep walk on ISO 1.27.x: 1 intervention (R-505), STOP before publish
gates / gates (push) Successful in 17s
ISO 1.27.1 gated PASS and proven live: first-boot console Felhom-only, pvebanner masked across a proven reboot. Hub v0.113.0 hand-over copy live (R-497 closed). Full first hour walked again on customer tester-1 (three disks + one disk): deploy, use, backup, removal, byte-identical restore, power cut, typo all PASS. The tunnel gives a fresh box no routes: 12/12 503 from DooPlex (R-505); the record has no e-mail (R-508). Rows R-507, R-508 filed; R-496/R-495 fixed/answered awaiting publish; day-0 A.1 no longer claims the controller creates hostnames (R-506). NOT PUBLISHED. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
6fd8c87516 |
doorstep: console is Felhom's (ISO 1.27.0 source), passphrase hand-over copy (hub 0.113.0 source), rulings
gates / gates (push) Successful in 19s
Phase 0: the public ISO never auto-installs by construction (no answer.toml, G1); the operator re-affirmed the interactive installer 2026-09-14. - felhom-bootstrap.sh: mask pvebanner.service, write a Hungarian /etc/issue (no :8006 admin URL); pairing banner names the Tulajdonosi jelmondat and paints through the CONSOLE_DEV seam (R-496). Harness: 8 checks, red first; fake hub now sends a pairing code (the banner was never tested, R-502). - hub: created flash + Credentials block tell the operator to hand the phrase over; the self-bind mail names the operator (R-497). Tests red first. - iso-release-gate G14-G16; domain ruling in 01-topology + CONTEXT; R-494 narrowed to P3; R-502..R-504 filed; volunteer guide and day-0 A.2 aligned. ISO_VERSION 1.27.0 (not built, not published). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
5e8a82c3c4 |
night 2026-09-13/14: first "be a customer" rotation (adventurelog) — 7 defects found, 13 rows closed
gates / gates (push) Successful in 18s
New runbooks/nightly-rotation.md; observations_gate.py reads every section (R-471); target-selection.md names real paths (R-461); R-93 carries the fact that drill-r50 is gone. Register: R-473/R-474/R-466/R-471/R-453/R-461 and v0.240.0's R-477/R-478/R-480/R-482/R-484/R-485/R-486 closed; R-481, R-483, R-487, R-488, R-489 opened. 09 §6.1, 07 §6, CONTEXT, STATUS note. Evidence: audits/nightly-2026-09-13-adventurelog/, audits/v0240-2026-09-13/. |
||
|
|
5ef0f52bcd |
Slice 4 shipped (R-448/R-443/R-439 CLOSED, proven live); R-472..R-476; the floor-between-bakes claim corrected
gates / gates (push) Successful in 19s
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure. Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/). Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the runbook, STATUS, CONTEXT, R-468 and the gate docstring. Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
ae59c31a84 |
R-459 CLOSED (MariaDB converts itself, proven by harness + live), golden 0.236.0 (R-467), the golden waiver (R-468)
Operator rulings 2026-09-13, both shipped the same day: - MariaDB finishes its own conversion (catalog eec1228/bd32830/3525e35). Harness E3/E3b `proven` with engine_state_after "already upgraded to 12.3.3-MariaDB [exit=1]", the skip line gone, C3 still `failed`; landed on demo-hp through the real 15-min cycle, nothing recreated, one deliberate restart logged "MariaDB upgrade not required" with the app serving. Evidence: documentation/audits/r459-close-2026-09-13/. The engine-major rule + gate keep every engine inside its major until Slice 4 (R-448) — removal tracked as R-469. - Goldens on a cadence, not per release. golden_currency_gate.py reads a dated waiver (documentation/tests/golden-waiver.yml, <= 14 days, row-bound): valid + BEHIND -> loud advisory, exit 0; expired -> red again naming the date; UNRECORDED (R-385) never covered; malformed -> 2, never 0. Tests cases 5-15 incl. the R-421 decoy; red-proof old-vs-new on the real behind tree. R-242's vouch half stays open. Cadence in RUNBOOK-manual-build.md §4.2 + the checklist. - Golden 0.236.0 baked, round-tripped, vouched, floor raised 0.232.0 -> 0.236.0 (documentation/tests/golden-0.236.0-2026-09-13/) — the last per-release bake; the waiver was issued AFTER it landed. No --no-verify anywhere in this session. Rows: R-459 CLOSED, R-467 CLOSED, R-242 narrowed; R-468/R-469/R-470/R-471 opened. 09 §3 gains decisions 5 and 6; STATUS items 11 and 12 closed; CONTEXT records the cadence ruling. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
574f5df107 |
the decoy sweep: 29 gates read, 16 fooled, 10 fixed - and a gate that refuses the next one (R-421)
gates / gates (push) Failing after 17s
THE CLASS, now a row: an instrument that matches a LABEL rather than the fact it names. Five instances - R-410, R-400, R-378, R-419, R-94 - and EVERY ONE was found by accident, by someone looking at something else. The gates enforce every other rule in this project, including the rule that findings must be written down rather than left in prose. Nothing had ever checked the gates. METHOD, and it is the transferable part: for each gate, construct the label WITHOUT the fact - a directory with the right name and no bake log, a handler case that exists only in a comment, a note whose prose mentions the marker it lacks - run the gate, record what it says. No verdict was reached by reading. Reading is how all five hid. RESULT: 29 distinct scripts (35 registrations; three are shared across three runners). 19 sound, 4 holes left OPEN with rows, 6 that no plausible decoy could be built for and are named UNTESTED rather than called sound. A gate nobody tried to fool is UNKNOWN. SCOPE IS A FACT TOO - the largest single cause, and mundane. Eight gates decided what to look at with os.listdir, one level. Every one was green AND CORRECT today, and every one would have gone blind the moment anyone added a subdirectory. mojibake and docker-v already used os.walk, caught the identical planted file, and are the control that proves the cause was the listing and not the decoy. IN THIS REPO: hub-confirm and manifest-bearer now walk. observations_gate (R-419, CLOSED) requires a marker at a line start or after a sentence boundary and strips inline code spans - a note SAYING it carries no marker no longer satisfies the marker test. closed-register now CONVICTS on a row it cannot parse instead of warning: FOUR rows were in that state, TWO of them written by the session that closed them the day before, and every one was exempt from the only check that reads that file. The rows were repaired first and the conviction added second - registering a failing gate refuses every push. THE META-GATE: decoy_coverage_gate.py refuses a gate registered without a decoy or a named exemption. It convicted ITSELF the moment it was registered, which is how it came to have one. Coverage is a DECLARATION the gate AST-parses, never a grep - searching a test file for a gate's name would be the very shape this sweep exists to find. The 20 uncovered gates are listed by name (R-426). NOT FIXED, each with a row and a decoy asserting TODAY's behaviour so the fix must be deliberate: R-422 reuse-refs (only 7 extensions; a rotted .md citation is invisible), R-423 site (PAGES is a hardcoded list of 7), R-424 one-register (a defect parked as `idea`), R-425 offbox-rename (fixed FILES list). R-427: closed_register_gate checks ONE direction - twelve open rows carry a closed verdict and were NOT moved, because telling finished from partly-finished is a judgement and R-378 is the record of a machine getting it wrong. FIVE DECOYS WITHDRAWN AS ILLEGITIMATE, mine, named in the audit. A decoy nobody would write proves nothing, and manufacturing a finding to fill a row is worse than an honest NO. No product code. No version bump. No image. No golden owed. All four runners green. Register: OPEN 172 -> 178, CLOSED 160 -> 161. |
||
|
|
1e6c387a0b |
R-404 CLOSED with the ruling; R-417 CLOSED by cause removal; R-418/419/420 filed
gates / gates (push) Successful in 17s
THE RULING WAS NEITHER OPTION AS FRAMED. Both offered answers - narrow the gate, or leave it and write waivers - argued about the gate, and the gate was never the problem. DIAGNOSIS, from live source: golden_currency_gate.py never looks at the push. It compares the controller's newest CHANGELOG heading against this repo's bake evidence and returns the same verdict whatever you are pushing - correct for a standing invariant, wrong as a push gate. And controller_gates.py had NO golden-currency entry at all. So the repo where a release happens never checked, and the repo that cannot create the debt was refused on every push. 18 of the last 24 pushes here touched no code - measured, and the new classifier agrees EXACTLY - most of them by construction, because the controller's code is in one repo and its register lives in this one. SIX of those 18 were bake records, so the push that PAYS the debt is itself documents-only: the gate was blocking its own cure. Not the waiver its docstring prescribes: that clause was written for a release nobody wants a golden for. R-417 was a release we DID want a golden for, on a night the runbook forbade baking. A waiver would have recorded a lie. RULING: block the push that can create the debt, notify the push that cannot. The gate's logic, exit codes and wording are BYTE-IDENTICAL. Only the consequence changed, for one gate, on one kind of push, with a loud ADVISORY block so nothing goes quiet. R-242's vouch half is amended in place to say it is UNTOUCHED and still open - a baked-but-unvouched golden still passes both the gate and the new notice. Do not read R-404's closure as closing it. FILED: R-418 - this runner's docstring listed ELEVEN gates while THIRTEEN were registered; one-register and closed-register ran undocumented since 2026-08-24. Enumeration fixed here, the correspondence is still unenforced. R-419 - observations_gate.py accepts an item whose body merely CONTAINS "NOT-A-FINDING", even in prose disclaiming it; found by accident when a planted test observation passed and my live validation proved nothing. R-420 - controller_gates.py could not express a non-blocking gate at all before today. Register: OPEN 171 -> 172, CLOSED 158 -> 160. |
||
|
|
6e550aedd3 |
R-87 put back in the register; closed_register_gate.py is the 12th gate (R-405, R-406)
gates / gates (push) Failing after 17s
Records and process only. No machine contacted. No product code, no version bump, no build, no deploy. R-87 was moved into CLOSED-ITEMS.md by the 2026-08-22 compression sweep |
||
|
|
c30430c530 |
skills: five process-domain skills + check_skills.py
gates / gates (push) Failing after 15s
The four existing skills cover the product; nothing covered how work is reported. Two rules this project has paid for — check the artifact rather than the report, and do not state a claim more firmly than the evidence allows — lived only in the operator's head and in chat, where Claude Code never read them. - felhom-evidence five confidence tiers, artifact-over-report - felhom-diagnosis no hypothesis until a command has been seen red - felhom-plain-language ASD-STE100, two options, the re-pitch - felhom-handoff the note goes to a FILE, not the conversation - felhom-doc-authoring the pointer decides whether material is reached scripts/check_skills.py asserts what decides whether a skill is EVER reached: frontmatter parses, name == directory, description and body non-empty, under 150 lines, installed copy still samefile()s into the repo. install_skills.py globs and never reads the file, so a missing description installs perfectly and then silently never loads. It convicted on its first run: felhom-build-deploy is 179 lines. NOT trimmed here (pre-existing skills are out of scope, and trimming a deploy skill without exercising its commands is how a wrong command reaches a live host) — a named single-entry GRANDFATHERED exception, WARNed every run, R-394. A new skill over the limit is convicted. Red-proof run and seen failing: description removed from felhom-evidence -> exit 1, "frontmatter field 'description' is missing or empty". Restored, tree clean. skills/SOURCES.md records both MIT upstreams, that these are adaptations not copies, and the six pieces deliberately EXCLUDED with reasons. Register: R-392 (no architecture doc covers the two-AI workflow), R-393 (decision-log skill deferred, with the reason), R-394. |
||
|
|
ef6ac6fe74 |
One register, enforced by a gate; closed work compressed into siblings (R-376..R-378)
gates / gates (push) Successful in 16s
Records and process only. No machine contacted. ONE REGISTER (operator ruling). 17 roadmap rows moved into OPEN-ITEMS.md keeping their identifiers, evidence and original filing dates - the oldest R-10, filed 2026-07-15, 38 days. 15 ideas stay in ROADMAP.md, which is their home; the gate exempts them by their own state word. 59 already-closed rows stay as history. Sorting rule recorded in the roadmap header: does the item assert something about the shipped product a reader could check and find false? scripts/one_register_gate.py, wired as the 11th gate. Control run: baseline passes, a planted open roadmap-only row is convicted by name, removing it passes with the file byte-identical, and a planted `idea` row is correctly exempt. Its four residual holes are in its docstring. The gate earned its keep immediately: it caught R-103, a READY finding my hand-sort mis-read as done because my regex matched the whole row where the body contains "shipped" - the gate matches the state cell. It also caught R-203 and R-163, recorded closed in the register and still open in the roadmap; the roadmap copies are marked SUPERSEDED with the register's verdict. HOUSEKEEPING. OPEN-ITEMS 672,376 -> 327,109 bytes (-51%); ROADMAP 239,306 -> 78,110 (-67%). Closed work compressed to 17% into CLOSED-ITEMS.md and ROADMAP-HISTORY.md; every entry names the commit whose git show returns the full original text. Rule-sentences are kept verbatim under "Reasoning kept" rather than judged entry by entry - 25 carry one. CONTEXT.md deliberately NOT compressed and the disagreement is argued in the report: 86% of it is standing rulings still in force, this prompt's own 3.4 says the log is never edited, and it has no per-ruling delimiter. Filed as R-377 - the problem is navigational, not volumetric. The hot/bulk placement decision was NEVER recorded as a decision anywhere - established, not assumed. Now marked [DESIGN] with a pointer honest about having no original date, given a decision-log entry that records what was rejected, and the [DESIGN]/[FACT] legend carried from 1 of 8 architecture documents to 8 of 8. Existing statements deliberately left unmarked (R-376). PROMPT-TEMPLATE gains N.7: compress what you closed, rehome live reasoning before it goes, state the register's size before and after. Ceiling R-375 -> R-378. |
||
|
|
ca543b8f69 |
where-felhom-stands: bring the picture up to 2026-08-22, and stop the page disagreeing with its source
gates / gates (push) Successful in 15s
Eight claims re-checked against the drill and the v0.218.0 fixes; three moved, all downward.
backup.offsite walked -> partial. "18 snapshots, daily, unbroken" was true on
2026-08-09 and false by 2026-08-21: the next snapshot after that date
was put there by hand, twelve days later. The rebuild lost the target
and the per-app switches came back off, so a run reported "backup OK:
0 app(s) backed up".
fail.wiped-reinstalled.data walked -> partial. A real reinstall orphans BOTH off-premises tiers:
restic silently for 12 days (R-193), and the PBS archives from before
the reinstall cannot be opened by the rebuilt box at all (R-366).
backup.fill-warning walked -> partial. The warning fires correctly, but the watcher runs
once a day, so a filesystem that fills at 03:31 goes unannounced for
~24 h. Watched silent while a volume sat at 99%.
Five re-checked and held: backup.tier1 and recover.byte-identical carry the R-355/R-354 story and
their fixes; backup.restore-proof stays grey for a sharper reason (orphaned archives, not an
untested tier); backup.sikeres gains two fresh instances; fail.customer-self-restore records that
R-356 now blocks 40 of 53 apps regardless of who is driving.
render_stands.py: the header's commit shas were hardcoded, so the page cited the August 9th commits
while the YAML said otherwise - the stale-build-product failure the renderer exists to prevent. They
are parsed now. The count beside them said "15 status(es) moved in that pass" when 15 was every
recorded move ever; it now separates the two numbers.
check_stands passes, and was itself proven able to convict first: a claim marked `missing` flipped to
`walked` in a scratch copy fired rule 5 by name (use.dlna), and the real file still passes.
|
||
|
|
104ef34f57 |
gates: the today-override announces itself; malformed no longer swallowed
gates / gates (push) Successful in 15s
Part 6 of the hub-blindness task, separable and done rather than dropped. Both due_checks_gate.py and instructions_gate.py read FELHOM_GATE_TODAY so their suites can control "today", and neither said so. A shell that still has it exported -- exactly what a session doing gate-test work leaves behind -- made both gates evaluate against a fabricated date and pass in SILENCE. That is this project's own named failure class: an instrument that can quietly return the wrong answer is not a measurement. The seam is legitimate and stays; the silence was the defect. Both now print a loud line naming the variable, its value, and that the real date is being ignored, before any verdict. And instructions_gate.py no longer swallows a MALFORMED override: it used to fall through to the real date without a word while due_checks_gate.py already exited 2 on the same input -- one variable, two gates, disagreeing about what a mistake means. Both exit 2 now. Tests extended in both suites (42 and 73 assertions, green). Red-proof: the announcement was deleted from due_checks_gate.py and its two assertions were seen failing, then reverted. Also adds REPORT-hub-blindness.md (topic-suffixed; the shared REPORT.md is left alone per the parallel-session rule). |
||
|
|
0a5e9b14dc |
due-checks gate (R-341), floor raise recorded (R-343), snapshot coverage (R-342)
gates / gates (push) Successful in 14s
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as prose in a register row; nothing read those dates and nothing would have objected when they passed. The dates now live in a DUE-CHECKS block INSIDE OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the pre-push hook and CI. exit 0 nothing due (prints pending count + nearest date; empty block too) exit 1 a row is due/overdue (due <= today, UTC -- due TODAY counts), or a row names an item with no R-row exit 2 block absent/duplicated/unparseable -- INCONCLUSIVE, never 0 It REFUSES rather than warns, and its docstring states the limitation: it is NOT a scheduler, it fires on the next push, not on the date. 37 tests. BOTH red-proofs run and reverted -- and the first one earned its keep by catching a hollow assertion of MINE rather than confirming the gate: flipping <= to < left a due-today row in neither bucket, min() raised on an empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the boundary was wrong. An exit code cannot tell a verdict from a crash. The test now asserts the conviction banner and the absence of a traceback, and the gate returns 2 rather than crashing if that partition breaks again. PART 3 — the floor raise, and the premise was WRONG. Read back from the store (not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero per-customer overrides, no "managed floor HELD" line. But read 5 shows the raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save, exactly the immediate action publish-train rule 2 documents. No error events followed; it restarted clean. R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five reads clean and no directive served. It went well, but a record calling it inert when it moved a customer box is what misleads the next reader. The row also states why the floor was behind -- rule 2 policy, not drift, earned by the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor (store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267, fails open at :78-82) rather than asserting them. Two boxes are below the floor and neither reports: drill-r50 (blocked, powered off) and peti-felhom (host row deleted). peti-felhom was NOT contacted -- its row records that a report from a deleted host 401s and is not persisted, so the raise cannot reach it. PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a separate Volume that snapshots exclude, so a rollback restores software state and NOT the datastore. Fine for that upgrade; the safeguard for any future procedure that could touch the datastore does not exist and is Viktor's call. Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling 200). Capability map deliberately unchanged; no row cites a floor or golden version. repo_gates.py fully green, 10/10. |
||
|
|
fc737b0fc0 |
installer v1.28.0: the removal genuinely reverses the installation (R-316)
gates / gates (push) Successful in 13s
v1.27.0's fix worked exactly once per machine. Measured on drill-r50 from virgin, on the PUBLISHED v1.27.0, before anything was changed: cycle 1 recorded 'no' and freed :53; cycle 2 recorded 'yes' and left dnsmasq running on 0.0.0.0:53; cycle 3 refused, exit 1. Every box already in the field is at cycle 2, and a reinstall onto a machine that has had Felhom is cycle 2 by definition. Why cycle 2 says yes: the preflight's ownership question is dpkg-query package presence and nothing else - not the absence of a record. Stopping the unit and leaving the package made our own package read as the household's one cycle later. Now the uninstall removes the package when the record says we installed it. Order unchanged and load-bearing: read the record, act, then delete the state file that holds it. TWO packages are recorded, because dnsmasq ships the unit and dnsmasq-base ships /usr/sbin/dnsmasq, and each is taken back only if we added it. The dependency check is a SIMULATION, not a guess: apt-get -s purge is asked what it would remove and the purge proceeds only if that set is a subset of ours; otherwise stop+disable, naming the package that blocked it. Never interactive, never fatal, and the success is re-queried rather than read off an exit code. Watched: three fixed cycles -> install 3 PASSES; a household resolver untouched; a dependent package not purged and named; no record -> untouched with the command named. Red-proofs with the mutation asserted applied: remove the purge -> cycle 3 refuses in those exact words; remove the ownership check -> a household resolver is purged; infer ownership -> the guess is taken. Also: R-317 (the agent stats a path dnsmasq-base owns to decide whether to install dnsmasq - pre-existing, now reachable), R-318 (no honest ownership marker exists for existing boxes; the preflight message is the mechanism), and the status page's decisions section rewritten to say what each decision costs and what doing nothing selects. |
||
|
|
1d5f2b8bb6 |
DRILL: the retained key works, and the customer cannot reach it
gates / gates (push) Successful in 23s
Three verdicts, kept separate because collapsing them is how this assumption survived a week. (a) The material IS retained. host_escrow_superseded id 11 is the first retained row in fleet history to carry identity_blob (572 B), byte-identical to the pre-supersession row (sha256 a10032341c8584ed...). (b) The retained material DOES open the old store. Unsealed with the old recovery code it yielded a password byte-identical to the pre-change one, and restored three planted files byte-identical from a store the box itself could no longer open - including a Hungarian accented filename verified as raw bytes. Negative control ran first and failed closed. (c) The customer has NO route, and is misinformed. ListSupersededEscrow has zero production callers; the recovery path selects FROM host_escrow. Asked with the code that had just worked by hand, the product answered "the recovery code did not open the sealed bundle". A valid code for retained history is reported as a bad code - the R-224 class again. R-304, rank 1. Both installer faults were watched happening first, so installer-v1.27.0 is now published (tag + both webpage.yaml refs). Pre-fix: the box came up on controller 0.98.3 against a vouched 0.213.0, below the floor and below the version carrying the recovery screen; and our own uninstall left dnsmasq on 0.0.0.0:53 so our own next install refused. R-297 and R-300 CLOSED. Also filed R-305 (the dnsmasq fix fires once per machine - the leftover returns on the second reinstall, proven), R-306 (--preflight-only writes state it says it does not), R-307 (a live abandon countdown on demo-felhom, firing 2026-08-24 - operator decision), R-308 (stored controller password stale), R-309 (the day-0 runbook's publication claim has been false since R-110), R-310 (two edges). Ceiling R-303 -> R-310. Capability map moved: the retention claim is now marked operator-only. Phase A logs did not survive the intermediate revert; recorded. |
||
|
|
125aec1be2 |
R-300: uninstall no longer leaves dnsmasq blocking the next install
gates / gates (push) Successful in 18s
Removing the snippet and restarting left dnsmasq enabled and unconstrained on 0.0.0.0:53, so the next byo install's preflight refused and the customer went debugging a home network that was never at fault. Ownership is recorded at preflight (the only moment it is a fact - the package is installed by the agent, not this script) and honoured at removal. Boxes already in the field carry no record and fail safe to restart-only, with the reason and the command logged; the preflight message covers them instead. Not observed live - no installer-v1.27.0 tag is cut. Files R-299..R-301. |
||
|
|
eb600872f2 |
R-297: installer compares a local golden against the manifest before using it
Step 7 short-circuited on any local golden archive with no version compare, no digest and no warning, so the manifest sha256 was consulted only on the fetch path. Local discovery is newest-by-filename: correct by recency, never by verification. A box could reinstall from a stale archive and come back below the version where the offsite recovery screen exists. Digest first, then the baked controller tag. An auto-discovered mismatch re-fetches the vouched golden; an operator-named mismatch refuses. An unreadable manifest refuses rather than passing. Not published: installer-v1.26.0 is deliberately not cut until a fresh install has been observed taking a stale local golden on drill-r50. Also files R-295..R-298. |
||
|
|
3ca9a7bbe6 |
R-242: build the gate the rule described — a release without a golden now fails the push
gates / gates (push) Failing after 13s
R-242 was filed 2026-08-07 as a mechanism-less rule and RECURRED WITHIN A DAY: controller v0.206.0 shipped the R-241 fixes while the vouched golden still carried 0.205.0, so a machine installed this morning would have received neither. Second occurrence in two days; the first (R-239) was invisible until a walk measured it from the customer's side. SHOWN FAILING FIRST, against today's state, before anything was baked - that is the gate's red-proof and the whole point of building it before the bake: newest released controller : 0.206.0 newest golden baked : 0.205.0 GOLDEN CURRENCY GATE FAILED ... A machine installed right now would receive v0.205.0 - the release is written, tested and pushed, and NOT delivered. Entry point exits 1; summary reports CONVICTED: golden-currency. *** THIS PUSH USED --no-verify, to push past the gate's OWN conviction. *** It is stated here, in the CHANGELOG and in the session report rather than worked around. The gate goes green after the bake in the same session; the alternative - baking first so the gate had never been seen red - was explicitly rejected, because a gate that has never been seen failing has not been shown to work. IT IS --fast, AND THAT FORCED THE DESIGN. Both the pre-push hook and CI run repo_gates.py --fast, which by contract selects only gates touching no network. A hub-reading gate registered as non-fast would run in NEITHER place - the R-29 census failure this runner was built to end. SO IT CHECKS THE BAKE, NOT THE VOUCH. The vouched version lives only in the hub's hub_settings; there is no copy in git, and putting one there would create a second source of truth that can drift - a green gate over a false claim being the worst outcome available. A bake without a vouch still passes. That gap is real, is stated in the docstring, and stays on R-242 rather than being hidden. The recurrence this gate exists for was a missing BAKE. IT COMPARES VERSIONS, NOT BEHAVIOUR, so a release that changed nothing customer-visible also trips it. Accepted deliberately: judging "customer-visible" by hand is what failed twice, and the cost of a false trip is one bake. A waiver belongs in the register, never in a habit of bypassing. Inconclusive (exit 2) on an absent controller clone or an unparseable header: not knowing is never a pass. |
||
|
|
15fa5273ba |
gate: check 7 (register citations) + content WARNings on the memory index
gates / gates (push) Successful in 8s
Check 7 catches "cites a register item and calls it open when it is not" -- the R-168 class, four files, one self-contradicting. Trigger is an openness CLAIM, not any citation: policing every mention would fire on ~30 legitimate provenance citations and the gate would be switched off. Deliberate deviation from the task's literal wording, to keep it alive. Two bugs found by the check's own red-proofs, both of which would have shipped: - the state marker is not self-closing (**SHIPPED - text**), so the first parser read R-168 itself as OPEN -- a gate that cannot convict its founding case is decoration; - the CLOSED exemption was line-wide, so "shipped" in a title pardoned "OPEN R-25b". Check 6 gains WARN-only content classes on MEMORY.md. Link targets are stripped first: the earlier scan reported three expired statements, all three false (dates in filenames), while missing the one real expired claim, whose deadline was written ~08-02 with no ISO date. 39 -> 60 assertions. All four runners green. |
||
|
|
f65ea89a24 |
workspace: version the root CLAUDE.md + InstructionsLoaded hook, and report which rules fire (R-229)
gates / gates (push) Successful in 8s
install_workspace.py lays down the two things that shaped every session while existing on one host only. Unlike install_skills.py the targets are LIVE CONFIG, so: timestamped backup before every write, settings.json MERGED (this script owns exactly one key), a diverged CLAUDE.md reported rather than silently resolved, and an unparseable settings.json refused outright. Proven: all 7 top-level settings keys survived byte-identically, and run 2 wrote nothing. rules_report.py surfaces the column that matters -- rules that have NEVER fired, which are mis-globbed or dead. 6 of 9 on first run. The hook now self-rotates at 5 MB. The memory store is BACKED UP, NOT COMMITTED (auto-written, may name hosts/paths): added to dooplex-backup.service's User Data component. /opt/backup/scripts/ is itself unversioned host state -- filed, not fixed here. |
||
|
|
f27aed87cd |
gate: instructions_gate check 6 — the auto-memory index (R-229)
gates / gates (push) Successful in 15s
MEMORY.md is the larger half of what loads before a word is typed (8.4k tokens vs the root CLAUDE.md's 6.6k) and is the one instruction file nobody hand-edits, so nothing was watching it. Three deliberately different outcomes, each pinned by a test: over-ceiling FAILS (auto-memory drops content past the limit with no error), an orphan WARNS (the store is outside git), and an absent store PASSES while PRINTING its reason -- asserted on the reason text, because a pass with no reason is indistinguishable from a gate that stopped running. 39 assertions (was 20). Red-proof run against the real store, not a fixture. |
||
|
|
3a9dd81e18 |
docs+gate: felhom.eu/CLAUDE.md becomes core + path-scoped rules; instructions gate registered (R-229)
gates / gates (push) Successful in 8s
227 -> 115 effective lines, split into .claude/rules/{hub,website,manifests,docs}.md, and
repo_gates.py gains gate 6. Trim first, register second: a registered-but-failing gate refuses
every push through the pre-push hook, which is why this repo -- the one that OWNS the gate --
was the only one not running it.
Register discipline and the R-110 installer fence deliberately stayed in the core; both have
triggers no fixed glob covers, and scoping them would have rebuilt the failure class they exist
to prevent.
Scoping proven from the InstructionsLoaded hook log in two fresh sessions, not from frontmatter.
|
||
|
|
c21bcf84f7 |
docs+gate: instruction files cannot silently regrow (R-229)
gates / gates (push) Successful in 7s
New shared scripts/instructions_gate.py, registered in controller_gates.py and agent_gates.py, never copied into a sibling repo (the reuse_refs_check.py precedent). 20 fixture tests, all asserting the effect: exit code AND that the message names the file and the reason. It is a consistency gate, not a budget gate, and the failure message says so. A /context reading measured the instruction files at 15k tokens against 869k free in a 1M window -- space is not the constraint, and a future reader must not re-derive the wrong reason. The 200-line ceiling is adherence guidance; a file nobody can hold in their head is where contradictions hide, and five were found here. Checks run against effective text (HTML comments stripped, because they are stripped before injection): the line ceiling; every .claude/rules/*.md declares paths: or an explicit unconditional: true; no component version literal; no TEMPORARY block carrying a past date; and the workspace-root CLAUDE.md is byte-identical to its versioned copy -- the live file sits outside any git repo, so that copy is its only version-controlled record. Two traps recorded so they are not reintroduced: a bare \d+\.\d+\.\d+ matches the first three octets of every IPv4 (the gate excludes dotted quads, or it fails on 192.168.0.180 in the agent's own file); and unconditional: true is NOT a Claude Code feature but this project's own marker. Workspace-root CLAUDE.md 208 -> 182 lines (142 effective), copy kept identical. The nine-instance invariant table moved into the felhom-testing skill, which triggers when writing or reviewing a test; all three directive bullets stayed in the core. felhom.eu/CLAUDE.md got surgical corrections only and is knowingly still over the ceiling at 227 effective lines -- closing it needs the restructure R-229 defers, said plainly rather than quietly absorbed. CONTEXT.md gains standing ruling S-35. OPEN-ITEMS.md gains R-229. Docs only -- no Go, no version bump, nothing built or deployed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JJc8sAGRWmavP3rMtdpkr2 |
||
|
|
51871a7ea6 |
installer 1.25.0: the off-site tier stops asking to prune (R-191)
gates / gates (push) Successful in 8s
Every weekly off-site run uploaded successfully and then failed the job on a prune the box's token is deliberately refused — R-89 moved off-site pruning server-side to ep0 and box tokens stay write-only. The 2026-07-26 'two weeks' ruling was not reversed; where it is enforced moved, and keep_last: 2 did not follow. Now 0, which the agent's existing guard already reads as 'never prune from the box'. Verified read-only on ep0 before changing it: both namespaces have a prune job at 03:30 keep-last 2 that has run every day since 2026-07-27 — 18 tasks, all OK, the newest keeping exactly two. Without that check this would have traded a weekly false alarm for unbounded growth. A gate asserts the offsite tier carries no client-side prune. The local tier is untouched. |
||
|
|
688470c945 |
installer 1.24.0: a PRE-EXISTING backup target is granted too (R-185)
gates / gates (push) Successful in 7s
configure_backup_target has two arms and only one granted. Case A creates the
storage and grants in the same breath; the Scenario-F arm ('the target already
exists') returned without granting. A box whose felhom-backup pre-dated the
install therefore pointed local_backup_target at a storage its own token could
not read — measured on BOTH demo boxes: {"data":[]} through the token while root
lists three archives. That tier was never restore-tested and nothing said so,
because an empty listing is also what a brand-new tier returns.
The reuse arm now ensures the ACL through the same guarded wrapper. Scenario F is
unviolated: the storage DEFINITION is untouched, and pveum acl modify is
idempotent. BACKUP_TARGET_ID is deliberately NOT added to PVE_STORAGES — that
list is granted a step before the target is resolved, and --acl-storages entries
are preflight-checked for existence; the comment now says so.
A gate asserts it: every arm that resolves the target must also grant on it.
Red-proved by reverting the arm.
|
||
|
|
bee6848458 |
installer v1.23.0 — publishing becomes an act, not a side-effect (R-110, R-183)
gates / gates (push) Successful in 8s
Two channels moved off main in the same change, because either one left behind makes the other cosmetic. Channel 1 — the served script. webpage.yaml git-synced /scripts/ from --branch=main every 30s and nginx served that tree, so pushing this file WAS publishing it: within half a minute it was what every new machine downloaded and ran as root, with no staging and no rollback but another push. The sync is now SPLIT: the website keeps tracking main at the same cadence (a copy edit must never need a release) and /scripts/ tracks the tag installer-v<SCRIPT_VERSION>. PROVEN before the manifest was touched: git-sync v4.4.0 follows a tag AND notices a MOVED one — measured on a throwaway sync against this repo, "update required ... local:<old> remote:<new>" -> "updated successfully", within one period. The moved-tag half is what the publish model rests on. Channel 2 — the sixteen files fetched at run time. fetch_raw pulled from $AGENT_REPO/raw/branch/main; it now pulls raw/tag/v$ART_AGENT_VER. That is a correctness fix, not only a channel one (R-183): a fresh install fetched the vouched agent BINARY while taking its unit file, sudoers and guarded wrappers from whatever main held. Two refs, one install, nothing compared them. Their correct ref was never SCRIPT_VERSION — they do not live in this repo. No fallback to a branch: a vouched version whose tag is missing fails loudly rather than quietly serving main. Channel 3 — the URL — needed no change, recorded rather than left silent: https://felhom.eu/scripts/felhom-host-install.sh never carried a ref, so both producers follow the tag with no edit. No hub change, no hub version bump. Gate 6 in hostinstall_gates.py pins all three structurally with no network, so it stays in --fast and runs in CI. It deliberately does NOT assert "a tag exists for the current SCRIPT_VERSION": that would go red on the very push that bumps the version, before publishing — and publishing being separate is the ruling. |
||
|
|
aa62449694 |
R-178 CLOSED: both demo boxes reinstalled from the merged golden and proven
gates / gates (push) Successful in 8s
Two boxes, two DIFFERENT supply paths, so the session proved the disk shape and the delivery route rather than one of them twice. demo-hp (layout proof, --golden <local volid>): mp0 at /var/lib/felhom, backup=1, 70G, no mp1; /var/lib/docker and /mnt/sys_drive both real mounts of its subdirectories via fstab; one df figure and one device id (64519) on all three paths; reboots 3/3 with the binds surviving each. demo-felhom (pipeline proof, --force-gitea-golden): fetch_verify succeeding against the vouched manifest for BOTH artifacts -- 'verified sha256 54e2a4c431daf580... matches the hub manifest' for the golden, a7763d31... for the agent. 250G single volume, grep -c '^mp1:' = 0, reboots 3/3. Journey proven on both, endpoint-level: claim -> deploy -> back up -> restore, with a planted marker returning byte-identical on each box. Ceiling measured gone: 65 GiB and 233 GiB available to a recovery unit, against 19 and 45. R-165 -> IMPLEMENTED, not PROVEN-LIVE, on the operator's ruling. B2, which that row records as the bulkhead's replacement, fired live for the first time and does refuse per app, delete nothing and alert -- but it is checked only in captureAllRecoveryUnits while runVolumeDumps writes the bulk unguarded, and its 'the previous unit is untouched' claim was measured false (182,272 B dump replaced by 2,147,666,432 B under a manifest still dated 06:34:26). -> R-181. New: R-179 (uninstall leaves NAS network-storage units), R-180 (--archive-storage not cross-checked against the ACL grant; 403 at step 8/8 after root@pam is rotated), R-181. Third instance of R-115 recorded (agent 0.120.0 unpublished). No code written, no version bumps -- this was a runbook. |
||
|
|
e3525e62ac |
host-install: one data volume, derived from the disk (R-165)
gates / gates (push) Successful in 7s
felhom-agent v0.120.0 merges the two data volumes into one, and step_grows computed two numbers while the install call passed both — so this had to change with the agent or every install would have provisioned a half-sized box. The 80/20 split is summed (226 = 184+42), so a standard appliance keeps exactly the 250 G it had, no longer split by a wall. The size still comes from the physical disk: step_grows already read the thin pool's free space, and the merge only collapsed its two outputs into one. --sysdata-grow is deprecated but still honoured, because the agent folds a hand-passed value in rather than dropping it. |
||
|
|
c718aad1bc |
docs: R-168 SHIPPED, R-29 CLOSED on the demonstrated alarm, R-169 minted
gates / gates (push) Successful in 7s
SPIKE-ci-runner-2026-08-02.md: all six probes with method, measurement and ruling; none STOPped. P2 (stock image has git but no python3) and P6 (a runner that loses its state re-registers and orphans the old record) changed the design; P5 (a failed run signals NOTHING) is why the alarm exists at all. R-168 SHIPPED with its evidence. R-29 CLOSED — on the demonstrated alarm and not on a green run, as required: the class it opened is answered at both ends, the hook refusing locally and CI catching a --no-verify bypass and emailing. R-161 noted: its automatic half now exists for the STATIC gate, while its original scope, the runtime gate, is deliberately still not automatic and should stay that way. NEW R-169 (grep established R-168 was the highest in use): CI can only report, because there is no gate in the road. Making it blocking needs branch protection plus a PR workflow, both of which change how the operator works — so it is theirs to decide, and the row states the cost honestly rather than recommending it. CONTEXT gains S-8 (CI detects, does not block, and why that is structural), S-9 (a detector that tells no one is not finished, plus the curl and Cloudflare-1010 traps), S-10 (the runner is unprivileged because DooPlex is Tier 2), S-11 (CI reproduces the sibling layout). CLAUDE.md gains the rule earned by red-proofing: a go test -run pattern that matches no test prints ok and exits 0, and an instrument that can silently drop results is not a measurement. |
||
|
|
4707be755c |
docs: R-94 closed, R-29 leg (a) closed + leg (b) half, R-168 minted
hub/CHANGELOG v0.87.0 + scripts/CHANGELOG gate-enforcement entry. CONTEXT gains S-6 (the hub renders no host-install version and the gate pins its absence) and S-7 (gates run from one entry point per repo; reuse_refs_check was fixed rather than the REUSE.md convention, with both rejected alternatives recorded). OPEN-ITEMS: R-94 CLOSED all three legs, leg (a) by DELETION with its reason; R-29 leg (a) CLOSED and leg (b) HALF-SHIPPED with the census result written into the row (13 gates; every gate a CLAUDE.md names was green, two of the four unnamed were red); R-161 gains its successor pointer. NEW R-168 (grep established R-167 was the highest in use): Gitea Actions runner — measured 2026-08-02 as Gitea 1.26.2, Actions enabled on all four repos, 0 runners, 0 workflow runs, 0 branch protections, and the consequence that trunk-based direct-to-main pushes leave no merge for a status check to gate, so CI here can detect but not block. BLOCKED on a spike over host-mode vs privileged DinD on DooPlex and whether the workflow can avoid JavaScript actions. ROADMAP: R-94 collapsed to its one-liner, R-29 updated, R-168 added. |
||
|
|
f2fc76ec4b |
ISO v1.26.1 PUBLISHED — both entries proven, round trip verified
Live: https://iso.felhom.eu/felhom-installer-1.26.1-pve9.2-1.iso sha256 f3cc86d5f0ec68bba4155c994b4fa84e208d50209bb6e815636c99e5441059a6, 1705322496 bytes. PART 5 PASSED ON BOTH MENU ENTRIES, four observables each: Graphical spikegfx.felhom.eu pairing code J7N-2DA TerminalUI spikesix.felhom.eu pairing code ZY5-YY4 Both: manual install, own disk, own password, real completion signal, and the journal's 'not bound yet — polling every 30s ... normal waiting state, not an error'. Spike 4 had REASONED the graphical path follows from shared Install.pm; it is now measured. PART 6: G1-G10 + G13 all PASS against the uploaded file. G4's single hit is felhom-bootstrap.sh:480's substring TEST ('$envtext' != *FELHOM_RETRIEVAL_PASSPHRASE=*), not a value — my own regex matched the glob's asterisk. PART 7: uploaded via rclone in a container configured ENTIRELY by environment variables, so no credential file was ever written. Round trip verified from the public URL — not the local file. Bucket stays private: unauthenticated GET to the S3 endpoint 400, custom domain has no index (404). CORRECTED BEFORE UPLOAD: the generated manifest described a single automated entry with a 5s timeout and listed Graphical/Terminal UI as 'menu-removed'. Generator fixed, sidecar regenerated, and the ISO verified byte-identical before and after — the published file IS the file Part 5 validated. Hub-side cleared: appliances 16, 17, 18 discarded (303 each); zero rows remain. The endpoint is /appliances/<id>/discard, POST only (server.go:345) — not /delete. Teardown: VMs purged, spike5 storage removed, demo-hp back to 6.6G, drill-r50 and 9201 untouched. Still open and named: OPEN-ITEMS/ROADMAP dispositions for R-128/R-154/R-155 are not written; the .deb is not byte-reproducible (G7 sub-clause); before-network stub unreached; Secure Boot and real hardware not exercised. |
||
|
|
61e9b55737 |
SPIKE 4: a .deb in the ISO DOES deliver on an interactive install
Findings only — no script, profile or build file changed; no release ISO built, nothing published. documentation/audits/SPIKE-universal-iso-4-2026-07-31.md MEASURED, with a control, and the negative control is in the SAME box. One ISO (15 GRUB entries), a trivial probe .deb injected into /proxmox/packages/, two qm-created VMs on demo-hp (400 interactive / 401 automated control) on a scratch dir storage at the /mnt/nvme-1tb mount ROOT. Interactive (Terminal UI) install: - package installed (ii felhom-spike4-probe 0.0.1) - postinst RAN (marker + content intact) - it enabled a systemd unit, and that unit FIRED ON FIRST BOOT (uptime 7.98s, pid1=systemd) - while on the same machine proxmox-first-boot is NOT installed and /var/lib/proxmox-first-boot does not exist — Spike 3's negative reproduced, not assumed. Postinst environment (identical both paths): pid1=unconfigured.sh, NO running systemd, but 'systemctl enable' SUCCEEDS; /proc+/sys mounted; network+DNS happened to be up (inherited from the installer's DHCP — must NOT be relied on). Constraints: never systemctl start/daemon-reload, never require network, never fail, do the real work in the unit at first boot. Repack preserves it, but a naive 'xorriso -boot_image any replay' fails with 'Overlapping MBR partition entries' — iso-repack.sh:270-292 already documents that exact failure and its fix. R-153 RETRACTED into R-94 leg (b): OPEN-ITEMS.md:15 carries it verbatim at READY (XS), and R-29 says explicitly 'do not mint a new ID for a new instance'. Spike 3's further claim that the drift leaves the generator 'three minor versions stale' was FALSE and is corrected — R-94 retracts that exact reading; the served script is always main, so 1.22.0 is what every install already gets. No new R-rows opened. |
||
|
|
bb29186d62 |
SPIKE 3: [first-boot] does NOT fire on an interactive install
Findings only — no script, profile or build file changed; no release ISO built, nothing published. documentation/audits/SPIKE-universal-iso-3-2026-07-31.md MEASURED with a control from the SAME image (one ISO, 15 GRUB entries): - Automated entry -> hook fires: ttyS0 marker, marker file, /var/lib/proxmox-first-boot/proxmox-first-boot (0700), activation symlink, unit active. - Terminal-UI entry, normal manual install -> ALL absent, and the proxmox-first-boot PACKAGE is not installed at all. A whole-filesystem grep for the marker returns nothing. Mechanism cited: Config.pm:118 defaults first_boot.enabled=0 and set_first_boot_opt is never called in the Perl tree; Install.pm:746 returns early without it; Install.pm:1360 skips the package. proxinstall (graphical) has ZERO occurrences of first-boot. [first-boot] is an automated-installer feature, unavailable on every interactive path by construction. R-154. A delivery mechanism DOES exist and is UNTESTED: Install.pm:1343-1372 unpacks every .deb in the ISO's /proxmox/packages/ into the target on every path (fixed skip-list), then dpkg --configure -a runs postinsts (:1378) — how PVE ships first-boot itself. Read from source, not measured. Q5: the public image should carry NO answer.toml at all — that removes the baked root hash, the disk profile and the whole Spike 1-2 problem space, and makes it a one-line release gate. But iso-repack.sh:100-106 refuses an ISO without auto-installer-mode.toml. R-155. Incidental R-153: hub hostInstallVersion=1.19.0 vs SCRIPT_VERSION=1.22.0; hostinstall_gates.py detects it and exits 1 — the gate works, nothing runs it. Q3 (real stub at before-network) was NOT reached and is recorded as not reached. |
||
|
|
19c932a693 |
SPIKE 2 complete: locked root closes the PVE web UI; before-network gives a measured zero window
Findings only — no script, profile or build file changed; no ISO built, nothing published.
documentation/audits/SPIKE-universal-iso-2-2026-07-31.md
Both Tier 0 boxes went offline mid-session (provider cable fault; four routes tried, no Tier 2
fallback used) and returned. All three scenarios then ran to completion on real PVE, each signalled
by reboot-mode='power-off' rather than a disk hash.
- A LOCKED ROOT CLOSES THE PVE WEB INTERFACE. Measured at the exact endpoint the UI uses
(POST /api2/json/access/ticket, root@pam) WITH A WORKING CONTROL: known-password install returns
HTTP 200 + ticket; locked install returns 401 for every password and none can exist.
passwd -S root = L, shadow = literal-asterisk, PVE uses the stock PAM stack.
- GRUB recovery mode is also closed ('the root account is locked') — but init=/bin/bash still gives
an unauthenticated root@(none):/#. A locked box is recoverable, operator-only, at the console.
The installed GRUB has NO password, so locking root is not a physical-security measure. R-152.
- before-network MEASURED (A/B, same image): the hook RUNS (marker, uptime 6.58s) with entropy 256,
writable /etc, all binaries and openssl_rand_len=32, while ip_global is EMPTY and
listen_22_8006 = 0. fully-up is the converse: sshd+pveproxy active, 3 listening. Zero window.
- R-148: answer.toml.tmpl:27 justifies fully-up with a pvesh/pct dependency the stub does not have
(grep rc=1) — it blocked the ordering now measured as the fix.
- R-149 three ordering values; R-150 Condition-guarded hooks skip silently; R-151 demo-felhom built
from an uncommitted profile.
Three probes failed and are recorded as failed: a container probe that ran as uid 0, a GRUB probe
that missed the 1-second menu timeout, and a kernel-line edit one line off (caught by a pre-typing
verification screendump). The interim 'Layer 1 teardown INCOMPLETE' is corrected — the fixture had
never landed, because the staging mkdir was in the SSH call that timed out.
|
||
|
|
5bdd8372f8 |
SPIKE 2: before-network gives a zero window by construction; locked root closes sulogin
Findings only — no script, profile or build file changed; no ISO built, nothing published. documentation/audits/SPIKE-universal-iso-2-2026-07-31.md BOTH Tier 0 boxes went offline mid-session (remote site, 12:28 CEST; four routes tried, our tailscale pod healthy). Q1/Q2/Q3 each keep a part needing a nested VM: those are BLOCKED, not answered. DooPlex was NOT used as a fallback — Tier 2, and this task did not authorise it. Established without them: - STRUCTURAL: ordering='before-network' maps to proxmox-first-boot-network-pre.service (Before=network-pre.target, Type=oneshot) — it completes before ANY interface is configured, so a rotation there has a zero-length window BY CONSTRUCTION, not by being fast. - R-148: the stub does not need 'fully-up'. stub-first-boot.sh has no pvesh/pct/pveum/qm call (grep rc=1); that usage is in felhom-bootstrap.sh under its own After=network-online unit. answer.toml.tmpl:27 justifies the current ordering with a dependency that does not exist. - R-149: the ordering enum has THREE values (before-network, network-online, fully-up), not two. - MECHANISM (container, not PVE): locked root closes sulogin — 'the root account is locked' for both '*' and '!', with a working control. So 'discard' and 'lock' are the SAME outcome for recovery, making the escrow decision binary. - R-150: all four proxmox-first-boot-* units are Condition-guarded; a failed condition is a SKIP, so a hook that never ran looks identical to one that succeeded. - R-151: demo-felhom was installed from an UNCOMMITTED profile — a Tier 0 reference box is not reproducible from main. - Q4: four gates in iso-repack.sh enforce the single-entry menu; default/timeout already settable. The first mechanism probe was invalid (uid 0 bypassed pam_unix; sulogin had no tty) and a teardown error (shredding the control plaintext) are both recorded as failures, not massaged. demo-hp teardown is INCOMPLETE and named as such; the command is recorded, not claimed done. |
||
|
|
ea00976403 |
SPIKE: a universal ISO needs a different disk strategy and a locked root
Findings only — no script, profile or build file changed; no ISO built, nothing published. documentation/audits/SPIKE-universal-iso-2026-07-31.md - R-139 (HIGH): a disk filter matching >1 device does NOT fail safe. Observed in a nested VM — the installer silently picked one of two matching disks and wiped it; validate-answer accepts such an answer. The 'filter did not match any devices' guard covers the ZERO-match case only. - No udev property distinguishes an internal system disk from external media. Measured on demo-felhom with its 1TB external attached: ID_BUS='ata' for BOTH, lsblk RM=0 for both, and device-info exposes no removability property. demo-hp's NVMe carries no ID_BUS/ID_TYPE at all. - R-141 (HIGH): the answer schema makes a root credential mandatory, but root-password-hashed='*' validates AND installs to completion. [first-boot].ordering accepts 'before-network', the only ordering that closes the exposure window structurally. - Q3: prepare-iso leaves grub.cfg byte-identical to stock (15 entries, automated AND interactive) — a two-entry menu is purely a Felhom grub.cfg.tmpl change. - R-129 resolved: demo-hp's key is the operator's own, added post-install; demo-felhom's IS baked by an uncommitted profile. The reachable-before-rotation measurement FAILED twice and is recorded as failed, not inferred. Opens R-139..R-147; restates R-128. |
||
|
|
f6aed82940 |
host-install v1.22.0 — E-2 Part 2: new boxes get a real backup target, or are told they do not
Every box installed before this got local_backup_target "local" -- the vzdump
target on the SAME physical device as the guest, so a drive failure took the
guest and its only local backup together. E-1 fixed two machines by hand; this
fixes the installer.
Case A: an eligible secondary drive is already mounted -> create felhom-backup on
that drive's own mountpoint via the felhom-backup-target-apply wrapper (create +
grant) and point the primary tier at it.
Case B: system drive only -> the target stays on the system drive and this is
RECORDED AS DEGRADED, not as normal. The install still succeeds: a single-drive
appliance is a valid product, it just cannot survive drive loss.
Phase 0 inverts the emphasis: the installer has NO drive-enrollment step, so on a
fresh appliance Case A almost never fires. The common case is Case B with the
drive arriving later through the wizard (Part 3). Case A covers the reinstall
shape where an agent-generated .mount unit already brings the drive up by fs-UUID.
Eligibility suggests and refuses the absurd, never decides by transport: the
reference backup drive is an external USB HDD and BOTH demo boxes report
removable=0, so a transport rule disqualifies the reference drive and a removable
rule finds no candidate at all.
Scenario F: an already-configured box is never corrected -- an early return plus
setdefault, both load-bearing.
Proofs (installer-logic-tested against extracted functions with stubbed
pvesm/wrapper; NOT install-tested, no reinstall was performed):
A -> create + grant, resolved felhom-backup
B -> DEGRADED warnings, resolved local, rc=0 (install not failed)
F -> skipped, 0 wrapper calls
F red-proof (guard removed) -> 2 wrapper calls, i.e. it would have "corrected"
a correct box
|
||
|
|
b4c528801a |
host-install 1.21.0: F-LEAK — grant FelhomAgentGuest on the scratch VMID band
A failed restore-test's scratch guest never joins the felhom pool, so the pool-scoped grant cannot reach it and teardown 403s. Ten path-scoped /vms/<id> grants reach exactly the scratch band and nothing else. Removal path + verify step extended. |
||
|
|
adf1d1e619 |
R-82 Slice D/E: installer default 1.20.0 + architecture docs brought current
Slice D.1 — host-install 1.20.0: a FRESH box defaults to local-daily + offsite-weekly (felhom-pbs, 604800s, keep_last=2). setdefault semantics proven both ways: fresh gets the tier, an UPGRADE preserves the existing backup block verbatim — so an in-place upgrade can never silently start writing to an offsite datastore. Existing boxes are migrated explicitly. Slice E: - 07-backup-architecture.md: honest status header per CONTEXT ruling S-2, with an explicit STALE-outside-the-PBS-tier verdict (the controller tiers were last verified 41 controller versions ago). The PBS row claimed 'PBS on DooPlex' (the retired spike store) with no cadence; it now names felhom-pbs -> felhom-offsite on ep0 over wg-felhom, weekly, keep_last=2. NOT marked ratified — that is Viktor's review of the section 10 list. Discharges R-83. - 06-offsite-connectivity.md: the target-split remaining-work note collapsed (shipped), and records HOW S4.1's tier-aware timeout silently regressed — the mechanism was never removed, its INPUT changed when local_backup_target was retargeted to 'local'. Also notes S4.1 already diagnosed the teardown 403 as a phantom (a timeout consequence, not an ACL gap). - capability map: new row for recurring offsite backups actually LANDING, as distinct from the existing row proving ACTIVATION. IMPLEMENTED, not PROVEN-LIVE — the restore round-trip has not completed under the fixed code. - ROADMAP: R-82 SHIPPED with its remaining gate named, R-83 DISCHARGED, R-84 left open. - CONTEXT + REPORT: the arc, including the mid-arc correction I had to make. |
||
|
|
485321f694 |
R-50 Phase A: host-install v1.19.0 island default + hub version sync
- felhom-host-install v1.19.0: portless vmbr9 island bridge, appliance binds local_api on 169.254.253.1:8443, writes island_bridge/island_guest_addr, pins lan_resolver.host_ip to the LAN IP (Finding-1). --no-island opt-out. - hub hostInstallVersion 1.16.0 -> 1.19.0 (F-1 sync). hostinstall_gates PASS. - Pairs with agent v0.96.0 (attaches guest net1). byo unchanged. Coupling: island install requires agent >= 0.96.0 (vouch first). |