Files
felhom.eu/REPORT.md
T

11 KiB
Raw Blame History

REPORT — instruction files kept true; vaultwarden on the ladder; the burn-down continued (2026-10-06 afternoon)

Part Result
A — the instruction-file rule (R-891, R-469, sweep) done — rule in all four copies of unprompted-work.md §5 and PROMPT-TEMPLATE.md §9 rule 9; R-891 fixed and closed; R-469 re-read: NOT removable, narrowed (an operator question); sweep: 26 factual edits in four repos (list below). No permission prompt or refusal came up.
B — vaultwarden on the update ladder (R-890) done — the test-box admin seed built (catalog 6b4877d, red-proved); 1.36.0-alpine → 1.37.4-alpine proven on bench 9401 and on 9202; written by --write-ladder (catalog ecb8552); R-890 closed
C — the burn-down 3 closed (R-644, R-763, R-764), 2 stopped with the reason written (R-762 medium, R-717 needs a design)
Rows before Rows after Opened Closed
142 137 0 5 (R-890, R-891, R-644, R-763, R-764)

Baselines and rulings

Verified at the start: felhom.eu 7d0dffcf34, controller 36088fd82e (v0.300.0), agent cefdc731a4 (v0.149.0), catalog 65130c6c03; register 142. Read: every repo's CLAUDE.md and .claude/rules/, REPORT.md, R-890, R-891, R-469. The two rulings were recorded first as 09 §3 decisions 149 and 150, with the reviewer's error recorded in 150. Architecture read: 09-update-architecture.md §3 (decisions 16, 35, 42, 146), §6.4 part 4 (the ladder writer), §6.5 (the drill catalog).

Part A

The rule — .claude/rules/unprompted-work.md §5 „Instruction files", the operator's text verbatim plus the permission-check sentence; identical in felhom.eu, felhom-controller, app-catalog-felhom.eu and the workspace root's unversioned copy (md5 c1e6c881… for all four). documentation/PROMPT-TEMPLATE.md §9, rule 9. felhom-agent has no copy of the shared rule file (it never had one); adding one is a rule change, left for the operator.

R-469 — not met. R-463 closed (8 of 11 PostgreSQL apps converted by the box; zipline, adventurelog, immich stay on 16 by decision 42), but the engine gate now enforces decision 35: a PostgreSQL major passes only with a two-venue ladder entry carrying engine_conversion. Removing it would let an unproven major ship, and the image refuses to start on the old datadir. That is a loosened fence. Nothing is left for CC to build; the row asks the operator to close it (CC's pick) or keep it.

Every instruction-file edit (before → after, why). All factual; none loosens a rule.

felhom.eu

  1. CLAUDE.md „Gates — ONE entry point": a list of ten gates + „--fast … today that is all of them" → the GATES table is the list (nineteen gates); --fast skips the full-run-only gates and names them (today iso-bootstrap). Why: repo_gates.py imported: 19 gates, iso-bootstrap not fast. R-891.
  2. CLAUDE.md design pointer architecture/01..05-*.md → 01..11-*.md (00 = the capability map). Why: 06–11 exist.
  3. .claude/rules/docs.md „the locked design" 01..05 → 01..11. Same reason.
  4. .claude/rules/website.md: the list of what site_gates.py asserts gains „no embedded <style> blocks". Why: its gate 7 (site_gates.py docstring).
  5. .claude/rules/unprompted-work.md: §5 added; „Same wording lives in all three repos'" → names the three repos and the workspace root. Why: four copies exist.
  6. documentation/runbooks/workspace-CLAUDE.md (the workspace root CLAUDE.md links to it): standing rule 3 carried the R-286 sentence group twice in a row; the second copy removed.

felhom-agent 7. CLAUDE.md gates: „reuse_refs_check and instructions_gate … today that is all of them" → every gate in GATES (five: three shared, published, release-complete); --fast skips published. Why: agent_gates.py:52-60. 8. CLAUDE.md decoy paragraph: scripts/decoy_coverage_gate.py, documentation/audits/AUDIT-gate-decoys-… → with felhom.eu/. Why: neither exists in this repo. 9. .claude/rules/health-checks.md (comment): „The single source is now felhom.eu/CLAUDE.md 'Code quality rules'" → the copies are felhom.eu hub.md and the controller's gates.md. Why: that section holds no health-check rule.

felhom-controller 10. .claude/rules/gates.md: python3 controller/scripts/controller_gates.py (from controller/) → python3 scripts/controller_gates.py from controller/ (or the long path from the repo root). Why: the old path does not exist from controller/; the runner's usage line and CI use the new one. 11. .claude/rules/gates.md: the list of ten local gates → a pointer to GATES (fourteen local + three shared; the advisory golden-notice). Why: controller_gates.py:71-101. 12. .claude/rules/gates.md: the shared list gains observations_gate.py. Why: controller_gates.py:87. 13. .claude/rules/gates.md (two lines): decoy gate and audit paths with felhom.eu/. 14. CLAUDE.md Commands table: the same command fix as 10. 15. CLAUDE.md: „Protected stacks (traefik, cloudflared, felhom-controller)" → „…, and always samba". Why: internal/config/config.go alwaysProtectedStacks. 16. CLAUDE.md design pointer architecture/01/02/03-*.md → architecture/NN-*.md (01–03; 07, 09, 10). 17. .claude/rules/unprompted-work.md: as 5.

app-catalog-felhom.eu 18. CLAUDE.md: „each holding exactly docker-compose.yml + .felhom.yml" → without „exactly", plus steps/<StepKey>.yml for a superseded ladder step. Why: 20 templates have steps/. 19. „Only the two template files sync" → „Only the template files sync". 20. „60 checks in 10 groups" → 61. Why: NEW-APP-CHECKLIST.md row-count table. 21. „runs all four gates … scopes the two gates that accept scoping" → every gate in GATES (fourteen); scopes every gate that accepts an app scope (eight). Why: catalog_gates.py:88-131. 22. „CI was rejected for now … R-161 stays open at reduced scope" → CI exists and re-runs --fast; R-161 closed 2026-10-05. 23. „--fast … is gate 1 and the engine-major gate … the other two" → every gate except image-resolvable and volume-persistence. 24. The paragraph „There is no gate for this yet and that is a known gap" removed. Why: check-catalog-since.py is a registered gate (R-452, closed); the same section already names it. 25. „The four MariaDB services (…)" → the five (grimmory added), noting the gate judges by glob; „The eleven PostgreSQL services" → twelve, postgis and immich's image counted. Why: templates/*/docker-compose.yml. 26. Decoy gate and audit paths with felhom.eu/; .claude/rules/unprompted-work.md as 5.

Sweep notes not acted on: word-for-word repeated paragraphs (R-421 in three CLAUDE.md files; „Presence is not success" in the controller's backup-paths.md) — repetition, not a wrong fact.

Part B — R-890

Built (catalog 6b4877d): box_admin_seed_allowed(w, app) allows the admin seed on a box walk only when all five hold — not the bench venue; FELHOM_BOX_ADMIN_SEED=1; the walk targets demo-hp guest 9202 (never 9201, a demo box's household-shaped guest); THIS run installed the app (box_walk.DEPLOYED_THIS_RUN, set on the deploy's 202); the box's controller.yaml names the drill catalog. The invite runs inside the box: the token is read from the container's environment into a shell variable and handed to curl on stdin; the admin cookie is in a 0600 file that is shredded (vaultwarden marks it Secure, so it is sent by hand over the container's plain-http address); only HTTP codes come back. The Tester 1 box is allowed by the ruling, but the walk has no route to it (box_walk.py drives demo-hp guests only) — written in the code, not built. Tests BoxAdminSeedGuard (6); red-proof: a guard returning (True, "") fails three tests (test_each_condition_alone_refuses, test_a_household_shaped_box_is_never_seeded, and the bench's test_the_box_walk_never_signs_in_as_admin).

Proven:

  • Bench LXC 9401 (harness v5, 600 s memory watch): verdict proven; seed read back before and after; healthy in 30.9 s; anon peak 16.3 % (cgroup 22.3 %); 0 kills, 0 restarts; no file changed.
  • Box 9202 (drill catalog): self-registration 400; the in-box admin sign-in and invite → (200, session yes, 200); invited registration 200; seed read back; guarded Update backing-up → pulling → copying → verifying → done in 12.3 s; seed read back; badge „Frissítés elérhető — 80 napja" → „Naprakész"; removed through the product.
  • upgrade-test.py --write-ladder from both verdicts → vaultwarden/server:1.37.4-alpine, catalog_since 2026-10-06, the first ladder entry (catalog ecb8552); check-image-resolvable.py vaultwarden OK.
  • Who gets it: no box reports vaultwarden today (hub DB copy with its WAL: last vaultwarden telemetry 2026-09-18, demo-hp; control: the latest telemetry row of any app 13:51 today; the copy shredded). Read directly: demo-hp 9201 and demo-felhom 9201 run no vaultwarden container.

Part C

  • R-644 closed: no gokapi on 9202 (no container, no volume; control: the six running containers are listed).
  • R-763 closed (fix from 2026-10-05, proven now): wger installed fresh on 9202; its own process reads ALLOW_REGISTRATION False, ALLOW_GUEST_USERS False; straight at the app, a stranger's sign-up GET and a CSRF-valid POST both redirect to the features page (control: the same CSRF token passes on the login form, 200); three anonymous dashboard visits → login; users 1 before and after (2 → 4 before the fix).
  • R-764 closed (proven now): mail off → healthy, 0 restarts, console backend; mail on (the toggle's injection) → settings load, smtp backend, port 2526, TLS off. Not measured: a real reset mail (9202 has no app-mail).
  • R-762 stopped: still 404 on wger 2.7; the image has no whitenoise and runs Django's runserver, so it needs a second container for /static and /media — medium. Note added to the row.
  • R-717 stopped — needs a design: the controller's after_setup command form runs only when the lock is set; the household's 15-minute window cannot reopen a database switch. Note added to the row.

Said plainly

  • A merge conflict was committed into the DRILL catalog (vaultwarden's .felhom.yml, by a commit -am after a failed merge). Found at once by its conflict markers and fixed in the next drill commit; the drill was then reset to live. The live catalog was never touched by it.
  • My first wger user-count command had a quoting bug (it printed a Python NameError); the reads were repeated through a file. The stranger checks were not affected.

CI, last commit of every repo

catalog ecb8552 → 1433 success (and 6b4877d → 1430 success); agent 3e8ebeb → 1431 success; controller 3124278 → 1432 success; felhom.eu a29ee9dd → 1434 (checked before the next push), and this commit (checked after the push).

Teardown

Machine: 9202 back on the live catalog (controller.yaml restored byte-identical), vaultwarden and wger removed through the product (no container, no volume), the same six containers as at the start. Bench 9401: 0 containers, .env gone, stopped. Drill: reset to live (ecb8552 = ecb8552). Host: nothing else. Hub: nothing changed (one read-only DB copy, shredded). Scratch secrets shredded at the end.