c033b3b617af2e1c2df3d6aa3140331fd0725166
96 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3738dfc548 |
hub v0.116.0: every new customer starts WITH the off-site copy (operator ruling)
gates / gates (push) Successful in 19s
Off-site is ON by default for a new customer — shared, 100 GB soft quota prefilled, the checkbox kept so an operator can opt a customer out. The reason is this repo's own [FACT]: the whole-guest tiers do not carry the data drive and a Tier-1 unit has no file leg, so with this unticked a one-drive box keeps NO copy of the household's own files. Measured on a fresh box the same day. The quota is prefilled because the fill warning only fires when quota_gb > 0. Also registers controller v0.244.0's app_deploy_started / app_deploy_failed in both allowedEventTypes and customerMessages, per the rule that the two move together. Red-proofed: dropping the default fails the new render test. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
4c4e3b3a3f |
docs: supervisor (03), node_* ruling (08, CONTEXT), per-tier page + tier skip (07), self-bind triggers + PBS-DR lifecycle (05), settings after install (02), park + No TLS Verify runbooks, volunteer prerequisites
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
a4d684412b |
R-505: tunnel had no published route; operator added *.enkicsifelhom.hu -> https://traefik, verified with a throwaway connector; day-0 A.1 names the exact route
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
65790672d5 |
doorstep walk on ISO 1.27.x: 1 intervention (R-505), STOP before publish
gates / gates (push) Successful in 17s
ISO 1.27.1 gated PASS and proven live: first-boot console Felhom-only, pvebanner masked across a proven reboot. Hub v0.113.0 hand-over copy live (R-497 closed). Full first hour walked again on customer tester-1 (three disks + one disk): deploy, use, backup, removal, byte-identical restore, power cut, typo all PASS. The tunnel gives a fresh box no routes: 12/12 503 from DooPlex (R-505); the record has no e-mail (R-508). Rows R-507, R-508 filed; R-496/R-495 fixed/answered awaiting publish; day-0 A.1 no longer claims the controller creates hostnames (R-506). NOT PUBLISHED. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
6fd8c87516 |
doorstep: console is Felhom's (ISO 1.27.0 source), passphrase hand-over copy (hub 0.113.0 source), rulings
gates / gates (push) Successful in 19s
Phase 0: the public ISO never auto-installs by construction (no answer.toml, G1); the operator re-affirmed the interactive installer 2026-09-14. - felhom-bootstrap.sh: mask pvebanner.service, write a Hungarian /etc/issue (no :8006 admin URL); pairing banner names the Tulajdonosi jelmondat and paints through the CONSOLE_DEV seam (R-496). Harness: 8 checks, red first; fake hub now sends a pairing code (the banner was never tested, R-502). - hub: created flash + Credentials block tell the operator to hand the phrase over; the self-bind mail names the operator (R-497). Tests red first. - iso-release-gate G14-G16; domain ruling in 01-topology + CONTEXT; R-494 narrowed to P3; R-502..R-504 filed; volunteer guide and day-0 A.2 aligned. ISO_VERSION 1.27.0 (not built, not published). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
38848ffbeb |
drill: a stranger's first hour on 0.242.0 — 1 intervention, not ready for a volunteer
gates / gates (push) Successful in 21s
Golden 0.242.0 baked, round-trip verified and vouched (cadence rule, R-468). Fresh box from the public ISO on demo-hp: landed on the vouched set, two apps deployed and used, backup, remove, byte-identical restore, power cut and code typo all PASS. Stopped for a volunteer by R-493 (no instructions) and R-494 (the setup mail's dashboard link has no DNS; intervention I1). R-493..R-500 filed. Capability map: first-hour row added (PARTIAL), journey row scoped. Stopgap Hungarian volunteer guide written. Hub teardown layer pending. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
41590f8ee6 |
second night: scratch guest 9202 built (R-481 CLOSED, persists); controller v0.242.0 delivered (R-487 R-491 R-490 R-476 R-456 CLOSED, R-489 re-scoped); R-492 filed; rotation restarted from bentopdf; morning note
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
5e8a82c3c4 |
night 2026-09-13/14: first "be a customer" rotation (adventurelog) — 7 defects found, 13 rows closed
gates / gates (push) Successful in 18s
New runbooks/nightly-rotation.md; observations_gate.py reads every section (R-471); target-selection.md names real paths (R-461); R-93 carries the fact that drill-r50 is gone. Register: R-473/R-474/R-466/R-471/R-453/R-461 and v0.240.0's R-477/R-478/R-480/R-482/R-484/R-485/R-486 closed; R-481, R-483, R-487, R-488, R-489 opened. 09 §6.1, 07 §6, CONTEXT, STATUS note. Evidence: audits/nightly-2026-09-13-adventurelog/, audits/v0240-2026-09-13/. |
||
|
|
f181efd6a7 |
hub v0.112.0: a floor carries a declared MinAgent past the golden (R-472)
gates / gates (push) Successful in 18s
Operator ruling 2026-09-13. Above the vouched golden, a floor saved with a declared MinAgent is served under the same agent comparison; an undeclared one is still held beyond the golden. The declaration is stored beside each floor as FLOOR=MINAGENT so it never carries to a later floor. Both floor forms require min_agent above the golden (flash floor_needs_min_agent, nothing stored). The Hosts page and the API log name the source. Vouch path and R-120 gate untouched. Scenarios A-E tested; red-proofs A and C in documentation/audits/rulings-r472-r475-2026-09-13/. |
||
|
|
5ef0f52bcd |
Slice 4 shipped (R-448/R-443/R-439 CLOSED, proven live); R-472..R-476; the floor-between-bakes claim corrected
gates / gates (push) Successful in 19s
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure. Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/). Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the runbook, STATUS, CONTEXT, R-468 and the gate docstring. Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
ae59c31a84 |
R-459 CLOSED (MariaDB converts itself, proven by harness + live), golden 0.236.0 (R-467), the golden waiver (R-468)
Operator rulings 2026-09-13, both shipped the same day: - MariaDB finishes its own conversion (catalog eec1228/bd32830/3525e35). Harness E3/E3b `proven` with engine_state_after "already upgraded to 12.3.3-MariaDB [exit=1]", the skip line gone, C3 still `failed`; landed on demo-hp through the real 15-min cycle, nothing recreated, one deliberate restart logged "MariaDB upgrade not required" with the app serving. Evidence: documentation/audits/r459-close-2026-09-13/. The engine-major rule + gate keep every engine inside its major until Slice 4 (R-448) — removal tracked as R-469. - Goldens on a cadence, not per release. golden_currency_gate.py reads a dated waiver (documentation/tests/golden-waiver.yml, <= 14 days, row-bound): valid + BEHIND -> loud advisory, exit 0; expired -> red again naming the date; UNRECORDED (R-385) never covered; malformed -> 2, never 0. Tests cases 5-15 incl. the R-421 decoy; red-proof old-vs-new on the real behind tree. R-242's vouch half stays open. Cadence in RUNBOOK-manual-build.md §4.2 + the checklist. - Golden 0.236.0 baked, round-tripped, vouched, floor raised 0.232.0 -> 0.236.0 (documentation/tests/golden-0.236.0-2026-09-13/) — the last per-release bake; the waiver was issued AFTER it landed. No --no-verify anywhere in this session. Rows: R-459 CLOSED, R-467 CLOSED, R-242 narrowed; R-468/R-469/R-470/R-471 opened. 09 §3 gains decisions 5 and 6; STATUS items 11 and 12 closed; CONTEXT records the cadence ruling. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
1d59353df4 |
provider questions: arm them against the Storage Box / Storage Share conflation (operator-found)
gates / gates (push) Successful in 18s
The operator noticed the "can I restore specific files from within a backup?" FAQ lives under storage-share, not storage-box, and asked which product it covers. It is Storage SHARE only — a managed Nextcloud — and it never mentions Storage Box. Its own text gives it away: Nextcloud's data cache, a database dump, the konsoleH web interface. The two products document OPPOSITE answers: Storage BOX (ours) "You can download individual files or entire directories as usual" Storage SHARE (not) "we only support restores for the full backup ZFS snapshot" That matters because a web search for the obvious phrasing surfaces the SHARE page and it reads like a definitive NO — so a support agent could answer Question 1 from the wrong page and push R-95 to the top of the register for no reason. Question 1 now names the product, quotes the Storage Box line, and states up front that we know what the Share FAQ says. A warning block at the head of the file tells the reader to check which product any full-snapshot-only answer is about before acting on it. Verified by grep: nothing in this repository ever leaned on the Share claim. The only vendor line cited anywhere is the Storage Box one. R-436 strengthened from the same source the operator supplied: the rclone-over-SSH restic backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner -- "we support the restic backend, which is provided by Rclone over SSH". And the same page settles that the docs cannot answer the caveat: neither its Rclone nor its Restic section mentions append-only at all, so nobody need re-read the documentation hoping for it. Unlooked-for corroboration: that page's port-23 command table matches, item for item, the help output measured live on our own sub-account. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM |
||
|
|
db38f4c800 |
hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta
gates / gates (push) Successful in 17s
R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.
was: "...still hold the older copy, so this is recoverable file-by-file; it is NOT
confirmed data loss. Check whether a deletion ran on the box before restoring."
now: "...still hold the older copy. The route back out of them is not yet established,
so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
before restoring anything, and check whether a deletion ran on the box."
It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.
Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.
R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.
THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.
Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.
R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.
Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
|
||
|
|
1f74427fd2 |
target-selection: a drill night will see the golden ADVISORY, and that is expected (R-417)
The instruction that produced R-417 was a prompt, not a file, so the next drill author would have met the same surprise. This is the durable home they actually read before picking a machine. Says what to expect (a loud ADVISORY on every documents push, all night), what still refuses (code pushes, and every other gate), and what the honest instrument is if a release must ship without a golden - a waiver row, never --no-verify. |
||
|
|
c30430c530 |
skills: five process-domain skills + check_skills.py
gates / gates (push) Failing after 15s
The four existing skills cover the product; nothing covered how work is reported. Two rules this project has paid for — check the artifact rather than the report, and do not state a claim more firmly than the evidence allows — lived only in the operator's head and in chat, where Claude Code never read them. - felhom-evidence five confidence tiers, artifact-over-report - felhom-diagnosis no hypothesis until a command has been seen red - felhom-plain-language ASD-STE100, two options, the re-pitch - felhom-handoff the note goes to a FILE, not the conversation - felhom-doc-authoring the pointer decides whether material is reached scripts/check_skills.py asserts what decides whether a skill is EVER reached: frontmatter parses, name == directory, description and body non-empty, under 150 lines, installed copy still samefile()s into the repo. install_skills.py globs and never reads the file, so a missing description installs perfectly and then silently never loads. It convicted on its first run: felhom-build-deploy is 179 lines. NOT trimmed here (pre-existing skills are out of scope, and trimming a deploy skill without exercising its commands is how a wrong command reaches a live host) — a named single-entry GRANDFATHERED exception, WARNed every run, R-394. A new skill over the limit is convicted. Red-proof run and seen failing: description removed from felhom-evidence -> exit 1, "frontmatter field 'description' is missing or empty". Restored, tree clean. skills/SOURCES.md records both MIT upstreams, that these are adaptations not copies, and the six pieces deliberately EXCLUDED with reasons. Register: R-392 (no architecture doc covers the two-AI workflow), R-393 (decision-log skill deferred, with the reason), R-394. |
||
|
|
0a5e9b14dc |
due-checks gate (R-341), floor raise recorded (R-343), snapshot coverage (R-342)
gates / gates (push) Successful in 14s
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as prose in a register row; nothing read those dates and nothing would have objected when they passed. The dates now live in a DUE-CHECKS block INSIDE OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the pre-push hook and CI. exit 0 nothing due (prints pending count + nearest date; empty block too) exit 1 a row is due/overdue (due <= today, UTC -- due TODAY counts), or a row names an item with no R-row exit 2 block absent/duplicated/unparseable -- INCONCLUSIVE, never 0 It REFUSES rather than warns, and its docstring states the limitation: it is NOT a scheduler, it fires on the next push, not on the date. 37 tests. BOTH red-proofs run and reverted -- and the first one earned its keep by catching a hollow assertion of MINE rather than confirming the gate: flipping <= to < left a due-today row in neither bucket, min() raised on an empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the boundary was wrong. An exit code cannot tell a verdict from a crash. The test now asserts the conviction banner and the absence of a traceback, and the gate returns 2 rather than crashing if that partition breaks again. PART 3 — the floor raise, and the premise was WRONG. Read back from the store (not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero per-customer overrides, no "managed floor HELD" line. But read 5 shows the raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save, exactly the immediate action publish-train rule 2 documents. No error events followed; it restarted clean. R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five reads clean and no directive served. It went well, but a record calling it inert when it moved a customer box is what misleads the next reader. The row also states why the floor was behind -- rule 2 policy, not drift, earned by the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor (store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267, fails open at :78-82) rather than asserting them. Two boxes are below the floor and neither reports: drill-r50 (blocked, powered off) and peti-felhom (host row deleted). peti-felhom was NOT contacted -- its row records that a report from a deleted host 401s and is not persisted, so the raise cannot reach it. PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a separate Volume that snapshots exclude, so a rollback restores software state and NOT the datastore. Fine for that upgrade; the safeguard for any future procedure that could touch the datastore does not exist and is Viktor's call. Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling 200). Capability map deliberately unchanged; no row cites a floor or golden version. repo_gates.py fully green, 10/10. |
||
|
|
4d6ec7c7bb |
hub v0.104.0: the guest network gets a reader (R-319), and the hub half of the naming (R-295)
gates / gates (push) Successful in 14s
Four paper debts and one fact given a reader. Hub-only — nothing to bake. A4 — the entry about "the tester's machine" named a risk correctly and labelled it in a way that invited deleting it. Established from the hub's own store: `peti-felhom` is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47 this morning. The prompt's premise conflated the two; the register now says which is which. A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out of STATUS's "Waiting on you", which is now empty. A3 — day0-install §C.1 said pushing the installer publishes it. It has not since R-110. Corrected, with the two manifest pins named and an outside-verification command; the one copy that repeated it (a dated audit, true when written) carries a superseded note. A5 — standing rule 5: evidence comes off the machine at the end of the phase that produced it, before any revert. Earned twice in three days on the same box at the same point (R-320). Four homes, plus what to do when it is already gone. R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll` mail kind so the mail names the page a REBUILT box actually shows („A szerver beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach. Naming only; the acceptance pin proves the secret is untouched. R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never drawn as healthy — three absences, three sentences. No alarm, deliberately. Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190. B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now. Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled` reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed. Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect). |
||
|
|
1c47e3b6fd |
golden 0.203.0 baked + published; runbook acceptance markers fixed (R-233)
gates / gates (push) Successful in 13s
Bake evidence: documentation/tests/golden-0.203.0-2026-08-06/ (bake.log + README). sha256 3039c6ffa7a5a8b2d959daddb2895c58b44de70f8d4f4a7e12ad4b1c0d61dc88, verified by an independent round-trip download and by reading /etc/felhom-controller-image out of the published archive itself. NOT vouched — the hub still serves 0.201.0. R-233: RUNBOOK-manual-build.md §4.1 named three pass markers, two of which the script cannot print (`overlay2 OK` does not exist; `mp1` stopped existing in build-golden.sh v3.0.0 under R-165), and a 404 pre-gate URL with the wrong filename, which would 404 for the wrong reason and pass even when the version already existed. A grep for an impossible string reads 0 forever and 0 is indistinguishable from failure. Markers re-captured from the real log; token handling moved off the command line into an in-VM runner script; a positive control is now required on the token-leak grep; the vouch step rewritten as the three-field change it is (golden_version + agent_version + min_agent). |
||
|
|
5ca5082e7c |
docs: close out the instruction arc — t740 corrected on evidence, registers, ledger, S-37
gates / gates (push) Successful in 8s
target-selection.md said demo-hp has no off-site tier. Measured first: pvesm list felhom-pbs on the box returns two snapshots in demo-hp's OWN namespace (2026-07-28, 2026-08-04) against ep0's felhom-offsite. The claim was TRUE WHEN WRITTEN and went stale when F10 resolved 2026-07-23. The measurement is kept in an HTML comment beside the corrected sentence. This file decides which machine may be destroyed, so the sentence was load-bearing, not cosmetic. R-229(b) CLOSED (agent 175 -> 99 eff). R-230(b) CLOSED (symlink, proven from fresh sessions). R-230(a) part-actioned -- three false statements fixed, WARN loop added, bulk ruling still owed. S-37: a claim in an instruction file is checked, not trusted. |
||
|
|
c21bcf84f7 |
docs+gate: instruction files cannot silently regrow (R-229)
gates / gates (push) Successful in 7s
New shared scripts/instructions_gate.py, registered in controller_gates.py and agent_gates.py, never copied into a sibling repo (the reuse_refs_check.py precedent). 20 fixture tests, all asserting the effect: exit code AND that the message names the file and the reason. It is a consistency gate, not a budget gate, and the failure message says so. A /context reading measured the instruction files at 15k tokens against 869k free in a 1M window -- space is not the constraint, and a future reader must not re-derive the wrong reason. The 200-line ceiling is adherence guidance; a file nobody can hold in their head is where contradictions hide, and five were found here. Checks run against effective text (HTML comments stripped, because they are stripped before injection): the line ceiling; every .claude/rules/*.md declares paths: or an explicit unconditional: true; no component version literal; no TEMPORARY block carrying a past date; and the workspace-root CLAUDE.md is byte-identical to its versioned copy -- the live file sits outside any git repo, so that copy is its only version-controlled record. Two traps recorded so they are not reintroduced: a bare \d+\.\d+\.\d+ matches the first three octets of every IPv4 (the gate excludes dotted quads, or it fails on 192.168.0.180 in the agent's own file); and unconditional: true is NOT a Claude Code feature but this project's own marker. Workspace-root CLAUDE.md 208 -> 182 lines (142 effective), copy kept identical. The nine-instance invariant table moved into the felhom-testing skill, which triggers when writing or reviewing a test; all three directive bullets stayed in the core. felhom.eu/CLAUDE.md got surgical corrections only and is knowingly still over the ceiling at 227 effective lines -- closing it needs the restructure R-229 defers, said plainly rather than quietly absorbed. CONTEXT.md gains standing ruling S-35. OPEN-ITEMS.md gains R-229. Docs only -- no Go, no version bump, nothing built or deployed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JJc8sAGRWmavP3rMtdpkr2 |
||
|
|
d30c2a51ed |
R-224..R-228 CLOSED: registers, capability map, campaign annotation, STATUS
gates / gates (push) Successful in 7s
Five closed in controller v0.202.0 + agent v0.126.0, each with its live or red-proof evidence in the row. Five explicitly still open and named as such rather than left to inference: R-214, R-220, R-221, R-213, R-202 — and R-220 is flagged as currently worked around BY HAND on the campaign venue, which is the only reason an app could be deployed there. The capability map's recovery row STAYS FAIL and says why: fixes are not a re-walk, nothing walked a customer end to end, and the customer-facing messages were NOT re-driven live because /recovery correctly retires itself once the old data is set aside — restoring that state is the reconfiguration the task forbade. The campaign document is ANNOTATED, not rewritten: it records what was true when it ran, and that is its value. workspace-CLAUDE.md gains comment-vs-code entry 9 — the escrow header said the errors were 'DISTINCT on purpose' and named THREE situations while a fourth was folded into one of them, and a green test named the defect and did not prevent it because it asserted a STRING one layer below the merge. ROADMAP needed no collapse — it carries no rows for these IDs. |
||
|
|
1a0f7db92f |
docs: CAMPAIGN-11 — the journey FAILED, R-198's retention PROVEN, six findings fixed
gates / gates (push) Successful in 8s
Registers and evidence for the campaign and its fix pass. OPEN-ITEMS: R-214..R-223. Six SHIPPED (R-215/216/217/218/219/222); three deliberately still open and each blocks a real flow (R-214 console banner, R-220 drives unenrollable after a rebuild, R-221 a rebuilt box cannot run the escrow ceremony); R-223 minted and WAITING-ON-OPERATOR (vouch agent 0.125.0). R-213 and R-202 untouched. Capability map: a new row for the customer's UNAIDED journey, recorded FAILED and staying failed until a re-walk passes — fixes are not a journey. The existing rebuild row is corrected where it said R-198's retention was unit-proven only: it was proven in production on the first supersession since the fix, identity_blob retained at 572 B byte-length exact. CLAUDE.md comment-vs-code table: eighth entry — ResolveManagedFloor, the first where the false invariant was a GUARD rather than a comment alone. STATUS: the headline is now "the backup promise is proved, the recovery journey is not", and the one thing waiting on the operator. |
||
|
|
b93ee06abc |
R-190 filed; two corrections to yesterday's R-185 record
gates / gates (push) Successful in 8s
CORRECTION 1 — the runbook annotation and the R-185 row both said the drift did not surface as a 403 because writes go through a root path. That is WRONG. demo-felhom's local-api backup jobs 403'd six times between 09:24 and 17:34 CEST on exactly that storage and privilege, and the hub raised whole_guest_backup_failed at the first with edge-triggering suppressing the rest. The impact was not only an unreadable tier: the agent's own whole-guest backups to it were failing. CORRECTION 2 — on this box the grant was LOST, not never issued. A vzdump by the agent's token to that storage completed OK at 04:44:50 the same morning; the first 403 is 09:24:56. Ruled out by measurement: a host reinstall (uptime 12 days), any pveum/ACL/user.cfg activity in syslog 04:00-10:00, any ACL entry in the cluster log. Correlated but not established: guest 9201 was reprovisioned nine minutes before the first failure. R-190 files the unexplained disappearance, and notes that the new store-grant probe detects the STATE but says nothing about the TRANSITION. |
||
|
|
e3187c86d5 |
docs: R-185 closed — the silence as well as the grant
gates / gates (push) Successful in 8s
- OPEN-ITEMS: R-185 closed with the measurement, the corrected root cause (the installer's Scenario-F reuse arm, not PVE_STORAGES), and the live sequence. Records that demo-hp carried the same drift and was fixed too. - capability map: the whole-guest row's HOST-tier half was OPTIMISTIC and now says so — that tier was not merely unproven, it was unprovable on both demo boxes, and every live proof cited was on the offsite tier. - vzdump-target-move runbook: its item 5 predicted this; annotated (not rewritten) with what actually happened — the create arm did grant, the reuse arm did not, and it surfaced as a silent unreadable tier rather than the 403 the item expected, because vzdump writes through a root path. - CONTEXT: S-21 (an empty listing cannot distinguish forbidden from newborn; the measured trap that an ungranted path answers with INHERITED privileges) and S-22 (the Scenario-F arm must finish the job). - STATUS: rewritten for the operator, back to one screen. |
||
|
|
323f45a5ef |
hub v0.91.0 — the staleness window learns each tier's own rhythm (R-86 Part 2)
gates / gates (push) Successful in 7s
Ships WITH agent v0.121.0, not after it. The agent now proves a tier once per
ARCHIVE GENERATION, so a weekly tier is proved weekly — in perfect health. The
flat 7-day restoreProvenStaleAfter derived its number from the 24h cadence R-86
removes, and a healthy weekly tier's proof age reaches EXACTLY 168h just before
its next proof: it sat ON the line, so any ordinary delay tipped it into a
nightly alarm about a working system.
restoreProvenWindow(tier, observed, ok):
- the tier's own archive interval, OBSERVED from reports the hub already holds
(pbs_snapshots + successful backups attributed by TARGET TYPE, slice A.4)
- x4 generations = the same tolerance the flat constant expressed
- floored at 7d (never tighter than before), capped at 12d (strictly inside the
2-week offsite retention)
- falls back to the DECLARED rhythm (26h host / 8d offsite — the thresholds the
backup-freshness checker already uses) when history is too short to observe
one; falling back to the FLOOR would recreate the false alarm on a fresh box
Kept: absence is UNKNOWN until the anchored window passes; the signal stays
edge-triggered; failed and stale remain distinct events. Every reason string now
states the window it was judged against (R-100's corollary).
Also backfills the missing v0.90.1 CHANGELOG entry (deployed since
|
||
|
|
e34b614e5b |
docs: R-182 closed, R-90 closed on measurement, R-86 unblocked, ep0 record corrected
gates / gates (push) Successful in 7s
R-182 CLOSED (controller v0.194.0 + hub v0.90.0/.1), proven live on demo-hp. The hub's notification_log for the run reads: two per-app failures RECORDED, one digest SENT naming both, and the customer channel SKIPPED with operator_only. Against the measured previous behaviour — two failures, one email naming one app, one leaving no trace anywhere. Scenario D proved itself on an event I had not planned: disk_critical alarmed on two filesystems, the second was collapsed by the cooldown, and that collapse is now visible WITH ITS KEY. Yesterday it would have left nothing at all. A gap the spec did not anticipate is recorded with its fix: the per-app event also fires from the periodic sweep, outside any run, so making it record-only would have created a NEW silence. The sweep emits a digest too, with no run_id, so it stays under the ordinary hourly cooldown. ep0: MEASURED on the box — 7757 MB (8 GB), 4 vCPU, and the 4 GiB swapfile SURVIVED the resize and is active (checked, because a resize is a stop/start). The 40 GB local disk is UNCHANGED, so no disk figure was touched anywhere. Five documents corrected — three of which the task's list did not name, found by searching. Two audit/evidence documents ANNOTATED, body untouched: they record what was true when written and that is their value. R-90 CLOSED. R-86 unblocked and re-ranked, stated honestly: 8 GB is comfortable, not unbounded — the original OOM was a 14.46 GB restore — so the restore-test cadence should still be paced, just not by fear of the endpoint. target-selection.md's "D-d did not name ep0 either way" is deliberately left standing. It is the operator's question, not CC's. STATUS.md 127 -> 83 lines, items rather than sentences. |
||
|
|
8360f940bf |
docs: seventh row in the shipped-guarantees table (R-181), versioned workspace CLAUDE.md
gates / gates (push) Successful in 7s
Syncs documentation/runbooks/workspace-CLAUDE.md with the workspace root file. Row 7: the B2 refusal claimed the previous unit was untouched; nothing-deleted held, untouched was measured false. Also records WHY it survived review: it passed a full green suite AND three of its own red-proofs, because every one of them asserted the mechanism inside captureAllRecoveryUnits and none asserted the consequence across the whole backup run. The test that would have caught it is the one the fix ships — fingerprint the tree before and after, and compare. |
||
|
|
e994bf35d2 |
STATUS.md: a plain-language operator page, and today's four decisions recorded
Documentation only — no code, no box, no build.
STATUS.md (repo root, 652 words / 67 lines): what works · what's broken ·
what we're working on · waiting on you · changed since. A VIEW of
OPEN-ITEMS.md, holding nothing of its own; not CONTEXT.md, and both files
now say why they stay separate. No R-n is the subject of a sentence —
identifiers are bracketed pointers only.
CONTEXT.md S-5 records the four operator decisions taken 2026-08-02
(D-a … D-d), none of them implemented:
D-a merge mp1 into mp0 rather than resize it — before any external
install, and D-c ships in the same step → R-165
D-b desired/observed app state in its own store, with the state-store
safety rule verbatim → R-166 (BLOCKED)
D-c customer fill warning + operator backup-failure alert → R-167
D-d only DooPlex and Peti's box are protected → target-selection.md
R-163 RE-FRAMED, not closed: the sizing question is withdrawn rather than
answered; the row survives as the record of the constraint until R-165
lands. R-156's papra referral RESOLVED — deployed nowhere, so the template
fix strands nothing; the docker ps evidence is recorded with its
provenance and its scope limit.
target-selection.md: two protected machines, everything else disposable.
ep0 is no longer Tier 2 but is not scratch (it holds the only off-premises
copy of real customer data) — flagged for explicit operator confirmation.
The demo-box backup-target fence drops from prohibition to stated cost,
because D-d spends that reference anyway.
CLAUDE.md gains an End-of-session checklist carrying the STATUS.md
maintenance rule and "a finding goes in OPEN-ITEMS.md first".
|
||
|
|
e9a74a0019 |
docs: remove a gate criterion that could never pass, and close three register rows
PART 1 — the release gate.
G7 required the packaged .deb to sha256-match the one built from committed source. That is
unsatisfiable BY CONSTRUCTION: dpkg-deb stamps the build time into every archive, so two builds of
byte-identical source differ. It was already failing when the 1.26.1 release ran it. A criterion
nobody can satisfy gets waived once and read as advisory ever after — which is how R-29's shelf of
never-run gates was built. Sub-clause dropped, reason recorded in G7's own note the way G6's
amendment was, so a future reader can restore it if SOURCE_DATE_EPOCH ever makes it meaningful.
RULING ASKED FOR — is payload integrity covered by G9 alone? NO, and G9 is widened rather than a new
criterion invented. The package ships TWO payload files (build-deb.sh:54-55); G9 checked only the
script. The systemd UNIT was covered by nothing: G7 covered the container, G8 covers the postinst
behaviourally, G13 covers directory presence. The unit is not incidental — its After=, its
ConditionPathExists= and its Restart= decide WHEN AND WHETHER day-0 runs at all, so a drifted unit
would have shipped silently. Same shape as the /etc/felhom miss that G13 exists to prevent: a check
that proved the thing present and said nothing about what it depended on. The check passes today.
G13 moved to sit after G12 — it was minted late and left between G10 and G11.
PART 2 — register dispositions. BASELINE DISCREPANCY, reported rather than worked around: only R-128
had a row. R-154 and R-155 had NO row in either file — minted in a spike document and never carried
across, which is R-123's class, not the drift the task described. Rows created, closed, with the
reasoning, because in all three cases the reasoning is the durable part:
R-128 closed by CORRECTING a false claim, not by making the assertion real — the coupling does not
exist and asserting it would invent a constraint. Flagged so nobody 'restores' it.
R-154 closed with the measurement and where it now lives in pushed source.
R-155 NARROWED, not deleted — unchanged for FELHOM_MENU=single, inapplicable to release. Flagged so
the guard is not later removed wholesale on the strength of 'R-155 closed it'.
Documentation only: no code, no build, no ISO, no upload, no box touched.
|
||
|
|
f2fc76ec4b |
ISO v1.26.1 PUBLISHED — both entries proven, round trip verified
Live: https://iso.felhom.eu/felhom-installer-1.26.1-pve9.2-1.iso sha256 f3cc86d5f0ec68bba4155c994b4fa84e208d50209bb6e815636c99e5441059a6, 1705322496 bytes. PART 5 PASSED ON BOTH MENU ENTRIES, four observables each: Graphical spikegfx.felhom.eu pairing code J7N-2DA TerminalUI spikesix.felhom.eu pairing code ZY5-YY4 Both: manual install, own disk, own password, real completion signal, and the journal's 'not bound yet — polling every 30s ... normal waiting state, not an error'. Spike 4 had REASONED the graphical path follows from shared Install.pm; it is now measured. PART 6: G1-G10 + G13 all PASS against the uploaded file. G4's single hit is felhom-bootstrap.sh:480's substring TEST ('$envtext' != *FELHOM_RETRIEVAL_PASSPHRASE=*), not a value — my own regex matched the glob's asterisk. PART 7: uploaded via rclone in a container configured ENTIRELY by environment variables, so no credential file was ever written. Round trip verified from the public URL — not the local file. Bucket stays private: unauthenticated GET to the S3 endpoint 400, custom domain has no index (404). CORRECTED BEFORE UPLOAD: the generated manifest described a single automated entry with a 5s timeout and listed Graphical/Terminal UI as 'menu-removed'. Generator fixed, sidecar regenerated, and the ISO verified byte-identical before and after — the published file IS the file Part 5 validated. Hub-side cleared: appliances 16, 17, 18 discarded (303 each); zero rows remain. The endpoint is /appliances/<id>/discard, POST only (server.go:345) — not /delete. Teardown: VMs purged, spike5 storage removed, demo-hp back to 6.6G, drill-r50 and 9201 untouched. Still open and named: OPEN-ITEMS/ROADMAP dispositions for R-128/R-154/R-155 are not written; the .deb is not byte-reproducible (G7 sub-clause); before-network stub unreached; Secure Boot and real hardware not exercised. |
||
|
|
a967da7d2c |
iso 1.26.1: ship /etc/felhom/ — the directory the bootstrap writes its state into
FIX for the Part-5 failure. felhom-bootstrap.sh writes the appliance token (:431), the pairing code
(:435) and .bootstrap-done into /etc/felhom/. The old stub-first-boot.sh created it explicitly
('install -d -m 0755 /etc/felhom /usr/local/sbin'); packaging dropped the env FILE correctly and the
DIRECTORY with it. Measured consequence on a real interactive install: the box registered at the hub,
could not persist its token, and polled 'HTTP 401 — still retrying' forever with no claim code.
- build-deb.sh now ships ./etc/felhom/ (0755, empty) and ASSERTS it, plus ./usr/local/sbin/ and
./lib/systemd/system/, as G13. RED-PROOFED: removing the install -d makes the build exit 3 with
'is not in the package (G13)', and restoring it goes green.
- The gate gains G13 with the reasoning: G7/G8/G9 all passed on the broken package. G9 proves the
payload is the right payload and says NOTHING about what the payload depends on.
ISO_VERSION -> 1.26.1.
|
||
|
|
01a8155c5a |
iso v1.26.0: the PUBLIC release image — no answer file, interactive install, day-0 by .deb
Design inputs: SPIKE-universal-iso-{1,2,3,4}-2026-07-31.md. Every choice below is a measurement.
NEW: scripts/iso/pkg/ — the felhom-bootstrap .deb, built from committed source.
Two files only (script + unit), NOT three: felhom-bootstrap.sh:91 reads /etc/felhom/bootstrap.env
only 'if [[ -r ]]', and its defaults at :95-96 are EXACTLY what the pairing env set
(build-felhom-iso.sh:257-258) — so shipping it would add a 0600 file to a public package to express
values the script already defaults to. NO dependencies: the binaries it calls run at FIRST BOOT,
not at postinst time, so SPIKE 4's open 'dpkg --configure -a' ordering question does not arise.
The postinst is structurally incapable of failing (no 'set -e', every statement guarded, ends
'exit 0'); build-deb.sh self-asserts G8/G9 and REFUSES to emit a package that violates them.
iso-repack.sh — two changes, both narrowing rather than deleting:
- R-155 guard: now applies to FELHOM_MENU=single ONLY. It protected the single-entry mode's promise
(one button labelled 'install' must not drop into a disk-picker); a release image carries no
auto-installer-mode.toml BY DESIGN (gate G1), so refusing it would be the guard firing on the
shape it describes rather than the one it prevents.
- the menu collapse now has a release mode: two INTERACTIVE entries, Graphical default, timeout 15.
Entry-count and banned-token gates are per-mode; the six-token list is UNCHANGED for single mode.
- .deb injection into /proxmox/packages/, with a skip-list collision check (a colliding name would
be dropped silently — the inert-payload class) and a post-remaster assertion that it landed in
final.iso, not merely in the extract tree.
build-felhom-iso.sh — --release: no profile, no root hash, no answer.toml, no prepare-iso at all.
Skipping prepare-iso is what removes the Automated entry by construction, since the stock grub.cfg
emits it only inside 'if [ -f auto-installer-mode.toml ]'.
R-128 RULING — FIXED, by correcting the claim rather than inventing an assertion for it. The comment
said ISO_VERSION 'aligns with SCRIPT_VERSION'; nothing evaluated it and the two had drifted. The
coupling does not exist: the ISO is frozen, felhom-host-install.sh is fetched at run time from main
(R-94/R-110), so an assertion would invent a constraint. Comment corrected, ISO_VERSION -> 1.26.0.
Release gate G6 AMENDED before the build, with its reasoning recorded in the runbook: the six-token
ban existed to keep users away from the manual installer, which the ruling makes the product.
'proxtui' (the TUI installer we ship) and 'nomodeset' (its graphics fallback) are dropped for
release images; proxdebug/Rescue Boot/memtest/fwsetup stay banned in both modes.
|
||
|
|
e787391c0a |
docs: the public ISO release gate, written BEFORE the first release image
A standard defined in advance cannot be rationalised afterwards, and this is the artifact that most needs one: once a file is on iso.felhom.eu and someone has downloaded it, it cannot be recalled. Twelve criteria, each checkable against the UPLOADED FILE rather than the build inputs, and each carrying the spike measurement that justifies it: - G1 no answer.toml / auto-installer-mode.toml — deletes the whole Spike 1-2 problem space and removes the Automated menu entry by construction rather than by a guard - G2/G3/G4 no root hash, no SSH key, no customer identity — the shared-credential classes - G5 credential scan by ENUMERATION against the stock ISO, not a pattern sweep (Spike 1 found /answer.toml precisely because the earlier recon grepped the wrong file) - G6 menu present, interactive default, timeout >= 10 (Spike 2 lost a probe to a 1-second menu), underscore timeout_style, and the banned-token safety gate kept unchanged - G7/G8 the felhom .deb present, and a postinst that cannot fail: no systemctl start/daemon-reload (no systemd runs in the installer chroot), no network use (the cable may be out), no 'set -e', ends 'exit 0' - G9 felhom-bootstrap.sh byte-identical to repo HEAD — the one frozen, drift-capable payload - G10 every build input committed (Spike 1: demo-felhom came from an uncommitted profile) - G11 published checksum AND a verified download round trip - G12 bucket Public Access stays Disabled Committed on its own, before any build. |
||
|
|
b4edc087fa |
Tester gate: golden re-baked to 0.188.0, fresh-install proof PASSED — a fresh box is safe to hand to a tester
§7.2 answer: YES. A real day-0 from the existing v1.25.0 ISO reached a claimable,
app-serving box in ~10 minutes unattended, and an app's data came back from the
drive with the guest's app.yaml gone — proven readable by the application over
its own TCP path, with a discriminator (PRE-BACKUP row = 1, POST-BACKUP row = 0).
Part 0: NO ISO rebuild needed, verified against the ISO on disk rather than from
source. It bakes only felhom-bootstrap.sh, its unit and the secret-free pairing
env (full-base64 match, 1 hit each) and 0 hits for any installer, controller or
golden marker. The installer is fetched at run time; the live URL is byte-identical
to repo HEAD (v1.22.0, six days newer than the ISO) and the fresh box ran it.
Part 1: baked 0.188.0 rather than the brief's 0.187.0 — 0.187.0 lacks D5, which
is the very claim Part 2 step 6 tests. Published (404 pre-gate with a 200 control;
anonymous download, 649310288 bytes, sha match), vouched, and consumed by a real
box. R-120's gate exercised BOTH ways: 0.185.1 refused with no write, 0.188.0
allowed — evaluated, not silently skipped.
Part 3: RUNBOOK-manual-build.md cited a "RECORDED" qemu line that is itself
labelled reconstructed and whose source says it was never saved. The real
invocation is now captured from this bake as §4.0, with the bake/publish/teardown
steps; the old entry is marked SUPERSEDED.
Teardown all three layers, hub disposition stated: VM destroyed, scratch storage
removed with space returned exactly, customer sess-g DELETED via full cascade.
sess-f deliberately left (R-131) with its command recorded.
Filed, none fixed: R-128 (false ISO_VERSION invariant comment), R-129 (demo-hp's
"no baked SSH key" is stale — key auth works), R-130 (HARD_MIN_LVM_GIB warns and
proceeds), R-131 (fourth orphaned scratch customer), R-132 (curl's %{redirect_url}
printed the hub operator password into a transcript — HUB_PW needs rotating).
|
||
|
|
1956e5d390 |
hub v0.84.0 — break-glass console credential on the host page
The credential existed and was not reachable when it was wanted. Every box has
had a strong random root@pam password since TASK G1, vaulted in the hub at day 0
and used for real during the sshd incident — but the only way to read it back was
a hand-written curl carrying the global operator key, a secret kept out-of-band.
In practice the PVE web console on a demo box felt locked.
The host page grows a Console access card: presence + username + set_at by
default, Reveal fetches the plaintext on demand for 60 s with a Copy button.
Masking clears the JS variable, and also fires on a second click and on
visibilitychange. A host with nothing vaulted says so, and says why.
The secret is NEVER rendered into the page, and that constraint shapes the
change. The render path uses a new store.GetHostRecoveryMeta whose struct and
SELECT both omit the secret column, so it is structurally incapable of carrying
one. The plaintext crosses the wire only in the response to POST
/hosts/{id}/reveal-recovery-credential (Cache-Control: no-store, CSRF-gated at
the ServeHTTP level; POST precisely so that gate applies and so no secret is
retrievable by URL alone). Deliberately NOT the customer page's data-secret
widget, which embeds the plaintext on every load.
A delivered reveal writes one recovery_credential_revealed event on the host's
customer timeline (info, source hub, Hungarian) via SaveEvent alone — no
dispatcher, nobody emailed, the log_tail_requested shape. Two reveals write two
events: the register records accesses, not states. A 404 is not an access. An
unbound host reveals fine and writes no event; the [INFO] hub line, carrying the
username and a length only, is then the record.
The global-key API path is untouched by design — it is the route for when the
hub UI itself is broken, and coupling it to the session layer would delete the
independence that makes it a fallback.
Recorded as a real trade: the hub session password alone now unlocks console root
fleet-wide, where retrieval previously also needed the global key. Accepted for a
single-operator, HU-geo-fenced hub that already stores these passwords in
plaintext at rest (CONTEXT.md ruling S-4). The plaintext-at-rest half is filed as
R-133 — every hub DB backup is a fleet-wide console-credential dump.
Tests 550 -> 559; four red-proofs (page leak, audit event, CSRF gate, route
order) each run, observed failing, and reverted. The route-order proof is a seam
test driving ServeHTTP: a handler-level test cannot see that defect, because the
handler is correct and simply never runs.
|
||
|
|
376365bb12 |
docs(target-selection): a fixture may prove a mechanism; only a fresh box may prove a path
The page said which machine is safe to break but not when reusing a test box is legitimate. That distinction is exactly what surfaced R-120: R-116's closing run deliberately did a real day-0 from the ISO instead of reusing the standing fixture, and the fresh box installed the golden's controller -- a release behind -- and showed the customer the wrong absent-target message. A fixture would have shown a controller nobody installs. Adds to the Tier 1 section: a reusable snapshot-reset fixture is the right default for MECHANISM work (payload capture, fix cycles, claims about code behaviour), while a fresh day-0 from the ISO is REQUIRED for any claim about the install path, the golden image, agent publish/vouch or first-boot state -- naming the drift family it exists to catch (R-111, R-115, R-120). Also: a fixture must record its provenance (which golden, agent and controller, and when), because a fixture whose versions drift silently is R-120's mechanism turned into a permanent installation -- worse than no fixture, since it produces confident wrong results quickly. Part 1 of the R-120 task, committed alone and before the bake. Docs only. |
||
|
|
e6b5fa1e63 |
docs: retract the expired NVMe fence, drop component versions from the inventory, fence acts in the template
Follow-up acting on the observations filed with runbooks/target-selection.md. operations/nodes.md - The demo-hp NVMe was documented "PRESENT AND UNENROLLED -- do not touch" and listed under "What is NOT enrolled here (deliberately)". Both are FALSE and had been for eight days: it was enrolled 2026-07-22 through the normal Tarhely flow and is now /mnt/nvme-1tb -- the enrolled user-data drive AND the felhom-backup target (verified live 2026-07-30: nvme0n1 -> /mnt/nvme-1tb, and dir: felhom-backup / path /mnt/nvme-1tb / is_mountpoint 1). The fence's own condition (join via Tarhely, not the installer, not by hand) was SATISFIED, so the prohibition expired with it -- while still contradicting the task specs that correctly sent drill-VM disks there. Retracted with its reason recorded, and the caution that IS still live kept (dir storage at the mountpoint ROOT, else exactMount fails and the storage reads disconnected forever). - Component versions REMOVED and a note explains why: agent/controller/hub versions change several times a day, so a number written in an inventory is wrong within hours and then read as fact -- and the fleet is not uniform (on 2026-07-30 the two boxes ran different agent AND different controller versions). Points at the authorities instead: hub /hosts + /configs, felhom-agent --version, docker ps. - Site addresses now say re-check rather than asserting one (the N100 read .162, not the recorded .147); records that LAN literals are unreachable from DooPlex while the boxes are away. Adds the target-selection pointer: this page is what the hardware IS, that page is what may be done to it. PROMPT-TEMPLATE.md -- the upstream generator of the defect - Section 12's "Do NOT touch [the untouchable]" asked the spec author to name a THING. Now asks for the forbidden ACT plus its REASON, with the demo-hp case as the worked example of how a bare object-fence over-reads. - Section 13 gains the positive counterpart, which was the actual gap: if a task needs a machine to break, NAME IT. Listing only what is off-limits leaves the most valuable unfenced machine as the residual choice. runbooks/workspace-CLAUDE.md (+ the untracked root copy re-synced, verified identical) - Host table gains a Blast radius column and the missing demo-hp row, notes felhotest as Connection refused, and points at target-selection.md. This is the file that loads FIRST every session, so leaving it with the old table would have undercut the whole fix. No code, no build, no deploy, no host reconfigured or renamed. |
||
|
|
699790b12d |
docs: write down which boxes are disposable (target selection by blast radius)
Nothing in the repo said which machines are safe to break. The host table gave access and role and stopped there, so a session needing a victim had to guess -- and the guessing inverted: the two boxes that exist to be broken were treated as sacred, and DooPlex (the recovery chain) got used because it was the only box no spec had fenced. New documentation/runbooks/target-selection.md -- one page, three tiers, and per machine what is freely permitted / needs care / forbidden, each carrying its REASON so a rule can be correctly narrowed later instead of ossifying. States the selection rule positively (start at Tier 0; a Tier 2 box only when a task says so explicitly; an absent fence is not permission) and that fences name ACTS, not machines -- demo-hp's over-subscribed local-lvm is one dangerous storage, not a dangerous box. CLAUDE.md: host table gains a Blast radius column, gains the missing demo-hp row (it was where the drill VMs ran and it was not in the table at all), and a pointer line to the new runbook. CORRECTION to the spec's problem statement: the designation was not missing. The 2026-07-25 operator ruling naming the t740 as drill+build VM host -- explicitly "moved off DooPlex" -- already existed in operations/nodes.md. It sat where no session reads at start, while the prohibitions were repeated in every task spec. The defect is reachability of the ruling, not its absence, and the R-116 drill on DooPlex contradicted a written ruling rather than filling a vacuum. CORRECTION to the R-116 record, same commit: the baseline claimed controller 0.186.0 on both demo boxes. Only felhom-pve was sampled and generalised; demo-hp re-checked directly runs 0.185.1, so the fleet is split and R-114's TargetAbsent branch is absent from demo-hp. Fixed in the audit table and REPORT-r116-diag. Docs only -- no code, no build, no deploy, no host reconfigured, no host renamed. |
||
|
|
d4c07873ca |
docs: correct the installer-channel record — R-94 retracted and re-scoped, R-110 opened
The 2026-07-29 R-94/E-2d finding was written from an unverified claim and was false. `felhom-bootstrap.sh:96` fetches the installer from the WEBSITE, not the hub; the website git-syncs /scripts/ from main on a 30s period; every install since 1.22.0 hit main this morning already runs 1.22.0. Confirmed by live fetch. - OPEN-ITEMS.md: merge the two duplicate R-94 rows into one, retract the false framing, re-scope to what it actually is (a drifting hand-synced constant plus two pieces of dead safety equipment), unblock it from E-2d. - OPEN-ITEMS.md: de-rank R-94 in the ranked list — the "high-consequence" reason was the false claim in its most load-bearing form. - OPEN-ITEMS.md: E-2d — the ISO is the STRONGER proof route, not an obstacle. Phase 0 question answered at source: PAIRING falls through to run_direct in the same invocation (:495-499), so it reaches the identical installer call. - ROADMAP.md:149: same retraction; the original diagnosis (a hand-synced constant in a second repo drifts every time the first ships) survives. - ROADMAP.md + OPEN-ITEMS.md: new R-110 — main is the installer's publish channel and there is no staging, tag, pinned path or rollback, for the one artifact that runs as root on a virgin box. Operator ruling, not a defect. - day0-install.md C.1: one sentence recording the same about the fetch URL. Documentation only. No version bump, no CHANGELOG entry, no code, no box touched. |
||
|
|
b5a73e050b |
Move the local whole-guest backup off the guest's own device (demo-hp + demo-felhom)
Supervised operational run. No code, no version bump, nothing deleted.
Primary backup tier on both demo boxes moved from `local` (a dir storage on
/var/lib/vz -- the SAME physical device as the guest) to `felhom-backup`, a dir
storage on each box's secondary drive:
demo-hp /mnt/nvme-1tb uuid:91d2dc2d-... archive 2,256,044,492 B
demo-felhom /mnt/hdd_1 uuid:47a3361a-... archive 5,957,878,962 B
Both proven end to end via the real UI path: archive lands on the secondary
drive (df delta matches the archive byte-for-byte), restore-test auto-selects it
and passes with mount_parity: ok, and freshness survives an agent restart with
an empty in-memory store -- so the age can only have come from the new storage.
Phase 0: the target is CONFIGURATION, not converged (the sole writer of
agent.json touches only escrow.pbs_storage_id and preserves unknown keys), so
the runbook's STOP did not fire. No consumer hardcodes "local" on the backup path.
Findings:
- F-1 the storage path must BE the mountpoint; a subdirectory fails exactMount
and the target reports disconnected permanently (observe.go:321)
- F-2 --is_mountpoint 1 is load-bearing; proven live, an unguarded storage on a
non-mounted path reports active with the ROOT filesystem's free space and
had already created dump/ on pve-root -- a silent retarget onto the very
device this change escapes
- F-3 FelhomAgentStore is granted per storage path; without it every backup
403s. felhom-host-install.sh must issue it for new installs
- R-109 (new) the DR recipe records no backup target, and each box now carries
two content=backup dir storages, one live and one frozen
- R-105 narrowed and TRACED: dr_recipe drives was [] fleet-wide because the
enrolled drives were never PVE storages, so isUserDataDrive never saw
them. Both boxes now populate drives; SMART on the backup drives too
Absent-drive behaviour today is fail-loudly with no silent retarget (PVE half
live-proven; agent half source-traced). That is NOT the intended fall-back-and-
alarm design -- filed as E-2 with the honest single-drive label.
Reported in full in the record: the agent was restarted with a felhom-pbs backup
in flight, producing a spurious tier failure. The backup had in fact succeeded
(PVE task OK, 6,264,034,053 B snapshot) and the spurious failure reached no
channel -- R-84 ground truth superseded it.
Outstanding: full drive-loss recovery (needs physical access) and the agent half
of the absent-drive behaviour.
|
||
|
|
f47b0a61d7 | R-101 + F-DIAG closed, F-OPS documented (manual-restore runbook) | ||
|
|
b505ee9125 |
R-100: offsite staleness counts from the last SUCCESS (hub v0.80.0)
isStale counted from last_run, written unconditionally on failure, so a nightly-failing tier read as fresh forever. Now anchored on last_success with an explicit legacy degrade (logged once) and the never-ran branch untouched. emitStale states the real reason. |
||
|
|
e168600148 |
docs: F-CRIT-1 + F-A1 shipped (controller v0.179.0); invariant rule
Both marked SHIPPED + PROVEN-LIVE in OPEN-ITEMS and the campaign doc. All three of Campaign 8's alarm findings are now closed (F-CRIT-1, F-CRIT-2, F-A1). Adds the standing rule earned by this arc to the versioned workspace CLAUDE.md: a comment asserting an invariant needs a test pinning it, or it is a wish — with all six shipped-false-guarantee instances catalogued, and the corollary that a test should assert the CONSEQUENCE (does the alarm fire?) not the MECHANISM (does suppression expire?). |
||
|
|
a0a1556ce6 |
docs: R-88b shipped; standing rule 4; READY rows re-ranked
R-88b closed (agent v0.105.0 + controller v0.178.0) — age_state gives 'unknown' its own representation, with empty meaning legacy rather than unknown so the first-backup valve keeps working on un-upgraded boxes. R-97 note updated: hub v0.79.0 (R-97c) replaced a FALSE operator-only comment with a real register — the comment claimed a guarantee the code did not provide. Standing rule 4 (R-96): a recommendation that is not followed gets one line saying why. Added to the live CLAUDE.md and this versioned copy — the live file is not in a git repo, so committing to it alone would leave the rule as durable as the chat it came from. READY re-ranked: R-95 now leads. |
||
|
|
9ea5675950 |
docs: sync workspace-CLAUDE.md with the live file, carrying R-96's three rules
The workspace root /mnt/5_hdd/felhom.eu/git/CLAUDE.md is NOT a git repo — this is its only version-controlled copy, and it had drifted since 2026-07-19. Committing the three standing rules to the live file alone would have left them exactly as undurable as the chat log they came from, which is the whole point of R-96. |
||
|
|
a31872ea24 |
docs(pbs): move PBS prune server-side, close the write proof, schedule GC
Supervised runbook execution. No code, no version bump. The felhom-pbs tier had reported `job errors` on EVERY demo-hp backup since the tier was created on 07-26, while the data landed correctly every time: `DatastoreBackup` grants Datastore.Backup but not Datastore.Prune, so the box's keep_last=2 prune was denied. Operator ruling: retention is a COMMERCIAL attribute owned by the hub; ep0 executes. Box tokens therefore stay write-only - a compromised box must not be able to delete its own offsite backups. No grant was widened and felhom-tenantsync.sh is unchanged (the ruling makes it correct). Increment 1: - boxes stop attempting prune. allowPBSPrune is DERIVED (`!t.Primary && t.KeepLast > 0`), so keep_last: 0 on the PBS tier disables both the --prune-backups value and the gate in one config edit, and the tier stays armed. Verified prune_pbs_allowed=false on both boxes with no tier REJECTED line. - per-namespace prune jobs on ep0, keep-last 2, daily 03:30 UTC (05:30 CEST), dry-run gated. demo-hp 3->2, demo-felhom untouched, chunk count unchanged (prune removes indexes, not chunks). Write proof CLOSED: 08:25:47 job errors -> 09:37:29 TASK OK, snapshot 2026-07-27T09:37:29Z, chunks 9787->9813, prune step absent entirely. Driven through POST /api/guest-backup/trigger (the UI path), not --selftest and not raw vzdump. Hub gauge evidence explicitly NOT satisfied - the delta is below its 0.1 GB display granularity. GC scheduled sun 04:30 UTC and deliberately NOT run: every chunk still carries a fresh atime from the migration copy, so a run today would reclaim nothing. verify-new enabled per operator ruling, turning an inert hub alarm live. Legacy demo-felhom-01 namespace deleted with its two ACL entries and its token (operator ruling, confirmed twice) so nothing dangles. R-89 records the target architecture and carries the unanswered parallel question: does the restic key on storage-box-pool-1 have DELETE rights? If so the daily app-data tier has the identical exposure and append-only is the equivalent answer. ep0 is Etc/UTC, not CEST - corrected in the record. |
||
|
|
2b24c70536 |
docs(ep0): hub PBS-DR capacity gauge verified correct after the volume move
The last open item from the datastore relocation. Hub operator UI (Offsite -> PBS DR) reports felhom-offsite (ep0) at 97.9 GB capacity, 12.6 GB used, 13% full - agreeing with the on-box df (98 G / 13 G / 13%). The gauge follows the datastore's CONFIGURED PATH, so the relocation required no hub-side change. RUNBOOK section 10.3 warned that a stale 37.2 GB reading would mean the gauge reads the wrong filesystem and would be a real bug worth a roadmap item - it does not, and there is no bug. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018ARoadHBf8rHoscfiqeVZn |
||
|
|
ad8057c4e3 |
docs(ep0): relocate the felhom-offsite PBS datastore onto the 100 GB volume
Supervised runbook execution. No code change, no version bump. felhom-offsite moved from ep0's 40 GB root disk (/srv/pbs-felhom) to a dedicated 100 GB Hetzner Cloud Volume (/mnt/pbs-datastore, ext4 -m 0, by-id fstab, relatime). Datastore NAME unchanged, so the PBS-DR descriptors, per-box storage ids, ACLs and namespaces are untouched. Capacity: 37.2 GB -> 98 GB total, 28.9% -> 13% used, headroom to the 80% warn 19 GB -> ~65 GB. This CLEARS the R-82 Phase 0 P0.3 STOP. Per-tenant encryption still precludes cross-customer dedup, so the slope is unchanged - the volume buys runway, not a better cost model. Verified: byte totals and chunk counts identical (9748), 7/7 snapshots across all three namespaces, backup:backup ownership, clean itemised dry-run, full verify job TASK OK with 0 errors, and a restore round-trip (source_tier pbs, pass true, mount_parity ok, clean teardown). Nothing deleted - the original 13 GB stays at /srv/pbs-felhom as the rollback until a new weekly backup lands. GC deliberately not run. Three findings recorded: - the `scratch` datastore points at a non-existent path (pre-existing; now logs ENOENT every start) - operator decision - the runbook's S6 guard test proves the wrong proposition: RequiresMountsFor re-mounts rather than refusing, so the test only bites when the device is genuinely unavailable (re-run that way, and the refusal was observed) - amendment recommended - S11: storage box u629193 has no live backup path, BUT ep0 carries an enabled sshfs mount unit against it that must be removed before the box is deleted Deviations: the volume arrived pre-formatted and mounted; S8 ran on demo-felhom rather than demo-hp (no SSH key for demo-hp); the window was contended by a stale in-memory 10-minute restore-test cadence whose config had already been reverted on disk. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018ARoadHBf8rHoscfiqeVZn |
||
|
|
9cfa619ec3 |
hub v0.74.0: allow local_api_endpoint_drift; R-77 docs + R-78/79/80
The allowlist entry is REQUIRED, not cosmetic: handleEvent 400s an unknown event_type, so controller v0.173.0's new drift alert would be silently inert without it. Shipped with the controller that emits it. Docs: - RUNBOOK-local-api-endpoint-drift.md — how to repair a drift, including the step everyone will want to skip (establish which value is CORRECT from what the agent is actually bound to, rather than assuming bootstrap.json wins) and what success looks like (SILENCE, not a "recovered" line, because a fresh controller's healthy first observation is not logged). Records both 2026-07-26 repairs. - ROADMAP: R-77 shipped; R-78 the local_api authority ruling, with the clobber-a-working-channel risk spelled out in BOTH directions so it is not resolved opportunistically; R-79 the whole-surface English-strings sweep; R-80 expected_backup_missed, flagged as likely outranking R-77 because 7.3 days of stale backup materially exceeds the ~1.5-day channel outage, so the causal link the DIAG hedged on cannot be the whole story. - Capability map: note against the drive-wizard row (every agent-backed capability rides this channel) that a silent drift class is now detected. NO row status flips — detection is not prevention. |
||
|
|
46504938ba |
R-50 Phase B: island migration runbook (B0) + drill validation (B1 PASS)
Idempotent LAN->island migration procedure with rollback table + abort criteria (firewall LAST). Validated verbatim on drill VM 300: rolled to r50pre, migrated, island /storage 200, LAN DNS held on the LAN IP (Finding-1 pin), apps healthy, hub reports 0.96.0. No rollback fired. |