091a4b744428ee7ca756001b48d8654c308fbbb1
240 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
091a4b7444 |
Correct the placement mis-framing, and file what we wrote down and never filed (R-368..R-375)
gates / gates (push) Successful in 16s
Documentation and survey only. No code, no machine contacted. THE CORRECTION. The 40 catalogue templates without a configurable path are not missing a choice: 01-topology-and-trust.md:150-152 classes each volume hot (DB/config/cache -> fast storage, ENFORCED) or bulk (media/files), and the 40 are all-hot apps. The deploy page has been saying so to the customer all along (deploy.html:624-625). SPEC-app-data-placement and R-352 are corrected in place with the framing MARKED, not deleted; every measurement stands. R-356 was re-checked and survives, strengthened - an absent HDD_PATH is the normal state, so reading it as "not installed" misreads a correct configuration. The disk claim, precisely: since R-165 there is ONE guest data volume with two binds, not two volumes (build-golden.sh:29-40, 99). A physical-disk failure losing data and first-tier copy together is REAL and is what the other tiers exist for. A full data volume stopping the OS is NOT real and was the overstated one. THE SWEEP. 113 survey-class documents examined, 14 statements of "not filed", 2 already filed. Its positive control convicted the sweep itself twice before it convicted the corpus - markdown bold broke the strongest pattern, and the reporter re-searched a truncated line - both false zeros of the exact class being hunted, and together worth 2 of the 14. THE HEADLINE. The gap the 2026-08-21 drill rediscovered WAS filed - as R-107, ROADMAP.md:122, M/READY, 2026-07-28 - and is absent from OPEN-ITEMS.md, which calls itself the single source of truth. OPEN-ITEMS and that rule both landed 2026-07-27; R-107 went to ROADMAP alone the day after. 72 ids live only in ROADMAP, 29 not done, some of them findings. Filed as R-369 (HIGH). Five more still-open gaps filed with their ages: R-371 (17d), R-372 (38d, the oldest), R-373 (20d), R-374 (14d), R-375 (4d). R-368 corrects Part 4: the storage default IS applied at deploy time via deploy.html:612 - the earlier "the deploy route never reads it" came from grepping Go and never the templates. R-370 records the process failure and is closed by the template change. PROMPT-TEMPLATE gains the two rules it lacked: name the architecture document for the area and say what it says (with a file->area map and the test "is this something we chose?"), and an enumerated gap becomes a register row in the same session - a ROADMAP row alone does not count. Ceiling R-367 -> R-375. |
||
|
|
877fcd2a38 |
R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed
gates / gates (push) Successful in 16s
Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with a negative control first — the same planted, hash-recorded fixture run through the same steps on both builds. R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the off-site copy or the restore; and because the same wrong name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label. Sweep proven able to convict before its count was trusted: one affected app of 53. R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit, before the database and inside the stopped window, and VolumesReplayed reaches the sentence. The half-false comment beside the skip is corrected and the half that still holds is named. Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b, verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are the operator's decision, and raising the floor is what puts this on demo-felhom, which is still on 0.217.0 and still has both defects. R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them (an existing guard), they are adoptable by hand, and doing it automatically would be a migration. Ceiling R-366 -> R-367. |
||
|
|
7064596c2e |
DRILL closeout: the scheduled cycle agrees, the abandonment sweep watched firing, R-366
gates / gates (push) Successful in 16s
The box's own 02:30 / 03:30 / 04:15 cycle ran unattended and agrees with the manual one on every measure. The one that matters: R-355 is not an artefact of manual triggering — the scheduled run again wrote paperless-ngx's PostgreSQL dump into a directory for a stack that does not exist, and again left the app's own unit recording db_dumps: null. Part 4.2's terminal deletion has now been observed. At 05:10 the sweep removed exactly the recorded set-aside store and left the live repository untouched; at 05:13 the hub dropped the sealed package that protected it and said so (event 3025). Both halves went together, three minutes apart, and the controller cleared its own state. demo-felhom's two preserved fixtures were verified untouched throughout. R-366 (HIGH) filed, found incidentally: the 21 August reinstall orphaned demo-hp's PBS whole-guest archives as well as its restic repo, so a rebuilt box loses BOTH off-premises tiers at once. The restore-test caught it and named the key mismatch precisely; it is merely called "a failed restore test" rather than "your older backups are unreadable". demo-hp is left HEALTHY, not broken. Ceiling R-353 -> R-366. |
||
|
|
f5a4fceeeb |
DRILL 2026-08-21: the off-site restore never replays named volumes (R-354..R-365)
gates / gates (push) Successful in 16s
Diagnostic only — no code changed, no version bumped, nothing deployed.
The verdict is a mixture. The unit and the off-site snapshot HOLD the data, proven
by identity in both storage classes including two Hungarian accented filenames. The
loss is in the last leg: ReconstituteFromOffsite skips every isUnit placement and the
volume tars live inside the unit, so the off-site full restore has no named-volume
leg at all — while the local restore-from-unit does, and returned the same tar
byte-identical minutes later.
Twelve rows opened, ceiling R-353 -> R-365. Three HIGH:
R-354 off-site restore never replays volume dumps
R-355 paperless-ngx's Postgres is dumped under a non-existent stack, so its unit
has no DB dump, no safety dump is taken, and the customer is told it has none
R-356 the off-site restore refuses for all 40 no-drive apps saying the running app
"is not installed", with a remedy those apps make impossible
R-353's instruction (2) is satisfied and annotated: the 40-class DOES reach the
off-site tier. Its instruction (1) stands and is now larger. R-329 confirmed still
live and now the only bad-severity emit fleet-wide.
Evidence: documentation/audits/DRILL-backup-truth-2026-08-21/evidence/
|
||
|
|
910fd91124 |
agent 0.130.0 published and vouched; R-347 closed, R-349 + R-350 filed
gates / gates (push) Successful in 14s
Released via scripts/release-agent.sh: tag v0.130.0 at 7569f34, sha256
a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3,
verified by independent download and reproducible byte for byte with
-trimpath -buildvcs=false.
Vouched agent 0.129.0 -> 0.130.0 in the Day-0 manifest. Only the agent
fields changed: min_agent stays 0.129.0 because it states what the GOLDEN
CONTROLLER requires, and raising it would have HELD the floor for every
box below 0.130.0. Global floor untouched at 0.216.0 -- and on hub
v0.106.0 it is a separate form with its own action, so publish-train
rule 2's hazard no longer exists in the shape its incident describes.
No --no-verify: the CHANGELOG heading was flipped only after the tag and
package existed, so release-complete passes on the real artifact.
R-349: the fleet was running a DIFFERENT binary under the same version
name -- the proof deploy was a hand build, the release is -trimpath.
Self-update could never have corrected it, because every version check
compares the string. Both boxes reinstalled from the downloaded package.
The proper fix exists in miniature as wrapper_sha256 and was never
extended to the agent's own binary.
R-350: I printed the hub password into the session transcript via
curl -w '%{redirect_url}' -- the hub answers 303 and curl re-attaches the
credential. Not in git, not in any committed file, not in the evidence
directory. Rotation is the operator's call.
ep0 closes at fd 17, ESTAB 0, CLOSE-WAIT 0 -- its t0 baseline -- and was
read-only for this entire arc.
|
||
|
|
57dd62b097 |
R-344 fixed and proven on both boxes: ep0 is back to fd 17 from 415
gates / gates (push) Successful in 14s
P1, outcome (i) in one second: replacing the agent on demo-hp released exactly its 199 established connections (ep0 fd 415 -> 216). CLOSE-WAIT stayed 0, so outcome (ii) does not exist and gets no row -- ep0 reaps on peer FIN correctly, and the 543 CLOSE-WAIT at the 08-18 wedge has another explanation. P2, 1.03 h (operator closed the >=4 h window early, so no daily rate is extrapolated): control +4, fixed +0, with each box making exactly 4 /snapshots and 4 /version calls. Same cadence, same work: 4 cycles -> 4 leaks vs 4 cycles -> 0. The fixed box's cycles are in ep0's log, so the zero is the fix and not a stopped agent. P3: the second box took ep0 from 220 to 17 fd in under two seconds. 17 is precisely the t0 baseline of 2026-08-18 09:51:22Z. Corrects a claim this session made earlier the same day: the accumulated descriptors did NOT need an ep0 proxy restart. They were held on both sides. ep0 was read-only throughout; its PID never changed. R-344 updated and left OPEN (unpublished is not delivered). R-336 re-scoped -- its old next-step would have fixed nothing while looking like a failed fix, and it is now a scaling row (~25 req/s at fifty customers). R-347 filed for the delivery gap (Viktor decides). R-348 filed: an agent restart blanks the reported backup list for ~18 h and the Store comment calls it unaffected -- blinds no alarm, checked not assumed. |
||
|
|
9299f85c4b |
SPIKE ep0 connections: CI green by run id (360/237, 19672e685)
gates / gates (push) Successful in 14s
|
||
|
|
19672e685e |
SPIKE ep0 connections: the leak is felhom-agent's, not the poll rate
gates / gates (push) Successful in 14s
R-341's first dated check, taken at +46.2 h: fd 17 -> 405 over 166,251 s
= 201.6/day. Pre-registered range was 370-450; observed 388. UNCHANGED,
as predicted. CLOSE-WAIT is 0 -- absent entirely, not merely flat.
Q1: exactly two peers, 194 each, no third party.
Q2: outcome (a). ep0 388 = 194 + 194 on the boxes, twice, and the four
new sockets carry the same source ports on both sides. 0 closed in 31 min.
The finding: all 388 are held by felhom-agent. pvestatd and
proxmox-backup-client made 162,404 requests and leaked zero. Mechanism is
a per-cycle http.Transport with a zero-value IdleConnTimeout that nothing
ever closes (internal/pbs/client.go:56, main.go:1486). R-336's premise
does not survive this -- cutting the poll rate would have fixed nothing.
Q3 NOT measured: Phase C held at STOP 1, prediction pre-registered first.
Part 0 captures the due-checks gate's first conviction on a real overdue
date (rc=1, names R-341, sole failure among 10 gates). Row cleared at
Part 4, after the result was recorded in R-341, not to make a push work.
New: R-344 (the transport leak), R-345 (hub/Makefile pushes :latest),
R-346 (ActiveEnterTimestamp reads 5h56m early -- NRestarts is still 0).
|
||
|
|
3e50902a98 |
RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
gates / gates (push) Failing after 14s
Both STOPs cleared by the operator. No code changed; documentation only.
STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.
STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.
STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.
UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.
SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.
CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
- the "~85/day, ~2 years of runway" figures were WRONG. They came from a
single 17-minute window with a delta of ONE descriptor. Real rate is
183-200/day over two independent windows; runway ~357 days, not 2 years.
- the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
543 CLOSE-WAIT. R-336's fix must target unreaped connections.
- "proxmox-backup-api" reported inactive during verification; that unit does
not exist. Bad query, not a fault, written down because it looked like one.
R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.
golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
|
||
|
|
435e044cf1 |
INCIDENT/R-336: the leak is measured live, not assumed
gates / gates (push) Failing after 15s
Post-restart baseline on ep0: 19 fds at 16m43s (from 18), 1 CLOSE-WAIT, Recv-Q 0. One descriptor per ~17 min is ~85/day, which agrees with the ~73/day implied independently by the failure itself (1016 sockets over 14 days of uptime). Two estimates of the same slope agreeing turns "the ceiling raise is mitigation, not a cure" from a plausible claim into a measured one, and puts the next ceiling at ~2 years instead of a fortnight. Recorded because standing rule 3 asks for a positive observable: this is it, and it fired. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016p1PTCzb8rF5G9Aa1qhpBN |
||
|
|
ebfd0967c1 |
INCIDENT + registers: ep0's PBS proxy served nobody for 9.5h (R-336..R-338)
gates / gates (push) Failing after 12s
Two whole_guest_backup_failed alerts at 04:30 and 04:32 CEST were one incident, and not on either customer box: ep0's proxmox-backup-proxy was active, holding its listening socket, and accepting nothing. Root cause: accept() returning EMFILE. The process held exactly 1024 fds -- its systemd-default soft RLIMIT_NOFILE -- of which 1016 were sockets and 547 connections sat in CLOSE-WAIT. The 1024-deep accept backlog had overflowed (Recv-Q 1025), so every client timed out. It was wedged from its own loopback too, which is what moved this from a network problem to a process problem. Fed by ~85k requests/day (a flat 3,538/hour) against an endpoint written to weekly, that leak reached the ceiling in 14 days of uptime. Fix: LimitNOFILE=65536 drop-ins for both PBS units, restart, verified from both boxes (200 in ~0.1s, felhom-pbs active), then re-drove the missed backups through the product path -- POST /backup?target=felhom-pbs on each agent's local API, not a hand-run vzdump. demo-felhom ct/9201/2026-08-18T03:57:43Z 4.10 GB 36.4s demo-hp ct/9201/2026-08-18T03:58:43Z 4.29 GB 41.5s Both host reports now carry felhom-pbs success=true, so the hub is green on the evidence rather than on a restart having been performed. No data lost, no backup skipped: the daily local tier was never affected and the PBS tier is weekly, so the window cost exactly one attempt. Evidence copied off ep0 BEFORE the restart, per standing rule 5. Filed: R-336 (the ~1 req/s poll rate is the real defect; the raised ceiling is mitigation, not a cure), R-337 (a status endpoint that trailed its own artifact by minutes then caught up -- WATCHING, downgraded from the defect I first wrote, because it self-corrected), R-338 (demo-hp is not on the R-50 island at all and nodes.md says it is; its local API is bound to the customer LAN). R-334 updated: still open, now one version wider (controller 0.216.0 vs golden 0.214.0). golden-currency is the only failing gate and is inherited -- it reads files this session did not touch -- so this push used --no-verify, stated per .claude/rules/gates.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016p1PTCzb8rF5G9Aa1qhpBN |
||
|
|
848de8153d |
docs(audits): first genuinely failing disk — the SMART PASSED trap
gates / gates (push) Successful in 12s
Commit the raw evidence from ST3000VX010 S/N Z6A07P2G (/dev/sdg on DooPlex), which went 8 -> 352 unreadable sectors 11-13 Aug while smart_status.passed stayed true throughout. - fixtures/smart-ST3000VX010-failing-2026-08-14.json: raw smartctl -a -j, verbatim - fixtures/smartd-history-sdg-2026-08-14.txt: 406 smartd journal lines, 11-14 Aug - DIAG-smart-passed-trap-2026-08-14.md: the mechanism (attrs 187/197/198 all carry thresh 0, so a normalized value that floors at 1 can never fail the overall verdict on unreadable sectors), the non-monotonic timeline, the three controller defects with locators, and the counterfactual: zero emails would have been sent. |
||
|
|
e0b56c976f |
REPORT + CONTEXT: the third name, the second door, and a number that answered a different question
gates / gates (push) Successful in 15s
Three rules carried forward. A name must separate on the STEM, not the noun — naming this secret after the act it is used in would have recreated the trap, because the other factor on the same page is the „Párosító kód". A guard is worth what its positive control is worth: this one's selftest convicted its own step-3 case and found a defect in the guard itself. And a suppression must rest on the machine's own declaration, then be checked for the SECOND door — recording the disabled state rather than deleting it is what let the deadline check skip it too. Yesterday's report is preserved to audits/ because it carries the only record of the self-heal verdict (Part C was dropped, so that reasoning is in no register row) — the rule written last night, applied to itself the first time it mattered. |
||
|
|
4d6ec7c7bb |
hub v0.104.0: the guest network gets a reader (R-319), and the hub half of the naming (R-295)
gates / gates (push) Successful in 14s
Four paper debts and one fact given a reader. Hub-only — nothing to bake. A4 — the entry about "the tester's machine" named a risk correctly and labelled it in a way that invited deleting it. Established from the hub's own store: `peti-felhom` is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47 this morning. The prompt's premise conflated the two; the register now says which is which. A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out of STATUS's "Waiting on you", which is now empty. A3 — day0-install §C.1 said pushing the installer publishes it. It has not since R-110. Corrected, with the two manifest pins named and an outside-verification command; the one copy that repeated it (a dated audit, true when written) carries a superseded note. A5 — standing rule 5: evidence comes off the machine at the end of the phase that produced it, before any revert. Earned twice in three days on the same box at the same point (R-320). Four homes, plus what to do when it is already gone. R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll` mail kind so the mail names the page a REBUILT box actually shows („A szerver beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach. Naming only; the acceptance pin proves the secret is untouched. R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never drawn as healthy — three absences, three sentences. No alarm, deliberately. Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190. B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now. Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled` reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed. Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect). |
||
|
|
2d05b29b82 |
REPORT: installer v1.28.0 published and verified live; evidence, and the Part 1 logs I lost again
gates / gates (push) Successful in 13s
|
||
|
|
c1319a91a8 |
Correct the drill's wall clock to the measured end time (17:45, not the estimated 17:55)
gates / gates (push) Successful in 19s
|
||
|
|
1d5f2b8bb6 |
DRILL: the retained key works, and the customer cannot reach it
gates / gates (push) Successful in 23s
Three verdicts, kept separate because collapsing them is how this assumption survived a week. (a) The material IS retained. host_escrow_superseded id 11 is the first retained row in fleet history to carry identity_blob (572 B), byte-identical to the pre-supersession row (sha256 a10032341c8584ed...). (b) The retained material DOES open the old store. Unsealed with the old recovery code it yielded a password byte-identical to the pre-change one, and restored three planted files byte-identical from a store the box itself could no longer open - including a Hungarian accented filename verified as raw bytes. Negative control ran first and failed closed. (c) The customer has NO route, and is misinformed. ListSupersededEscrow has zero production callers; the recovery path selects FROM host_escrow. Asked with the code that had just worked by hand, the product answered "the recovery code did not open the sealed bundle". A valid code for retained history is reported as a bad code - the R-224 class again. R-304, rank 1. Both installer faults were watched happening first, so installer-v1.27.0 is now published (tag + both webpage.yaml refs). Pre-fix: the box came up on controller 0.98.3 against a vouched 0.213.0, below the floor and below the version carrying the recovery screen; and our own uninstall left dnsmasq on 0.0.0.0:53 so our own next install refused. R-297 and R-300 CLOSED. Also filed R-305 (the dnsmasq fix fires once per machine - the leftover returns on the second reinstall, proven), R-306 (--preflight-only writes state it says it does not), R-307 (a live abandon countdown on demo-felhom, firing 2026-08-24 - operator decision), R-308 (stored controller password stale), R-309 (the day-0 runbook's publication claim has been false since R-110), R-310 (two edges). Ceiling R-303 -> R-310. Capability map moved: the retention claim is now marked operator-only. Phase A logs did not survive the intermediate revert; recorded. |
||
|
|
1d6f1c522d |
Rehearsal 2026-08-09 COMPLETE: data BYTE-IDENTICAL, journey needs a shell twice
gates / gates (push) Successful in 23s
The walk finished. All four planted files came back byte-identical out of snapshot 41c830db, including two Hungarian accented filenames verified as RAW NAME BYTES (NFC preserved) — the discriminator the Gate 0 positive control was built for, having been watched failing on an NFC->NFD rename that renders the same. Unlock 21s, restore 13.2s. It finished only because a terminal was available twice: - R-273 CLOSED. v0.128.0 was published as a package and never git-tagged, so every install died at 5/8. Tag pushed on operator instruction after an INDEPENDENT download proved the package sha equalled the vouched value; --resume then reached Day-0 SUCCESS in 3m49s on controller 0.210.0. The two guards that would stop the class recurring are still owed. - R-280 NEW, rank 1. A reinstalled box cannot re-attach its own data drive by any dashboard route: /api/disks/candidates returns empty because both lists are built from the UNCLAIMED-disk scan, and the drive is claimed precisely because it is also the backup target. Correct for "initialise", over-broad for "attach", which is non-destructive by definition. The restore page meanwhile says "Ez ket kattintas" and points at that empty list. Cleared by POSTing /mnt/sys_drive — an internal path no household could produce. Also new: R-281 the hub said NOTHING through the entire reinstall and the sealed-backup tripwire did not fire on a real unseal (positive control: 2 events all day fleet-wide); R-282 one code with three names and a mail pointing at a page the box does not show; R-283 hub reads "Claimed 18d ago" while the box serves its setup page; R-284 "almost full" over a 93%-free store. R-274 NARROWED by measurement rather than left as written: the resume path fetched the vouched golden correctly, because --resume skips the preflight that does local discovery. What survives is real — discovery is sort|tail -1 with no manifest comparison — but a FRESH install taking a stale golden is still not observed, and the row says so. Two of my own claims were refuted by test and are recorded as refuted, not quietly dropped: the leftover sudoers file is inert (sudo skips dotted names), and demo-hp's off-site tier was healthy all along. |
||
|
|
b1afbb8a4d |
Rehearsal 2026-08-09: the walk stops at P3 — R-273 blocks every install fleet-wide
gates / gates (push) Successful in 24s
P1 uninstall, P2 preflight, P3 install. The install FAILED at step 5/8 in 44s, and the two rank-1 findings are both on the setting-up path a tester's visit is made of. Eleven register rows minted (R-269..R-279); ceiling moves 268 -> 279. R-273 (RANK 1) — the hub vouches agent 0.128.0; that version was published as a Gitea PACKAGE but never git-tagged. Since R-183 the installer correctly pins its config fetches to raw/tag/v<vouched>, so every fresh install and every reinstall now 404s as root, mid-install. Measured: main 200, v0.127.0 200, v0.128.0 404. This is R-184 arriving; release-agent.sh:23 already documents the exact hazard. Existing boxes are fine (self-update takes the binary from the registry). NOT fixed here — publishing a release tag is outward-facing and the runbook says stop and report. One command unblocks it; it is in STATUS.md. R-272 (RANK 1) — Felhom's own uninstall leaves the condition that makes Felhom's own reinstall refuse. It installs dnsmasq at day-0, then on teardown removes the snippet and RESTARTS the daemon unconstrained (process start time lands inside the uninstall window), which grabs 0.0.0.0:53; the next preflight then refuses, and the message reads as though the owner's LAN DNS is at fault. R-274 — a local golden is adopted with no version and no sha check; the manifest vouch is consulted only on the fetch path. demo-hp's local copy is controller 0.192.0 against a vouched 0.210.0, and below the 0.200.0 where the recovery screen shipped. Not yet observed end-to-end (R-273 killed step 5 first). Also: R-275 orphaned credential backups + uid reuse, R-276 the wg tunnel outlives the uninstall, R-269/270/271 from the token rotation, R-277 three hub surfaces misreport a healthy off-site tier, R-278 demo-felhom six days unprotected, R-279 no operator-triggerable off-site run. Two hypotheses of mine were tested and REFUTED rather than shipped as findings: the leftover sudoers file is inert (sudo skips dotted filenames), and demo-hp's off-site tier was healthy all along - I had misread the hub and said so. STATUS.md records the three rulings §8.3 asked for, with the floor CORRECTED to its live value 0.200.0 and the count corrected to twenty. |
||
|
|
34646295dc |
Rehearsal 2026-08-09: pre-phase + Gate 0 recorded before the destructive walk
gates / gates (push) Successful in 29s
Venue demo-hp, operator-approved at STOP 1. Records the state that P1 destroys,
plus seven pre-walk findings, while they can still be checked against a live box.
R-268 CLOSED — the leaked per-guest local-API token is rotated and the rotation
is PROVEN in both directions (old refused, new accepted, channel up with a
positive observable). Rotating it surfaced three defects:
- an out-of-process rotation does NOT revoke the old token. The daemon serves
Lookup from a stale index and re-reads only on a MISS, so a superseded token
is a direct hit. Red-proved in a unit probe AND live on hardware; the shipped
RemintCoherence test passes only because it looks up the NEW token first.
- R-268's own recipe is incomplete: ensureLocalAPI returns early on a present
local_api block, so writing bootstrap.json is not enough — the controller
serves the old token from controller.yaml across restarts.
- the agent-channel alarm never closes: the UP branch does not notify from an
unseeded state, and the alarm's own remedy ("re-bootstrap") resets it.
Gate 0 complete: dataset planted in the Calibre library (coverage verified, not
assumed) with two Hungarian accented filenames; the comparator watched FAILING
three ways including an NFC->NFD rename that renders identically; off-site run
driven through the product's own button; restore point recorded by identity as
snapshot 41c830db, confirmed to carry all four files.
Also corrects the record: demo-hp's off-site tier is HEALTHY. Three hub surfaces
agreed it was absent and all three mislead — the panel showing 0 snapshots renders
the LOCAL tier, 162 KB rounds to 0.0 GB, and a two-day-old stuck event reads as
current. And the managed-update floor is live at 0.200.0, not 0.156.0.
|
||
|
|
b7fb2117ae |
CAMPAIGN 12 — the class sweep: golden 0.208.0 baked (awaiting vouch), R-256..R-263 filed, gating ranked
gates / gates (push) Successful in 20s
Part 1. Golden 0.208.0 baked on the drill VM, published and ROUND-TRIP VERIFIED — 656 150 362 B, sha256 ba668f59..5ffb82, and ./etc/felhom-controller-image read OUT of the downloaded archive says felhom-controller:0.208.0. Acceptance markers all green (overlay2 x1, mount points x2 rootfs+mp0, upload HTTP 201 x1, excluding/FATAL/mp1 x0), Result=success. Token file->file, read inside the VM; systemctl show grep = 0; committed-log grep = 0 WITH a control proving the grep works. Bake VM destroyed, drill disk restored to virgin. NOT VOUCHED — the campaign halts there deliberately. golden_currency_gate.py was correctly RED on arrival and is green after the bake. No --no-verify was needed anywhere in this session. Parts 2-4. Seven defect classes swept for siblings by class rather than by feature. Analysis only: no product code, nothing deployed, no machine touched beyond the bake VM. Eight new rows R-256..R-263 (ceiling moved from R-255), grouped by class in OPEN-ITEMS.md. C1 produced no new instance and has no row. The sharpest is R-260: the agent reports operator_key_configured every heartbeat, the hub has no field for it, so the check that answers "can the operator get into this box" returns ok for a box with no operator key installed. Every class states whether its method re-found the known instances, because a method that cannot re-find them has not been shown to work: C1 2/3 (verified by replaying the pre-fix templates), C2 2/2, C3 2/3 + 1 as fixed, C4 fix-pattern re-found, C5 re-found, C6 deadcode 0/2 and bespoke 1/2, C7 weakest and said so. Blind spots stated per class; seven suspicions investigated and DISPROVED, including two of my own methods. Part 4's ranking is in ROADMAP.md as G-1..G-8. Gate C5 (cross-repo tag reachability — cheap, --fast-eligible, would have caught every R-260 instance on the introducing commit). Do NOT gate C6: golang.org/x/tools/cmd/deadcode was measured against a PLANTED probe and is blind to unreachable METHODS on widely-used types, which is exactly the shape both known instances have. R-242's untouched half is recorded, not built: this bake demonstrated it, the currency gate flipping green the moment the evidence DIRECTORY existed, before the round trip finished and with no vouch near it. Correction the campaign owed its own brief: escrow_stale was described as closed; it is R-247 and READY. The live repo is the source. Sampled rather than swept, exactly: C7 60 of 2652 production invariant comments and NONE of the 1440 test comments (that half is owed); C2 19 of 221 refusals; C3/C4 controller only. No finding was reproduced live. STATUS.md is 100 lines against its 93-line one screen. |
||
|
|
71c43f87c2 |
SPIKE R-241 — ruled: a MINTING defect, not a screen-predicate defect
gates / gates (push) Successful in 6s
Read-only spike on the standing venue (demo-hp VM 324 finalwalk-appliance),
read before anything else touched the fleet. No product code written.
THE RULING. The recovery screen was telling the truth. There genuinely was
nothing recoverable under the key the box held, because the box minted that
key itself over the top of a sealed package it already knew the hub held.
Fixing the predicate would paper over a box quietly making its own history
unopenable.
Three measurements carry it:
1. WriteOffboxSecrets (offbox.go:411) mints on ONE input - does the file
exist. Its two neighbours in the same file both consult
GetHubEscrowIdentityPresent(). Same fact on three paths, used on two.
2. The flag was the PRECONDITION of the chain that reached the minting:
the retry job only logs when the declaration is live, and the venue
logged it at 02:48:03Z - 30 minutes and six ticks before the mint.
3. The box KNEW and discarded it: at 03:28:03Z, 35 minutes before the
customer looked, escrow_confirm.go:154 computed the exact
discriminator and logged the mismatch. Never persisted, never surfaced.
And the hub explicitly disclaims doing this - offsiteheal's package doc:
'credential automatic, key customer-present ... must not quietly widen'.
Also answered: Q2 (shape (b) is structurally unreachable - the escrow gate
sits upstream of the only orphan producer, with a positive control that the
scheduler was alive); Q7 (the 'create a new code' button does NOT destroy
the data - R-198's retention holds - but it converts a self-service recovery
into one needing an unbuilt read path, and it re-enables the screen while
invalidating the code that screen accepts).
New rows:
R-242 - a release that changes customer-visible behaviour is not delivered
until a golden carries it, and nothing enforces that. Recorded
with three proposed shapes; deliberately NOT built.
R-243 - a box in this state silently stops backing up and NO alarm fires:
three individually-correct exclusions leave one state unobserved.
Highest register ID moves R-241 -> R-243.
|
||
|
|
0c4411e54b |
R-201 re-walk: the data PASSES again, the journey still FAILS — two dead ends, down from four
gates / gates (push) Successful in 9s
Asked Campaign 11 Phase 1's question a second time, on the fixed build, on a
NEW appliance (VM 322, customer rewalk). The Campaign 11 venue was untouched.
THE DATA: PASS. All three sentinels byte-identical out of the pre-destruction
snapshot a7bc23bd in 23s through the customer's own restore flow — including a
12 MB binary and an accented Hungarian filename whose NAME BYTES are identical
too (verified as hex, not as rendered text).
THE JOURNEY: FAIL, two dead ends against Phase 1's four.
1. R-218's CONSUME half. The hub re-staged the credential at 11:44:57 saying
'the box re-consumes on its next cycle'; a full cycle ran at 11:55:46/54
(with a positive control that it ran) and it did not. A census of the
customer-reachable actions found none that fetches it. Only a command line
INSIDE THE GUEST moved it — 18s, confirming nothing was wrong with the
credential, target or key: only the trigger. R-218's row said SHIPPED and
over-claimed; it is corrected to REOPENED for the consume half.
2. R-220. Drives still unenrollable after a rebuild, needing a Proxmox-host
unmount; without it no app redeploys and the restore page stays empty.
Unaided RTO STILL UNDEFINED. Attended: +45s key placed, +24m12s tier up,
+30m13s data verified. The 30m must not be quoted as the customer number.
What passed and is new: the recovery screen appeared WITHOUT being sought,
answered all three questions with a seal date matching the hub exactly, the
emailed reset code worked first try, the unlock was a real 1.528s unseal, and
R-225's fix was seen working in the wild (unknown, not a false zero).
R-216 part 4 reproduced live: the reinstall downgraded the hand-installed agent
0.126.0 -> 0.125.0.
DELIVERY GAP recorded as owed and NOT conflated with the journey: a fresh
install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions,
neither carrying the fixes — installed by hand. Nothing was vouched.
Capability map row STAYS FAIL. Campaign 11 doc gets a dated ADDENDUM, not a
rewrite.
|
||
|
|
a1a542b9a7 |
RECON: what DooPlex's own backup actually protects (R-232) — read-only, nothing changed
gates / gates (push) Successful in 9s
Five sets, 14/14 successful runs in 14 days, and a file restored byte-identical from the data repo -- the first time a restore has ever been demonstrated here. What the failure matrix exposes: notify_failure is a no-op (NOTIFY_ON_FAILURE=true, webhook commented out) so a failed backup tells nobody; NOTHING leaves the box (Longhorn's target is DooPlex itself over NFS, the only outbound-looking cron pulls inbound for another project); the backup tree is a single writable path with no append-only; two sets are same-disk (.claude-memory and the postgres dumps, whose source sits inside the backup tree); Longhorn retains 1 generation; secrets/restic-repo has never held a snapshot because backup-secrets.sh never calls restic; and /opt/backup/docs/BACKUP-RESTORE.md does not exist though the unit advertises it. NOT a finding: the restic passphrase. On-box copy is on a different disk from the backups and the operator holds an offline copy out of band, so disk loss is recoverable. The residual is only that it is operator-held rather than system-held escrow. A stale lock in the data repo is MINE, from the previous task's verification. Fenced from clearing it; recorded with its impact and the one command that fixes it. |
||
|
|
5ca5082e7c |
docs: close out the instruction arc — t740 corrected on evidence, registers, ledger, S-37
gates / gates (push) Successful in 8s
target-selection.md said demo-hp has no off-site tier. Measured first: pvesm list felhom-pbs on the box returns two snapshots in demo-hp's OWN namespace (2026-07-28, 2026-08-04) against ep0's felhom-offsite. The claim was TRUE WHEN WRITTEN and went stale when F10 resolved 2026-07-23. The measurement is kept in an HTML comment beside the corrected sentence. This file decides which machine may be destroyed, so the sentence was load-bearing, not cosmetic. R-229(b) CLOSED (agent 175 -> 99 eff). R-230(b) CLOSED (symlink, proven from fresh sessions). R-230(a) part-actioned -- three false statements fixed, WARN loop added, bulk ruling still owed. S-37: a claim in an instruction file is checked, not trusted. |
||
|
|
92a076c239 |
docs: backup verified IN the repository — snapshot b587f775, 118/118 + 40/40 (R-229)
gates / gates (push) Successful in 8s
58,158 files / 405.865 GiB in 27:56; 994 MiB added, 77 MiB stored -- the one-time full re-read cost I/O, not storage, as predicted. Also records a false alarm I raised against my own instrument: `restic ls <snapshot> <path>` reported 0 of 40 archived memories because restic 0.18.0's path filter does not recurse into subdirectories. The unfiltered listing shows all 40. An instrument that can drop results silently is not a measurement -- and this one was measuring a safety property. |
||
|
|
a0b50eb48f |
docs: R-168 is CLOSED — correct the stale "CI is still owed" sentence carried into the trim (R-229)
gates / gates (push) Successful in 8s
Found by the checklist's own last item (confirm your push's CI run by run ID). Corrected here and in the three sibling repos. Ledger J2 records the two lessons: a trim is a VOLUME operation and carries stale claims forward unless each is re-checked, and the gate cannot catch this class -- "this register item is closed" is not mechanically checkable from the instruction file. |
||
|
|
9c10c5d600 |
docs: ledger, registers and standing ruling S-36 for the part-2 instruction/memory work (R-229)
gates / gates (push) Successful in 8s
R-229 legs (a) and (c) CLOSED. R-230 opened for what this deliberately left -- the ruling owed on auto-written staleness, the symlink decision, the spec-as-failing-test pilot. R-231 opened for /opt/backup/scripts/ being unversioned host state, found while adding the memory store to the backup set: the same class this task was closing, inside the fix for it. S-36: the auto-memory store is backed up, never committed; no memory file is ever deleted; a rule file created mid-session is NOT live in that session. Second session in this repo, so REPORT-<topic>.md per CLAUDE.md; shared REPORT.md untouched. |
||
|
|
c21bcf84f7 |
docs+gate: instruction files cannot silently regrow (R-229)
gates / gates (push) Successful in 7s
New shared scripts/instructions_gate.py, registered in controller_gates.py and agent_gates.py, never copied into a sibling repo (the reuse_refs_check.py precedent). 20 fixture tests, all asserting the effect: exit code AND that the message names the file and the reason. It is a consistency gate, not a budget gate, and the failure message says so. A /context reading measured the instruction files at 15k tokens against 869k free in a 1M window -- space is not the constraint, and a future reader must not re-derive the wrong reason. The 200-line ceiling is adherence guidance; a file nobody can hold in their head is where contradictions hide, and five were found here. Checks run against effective text (HTML comments stripped, because they are stripped before injection): the line ceiling; every .claude/rules/*.md declares paths: or an explicit unconditional: true; no component version literal; no TEMPORARY block carrying a past date; and the workspace-root CLAUDE.md is byte-identical to its versioned copy -- the live file sits outside any git repo, so that copy is its only version-controlled record. Two traps recorded so they are not reintroduced: a bare \d+\.\d+\.\d+ matches the first three octets of every IPv4 (the gate excludes dotted quads, or it fails on 192.168.0.180 in the agent's own file); and unconditional: true is NOT a Claude Code feature but this project's own marker. Workspace-root CLAUDE.md 208 -> 182 lines (142 effective), copy kept identical. The nine-instance invariant table moved into the felhom-testing skill, which triggers when writing or reviewing a test; all three directive bullets stayed in the core. felhom.eu/CLAUDE.md got surgical corrections only and is knowingly still over the ceiling at 227 effective lines -- closing it needs the restructure R-229 defers, said plainly rather than quietly absorbed. CONTEXT.md gains standing ruling S-35. OPEN-ITEMS.md gains R-229. Docs only -- no Go, no version bump, nothing built or deployed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JJc8sAGRWmavP3rMtdpkr2 |
||
|
|
d30c2a51ed |
R-224..R-228 CLOSED: registers, capability map, campaign annotation, STATUS
gates / gates (push) Successful in 7s
Five closed in controller v0.202.0 + agent v0.126.0, each with its live or red-proof evidence in the row. Five explicitly still open and named as such rather than left to inference: R-214, R-220, R-221, R-213, R-202 — and R-220 is flagged as currently worked around BY HAND on the campaign venue, which is the only reason an app could be deployed there. The capability map's recovery row STAYS FAIL and says why: fixes are not a re-walk, nothing walked a customer end to end, and the customer-facing messages were NOT re-driven live because /recovery correctly retires itself once the old data is set aside — restoring that state is the reconfiguration the task forbade. The campaign document is ANNOTATED, not rewritten: it records what was true when it ran, and that is its value. workspace-CLAUDE.md gains comment-vs-code entry 9 — the escrow header said the errors were 'DISTINCT on purpose' and named THREE situations while a fourth was folded into one of them, and a green test named the defect and did not prevent it because it asserted a STRING one layer below the merge. ROADMAP needed no collapse — it carries no rows for these IDs. |
||
|
|
9c1d05d360 |
CAMPAIGN-11: hygiene, what-did-not-run, venue end state, and the session report
gates / gates (push) Successful in 8s
The recovery codes are shredded with the plant->find->shred->fail-to-find
control the brief asks for, and THE CONTROL PAID FOR ITSELF ON ITS FIRST RUN:
it found the Phase 0 code in ~/.config/credentials as R_CAMPAIGN_11 — a copy
this session did not create and would never have looked for. Without it, a
'codes shredded' claim would have been false. That key was removed from the
shared file with a verified diff (every other line identical, nine keys intact)
and HUB_PW re-tested at hub:200.
Consequence stated plainly rather than left to be discovered:
/home/felhom-repo.orphaned-20260805 (12 535 KB, the three Phase 0 sentinels) is
now permanently unopenable — which is what the set-aside screen promises, and
teardown removes it anyway.
Venue left WORKING and said so: ONLINE, 4 containers healthy, backup target not
degraded, off-site on 2 snapshots. Two things a future session needs: the raw
/mnt/{adatok,mentes} mounts are deliberately left unmounted (R-220's
workaround), and the appliance root credential was shredded — re-fetch it from
the hub.
REPORT-campaign11-phase24.md rather than REPORT.md, per the repo's
parallel-session rule.
No product code changed. No version bumped.
|
||
|
|
df6081e60b |
CAMPAIGN-11 Phase 4: the soak passed, with two absences answered honestly
gates / gates (push) Successful in 8s
All five daily jobs fired exactly once and completed on time. The 04:15 offsite run produced snapshot_count 1 -> 2, last_status ok, 24.9s — unprompted, on a box rebuilt twice and set aside four hours earlier. Nothing on the must-not list fired, including offsiteheal: that is R-218's NEGATIVE control passing (a box with a target must not declare), which does not substitute for its positive half, still owed. Investigated and DISPROVED: tier2-backup completing in 118ms looked like a silent no-op; the copy is real and on the backup drive (818.5 KB verified). Correction to my own pre-registration: backup_run_digest is a TEST FILENAME, not an event type. The real one is backup_run_failures, a failures digest whose silence on a clean night is correct — reporting it as a miss would have been a finding I invented. What survives: the offsite run emits no hub event while both lesser tiers announce success; failures and staleness are covered elsewhere, so it is a consistency wrinkle, recorded not filed. Two absences answered rather than assumed: the restore-test's silence was PRE-REGISTERED as correct (6h eval, 24h settle); the agent's whole-guest tier is left explicitly UNRESOLVED because routine local-api calls are not logged at INFO — a 5h search returns 0 on a box that demonstrably served such calls, so the instrument cannot answer. No product code changed. |
||
|
|
95e2fba8ca |
CAMPAIGN-11 doc: invariants, RTO, teardown-owed, hygiene
gates / gates (push) Successful in 7s
Invariants across the phase boundaries: I1/I4(product)/I5/I7 held; I3 breached twice (R-227's raw Bad Gateway, and R-220's refusal naming an impossible action, reproduced live a third time); I6 breached twice (R-224, R-225). I2 is recorded as untested rather than passed, because F10 could not be injected. I4 was breached by the HARNESS, not the product — sourcing ~/.config/credentials echoed two demo-box recovery codes into the transcript. RTO unchanged and explicitly not re-measured: Phase 2's faults are not a re-walk. Phase 1's 61 attended minutes stand; the unaided number stays undefined. Teardown recorded as OWED with all three layers named plus the off-site side — including that the campaign's sub-account now holds TWO repos, the fresh one and the 12 535 KB set-aside history, and that demo-felhom/demo-hp namespaces on ep0 must not be touched. |
||
|
|
ce6fafa14f |
CAMPAIGN-11: R-228 (set-aside history invisible) + campaign doc F7/F9/F10/F11 sections
The move-aside is correct and verified byte-for-byte; what follows it is not. settings.json records orphaned_renamed_to and OrphanedRenamedTo has zero references in any template or handler — 12.5 MB of deliberately retained customer data at a path the box knows and never shows. Fifth instance of the project's seam-built-but-never-wired pattern. Campaign document gains the F7/F9/F10/F11 write-ups: F7 split (set-aside PASS, afterwards FAIL), F9 partial (R-215's gate proven live on a narrower shape), F10 NOT INJECTED (harness — three self-healing attempts), F11 PASS. No product code changed. |
||
|
|
04ac465da6 |
CAMPAIGN-11 Phases 2+4: the fault journal, the campaign document, and hub v0.97.1's missing heading
gates / gates (push) Successful in 8s
Phase 2 (eleven injected faults) and the §4 positives that were owed. - §4.1 MEASURED, twice: the box's rendered GetFloor() is 0.200.0, and a cold-started controller logs "settle-gate: GO — at/above floor 0.200.0" against the same line reading "floor still unknown" while the hold was in force. Also corrects the brief's plan: SetFloor's line is u.dbg(), gated on logging.level=debug and written to the logger, so it can NEVER reach the debug ring — a restart alone would not have produced it. - §4.2 still NOT measured, deliberately: the venue has an off-site target, so needsOffsiteCredential correctly returns false. Recorded, not inferred from the unit test. - F1 PARTIAL, F2 PASS, F3 FAIL, F4 FAIL, F5 PASS, F6 PASS, F8 PARTIAL. F3+F4: a hub outage and a stopped agent are both rendered as "this code does not open your package", in 0.056 s and 0.030 s — no unseal attempted. The agent's own err field distinguishes them exactly and it is discarded at the HTTP boundary; the R-216 capability gate answers source=version and cannot see reachability. - R-217's fix HOLDS under exactly its fault (F5), verified with the false-claim strings absent and accented positive controls present. hub/CHANGELOG.md: v0.97.1 had no heading of its own — the change was written into the v0.97.0 entry while the deployed tag is 0.97.1. Given its own entry, marked as added retroactively. Second occurrence of the class (agent 0.90.1). Evidence: documentation/tests/campaign11-evidence-2026-08-05/journal-phase24.md No product code changed. |
||
|
|
c917251eeb |
docs: P6 EXECUTED — containerd store moved to SSD2; reboot deferred (R-209, R-209a)
gates / gates (push) Successful in 7s
Operator ruled "proceed" after reading the pre-analysis; the storageReserved condition CC made the move conditional on was applied with it. The pre-analysis is retained verbatim as §11-bis so the ruling can be audited against what it was given. MOVED WITH ZERO LOSS, verified BEFORE the original was touched on four independent observables: 550,891 = 550,891 filesystem entries; 448 = 448 trusted.overlay xattrs (load-bearing — overlayfs stacking rides them, which is why rsync -X is not optional); 37,243 = 37,243 hardlinks; byte-identical meta.db sha256. Verified again after: identical image/tag/volume ID SETS, cache 2.782 GB / 38 records, ceiling still 30 GiB, pg 4 DBs / 31 tables / 175,135,767 B, redis DBSIZE 2437. End-to-end proof was a real build on the relocated store, rc=0 — not an inventory count. k3s was never at risk, and that was established BEFORE stopping anything: it runs a separate containerd, so Gitea, the registry, the hub, PBS, Longhorn and ~160 pods stayed up. Only the two jarr-* dev containers were affected. TRAP FOUND WHILE PROVING THE GUARD, and it is the reusable part: RequiresMountsFor on a path with NO mount unit is a SILENT NO-OP — containerd started normally against an absent-but-unmounted path. A typo'd guard buys nothing and says nothing: the built-but-never-wired shape again. The guard was therefore verified positively at the unit level (Requires= AND After=mnt-ssd_2.mount on both units), and refusal proven with a genuinely absent DEVICE via a temporary synthetic .mount unit — /mnt/ssd_2 hosts 12 live Longhorn replicas and must never be unmounted, and editing fstab on a production host risks emergency mode at boot. Result: "Job containerd.service/start failed with result 'dependency'", is-active inactive. It refused rather than creating an empty store on the root disk. storageReserved on SSD2 0 -> 80 GB; still Schedulable=True at 76.34%. NOT PROVEN, AND NOT GLOSSED (R-209a): the operator ruled DO NOT REBOOT, so the move has not survived a reboot — which is this task's own standard for validated. The MECHANISM is proven; the CONSEQUENCE (that a real boot mounts /mnt/ssd_2 before containerd starts, in this host's actual ordering) is not, and mount-ordering reasoning is exactly the class this project has been burned by. So: the pre-move tree stays MOVED ASIDE, not deleted (34.3 GB, which is why / sits at 54% and not lower), and felhom-store-postboot-check.service — oneshot, enabled, dry-run PASS at install — validates it automatically at the next boot, asserting positively that /var/lib/containerd does NOT exist. Read /var/log/felhom-store-postboot-check.log after the next boot; on PASS the rollback copy can go and ~34 GB returns. Gates: repo_gates.py --fast — all OK, rc=0 (run separately from this commit). |
||
|
|
ee9d9bf203 |
docs: SPIKE — DooPlex build-cache containment; R-205..R-211 (2026-08-05)
gates / gates (push) Successful in 7s
Spike output only; no production Go code. The one shipped change rides in homelab-manifests 6808a4b (R-205, the monitoring rule). VERDICT: mechanism confirmed, with one correction and one refutation. - CONFIRMED: builder.gc IS honoured under the containerd worker and DOES evict. Proven by naming a 440 MB `go mod download` record present at build N and absent by N+2 — not by absence of an error. - CORRECTED: honoured ONLY in the `policy` array form. The flat form is SILENTLY ignored — daemon starts, logs nothing, keeps its defaults. `dockerd --validate` returned "configuration OK" for a bogus key AND for a config that then crashed the daemon. The oracle is `docker buildx inspect`. - REFUTED: Docker's `data-root` would NOT move the cache — it moves 0.62 GB. The 181.4 GB belongs to the system containerd (`root` in /etc/containerd/config.toml). P3 (operator-approved) executed: prune claimed 156.9 GB, the filesystem returned 150.35 GB (the 6.5 GB gap is layers shared with images), SYNCHRONOUSLY — / went 86% -> 53% used, and Longhorn's default disk went Schedulable=False (DiskPressure) -> Schedulable=True (18.85% -> 50.32%). P7 root-caused the largest item and it is NOT the cap: all 208 `go mod download` records had Usage count 1. Isolated by controlled builds — same VERSION build-arg -> CACHED, new VERSION -> executed, byte-identical tree. `ARG VERSION`/`ARG GIT_COMMIT` sit ABOVE the module-download step, and a RUN's cache key includes the stage environment. Both Dockerfiles have it. One line each to fix -> R-208. P6 NOT EXECUTED — stops at the operator, as specified. Pre-analysis: the move is safe as measured (+38.8 pp above the 25% floor) but SSD2 is the only Longhorn disk with storageReserved=0 and is overcommitted 6.9x; at full inflation the move lands 12 pp BELOW the floor. The prune removed the move's urgency, so CC recommends against it unless ~80 GB is reserved on SSD2 -> R-209. Register: R-205 (CLOSED, shipped), R-206 (Ansible: cap + prune + narrowed Docker ban), R-207 (DRY_RUN guard), R-208 (ARG ordering), R-209/R-210 (operator), R-211 (Prometheus has no config-reloader — rules changes have never applied until something restarted the pod; found while verifying R-205). Gates: repo_gates.py --fast — all OK, rc=0 (run separately from this commit). |
||
|
|
2a7ac03c47 |
R-201 PASSED: a customer's file survived a machine rebuild and came back byte-identical
gates / gates (push) Successful in 7s
|
||
|
|
b228fd102d |
R-201 night run: the off-site key IS recoverable after a real rebuild (proven); the verdict is blocked by R-204
gates / gates (push) Successful in 6s
|
||
|
|
73fb595e38 |
R-203 shipped: the app and its backup agree, and 'ok' means it — R-201 unblocked
gates / gates (push) Successful in 7s
|
||
|
|
a0c4b607a6 |
drill R-201: prepared and HALTED before the wipe — a mandatory data directory was missing from the off-site snapshot while the run said ok (R-203)
gates / gates (push) Successful in 8s
|
||
|
|
3f2b7bc023 |
RECON: trace the offsite DR chain link by link — it does not join up (R-198..R-201)
gates / gates (push) Successful in 7s
Read-only recon of the escrow -> recovery chain, from a dead node to an open
repository. No production code, no build, no version bump.
Headline: the hub's superseded-escrow retention does NOT retain the offsite
repository password. host_escrow_superseded has no identity_blob column and
demoteCurrentEscrowTx copies only the K-escrow blob, so what survives a
supersession is the PBS datastore key, not the restic repo password. The next
escrow ceremony -- which the system tells a rebuilt box's customer to run --
destroys the last copy. Both demo boxes crossed that line on 2026-08-04.
Also established:
- the hub's blob-serving endpoints (re-enroll / restore-directive) have zero
callers anywhere: agent, hub UI, scripts, runbooks (R-199)
- POST /backup/offbox/inject-password is routed and handled but no template
contains the form (R-200)
- nothing in the recovery path has ever been exercised; the one live
round-trip proof (2026-06-10) predates the ResticRepoPassword field (R-201)
- a fail-closed mint refusal IS implementable: the report ACK already carries
escrow{identity_blob_present, restic_pw_sha256} and the controller discards
it whenever no offbox target exists
Corrections: yesterday's spike annotated (candidate (b) overturned in part --
unattended recovery is impossible, customer-present is not); capability-map
retention claim struck through and replaced with what the code does.
Deliverable: documentation/audits/RECON-offsite-dr-chain-2026-08-04.md
Register: new R-198..R-201; R-193 and R-192 updated; STATUS.md refreshed.
|
||
|
|
7fff45d688 |
R-195: a customer with no machine ever bound does not alarm (hub v0.92.0) + R-193/R-192 spike
gates / gates (push) Successful in 7s
Part 4 (ships): `david` — a prospective customer with hosts=0, host_deletions=0, reports=0 — e-mailed an expected_dbdump_missed ERROR at 03:00 UTC three mornings running. The existing down-skip could never cover it: it reads the staleness checker's state, which is seeded from a query over the `reports` table, so a customer that never reported has no state at all and GetState() returns "" rather than "down". store.HasEverBoundHost (hosts row OR host_deletions tombstone) is consulted once per customer at the top of the deadline loop. The discriminator is "was a host EVER bound", never "has a report arrived" — a box installed and never heard from is a real fault and keeps alarming. Fail-OPEN on a read error. Red-proof observed: removing the guard fails with `got [expected_dbdump_missed]`, verbatim the event david sent. Parts 0-3 (spike, NO production code for R-193/R-192): audits/SPIKE-offsite-credential-recovery-2026-08-04.md establishes that the one-shot provider password is the RECOVERABLE secret and the restic repository password is the irreplaceable one — and that a guest rebuild mints a fresh one, orphaning the previous off-site history. Measured without touching a box, by comparing host_escrow.restic_pw_sha256 against host_escrow_superseded: BOTH demo boxes changed (demo-hp 15 snapshots / 40.9 MB, demo-felhom 36 snapshots / 1.14 GB). demo-felhom's "lucky" 76-second recovery restored delivery and not the repository, silently, for 13h. ReissueCredentials does NOT rotate the restic password (R-39's record and two hub comments are wrong -> R-196); candidate (b) is not implementable against a zero-knowledge escrow; candidate (a) already exists as F3 and is wired to the wrong event. Ends in ranked options and an unanswered question for the operator. R-195 SHIPPED; R-196 + R-197 filed; R-192 + R-193 updated, neither closed. |
||
|
|
e34b614e5b |
docs: R-182 closed, R-90 closed on measurement, R-86 unblocked, ep0 record corrected
gates / gates (push) Successful in 7s
R-182 CLOSED (controller v0.194.0 + hub v0.90.0/.1), proven live on demo-hp. The hub's notification_log for the run reads: two per-app failures RECORDED, one digest SENT naming both, and the customer channel SKIPPED with operator_only. Against the measured previous behaviour — two failures, one email naming one app, one leaving no trace anywhere. Scenario D proved itself on an event I had not planned: disk_critical alarmed on two filesystems, the second was collapsed by the cooldown, and that collapse is now visible WITH ITS KEY. Yesterday it would have left nothing at all. A gap the spec did not anticipate is recorded with its fix: the per-app event also fires from the periodic sweep, outside any run, so making it record-only would have created a NEW silence. The sweep emits a digest too, with no run_id, so it stays under the ordinary hourly cooldown. ep0: MEASURED on the box — 7757 MB (8 GB), 4 vCPU, and the 4 GiB swapfile SURVIVED the resize and is active (checked, because a resize is a stop/start). The 40 GB local disk is UNCHANGED, so no disk figure was touched anywhere. Five documents corrected — three of which the task's list did not name, found by searching. Two audit/evidence documents ANNOTATED, body untouched: they record what was true when written and that is their value. R-90 CLOSED. R-86 unblocked and re-ranked, stated honestly: 8 GB is comfortable, not unbounded — the original OOM was a 14.46 GB restore — so the restore-test cadence should still be paced, just not by fear of the endpoint. target-selection.md's "D-d did not name ep0 either way" is deliberately left standing. It is the operator's question, not CC's. STATUS.md 127 -> 83 lines, items rather than sentences. |
||
|
|
7406ac7bbf |
audits: R-165 Phase 0 — P1 and P2 measured, nothing changed
gates / gates (push) Successful in 7s
P1 PASS: a pre-merge archive (mp0+mp1, confirmed from its own vzdump log)
restore-tests clean on demo-hp with mount_parity ok in 84s. mountParity was
not touched. Limit stated: run with the current agent because the merged one
does not exist until after the STOP; the comparison is archive-vs-its-own-
restore and never consults the host layout, so it carries provided Part 2
honours its constraint not to touch the restore path. Re-run after Part 2.
P2: all three probed variants are mechanically clean — both paths writable,
ONE df figure, dockerd 3/3 reboots, /mnt propagation, and a container's
statfs(/) reporting the merged volume. The task's flagged ordering risk for
V-b did not materialise. They are separated by SCOPING instead:
V-a container sees /mnt = 8.0K but customer data sits inside Docker's
data-root, so clearing /var/lib/docker destroys every local unit
V-b container sees Docker's ENTIRE data-root under /mnt (17.9M on an
empty box), making the bootstrap's own scoping comment false
V-c neutral mount at /var/lib/felhom, both paths binds — breaks neither
V-c was probed because the measurements showed each named variant violates
a different documented invariant. It is offered as a measured option for the
operator, NOT adopted.
Teardown all three layers: 9401/9402/9403 destroyed, 5.19 GB returned, and
the hub registers verified unchanged (5 customers, 4 hosts).
|
||
|
|
41dbecb264 |
docs: R-167 + R-158 CLOSED, R-165 SPIKED, R-174..R-177 filed
gates / gates (push) Successful in 8s
R-167/R-158 shipped and proven live (controller v0.191.x, hub v0.89.0): two new capability-map rows PROVEN-LIVE with live citations, and 07-backup-architecture.md §7.5's closing claim "nothing warns when an app crosses the line" is now false and rewritten (S-1: an architectural contract changed in the same session). §7.5 also gains the caveat that its size bound is ONE BOX'S, not the fleet's. Part 3 SPIKE (audits/SPIKE-r165-mp1-merge-2026-08-02.md): M1-M5 measured, NO layout touched. Three findings the merge session must not re-derive: "the layout" is not one thing (200G/50G vs 50G/20G vs 16G/8G); mp1 is a BULKHEAD and not only a ceiling, so after the merge an overflow reaches /var/lib/docker; the golden fails closed on the split in four places. D-a's condition (1) is currently SATISFIED — no external box is in the hub's register, and both demo boxes are Tier 0 and reinstallable. Recommendation given, choice NOT made — it ends at the operator's ruling. CONTEXT.md S-11 (D-c's routing, and why R-158's own backup_failed proposal was overruled) and S-12 (the monitoring landed BEFORE the merge). STATUS.md gains the plain-language section and the merge decision, with two older entries trimmed so the page did not grow. New rows R-174 (closed same session), R-175, R-176, R-177; each ID grepped free before minting. |
||
|
|
ad28699761 |
docs: R-157 A / R-170 / R-171 closed — boot recovery finished
gates / gates (push) Successful in 7s
Controller v0.190.0. Docs only here; no hub change, no hub version bump. - audits/DIAG-bootrecon-drive-absent-2026-08-02.md — NEW. The Part 0 diagnosis, including the run that produced a FALSE NEGATIVE and the mechanism behind it (the agent re-binds an unmounted drive within ~60s, so the drive gate's startup reconcile restarted the apps one second before the sweep looked). Records that the write hazard was blocked only by an ACCIDENTAL filesystem permission that no code owns and no test pins. - architecture/02 §0a — the boot-recovery contract (S-1): both gates read desired state; the sweep observes a SETTLED fleet and each sample must refresh first; nothing is started without asking, fail-safe. Plus the durable warning: Manager.StartStack has no gate of its own. - 00-capability-map — the boot-recovery row, with the repeat count cited per N.5 (6 of 6 hard resets) rather than a bare PROVEN-LIVE. - OPEN-ITEMS / ROADMAP — R-157 CLOSED (both mechanisms), R-170 CLOSED, R-171 NEW and closed the same session, marked a regression from v0.189.0. - STATUS.md — the power-cut line moved from "What's broken" to "What works right now" with its repeat count; one dated bullet in the change log. - CONTEXT.md S-13 — the lessons worth carrying: "it didn't happen this time" is not a disproof; widening a window makes previously-unreachable overlaps reachable; and a settle detector is only as good as the freshness of what it samples — the fix's own defect, found live rather than by review. |
||
|
|
c718aad1bc |
docs: R-168 SHIPPED, R-29 CLOSED on the demonstrated alarm, R-169 minted
gates / gates (push) Successful in 7s
SPIKE-ci-runner-2026-08-02.md: all six probes with method, measurement and ruling; none STOPped. P2 (stock image has git but no python3) and P6 (a runner that loses its state re-registers and orphans the old record) changed the design; P5 (a failed run signals NOTHING) is why the alarm exists at all. R-168 SHIPPED with its evidence. R-29 CLOSED — on the demonstrated alarm and not on a green run, as required: the class it opened is answered at both ends, the hook refusing locally and CI catching a --no-verify bypass and emailing. R-161 noted: its automatic half now exists for the STATIC gate, while its original scope, the runtime gate, is deliberately still not automatic and should stay that way. NEW R-169 (grep established R-168 was the highest in use): CI can only report, because there is no gate in the road. Making it blocking needs branch protection plus a PR workflow, both of which change how the operator works — so it is theirs to decide, and the row states the cost honestly rather than recommending it. CONTEXT gains S-8 (CI detects, does not block, and why that is structural), S-9 (a detector that tells no one is not finished, plus the curl and Cloudflare-1010 traps), S-10 (the runner is unprivileged because DooPlex is Tier 2), S-11 (CI reproduces the sibling layout). CLAUDE.md gains the rule earned by red-proofing: a go test -run pattern that matches no test prints ok and exits 0, and an instrument that can silently drop results is not a measurement. |
||
|
|
482af37b7d |
Campaign 10 closeout Part 2: teardown complete — five layers, each verified gone
Evidence-survival check FIRST: HEAD == origin/main ==
|
||
|
|
7efb7a53d3 |
Campaign 10 closeout Part 1: Q1 lowers R-158's rank; Q2 clears ValidateDump and kills C2's gate
Q1 - what the customer sees when a backup refuses for lack of space. The failure
IS customer-visible: /backups renders "Adatmentés sikertelen" with a cross mark.
It is absent from the dashboard, the launcher, the app detail page, and - the one
worth fixing - from /backups/apps, the per-app page where you would naturally ask
whether a given app is backed up.
Point 5 measured across three runs: it retries, stays failed while constrained
(marker persists, unit mtime unchanged at 07:30:50), and clears on recovery with a
fresh unit at 07:37:38. /backups/apps reading "Utolsó: 3 perce" tracks the unit's
REAL mtime, not the failed run, so it is honest about the age of the last good
unit rather than claiming a fresh one. Explicitly NOT the R-156 family.
So R-158 is a NOTIFICATION GAP, not a silent-failure defect, and ranks BELOW
R-157 - whose mechanism B leaves a deployed app not running while deadapp reports
"0 currently down", silent on every channel.
Q2 - ValidateDump was right and no bad dumps are shipping. The live DB genuinely
had zero accounts (only _prisma_migrations 129, cc_proof 82, instance_settings 1).
An empty table proves nothing, so an account was SEEDED as the task required: the
warning then stopped entirely and the dump provably contained the rows (c10acct 1,
c10user 2; 102766 -> 103029 bytes).
But that kills C2's proposed ordering. A fresh appliance legitimately has zero
accounts, so gating on "accounts has rows" would block the backups of every new
customer until someone registers. The validator's fact is right; its inference
("may predate the customer's data") is wrong - there was no data to predate. The
chain is therefore longer: a sound predicate first (compare the dump against the
LIVE db, per-table counts, not an absolute expectation), then warn->gate, then the
tar-drop. Until then the DB volume tar stays load-bearing - not because dumps are
bad, but because nothing can yet prove one is good.
No new R-n; register grepped. Nothing fixed. Part 2 (teardown) follows.
|