dfd854474e92ae5dfb791b8ab45e76867b514092
129 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
dfd854474e |
drill 0.243.0 complete: 0 interventions, the connect e-mail proven, NOT ready for a volunteer
gates / gates (push) Successful in 21s
The automatic connect e-mail is proven with a real mailbox: the host record was deleted at 12:22:59Z and the mail reached the customer at 12:23:00Z, one second later, with selfbind_link_sent (host delete) on the timeline. The requirement was two minutes. The hub refuses to delete an ONLINE host with no override, so the record had to fall stale first — that wait is part of the proof. Interventions: 0. Every P1 fix this drill set out to prove held on a fresh box. The verdict is still no, for a new reason: a one-drive box with no off-site tier keeps none of the household's own files in any backup, the page says otherwise, and the restore that should save them makes it worse (R-537, R-538). Teardown, three layers, stated. Customer tester-1 kept; RESET never used; nothing on the off-site server written or removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
328c3fc23c |
BIGNIGHT morning note (STATUS), topic report, capability-map annotations; teardown layer 1 done, layer 3 pending
gates / gates (push) Successful in 19s
|
||
|
|
65790672d5 |
doorstep walk on ISO 1.27.x: 1 intervention (R-505), STOP before publish
gates / gates (push) Successful in 17s
ISO 1.27.1 gated PASS and proven live: first-boot console Felhom-only, pvebanner masked across a proven reboot. Hub v0.113.0 hand-over copy live (R-497 closed). Full first hour walked again on customer tester-1 (three disks + one disk): deploy, use, backup, removal, byte-identical restore, power cut, typo all PASS. The tunnel gives a fresh box no routes: 12/12 503 from DooPlex (R-505); the record has no e-mail (R-508). Rows R-507, R-508 filed; R-496/R-495 fixed/answered awaiting publish; day-0 A.1 no longer claims the controller creates hostnames (R-506). NOT PUBLISHED. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
38848ffbeb |
drill: a stranger's first hour on 0.242.0 — 1 intervention, not ready for a volunteer
gates / gates (push) Successful in 21s
Golden 0.242.0 baked, round-trip verified and vouched (cadence rule, R-468). Fresh box from the public ISO on demo-hp: landed on the vouched set, two apps deployed and used, backup, remove, byte-identical restore, power cut and code typo all PASS. Stopped for a volunteer by R-493 (no instructions) and R-494 (the setup mail's dashboard link has no DNS; intervention I1). R-493..R-500 filed. Capability map: first-hour row added (PARTIAL), journey row scoped. Stopgap Hungarian volunteer guide written. Hub teardown layer pending. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
681c3d6a6d |
docs: rulings 7 and 8 shipped and proven live (R-470/R-472/R-475 CLOSED); R-477..R-480 opened
gates / gates (push) Successful in 21s
Hub v0.112.0 serves a floor above the golden with a declared MinAgent;
controller v0.239.0 reached both demo boxes by that floor in 14 s and 15 s
and updates on any backup tier. 09 §3 decisions 7 and 8, §6/§6.1; 07 §6
line; capability map row; STATUS items 15/16 done and the cadence line
corrected; CONTEXT; register: R-470/R-472/R-475 compressed to CLOSED-ITEMS
(full text at
|
||
|
|
5ef0f52bcd |
Slice 4 shipped (R-448/R-443/R-439 CLOSED, proven live); R-472..R-476; the floor-between-bakes claim corrected
gates / gates (push) Successful in 19s
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure. Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/). Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the runbook, STATUS, CONTEXT, R-468 and the gate docstring. Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
4b2e5608c2 |
R-442 CLOSED (controller v0.236.0): "delete my data too" deletes it or refuses; R-465..467 opened
gates / gates (push) Failing after 18s
- OPEN-ITEMS: R-442 row removed; R-465 (six remaining Paths.HDDPath readers — audit), R-466 (recovery-unit residue after "delete backups"), R-467 (v0.236.0 owes a golden). - CLOSED-ITEMS: R-442 compressed, reasoning kept, fleet shape now established. - 00-capability-map: lifecycle row narrowed (Campaign 3 proved remove removes the APP, not the data) and re-proven from audits/R442-2026-09-13/. - STATUS: item 13 in plain language; item 7 closed (ruled 2026-09-02, 09 §3). - audits/R442-2026-09-13/: live evidence (A, C, D bodies, controls, log window, teardown). Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
417df06f35 |
slice 3 docs: the ruling, the shipped mechanism, and four rows closed
gates / gates (push) Successful in 17s
09-update-architecture.md gains the fourth dated operator ruling (2026-09-06, Option 1) and its section 5 is rewritten from a proposed shape into the shipped one: the pin, the stored definition, the render table, the four writers, the startup ordering, and the trap this slice set for slice 2 - the live compose file is now the frozen one, so a badge comparing against it would answer Naprakesz on exactly the apps that are behind. 02-controller-module-map.md said 'copy compose + .felhom.yml'. That stopped being true today, so it is corrected, and the two sections describing the old seam now carry a banner saying they describe v0.234.0 and below - kept because every box under v0.235.0 still behaves that way and because they are the measured account of why it changed. R-447, R-441, R-438 and R-455 closed and compressed into CLOSED-ITEMS; R-458 opened for the .felhom.yml asymmetry, with what would settle it by measurement. Live evidence: two real catalog pushes travelling the real 15-minute cycle, both reverted, the tree byte-identical afterwards. The restart that used to take 18.3 seconds and pull a new image now takes 0.1 seconds and pulls nothing. |
||
|
|
bc47dd4ef9 |
v0.234.0: a known limitation written on 2026-09-02 was a defect by the next morning
gates / gates (push) Successful in 18s
The operator looked at demo-felhom and found OpenGist - up 15 hours, running exactly the catalog pin, showing no badge at all. 09-update-architecture.md had recorded that as an accepted limitation the day before: 'the fleet view fills in gradually'. On a quiet box gradually means never, and a feature that fills itself in on an event nobody triggers is, on the quiet installations, not shipped. That limitation row is now struck with the reason kept. The living document gains slice 1b, the two admission rules of the backfill (it never overwrites, and it refuses to seed a partial observation because the badge reads a service-count mismatch as BEHIND), and the note that the same field having two writers with two different admission rules is deliberate. Live evidence added: all nine apps already had records by the time 0.234.0 was ready, so the natural fleet state could no longer exercise the new code - said plainly rather than papered over. The pre-0.233.0 shape was recreated on demo-hp by stripping two records; the backfill re-seeded exactly those two with digests matching independently-read ground truth and left the other seven alone. The refusal half was deliberately NOT staged live: it needs a degraded app, and manufacturing one risks the false-customer-email class that already cost 61 mails (R-330). Unit-tested with a red-proof, and recorded as unproven-live. R-457: a test that hardcodes a date and asserts an age derived from it is green only on the day it is written. Mine was, and it went red overnight. Six other files carry both a date literal and time.Now() - named as candidates, not accused. |
||
|
|
0705942783 |
the badge IS proven live, and the 'stale password' finding was mine, not the box's
gates / gates (push) Successful in 16s
I reported that the vaulted dashboard password no longer worked on either demo box, and quoted the controller's own 'Failed login' as the discriminator. The password was fine. ~/.config/credentials quotes its values with SINGLE quotes and my sed stripped only double quotes, so the quote characters went out as part of the password. The operator corrected it in one line; one retry returned 302. The instrumentation lesson is the finding and R-453 now carries it: 'Failed login' separates wrong-password from wrong-Host-header, and that is ALL it separates. It cannot tell a wrong password from wrong password HANDLING, and I read it as if it could. This is the second time this file's quoting has produced a confident wrong verdict, so the fix is one shared extraction helper, not a resolution to be careful. With the session recovered, the badge is validated on live pages: Naprakesz twice on /stacks and on /apps/bookstack; NO badge at all on /apps/docmost (a deployed app with no record - absent is UNKNOWN, not current); and 'Frissites elerheto - 52 napja' on both surfaces, the age being real arithmetic on bentopdf's catalog_since. The behind state was staged by editing one compose tag, with no restart and no up -d, and reverted byte-identically (sha256 equal, diff empty, container never touched). Capability-map row upgraded to PROVEN-LIVE with the one unexercised badge state named. STATUS item 9 now needs nothing from the operator. |
||
|
|
6035dfcc3a |
09-update-architecture.md: the update path finally has a document, and it is a living one
gates / gates (push) Successful in 17s
R-438's document half. It records how an update works AS MEASURED, quotes the RestartStack comment that proves the restart half was CHOSEN (a design decision is not a defect), carries the three operator rulings of 2026-09-02, strikes the word 'rollback' (once a migration has run the old image will not start), states the target shape, and lists the seven slices with a status each. R-438 and R-440 amended and BOTH STAY OPEN: the mechanism is documented, not changed. Nothing closed, so CLOSED-ITEMS.md is untouched. Eight new register rows, 194 -> 202: R-446 (Naprakesz can be false for the 23 floating pins), R-447..R-451 (one per remaining slice, with a rank and an owner), R-452 (no gate enforces catalog_since - the runner fetches at --depth 1), and R-453 (the vaulted dashboard password is stale on BOTH demo boxes, which is what stopped the badge render from being validated live). Live evidence for slices 1 and 2 in documentation/tests/. The record is PROVEN LIVE through the boot reconciler on demo-hp - one entry per compose service, digests matching ground truth read independently. The badge RENDER is not, and the five attempts are listed rather than summarised. |
||
|
|
56c7e373a3 |
SPIKE: what an app update actually does, and which other paths do it too
gates / gates (push) Successful in 18s
THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has already moved, and the Restart button does it — 18.3s with a network pull when the target image is absent, 0.5s when present, against a negative control that did not even recreate the container. The boot reconciler does the same thing unattended when an app fails to come back (bootrecon.go:269 -> StartStack). AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades nothing. Docker restores the old containers and the reconciler logs 'no boot-orphaned apps (nothing to start)'. AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run, the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported'. 'Rollback' is the wrong word for this arc and is struck. Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator confirmation. Peti's box was never contacted. No production code written in any repo: felhom-controller is at 960d29b0612c before and after, tree clean, and build/vet/test are green — run at the end to prove exactly that. Register: R-438/439/440 updated with live evidence; R-441..R-445 opened (restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB left behind; update reports success over a broken app; no fleet fstrim; hub telemetry outlives the app). Capability map gains three measured rows. The mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR, not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision. Five claims in the brief are named as wrong, including two of my own method. |
||
|
|
30681764cb |
hub v0.111.0: notice a deletion within a day (R-431); correct R-429; re-scope R-95
gates / gates (push) Successful in 17s
THE RECORD WAS TELLING A WORSE STORY THAN THE TRUTH FOR TWO MONTHS, and my own probe is why.
R-429 CORRECTED. Yesterday's spike searched for a directory called `.snapshots`. The vendor documents
the path as /.zfs/snapshot. The probe's CONTROLS were sound and its SUBJECT was wrong, so "not found"
was true and meant nothing. Re-probed at the documented path on both boxes, with controls:
- /.zfs lists (shares, snapshot) from inside the jail;
- a write into /.zfs/snapshot is REFUSED - dest open ...: Failure - while the identical write to
the account home SUCCEEDS and was cleaned up.
That is the append-only property PROVEN rather than cited, and it is the sentence the whole re-scope
rests on. Seven daily snapshots are confirmed in the panel. The mitigation works. What remains true,
and was always the actual finding: the row claiming it had no R-number, its "confirm tomorrow" went
36 days unanswered, and the DUE-CHECKS block built for that class was empty. THE FINDING WAS NEVER THE
SNAPSHOTS - IT WAS THAT NOBODY COULD TELL.
R-95 RE-SCOPED: the box can delete its LIVE repository but cannot write to the daily snapshots of it,
so a deletion costs at most one day plus a per-file recovery - not open-ended loss. The ranking is
Viktor's; it has been #1 since July on the old story.
R-432 FILED: a sub-account sees /.zfs/snapshot EMPTY while the same box holds seven snapshots, so
per-file recovery is operator-only today. One panel read settles whether a NAMED snapshot can still be
entered, which would make it product-reachable.
R-431 SHIPPED. Third signal in OffsiteChecker. On the hub deliberately: a detector on the box is one
the deletion can silence. Threshold REASONED, not invented - over 12 898 reports every decrease lands
on ZERO and predates stats_known, and in the stats_known window there are none, so observed churn gave
nothing to calibrate against. Retention cannot halve a total; a mass deletion goes to ~0. Hence: more
than half, and at least 5. Guarded by StatsKnown (R-331), the declared State (R-204) and run success
(R-100 - whose lesson lives in this very file).
ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms.
Three red-proofs run. The escalation one only became real after the first version was found HOLLOW -
it re-swept the same report, so the baseline had already moved and the latch was never consulted.
07 row 10's status is NOT moved: the write-refusal is measured, but the recovery ROUTE has never been
walked, which is what PARTIAL means.
|
||
|
|
f41a1a0ad8 |
R-411/408/407, R-414, R-412a leg 1, R-410, R-406 CLOSED; determination + live evidence
gates / gates (push) Failing after 18s
Part 2.1's determination is the first artifact: the scratch resolver was consciously OUT OF SCOPE for R-356, not excluded on state-only grounds - established from R-356's own commit 08eb1a6, whose tests say 'the prepared scratch still resolves ... only the DESTINATION moves'. So 07 section 6.3's rule applies and now has a FOURTH consumer, and the section says so. Live evidence: the collision rerun on demo-hp with the sampler positively controlled first (12 locks=1 across a real check, 4 locks=0 quiet), showing unlock --remove-all 0 times where the drill saw it twice; and the proof reaching verdict pass on demo-felhom - the box that could not run it at all - recorded where last_proof_result had been ABSENT every night. Capability map: the off-site proof row now records that the nightly firing IS proven (it ran unattended at 05:30 on demo-hp) and that a driveless box can now be proved. Register: six rows closed and compressed. OPEN 176 -> 170, CLOSED 152 -> 158. |
||
|
|
7ee25925f9 |
R-87 CLOSED: live evidence, capability row, architecture verdict, registers
gates / gates (push) Failing after 17s
Controller v0.231.0 + hub v0.110.0, both deployed and verified on demo-hp. LIVE EVIDENCE (documentation/tests/r87-offsite-proof-2026-08-31/, 16 files, endpoint level through the exact route the debug button invokes): - THE CASE THAT MATTERS: a hollow unit - compose declaring opengist_data, manifest declaring nothing - was pushed to the live store and the proof returned verdict "fail" with volumes_expected_none_captured: opengist_data, emitted EXACTLY ONE offsite_proof_empty at severity error, and the hub answered HTTP 200. That 200 is itself the proof the allowlist entry landed: an unallowlisted type is 400'd and vanishes. - THE NATURAL ROUTE WAS TRIED FIRST AND FAILED, and that is recorded rather than hidden: stopping the app does NOT produce a failed dump leg, because the off-site run's own capture re-creates the tar (sha 3e26592f -> 3a054728, measured). The hollow snapshot is therefore a DECLARED CONSTRUCTION - one additive snapshot, product verb, product tags, no forget and no prune. State restored: the product's own run made a healthy snapshot the newest again and the proof then passed opengist. - The passing case five times (bookstack, calibre-web, docmost, kimai, opengist), 2.2-4.0s each, matching the spike's measured band. - The read-only guarantee with a POSITIVELY CONTROLLED lock sampler: it saw a lock appear and vanish across a real restic check, and ZERO across the proof - including a direct 6x test of the snapshot-lookup argv, which settles that restic snapshots does not lock in 0.14.0 either. - Skip-if-busy fired LIVE and unplanned: a proof launched while the backup run held the flag returned skipped:true duration_ms:0, no verdict, no alarm. - The customer's own verification copies were untouched throughout, which is the safety property the separate proof root exists for. ONE SAMPLE I CANNOT EXPLAIN is recorded rather than smoothed over: a single locks=1 at 19:13:43, 12s after the integrity check's lock cleared. Two independent tests exclude the proof; I did not establish what it was. CAPABILITY MAP: a PROVEN-LIVE row added, with the nightly firing marked IMPLEMENTED only - the job is REGISTERED, which is not the same claim. 07 section 8 MATRIX ROW 4 WAS NOT MOVED, deliberately, and section 10.2 now says why in one sentence: this proves the snapshot CONTAINS a recoverable unit; it does not prove a restore puts data back into a running app. Without that sentence the new green tick reads as covering the drill. REGISTER: R-87 CLOSED and compressed into CLOSED-ITEMS.md. OPEN 172 -> 171, CLOSED 151 -> 152. No new rows minted. R-408 and R-409 stay open and are referenced by this work. golden-currency is RED and it is a DECLARED, EXPECTED debt: v0.231.0 is released and the newest golden carries 0.230.0. The fleet is on 0.230.0; demo-felhom does not have this job. A golden carrying 0.231.0 is OWED and it is Viktor's call (R-242). This push uses --no-verify for that reason - bypass #8. |
||
|
|
dddcc808be |
R-403 CLOSED (controller v0.230.0), R-404 filed as a decision for Viktor
gates / gates (push) Failing after 17s
07-backup-architecture gains section 8.2, placed beside row 5 on purpose: the derived-copy rebuild rule is UNCHANGED and section 8.2 names the single exception, so a future reader who finds RunTier2 skipping a leg does not fix it back. It carries the measurement (120 082 104 B -> 7 036 B on the shipped v0.229.0), the four-case table, why hollow is a manifest question and not a size question, why the data legs are deliberately not guarded, and why the capture job is not guarded either. 00-capability-map: the Tier-2 row's status does NOT move, stated explicitly rather than left ambiguous. R-403 removes a way the route could be DESTROYED between uses; it does not change what the route can be relied on for. Register: R-403 CLOSED and compressed into CLOSED-ITEMS (594 -> 593 open lines). R-404 FILED as a DECISION and deliberately not acted on - should a documents-only push be subject to the golden-currency gate, now that it has been correctly bypassed six times? Both sides stated, plus what happens if Viktor does nothing. The gate was NOT changed. R-242: seventh conviction, and the FIRST where the day-0 ground does not apply - R-403 is a defect in the nightly Tier-2 copy, which a day-0 box starts running on its first night. This push uses git push --no-verify, declared here and in felhom-controller/REPORT.md. A golden carrying 0.230.0 is owed and is more urgent than the previous six. STATUS: the R-403 item moves out of 'Broken' into what works, in plain words; the delivery item now says a golden is owed and that the fleet carries the defect; R-404 goes into the decide section with its do-nothing outcome. Drill evidence: nine phase logs, including the two things that went wrong (a repair whose rsync was not installed in the guest and silently did nothing, and a session that expired mid-run so a POST did nothing). |
||
|
|
c2de785bf2 |
R-102 + R-103 CLOSED (controller v0.229.0) — architecture, register, STATUS, drill evidence
gates / gates (push) Failing after 17s
07-backup-architecture: 6.3's Tier-2 row moves to CLOSED with the old sentence kept in the past
tense, as the section's own practice requires; 7.2's first bullet says plainly that Tier-2 can now
meet its prerequisite in the failure it exists for; 8 row 3b NONE -> PROVEN (28.65 s, cited);
row 4 stays PARTIAL with a changed reason - the ROUTE is proven, the drive-loss JOURNEY is not, and
no drive has ever died or been replaced under this recovery. 8.1's blanks updated per row.
6.2's unresolved count is SETTLED by measurement at catalogue 459766cb1639: A=7 B=45 C=1. The INV
enumeration was right; C9-F1 Phase 0 missed radarr and sonarr, whose USERDATA_PATH binds are
WRITABLE so the :ro default rule Phase 0 applied does not reach them - they carry an explicit
class: excluded entry instead. Class C is bentopdf. No catalogue file was changed.
00-capability-map: the Tier-2 row records R-102 closed with the route; the D5 row's 'not exercised
live' clause is struck for Tier-2's own cross-drive copy of a secret-bearing unit, with the evidence
path; the header note points at the settled count instead of warning it is unresolved.
Register: R-102, R-103 and their C9-F4 / C9-F1b aliases closed and compressed into CLOSED-ITEMS
(596 -> 593 lines, each naming git show
|
||
|
|
1623a4d5b5 |
golden 0.228.0 — baked, published, round-trip verified, vouched, floor raised
gates / gates (push) Successful in 19s
Bake: build-golden.sh v3.0.0 in the drill VM, reverted to virgin and cold-booted.
GOLDEN_SHA256 76a3a98b9e7cc23bf8ae51b38a6272f576df285cb34cd22235ac3f06a31e53ec,
658 079 744 B. All four acceptance markers counted 1; excluding/FATAL/mp1 counted 0.
The 404 pre-gate was proven to work before its 404 was believed — the target URL
404'd while the existing 0.227.1 package 200'd on the same command.
The evidence is the ROUND TRIP: the downloaded bytes match the bake's size and
sha, and ./etc/felhom-controller-image read OUT of the downloaded archive says
gitea.dooplex.hu/admin/felhom-controller:0.228.0.
Vouch: three fields together — golden_version 0.228.0, agent_version 0.130.0,
min_agent 0.129.0 (read from the controller CHANGELOG header, not assumed);
wrapper_sha256 carried through explicitly. agent >= min_agent, so not the R-216
shape. Verified by RE-READING the manifest, never the flash. R-120 gate passed.
Floor raised 0.227.1 -> 0.228.0 (impact preview {"below":3,"valid":true}).
The floor is ACTING: demo-felhom self-updated 0.227.1 -> 0.228.0 in under a
minute and re-registered offsite-integrity by itself. Both demo boxes now
re-read their whole off-site store on the weekly check.
Token hygiene: file->file scp, runner script inside the VM, unit properties
grepped 0. The leak grep on the committed log was proven with a planted copy
(1) before its 0 was believed. Teardown: guest 9100 purged, secrets shredded
after the log was copied out, VM off, disk reverted to virgin.
golden_currency_gate.py red -> green. All 12 felhom.eu gates OK.
|
||
|
|
77a5a1154b |
docs: controller v0.228.0 — R-399/R-400 closed, R-401/R-402 filed
gates / gates (push) Failing after 19s
STATUS.md: header said 2026-08-23 over a 2026-08-30 body, and two "Waiting on you" items were both numbered 4 — both fixed. R-399 leaves that section (decided and shipped); the depth change is stated in plain words and the remaining items each say what happens if Viktor does nothing. 00-capability-map.md: the off-site verification row now carries its DEPTH, and its live citation is the 2026-08-31 run at 100%. The weekly firing at the new depth stays IMPLEMENTED, not PROVEN-LIVE. 07-backup-architecture.md §10.2: R-399 recorded closed, with the one sentence that stops it being turned back down — the structure check PASSED a size-preserving pack corruption. R-87 untouched and still OPEN. Register: R-399 and R-400 compressed into CLOSED-ITEMS.md with their reasoning kept and 300d7e8 named as the commit holding the originals. R-401 filed with a TRIGGER (the slow-check WARN firing) rather than a date. R-402 filed: the integrity verdict and its depth are on the wire and no hub surface reads either. OPEN 166 -> 165, CLOSED 148 -> 150. wire_contract_gate.py: offsite.last_integrity_depth allowlisted WITH ITS REASON beside its sibling last_integrity_ok, both to be deleted together when a hub surface is built (R-402). |
||
|
|
99af997ab9 |
R-359/R-397 closed, R-398 corrected, R-399/R-400 filed with measured numbers
gates / gates (push) Failing after 18s
THE MEASUREMENT IS THE STORY, and it re-frames the row it was filed under. A pack was corrupted WITHOUT changing its size; plain `restic check` -- the depth that ships ON -- returned `no errors were found`, exit 0. Only --read-data caught it. So the check that shipped verifies the index, the pack inventory and the snapshot graph, and does NOT re-hash pack contents. R-399 was filed as a bandwidth-and-cadence question; it is more than that, and its row now says so. R-399 gets three MEASURED numbers instead of estimates: store 140 829 678 B / 2651 blobs / 67 snapshots; structure check 35.0 s; curve 10% 35.9 s, 50% 37.3 s, 100% 39.2 s. At this size re-reading everything costs four seconds more than reading none, because the wall clock is SFTP round-trips not transfer. The row states the limit too: these do NOT extrapolate. R-400: the sweep the task asked for found EIGHT dead debug buttons, not one. 24 endpoints referenced in debug.html, 17 dispatched. Single dispatcher, exact match, default NotFound -- so they 404. A third of a debug page does nothing, on the surface an operator reaches for when something is already wrong. R-398 is CORRECTED AND LEFT OPEN, not closed. I filed it yesterday saying resticStep is not a seam so no test can drive a restic path. The layer below it has been injectable since the off-site tier shipped. The row survives as the record that the seam EXISTS so nobody re-files it. 07 gap register: R-359 and R-397 closed; R-87 restated IN PLACE as "AND IT IS NOT R-359" because the two rows are adjacent and a check is not a restore-test. 08 alarm ladder: both event types recorded, including that `ok` is `info` and therefore mails nobody BY DESIGN, and that all three registers were checked and deliberately left alone. 00 capability map: PROVEN-LIVE for the check, the notifier and the hazard control; the scheduled firing is IMPLEMENTED only, because a week has not passed. wire_contract_gate: `offsite.last_integrity_ok` allowlisted WITH A REASON. The gate was right -- the controller emits a field no hub struct can decode. Building the display is a hub change and R-331 ruled that class the operator's decision; the entry says to delete it when a surface exists. This push used `git push --no-verify`. golden-currency is CONVICTED and right: 0.227.1 is released and the golden carries 0.226.1. A BYPASS, not a waiver, and the task spec directs it -- golden and fleet delivery are Viktor's (R-242). It is item 3 under "Waiting on you". Register 163 -> 165 -> 163. |
||
|
|
e027b5d999 |
Register + architecture for controller v0.226.0 (R-353/357/358/360/396), and R-395 fixed
gates / gates (push) Failing after 17s
Closes R-353, R-357, R-358 and R-360 with their shipping version and evidence path, and files two new rows. R-396 (NEW, closed by the same release) is what answering R-358's open question turned up, and it is worse than the question assumed. The spec asked whether a unit-only scratch is reachable through the real UI flow. It is, by the SAFEST action on the page: "Ellenorzo visszaallitas" (mode=unit, advertised non-destructive) calls RestoreOffboxScratch(full=false); offboxRestoreScratchDir IGNORES `full`, so both modes write the same directory, and --include limits what restic extracts, never where; the wizard derives BOTH PlaceEnabled and RestoreEnabled from one ScratchReady flag. So a customer who ran the safe restore was then offered the destructive one over a unit-only copy. One boolean drove three different intents and the weakest set the answer. R-395 (filed by the spec) is fixed in this commit, not just recorded. STATUS.md said golden 0.223.0 / floor 0.222.0 in one block and demo-hp 0.219.0 / floor 0.218.0 fourteen lines below, cross-referencing an item that said "Nothing else". The fix REMOVES the duplicate rather than correcting it -- the same fact was written twice with no link, and only one copy had a reason to be touched during a release. "What works" now points at the item above instead of restating a version. 07-backup-architecture: four rows added to the 10.2 gap register plus R-396. Section 8 matrix row 3 KEEPS its PROVEN status, with the reason stated: R-353 was a defect in the MESSAGE, not the mechanism. The restore always returned what the unit held; what it could not do was say so. A status that measures whether data comes back must not move because a status line was wrong. 00-capability-map: one new row, and it splits what is claimed. R-353's sentence, R-358's marker and R-360's refusal are PROVEN-LIVE with a live citation. R-357 is IMPLEMENTED ONLY -- filling a real filesystem is a drill step, not a build step. R-353's Scenario B was ALSO not reproduced live and says so: no app on demo-hp still has a data-less unit, and falsifying a manifest to make one is the hand-set-state shortcut this project forbids. This push used `git push --no-verify`. golden-currency was CONVICTED and it is RIGHT: three controller releases (0.224.0, 0.225.0, 0.226.0) and the golden still carries 0.223.0. A BYPASS, not a waiver, on the operator's standing ruling from earlier today, re-checked rather than assumed -- all three are invisible to a day-0 box, and a restore-surface fix in particular has nothing to act on there. The ground expires the moment a release changes first-boot behaviour. Tracked on R-242; ONE bake carrying 0.226.0 covers all three. |
||
|
|
55274d5ef3 |
R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s
The gate failed only on `released > baked`, so it could catch a forgotten bake and nothing else. A golden AHEAD of the record passed silently - and that is how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG heading still read v0.221.0, with every gate green. Reproduced on the real history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0. The gate now asks whether the version being shipped is WRITTEN DOWN: the baked version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG. Membership rather than `baked > released` deliberately - a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2) preserved; every refusal names a reason and a route. Red-proofed both directions: old gate/old record exit 0, new gate/old record exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0. 08-alarm-ladder.md is new, and its absence was itself the finding: no document owned "when does a broken app raise an alarm?". The rules lived as comments in four packages, each locally correct, with the ordering between them legible only by reading one function top to bottom - which is how R-384 survived review. R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed closed. R-386 filed OPEN: a single-container app stopped out of band raises no alarm, and a comment claims the opposite - measured live, 9 scans, 0 events, against a positive control from the same box 17 minutes earlier. Not fixed here. Golden 0.222.0 baked and published; vouching is the operator's act. |
||
|
|
1eb64bec51 |
R-361 docs: the [FACT], the negative that cancelled Part 2, R-383/R-384, golden 0.221.1
gates / gates (push) Successful in 17s
07-backup-architecture.md gains a dated [FACT] on R-361 - a comment asserting an invariant the code did not have, for four months - and a [DESIGN] on the db_dumps decision INCLUDING the trap it created: a stable list lets the already-current early return fire, so per-capture housekeeping must sit above it. 00-capability-map.md records the NEGATIVE from Part 3 so it is not re-derived: a held app does NOT raise the dead-app alarm. It aggregates to unhealthy, which IsDownState excludes. Measured on the shipped build with the scans demonstrably running over it. No suppression was built and no row opened. R-383: the double-failure message names an undo copy that is not there - R-361's own class, one surface over, observed on both 0.220.2 and 0.221.1. R-384: an app whose database has died reads unhealthy and raises no alarm. R-361 closed and compressed. OPEN-ITEMS 325236 -> 327266 bytes. Golden 0.221.1 baked, published and round-trip verified. The golden-currency gate blocked this push and that block is not circular, so it was satisfied rather than bypassed - no --no-verify anywhere in this session. |
||
|
|
4e488321bf |
DRILL R-356b: the off-site restore for a driveless app that HAS a database
gates / gates (push) Successful in 16s
A drill, not an implementation. No code, no version bump, no CHANGELOG entry. Ten of the forty driveless apps carry a database; I re-measured that count and got 10. For those ten the restore is a five-leg operation that never ran at all until this week, because R-356 refused before any of it started. Walked end to end on demo-hp for both engines - docmost (Postgres 16) and bookstack (MariaDB 12.3) - each deployed for the drill, planted through the app's own interface, destroyed for real, restored through the endpoint the UI posts to. Q1 does it complete: YES. All five legs ran and succeeded, 32s / 25s. Accented names byte-identical both directions. Q2 which leg won: the SQL DUMP. Three-way discriminator returned the altered dump's value. This confirms R-164's F17 ordering on the OFF-SITE path; R-164 only ever cited the local one. Scratch-only mutation; store proved unmutated. Q3 does a failure tell the truth: partly, and two defects. Filed R-379 (HIGH, the undo copy is valid, named, and unappliable by any product action - proven by applying it by hand on both engines), R-380 (HIGH, a failed MariaDB replay leaves a partial database behind an app reporting healthy, where Postgres crash-loops visibly), R-381 (MEDIUM, the failure message pastes engine stderr including customer table rows into the Hungarian surface), R-382 (LOW, the summary log omits the volume count it already has). H1, H2 and H4 did NOT fire and that is recorded. H3 fired in a shape nobody predicted: not a quiet success, but a loud error over a silent inconsistency. R-361 reproduced independently on a second app. restic check: no errors, 29 snapshots. A flaw in the drill's own planting - a double-escaped accented title - was caught by reading stored bytes as hex, recorded, and re-measured in Phase 1b. Register 325236 -> 330683 bytes. Nothing dropped. Teardown: two apps retained with reason, no pvesm before-snapshot taken (said plainly), no hub-side record created. |
||
|
|
ef6ac6fe74 |
One register, enforced by a gate; closed work compressed into siblings (R-376..R-378)
gates / gates (push) Successful in 16s
Records and process only. No machine contacted. ONE REGISTER (operator ruling). 17 roadmap rows moved into OPEN-ITEMS.md keeping their identifiers, evidence and original filing dates - the oldest R-10, filed 2026-07-15, 38 days. 15 ideas stay in ROADMAP.md, which is their home; the gate exempts them by their own state word. 59 already-closed rows stay as history. Sorting rule recorded in the roadmap header: does the item assert something about the shipped product a reader could check and find false? scripts/one_register_gate.py, wired as the 11th gate. Control run: baseline passes, a planted open roadmap-only row is convicted by name, removing it passes with the file byte-identical, and a planted `idea` row is correctly exempt. Its four residual holes are in its docstring. The gate earned its keep immediately: it caught R-103, a READY finding my hand-sort mis-read as done because my regex matched the whole row where the body contains "shipped" - the gate matches the state cell. It also caught R-203 and R-163, recorded closed in the register and still open in the roadmap; the roadmap copies are marked SUPERSEDED with the register's verdict. HOUSEKEEPING. OPEN-ITEMS 672,376 -> 327,109 bytes (-51%); ROADMAP 239,306 -> 78,110 (-67%). Closed work compressed to 17% into CLOSED-ITEMS.md and ROADMAP-HISTORY.md; every entry names the commit whose git show returns the full original text. Rule-sentences are kept verbatim under "Reasoning kept" rather than judged entry by entry - 25 carry one. CONTEXT.md deliberately NOT compressed and the disagreement is argued in the report: 86% of it is standing rulings still in force, this prompt's own 3.4 says the log is never edited, and it has no per-ruling delimiter. Filed as R-377 - the problem is navigational, not volumetric. The hot/bulk placement decision was NEVER recorded as a decision anywhere - established, not assumed. Now marked [DESIGN] with a pointer honest about having no original date, given a decision-log entry that records what was rejected, and the [DESIGN]/[FACT] legend carried from 1 of 8 architecture documents to 8 of 8. Existing statements deliberately left unmarked (R-376). PROMPT-TEMPLATE gains N.7: compress what you closed, rehome live reasoning before it goes, state the register's size before and after. Ceiling R-375 -> R-378. |
||
|
|
d895d9f7dd |
STATUS + capability map: narrow the end-to-end off-site claim to the leg it was proven on
gates / gates (push) Successful in 16s
The 2026-08-04 row claimed "a customer's file survives a machine rebuild and comes back — the whole off-site story, end to end". Tonight's drill shows that holds for the declared-userdata leg of a drive-declaring app and for nothing else: the off-site restore has no named-volume leg (R-354), and refuses outright for the 40 apps that declare no data drive (R-356). Since that class keeps ALL its data in named volumes, the end-to-end story is unproven there and disproven for the volume leg generally. The escrow/key half of the row is untouched and still stands. STATUS.md also corrects the fleet pair it still named (0.214.0/0.129.0 -> 0.217.0/0.130.0) and records that demo-hp's off-site had been silent since 9 August. |
||
|
|
ab2262c91c |
hub v0.106.0: report loss of visibility into the off-site stores (R-339)
gates / gates (push) Successful in 14s
THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.
REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).
THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.
Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
- the unreachable event REPEATS rather than escalating once. The band shape
would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
mail is missable. It leans on the dispatcher's 1 h operator cooldown to
become an hourly "still blind" heartbeat.
- ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
which means we reached it. Counting it would alert for days on a healthy
pre-update endpoint.
- born-blind is reported: the counter is not gated on having a snapshot, so
a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
than zero-valued -- a fabricated timestamp reads as "it was fine until
then".
Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).
Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.
Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
|
||
|
|
767960bb11 |
docs: disk-health phase 1 — capability map, roadmap arc, register rows R-328..R-333
gates / gates (push) Successful in 14s
- capability-map: the disk-failure scenario no longer says a failing disk has never been seen. Healthy path + delivery + the severity wire stay PROVEN-LIVE; the new Hiba-from-counters path is IMPLEMENTED and explicitly NOT proven-live (R-332), because it has only ever run against the fixture's values. - ROADMAP R-73: phase 1 shipped; the genuinely hub-side half splits into R-330 (phase 2, a declared wire change under G-1) and R-331 (phase 3, growth-rate detection and retiring the static 64). Its premise 'no demo hardware exposes real SMART' is retired — a real failing drive is now committed as a fixture. - register: R-328 (the severity drop, CLOSED and proven live side by side), R-329 (app_start_failed has the same defect, needs a decision first), R-330/R-331 (phases 2 and 3), R-332 (the Fail path has never fired on real hardware, WATCHING), R-333 (NVMe temperature bands measured 2 degrees from tripping on a healthy drive; and the agent's smartctl has no -n standby). |
||
|
|
4d6ec7c7bb |
hub v0.104.0: the guest network gets a reader (R-319), and the hub half of the naming (R-295)
gates / gates (push) Successful in 14s
Four paper debts and one fact given a reader. Hub-only — nothing to bake. A4 — the entry about "the tester's machine" named a risk correctly and labelled it in a way that invited deleting it. Established from the hub's own store: `peti-felhom` is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47 this morning. The prompt's premise conflated the two; the register now says which is which. A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out of STATUS's "Waiting on you", which is now empty. A3 — day0-install §C.1 said pushing the installer publishes it. It has not since R-110. Corrected, with the two manifest pins named and an outside-verification command; the one copy that repeated it (a dated audit, true when written) carries a superseded note. A5 — standing rule 5: evidence comes off the machine at the end of the phase that produced it, before any revert. Earned twice in three days on the same box at the same point (R-320). Four homes, plus what to do when it is already gone. R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll` mail kind so the mail names the page a REBUILT box actually shows („A szerver beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach. Naming only; the acceptance pin proves the secret is untouched. R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never drawn as healthy — three absences, three sentences. No alarm, deliberately. Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190. B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now. Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled` reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed. Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect). |
||
|
|
6362bb6cb6 |
hub v0.103.0 — a host can read the packages we kept for it (R-311)
ListSupersededEscrow had zero production callers for nineteen days. It is the only reader of a retained identity_blob, so the retention shipped in v0.93.0 was material the product could not reach - proven on the fixture 2026-08-12, where a code that opens a retained package was answered as a code that opened nothing. New GET /api/v1/hosts/<id>/escrow/retained: self-scoped exactly as the current-row GET, same recovery-mode gate, same audit event written BEFORE the bytes leave, capped at 16. Rows with a NULL identity_blob are WITHHELD and returned as unopenable_count. They retain the PBS key, not the repository password, so they can never open what the caller is asking about; serving them would have the agent try packages that cannot succeed and would let the screen claim an earlier package is openable on exactly the boxes the original defect hurt. The count is returned because their existence is load-bearing and underivable. The trade, stated rather than waved through: the hub still cannot read any of it - sealed bytes in, sealed bytes out, no decrypt path, no recovery code ever held. What widens is volume, bounded by self-scope, the recovery-mode gate and the cap. The response is a NAMED TYPE, not a map, so the wire-contract gate can resolve it; the wire is declared as a fourth ROOT and the gate now checks 182 tags rather than 174. A positive control shows that check is name-presence, not decodability - filed as R-315 rather than reported as coverage. Six tests through the real endpoint; four red-proofs asserted applied. |
||
|
|
1d5f2b8bb6 |
DRILL: the retained key works, and the customer cannot reach it
gates / gates (push) Successful in 23s
Three verdicts, kept separate because collapsing them is how this assumption survived a week. (a) The material IS retained. host_escrow_superseded id 11 is the first retained row in fleet history to carry identity_blob (572 B), byte-identical to the pre-supersession row (sha256 a10032341c8584ed...). (b) The retained material DOES open the old store. Unsealed with the old recovery code it yielded a password byte-identical to the pre-change one, and restored three planted files byte-identical from a store the box itself could no longer open - including a Hungarian accented filename verified as raw bytes. Negative control ran first and failed closed. (c) The customer has NO route, and is misinformed. ListSupersededEscrow has zero production callers; the recovery path selects FROM host_escrow. Asked with the code that had just worked by hand, the product answered "the recovery code did not open the sealed bundle". A valid code for retained history is reported as a bad code - the R-224 class again. R-304, rank 1. Both installer faults were watched happening first, so installer-v1.27.0 is now published (tag + both webpage.yaml refs). Pre-fix: the box came up on controller 0.98.3 against a vouched 0.213.0, below the floor and below the version carrying the recovery screen; and our own uninstall left dnsmasq on 0.0.0.0:53 so our own next install refused. R-297 and R-300 CLOSED. Also filed R-305 (the dnsmasq fix fires once per machine - the leftover returns on the second reinstall, proven), R-306 (--preflight-only writes state it says it does not), R-307 (a live abandon countdown on demo-felhom, firing 2026-08-24 - operator decision), R-308 (stored controller password stale), R-309 (the day-0 runbook's publication claim has been false since R-110), R-310 (two edges). Ceiling R-303 -> R-310. Capability map moved: the retention claim is now marked operator-only. Phase A logs did not survive the intermediate revert; recorded. |
||
|
|
4a4a1e245a |
R-265 CI timeout + golden 0.210.0 baked; R-221/R-259/R-258 closed, R-266 minted, G-3 unblocked
gates / gates (push) Successful in 32s
Four defects of one family, all shipped today: something the box already knows, thrown away or drawn
as its opposite. Agent v0.128.0, controller v0.210.0. NO HUB CODE, no hub bump, no ArgoCD sync.
R-265 (this repo). timeout-minutes: 5 on the gates job — every honest run in the observed session
finished in 18-34s, so this is ~9x the slowest and far under whatever reaped run 264 at 834s with no
log. The alarm mail now carries Elapsed (start stamp via $GITHUB_ENV; an absent stamp prints
"unknown (no start stamp)", never a bogus 1.7-billion-second figure) and its "names itself in the run
log" sentence is qualified so it cannot mislead when there is no log.
⚠ THE UNKNOWN IS NOT CLOSED. Whether the if: failure() alarm fires for a REAPED job is still
unverified. The timeout makes the reap unreachable in practice; it does not answer what happens in
one. Demonstrating it means deliberately hanging a run on main, which would leave the branch red for
a parallel session. Said in the workflow comment, the changelog, R-265 and the report — none of them
claiming it is answered.
GOLDEN 0.210.0 baked, published, round-trip verified, NOT VOUCHED. The currency gate went red the
moment the controller was bumped — correct — and is closed by the bake, never --no-verify. No
--no-verify anywhere this session.
⚠ THE AGENT WAS NOT PUBLISHED UNTIL THIS SESSION CHECKED, AND IT MATTERED. R-221's fix is in the
AGENT, and a fresh install takes its agent from the Day-0 manifest. The binary had been hand-deployed
to felhom-pve and never published, so agent_version 0.128.0 was not selectable and a fresh install
would have received 0.127.0 — the golden would have carried the controller fixes and NOT the one the
headline defect needed. Caught by checking each Day-0 value was FETCHABLE rather than assuming.
Published from the live-deployed bytes, sha-verified across the hop first.
Registers. R-221, R-259, R-258, R-265 CLOSED. R-266 MINTED (READY): the failed root statfs still
travels to the hub as a 0-of-0 disk; ranked LOW because it is the quiet direction — it can only miss
a true alarm, never raise a false one — and it is now a two-repo wire change governed by G-1's gate.
Highest ID moved R-265 -> R-266.
CONTEXT S-39 rules the convention this project was missing: "we do not know" is never drawn as
"fine", and the codebase has ONE way of saying it — an explicit ...Known bool companion checked in
the template. ROADMAP G-3 was explicitly blocked on that decision and is unblocked; what remains
there is a survey-and-convert of existing sites, not the gate.
Capability map row 93 CHECKED and it was NOT claiming something untrue — it is about the operator
notification path. But its narrative ("the page you open to ask whether ONE app is backed up")
invites the wrong reading, and the adjacent thing WAS false until v0.210.0, so the row now records
that the two halves disagreed and only the operator half was true.
Six red-proofs across the two code repos, each with the mutation asserted applied. The one that
matters: Part 1 Scenario A FAILED against today's tree, with the intended message.
Part 1's operator-present live validation is OWED and is the session's STOP.
repo_gates --fast: all 8 OK.
|
||
|
|
b080ecf411 |
hub v0.99.0 — the hub can see whether the operator can get in (R-260); G-1 gate closes, R-247 closes
oobDegraded tested five things and the sixth never arrived. The agent has emitted `operator_key_configured` on every heartbeat since v0.72.0 — the SAME version that introduced the `oob` stanza carrying it — and store.HostOOBRow mirrored five of the agent's eight OOB fields. With no field for it, encoding/json discarded the fact on arrival, so a box with felhom-sshd active, reachable, a valid config and a configured peer reported `ok` with NO OPERATOR KEY INSTALLED AT ALL. Not a wrong answer: an answer to a question nobody was asking. `operator_peer_configured`, which the hub did read, only says the peer IP is in desired-state — that OOB is MEANT to work, not that entry is possible. Now decoded: operator_key_configured, plus wg_handshake_age_s and healed_at. The last two ride the ALERT TEXT and are deliberately NOT in the predicate — widening a check beyond the fact that is now arriving is how a check stops being read. SCENARIO F, decided on a measurement rather than a preference. operator_key_configured decodes as a POINTER: nil = the agent never said, reported distinctly and never as ok. The version gate was rejected because the field and its stanza shipped in the SAME agent version (v0.72.0), so a stanza without the field cannot come from any released agent; the fleet is 0.113.0/0.127.0 and the vouched floor is 0.127.0. Handled explicitly anyway and pinned, because "cannot happen" is a claim this project has been burned by. THE MESSAGE NAMES THE FAULT. oobDegradedReason is the single source for both predicate and text, so the alert can never name a different fault from the one that fired. The old form derived it separately and had a vocabulary of two — unreachable, or config invalid — with no way to say the key is missing. The operator reads this at 07:00. TESTS DRIVE THE DECODE BOUNDARY. Every hub OOB test before this built a HostOOBRow by hand, and a test written that way CANNOT SEE A FIELD THAT NEVER DECODES — which is how this held a green suite for five weeks. The pre-existing fixture oobReport() also omitted the field, so those scenarios ran against a report shape no released agent produces (same family as R-262). Both fixed. Red-proofs, 8 expected outcomes and 0 wrong, each with the mutation asserted applied: dropping the field returns the false ok; an unconditional check alerts a healthy box; unknown-as-ok restores the silent pass. G-1 CLOSED — scripts/wire_contract_gate.py shipped as ranked, built BEFORE the fixes and seen failing on 40 fields (documentation/tests/wire-contract-gate-2026-08-08/BEFORE.md). Two instrument defects the control caught first: a substring false negative (grep -F healed_at matched privsep_healed_at) and treating dr_recipe as wholly opaque when its top-level sections ARE decoded through an allow-list that already cost offsite_restic (R-122). The prompt for this session said "465 emitted tags, eight unreachable". Checked against the repo: R-260 said "at least eight DECISION-BEARING facts", never eight tags. The real count is 40. R-260 CLOSED (class gated, sharpest instance fixed). R-247 CLOSED (controller v0.209.0). R-264 MINTED and OPEN — the 21 facts with no consumer, allowlisted with reasons so that gating the class could not be mistaken for deciding them. Still open and named: R-246, R-255..R-259, R-261..R-263, and C7's test-comment half. Capability map checked: it claims OOB access is implemented, never monitored, so no row was untrue; what was untrue sat one layer down and the row now records it. repo_gates --fast: all 8 OK. go build/vet/test green in hub, run separately from this commit. |
||
|
|
c1dec41328 |
walk5 venue TORN DOWN — census 168 rows -> 67, and R-244 grew by 30 as predicted
gates / gates (push) Failing after 13s
Operator-confirmed. Stopped under a name guard (demo-hp carries its own 9201), aged past the hub's stale_threshold read from the DEPLOYED ConfigMap (30m), and polled delete-impact until deletable:true — treating an empty response as retry, never as success. Cascade + qm destroy --purge, guarded a second time. Every layer verified absent against a positive control that must survive and does: VM 300 drill-r50 and demo-hp's own guest 9201 still there; ep0 namespaces demo-felhom + demo-hp still there; wg peers .2 .3 .4 .250 still on the live wg0; hub rows for demo-felhom, demo-hp, peti-felhom untouched. 16.64 GiB returned against 17 G measured. RECORDED FOR THE NEXT TEARDOWN: the WG peer is removed on a ~5-minute SCHEDULE, not by the cascade. Immediately after the delete the hub row was gone while 10.77.0.5 was still on ep0's live wg0; wgsync had last run 37 seconds before the cascade, and the next push (4 peers) removed it, verified on the live interface at 16:57:07Z. The previous ledger checked this after it had already converged, so it read as instantaneous — a teardown that checks too soon would file a false finding. R-244 grew by 30 rows (app_log_issues), PREDICTED in the pre-run enumeration rather than discovered afterwards. Running total across torn-down venues ~101. Nothing here claims a clean teardown. Storage Box layer evidenced from the hub's own deprovision log: the HETZNER_API token in ~/.config/credentials cannot see box 611421 (subaccounts -> 404, storage_boxes -> 200 with 0 entries) — it is scoped to another project. |
||
|
|
3f4fb3825f |
R-201 CLOSED — the unaided recovery journey passes, both halves, on the fifth walk
gates / gates (push) Successful in 17s
Capability map: the unaided-recovery row turns FAILED -> PROVEN-LIVE, scoped, with what it still does not claim stated in the row itself: shape (c) did not fire positively (with the mint guard holding there is no local key, so the offer comes from shape (a)); and 'unaided' here means possible-without-a-shell, not obvious, because two obstacles are unsignposted. OPEN-ITEMS: R-201 closed with its evidence. Five new rows R-249..R-253 (the retrieval passphrase in page HTML; the host-key scan ladder vs AAAA settle; the listing's per-tag rows; the two unsignposted restore steps). R-243 annotated rather than re-filed: on a REBUILD offsite_delivery_stuck does not skip, so the row's gap is narrower than it reads. STATUS.md: headline changed, and trimmed 97 -> 92 lines rather than extended, per its own header. Teardown recorded as OWED with its before-measurements, the stop-and-age gate, and the positive controls that must survive. |
||
|
|
9657334fb7 |
R-241 FIXED: registers, capability map, STATUS, hub CHANGELOG v0.98.0
gates / gates (push) Successful in 14s
R-241 closed against controller v0.206.0 + hub v0.98.0, following the spike's ruling rather than the obvious reading. The row records what the fix does AND the two real bugs the tests caught rather than review - a missing t.Enabled (caught by an EXISTING test) and a missing falling-edge sync that reintroduced the very defect the epoch exists to fix. R-243 UPDATED, not closed: the STATE it describes can no longer be entered (the mint guard), and what replaces it is VISIBLE rather than silent - the box declares awaiting_recovery_key and the customer is offered the screen. But the ALARM GAP is untouched, for the same three reasons, so a box whose customer never acts still stops backing up with no operator signal. The remaining work is an operator-side signal for a box held past some age, deliberately not bundled into R-241's fix. R-245 NEW - WAITING-ON-OPERATOR, recorded and NOT built: should an undecided customer be auto-abandoned after 30 days? The operator's proposal is recorded WITH the reasoning against it, so the decision can be revisited properly: a reinstall implies a person, so nobody is absent; a customer who cannot find their code gets in touch, which is why the operator LEVERS were the thing worth building; the cost is the customer's own storage allowance; and the real harm is QUOTA, which is a condition, not a calendar. If it is ever built, build it to trigger on the harm with a dated warning, never on a date alone. The capability map's recovery-journey row STAYS FAIL. These are fixes, not a walk - nothing here walked a customer end to end, and the row goes green only when one completes with no operator intervention AND a byte-identical sentinel. R-214, R-202 and R-240 are still open. STATUS compressed rather than extended, per its own one-screen rule, and the "rebuilding throws away the off-site history" line corrected: the cause is fixed, so leaving it as a live defect would be false. hub CHANGELOG v0.98.0 for the superseded-package purge. Highest register ID moves R-244 -> R-245. |
||
|
|
db578cd44d |
R-239 CLOSED (golden 0.205.0 vouched); R-241 ruling into the map and STATUS
gates / gates (push) Successful in 14s
R-239: the operator approved the vouch this session. golden_version 0.203.0 -> 0.205.0 (+ derived sha); agent_version and min_agent both stayed 0.127.0, because the new golden's MinAgent is also 0.127.0 - so in the event it was a ONE-field change, not three. wrapper_sha256 was carried through explicitly: the handler reads it from the form and CLEARS it when omitted. Verified from the stored hub_settings (WAL-aware copy), not from the flash. The R-120 gate passed exactly - the newest controller the fleet reports is 0.205.0, so a 0.204.0 golden would have been refused. R-241: the capability map's recovery-journey row and STATUS carry the spike's ruling - a MINTING defect, not a screen-predicate defect. The row stays FAIL: delivery is not a journey, and R-241 is diagnosed, not fixed. R-242 and R-243 surfaced in STATUS in plain language. |
||
|
|
2228c0bff6 |
final walk COMPLETE — data PASS, journey FAIL; R-241 filed
gates / gates (push) Successful in 15s
THE DATA: PASS. All three sentinels byte-identical out of snapshot f5c53b03, including the 12 MB binary and the accented Hungarian filename whose NAME BYTES are identical too. Disk -> restic -> SFTP -> Storage Box -> rebuilt machine -> disk, intact. THE JOURNEY: FAIL, and further from the line than the previous walk. The claim worked first try (302 in 0.164s). Then: / lands on the launcher with no recovery pointer, /recovery 302s away, and the remote page offers to CREATE a new recovery code — which would orphan the history the customer's code protects. There is no field anywhere to enter the code they hold. The operator's documented remedy also refuses, correctly and fail-closed. Recovery needed three guest command lines. R-241 — and the cause is a success this same walk proved six hours earlier. OffsiteRecoveryOffer() shows the screen only when (a) there is NO repository password (pristine rebuild) or (b) one exists but the history will not open under it. Overnight the credential self-heal collected the staged credential and applied the tier, writing a FRESH key at 03:18Z — so (a) is false; and (b) is unreachable because orphan detection needs a run, and runs are blocked by escrow_state=pending. The gap is self-locking. Measured keys: on-disk 9b4a9a9d... vs recovered-from-R 30ef574f... This is R-218's shape one level up: succeeding at the self-heal stopped the box OFFERING the recovery it still needed. Registers: R-201 moved to its outcome; R-241 filed; capability map's recovery row stays FAIL with both halves and the cause named; STATUS rewritten for the operator. Highest ID R-238 -> R-241. The venue is left with the recovered key in place and the self-heal key moved aside, never deleted. Teardown still owed. |
||
|
|
feed748325 |
R-234 root-caused and CLOSED; R-218's state field corrected
gates / gates (push) Successful in 18s
R-234 was filed as "toggling an app on leaves it without a bundle, so the first run skips it". Measured on demo-hp: that state does not survive a run — the off-site run's own pre-dump phase calls captureAllRecoveryUnits for every DEPLOYED stack, through admitApp, before the push, and a unit moved aside was RECREATED. The actual cause was the single-flight: the manual run was dropped because an earlier one was still going, runOffboxBackup returned nil, the handler had already answered "A tavoli mentes elindult", and the card then showed the PREVIOUS run's green verdict. Fixed in controller v0.205.0 and proven live on demo-hp: a second request while one is in flight now says "Mar fut egy tavoli mentes — ez a keres nem inditott ujat. A most lathato eredmeny meg a korabbi futase", as a flash_error. Independently, and a real gap on its own: a run that skipped an app the customer selected is now `incomplete`, not `ok`. Selected+deployed with no unit counts; selected-but-undeployed is named with what to do but does NOT count, because a box left amber by an app somebody removed is a status nobody reads. R-218's state field read REOPENED while the same row's body already recorded the fix shipped in v0.203.0 and proven live. Corrected to CLOSED, keeping the over-claim history — it is why the row is worded as it is. Capability map: the off-site capture row's `incomplete` sentence widened to cover a whole-app skip, and it still does not claim a newly-selected app is protected by the next run — for a deployed app it is, for an undeployed one the card says so. Still open, deliberately: R-213, R-202, R-214, R-235. |
||
|
|
190c432f3a |
capability map: the recovery journey stays FAIL, with the two dead ends closed and vouched
gates / gates (push) Successful in 8s
R-218 and R-220 shipped (controller v0.203.0 / agent v0.127.0) and are proven live on a genuinely rebuilt box; golden 0.203.0 + agent 0.127.0 + min_agent 0.127.0 are vouched, so the delivery gap the re-walk recorded is gone. The row stays FAIL because the walk did not finish: it stopped at R-237 (the restore list was keyed on installed-and-toggled apps), now fixed in v0.204.0 and proven live — but no sentinel was restored, so the data half is unproven in either direction for that venue. R-238 reclassified (harness artifact, real residue fixed); R-236 withdrawn. |
||
|
|
0c4411e54b |
R-201 re-walk: the data PASSES again, the journey still FAILS — two dead ends, down from four
gates / gates (push) Successful in 9s
Asked Campaign 11 Phase 1's question a second time, on the fixed build, on a
NEW appliance (VM 322, customer rewalk). The Campaign 11 venue was untouched.
THE DATA: PASS. All three sentinels byte-identical out of the pre-destruction
snapshot a7bc23bd in 23s through the customer's own restore flow — including a
12 MB binary and an accented Hungarian filename whose NAME BYTES are identical
too (verified as hex, not as rendered text).
THE JOURNEY: FAIL, two dead ends against Phase 1's four.
1. R-218's CONSUME half. The hub re-staged the credential at 11:44:57 saying
'the box re-consumes on its next cycle'; a full cycle ran at 11:55:46/54
(with a positive control that it ran) and it did not. A census of the
customer-reachable actions found none that fetches it. Only a command line
INSIDE THE GUEST moved it — 18s, confirming nothing was wrong with the
credential, target or key: only the trigger. R-218's row said SHIPPED and
over-claimed; it is corrected to REOPENED for the consume half.
2. R-220. Drives still unenrollable after a rebuild, needing a Proxmox-host
unmount; without it no app redeploys and the restore page stays empty.
Unaided RTO STILL UNDEFINED. Attended: +45s key placed, +24m12s tier up,
+30m13s data verified. The 30m must not be quoted as the customer number.
What passed and is new: the recovery screen appeared WITHOUT being sought,
answered all three questions with a seal date matching the hub exactly, the
emailed reset code worked first try, the unlock was a real 1.528s unseal, and
R-225's fix was seen working in the wild (unknown, not a false zero).
R-216 part 4 reproduced live: the reinstall downgraded the hand-installed agent
0.126.0 -> 0.125.0.
DELIVERY GAP recorded as owed and NOT conflated with the journey: a fresh
install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions,
neither carrying the fixes — installed by hand. Nothing was vouched.
Capability map row STAYS FAIL. Campaign 11 doc gets a dated ADDENDUM, not a
rewrite.
|
||
|
|
d30c2a51ed |
R-224..R-228 CLOSED: registers, capability map, campaign annotation, STATUS
gates / gates (push) Successful in 7s
Five closed in controller v0.202.0 + agent v0.126.0, each with its live or red-proof evidence in the row. Five explicitly still open and named as such rather than left to inference: R-214, R-220, R-221, R-213, R-202 — and R-220 is flagged as currently worked around BY HAND on the campaign venue, which is the only reason an app could be deployed there. The capability map's recovery row STAYS FAIL and says why: fixes are not a re-walk, nothing walked a customer end to end, and the customer-facing messages were NOT re-driven live because /recovery correctly retires itself once the old data is set aside — restoring that state is the reconfiguration the task forbade. The campaign document is ANNOTATED, not rewritten: it records what was true when it ran, and that is its value. workspace-CLAUDE.md gains comment-vs-code entry 9 — the escrow header said the errors were 'DISTINCT on purpose' and named THREE situations while a fourth was folded into one of them, and a green test named the defect and did not prevent it because it asserted a STRING one layer below the merge. ROADMAP needed no collapse — it carries no rows for these IDs. |
||
|
|
453e4503a9 |
capability map: the recovery row stays FAIL — what Phases 2+4 add, and what they do not
gates / gates (push) Successful in 9s
Says explicitly that these faults are NOT a re-walk, so the row cannot go green on them. What they add: the BACKUP promise strengthened (the offsite tier ran itself at 04:15 on a twice-rebuilt box, snapshot_count 1->2; all five daily jobs fired once; nothing on the must-not list fired), R-217 and R-215 proven live under exactly their faults, and the set-aside proved not to delete (12 535 KB byte-exact at the far end). What they do NOT add: any progress on the JOURNEY. R-224 is Phase 1's headline defect relocated from the version channel to the transport — a hub outage and a stopped agent are both reported as a bad recovery code, in 0.056s and 0.030s, with no unseal attempted. Plus R-225, R-226, R-228. Also records that §4.1 is now MEASURED rather than deduced, and §4.2's positive half still is not. |
||
|
|
1a0f7db92f |
docs: CAMPAIGN-11 — the journey FAILED, R-198's retention PROVEN, six findings fixed
gates / gates (push) Successful in 8s
Registers and evidence for the campaign and its fix pass. OPEN-ITEMS: R-214..R-223. Six SHIPPED (R-215/216/217/218/219/222); three deliberately still open and each blocks a real flow (R-214 console banner, R-220 drives unenrollable after a rebuild, R-221 a rebuilt box cannot run the escrow ceremony); R-223 minted and WAITING-ON-OPERATOR (vouch agent 0.125.0). R-213 and R-202 untouched. Capability map: a new row for the customer's UNAIDED journey, recorded FAILED and staying failed until a re-walk passes — fixes are not a journey. The existing rebuild row is corrected where it said R-198's retention was unit-proven only: it was proven in production on the first supersession since the fix, identity_blob retained at 572 B byte-length exact. CLAUDE.md comment-vs-code table: eighth entry — ResolveManagedFloor, the first where the false invariant was a GUARD rather than a comment alone. STATUS: the headline is now "the backup promise is proved, the recovery journey is not", and the one thing waiting on the operator. |
||
|
|
f45b1f6761 |
docs: R-193 CLOSED (the recovery screen); R-213 minted for the put-back
gates / gates (push) Successful in 7s
- OPEN-ITEMS: R-193 CLOSED with both 2026-08-05 rulings (unlocking and restoring are separate; 'I do not want the old data' moves the store aside after a double confirmation), and the shape-(b) reasoning — WriteOffboxSecrets auto-generates a repository password on re-apply, so the literal 'fresh data area' trigger would have opened a window that closes by itself. - R-213 MINTED (R-212 was and still is the highest, re-checked for the second writer): putting files back in place, with the live-versus-backup comparison named as its requirement. Not started, deliberately. - capability map: the 'needs someone who knows to look' qualifier is GONE; what remains is stated narrowly — no correct-code run through the page, the put-back is out of scope, and the journey has not been re-walked end to end. - 07-backup-architecture 7.0: a fifth row, and where the screen deliberately stops. - CONTEXT: standing ruling S-34. - STATUS: the headline change and the two things still owed as proof. No hub change and no hub bump. |
||
|
|
4faebe2926 |
docs: R-204 ALL FOUR items closed; R-193 credential half; R-192 by replacement; R-212 filed
gates / gates (push) Successful in 8s
- OPEN-ITEMS: R-204 all four CLOSED with both 2026-08-05 rulings recorded (the declared-state trigger and its four-meanings-of-absence reasoning; the recovery preview's dashboard-password exposure accepted as metadata, not content). R-193's credential half CLOSED, screen + deletion still open. R-192 CLOSED by REPLACEMENT. R-202 untouched. - R-212 MINTED (R-211 was the highest, grepped): the orphaned-ciphertext deletion HALTED at its STOP because the measured paths do not match the register — three set-aside stores totalling ~1.45 GB, and the thing that is exactly 1.2 GB is demo-felhom's LIVE repo. Nothing was deleted. - capability map: all four interventions closed; the row KEEPS a qualifier for a new reason — no step needs an operator, but there is no customer-facing recovery screen, and the journey has not been re-walked end to end. - 07-backup-architecture 7.0: the four-step table updated; the declaration-vs- inference reasoning and the credential-automatic/key-customer-present split. - CONTEXT: standing ruling S-33. - STATUS: the headline change and the deletion STOP. - REPORT-r204-item4.md rather than REPORT.md: a parallel session is active in this shared clone. |
||
|
|
0dbd954fec |
docs: R-196 closed, R-204 items 1-3 closed, item 4 open (R-193)
gates / gates (push) Successful in 7s
- OPEN-ITEMS: R-196 CLOSED; R-204 items 1-3 CLOSED with item 4 named and its dependency stated. Header restates that R-202, the 1.2 GB ciphertext deletion and R-198's still-unit-proven retention all REMAIN OPEN. - capability map: the recovery row keeps its 'with a person present' qualifier, names which crutch remains, and cites the three now gone. - 07-backup-architecture: new 7.0 - what a customer can and cannot do ALONE, the four steps in a table with status. This is the section a future reader will use to answer that question. - CONTEXT: standing ruling S-32, superseding S-31 steps 2-5. - STATUS: rewritten to one screen per its own header; removes a corrupted half-overwritten section left from the drill session. - ROADMAP: R-196 and R-204 collapsed. |
||
|
|
2a7ac03c47 |
R-201 PASSED: a customer's file survived a machine rebuild and came back byte-identical
gates / gates (push) Successful in 7s
|
||
|
|
b228fd102d |
R-201 night run: the off-site key IS recoverable after a real rebuild (proven); the verdict is blocked by R-204
gates / gates (push) Successful in 6s
|
||
|
|
73fb595e38 |
R-203 shipped: the app and its backup agree, and 'ok' means it — R-201 unblocked
gates / gates (push) Successful in 7s
|