5e8a82c3c47b071112c51e8a355fdafdc42d7dc1
125 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5e8a82c3c4 |
night 2026-09-13/14: first "be a customer" rotation (adventurelog) — 7 defects found, 13 rows closed
gates / gates (push) Successful in 18s
New runbooks/nightly-rotation.md; observations_gate.py reads every section (R-471); target-selection.md names real paths (R-461); R-93 carries the fact that drill-r50 is gone. Register: R-473/R-474/R-466/R-471/R-453/R-461 and v0.240.0's R-477/R-478/R-480/R-482/R-484/R-485/R-486 closed; R-481, R-483, R-487, R-488, R-489 opened. 09 §6.1, 07 §6, CONTEXT, STATUS note. Evidence: audits/nightly-2026-09-13-adventurelog/, audits/v0240-2026-09-13/. |
||
|
|
681c3d6a6d |
docs: rulings 7 and 8 shipped and proven live (R-470/R-472/R-475 CLOSED); R-477..R-480 opened
gates / gates (push) Successful in 21s
Hub v0.112.0 serves a floor above the golden with a declared MinAgent;
controller v0.239.0 reached both demo boxes by that floor in 14 s and 15 s
and updates on any backup tier. 09 §3 decisions 7 and 8, §6/§6.1; 07 §6
line; capability map row; STATUS items 15/16 done and the cadence line
corrected; CONTEXT; register: R-470/R-472/R-475 compressed to CLOSED-ITEMS
(full text at
|
||
|
|
5ef0f52bcd |
Slice 4 shipped (R-448/R-443/R-439 CLOSED, proven live); R-472..R-476; the floor-between-bakes claim corrected
gates / gates (push) Successful in 19s
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure. Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/). Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the runbook, STATUS, CONTEXT, R-468 and the gate docstring. Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
ae59c31a84 |
R-459 CLOSED (MariaDB converts itself, proven by harness + live), golden 0.236.0 (R-467), the golden waiver (R-468)
Operator rulings 2026-09-13, both shipped the same day: - MariaDB finishes its own conversion (catalog eec1228/bd32830/3525e35). Harness E3/E3b `proven` with engine_state_after "already upgraded to 12.3.3-MariaDB [exit=1]", the skip line gone, C3 still `failed`; landed on demo-hp through the real 15-min cycle, nothing recreated, one deliberate restart logged "MariaDB upgrade not required" with the app serving. Evidence: documentation/audits/r459-close-2026-09-13/. The engine-major rule + gate keep every engine inside its major until Slice 4 (R-448) — removal tracked as R-469. - Goldens on a cadence, not per release. golden_currency_gate.py reads a dated waiver (documentation/tests/golden-waiver.yml, <= 14 days, row-bound): valid + BEHIND -> loud advisory, exit 0; expired -> red again naming the date; UNRECORDED (R-385) never covered; malformed -> 2, never 0. Tests cases 5-15 incl. the R-421 decoy; red-proof old-vs-new on the real behind tree. R-242's vouch half stays open. Cadence in RUNBOOK-manual-build.md §4.2 + the checklist. - Golden 0.236.0 baked, round-tripped, vouched, floor raised 0.232.0 -> 0.236.0 (documentation/tests/golden-0.236.0-2026-09-13/) — the last per-release bake; the waiver was issued AFTER it landed. No --no-verify anywhere in this session. Rows: R-459 CLOSED, R-467 CLOSED, R-242 narrowed; R-468/R-469/R-470/R-471 opened. 09 §3 gains decisions 5 and 6; STATUS items 11 and 12 closed; CONTEXT records the cadence ruling. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
4b2e5608c2 |
R-442 CLOSED (controller v0.236.0): "delete my data too" deletes it or refuses; R-465..467 opened
gates / gates (push) Failing after 18s
- OPEN-ITEMS: R-442 row removed; R-465 (six remaining Paths.HDDPath readers — audit), R-466 (recovery-unit residue after "delete backups"), R-467 (v0.236.0 owes a golden). - CLOSED-ITEMS: R-442 compressed, reasoning kept, fleet shape now established. - 00-capability-map: lifecycle row narrowed (Campaign 3 proved remove removes the APP, not the data) and re-proven from audits/R442-2026-09-13/. - STATUS: item 13 in plain language; item 7 closed (ruled 2026-09-02, 09 §3). - audits/R442-2026-09-13/: live evidence (A, C, D bodies, controls, log window, teardown). Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
d6837d98ee |
SPIKE R-459: the skipped MariaDB conversion is stable, and the trade it implied does not exist
gates / gates (push) Successful in 20s
Outcome A, qualified. Not B and not C. It does not degrade: 5 of 5 restarts of 12.3 on an 11.6 datadir, readback passed every time, mariadb_upgrade_info unchanged, the entrypoint line never escalated past [Note]. It also never heals - the engine answers 'Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!' on every start and will forever. The trade R-459 was expected to produce is not real. Converting properly SUCCEEDS across the multi-major jump, takes 7 seconds, backs up the system database unasked - and putting 11.6 back afterwards STILL starts and serves the data. So the operator is being handed a cheap correction, not a choice between a correct engine and a reversible one. The exit-code polarity was measured rather than read: 0 means the upgrade IS needed, 1 means it is not. Assuming either the flag name or the polarity would have inverted the headline. And run without credentials the same command returns a confident-looking FATAL ERROR that is an auth failure. R-464: after converting and going back, the entrypoint prints 'MariaDB upgrade not required' on a state the same engine calls an unsupported downgrade. The obvious cheap instrument for R-459 would have been to grep for that line, and it would have reported fine for the broken case. R-463: the PostgreSQL analogue, deliberately NOT measured here. 11 templates, 8 on postgres:16-alpine, register grep for pg_upgrade returns zero. The two engines fail in OPPOSITE directions - MariaDB skips quietly, Postgres refuses to start - so that one cannot hide; it presents as eight apps down at once. No template changed. Teardown all three layers, hub checked rather than asserted, local-lvm 30.53 percent before and after. |
||
|
|
a1a6c73fe1 |
SPIKE: an upgrade test that runs again — and a real defect in our own bookstack template
gates / gates (push) Successful in 19s
R-449. Until today one app upgrade out of 53 had ever been measured, by hand, and the whole update arc was designed against that single data point. C3 first: the negative control, whose TO image exits immediately, came back failed. That is what makes the greens mean anything, and it cost 556s because a negative is only honest if it waits out the full settle window. Seven edges, three apps. All five real catalog upgrades kept the customer's data. The finding that changes an assumption the arc was carrying: whether an upgrade can be UNDONE is a property of the individual APP, not of upgrades. Docmost refuses - 'corrupted migrations: previously executed migration 20260213T085259-notifications is missing' - and privatebin does not. That reproduces the Nextcloud result on a second app by a DIFFERENT mechanism, so the struck word 'rollback' now rests on two measurements instead of one. The finding nobody was looking for, R-459: our own bookstack template moves MariaDB across a major and sets no MARIADB_* env at all, so the engine logs that the datadir upgrade it requires is being skipped, and serves anyway. The cause is assigned rather than guessed - the app half alone produces no upgrade line, both edges that move the engine produce it - which is exactly what decomposing E3 into E3a and E3b was for. It also explains why E3's abort looked like it worked: the datadir was never converted. Whether that ever breaks is NOT established, and the row says so. Also opened: R-460 (bookstack's file half cannot be seeded headlessly), R-461 (target-selection.md names a venue that does not exist and fences a VM that is gone), R-462 (the widening, costed with this run's real numbers - and the cost is dominated by fixtures, which do not amortise). Teardown all three layers, hub checked rather than asserted. local-lvm read 30.50 percent before and after. The capability map was deliberately NOT edited: this measured apps, not the product. |
||
|
|
417df06f35 |
slice 3 docs: the ruling, the shipped mechanism, and four rows closed
gates / gates (push) Successful in 17s
09-update-architecture.md gains the fourth dated operator ruling (2026-09-06, Option 1) and its section 5 is rewritten from a proposed shape into the shipped one: the pin, the stored definition, the render table, the four writers, the startup ordering, and the trap this slice set for slice 2 - the live compose file is now the frozen one, so a badge comparing against it would answer Naprakesz on exactly the apps that are behind. 02-controller-module-map.md said 'copy compose + .felhom.yml'. That stopped being true today, so it is corrected, and the two sections describing the old seam now carry a banner saying they describe v0.234.0 and below - kept because every box under v0.235.0 still behaves that way and because they are the measured account of why it changed. R-447, R-441, R-438 and R-455 closed and compressed into CLOSED-ITEMS; R-458 opened for the .felhom.yml asymmetry, with what would settle it by measurement. Live evidence: two real catalog pushes travelling the real 15-minute cycle, both reverted, the tree byte-identical afterwards. The restart that used to take 18.3 seconds and pull a new image now takes 0.1 seconds and pulls nothing. |
||
|
|
bc47dd4ef9 |
v0.234.0: a known limitation written on 2026-09-02 was a defect by the next morning
gates / gates (push) Successful in 18s
The operator looked at demo-felhom and found OpenGist - up 15 hours, running exactly the catalog pin, showing no badge at all. 09-update-architecture.md had recorded that as an accepted limitation the day before: 'the fleet view fills in gradually'. On a quiet box gradually means never, and a feature that fills itself in on an event nobody triggers is, on the quiet installations, not shipped. That limitation row is now struck with the reason kept. The living document gains slice 1b, the two admission rules of the backfill (it never overwrites, and it refuses to seed a partial observation because the badge reads a service-count mismatch as BEHIND), and the note that the same field having two writers with two different admission rules is deliberate. Live evidence added: all nine apps already had records by the time 0.234.0 was ready, so the natural fleet state could no longer exercise the new code - said plainly rather than papered over. The pre-0.233.0 shape was recreated on demo-hp by stripping two records; the backfill re-seeded exactly those two with digests matching independently-read ground truth and left the other seven alone. The refusal half was deliberately NOT staged live: it needs a degraded app, and manufacturing one risks the false-customer-email class that already cost 61 mails (R-330). Unit-tested with a red-proof, and recorded as unproven-live. R-457: a test that hardcodes a date and asserts an age derived from it is green only on the day it is written. Mine was, and it went red overnight. Six other files carry both a date literal and time.Now() - named as candidates, not accused. |
||
|
|
0705942783 |
the badge IS proven live, and the 'stale password' finding was mine, not the box's
gates / gates (push) Successful in 16s
I reported that the vaulted dashboard password no longer worked on either demo box, and quoted the controller's own 'Failed login' as the discriminator. The password was fine. ~/.config/credentials quotes its values with SINGLE quotes and my sed stripped only double quotes, so the quote characters went out as part of the password. The operator corrected it in one line; one retry returned 302. The instrumentation lesson is the finding and R-453 now carries it: 'Failed login' separates wrong-password from wrong-Host-header, and that is ALL it separates. It cannot tell a wrong password from wrong password HANDLING, and I read it as if it could. This is the second time this file's quoting has produced a confident wrong verdict, so the fix is one shared extraction helper, not a resolution to be careful. With the session recovered, the badge is validated on live pages: Naprakesz twice on /stacks and on /apps/bookstack; NO badge at all on /apps/docmost (a deployed app with no record - absent is UNKNOWN, not current); and 'Frissites elerheto - 52 napja' on both surfaces, the age being real arithmetic on bentopdf's catalog_since. The behind state was staged by editing one compose tag, with no restart and no up -d, and reverted byte-identically (sha256 equal, diff empty, container never touched). Capability-map row upgraded to PROVEN-LIVE with the one unexercised badge state named. STATUS item 9 now needs nothing from the operator. |
||
|
|
6035dfcc3a |
09-update-architecture.md: the update path finally has a document, and it is a living one
gates / gates (push) Successful in 17s
R-438's document half. It records how an update works AS MEASURED, quotes the RestartStack comment that proves the restart half was CHOSEN (a design decision is not a defect), carries the three operator rulings of 2026-09-02, strikes the word 'rollback' (once a migration has run the old image will not start), states the target shape, and lists the seven slices with a status each. R-438 and R-440 amended and BOTH STAY OPEN: the mechanism is documented, not changed. Nothing closed, so CLOSED-ITEMS.md is untouched. Eight new register rows, 194 -> 202: R-446 (Naprakesz can be false for the 23 floating pins), R-447..R-451 (one per remaining slice, with a rank and an owner), R-452 (no gate enforces catalog_since - the runner fetches at --depth 1), and R-453 (the vaulted dashboard password is stale on BOTH demo boxes, which is what stopped the badge render from being validated live). Live evidence for slices 1 and 2 in documentation/tests/. The record is PROVEN LIVE through the boot reconciler on demo-hp - one entry per compose service, digests matching ground truth read independently. The badge RENDER is not, and the five attempts are listed rather than summarised. |
||
|
|
56c7e373a3 |
SPIKE: what an app update actually does, and which other paths do it too
gates / gates (push) Successful in 18s
THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has already moved, and the Restart button does it — 18.3s with a network pull when the target image is absent, 0.5s when present, against a negative control that did not even recreate the container. The boot reconciler does the same thing unattended when an app fails to come back (bootrecon.go:269 -> StartStack). AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades nothing. Docker restores the old containers and the reconciler logs 'no boot-orphaned apps (nothing to start)'. AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run, the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported'. 'Rollback' is the wrong word for this arc and is struck. Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator confirmation. Peti's box was never contacted. No production code written in any repo: felhom-controller is at 960d29b0612c before and after, tree clean, and build/vet/test are green — run at the end to prove exactly that. Register: R-438/439/440 updated with live evidence; R-441..R-445 opened (restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB left behind; update reports success over a broken app; no fleet fstrim; hub telemetry outlives the app). Capability map gains three measured rows. The mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR, not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision. Five claims in the brief are named as wrong, including two of my own method. |
||
|
|
db38f4c800 |
hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta
gates / gates (push) Successful in 17s
R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.
was: "...still hold the older copy, so this is recoverable file-by-file; it is NOT
confirmed data loss. Check whether a deletion ran on the box before restoring."
now: "...still hold the older copy. The route back out of them is not yet established,
so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
before restoring anything, and check whether a deletion ran on the box."
It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.
Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.
R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.
THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.
Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.
R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.
Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
|
||
|
|
17d92e71a1 |
REPORT + CONTEXT + STATUS: the record corrected, and the thing worth building shipped
gates / gates (push) Successful in 18s
Opens with Part 1's answer because everything reads differently after it: a customer's own account can reach the snapshot DOOR and is REFUSED writes to it, but sees the tree EMPTY. The write-refusal is the load-bearing sentence of the whole R-95 re-scope and it is now PROVEN rather than cited - the control write to the account home succeeded and was cleaned up, the write into /.zfs/snapshot returned `dest open ...: Failure`, and nothing was left behind. Identical on both boxes. storage-box-pool-1 IS u629488, so the emptiness is per-sub-account filtering rather than absence - which means recovery is an operator act in a browser today (R-432), and that decides whether R-95's remedy can ever be product-driven. STATUS carries two items for Viktor in plain words: read one snapshot name off the panel (two minutes, and it may make recovery product-reachable), and IGNORE the alarm mail he received today - the live firing was required to prove delivery and nothing was deleted. Five of my own mistakes are named, including the one that matters most: my first escalation-only test was HOLLOW and its red-proof PASSED. It re-swept the same report, so the baseline had already moved and the latch was never consulted. That is why red-proofs are run. |
||
|
|
0476a8d8e6 |
SPIKE R-95: the safety net cannot be seen from the box - and that RAISES the urgency
gates / gates (push) Successful in 17s
READ-ONLY STUDY. No code, no version, no image, no golden. No delete verb was issued against any live store. ep0, DooPlex and Peti's box were not touched at all. Q1 FIRST, AND IT DID NOT GO THE EXPECTED WAY. The brief supposed a working seven-day snapshot net might bound the worst case. Measured on BOTH boxes over their own SFTP credential, with a positive and a negative control on each: NO .snapshots is visible to either sub-account - not in the account home, not inside the repo - and the account is jailed at /. Either none exist or a sub-account cannot see them, and the second is not a reprieve: a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel. The register's claim rests on nothing that was checked. The row has NO R-number, so nothing can cite it; its "confirm tomorrow" was 2026-07-27, 36 days ago; and the DUE-CHECKS block built for exactly this (R-341) is EMPTY. R-95's word "ARMED" is withdrawn pending R-429. The confirming field is a Hetzner API field, so this spike STOPPED at the section 11-D fence and left it for Viktor - ten minutes in the panel, and it re-ranks everything. Q2: TEN verbs, not nine. `check` was missing from the brief's list; `dump` is not a verb (it is a progress phase constant) and was withdrawn. There are TWO `forget --prune` sites - offbox.go:1388 AND offbox.go:1759 - and disarming one without the other reproduces R-191 exactly. Q3 (documented, from our own API mirror): AccessSettings has five booleans and `readonly` is the only permission axis. No append-only. So the PBS shape does NOT transfer - PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing. Q5 (measured, faithful append-only model, both controls passed first): withdrawing delete does NOT wedge the store - restic treats a dead owner's lock as stale and proceeds. The constraint everyone feared is not the blocker. But `unlock --remove-all` printed "successfully removed locks" while the lock was still there, and resticStep's crash-lock self-heal is built on that call - R-430. Q6 (measured, with a control): restic 0.14.0 DOES speak rest:. Append-only is a rest-server flag, not a restic one. Q7 (measured): detection is nearly free. snapshot_count already reaches the hub and the hub APPENDS reports, so the history to compare against is already on disk. RECOMMENDATION: answer Q1 today (Viktor, ten minutes), then build detection, then move retention off the box. Defer the transport change until Q1 is answered. Register: OPEN 179 -> 181. Filed R-429, R-430; R-95 updated and kept OPEN. 07 row 10's status is deliberately UNCHANGED. |
||
|
|
1e6c387a0b |
R-404 CLOSED with the ruling; R-417 CLOSED by cause removal; R-418/419/420 filed
gates / gates (push) Successful in 17s
THE RULING WAS NEITHER OPTION AS FRAMED. Both offered answers - narrow the gate, or leave it and write waivers - argued about the gate, and the gate was never the problem. DIAGNOSIS, from live source: golden_currency_gate.py never looks at the push. It compares the controller's newest CHANGELOG heading against this repo's bake evidence and returns the same verdict whatever you are pushing - correct for a standing invariant, wrong as a push gate. And controller_gates.py had NO golden-currency entry at all. So the repo where a release happens never checked, and the repo that cannot create the debt was refused on every push. 18 of the last 24 pushes here touched no code - measured, and the new classifier agrees EXACTLY - most of them by construction, because the controller's code is in one repo and its register lives in this one. SIX of those 18 were bake records, so the push that PAYS the debt is itself documents-only: the gate was blocking its own cure. Not the waiver its docstring prescribes: that clause was written for a release nobody wants a golden for. R-417 was a release we DID want a golden for, on a night the runbook forbade baking. A waiver would have recorded a lie. RULING: block the push that can create the debt, notify the push that cannot. The gate's logic, exit codes and wording are BYTE-IDENTICAL. Only the consequence changed, for one gate, on one kind of push, with a loud ADVISORY block so nothing goes quiet. R-242's vouch half is amended in place to say it is UNTOUCHED and still open - a baked-but-unvouched golden still passes both the gate and the new notice. Do not read R-404's closure as closing it. FILED: R-418 - this runner's docstring listed ELEVEN gates while THIRTEEN were registered; one-register and closed-register ran undocumented since 2026-08-24. Enumeration fixed here, the correspondence is still unenforced. R-419 - observations_gate.py accepts an item whose body merely CONTAINS "NOT-A-FINDING", even in prose disclaiming it; found by accident when a planted test observation passed and my live validation proved nothing. R-420 - controller_gates.py could not express a non-blocking gate at all before today. Register: OPEN 171 -> 172, CLOSED 158 -> 160. |
||
|
|
63eff21a5c |
golden 0.232.0 baked, vouched, floor raised — and it carries TWO releases
gates / gates (push) Successful in 17s
GOLDEN_SHA256 5f8a53ed5b19a6cb2006298ce6239f6fca2b990cc3ef6eada89f602801ca91b8, 657 494 489 B. 0.231.0 was never baked, so the fleet went 0.230.0 -> 0.232.0. THE CHECK THE 0.230.0 BAKE SKIPPED, AND THIS ONE DID NOT: the bake script's fingerprint was compared ACROSS THE HOP - 7b0fb5cf...73b6a1 on DooPlex and inside the VM. The previous bake recorded only the DooPlex-side hash and said so; this one is a measurement. Three independent readers agreed before anything was vouched: the bake's own print, the round trip of the published bytes (HTTP 200, 657494489 B, same sha, hashed from what was downloaded), and the hub's Day-0 dropdown reading Gitea on a different code path. The delivered artifact names its own controller - ./etc/felhom-controller-image reads felhom-controller:0.232.0 - with 19382 entries under var/lib/felhom/docker/. Both pre-gates were shown able to see something before their zeroes were believed, and the manifest was RE-READ after vouching rather than trusted from the 303 flash. DELIVERY WAS ACTUALLY EXERCISED. Both boxes had been hand-deployed during validation, so the floor had nothing to move. Rather than report delivery untested, demo-felhom was rolled back to 0.231.0 and the chain run for real - it moved itself in ~20s: 10:47:16 controller-swap: image file written, restarting bootstrap target=...0.232.0 10:47:26 controller-swap: new controller healthy target=...0.232.0 And this is the first bake golden_currency_gate.py actually gates: it now reads the GOLDEN_SHA256 line out of the bake log rather than matching a directory name (R-410, shipped hours earlier the same day). All 13 felhom.eu gates are green, golden-currency included, for the first time since v0.230.0 was released. |
||
|
|
f8f9ffdf2b |
STATUS: R-414 is the morning's first item - the new nightly check cannot run on demo-felhom
gates / gates (push) Failing after 17s
|
||
|
|
7ee25925f9 |
R-87 CLOSED: live evidence, capability row, architecture verdict, registers
gates / gates (push) Failing after 17s
Controller v0.231.0 + hub v0.110.0, both deployed and verified on demo-hp. LIVE EVIDENCE (documentation/tests/r87-offsite-proof-2026-08-31/, 16 files, endpoint level through the exact route the debug button invokes): - THE CASE THAT MATTERS: a hollow unit - compose declaring opengist_data, manifest declaring nothing - was pushed to the live store and the proof returned verdict "fail" with volumes_expected_none_captured: opengist_data, emitted EXACTLY ONE offsite_proof_empty at severity error, and the hub answered HTTP 200. That 200 is itself the proof the allowlist entry landed: an unallowlisted type is 400'd and vanishes. - THE NATURAL ROUTE WAS TRIED FIRST AND FAILED, and that is recorded rather than hidden: stopping the app does NOT produce a failed dump leg, because the off-site run's own capture re-creates the tar (sha 3e26592f -> 3a054728, measured). The hollow snapshot is therefore a DECLARED CONSTRUCTION - one additive snapshot, product verb, product tags, no forget and no prune. State restored: the product's own run made a healthy snapshot the newest again and the proof then passed opengist. - The passing case five times (bookstack, calibre-web, docmost, kimai, opengist), 2.2-4.0s each, matching the spike's measured band. - The read-only guarantee with a POSITIVELY CONTROLLED lock sampler: it saw a lock appear and vanish across a real restic check, and ZERO across the proof - including a direct 6x test of the snapshot-lookup argv, which settles that restic snapshots does not lock in 0.14.0 either. - Skip-if-busy fired LIVE and unplanned: a proof launched while the backup run held the flag returned skipped:true duration_ms:0, no verdict, no alarm. - The customer's own verification copies were untouched throughout, which is the safety property the separate proof root exists for. ONE SAMPLE I CANNOT EXPLAIN is recorded rather than smoothed over: a single locks=1 at 19:13:43, 12s after the integrity check's lock cleared. Two independent tests exclude the proof; I did not establish what it was. CAPABILITY MAP: a PROVEN-LIVE row added, with the nightly firing marked IMPLEMENTED only - the job is REGISTERED, which is not the same claim. 07 section 8 MATRIX ROW 4 WAS NOT MOVED, deliberately, and section 10.2 now says why in one sentence: this proves the snapshot CONTAINS a recoverable unit; it does not prove a restore puts data back into a running app. Without that sentence the new green tick reads as covering the drill. REGISTER: R-87 CLOSED and compressed into CLOSED-ITEMS.md. OPEN 172 -> 171, CLOSED 151 -> 152. No new rows minted. R-408 and R-409 stay open and are referenced by this work. golden-currency is RED and it is a DECLARED, EXPECTED debt: v0.231.0 is released and the newest golden carries 0.230.0. The fleet is on 0.230.0; demo-felhom does not have this job. A golden carrying 0.231.0 is OWED and it is Viktor's call (R-242). This push uses --no-verify for that reason - bypass #8. |
||
|
|
2263245cf2 |
golden 0.230.0 baked, vouched, floor raised - demo-felhom moved itself off the R-403 build (R-410 filed)
gates / gates (push) Successful in 17s
GOLDEN_SHA256 9287f7cef5f13166276e8406005e3f28004004510c5184f1c1c7377f7aafad2e, 657 873 700 B. Evidence documentation/tests/golden-0.230.0-2026-08-31/. WHY IT WAS OWED: the newest golden was 0.229.0, which IS the build R-403 says deletes a good copy. Every fresh install and the whole fleet floor still carried it. golden_currency_gate.py had been red across |
||
|
|
130f7a6eba |
R-87 SPIKE: measured, do not build it as written (R-407..R-409 filed)
gates / gates (push) Failing after 17s
Spike. NO production code. No version bump, no build, no deploy, no golden.
felhom-controller and felhom-agent were READ ONLY. The fleet stays on v0.230.0.
Q1 restic is 0.14.0 (go1.19.8, bookworm 12.15) - the four source comments asserting
it are CONFIRMED, not corrected.
Q2 --verify DOES exist and is NOT a content check. Red-proof: one byte changed in a
restored 160 MB tar with size and mtime preserved passed clean, rc=0. Verify took
131 ms on a 213 MB / 7-file tree, which cannot be hashing. A size or mtime mismatch
causes a silent re-download, not a failure. Controls: --target 1 hit, four post-0.14
flags and a nonsense string 0 hits each. Neither --verify nor --no-lock appears
anywhere in the controller source.
Q3 no reference for "correct" exists. restic ls --json carries no content hash in
0.14.0, and the unit manifest hashes 4918 B of a 213231242 B unit - 0.0023 percent,
the config files and not the dumps or the tars. R-409.
Q4 it is CHEAP. All 8 apps / 774378123 B logical restored back to back in 25 s, against
40257 ms for the weekly 100 percent check beside it. Individual restores 2253-3978 ms
regardless of size: cost is per-snapshot round-trip plus ~1 s per 200 MB. Peak scratch
is the app's full logical size. The 1.1 MB restic cache is index only and hides nothing
(--no-cache 5423 ms vs cached 3198 ms, trees byte-identical).
Q5 skip-if-busy stays right at 25 s against a 2m52s nightly backup. But
RestoreOffboxScratch takes NO acquireRunning, while offbox_integrity.go:28 asserts
every off-site operation does. R-408.
Q6 observed with a positively-controlled lock sampler: restic restore takes NO lock;
restic check DOES (locks 0 -> 1 for nine samples -> 0 across the check, zero across two
restores). The product writes anyway - unlockStale runs `restic unlock`, a delete verb,
before every restore (offbox_restore.go:289). The task's lead was right in direction and
wrong in mechanism. R-95's constraint IS satisfiable: --no-lock plus skipping unlockStale
writes nothing, and both mechanisms exist unused. Neither was fixed - the task forbids it.
offbox_integrity.go:255's "It NEVER writes to the repository" is R-407.
Q7 THE DECIDING ONE: of R-353/354/356/358/403 an unattended scratch-restore would have
caught ONE (R-356). The value is elsewhere, and the weekly check structurally cannot
reach it: `check` proves the stored bytes are the stored bytes, never that we stored the
RIGHT thing. A hollow unit backs up, checks at 100 percent and restores cleanly and
recovers nothing - R-403, measured in bytes on 31 August.
RECOMMENDATION: option C, the NARROW test - one app a night, restored to scratch, checked
against its own manifest.json through the existing unitCarriesData, scratch deleted, the
SNAPSHOT recorded as the proof. Options A (do not build) and B (scheduled attended drill)
considered explicitly; B is weakest because it is what already happens. R-87 should be
RE-SCOPED, not built as written, and that is Viktor's call - the row stays open carrying
the verdict and STATUS.md item 4 asks it in plain words.
Also corrected in 07-backup-architecture.md: matrix rows 4 and 10 both said "the depth
that ships ON does not re-read pack contents (R-399)". R-399 CLOSED in v0.228.0 and the
depth is 100 percent. Two stale cells, fixed, and the spike verdict added beside them.
Row 4's verdict is UNCHANGED by the spike and now says so.
Teardown: all three layers, none of them "nothing was created" - 6 files on the PVE host,
9 in the guest, 5 plus 2 run-flags in the container, all removed and verified empty. The
four scratch directories this session's restores created were removed; three that
pre-date the session were left alone. Two state changes recorded rather than hidden: the
control integrity run recorded its verdict (depth structure -> 100%, due-ness +7 days),
and four restores appear in the controller log. Nothing was written to the off-site
repository by hand.
Evidence: documentation/audits/evidence-spike-restic-restore-2026-08-31/ - 31 files,
every one pulled off the box BEFORE teardown (R-320).
golden-currency is RED at this commit and was already red at
|
||
|
|
dddcc808be |
R-403 CLOSED (controller v0.230.0), R-404 filed as a decision for Viktor
gates / gates (push) Failing after 17s
07-backup-architecture gains section 8.2, placed beside row 5 on purpose: the derived-copy rebuild rule is UNCHANGED and section 8.2 names the single exception, so a future reader who finds RunTier2 skipping a leg does not fix it back. It carries the measurement (120 082 104 B -> 7 036 B on the shipped v0.229.0), the four-case table, why hollow is a manifest question and not a size question, why the data legs are deliberately not guarded, and why the capture job is not guarded either. 00-capability-map: the Tier-2 row's status does NOT move, stated explicitly rather than left ambiguous. R-403 removes a way the route could be DESTROYED between uses; it does not change what the route can be relied on for. Register: R-403 CLOSED and compressed into CLOSED-ITEMS (594 -> 593 open lines). R-404 FILED as a DECISION and deliberately not acted on - should a documents-only push be subject to the golden-currency gate, now that it has been correctly bypassed six times? Both sides stated, plus what happens if Viktor does nothing. The gate was NOT changed. R-242: seventh conviction, and the FIRST where the day-0 ground does not apply - R-403 is a defect in the nightly Tier-2 copy, which a day-0 box starts running on its first night. This push uses git push --no-verify, declared here and in felhom-controller/REPORT.md. A golden carrying 0.230.0 is owed and is more urgent than the previous six. STATUS: the R-403 item moves out of 'Broken' into what works, in plain words; the delivery item now says a golden is owed and that the fleet carries the defect; R-404 goes into the decide section with its do-nothing outcome. Drill evidence: nine phase logs, including the two things that went wrong (a repair whose rsync was not installed in the guest and silently did nothing, and a session that expired mid-run so a POST did nothing). |
||
|
|
83ff9e8e38 |
golden 0.229.0 baked, vouched, floor raised — R-242's sixth debt PAID the same day
gates / gates (push) Successful in 16s
GOLDEN_SHA256 39aa886df77b21757aef3b298a389343dc0df5134bb0f14e8f92a451d7bdae87, 656 864 331 B.
The evidence is the ROUND TRIP, not the build log: the published bytes were downloaded back and
match the bake on both size and sha, and ./etc/felhom-controller-image read OUT of the downloaded
archive says felhom-controller:0.229.0 - the delivered artifact naming the controller it will start.
A THIRD independent reader agreed before anything was vouched: the hub's own Day-0 dropdown read the
same sha straight from Gitea, a different code path.
Both pre-gates were proven able to see something before their negative results were believed - the
404 pre-gate against a 200 from 0.228.0, and the token-leak grep against a seeded throwaway copy.
Acceptance markers counted on the COMMITTED log: 1/1/1/1 present, 0/0/0 absent.
The vouch is a three-field change, checked rather than assumed: MinAgent 0.129.0 read from the
golden's controller CHANGELOG header, agent_version 0.130.0 >= min_agent 0.129.0 (not the R-216
shape), agent_sha256 and wrapper_sha256 carried through explicitly because the handler clears a
field it is not sent. Verified by re-reading the manifest, never by trusting the flash. The R-120
gate PASSED rather than being bypassed - fleet newest 0.229.0, golden 0.229.0.
The floor is proven ACTING, not merely set: demo-felhom self-updated 0.228.0 -> 0.229.0 and logged
settle-gate GO at/above floor 0.229.0. Nobody deployed to that box. Both demo machines now carry the
Tier-2 unit restore.
golden_currency_gate.py went red -> green; the --no-verify bypass declared on
|
||
|
|
c2de785bf2 |
R-102 + R-103 CLOSED (controller v0.229.0) — architecture, register, STATUS, drill evidence
gates / gates (push) Failing after 17s
07-backup-architecture: 6.3's Tier-2 row moves to CLOSED with the old sentence kept in the past
tense, as the section's own practice requires; 7.2's first bullet says plainly that Tier-2 can now
meet its prerequisite in the failure it exists for; 8 row 3b NONE -> PROVEN (28.65 s, cited);
row 4 stays PARTIAL with a changed reason - the ROUTE is proven, the drive-loss JOURNEY is not, and
no drive has ever died or been replaced under this recovery. 8.1's blanks updated per row.
6.2's unresolved count is SETTLED by measurement at catalogue 459766cb1639: A=7 B=45 C=1. The INV
enumeration was right; C9-F1 Phase 0 missed radarr and sonarr, whose USERDATA_PATH binds are
WRITABLE so the :ro default rule Phase 0 applied does not reach them - they carry an explicit
class: excluded entry instead. Class C is bentopdf. No catalogue file was changed.
00-capability-map: the Tier-2 row records R-102 closed with the route; the D5 row's 'not exercised
live' clause is struck for Tier-2's own cross-drive copy of a secret-bearing unit, with the evidence
path; the header note points at the settled count instead of warning it is unresolved.
Register: R-102, R-103 and their C9-F4 / C9-F1b aliases closed and compressed into CLOSED-ITEMS
(596 -> 593 lines, each naming git show
|
||
|
|
1623a4d5b5 |
golden 0.228.0 — baked, published, round-trip verified, vouched, floor raised
gates / gates (push) Successful in 19s
Bake: build-golden.sh v3.0.0 in the drill VM, reverted to virgin and cold-booted.
GOLDEN_SHA256 76a3a98b9e7cc23bf8ae51b38a6272f576df285cb34cd22235ac3f06a31e53ec,
658 079 744 B. All four acceptance markers counted 1; excluding/FATAL/mp1 counted 0.
The 404 pre-gate was proven to work before its 404 was believed — the target URL
404'd while the existing 0.227.1 package 200'd on the same command.
The evidence is the ROUND TRIP: the downloaded bytes match the bake's size and
sha, and ./etc/felhom-controller-image read OUT of the downloaded archive says
gitea.dooplex.hu/admin/felhom-controller:0.228.0.
Vouch: three fields together — golden_version 0.228.0, agent_version 0.130.0,
min_agent 0.129.0 (read from the controller CHANGELOG header, not assumed);
wrapper_sha256 carried through explicitly. agent >= min_agent, so not the R-216
shape. Verified by RE-READING the manifest, never the flash. R-120 gate passed.
Floor raised 0.227.1 -> 0.228.0 (impact preview {"below":3,"valid":true}).
The floor is ACTING: demo-felhom self-updated 0.227.1 -> 0.228.0 in under a
minute and re-registered offsite-integrity by itself. Both demo boxes now
re-read their whole off-site store on the weekly check.
Token hygiene: file->file scp, runner script inside the VM, unit properties
grepped 0. The leak grep on the committed log was proven with a planted copy
(1) before its 0 was believed. Teardown: guest 9100 purged, secrets shredded
after the log was copied out, VM off, disk reverted to virgin.
golden_currency_gate.py red -> green. All 12 felhom.eu gates OK.
|
||
|
|
77a5a1154b |
docs: controller v0.228.0 — R-399/R-400 closed, R-401/R-402 filed
gates / gates (push) Failing after 19s
STATUS.md: header said 2026-08-23 over a 2026-08-30 body, and two "Waiting on you" items were both numbered 4 — both fixed. R-399 leaves that section (decided and shipped); the depth change is stated in plain words and the remaining items each say what happens if Viktor does nothing. 00-capability-map.md: the off-site verification row now carries its DEPTH, and its live citation is the 2026-08-31 run at 100%. The weekly firing at the new depth stays IMPLEMENTED, not PROVEN-LIVE. 07-backup-architecture.md §10.2: R-399 recorded closed, with the one sentence that stops it being turned back down — the structure check PASSED a size-preserving pack corruption. R-87 untouched and still OPEN. Register: R-399 and R-400 compressed into CLOSED-ITEMS.md with their reasoning kept and 300d7e8 named as the commit holding the originals. R-401 filed with a TRIGGER (the slow-check WARN firing) rather than a date. R-402 filed: the integrity verdict and its depth are on the wire and no hub surface reads either. OPEN 166 -> 165, CLOSED 148 -> 150. wire_contract_gate.py: offsite.last_integrity_depth allowlisted WITH ITS REASON beside its sibling last_integrity_ok, both to be deleted together when a hub surface is built (R-402). |
||
|
|
db0812b6f2 |
Golden 0.227.1 baked, vouched, floor raised — and the floor delivered the new job by itself
gates / gates (push) Successful in 16s
Second full delivery of the day. golden_currency_gate.py went red -> green on the same command, so the --no-verify bypass declared on the previous push is now historical rather than standing. GOLDEN_VERSION 0.227.1 GOLDEN_SHA256 66754491dc9bd0130ef8ded9562f63c53a5ffdcfd91baa551141e55fa083ea32 size 657 403 203 B baked gitea.dooplex.hu/admin/felhom-controller:0.227.1 MinAgent 0.129.0 (read from the controller CHANGELOG header, not assumed) THE EVIDENCE IS THE ROUND TRIP. The published bytes were downloaded back -- size and sha256 identical to what the bake reported -- and ./etc/felhom-controller- image was read OUT of the downloaded archive: felhom-controller:0.227.1. That is the delivered artifact naming the controller it will start, from the bytes a customer's box would actually fetch. Markers counted: docker OK (overlay2 = 1, mount point rootfs = 1, mp0 = 1, upload OK (HTTP 201) = 1; excluding = 0, FATAL = 0, mp1 = 0. 404 pre-gate passed before the run and the script's own pre-delete agreed, so nothing was overwritten. Three-field vouch, all three checked: agent_version 0.130.0 >= min_agent 0.129.0 (NOT the R-216 shape), wrapper_sha256 carried through explicitly because the handler clears it when omitted. Verified by RE-READING the manifest rather than trusting the flash. The R-120 gate on that POST passed on its own terms rather than being worked around. AND THE LINE WORTH KEEPING. demo-felhom self-updated 0.226.1 -> 0.227.1 in under 30 seconds and then logged: [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST A box nobody deployed to now runs today's off-site integrity check on its own schedule. That is a floor DELIVERING rather than merely recording, observed instead of assumed -- and it is the strongest evidence R-242 has carried. Token hygiene: file->file, read inside the VM by a runner script, never on a command line (systemctl show ... grep -c -F token = 0). The leak grep on the committed log was PROVEN TO WORK before its 0 was believed. Teardown: guest 9100 destroyed --purge, secrets shredded AFTER the log was copied out, VM powered off, disk reverted to virgin. R-242 now records the cadence as MEASURED: five convictions and two full bakes in one day. Every bypass declared, every debt paid -- and the pattern the row exists to name is exactly that a release and its delivery are separate acts. Its other half stays open: nothing gates the VOUCH itself. |
||
|
|
99af997ab9 |
R-359/R-397 closed, R-398 corrected, R-399/R-400 filed with measured numbers
gates / gates (push) Failing after 18s
THE MEASUREMENT IS THE STORY, and it re-frames the row it was filed under. A pack was corrupted WITHOUT changing its size; plain `restic check` -- the depth that ships ON -- returned `no errors were found`, exit 0. Only --read-data caught it. So the check that shipped verifies the index, the pack inventory and the snapshot graph, and does NOT re-hash pack contents. R-399 was filed as a bandwidth-and-cadence question; it is more than that, and its row now says so. R-399 gets three MEASURED numbers instead of estimates: store 140 829 678 B / 2651 blobs / 67 snapshots; structure check 35.0 s; curve 10% 35.9 s, 50% 37.3 s, 100% 39.2 s. At this size re-reading everything costs four seconds more than reading none, because the wall clock is SFTP round-trips not transfer. The row states the limit too: these do NOT extrapolate. R-400: the sweep the task asked for found EIGHT dead debug buttons, not one. 24 endpoints referenced in debug.html, 17 dispatched. Single dispatcher, exact match, default NotFound -- so they 404. A third of a debug page does nothing, on the surface an operator reaches for when something is already wrong. R-398 is CORRECTED AND LEFT OPEN, not closed. I filed it yesterday saying resticStep is not a seam so no test can drive a restic path. The layer below it has been injectable since the off-site tier shipped. The row survives as the record that the seam EXISTS so nobody re-files it. 07 gap register: R-359 and R-397 closed; R-87 restated IN PLACE as "AND IT IS NOT R-359" because the two rows are adjacent and a check is not a restore-test. 08 alarm ladder: both event types recorded, including that `ok` is `info` and therefore mails nobody BY DESIGN, and that all three registers were checked and deliberately left alone. 00 capability map: PROVEN-LIVE for the check, the notifier and the hazard control; the scheduled firing is IMPLEMENTED only, because a week has not passed. wire_contract_gate: `offsite.last_integrity_ok` allowlisted WITH A REASON. The gate was right -- the controller emits a field no hub struct can decode. Building the display is a hub change and R-331 ruled that class the operator's decision; the entry says to delete it when a surface exists. This push used `git push --no-verify`. golden-currency is CONVICTED and right: 0.227.1 is released and the golden carries 0.226.1. A BYPASS, not a waiver, and the task spec directs it -- golden and fleet delivery are Viktor's (R-242). It is item 3 under "Waiting on you". Register 163 -> 165 -> 163. |
||
|
|
4f875174fe |
Golden 0.226.1 baked, vouched, and the fleet floor raised — the debt is paid
gates / gates (push) Successful in 17s
golden_currency_gate.py had been CONVICTED four times today across three controller releases. One bake covers all three, and the gate went red -> green on the same command, which is its proof that it measures something real. THE THREE DECLARED BYPASSES ARE NOW HISTORICAL RATHER THAN STANDING. GOLDEN_VERSION 0.226.1 GOLDEN_SHA256 70ed8e9377dec22a9b493e55f222b0e25a49d7f3caec8c506e0412fd6baefe69 size 657 197 592 B baked gitea.dooplex.hu/admin/felhom-controller:0.226.1 MinAgent 0.129.0 (read from the controller CHANGELOG header, not assumed) THE EVIDENCE IS THE ROUND TRIP, NOT THE BUILD LOG. The published bytes were downloaded back -- size and sha256 both identical to what the bake reported -- and ./etc/felhom-controller-image was read OUT of the downloaded archive: `felhom-controller:0.226.1`. That is the delivered artifact naming the controller it will start, from the bytes a customer's box would actually fetch. Acceptance markers counted, not eyeballed, each string captured from this run's own log rather than paraphrased from the runbook (two of the three the runbook named until R-233 could not match anything the script prints): docker OK (overlay2 = 1, including mount point rootfs = 1, mp0 = 1, upload OK (HTTP 201) = 1; excluding = 0, FATAL = 0, mp1 = 0. The 404 pre-gate passed before the run, so nothing was overwritten. THE VOUCH IS A THREE-FIELD CHANGE AND ALL THREE WERE CHECKED: agent_version 0.130.0 >= min_agent 0.129.0, so NOT the R-216 shape; wrapper_sha256 carried through explicitly because the handler clears it when omitted. Verified by RE-READING the manifest rather than trusting the flash -- golden option 0.226.1 SELECTED, all four shas matching. THE FLOOR IS PROVEN ACTING, NOT MERELY SET. demo-felhom self-updated within 30 seconds: "[selfupdate] Post-update startup: update successful (0.225.0 -> 0.226.1)". Both demo machines now run 0.226.1 and only one of them was deployed to by hand. Token hygiene: copied file->file, read inside the VM by a runner script, never on a command line (systemctl show ... | grep -c -F token = 0). THE LEAK GREP ON THE COMMITTED LOG WAS PROVEN TO WORK BEFORE ITS 0 WAS BELIEVED -- a throwaway copy with the token appended grepped 1, was shredded, and only then was the real log's 0 taken as evidence. Teardown: build guest 9100 destroyed --purge, secrets shredded AFTER the log was copied out (standing rule 5), VM powered off, disk reverted to virgin. R-242's OTHER half is untouched and still open: nothing gates the VOUCH itself. |
||
|
|
e027b5d999 |
Register + architecture for controller v0.226.0 (R-353/357/358/360/396), and R-395 fixed
gates / gates (push) Failing after 17s
Closes R-353, R-357, R-358 and R-360 with their shipping version and evidence path, and files two new rows. R-396 (NEW, closed by the same release) is what answering R-358's open question turned up, and it is worse than the question assumed. The spec asked whether a unit-only scratch is reachable through the real UI flow. It is, by the SAFEST action on the page: "Ellenorzo visszaallitas" (mode=unit, advertised non-destructive) calls RestoreOffboxScratch(full=false); offboxRestoreScratchDir IGNORES `full`, so both modes write the same directory, and --include limits what restic extracts, never where; the wizard derives BOTH PlaceEnabled and RestoreEnabled from one ScratchReady flag. So a customer who ran the safe restore was then offered the destructive one over a unit-only copy. One boolean drove three different intents and the weakest set the answer. R-395 (filed by the spec) is fixed in this commit, not just recorded. STATUS.md said golden 0.223.0 / floor 0.222.0 in one block and demo-hp 0.219.0 / floor 0.218.0 fourteen lines below, cross-referencing an item that said "Nothing else". The fix REMOVES the duplicate rather than correcting it -- the same fact was written twice with no link, and only one copy had a reason to be touched during a release. "What works" now points at the item above instead of restating a version. 07-backup-architecture: four rows added to the 10.2 gap register plus R-396. Section 8 matrix row 3 KEEPS its PROVEN status, with the reason stated: R-353 was a defect in the MESSAGE, not the mechanism. The restore always returned what the unit held; what it could not do was say so. A status that measures whether data comes back must not move because a status line was wrong. 00-capability-map: one new row, and it splits what is claimed. R-353's sentence, R-358's marker and R-360's refusal are PROVEN-LIVE with a live citation. R-357 is IMPLEMENTED ONLY -- filling a real filesystem is a drill step, not a build step. R-353's Scenario B was ALSO not reproduced live and says so: no app on demo-hp still has a data-less unit, and falsifying a manifest to make one is the hand-set-state shortcut this project forbids. This push used `git push --no-verify`. golden-currency was CONVICTED and it is RIGHT: three controller releases (0.224.0, 0.225.0, 0.226.0) and the golden still carries 0.223.0. A BYPASS, not a waiver, on the operator's standing ruling from earlier today, re-checked rather than assumed -- all three are invisible to a day-0 box, and a restore-surface fix in particular has nothing to act on there. The ground expires the moment a release changes first-boot behaviour. Tracked on R-242; ONE bake carrying 0.226.0 covers all three. |
||
|
|
ebdc04601d |
docs(hub v0.108.0): the delivery grain, the cooldown ruling, and gate 11's first subject
gates / gates (push) Successful in 15s
The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or coarse, and why the default is coarse. CONTEXT records two rulings: the grain is allow-listed rather than inferred from the payload, with crossdrive_failed as the proof that a payload rule would have been wrong; and a finding recorded only in REPORT.md has a lifetime of one session. R-389 closed and compressed, keeping its rules and naming the commit whose git show returns the full text. R-390 and R-391 left open. REPORT.md is gate 11's first real subject and passes: six observations, two FILED, four NOT-A-FINDING with their reasons. Three of those declarations are things a tidier report would have omitted - the gate's own spec would have passed the item it was built to catch, the burst has no ceiling, and ArgoCD said "successfully rolled out" while still running the old image. STATUS carries forward the one thing outstanding: the controller floor still reads 0.222.0 while the golden reads 0.223.0. |
||
|
|
2f7c9a6ce5 |
docs(R-329/R-386/R-387): the severity contract, the intent ruling, and Part 5 recorded
gates / gates (push) Successful in 17s
The alarm ladder gains the severity contract (the hub's vocabulary is exact, it coerces silently, and three things now hold it) and the intent test with its three-way ruling on unknown. Both marked [DESIGN] with the live measurements. Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy, verbatim, marked plainly as direction rather than current behaviour, with the 12 -> 15 toggle growth as the argument. Filed as R-388, a product decision. R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the full-text commit named. R-387 filed closed - including WHY the dispatcher branch was kept rather than deleted, which is evidence (three monitor checkers call ProcessEvent directly) and not caution. The drill record names three things that had to be re-run: an inert red-proof mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A NOT proving the customer gate because demo-hp has no prefs row at all. Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B. |
||
|
|
55274d5ef3 |
R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s
The gate failed only on `released > baked`, so it could catch a forgotten bake and nothing else. A golden AHEAD of the record passed silently - and that is how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG heading still read v0.221.0, with every gate green. Reproduced on the real history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0. The gate now asks whether the version being shipped is WRITTEN DOWN: the baked version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG. Membership rather than `baked > released` deliberately - a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2) preserved; every refusal names a reason and a route. Red-proofed both directions: old gate/old record exit 0, new gate/old record exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0. 08-alarm-ladder.md is new, and its absence was itself the finding: no document owned "when does a broken app raise an alarm?". The rules lived as comments in four packages, each locally correct, with the ordering between them legible only by reading one function top to bottom - which is how R-384 survived review. R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed closed. R-386 filed OPEN: a single-container app stopped out of band raises no alarm, and a comment claims the opposite - measured live, 9 scans, 0 events, against a positive control from the same box 17 minutes earlier. Not fixed here. Golden 0.222.0 baked and published; vouching is the operator's act. |
||
|
|
1eb64bec51 |
R-361 docs: the [FACT], the negative that cancelled Part 2, R-383/R-384, golden 0.221.1
gates / gates (push) Successful in 17s
07-backup-architecture.md gains a dated [FACT] on R-361 - a comment asserting an invariant the code did not have, for four months - and a [DESIGN] on the db_dumps decision INCLUDING the trap it created: a stable list lets the already-current early return fire, so per-capture housekeeping must sit above it. 00-capability-map.md records the NEGATIVE from Part 3 so it is not re-derived: a held app does NOT raise the dead-app alarm. It aggregates to unhealthy, which IsDownState excludes. Measured on the shipped build with the scans demonstrably running over it. No suppression was built and no row opened. R-383: the double-failure message names an undo copy that is not there - R-361's own class, one surface over, observed on both 0.220.2 and 0.221.1. R-384: an app whose database has died reads unhealthy and raises no alarm. R-361 closed and compressed. OPEN-ITEMS 325236 -> 327266 bytes. Golden 0.221.1 baked, published and round-trip verified. The golden-currency gate blocked this push and that block is not circular, so it was satisfied rather than bypassed - no --no-verify anywhere in this session. |
||
|
|
a8caa0fdde |
R-379/R-380 docs: the failure ladder, the drill record, register housekeeping
gates / gates (push) Successful in 17s
07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay -> rollback -> hold, including why no engine flag closes it: --single-transaction makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the fix and the flag is a belt. Drill record for the live walk, including the TWO defects the walk found in the fix itself (a rollback into a re-created container; an operator route that cleared the file while the running controller kept refusing) and the ONE red-proof that PASSED, which is reported rather than omitted. R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes. STATUS.md restates the outcome and names the next operator step. |
||
|
|
8c9f1b798b |
golden 0.219.0 baked, published and round-trip verified (NOT vouched)
gates / gates (push) Successful in 18s
Baked in the drill VM per RUNBOOK-manual-build.md 4.0/4.1, carrying controller v0.219.0 (R-356). GOLDEN_VERSION 0.219.0 GOLDEN_SHA256 67b46f78f8ed9c7b1876265ab1bde9ec6798897898b1836acece9f3864a2aeb6 656832571 bytes All five pass markers matched, both negative controls at 0. Verified by ROUND TRIP - the published object downloaded again and its sha recomputed - not by the number the script printed. Both token-leak greps were proved able to convict before their zeros were believed: planted copy grepped 1, shredded, then the 0 accepted. Teardown complete: guest 9100 purged, four secret/script files shredded after the log was copied out, qemu exited, disk reverted to virgin. The revert first refused while qemu held the image, which is the runbook's own no-holder proof. NOT vouched - that is a three-field operator save (golden_version 0.219.0, agent_version 0.130.0, min_agent 0.129.0). |
||
|
|
c297b9f85e |
R-356 docs: correct R-107 in the architecture, record the design, refresh STATUS, compress the register
gates / gates (push) Failing after 17s
07-backup-architecture.md: three places said no offsite action unpacks the named-volume tars. R-107 closed in controller v0.218.0; all three corrected with a dated [FACT], the old sentence kept in the past tense. R-102 is NOT closed and the correction says so explicitly. New [DESIGN] paragraph in 6.3: the restore destination is resolved by the same rule as the capture destination, and the wrong-disk refusal applies to apps that have a drive to get wrong. Carries the 13/40 measurement. STATUS.md was internally contradictory - nothing waiting, and one decision waiting, for something the same page recorded as shipped. 218 -> 102 lines; the deciding section now says what happens if nothing is done. R-356 compressed into CLOSED-ITEMS.md; OPEN-ITEMS 327109 -> 325236 bytes. Drill record and 16 evidence files for the live walk on demo-hp. |
||
|
|
877fcd2a38 |
R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed
gates / gates (push) Successful in 16s
Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with a negative control first — the same planted, hash-recorded fixture run through the same steps on both builds. R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the off-site copy or the restore; and because the same wrong name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label. Sweep proven able to convict before its count was trusted: one affected app of 53. R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit, before the database and inside the stopped window, and VolumesReplayed reaches the sentence. The half-false comment beside the skip is corrected and the half that still holds is named. Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b, verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are the operator's decision, and raising the floor is what puts this on demo-felhom, which is still on 0.217.0 and still has both defects. R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them (an existing guard), they are adoptable by hand, and doing it automatically would be a migration. Ceiling R-366 -> R-367. |
||
|
|
d895d9f7dd |
STATUS + capability map: narrow the end-to-end off-site claim to the leg it was proven on
gates / gates (push) Successful in 16s
The 2026-08-04 row claimed "a customer's file survives a machine rebuild and comes back — the whole off-site story, end to end". Tonight's drill shows that holds for the declared-userdata leg of a drive-declaring app and for nothing else: the off-site restore has no named-volume leg (R-354), and refuses outright for the 40 apps that declare no data drive (R-356). Since that class keeps ALL its data in named volumes, the end-to-end story is unproven there and disproven for the volume leg generally. The escrow/key half of the row is untouched and still stands. STATUS.md also corrects the fleet pair it still named (0.214.0/0.129.0 -> 0.217.0/0.130.0) and records that demo-hp's off-site had been silent since 9 August. |
||
|
|
910fd91124 |
agent 0.130.0 published and vouched; R-347 closed, R-349 + R-350 filed
gates / gates (push) Successful in 14s
Released via scripts/release-agent.sh: tag v0.130.0 at 7569f34, sha256
a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3,
verified by independent download and reproducible byte for byte with
-trimpath -buildvcs=false.
Vouched agent 0.129.0 -> 0.130.0 in the Day-0 manifest. Only the agent
fields changed: min_agent stays 0.129.0 because it states what the GOLDEN
CONTROLLER requires, and raising it would have HELD the floor for every
box below 0.130.0. Global floor untouched at 0.216.0 -- and on hub
v0.106.0 it is a separate form with its own action, so publish-train
rule 2's hazard no longer exists in the shape its incident describes.
No --no-verify: the CHANGELOG heading was flipped only after the tag and
package existed, so release-complete passes on the real artifact.
R-349: the fleet was running a DIFFERENT binary under the same version
name -- the proof deploy was a hand build, the release is -trimpath.
Self-update could never have corrected it, because every version check
compares the string. Both boxes reinstalled from the downloaded package.
The proper fix exists in miniature as wrapper_sha256 and was never
extended to the agent's own binary.
R-350: I printed the hub password into the session transcript via
curl -w '%{redirect_url}' -- the hub answers 303 and curl re-attaches the
credential. Not in git, not in any committed file, not in the evidence
directory. Rotation is the operator's call.
ep0 closes at fd 17, ESTAB 0, CLOSE-WAIT 0 -- its t0 baseline -- and was
read-only for this entire arc.
|
||
|
|
57dd62b097 |
R-344 fixed and proven on both boxes: ep0 is back to fd 17 from 415
gates / gates (push) Successful in 14s
P1, outcome (i) in one second: replacing the agent on demo-hp released exactly its 199 established connections (ep0 fd 415 -> 216). CLOSE-WAIT stayed 0, so outcome (ii) does not exist and gets no row -- ep0 reaps on peer FIN correctly, and the 543 CLOSE-WAIT at the 08-18 wedge has another explanation. P2, 1.03 h (operator closed the >=4 h window early, so no daily rate is extrapolated): control +4, fixed +0, with each box making exactly 4 /snapshots and 4 /version calls. Same cadence, same work: 4 cycles -> 4 leaks vs 4 cycles -> 0. The fixed box's cycles are in ep0's log, so the zero is the fix and not a stopped agent. P3: the second box took ep0 from 220 to 17 fd in under two seconds. 17 is precisely the t0 baseline of 2026-08-18 09:51:22Z. Corrects a claim this session made earlier the same day: the accumulated descriptors did NOT need an ep0 proxy restart. They were held on both sides. ep0 was read-only throughout; its PID never changed. R-344 updated and left OPEN (unpublished is not delivered). R-336 re-scoped -- its old next-step would have fixed nothing while looking like a failed fix, and it is now a scaling row (~25 req/s at fifty customers). R-347 filed for the delivery gap (Viktor decides). R-348 filed: an agent restart blanks the reported backup list for ~18 h and the Store comment calls it unaffected -- blinds no alarm, checked not assumed. |
||
|
|
19672e685e |
SPIKE ep0 connections: the leak is felhom-agent's, not the poll rate
gates / gates (push) Successful in 14s
R-341's first dated check, taken at +46.2 h: fd 17 -> 405 over 166,251 s
= 201.6/day. Pre-registered range was 370-450; observed 388. UNCHANGED,
as predicted. CLOSE-WAIT is 0 -- absent entirely, not merely flat.
Q1: exactly two peers, 194 each, no third party.
Q2: outcome (a). ep0 388 = 194 + 194 on the boxes, twice, and the four
new sockets carry the same source ports on both sides. 0 closed in 31 min.
The finding: all 388 are held by felhom-agent. pvestatd and
proxmox-backup-client made 162,404 requests and leaked zero. Mechanism is
a per-cycle http.Transport with a zero-value IdleConnTimeout that nothing
ever closes (internal/pbs/client.go:56, main.go:1486). R-336's premise
does not survive this -- cutting the poll rate would have fixed nothing.
Q3 NOT measured: Phase C held at STOP 1, prediction pre-registered first.
Part 0 captures the due-checks gate's first conviction on a real overdue
date (rc=1, names R-341, sole failure among 10 gates). Row cleared at
Part 4, after the result was recorded in R-341, not to make a push work.
New: R-344 (the transport leak), R-345 (hub/Makefile pushes :latest),
R-346 (ActiveEnterTimestamp reads 5h56m early -- NRestarts is still 0).
|
||
|
|
ab2262c91c |
hub v0.106.0: report loss of visibility into the off-site stores (R-339)
gates / gates (push) Successful in 14s
THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.
REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).
THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.
Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
- the unreachable event REPEATS rather than escalating once. The band shape
would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
mail is missable. It leans on the dispatcher's 1 h operator cooldown to
become an hourly "still blind" heartbeat.
- ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
which means we reached it. Counting it would alert for days on a healthy
pre-update endpoint.
- born-blind is reported: the counter is not gated on having a snapshot, so
a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
than zero-valued -- a fabricated timestamp reads as "it was fine until
then".
Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).
Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.
Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
|
||
|
|
0a5e9b14dc |
due-checks gate (R-341), floor raise recorded (R-343), snapshot coverage (R-342)
gates / gates (push) Successful in 14s
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as prose in a register row; nothing read those dates and nothing would have objected when they passed. The dates now live in a DUE-CHECKS block INSIDE OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the pre-push hook and CI. exit 0 nothing due (prints pending count + nearest date; empty block too) exit 1 a row is due/overdue (due <= today, UTC -- due TODAY counts), or a row names an item with no R-row exit 2 block absent/duplicated/unparseable -- INCONCLUSIVE, never 0 It REFUSES rather than warns, and its docstring states the limitation: it is NOT a scheduler, it fires on the next push, not on the date. 37 tests. BOTH red-proofs run and reverted -- and the first one earned its keep by catching a hollow assertion of MINE rather than confirming the gate: flipping <= to < left a due-today row in neither bucket, min() raised on an empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the boundary was wrong. An exit code cannot tell a verdict from a crash. The test now asserts the conviction banner and the absence of a traceback, and the gate returns 2 rather than crashing if that partition breaks again. PART 3 — the floor raise, and the premise was WRONG. Read back from the store (not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero per-customer overrides, no "managed floor HELD" line. But read 5 shows the raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save, exactly the immediate action publish-train rule 2 documents. No error events followed; it restarted clean. R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five reads clean and no directive served. It went well, but a record calling it inert when it moved a customer box is what misleads the next reader. The row also states why the floor was behind -- rule 2 policy, not drift, earned by the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor (store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267, fails open at :78-82) rather than asserting them. Two boxes are below the floor and neither reports: drill-r50 (blocked, powered off) and peti-felhom (host row deleted). peti-felhom was NOT contacted -- its row records that a report from a deleted host 401s and is not persisted, so the raise cannot reach it. PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a separate Volume that snapshots exclude, so a rollback restores software state and NOT the datastore. Fine for that upgrade; the safeguard for any future procedure that could touch the datastore does not exist and is Viktor's call. Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling 200). Capability map deliberately unchanged; no row cites a floor or golden version. repo_gates.py fully green, 10/10. |
||
|
|
7d81681d6e |
golden 0.216.0: baked, published, vouched — gates green again
gates / gates (push) Successful in 13s
Closes the two-release day-0 gap that has been convicting CI since 2026-08-14. Run against RUNBOOK-manual-build.md 4.0 + 4.1. GOLDEN_VERSION = 0.216.0 GOLDEN_SHA256 = ac004dc90d8cefccc5448377892f9cff3a4c3e1e27d0e11129120e38ac31c34b archive = 656,970,239 bytes, controller image 0.216.0 template = debian-13-standard_13.6-1_amd64.tar.zst (listed live, not reused) Baselines re-read on the machine and all four matched the sheet: controller v0.216.0, its MinAgent 0.129.0, agent v0.129.0, previous golden 0.214.0. The published agent artifact for the vouched agent_version was confirmed present in the package registry rather than inferred from a CHANGELOG, and the R-216 check passed on the machine: MinAgent is EQUAL to, not above, the newest published agent. Verified beyond the script's own claim: the artifact was downloaded back out of Gitea and hashed, and it matches GOLDEN_SHA256 exactly. A script printing a digest and the registry serving those bytes are two different claims. Pass markers (corrected post-R-233 list) all present, quoted with line numbers in pass-markers.txt; excluding/FATAL absent; there is no mp1. Token never reached a command line: copied file->file, read inside the VM by the runner. systemctl show grep = 0. Token-leak grep on the COMMITTED log run with its positive control FIRST -- seeded copy 1, real log 0 -- because a grep -c that matches nothing also returns 0. Teardown: guest destroyed and purged, token/runner/script/log shredded AFTER the log was copied out, qemu exit confirmed with ps -eo comm (not pgrep -f), disk reverted to virgin. Vouched by the operator; verified by reading the hub's own store: golden 0.216.0 / agent 0.129.0 / min_agent 0.129.0, and the hub's recorded sha256 matches the independently downloaded artifact. That check was necessary because golden_currency_gate.py says of itself that it checks the BAKE, not the vouch. repo_gates.py --fast now rc=0, all nine gates OK -- first fully green run since 2026-08-14. Capability map deliberately NOT changed: the day-0 row cites drill documents, and the map's only golden literal is a dated historical citation on the recovery-journey row which bumping would falsify. R-334 is closed in a follow-up commit quoting this push's CI run id, since closing it without one would leave the ambiguity a third time. |
||
|
|
3e50902a98 |
RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
gates / gates (push) Failing after 14s
Both STOPs cleared by the operator. No code changed; documentation only.
STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.
STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.
STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.
UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.
SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.
CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
- the "~85/day, ~2 years of runway" figures were WRONG. They came from a
single 17-minute window with a delta of ONE descriptor. Real rate is
183-200/day over two independent windows; runway ~357 days, not 2 years.
- the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
543 CLOSE-WAIT. R-336's fix must target unreaped connections.
- "proxmox-backup-api" reported inactive during verification; that unit does
not exist. Bad query, not a fault, written down because it looked like one.
R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.
golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
|
||
|
|
ebfd0967c1 |
INCIDENT + registers: ep0's PBS proxy served nobody for 9.5h (R-336..R-338)
gates / gates (push) Failing after 12s
Two whole_guest_backup_failed alerts at 04:30 and 04:32 CEST were one incident, and not on either customer box: ep0's proxmox-backup-proxy was active, holding its listening socket, and accepting nothing. Root cause: accept() returning EMFILE. The process held exactly 1024 fds -- its systemd-default soft RLIMIT_NOFILE -- of which 1016 were sockets and 547 connections sat in CLOSE-WAIT. The 1024-deep accept backlog had overflowed (Recv-Q 1025), so every client timed out. It was wedged from its own loopback too, which is what moved this from a network problem to a process problem. Fed by ~85k requests/day (a flat 3,538/hour) against an endpoint written to weekly, that leak reached the ceiling in 14 days of uptime. Fix: LimitNOFILE=65536 drop-ins for both PBS units, restart, verified from both boxes (200 in ~0.1s, felhom-pbs active), then re-drove the missed backups through the product path -- POST /backup?target=felhom-pbs on each agent's local API, not a hand-run vzdump. demo-felhom ct/9201/2026-08-18T03:57:43Z 4.10 GB 36.4s demo-hp ct/9201/2026-08-18T03:58:43Z 4.29 GB 41.5s Both host reports now carry felhom-pbs success=true, so the hub is green on the evidence rather than on a restart having been performed. No data lost, no backup skipped: the daily local tier was never affected and the PBS tier is weekly, so the window cost exactly one attempt. Evidence copied off ep0 BEFORE the restart, per standing rule 5. Filed: R-336 (the ~1 req/s poll rate is the real defect; the raised ceiling is mitigation, not a cure), R-337 (a status endpoint that trailed its own artifact by minutes then caught up -- WATCHING, downgraded from the defect I first wrote, because it self-corrected), R-338 (demo-hp is not on the R-50 island at all and nodes.md says it is; its local API is bound to the customer LAN). R-334 updated: still open, now one version wider (controller 0.216.0 vs golden 0.214.0). golden-currency is the only failing gate and is inherited -- it reads files this session did not touch -- so this push used --no-verify, stated per .claude/rules/gates.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016p1PTCzb8rF5G9Aa1qhpBN |
||
|
|
b03a105375 |
hub v0.105.0: the third name, a machine told to be quiet, and a guard for the hub's own words
gates / gates (push) Successful in 17s
Hub only. No controller change, no agent change, no wire change — nothing to bake. demo-hp untouched: the operator is re-deploying it this evening. R-323 — the five-word phrase is „Tulajdonosi jelmondat". It was „Visszaállító jelszó": one word from the name retired last week, and false besides — it restores nothing, it proves the account owns the box being bound. Five sites, all in the hub; felhom-controller and felhom-agent carry the name nowhere, so no halt and no bake. Both suggested names were rejected with reasons: „Fiókjelszó" would collide with the dashboard login (a DIFFERENT real secret), and „Összekötési jelszó" would leave the two factors on this page separated only by kód-versus-jelszó — the exact shape being removed, since the other factor is the „Párosító kód". The chosen name differs on both axes, stem and noun. Naming only; the acceptance pin drives the real handler. R-324 — the hub's customer copy is under a guard for the first time. Retired names banned across all 95 hub files; retrieval stems registered in four declared customer surfaces. The selftest found a defect in its own instrument on the first run. One shared vocabulary in scripts/, drift-checked into the controller gate rather than copied (R-325 removes the scaffold). R-321 — a machine we told to be quiet is no longer reported as dead, and it was two doors, not one: because the state is RECORDED rather than deleted, the morning deadline check can skip it too. A deleted state returns "", which is not "down" — R-195's shape returning through a second door. The clock runs from the report the hub can see, so re-enabling starts it there and emits no recovery for an outage that never happened. Three red-proofs; the one that matters showed a genuinely dead machine sitting at "disabled" when the suppression was made unconditional. R-326 — "which claims are unproven" is answerable by a command now. The nine I have been repeating was the count of claims the 9 August pass DOWNGRADED, not the count of unproven ones. The real figures: 55 claims, 23 walked, 32 not — and only 6 of those 32 cite evidence. Its first run found a stale claim (R-327). |
||
|
|
4d6ec7c7bb |
hub v0.104.0: the guest network gets a reader (R-319), and the hub half of the naming (R-295)
gates / gates (push) Successful in 14s
Four paper debts and one fact given a reader. Hub-only — nothing to bake. A4 — the entry about "the tester's machine" named a risk correctly and labelled it in a way that invited deleting it. Established from the hub's own store: `peti-felhom` is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47 this morning. The prompt's premise conflated the two; the register now says which is which. A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out of STATUS's "Waiting on you", which is now empty. A3 — day0-install §C.1 said pushing the installer publishes it. It has not since R-110. Corrected, with the two manifest pins named and an outside-verification command; the one copy that repeated it (a dated audit, true when written) carries a superseded note. A5 — standing rule 5: evidence comes off the machine at the end of the phase that produced it, before any revert. Earned twice in three days on the same box at the same point (R-320). Four homes, plus what to do when it is already gone. R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll` mail kind so the mail names the page a REBUILT box actually shows („A szerver beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach. Naming only; the acceptance pin proves the secret is untouched. R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never drawn as healthy — three absences, three sentences. No alarm, deliberately. Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190. B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now. Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled` reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed. Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect). |
||
|
|
fc737b0fc0 |
installer v1.28.0: the removal genuinely reverses the installation (R-316)
gates / gates (push) Successful in 13s
v1.27.0's fix worked exactly once per machine. Measured on drill-r50 from virgin, on the PUBLISHED v1.27.0, before anything was changed: cycle 1 recorded 'no' and freed :53; cycle 2 recorded 'yes' and left dnsmasq running on 0.0.0.0:53; cycle 3 refused, exit 1. Every box already in the field is at cycle 2, and a reinstall onto a machine that has had Felhom is cycle 2 by definition. Why cycle 2 says yes: the preflight's ownership question is dpkg-query package presence and nothing else - not the absence of a record. Stopping the unit and leaving the package made our own package read as the household's one cycle later. Now the uninstall removes the package when the record says we installed it. Order unchanged and load-bearing: read the record, act, then delete the state file that holds it. TWO packages are recorded, because dnsmasq ships the unit and dnsmasq-base ships /usr/sbin/dnsmasq, and each is taken back only if we added it. The dependency check is a SIMULATION, not a guess: apt-get -s purge is asked what it would remove and the purge proceeds only if that set is a subset of ours; otherwise stop+disable, naming the package that blocked it. Never interactive, never fatal, and the success is re-queried rather than read off an exit code. Watched: three fixed cycles -> install 3 PASSES; a household resolver untouched; a dependent package not purged and named; no record -> untouched with the command named. Red-proofs with the mutation asserted applied: remove the purge -> cycle 3 refuses in those exact words; remove the ownership check -> a household resolver is purged; infer ownership -> the guess is taken. Also: R-317 (the agent stats a path dnsmasq-base owns to decide whether to install dnsmasq - pre-existing, now reachable), R-318 (no honest ownership marker exists for existing boxes; the preflight message is the mechanism), and the status page's decisions section rewritten to say what each decision costs and what doing nothing selects. |