Commit Graph

294 Commits

Author SHA1 Message Date
admin 75d783ee2c BIGNIGHT F6: drive lost during backup — skipped apps, success:true, cooldown silenced mails; next run honest; R-519/R-521 amended; alarm table
gates / gates (push) Successful in 19s
2026-09-14 22:50:36 +02:00
admin 4ac6643819 BIGNIGHT F5: drive back as sdc, apps restart alone in 91s, data sha equal
gates / gates (push) Successful in 19s
2026-09-14 22:29:58 +02:00
admin ad8d072752 BIGNIGHT: alarm truth table draft (Phase 3-4, F1-F4)
gates / gates (push) Successful in 21s
2026-09-14 22:08:36 +02:00
admin 75b31c4562 BIGNIGHT F3/F4: drive unplug honest on screen, system apps unaffected; R-520, R-521; R-516 amended
gates / gates (push) Successful in 20s
2026-09-14 22:06:08 +02:00
admin c4ebd9cc63 BIGNIGHT F2 power cut during backup: guard restarts the stopped app, true alarm; R-519 (torn point dated by its newest part); R-517 amended
gates / gates (push) Successful in 20s
2026-09-14 21:51:40 +02:00
admin 97c9b01d1f BIGNIGHT F1 power cut: PASS, all 12 back on the same images in 4m03s, no false alarm
gates / gates (push) Successful in 20s
2026-09-14 21:40:00 +02:00
admin e46e525ef6 BIGNIGHT phase 4 done: tiers run, guarded update PASS, catalog reverted; box logs
gates / gates (push) Successful in 18s
2026-09-14 21:18:02 +02:00
admin 6aaa3a4a34 BIGNIGHT phase 4: R-517 (P1, backup page claims a failed PBS tier is current and present), R-518 (apps down ~8 min on 'a few seconds'); backup evidence
gates / gates (push) Successful in 20s
2026-09-14 21:16:16 +02:00
admin e0366f1f05 BIGNIGHT phase 3 done: 12 apps seeded and used, 0 interventions; R-515, R-516 filed; box logs
gates / gates (push) Successful in 20s
2026-09-14 20:58:22 +02:00
admin af256795b5 BIGNIGHT phase 3: R-514 filed (paperless OOM on a 20-document upload, silent); immich, jellyfin, mealie evidence
gates / gates (push) Successful in 19s
2026-09-14 20:47:01 +02:00
admin 5e8bff6808 BIGNIGHT: R-513 filed (P1 security: FileBrowser admin/admin on every box; demo-hp login page public)
gates / gates (push) Successful in 21s
2026-09-14 20:32:03 +02:00
admin 8a12c9a1bc BIGNIGHT phase 3: immich + vaultwarden seeded; R-512 filed (vaultwarden open signup, read-only control)
gates / gates (push) Successful in 18s
2026-09-14 20:27:52 +02:00
admin 3a7bbd2f6b BIGNIGHT phase 3: bookstack, docmost, privatebin, gokapi, nextcloud seeded and used; evidence
gates / gates (push) Successful in 20s
2026-09-14 20:25:27 +02:00
admin 4d92127f1a BIGNIGHT phase 2: claim with the mailed code, gate FAIL (R-510), R-511 filed (DR tier stuck), drive enrolled, box logs
gates / gates (push) Successful in 19s
2026-09-14 20:13:51 +02:00
admin 39a2627f40 BIGNIGHT: R-510 filed (tester-1 route lacks noTLSVerify -> 502) before intervention I2; bind + reenroll-mail evidence
gates / gates (push) Successful in 19s
2026-09-14 20:09:29 +02:00
admin d9522c38fb BIGNIGHT: R-509 filed (no self-bind mail for an existing customer) before intervention I1; journal + screens so far
gates / gates (push) Successful in 19s
2026-09-14 19:55:33 +02:00
admin 6e4d372720 doorstep teardown complete: host tester-1-8603a2 deleted, ep0 peer gone, customer kept; ep0 DR data retained for ruling
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 18:52:09 +02:00
admin 65790672d5 doorstep walk on ISO 1.27.x: 1 intervention (R-505), STOP before publish
gates / gates (push) Successful in 17s
ISO 1.27.1 gated PASS and proven live: first-boot console Felhom-only,
pvebanner masked across a proven reboot. Hub v0.113.0 hand-over copy live
(R-497 closed). Full first hour walked again on customer tester-1 (three
disks + one disk): deploy, use, backup, removal, byte-identical restore,
power cut, typo all PASS. The tunnel gives a fresh box no routes: 12/12
503 from DooPlex (R-505); the record has no e-mail (R-508). Rows R-507,
R-508 filed; R-496/R-495 fixed/answered awaiting publish; day-0 A.1 no
longer claims the controller creates hostnames (R-506). NOT PUBLISHED.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 18:25:01 +02:00
admin 8c7f882d1c drill 0242: teardown complete in three layers; R-501 filed
gates / gates (push) Successful in 19s
Machine: VM 330 destroyed with its disks. Host: ISO and scratch removed,
~8.5 GiB returned on nvme-scratch. Hub: customer drill0242 DELETED via the
cascade (journal #17), pages 404; its ep0 WireGuard peer dropped at the next
full-list push. Evidence pulled before the destroy. R-501: the documented
CI-check recipe reads only the last jobs page, which is not in id order.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 16:40:59 +02:00
admin 38848ffbeb drill: a stranger's first hour on 0.242.0 — 1 intervention, not ready for a volunteer
gates / gates (push) Successful in 21s
Golden 0.242.0 baked, round-trip verified and vouched (cadence rule, R-468).
Fresh box from the public ISO on demo-hp: landed on the vouched set, two apps
deployed and used, backup, remove, byte-identical restore, power cut and code
typo all PASS. Stopped for a volunteer by R-493 (no instructions) and R-494
(the setup mail's dashboard link has no DNS; intervention I1). R-493..R-500
filed. Capability map: first-hour row added (PARTIAL), journey row scoped.
Stopgap Hungarian volunteer guide written. Hub teardown layer pending.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 16:17:47 +02:00
admin 41590f8ee6 second night: scratch guest 9202 built (R-481 CLOSED, persists); controller v0.242.0 delivered (R-487 R-491 R-490 R-476 R-456 CLOSED, R-489 re-scoped); R-492 filed; rotation restarted from bentopdf; morning note
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 23:06:21 +02:00
admin 72ee053a9e evening: rules file in place; R-483 CLOSED (catalog, operator-confirmed); R-479 CLOSED (controller v0.241.0, proven live); R-481 blocked on the one-host-one-customer model; R-491 filed
gates / gates (push) Successful in 18s
2026-09-13 21:57:47 +02:00
admin 321770d9d6 R-452 CLOSED: the catalog-since gate (hook-enforced); 09 §8.2 limitation lifted; STATUS note updated
gates / gates (push) Successful in 18s
2026-09-13 19:42:42 +02:00
admin 5e8a82c3c4 night 2026-09-13/14: first "be a customer" rotation (adventurelog) — 7 defects found, 13 rows closed
gates / gates (push) Successful in 18s
New runbooks/nightly-rotation.md; observations_gate.py reads every section
(R-471); target-selection.md names real paths (R-461); R-93 carries the
fact that drill-r50 is gone. Register: R-473/R-474/R-466/R-471/R-453/R-461
and v0.240.0's R-477/R-478/R-480/R-482/R-484/R-485/R-486 closed; R-481,
R-483, R-487, R-488, R-489 opened. 09 §6.1, 07 §6, CONTEXT, STATUS note.
Evidence: audits/nightly-2026-09-13-adventurelog/, audits/v0240-2026-09-13/.
2026-09-13 19:38:00 +02:00
admin 681c3d6a6d docs: rulings 7 and 8 shipped and proven live (R-470/R-472/R-475 CLOSED); R-477..R-480 opened
gates / gates (push) Successful in 21s
Hub v0.112.0 serves a floor above the golden with a declared MinAgent;
controller v0.239.0 reached both demo boxes by that floor in 14 s and 15 s
and updates on any backup tier. 09 §3 decisions 7 and 8, §6/§6.1; 07 §6
line; capability map row; STATUS items 15/16 done and the cadence line
corrected; CONTEXT; register: R-470/R-472/R-475 compressed to CLOSED-ITEMS
(full text at 2f5d3af), R-477..R-480 opened, R-474 reproduced a third
time. OPEN-ITEMS 431689 -> 432156 bytes, CLOSED-ITEMS 118051 -> 120598.
Evidence: documentation/audits/rulings-r472-r475-2026-09-13/.
2026-09-13 17:54:23 +02:00
admin f181efd6a7 hub v0.112.0: a floor carries a declared MinAgent past the golden (R-472)
gates / gates (push) Successful in 18s
Operator ruling 2026-09-13. Above the vouched golden, a floor saved with a
declared MinAgent is served under the same agent comparison; an undeclared
one is still held beyond the golden. The declaration is stored beside each
floor as FLOOR=MINAGENT so it never carries to a later floor. Both floor
forms require min_agent above the golden (flash floor_needs_min_agent,
nothing stored). The Hosts page and the API log name the source.

Vouch path and R-120 gate untouched. Scenarios A-E tested; red-proofs A
and C in documentation/audits/rulings-r472-r475-2026-09-13/.
2026-09-13 16:47:11 +02:00
admin 5ef0f52bcd Slice 4 shipped (R-448/R-443/R-439 CLOSED, proven live); R-472..R-476; the floor-between-bakes claim corrected
gates / gates (push) Successful in 19s
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the
proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure.
Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/).

Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched
golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the
runbook, STATUS, CONTEXT, R-468 and the gate docstring.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 12:30:17 +02:00
admin ae59c31a84 R-459 CLOSED (MariaDB converts itself, proven by harness + live), golden 0.236.0 (R-467), the golden waiver (R-468)
Operator rulings 2026-09-13, both shipped the same day:
- MariaDB finishes its own conversion (catalog eec1228/bd32830/3525e35). Harness E3/E3b `proven`
  with engine_state_after "already upgraded to 12.3.3-MariaDB [exit=1]", the skip line gone, C3
  still `failed`; landed on demo-hp through the real 15-min cycle, nothing recreated, one deliberate
  restart logged "MariaDB upgrade not required" with the app serving. Evidence:
  documentation/audits/r459-close-2026-09-13/. The engine-major rule + gate keep every engine
  inside its major until Slice 4 (R-448) — removal tracked as R-469.
- Goldens on a cadence, not per release. golden_currency_gate.py reads a dated waiver
  (documentation/tests/golden-waiver.yml, <= 14 days, row-bound): valid + BEHIND -> loud advisory,
  exit 0; expired -> red again naming the date; UNRECORDED (R-385) never covered; malformed -> 2,
  never 0. Tests cases 5-15 incl. the R-421 decoy; red-proof old-vs-new on the real behind tree.
  R-242's vouch half stays open. Cadence in RUNBOOK-manual-build.md §4.2 + the checklist.
- Golden 0.236.0 baked, round-tripped, vouched, floor raised 0.232.0 -> 0.236.0
  (documentation/tests/golden-0.236.0-2026-09-13/) — the last per-release bake; the waiver was
  issued AFTER it landed. No --no-verify anywhere in this session.

Rows: R-459 CLOSED, R-467 CLOSED, R-242 narrowed; R-468/R-469/R-470/R-471 opened. 09 §3 gains
decisions 5 and 6; STATUS items 11 and 12 closed; CONTEXT records the cadence ruling.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 10:14:37 +02:00
admin 4b2e5608c2 R-442 CLOSED (controller v0.236.0): "delete my data too" deletes it or refuses; R-465..467 opened
gates / gates (push) Failing after 18s
- OPEN-ITEMS: R-442 row removed; R-465 (six remaining Paths.HDDPath readers — audit),
  R-466 (recovery-unit residue after "delete backups"), R-467 (v0.236.0 owes a golden).
- CLOSED-ITEMS: R-442 compressed, reasoning kept, fleet shape now established.
- 00-capability-map: lifecycle row narrowed (Campaign 3 proved remove removes the APP,
  not the data) and re-proven from audits/R442-2026-09-13/.
- STATUS: item 13 in plain language; item 7 closed (ruled 2026-09-02, 09 §3).
- audits/R442-2026-09-13/: live evidence (A, C, D bodies, controls, log window, teardown).

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 09:07:08 +02:00
admin d6837d98ee SPIKE R-459: the skipped MariaDB conversion is stable, and the trade it implied does not exist
gates / gates (push) Successful in 20s
Outcome A, qualified. Not B and not C.

It does not degrade: 5 of 5 restarts of 12.3 on an 11.6 datadir, readback passed
every time, mariadb_upgrade_info unchanged, the entrypoint line never escalated
past [Note]. It also never heals - the engine answers 'Major version upgrade
detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!' on every start
and will forever.

The trade R-459 was expected to produce is not real. Converting properly SUCCEEDS
across the multi-major jump, takes 7 seconds, backs up the system database
unasked - and putting 11.6 back afterwards STILL starts and serves the data. So
the operator is being handed a cheap correction, not a choice between a correct
engine and a reversible one.

The exit-code polarity was measured rather than read: 0 means the upgrade IS
needed, 1 means it is not. Assuming either the flag name or the polarity would
have inverted the headline. And run without credentials the same command returns
a confident-looking FATAL ERROR that is an auth failure.

R-464: after converting and going back, the entrypoint prints 'MariaDB upgrade
not required' on a state the same engine calls an unsupported downgrade. The
obvious cheap instrument for R-459 would have been to grep for that line, and it
would have reported fine for the broken case.

R-463: the PostgreSQL analogue, deliberately NOT measured here. 11 templates, 8
on postgres:16-alpine, register grep for pg_upgrade returns zero. The two engines
fail in OPPOSITE directions - MariaDB skips quietly, Postgres refuses to start -
so that one cannot hide; it presents as eight apps down at once.

No template changed. Teardown all three layers, hub checked rather than asserted,
local-lvm 30.53 percent before and after.
2026-09-06 17:42:19 +02:00
admin a1a6c73fe1 SPIKE: an upgrade test that runs again — and a real defect in our own bookstack template
gates / gates (push) Successful in 19s
R-449. Until today one app upgrade out of 53 had ever been measured, by hand, and
the whole update arc was designed against that single data point.

C3 first: the negative control, whose TO image exits immediately, came back
failed. That is what makes the greens mean anything, and it cost 556s because a
negative is only honest if it waits out the full settle window.

Seven edges, three apps. All five real catalog upgrades kept the customer's data.

The finding that changes an assumption the arc was carrying: whether an upgrade
can be UNDONE is a property of the individual APP, not of upgrades. Docmost
refuses - 'corrupted migrations: previously executed migration
20260213T085259-notifications is missing' - and privatebin does not. That
reproduces the Nextcloud result on a second app by a DIFFERENT mechanism, so the
struck word 'rollback' now rests on two measurements instead of one.

The finding nobody was looking for, R-459: our own bookstack template moves
MariaDB across a major and sets no MARIADB_* env at all, so the engine logs that
the datadir upgrade it requires is being skipped, and serves anyway. The cause is
assigned rather than guessed - the app half alone produces no upgrade line, both
edges that move the engine produce it - which is exactly what decomposing E3 into
E3a and E3b was for. It also explains why E3's abort looked like it worked: the
datadir was never converted. Whether that ever breaks is NOT established, and the
row says so.

Also opened: R-460 (bookstack's file half cannot be seeded headlessly), R-461
(target-selection.md names a venue that does not exist and fences a VM that is
gone), R-462 (the widening, costed with this run's real numbers - and the cost is
dominated by fixtures, which do not amortise).

Teardown all three layers, hub checked rather than asserted. local-lvm read 30.50
percent before and after. The capability map was deliberately NOT edited: this
measured apps, not the product.
2026-09-06 11:48:57 +02:00
admin 56c7e373a3 SPIKE: what an app update actually does, and which other paths do it too
gates / gates (push) Successful in 18s
THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has
already moved, and the Restart button does it — 18.3s with a network pull when
the target image is absent, 0.5s when present, against a negative control that
did not even recreate the container. The boot reconciler does the same thing
unattended when an app fails to come back (bootrecon.go:269 -> StartStack).

AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades
nothing. Docker restores the old containers and the reconciler logs 'no
boot-orphaned apps (nothing to start)'.

AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run,
the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2)
is higher than the docker image version (31.0.14.1) and downgrading is not
supported'. 'Rollback' is the wrong word for this arc and is struck.

Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator
confirmation. Peti's box was never contacted. No production code written in any
repo: felhom-controller is at 960d29b0612c before and after, tree clean, and
build/vet/test are green — run at the end to prove exactly that.

Register: R-438/439/440 updated with live evidence; R-441..R-445 opened
(restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB
left behind; update reports success over a broken app; no fleet fstrim; hub
telemetry outlives the app). Capability map gains three measured rows. The
mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR,
not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision.

Five claims in the brief are named as wrong, including two of my own method.
2026-09-01 21:35:32 +02:00
admin 10c223bdfe DRILL R-95: the recovery route does not exist — stopped before the destructive phase
gates / gates (push) Successful in 19s
The drill was to delete demo-hp's off-site history and get it back out of a Storage
Box snapshot, filling row 10's blank RTO. Phase 1 found there is nothing to get it
back from: no snapshot is reachable from a sub-account BY ANY NAME.

Measured, read-only, no delete verb issued against any live store:
- 777,600 exact names in the vendor form YYYY-MM-DDTHH-MM-SS, nine full days at
  second granularity, plus 126 alternative shapes -> ZERO hits.
- The control is what makes that mean anything: the identical 600-name batch shape
  with one real path appended returned it, 6 of 6.
- Structural cause: /home (u629488-sub3) is st_dev 0,82; /.zfs/snapshot is st_dev
  0,276; /home/.zfs does not exist. A snapshot under /.zfs/snapshot belongs to a
  different dataset than the one holding felhom-repo.
- Three tools agree with controls in the same run: SFTP, the port-23 shell,
  rsync --list-only.

So yesterday's re-scope splits: clause (a) "the box cannot write into the snapshot
area" STANDS and is re-confirmed; clause (b) "the rest is recoverable file by file"
is NOT SUPPORTED. STOPPED before Phase 2 on the operator's ruling — with no recovery
leg the deletion would have destroyed real history to buy only an alarm test that
could not fire at the specified size. Store verified untouched at 69 snapshots.

R-432 ANSWERED (negatively; its panel-read next step withdrawn as unnecessary).
R-433 no snapshot reachable by any name — decides R-95's remedy and its rank.
R-434 the drop alarm's text promises a file-by-file recovery that cannot be performed.
R-435 the drop detector is blind to a single-app deletion (>50% of 69 needed, ~9 given).
R-436 LEAD: the provider offers `rclone serve restic --stdio` and restic 0.14.0 speaks
      `rclone:` (measured, controlled) — real prevention may need no new machine, IF
      the vendor pins --append-only. Ask before building.

07 §8 row 10: text corrected, status NOT moved, RTO still blank.
No code, no version bump, no image, no golden. REPORT.md's only copy of the R-331
report preserved as REPORT-r331-backup-card.md before overwrite.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 16:55:50 +02:00
admin 0476a8d8e6 SPIKE R-95: the safety net cannot be seen from the box - and that RAISES the urgency
gates / gates (push) Successful in 17s
READ-ONLY STUDY. No code, no version, no image, no golden. No delete verb was issued against any
live store. ep0, DooPlex and Peti's box were not touched at all.

Q1 FIRST, AND IT DID NOT GO THE EXPECTED WAY. The brief supposed a working seven-day snapshot net
might bound the worst case. Measured on BOTH boxes over their own SFTP credential, with a positive
and a negative control on each: NO .snapshots is visible to either sub-account - not in the account
home, not inside the repo - and the account is jailed at /. Either none exist or a sub-account
cannot see them, and the second is not a reprieve: a snapshot the box cannot see is one the box
cannot restore from, so recovery would be an operator act at the Hetzner panel.

The register's claim rests on nothing that was checked. The row has NO R-number, so nothing can cite
it; its "confirm tomorrow" was 2026-07-27, 36 days ago; and the DUE-CHECKS block built for exactly
this (R-341) is EMPTY. R-95's word "ARMED" is withdrawn pending R-429. The confirming field is a
Hetzner API field, so this spike STOPPED at the section 11-D fence and left it for Viktor - ten
minutes in the panel, and it re-ranks everything.

Q2: TEN verbs, not nine. `check` was missing from the brief's list; `dump` is not a verb (it is a
progress phase constant) and was withdrawn. There are TWO `forget --prune` sites - offbox.go:1388
AND offbox.go:1759 - and disarming one without the other reproduces R-191 exactly.

Q3 (documented, from our own API mirror): AccessSettings has five booleans and `readonly` is the
only permission axis. No append-only. So the PBS shape does NOT transfer - PBS is a server that can
refuse; a Storage Box is a filesystem that runs nothing.

Q5 (measured, faithful append-only model, both controls passed first): withdrawing delete does NOT
wedge the store - restic treats a dead owner's lock as stale and proceeds. The constraint everyone
feared is not the blocker. But `unlock --remove-all` printed "successfully removed locks" while the
lock was still there, and resticStep's crash-lock self-heal is built on that call - R-430.

Q6 (measured, with a control): restic 0.14.0 DOES speak rest:. Append-only is a rest-server flag,
not a restic one.

Q7 (measured): detection is nearly free. snapshot_count already reaches the hub and the hub APPENDS
reports, so the history to compare against is already on disk.

RECOMMENDATION: answer Q1 today (Viktor, ten minutes), then build detection, then move retention off
the box. Defer the transport change until Q1 is answered.

Register: OPEN 179 -> 181. Filed R-429, R-430; R-95 updated and kept OPEN. 07 row 10's status is
deliberately UNCHANGED.
2026-09-01 13:55:35 +02:00
admin 574f5df107 the decoy sweep: 29 gates read, 16 fooled, 10 fixed - and a gate that refuses the next one (R-421)
gates / gates (push) Failing after 17s
THE CLASS, now a row: an instrument that matches a LABEL rather than the fact it names. Five
instances - R-410, R-400, R-378, R-419, R-94 - and EVERY ONE was found by accident, by someone
looking at something else. The gates enforce every other rule in this project, including the rule
that findings must be written down rather than left in prose. Nothing had ever checked the gates.

METHOD, and it is the transferable part: for each gate, construct the label WITHOUT the fact - a
directory with the right name and no bake log, a handler case that exists only in a comment, a note
whose prose mentions the marker it lacks - run the gate, record what it says. No verdict was reached
by reading. Reading is how all five hid.

RESULT: 29 distinct scripts (35 registrations; three are shared across three runners). 19 sound, 4
holes left OPEN with rows, 6 that no plausible decoy could be built for and are named UNTESTED rather
than called sound. A gate nobody tried to fool is UNKNOWN.

SCOPE IS A FACT TOO - the largest single cause, and mundane. Eight gates decided what to look at with
os.listdir, one level. Every one was green AND CORRECT today, and every one would have gone blind the
moment anyone added a subdirectory. mojibake and docker-v already used os.walk, caught the identical
planted file, and are the control that proves the cause was the listing and not the decoy.

IN THIS REPO: hub-confirm and manifest-bearer now walk. observations_gate (R-419, CLOSED) requires a
marker at a line start or after a sentence boundary and strips inline code spans - a note SAYING it
carries no marker no longer satisfies the marker test. closed-register now CONVICTS on a row it
cannot parse instead of warning: FOUR rows were in that state, TWO of them written by the session
that closed them the day before, and every one was exempt from the only check that reads that file.
The rows were repaired first and the conviction added second - registering a failing gate refuses
every push.

THE META-GATE: decoy_coverage_gate.py refuses a gate registered without a decoy or a named exemption.
It convicted ITSELF the moment it was registered, which is how it came to have one. Coverage is a
DECLARATION the gate AST-parses, never a grep - searching a test file for a gate's name would be the
very shape this sweep exists to find. The 20 uncovered gates are listed by name (R-426).

NOT FIXED, each with a row and a decoy asserting TODAY's behaviour so the fix must be deliberate:
R-422 reuse-refs (only 7 extensions; a rotted .md citation is invisible), R-423 site (PAGES is a
hardcoded list of 7), R-424 one-register (a defect parked as `idea`), R-425 offbox-rename (fixed
FILES list). R-427: closed_register_gate checks ONE direction - twelve open rows carry a closed
verdict and were NOT moved, because telling finished from partly-finished is a judgement and R-378
is the record of a machine getting it wrong.

FIVE DECOYS WITHDRAWN AS ILLEGITIMATE, mine, named in the audit. A decoy nobody would write proves
nothing, and manufacturing a finding to fill a row is worse than an honest NO.

No product code. No version bump. No image. No golden owed. All four runners green.
Register: OPEN 172 -> 178, CLOSED 160 -> 161.
2026-09-01 12:39:45 +02:00
admin f41a1a0ad8 R-411/408/407, R-414, R-412a leg 1, R-410, R-406 CLOSED; determination + live evidence
gates / gates (push) Failing after 18s
Part 2.1's determination is the first artifact: the scratch resolver was consciously OUT OF
SCOPE for R-356, not excluded on state-only grounds - established from R-356's own commit
08eb1a6, whose tests say 'the prepared scratch still resolves ... only the DESTINATION
moves'. So 07 section 6.3's rule applies and now has a FOURTH consumer, and the section
says so.

Live evidence: the collision rerun on demo-hp with the sampler positively controlled first
(12 locks=1 across a real check, 4 locks=0 quiet), showing unlock --remove-all 0 times where
the drill saw it twice; and the proof reaching verdict pass on demo-felhom - the box that
could not run it at all - recorded where last_proof_result had been ABSENT every night.

Capability map: the off-site proof row now records that the nightly firing IS proven (it ran
unattended at 05:30 on demo-hp) and that a driveless box can now be proved.

Register: six rows closed and compressed. OPEN 176 -> 170, CLOSED 152 -> 158.
2026-09-01 10:36:54 +02:00
admin 22e1c95e6a the golden gate reads a fact not a name (R-410); R-133's collision resolved (R-406, R-416)
gates / gates (push) Failing after 17s
R-410. golden_currency_gate.py matched EVIDENCE_RE against os.listdir and read nothing
inside, so `mkdir documentation/tests/golden-9.9.9-2026-01-01` turned it green with no bake
behind it - noticed while the 0.230.0 bake was running, when the evidence directory existed
before the bake finished. It now reads the GOLDEN_SHA256= line out of that directory's bake
log: a directory name is a label, that line is a fact only a completed publish produces.
Still offline, still --fast, one file read. Directories that look right and hold nothing are
printed by name rather than silently ignored, so a half-finished bake is visible.

test_golden_currency_gate.py ships the red-proof with a POSITIVE CONTROL, without which
"it fails on an empty directory" would be satisfied by a gate that fails on everything:
  CASE 1 empty directory -> rejected and named
  CASE 2 log with no GOLDEN_SHA256 -> rejected
  CASE 3 real bake log -> counted, and its sha read      <- the control
  CASE 4 the tree is left byte-identical
Red-proofed: reverting the gate to name-matching fails cases 1, 2 and 3.

R-406. Citations MEASURED before choosing, which is what the row asked for: hub-uniqueness
had 3 references (all inside one audit doc), plaintext-break-glass had 5 (CONTEXT.md,
break-glass.md, hub/CHANGELOG.md, the capability map, a spike). The FEWER-cited one moved -
hub uniqueness is now R-415 - and all three citations were rewritten to "R-415 (was R-133)"
rather than silently swapped.

THIS IS THE OPPOSITE OF THE TASK'S LITERAL INSTRUCTION, which said renumber the second row on
the stated ground that "the older number has the longer reference trail". Measured, that
ground points the other way. The principle was followed and the letter was not, and the row
says so rather than leaving an unexplained diff.

R-416 filed: the within-register duplicate rule was deliberately NOT added in the same commit
that removed its only subject - a guard whose red-proof can only be a planted fixture is not
this project's standard. Now that the register is clean it can ship with the next real
duplicate as its first subject.
2026-09-01 10:19:55 +02:00
admin cee8f70e98 soak drill COMPLETE: 7 phases, 4 findings, the observer found the worst one (R-414)
gates / gates (push) Failing after 17s
Overnight soak 22:39->06:10 CEST. demo-hp the victim, demo-felhom the untouched observer.
No production code, no golden, no version bump. Report at
documentation/audits/DRILL-soak-2026-08-31/REPORT.md.

VERDICTS: 1 lock-collision FAIL, 2 guard-interactions PASS-with-one-defect, 3 R-357 PASS,
4 proof-edges PASS, 5 mutated-cycle PASS, 6 observer FAIL, 7 teardown PASS.

R-414 - THE MOST VALUABLE FINDING, AND ONLY AN UNTOUCHED BOX COULD HAVE FOUND IT. On
demo-felhom the nightly proof fired for the first time unattended at 05:30 and REFUSED:
"nowhere to restore to - nincs regisztralt adatmeghajto". Cause established, not inferred:
storage_paths is EMPTY, so there is no path to put a scratch on. It will fail this way every
night forever with only a WARN, and because the error path reaches no verdict,
last_proof_result stays ABSENT - which is also what a pre-0.231.0 controller sends. The hub
cannot tell "never ran" from "not deployed": the StatsKnown trap one level up. The box is
NOT unprotected; its off-site backup ran fine in 46.9s. It is the PROOF that cannot run.

R-411 - measured, not reasoned: restic stats TAKES A LOCK; a customer full-restore runs it
while holding no acquireRunning; the integrity check is therefore not blocked, meets that
lock and escalates to unlock --remove-all. The sampler caught "restore ..." and
"unlock --remove-all" in the SAME sample. Contained: the check was classified unreachable,
not damage, so no false alarm.

R-412 - CORRECTED from HIGH to LOW. I filed it on a mechanism I had not finished measuring.
The off-site run has its own pre-push dump leg, so a hollow unit is REPAIRED before it
ships - proven on two apps and confirmed by pulling the snapshot back out of the store.
What survives is a narrow race, plus a success line over a backup holding no data.

R-413 - the R-87 proof caught a product-produced hollow snapshot unattended, and the
nightly job fired on its own schedule at 05:30 for the first time (bentopdf PASSED on
9d002b38 in 2.315s). Both were listed "not yet live-validated" yesterday.

R-403 mirror guard PROVEN live, with a negative control: it fired when a unit was hollow
("The copy was PRESERVED rather than replaced with an empty one") and skipped 0 legs at
teardown when every unit was sound.

R-357 PASS at last, six days owed: a real full filesystem, refused BEFORE StopStack, app
never stopped, live data byte-identical, and it worked once the space came back.

Phase 4 built the false-alarm control the whole R-87 design rests on: bentopdf is the only
template of 53 with neither a database nor a named volume. It passes silently.

EIGHT of my own instrument errors are named in the report, each caught by its own control -
including a time guard that fired an injection four hours early, and filing R-412 at the
wrong severity.

Teardown clean on all three layers of both boxes; both healthy on 0.231.0.
OWED: a golden for 0.231.0, and a decision on keeping bentopdf.
2026-09-01 06:07:13 +02:00
admin ab8b884763 soak phase 5: R-403 guard PROVEN live; R-412 CORRECTED down after measuring the mechanism
gates / gates (push) Failing after 18s
R-403 MIRROR GUARD - PASS, forced after the natural test evaporated. privatebin was
injected hollow at 23:34 to meet the 03:30 mirror; the 02:30 db-dump re-made its tar, so
by 03:30 the primary was complete and the guard had nothing to refuse. Forced instead on
calibre-web through the real Tier-2 path: the guard fired and named itself -
"unit leg SKIPPED ... The copy was PRESERVED rather than replaced with an empty one
(R-403). The other legs continue." Secondary byte-identical, 23 files, tar sha d7e7f422.

R-412 CORRECTED, AND I OVERSTATED IT WHEN I FILED IT. The first wording claimed the hollow
unit sits in the store for a whole cycle because the volume-dump leg runs only on the
backup schedule. That is WRONG. The 04:15 off-site run has its OWN pre-push dump leg -
"Stopping calibre-web for safe volume dump", "Volume dump: ... -> 877.5 KB" - so a unit
that is hollow when a run starts is REPAIRED before it is pushed. Measured twice: opengist
and calibre-web both went in hollow and came out complete, and the snapshot pulled back
from the store (6fee3b5a) holds the volume tar and all 17 userdata files.

What remains real is narrower: the one hollow snapshot that DID reach the store was created
when the unit was destroyed INSIDE a run that had already completed that app's dump leg.
The race is real and was observed, and "backed up opengist (... 0 mandatory path(s))" is a
success line over a backup holding none of the app's data either way. Severity HIGH -> LOW,
with the correction stated in the row rather than quietly rewritten.

Phase 6 interim: the observer is clean so far - db-dump 674ms, tier2-backup 3ms (a no-op,
cause to be established not assumed), zero ERROR/WARN since 23:00.

Also recorded: two of Phase 5's four injections were NOT performed, with the reasons
established rather than asserted - there is no endpoint that reaches SetDisconnected and a
hand-set flag would be reverted by the live monitor before 04:15; and a corrupted manifest
provably never reaches the store because the capture rewrites it first.
2026-09-01 04:21:15 +02:00
admin 585ed654b4 soak drill phases 0-4: R-408's hazard is REAL and reachable; R-357 finally live (R-411..R-413)
gates / gates (push) Failing after 17s
Overnight soak on demo-hp, phases 0-4 of 7. Evidence
documentation/audits/DRILL-soak-2026-08-31/. No production code written, no golden,
no version bump - findings only.

PHASE 1 FAIL - R-411. The R-408 hazard was only ever reasoned about; tonight it was
measured through the product's own endpoints. restic stats TAKES A LOCK (clean-room:
nothing else running, 4x stats, sampler reads locks=1). A customer full-restore runs
stats in its preparation while holding NO acquireRunning, so the integrity check is not
blocked, runs, meets that lock, and resticStep escalates to unlock --remove-all - the
sampler caught "restore ..." and "unlock --remove-all" in the SAME sample at 20:50:51.
The log calls it "a stale exclusive lock left by a previous crash"; there was no crash.
THE CUSTOMER-FACING CONSEQUENCE IS CONTAINED and that is R-359 working: the check was
classified Unreachable, NOT damage, so no alarm and due-ness held. The opposite direction
is FENCED - five restores fired into a running check at 5/15/25/35/40s were all refused by
restoreOpBlocked with zero restic invoked.

PHASE 2 FAIL - R-412, and it is the natural instance of R-403's shape that yesterday's
session had to hand-build. A lost primary unit is rebuilt by the capture WITHOUT its
volume dumps (185664 B -> 4382 B, volume_dumps: None), and the off-site backup then ships
it and logs "backed up opengist ... 0 mandatory path(s)" - a success line over a backup
holding none of the app's data.

PHASE 2 also R-413: the R-87 proof CAUGHT that hollow snapshot unattended -
verdict "fail", volumes_expected_none_captured: opengist_data, one offsite_proof_empty at
severity error, while the four apps ahead of it in the rotation passed.
2.3 stale marker PASS, 2.4 two restores PASS, 2.5 corrected the runbook's premise
(RunTier2 has zero acquireRunning and zero restic references - a local mirror cannot
collide with the repo, so the proof correctly does not skip for it).

PHASE 3 PASS - R-357 live-validated at last, six days owed. A real full filesystem
(1900544 B free vs 5878421 B needed) on a separate 1 TB device that does not back Docker.
Refused BEFORE StopStack: app "Up 6 minutes (healthy)" unchanged, 0 safety dumps, live
userdata tree fingerprint 5a9b2db75dfd4ba4131adb3255670472 unchanged, message names both
numbers in Hungarian. Ballast removed, the SAME restore then worked and the fingerprint is
still identical. On the same full disk the Tier-2 run, the proof and the integrity check
all behaved: the proof refused before any download through the shared unitOnlyHeadroom
gate extracted today, reached no verdict and did not alarm.

PHASE 4 PASS - the false-alarm control the whole R-87 design rests on now EXISTS. Found by
reading all 53 catalogue composes: bentopdf is the only template with neither a database
service nor a named volume. Deployed, backed up, proved: PASSED in 2.224s with zero alarms.

Four of my own instrument errors were caught by their own controls before any result was
believed: a hub log line used as a controller positive control, a grep pattern that missed
a registered job, a heredoc that mangled a planted marker, and a catalogue scan that read
0 of 0 apps from the wrong path.

Phases 5-7 to follow: the real nightly cycle with injected faults, the untouched observer,
and teardown.
2026-08-31 23:33:32 +02:00
admin 130f7a6eba R-87 SPIKE: measured, do not build it as written (R-407..R-409 filed)
gates / gates (push) Failing after 17s
Spike. NO production code. No version bump, no build, no deploy, no golden.
felhom-controller and felhom-agent were READ ONLY. The fleet stays on v0.230.0.

Q1 restic is 0.14.0 (go1.19.8, bookworm 12.15) - the four source comments asserting
it are CONFIRMED, not corrected.

Q2 --verify DOES exist and is NOT a content check. Red-proof: one byte changed in a
restored 160 MB tar with size and mtime preserved passed clean, rc=0. Verify took
131 ms on a 213 MB / 7-file tree, which cannot be hashing. A size or mtime mismatch
causes a silent re-download, not a failure. Controls: --target 1 hit, four post-0.14
flags and a nonsense string 0 hits each. Neither --verify nor --no-lock appears
anywhere in the controller source.

Q3 no reference for "correct" exists. restic ls --json carries no content hash in
0.14.0, and the unit manifest hashes 4918 B of a 213231242 B unit - 0.0023 percent,
the config files and not the dumps or the tars. R-409.

Q4 it is CHEAP. All 8 apps / 774378123 B logical restored back to back in 25 s, against
40257 ms for the weekly 100 percent check beside it. Individual restores 2253-3978 ms
regardless of size: cost is per-snapshot round-trip plus ~1 s per 200 MB. Peak scratch
is the app's full logical size. The 1.1 MB restic cache is index only and hides nothing
(--no-cache 5423 ms vs cached 3198 ms, trees byte-identical).

Q5 skip-if-busy stays right at 25 s against a 2m52s nightly backup. But
RestoreOffboxScratch takes NO acquireRunning, while offbox_integrity.go:28 asserts
every off-site operation does. R-408.

Q6 observed with a positively-controlled lock sampler: restic restore takes NO lock;
restic check DOES (locks 0 -> 1 for nine samples -> 0 across the check, zero across two
restores). The product writes anyway - unlockStale runs `restic unlock`, a delete verb,
before every restore (offbox_restore.go:289). The task's lead was right in direction and
wrong in mechanism. R-95's constraint IS satisfiable: --no-lock plus skipping unlockStale
writes nothing, and both mechanisms exist unused. Neither was fixed - the task forbids it.
offbox_integrity.go:255's "It NEVER writes to the repository" is R-407.

Q7 THE DECIDING ONE: of R-353/354/356/358/403 an unattended scratch-restore would have
caught ONE (R-356). The value is elsewhere, and the weekly check structurally cannot
reach it: `check` proves the stored bytes are the stored bytes, never that we stored the
RIGHT thing. A hollow unit backs up, checks at 100 percent and restores cleanly and
recovers nothing - R-403, measured in bytes on 31 August.

RECOMMENDATION: option C, the NARROW test - one app a night, restored to scratch, checked
against its own manifest.json through the existing unitCarriesData, scratch deleted, the
SNAPSHOT recorded as the proof. Options A (do not build) and B (scheduled attended drill)
considered explicitly; B is weakest because it is what already happens. R-87 should be
RE-SCOPED, not built as written, and that is Viktor's call - the row stays open carrying
the verdict and STATUS.md item 4 asks it in plain words.

Also corrected in 07-backup-architecture.md: matrix rows 4 and 10 both said "the depth
that ships ON does not re-read pack contents (R-399)". R-399 CLOSED in v0.228.0 and the
depth is 100 percent. Two stale cells, fixed, and the spike verdict added beside them.
Row 4's verdict is UNCHANGED by the spike and now says so.

Teardown: all three layers, none of them "nothing was created" - 6 files on the PVE host,
9 in the guest, 5 plus 2 run-flags in the container, all removed and verified empty. The
four scratch directories this session's restores created were removed; three that
pre-date the session were left alone. Two state changes recorded rather than hidden: the
control integrity run recorded its verdict (depth structure -> 100%, due-ness +7 days),
and four restores appear in the controller log. Nothing was written to the off-site
repository by hand.

Evidence: documentation/audits/evidence-spike-restic-restore-2026-08-31/ - 31 files,
every one pulled off the box BEFORE teardown (R-320).

golden-currency is RED at this commit and was already red at dddcc80. Pre-existing, not
this session's debt. Second --no-verify push of the day for that reason; R-404's count
goes six -> seven and its row says so.

Ceiling R-406 -> R-409.
2026-08-31 15:59:09 +02:00
admin dddcc808be R-403 CLOSED (controller v0.230.0), R-404 filed as a decision for Viktor
gates / gates (push) Failing after 17s
07-backup-architecture gains section 8.2, placed beside row 5 on purpose: the derived-copy rebuild
rule is UNCHANGED and section 8.2 names the single exception, so a future reader who finds RunTier2
skipping a leg does not fix it back. It carries the measurement (120 082 104 B -> 7 036 B on the
shipped v0.229.0), the four-case table, why hollow is a manifest question and not a size question,
why the data legs are deliberately not guarded, and why the capture job is not guarded either.

00-capability-map: the Tier-2 row's status does NOT move, stated explicitly rather than left
ambiguous. R-403 removes a way the route could be DESTROYED between uses; it does not change what
the route can be relied on for.

Register: R-403 CLOSED and compressed into CLOSED-ITEMS (594 -> 593 open lines). R-404 FILED as a
DECISION and deliberately not acted on - should a documents-only push be subject to the
golden-currency gate, now that it has been correctly bypassed six times? Both sides stated, plus
what happens if Viktor does nothing. The gate was NOT changed.

R-242: seventh conviction, and the FIRST where the day-0 ground does not apply - R-403 is a defect
in the nightly Tier-2 copy, which a day-0 box starts running on its first night. This push uses
git push --no-verify, declared here and in felhom-controller/REPORT.md. A golden carrying 0.230.0 is
owed and is more urgent than the previous six.

STATUS: the R-403 item moves out of 'Broken' into what works, in plain words; the delivery item now
says a golden is owed and that the fleet carries the defect; R-404 goes into the decide section with
its do-nothing outcome.

Drill evidence: nine phase logs, including the two things that went wrong (a repair whose rsync was
not installed in the guest and silently did nothing, and a session that expired mid-run so a POST
did nothing).
2026-08-31 14:39:29 +02:00
admin 66156c619f R-403 drill evidence + the credential reader that ends a three-time mistake
gates / gates (push) Successful in 16s
The drill: the loss reproduced on the shipped v0.229.0 before anything was built. 120 082 104 B ->
7 036 B in one Tier-2 run, recorded as a success. Phases 1a (before), 1b (the hollow primary,
produced through the R-102 restore path exactly as the 2026-08-31 observation was), 1c (the loss),
1d (repair).

scripts/read_credential.py is Part 4's rider, and it exists because a note did not work three times:
2026-07-20 a Failed login was diagnosed as a stale password and written into memory; 2026-08-31 the
same misreading recurred and was caught; 2026-08-31, hours later, it recurred AGAIN and rewrote a
live box's password hash. Between them the project already had a memory file stating the rule, a
worked recipe in it, and a session report describing the mistake. The rule now lives in the code
path: one matching quote pair is unwrapped, the result is REFUSED if it still carries a quote, and
--expect-length gives the caller a second opinion. The value goes file->file at 0600 and stdout gets
only its length. test_read_credential.py asserts each refusal by its reason, with a positive control
before believing the not-in-stdout result.

Red-proof E1: remove the final quote assertion -> three cases fail by name.
2026-08-31 14:02:26 +02:00
admin c2de785bf2 R-102 + R-103 CLOSED (controller v0.229.0) — architecture, register, STATUS, drill evidence
gates / gates (push) Failing after 17s
07-backup-architecture: 6.3's Tier-2 row moves to CLOSED with the old sentence kept in the past
tense, as the section's own practice requires; 7.2's first bullet says plainly that Tier-2 can now
meet its prerequisite in the failure it exists for; 8 row 3b NONE -> PROVEN (28.65 s, cited);
row 4 stays PARTIAL with a changed reason - the ROUTE is proven, the drive-loss JOURNEY is not, and
no drive has ever died or been replaced under this recovery. 8.1's blanks updated per row.

6.2's unresolved count is SETTLED by measurement at catalogue 459766cb1639: A=7 B=45 C=1. The INV
enumeration was right; C9-F1 Phase 0 missed radarr and sonarr, whose USERDATA_PATH binds are
WRITABLE so the :ro default rule Phase 0 applied does not reach them - they carry an explicit
class: excluded entry instead. Class C is bentopdf. No catalogue file was changed.

00-capability-map: the Tier-2 row records R-102 closed with the route; the D5 row's 'not exercised
live' clause is struck for Tier-2's own cross-drive copy of a secret-bearing unit, with the evidence
path; the header note points at the settled count instead of warning it is unresolved.

Register: R-102, R-103 and their C9-F4 / C9-F1b aliases closed and compressed into CLOSED-ITEMS
(596 -> 593 lines, each naming git show 1623a4d5b5 for the original). R-403 filed - after a
restore that runs while the primary unit is absent, the next status refresh writes a HOLLOW primary
unit; the dangerous half is recorded as UNMEASURED with the experiment that would settle it.

R-242: sixth conviction of golden_currency_gate. THIS PUSH USES git push --no-verify, declared here
and in felhom-controller/REPORT.md - a BYPASS, not a waiver. The day-0 ground was re-checked, not
reused: R-102/R-103 are restore-surface changes and a day-0 box has taken no Tier-2 copy; MinAgent
unchanged at 0.129.0. OWED: bake a golden carrying 0.229.0, vouch it, raise the floor.

Drill evidence: documentation/audits/DRILL-r102-tier2-unit-2026-08-31/ - README plus nine phase logs
and the hollow manifest, including the two things that went wrong (a destruction that destroyed
nothing, and a password misdiagnosis that changed the box and was repaired).
2026-08-31 12:21:52 +02:00
admin e027b5d999 Register + architecture for controller v0.226.0 (R-353/357/358/360/396), and R-395 fixed
gates / gates (push) Failing after 17s
Closes R-353, R-357, R-358 and R-360 with their shipping version and evidence
path, and files two new rows.

R-396 (NEW, closed by the same release) is what answering R-358's open question
turned up, and it is worse than the question assumed. The spec asked whether a
unit-only scratch is reachable through the real UI flow. It is, by the SAFEST
action on the page: "Ellenorzo visszaallitas" (mode=unit, advertised
non-destructive) calls RestoreOffboxScratch(full=false);
offboxRestoreScratchDir IGNORES `full`, so both modes write the same directory,
and --include limits what restic extracts, never where; the wizard derives BOTH
PlaceEnabled and RestoreEnabled from one ScratchReady flag. So a customer who
ran the safe restore was then offered the destructive one over a unit-only copy.
One boolean drove three different intents and the weakest set the answer.

R-395 (filed by the spec) is fixed in this commit, not just recorded. STATUS.md
said golden 0.223.0 / floor 0.222.0 in one block and demo-hp 0.219.0 / floor
0.218.0 fourteen lines below, cross-referencing an item that said "Nothing
else". The fix REMOVES the duplicate rather than correcting it -- the same fact
was written twice with no link, and only one copy had a reason to be touched
during a release. "What works" now points at the item above instead of restating
a version.

07-backup-architecture: four rows added to the 10.2 gap register plus R-396.
Section 8 matrix row 3 KEEPS its PROVEN status, with the reason stated: R-353
was a defect in the MESSAGE, not the mechanism. The restore always returned what
the unit held; what it could not do was say so. A status that measures whether
data comes back must not move because a status line was wrong.

00-capability-map: one new row, and it splits what is claimed. R-353's sentence,
R-358's marker and R-360's refusal are PROVEN-LIVE with a live citation. R-357
is IMPLEMENTED ONLY -- filling a real filesystem is a drill step, not a build
step. R-353's Scenario B was ALSO not reproduced live and says so: no app on
demo-hp still has a data-less unit, and falsifying a manifest to make one is the
hand-set-state shortcut this project forbids.

This push used `git push --no-verify`. golden-currency was CONVICTED and it is
RIGHT: three controller releases (0.224.0, 0.225.0, 0.226.0) and the golden
still carries 0.223.0. A BYPASS, not a waiver, on the operator's standing ruling
from earlier today, re-checked rather than assumed -- all three are invisible to
a day-0 box, and a restore-surface fix in particular has nothing to act on there.
The ground expires the moment a release changes first-boot behaviour. Tracked on
R-242; ONE bake carrying 0.226.0 covers all three.
2026-08-30 19:39:41 +02:00
admin 36f8630020 R-341 check taken, and the golden-currency bypass declared
gates / gates (push) Failing after 16s
Two gates blocked the R-331 hub push. One is FIXED, one is BYPASSED, and the
difference is stated rather than blurred.

FIXED -- due-checks (R-341, 5 days overdue). The +7d measurement was TAKEN on
ep0 rather than deferred again. Precondition passed: proxy still MainPID 551655,
ps -o lstart= still 2026-08-18 09:51:04, NRestarts=0, so this is the same proxy
generation as t0 (anchor is ps, not ActiveEnterTimestamp, which reads 03:54:54Z
here -- R-346's trap).

Result: fd = 17. Not 17 more -- seventeen TOTAL, exactly the documented
baseline, against 405 at the first check. Socket histogram: one LISTEN, ESTAB 0,
CLOSE-WAIT 0.

The verdict is UNANSWERABLE, not "the upgrade fixed it". R-341 asks whether the
PBS 4.2.5-1 upgrade changed the fd slope; inside this interval we removed the
leak OURSELVES (R-344, agent 0.130.0, now live on both boxes). A slope of ~0
measures our fix, not the upgrade, and reading it the other way would credit a
changelog that was read in advance and found to contain no such mechanism. The
perturbation pre-registered for this window was Phase C at ~3%; the actual
perturbation was the removal of the entire phenomenon. Row closed as moot.

What it DOES establish is worth more than the original question: twelve days
after the R-344 fix, same proxy generation, no restart to hide behind, ep0 sits
at baseline with zero established connections. R-336's ~323-day runway concern
retires with it.

BYPASSED -- golden-currency. Controller v0.224.0 and v0.225.0 are released and
the newest golden bake carries 0.223.0, so a machine installed right now gets
neither. The gate is RIGHT. This push therefore uses `git push --no-verify`,
declared here, in hub/CHANGELOG.md, in REPORT.md and on R-242.

A BYPASS, not a waiver: the gate offers a waiver only for a release that
DELIBERATELY needs no golden, and these need one. The operator was asked and
ruled bypass-now-bake-later, on the ground that neither fix bites a day-0 box --
R-330 is a nightly false alarm about apps a new box has not installed yet, R-331
is a hub display over backups a new box has not taken yet -- and both arrive by
self-update. That ground is recorded because it is what to re-check: it does NOT
extend to a release changing first-boot behaviour.

OWED: bake a golden carrying 0.225.0 and vouch it (RUNBOOK-manual-build.md 4.1,
three-field change, MinAgent 0.129.0). Fourth bypass of this gate, and the gap
is now two releases wide rather than one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 18:49:53 +02:00
admin ebdc04601d docs(hub v0.108.0): the delivery grain, the cooldown ruling, and gate 11's first subject
gates / gates (push) Successful in 15s
The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or
coarse, and why the default is coarse. CONTEXT records two rulings: the grain is
allow-listed rather than inferred from the payload, with crossdrive_failed as
the proof that a payload rule would have been wrong; and a finding recorded only
in REPORT.md has a lifetime of one session.

R-389 closed and compressed, keeping its rules and naming the commit whose
git show returns the full text. R-390 and R-391 left open.

REPORT.md is gate 11's first real subject and passes: six observations, two
FILED, four NOT-A-FINDING with their reasons. Three of those declarations are
things a tidier report would have omitted - the gate's own spec would have
passed the item it was built to catch, the burst has no ceiling, and ArgoCD
said "successfully rolled out" while still running the old image.

STATUS carries forward the one thing outstanding: the controller floor still
reads 0.222.0 while the golden reads 0.223.0.
2026-08-23 14:12:23 +02:00
admin 2f7c9a6ce5 docs(R-329/R-386/R-387): the severity contract, the intent ruling, and Part 5 recorded
gates / gates (push) Successful in 17s
The alarm ladder gains the severity contract (the hub's vocabulary is exact, it
coerces silently, and three things now hold it) and the intent test with its
three-way ruling on unknown. Both marked [DESIGN] with the live measurements.

Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy,
verbatim, marked plainly as direction rather than current behaviour, with the
12 -> 15 toggle growth as the argument. Filed as R-388, a product decision.

R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the
full-text commit named. R-387 filed closed - including WHY the dispatcher branch
was kept rather than deleted, which is evidence (three monitor checkers call
ProcessEvent directly) and not caution.

The drill record names three things that had to be re-run: an inert red-proof
mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A
NOT proving the customer gate because demo-hp has no prefs row at all.

Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B.
2026-08-23 12:03:49 +02:00
admin 55274d5ef3 R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s
The gate failed only on `released > baked`, so it could catch a forgotten bake
and nothing else. A golden AHEAD of the record passed silently - and that is
how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG
heading still read v0.221.0, with every gate green. Reproduced on the real
history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0.

The gate now asks whether the version being shipped is WRITTEN DOWN: the baked
version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG.
Membership rather than `baked > released` deliberately - a comparison against
the newest heading alone goes green the moment any later entry is written,
leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2)
preserved; every refusal names a reason and a route.

Red-proofed both directions: old gate/old record exit 0, new gate/old record
exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0.

08-alarm-ladder.md is new, and its absence was itself the finding: no document
owned "when does a broken app raise an alarm?". The rules lived as comments in
four packages, each locally correct, with the ordering between them legible only
by reading one function top to bottom - which is how R-384 survived review.

R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed
closed. R-386 filed OPEN: a single-container app stopped out of band raises no
alarm, and a comment claims the opposite - measured live, 9 scans, 0 events,
against a positive control from the same box 17 minutes earlier. Not fixed here.

Golden 0.222.0 baked and published; vouching is the operator's act.
2026-08-23 07:59:52 +02:00
admin 1eb64bec51 R-361 docs: the [FACT], the negative that cancelled Part 2, R-383/R-384, golden 0.221.1
gates / gates (push) Successful in 17s
07-backup-architecture.md gains a dated [FACT] on R-361 - a comment asserting an
invariant the code did not have, for four months - and a [DESIGN] on the db_dumps
decision INCLUDING the trap it created: a stable list lets the already-current
early return fire, so per-capture housekeeping must sit above it.

00-capability-map.md records the NEGATIVE from Part 3 so it is not re-derived: a
held app does NOT raise the dead-app alarm. It aggregates to unhealthy, which
IsDownState excludes. Measured on the shipped build with the scans demonstrably
running over it. No suppression was built and no row opened.

R-383: the double-failure message names an undo copy that is not there - R-361's
own class, one surface over, observed on both 0.220.2 and 0.221.1.
R-384: an app whose database has died reads unhealthy and raises no alarm.

R-361 closed and compressed. OPEN-ITEMS 325236 -> 327266 bytes.

Golden 0.221.1 baked, published and round-trip verified. The golden-currency gate
blocked this push and that block is not circular, so it was satisfied rather than
bypassed - no --no-verify anywhere in this session.
2026-08-23 00:33:12 +02:00