Commit Graph

636 Commits

Author SHA1 Message Date
admin ab8b884763 soak phase 5: R-403 guard PROVEN live; R-412 CORRECTED down after measuring the mechanism
gates / gates (push) Failing after 18s
R-403 MIRROR GUARD - PASS, forced after the natural test evaporated. privatebin was
injected hollow at 23:34 to meet the 03:30 mirror; the 02:30 db-dump re-made its tar, so
by 03:30 the primary was complete and the guard had nothing to refuse. Forced instead on
calibre-web through the real Tier-2 path: the guard fired and named itself -
"unit leg SKIPPED ... The copy was PRESERVED rather than replaced with an empty one
(R-403). The other legs continue." Secondary byte-identical, 23 files, tar sha d7e7f422.

R-412 CORRECTED, AND I OVERSTATED IT WHEN I FILED IT. The first wording claimed the hollow
unit sits in the store for a whole cycle because the volume-dump leg runs only on the
backup schedule. That is WRONG. The 04:15 off-site run has its OWN pre-push dump leg -
"Stopping calibre-web for safe volume dump", "Volume dump: ... -> 877.5 KB" - so a unit
that is hollow when a run starts is REPAIRED before it is pushed. Measured twice: opengist
and calibre-web both went in hollow and came out complete, and the snapshot pulled back
from the store (6fee3b5a) holds the volume tar and all 17 userdata files.

What remains real is narrower: the one hollow snapshot that DID reach the store was created
when the unit was destroyed INSIDE a run that had already completed that app's dump leg.
The race is real and was observed, and "backed up opengist (... 0 mandatory path(s))" is a
success line over a backup holding none of the app's data either way. Severity HIGH -> LOW,
with the correction stated in the row rather than quietly rewritten.

Phase 6 interim: the observer is clean so far - db-dump 674ms, tier2-backup 3ms (a no-op,
cause to be established not assumed), zero ERROR/WARN since 23:00.

Also recorded: two of Phase 5's four injections were NOT performed, with the reasons
established rather than asserted - there is no endpoint that reaches SetDisconnected and a
hand-set flag would be reverted by the live monitor before 04:15; and a corrupted manifest
provably never reaches the store because the capture rewrites it first.
2026-09-01 04:21:15 +02:00
admin 585ed654b4 soak drill phases 0-4: R-408's hazard is REAL and reachable; R-357 finally live (R-411..R-413)
gates / gates (push) Failing after 17s
Overnight soak on demo-hp, phases 0-4 of 7. Evidence
documentation/audits/DRILL-soak-2026-08-31/. No production code written, no golden,
no version bump - findings only.

PHASE 1 FAIL - R-411. The R-408 hazard was only ever reasoned about; tonight it was
measured through the product's own endpoints. restic stats TAKES A LOCK (clean-room:
nothing else running, 4x stats, sampler reads locks=1). A customer full-restore runs
stats in its preparation while holding NO acquireRunning, so the integrity check is not
blocked, runs, meets that lock, and resticStep escalates to unlock --remove-all - the
sampler caught "restore ..." and "unlock --remove-all" in the SAME sample at 20:50:51.
The log calls it "a stale exclusive lock left by a previous crash"; there was no crash.
THE CUSTOMER-FACING CONSEQUENCE IS CONTAINED and that is R-359 working: the check was
classified Unreachable, NOT damage, so no alarm and due-ness held. The opposite direction
is FENCED - five restores fired into a running check at 5/15/25/35/40s were all refused by
restoreOpBlocked with zero restic invoked.

PHASE 2 FAIL - R-412, and it is the natural instance of R-403's shape that yesterday's
session had to hand-build. A lost primary unit is rebuilt by the capture WITHOUT its
volume dumps (185664 B -> 4382 B, volume_dumps: None), and the off-site backup then ships
it and logs "backed up opengist ... 0 mandatory path(s)" - a success line over a backup
holding none of the app's data.

PHASE 2 also R-413: the R-87 proof CAUGHT that hollow snapshot unattended -
verdict "fail", volumes_expected_none_captured: opengist_data, one offsite_proof_empty at
severity error, while the four apps ahead of it in the rotation passed.
2.3 stale marker PASS, 2.4 two restores PASS, 2.5 corrected the runbook's premise
(RunTier2 has zero acquireRunning and zero restic references - a local mirror cannot
collide with the repo, so the proof correctly does not skip for it).

PHASE 3 PASS - R-357 live-validated at last, six days owed. A real full filesystem
(1900544 B free vs 5878421 B needed) on a separate 1 TB device that does not back Docker.
Refused BEFORE StopStack: app "Up 6 minutes (healthy)" unchanged, 0 safety dumps, live
userdata tree fingerprint 5a9b2db75dfd4ba4131adb3255670472 unchanged, message names both
numbers in Hungarian. Ballast removed, the SAME restore then worked and the fingerprint is
still identical. On the same full disk the Tier-2 run, the proof and the integrity check
all behaved: the proof refused before any download through the shared unitOnlyHeadroom
gate extracted today, reached no verdict and did not alarm.

PHASE 4 PASS - the false-alarm control the whole R-87 design rests on now EXISTS. Found by
reading all 53 catalogue composes: bentopdf is the only template with neither a database
service nor a named volume. Deployed, backed up, proved: PASSED in 2.224s with zero alarms.

Four of my own instrument errors were caught by their own controls before any result was
believed: a hub log line used as a controller positive control, a grep pattern that missed
a registered job, a heredoc that mangled a planted marker, and a catalogue scan that read
0 of 0 apps from the wrong path.

Phases 5-7 to follow: the real nightly cycle with injected faults, the untouched observer,
and teardown.
2026-08-31 23:33:32 +02:00
admin 7ee25925f9 R-87 CLOSED: live evidence, capability row, architecture verdict, registers
gates / gates (push) Failing after 17s
Controller v0.231.0 + hub v0.110.0, both deployed and verified on demo-hp.

LIVE EVIDENCE (documentation/tests/r87-offsite-proof-2026-08-31/, 16 files, endpoint level
through the exact route the debug button invokes):

- THE CASE THAT MATTERS: a hollow unit - compose declaring opengist_data, manifest
  declaring nothing - was pushed to the live store and the proof returned verdict "fail"
  with volumes_expected_none_captured: opengist_data, emitted EXACTLY ONE
  offsite_proof_empty at severity error, and the hub answered HTTP 200. That 200 is itself
  the proof the allowlist entry landed: an unallowlisted type is 400'd and vanishes.
- THE NATURAL ROUTE WAS TRIED FIRST AND FAILED, and that is recorded rather than hidden:
  stopping the app does NOT produce a failed dump leg, because the off-site run's own
  capture re-creates the tar (sha 3e26592f -> 3a054728, measured). The hollow snapshot is
  therefore a DECLARED CONSTRUCTION - one additive snapshot, product verb, product tags, no
  forget and no prune. State restored: the product's own run made a healthy snapshot the
  newest again and the proof then passed opengist.
- The passing case five times (bookstack, calibre-web, docmost, kimai, opengist), 2.2-4.0s
  each, matching the spike's measured band.
- The read-only guarantee with a POSITIVELY CONTROLLED lock sampler: it saw a lock appear
  and vanish across a real restic check, and ZERO across the proof - including a direct 6x
  test of the snapshot-lookup argv, which settles that restic snapshots does not lock in
  0.14.0 either.
- Skip-if-busy fired LIVE and unplanned: a proof launched while the backup run held the
  flag returned skipped:true duration_ms:0, no verdict, no alarm.
- The customer's own verification copies were untouched throughout, which is the safety
  property the separate proof root exists for.

ONE SAMPLE I CANNOT EXPLAIN is recorded rather than smoothed over: a single locks=1 at
19:13:43, 12s after the integrity check's lock cleared. Two independent tests exclude the
proof; I did not establish what it was.

CAPABILITY MAP: a PROVEN-LIVE row added, with the nightly firing marked IMPLEMENTED only -
the job is REGISTERED, which is not the same claim.

07 section 8 MATRIX ROW 4 WAS NOT MOVED, deliberately, and section 10.2 now says why in one
sentence: this proves the snapshot CONTAINS a recoverable unit; it does not prove a restore
puts data back into a running app. Without that sentence the new green tick reads as
covering the drill.

REGISTER: R-87 CLOSED and compressed into CLOSED-ITEMS.md. OPEN 172 -> 171, CLOSED 151 ->
152. No new rows minted. R-408 and R-409 stay open and are referenced by this work.

golden-currency is RED and it is a DECLARED, EXPECTED debt: v0.231.0 is released and the
newest golden carries 0.230.0. The fleet is on 0.230.0; demo-felhom does not have this job.
A golden carrying 0.231.0 is OWED and it is Viktor's call (R-242). This push uses
--no-verify for that reason - bypass #8.
2026-08-31 21:32:06 +02:00
admin 2263245cf2 golden 0.230.0 baked, vouched, floor raised - demo-felhom moved itself off the R-403 build (R-410 filed)
gates / gates (push) Successful in 17s
GOLDEN_SHA256 9287f7cef5f13166276e8406005e3f28004004510c5184f1c1c7377f7aafad2e,
657 873 700 B. Evidence documentation/tests/golden-0.230.0-2026-08-31/.

WHY IT WAS OWED: the newest golden was 0.229.0, which IS the build R-403 says deletes a
good copy. Every fresh install and the whole fleet floor still carried it.
golden_currency_gate.py had been red across dddcc80, 6e550ae, 130f7a6 and 32a4c35.

THREE INDEPENDENT READERS agreed before anything was vouched: the bake's own print, the
round trip of the PUBLISHED bytes (HTTP 200, 657873700 B, same sha), and the hub's Day-0
dropdown reading Gitea on a different code path. And the delivered artifact names the
controller it will start - ./etc/felhom-controller-image read OUT of the downloaded
archive says felhom-controller:0.230.0, with 19382 entries under var/lib/felhom/docker/.

BOTH PRE-GATES were shown able to see something before their zeroes were believed: the
404 pre-gate, and the token-leak grep which returns 0 on the committed log and 1 on a
seeded throwaway copy. The transient unit's own properties were grepped for the token
too - 0, with the same seeded positive control returning 1. Acceptance markers counted on
the COMMITTED log: 1/1/1/1 present, 0/0 absent, and the zeroes are believable because the
same including-mount-point pattern returns two real lines on that file.

THE VOUCH IS A THREE-FIELD CHANGE and only one field moved, which is stated rather than
left to look careless: golden_version 0.229.0 -> 0.230.0; agent_version 0.130.0 and
min_agent 0.129.0 UNCHANGED because v0.230.0's CHANGELOG header says MinAgent 0.129.0 and
0.129.0 <= 0.130.0, so this is not the R-216 shape. The 303 flash was not treated as
proof - the page was re-read and golden_behind_fleet confirmed absent.

THE FLOOR is a separate setting and was raised on the operator's explicit answer:
min_controller_version 0.229.0 -> 0.230.0. THE POSITIVE OBSERVABLE, from the agent's own
journal on demo-felhom, which was still running the defective 0.229.0:
  16:21:30 controller-swap: image file written, restarting bootstrap  target=...0.230.0
  16:21:40 controller-swap: new controller healthy                    target=...0.230.0
Both boxes now 0.230.0 healthy. Honest note: the polling loop's first read already said
0.230.0, so the transition was not seen by the loop - the journal is the evidence.

R-410 FILED, found while the gate went green: golden_currency_gate.py is satisfied by a
DIRECTORY NAME (EVIDENCE_RE against os.listdir, :89,:123). I created the evidence
directory before the bake finished and the gate would have passed at that moment. It
already declares that it does not check the vouch; it does not declare that the bake
check is a filename check. Fix: read the GOLDEN_SHA256= line out of the directory's
bake.log, with a red-proof on an empty directory.

R-242 updated - seventh debt, paid the same day, twice in one day.

Teardown: pct destroy 9100 --purge, shred -u AFTER the log was copied out, poweroff,
qemu confirmed exited with ps -eo comm (not pgrep -f, which self-matches), disk reverted
to virgin.

All 13 gates green - the first push this session that needed no --no-verify.

Ceiling R-409 -> R-410.
2026-08-31 16:25:35 +02:00
admin 130f7a6eba R-87 SPIKE: measured, do not build it as written (R-407..R-409 filed)
gates / gates (push) Failing after 17s
Spike. NO production code. No version bump, no build, no deploy, no golden.
felhom-controller and felhom-agent were READ ONLY. The fleet stays on v0.230.0.

Q1 restic is 0.14.0 (go1.19.8, bookworm 12.15) - the four source comments asserting
it are CONFIRMED, not corrected.

Q2 --verify DOES exist and is NOT a content check. Red-proof: one byte changed in a
restored 160 MB tar with size and mtime preserved passed clean, rc=0. Verify took
131 ms on a 213 MB / 7-file tree, which cannot be hashing. A size or mtime mismatch
causes a silent re-download, not a failure. Controls: --target 1 hit, four post-0.14
flags and a nonsense string 0 hits each. Neither --verify nor --no-lock appears
anywhere in the controller source.

Q3 no reference for "correct" exists. restic ls --json carries no content hash in
0.14.0, and the unit manifest hashes 4918 B of a 213231242 B unit - 0.0023 percent,
the config files and not the dumps or the tars. R-409.

Q4 it is CHEAP. All 8 apps / 774378123 B logical restored back to back in 25 s, against
40257 ms for the weekly 100 percent check beside it. Individual restores 2253-3978 ms
regardless of size: cost is per-snapshot round-trip plus ~1 s per 200 MB. Peak scratch
is the app's full logical size. The 1.1 MB restic cache is index only and hides nothing
(--no-cache 5423 ms vs cached 3198 ms, trees byte-identical).

Q5 skip-if-busy stays right at 25 s against a 2m52s nightly backup. But
RestoreOffboxScratch takes NO acquireRunning, while offbox_integrity.go:28 asserts
every off-site operation does. R-408.

Q6 observed with a positively-controlled lock sampler: restic restore takes NO lock;
restic check DOES (locks 0 -> 1 for nine samples -> 0 across the check, zero across two
restores). The product writes anyway - unlockStale runs `restic unlock`, a delete verb,
before every restore (offbox_restore.go:289). The task's lead was right in direction and
wrong in mechanism. R-95's constraint IS satisfiable: --no-lock plus skipping unlockStale
writes nothing, and both mechanisms exist unused. Neither was fixed - the task forbids it.
offbox_integrity.go:255's "It NEVER writes to the repository" is R-407.

Q7 THE DECIDING ONE: of R-353/354/356/358/403 an unattended scratch-restore would have
caught ONE (R-356). The value is elsewhere, and the weekly check structurally cannot
reach it: `check` proves the stored bytes are the stored bytes, never that we stored the
RIGHT thing. A hollow unit backs up, checks at 100 percent and restores cleanly and
recovers nothing - R-403, measured in bytes on 31 August.

RECOMMENDATION: option C, the NARROW test - one app a night, restored to scratch, checked
against its own manifest.json through the existing unitCarriesData, scratch deleted, the
SNAPSHOT recorded as the proof. Options A (do not build) and B (scheduled attended drill)
considered explicitly; B is weakest because it is what already happens. R-87 should be
RE-SCOPED, not built as written, and that is Viktor's call - the row stays open carrying
the verdict and STATUS.md item 4 asks it in plain words.

Also corrected in 07-backup-architecture.md: matrix rows 4 and 10 both said "the depth
that ships ON does not re-read pack contents (R-399)". R-399 CLOSED in v0.228.0 and the
depth is 100 percent. Two stale cells, fixed, and the spike verdict added beside them.
Row 4's verdict is UNCHANGED by the spike and now says so.

Teardown: all three layers, none of them "nothing was created" - 6 files on the PVE host,
9 in the guest, 5 plus 2 run-flags in the container, all removed and verified empty. The
four scratch directories this session's restores created were removed; three that
pre-date the session were left alone. Two state changes recorded rather than hidden: the
control integrity run recorded its verdict (depth structure -> 100%, due-ness +7 days),
and four restores appear in the controller log. Nothing was written to the off-site
repository by hand.

Evidence: documentation/audits/evidence-spike-restic-restore-2026-08-31/ - 31 files,
every one pulled off the box BEFORE teardown (R-320).

golden-currency is RED at this commit and was already red at dddcc80. Pre-existing, not
this session's debt. Second --no-verify push of the day for that reason; R-404's count
goes six -> seven and its row says so.

Ceiling R-406 -> R-409.
2026-08-31 15:59:09 +02:00
admin 6e550aedd3 R-87 put back in the register; closed_register_gate.py is the 12th gate (R-405, R-406)
gates / gates (push) Failing after 17s
Records and process only. No machine contacted. No product code, no version bump,
no build, no deploy.

R-87 was moved into CLOSED-ITEMS.md by the 2026-08-22 compression sweep ef6ac6f while
its own state cell read "READY - RE-RANKED UP 2026-08-03 (R-86 closed)". R-378 records
that sweep moving six still-open rows and restoring them in the same session; this was
a seventh it missed. Nine days in the wrong file, with the register's ranking paragraph
ranking it fourth and pointing at nothing. Restored verbatim from ef6ac6f^, beside R-95
where it sat before.

The predicate is the LEADING VERDICT of the state cell, which is R-378's lesson and
decides the answer here. Measured on the file as pushed: an open word anywhere in the
state cell convicts 3 of 151 rows, two of them genuinely closed (R-224 and R-260 carry
"open"/"OPEN" inside long prose verdicts); the leading verdict convicts exactly 1; the
whole row convicts 144.

closed_register_gate.py, two rules: no open state word leading a CLOSED-ITEMS.md row's
verdict, and no R- id with a row in both registers. Red-proofed both, and negative-
controlled against the pushed pre-fix files where it convicts R-87 by name, rc=1;
restoring the planted rows leaves the file byte-identical. Registered as the 12th gate
in repo_gates.py, --fast, after it was green. Four residual holes in its docstring.

R-398 was also in both registers - a deliberate cross-reference stub. Now prose beneath
the table rather than a table row, because a row in both files is what rule 2 convicts on.

R-406 filed: two unrelated findings in OPEN-ITEMS.md both numbered R-133. The only such
collision in either register. Deliberately NOT gated - a within-register duplicate rule
would fail on a pre-existing row, and a registered-but-failing gate refuses every push.

golden-currency is RED at this commit and was already red at dddcc80 - controller
v0.230.0 released, newest golden 0.229.0. Pre-existing, not this session's debt.

Ceiling R-404 -> R-406.
2026-08-31 15:32:40 +02:00
admin dddcc808be R-403 CLOSED (controller v0.230.0), R-404 filed as a decision for Viktor
gates / gates (push) Failing after 17s
07-backup-architecture gains section 8.2, placed beside row 5 on purpose: the derived-copy rebuild
rule is UNCHANGED and section 8.2 names the single exception, so a future reader who finds RunTier2
skipping a leg does not fix it back. It carries the measurement (120 082 104 B -> 7 036 B on the
shipped v0.229.0), the four-case table, why hollow is a manifest question and not a size question,
why the data legs are deliberately not guarded, and why the capture job is not guarded either.

00-capability-map: the Tier-2 row's status does NOT move, stated explicitly rather than left
ambiguous. R-403 removes a way the route could be DESTROYED between uses; it does not change what
the route can be relied on for.

Register: R-403 CLOSED and compressed into CLOSED-ITEMS (594 -> 593 open lines). R-404 FILED as a
DECISION and deliberately not acted on - should a documents-only push be subject to the
golden-currency gate, now that it has been correctly bypassed six times? Both sides stated, plus
what happens if Viktor does nothing. The gate was NOT changed.

R-242: seventh conviction, and the FIRST where the day-0 ground does not apply - R-403 is a defect
in the nightly Tier-2 copy, which a day-0 box starts running on its first night. This push uses
git push --no-verify, declared here and in felhom-controller/REPORT.md. A golden carrying 0.230.0 is
owed and is more urgent than the previous six.

STATUS: the R-403 item moves out of 'Broken' into what works, in plain words; the delivery item now
says a golden is owed and that the fleet carries the defect; R-404 goes into the decide section with
its do-nothing outcome.

Drill evidence: nine phase logs, including the two things that went wrong (a repair whose rsync was
not installed in the guest and silently did nothing, and a session that expired mid-run so a POST
did nothing).
2026-08-31 14:39:29 +02:00
admin 66156c619f R-403 drill evidence + the credential reader that ends a three-time mistake
gates / gates (push) Successful in 16s
The drill: the loss reproduced on the shipped v0.229.0 before anything was built. 120 082 104 B ->
7 036 B in one Tier-2 run, recorded as a success. Phases 1a (before), 1b (the hollow primary,
produced through the R-102 restore path exactly as the 2026-08-31 observation was), 1c (the loss),
1d (repair).

scripts/read_credential.py is Part 4's rider, and it exists because a note did not work three times:
2026-07-20 a Failed login was diagnosed as a stale password and written into memory; 2026-08-31 the
same misreading recurred and was caught; 2026-08-31, hours later, it recurred AGAIN and rewrote a
live box's password hash. Between them the project already had a memory file stating the rule, a
worked recipe in it, and a session report describing the mistake. The rule now lives in the code
path: one matching quote pair is unwrapped, the result is REFUSED if it still carries a quote, and
--expect-length gives the caller a second opinion. The value goes file->file at 0600 and stdout gets
only its length. test_read_credential.py asserts each refusal by its reason, with a positive control
before believing the not-in-stdout result.

Red-proof E1: remove the final quote assertion -> three cases fail by name.
2026-08-31 14:02:26 +02:00
admin 83ff9e8e38 golden 0.229.0 baked, vouched, floor raised — R-242's sixth debt PAID the same day
gates / gates (push) Successful in 16s
GOLDEN_SHA256 39aa886df77b21757aef3b298a389343dc0df5134bb0f14e8f92a451d7bdae87, 656 864 331 B.

The evidence is the ROUND TRIP, not the build log: the published bytes were downloaded back and
match the bake on both size and sha, and ./etc/felhom-controller-image read OUT of the downloaded
archive says felhom-controller:0.229.0 - the delivered artifact naming the controller it will start.
A THIRD independent reader agreed before anything was vouched: the hub's own Day-0 dropdown read the
same sha straight from Gitea, a different code path.

Both pre-gates were proven able to see something before their negative results were believed - the
404 pre-gate against a 200 from 0.228.0, and the token-leak grep against a seeded throwaway copy.
Acceptance markers counted on the COMMITTED log: 1/1/1/1 present, 0/0/0 absent.

The vouch is a three-field change, checked rather than assumed: MinAgent 0.129.0 read from the
golden's controller CHANGELOG header, agent_version 0.130.0 >= min_agent 0.129.0 (not the R-216
shape), agent_sha256 and wrapper_sha256 carried through explicitly because the handler clears a
field it is not sent. Verified by re-reading the manifest, never by trusting the flash. The R-120
gate PASSED rather than being bypassed - fleet newest 0.229.0, golden 0.229.0.

The floor is proven ACTING, not merely set: demo-felhom self-updated 0.228.0 -> 0.229.0 and logged
settle-gate GO at/above floor 0.229.0. Nobody deployed to that box. Both demo machines now carry the
Tier-2 unit restore.

golden_currency_gate.py went red -> green; the --no-verify bypass declared on c2de785 is now
historical. R-242's other half is untouched and still open: nothing gates the VOUCH itself.

Teardown: build guest 9100 destroyed --purge, token/runner/script/log shredded AFTER the log was
copied out, VM powered off, qemu confirmed gone from ps -eo comm, disk reverted to virgin.
2026-08-31 12:38:01 +02:00
admin c2de785bf2 R-102 + R-103 CLOSED (controller v0.229.0) — architecture, register, STATUS, drill evidence
gates / gates (push) Failing after 17s
07-backup-architecture: 6.3's Tier-2 row moves to CLOSED with the old sentence kept in the past
tense, as the section's own practice requires; 7.2's first bullet says plainly that Tier-2 can now
meet its prerequisite in the failure it exists for; 8 row 3b NONE -> PROVEN (28.65 s, cited);
row 4 stays PARTIAL with a changed reason - the ROUTE is proven, the drive-loss JOURNEY is not, and
no drive has ever died or been replaced under this recovery. 8.1's blanks updated per row.

6.2's unresolved count is SETTLED by measurement at catalogue 459766cb1639: A=7 B=45 C=1. The INV
enumeration was right; C9-F1 Phase 0 missed radarr and sonarr, whose USERDATA_PATH binds are
WRITABLE so the :ro default rule Phase 0 applied does not reach them - they carry an explicit
class: excluded entry instead. Class C is bentopdf. No catalogue file was changed.

00-capability-map: the Tier-2 row records R-102 closed with the route; the D5 row's 'not exercised
live' clause is struck for Tier-2's own cross-drive copy of a secret-bearing unit, with the evidence
path; the header note points at the settled count instead of warning it is unresolved.

Register: R-102, R-103 and their C9-F4 / C9-F1b aliases closed and compressed into CLOSED-ITEMS
(596 -> 593 lines, each naming git show 1623a4d5b5 for the original). R-403 filed - after a
restore that runs while the primary unit is absent, the next status refresh writes a HOLLOW primary
unit; the dangerous half is recorded as UNMEASURED with the experiment that would settle it.

R-242: sixth conviction of golden_currency_gate. THIS PUSH USES git push --no-verify, declared here
and in felhom-controller/REPORT.md - a BYPASS, not a waiver. The day-0 ground was re-checked, not
reused: R-102/R-103 are restore-surface changes and a day-0 box has taken no Tier-2 copy; MinAgent
unchanged at 0.129.0. OWED: bake a golden carrying 0.229.0, vouch it, raise the floor.

Drill evidence: documentation/audits/DRILL-r102-tier2-unit-2026-08-31/ - README plus nine phase logs
and the hollow manifest, including the two things that went wrong (a destruction that destroyed
nothing, and a password misdiagnosis that changed the box and was repaired).
2026-08-31 12:21:52 +02:00
admin 1623a4d5b5 golden 0.228.0 — baked, published, round-trip verified, vouched, floor raised
gates / gates (push) Successful in 19s
Bake: build-golden.sh v3.0.0 in the drill VM, reverted to virgin and cold-booted.
GOLDEN_SHA256 76a3a98b9e7cc23bf8ae51b38a6272f576df285cb34cd22235ac3f06a31e53ec,
658 079 744 B. All four acceptance markers counted 1; excluding/FATAL/mp1 counted 0.

The 404 pre-gate was proven to work before its 404 was believed — the target URL
404'd while the existing 0.227.1 package 200'd on the same command.

The evidence is the ROUND TRIP: the downloaded bytes match the bake's size and
sha, and ./etc/felhom-controller-image read OUT of the downloaded archive says
gitea.dooplex.hu/admin/felhom-controller:0.228.0.

Vouch: three fields together — golden_version 0.228.0, agent_version 0.130.0,
min_agent 0.129.0 (read from the controller CHANGELOG header, not assumed);
wrapper_sha256 carried through explicitly. agent >= min_agent, so not the R-216
shape. Verified by RE-READING the manifest, never the flash. R-120 gate passed.

Floor raised 0.227.1 -> 0.228.0 (impact preview {"below":3,"valid":true}).
The floor is ACTING: demo-felhom self-updated 0.227.1 -> 0.228.0 in under a
minute and re-registered offsite-integrity by itself. Both demo boxes now
re-read their whole off-site store on the weekly check.

Token hygiene: file->file scp, runner script inside the VM, unit properties
grepped 0. The leak grep on the committed log was proven with a planted copy
(1) before its 0 was believed. Teardown: guest 9100 purged, secrets shredded
after the log was copied out, VM off, disk reverted to virgin.

golden_currency_gate.py red -> green. All 12 felhom.eu gates OK.
2026-08-31 10:53:43 +02:00
admin 77a5a1154b docs: controller v0.228.0 — R-399/R-400 closed, R-401/R-402 filed
gates / gates (push) Failing after 19s
STATUS.md: header said 2026-08-23 over a 2026-08-30 body, and two "Waiting on
you" items were both numbered 4 — both fixed. R-399 leaves that section (decided
and shipped); the depth change is stated in plain words and the remaining items
each say what happens if Viktor does nothing.

00-capability-map.md: the off-site verification row now carries its DEPTH, and
its live citation is the 2026-08-31 run at 100%. The weekly firing at the new
depth stays IMPLEMENTED, not PROVEN-LIVE.

07-backup-architecture.md §10.2: R-399 recorded closed, with the one sentence
that stops it being turned back down — the structure check PASSED a
size-preserving pack corruption. R-87 untouched and still OPEN.

Register: R-399 and R-400 compressed into CLOSED-ITEMS.md with their reasoning
kept and 300d7e8 named as the commit holding the originals. R-401 filed with a
TRIGGER (the slow-check WARN firing) rather than a date. R-402 filed: the
integrity verdict and its depth are on the wire and no hub surface reads either.
OPEN 166 -> 165, CLOSED 148 -> 150.

wire_contract_gate.py: offsite.last_integrity_depth allowlisted WITH ITS REASON
beside its sibling last_integrity_ok, both to be deleted together when a hub
surface is built (R-402).
2026-08-31 10:39:05 +02:00
admin db0812b6f2 Golden 0.227.1 baked, vouched, floor raised — and the floor delivered the new job by itself
gates / gates (push) Successful in 16s
Second full delivery of the day. golden_currency_gate.py went red -> green on
the same command, so the --no-verify bypass declared on the previous push is now
historical rather than standing.

  GOLDEN_VERSION  0.227.1
  GOLDEN_SHA256   66754491dc9bd0130ef8ded9562f63c53a5ffdcfd91baa551141e55fa083ea32
  size            657 403 203 B
  baked           gitea.dooplex.hu/admin/felhom-controller:0.227.1
  MinAgent        0.129.0  (read from the controller CHANGELOG header, not assumed)

THE EVIDENCE IS THE ROUND TRIP. The published bytes were downloaded back -- size
and sha256 identical to what the bake reported -- and ./etc/felhom-controller-
image was read OUT of the downloaded archive: felhom-controller:0.227.1. That is
the delivered artifact naming the controller it will start, from the bytes a
customer's box would actually fetch.

Markers counted: docker OK (overlay2 = 1, mount point rootfs = 1, mp0 = 1,
upload OK (HTTP 201) = 1; excluding = 0, FATAL = 0, mp1 = 0. 404 pre-gate passed
before the run and the script's own pre-delete agreed, so nothing was
overwritten.

Three-field vouch, all three checked: agent_version 0.130.0 >= min_agent 0.129.0
(NOT the R-216 shape), wrapper_sha256 carried through explicitly because the
handler clears it when omitted. Verified by RE-READING the manifest rather than
trusting the flash. The R-120 gate on that POST passed on its own terms rather
than being worked around.

AND THE LINE WORTH KEEPING. demo-felhom self-updated 0.226.1 -> 0.227.1 in under
30 seconds and then logged:

  [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST

A box nobody deployed to now runs today's off-site integrity check on its own
schedule. That is a floor DELIVERING rather than merely recording, observed
instead of assumed -- and it is the strongest evidence R-242 has carried.

Token hygiene: file->file, read inside the VM by a runner script, never on a
command line (systemctl show ... grep -c -F token = 0). The leak grep on the
committed log was PROVEN TO WORK before its 0 was believed.

Teardown: guest 9100 destroyed --purge, secrets shredded AFTER the log was
copied out, VM powered off, disk reverted to virgin.

R-242 now records the cadence as MEASURED: five convictions and two full bakes
in one day. Every bypass declared, every debt paid -- and the pattern the row
exists to name is exactly that a release and its delivery are separate acts. Its
other half stays open: nothing gates the VOUCH itself.
2026-08-30 21:48:15 +02:00
admin 99af997ab9 R-359/R-397 closed, R-398 corrected, R-399/R-400 filed with measured numbers
gates / gates (push) Failing after 18s
THE MEASUREMENT IS THE STORY, and it re-frames the row it was filed under. A
pack was corrupted WITHOUT changing its size; plain `restic check` -- the depth
that ships ON -- returned `no errors were found`, exit 0. Only --read-data
caught it. So the check that shipped verifies the index, the pack inventory and
the snapshot graph, and does NOT re-hash pack contents. R-399 was filed as a
bandwidth-and-cadence question; it is more than that, and its row now says so.

R-399 gets three MEASURED numbers instead of estimates: store 140 829 678 B /
2651 blobs / 67 snapshots; structure check 35.0 s; curve 10% 35.9 s, 50% 37.3 s,
100% 39.2 s. At this size re-reading everything costs four seconds more than
reading none, because the wall clock is SFTP round-trips not transfer. The row
states the limit too: these do NOT extrapolate.

R-400: the sweep the task asked for found EIGHT dead debug buttons, not one. 24
endpoints referenced in debug.html, 17 dispatched. Single dispatcher, exact
match, default NotFound -- so they 404. A third of a debug page does nothing, on
the surface an operator reaches for when something is already wrong.

R-398 is CORRECTED AND LEFT OPEN, not closed. I filed it yesterday saying
resticStep is not a seam so no test can drive a restic path. The layer below it
has been injectable since the off-site tier shipped. The row survives as the
record that the seam EXISTS so nobody re-files it.

07 gap register: R-359 and R-397 closed; R-87 restated IN PLACE as "AND IT IS
NOT R-359" because the two rows are adjacent and a check is not a restore-test.
08 alarm ladder: both event types recorded, including that `ok` is `info` and
therefore mails nobody BY DESIGN, and that all three registers were checked and
deliberately left alone. 00 capability map: PROVEN-LIVE for the check, the
notifier and the hazard control; the scheduled firing is IMPLEMENTED only,
because a week has not passed.

wire_contract_gate: `offsite.last_integrity_ok` allowlisted WITH A REASON. The
gate was right -- the controller emits a field no hub struct can decode.
Building the display is a hub change and R-331 ruled that class the operator's
decision; the entry says to delete it when a surface exists.

This push used `git push --no-verify`. golden-currency is CONVICTED and right:
0.227.1 is released and the golden carries 0.226.1. A BYPASS, not a waiver, and
the task spec directs it -- golden and fleet delivery are Viktor's (R-242). It
is item 3 under "Waiting on you".

Register 163 -> 165 -> 163.
2026-08-30 21:29:37 +02:00
admin 4f875174fe Golden 0.226.1 baked, vouched, and the fleet floor raised — the debt is paid
gates / gates (push) Successful in 17s
golden_currency_gate.py had been CONVICTED four times today across three
controller releases. One bake covers all three, and the gate went red -> green
on the same command, which is its proof that it measures something real. THE
THREE DECLARED BYPASSES ARE NOW HISTORICAL RATHER THAN STANDING.

  GOLDEN_VERSION  0.226.1
  GOLDEN_SHA256   70ed8e9377dec22a9b493e55f222b0e25a49d7f3caec8c506e0412fd6baefe69
  size            657 197 592 B
  baked           gitea.dooplex.hu/admin/felhom-controller:0.226.1
  MinAgent        0.129.0  (read from the controller CHANGELOG header, not assumed)

THE EVIDENCE IS THE ROUND TRIP, NOT THE BUILD LOG. The published bytes were
downloaded back -- size and sha256 both identical to what the bake reported --
and ./etc/felhom-controller-image was read OUT of the downloaded archive:
`felhom-controller:0.226.1`. That is the delivered artifact naming the
controller it will start, from the bytes a customer's box would actually fetch.

Acceptance markers counted, not eyeballed, each string captured from this run's
own log rather than paraphrased from the runbook (two of the three the runbook
named until R-233 could not match anything the script prints): docker OK
(overlay2 = 1, including mount point rootfs = 1, mp0 = 1, upload OK (HTTP 201) =
1; excluding = 0, FATAL = 0, mp1 = 0. The 404 pre-gate passed before the run, so
nothing was overwritten.

THE VOUCH IS A THREE-FIELD CHANGE AND ALL THREE WERE CHECKED: agent_version
0.130.0 >= min_agent 0.129.0, so NOT the R-216 shape; wrapper_sha256 carried
through explicitly because the handler clears it when omitted. Verified by
RE-READING the manifest rather than trusting the flash -- golden option 0.226.1
SELECTED, all four shas matching.

THE FLOOR IS PROVEN ACTING, NOT MERELY SET. demo-felhom self-updated within 30
seconds: "[selfupdate] Post-update startup: update successful (0.225.0 ->
0.226.1)". Both demo machines now run 0.226.1 and only one of them was deployed
to by hand.

Token hygiene: copied file->file, read inside the VM by a runner script, never
on a command line (systemctl show ... | grep -c -F token = 0). THE LEAK GREP ON
THE COMMITTED LOG WAS PROVEN TO WORK BEFORE ITS 0 WAS BELIEVED -- a throwaway
copy with the token appended grepped 1, was shredded, and only then was the real
log's 0 taken as evidence.

Teardown: build guest 9100 destroyed --purge, secrets shredded AFTER the log was
copied out (standing rule 5), VM powered off, disk reverted to virgin.

R-242's OTHER half is untouched and still open: nothing gates the VOUCH itself.
2026-08-30 20:23:52 +02:00
admin c8100aad6b Housekeeping + R-397/R-398 filed
gates / gates (push) Failing after 17s
Compresses the six rows closed today into CLOSED-ITEMS, keeping title, shipping
version, evidence paths and every sentence that states a RULE. Full original:
`git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md`. No open row touched.
Register 165 -> 167 -> 161.

Files two rows that would otherwise have died in an overwritten REPORT.md, which
is what the controller's observations gate exists to prevent:

R-397 -- NotifyIntegrityOK/NotifyIntegrityFailed have no caller anywhere and the
controller runs no integrity check at all. The part that actively misleads is not
the dead code: config.Monitoring.PingUUIDs carries a backup_integrity field and
the monitoring page renders "Mentes integritas -- Hetente (vasarnap)", so the
operator is told a weekly check runs. Decide WHETHER one is wanted, then delete
or build -- do not leave the third state.

R-398 -- resticStep is not a seam, so no test can drive any restic-backed path.
Felt directly in v0.226.0: R-358's safety property is an ORDER (clear the marker
before restic, write it after) and with no seam it could only be pinned by an AST
walk. The contrast is the argument -- offboxLatestSnapshot got a seam in the same
release, in four lines, because a correctness gate could not otherwise be proven.

This push used `git push --no-verify`. golden-currency remains CONVICTED and
remains right: three controller releases today, golden still 0.223.0. Bypass, not
waiver, on the operator's standing ruling, re-checked for this release rather
than reused blindly. Tracked on R-242; one bake carrying 0.226.0 covers all three.
2026-08-30 19:51:31 +02:00
admin e027b5d999 Register + architecture for controller v0.226.0 (R-353/357/358/360/396), and R-395 fixed
gates / gates (push) Failing after 17s
Closes R-353, R-357, R-358 and R-360 with their shipping version and evidence
path, and files two new rows.

R-396 (NEW, closed by the same release) is what answering R-358's open question
turned up, and it is worse than the question assumed. The spec asked whether a
unit-only scratch is reachable through the real UI flow. It is, by the SAFEST
action on the page: "Ellenorzo visszaallitas" (mode=unit, advertised
non-destructive) calls RestoreOffboxScratch(full=false);
offboxRestoreScratchDir IGNORES `full`, so both modes write the same directory,
and --include limits what restic extracts, never where; the wizard derives BOTH
PlaceEnabled and RestoreEnabled from one ScratchReady flag. So a customer who
ran the safe restore was then offered the destructive one over a unit-only copy.
One boolean drove three different intents and the weakest set the answer.

R-395 (filed by the spec) is fixed in this commit, not just recorded. STATUS.md
said golden 0.223.0 / floor 0.222.0 in one block and demo-hp 0.219.0 / floor
0.218.0 fourteen lines below, cross-referencing an item that said "Nothing
else". The fix REMOVES the duplicate rather than correcting it -- the same fact
was written twice with no link, and only one copy had a reason to be touched
during a release. "What works" now points at the item above instead of restating
a version.

07-backup-architecture: four rows added to the 10.2 gap register plus R-396.
Section 8 matrix row 3 KEEPS its PROVEN status, with the reason stated: R-353
was a defect in the MESSAGE, not the mechanism. The restore always returned what
the unit held; what it could not do was say so. A status that measures whether
data comes back must not move because a status line was wrong.

00-capability-map: one new row, and it splits what is claimed. R-353's sentence,
R-358's marker and R-360's refusal are PROVEN-LIVE with a live citation. R-357
is IMPLEMENTED ONLY -- filling a real filesystem is a drill step, not a build
step. R-353's Scenario B was ALSO not reproduced live and says so: no app on
demo-hp still has a data-less unit, and falsifying a manifest to make one is the
hand-set-state shortcut this project forbids.

This push used `git push --no-verify`. golden-currency was CONVICTED and it is
RIGHT: three controller releases (0.224.0, 0.225.0, 0.226.0) and the golden
still carries 0.223.0. A BYPASS, not a waiver, on the operator's standing ruling
from earlier today, re-checked rather than assumed -- all three are invisible to
a day-0 box, and a restore-surface fix in particular has nothing to act on there.
The ground expires the moment a release changes first-boot behaviour. Tracked on
R-242; ONE bake carrying 0.226.0 covers all three.
2026-08-30 19:39:41 +02:00
admin 36f8630020 R-341 check taken, and the golden-currency bypass declared
gates / gates (push) Failing after 16s
Two gates blocked the R-331 hub push. One is FIXED, one is BYPASSED, and the
difference is stated rather than blurred.

FIXED -- due-checks (R-341, 5 days overdue). The +7d measurement was TAKEN on
ep0 rather than deferred again. Precondition passed: proxy still MainPID 551655,
ps -o lstart= still 2026-08-18 09:51:04, NRestarts=0, so this is the same proxy
generation as t0 (anchor is ps, not ActiveEnterTimestamp, which reads 03:54:54Z
here -- R-346's trap).

Result: fd = 17. Not 17 more -- seventeen TOTAL, exactly the documented
baseline, against 405 at the first check. Socket histogram: one LISTEN, ESTAB 0,
CLOSE-WAIT 0.

The verdict is UNANSWERABLE, not "the upgrade fixed it". R-341 asks whether the
PBS 4.2.5-1 upgrade changed the fd slope; inside this interval we removed the
leak OURSELVES (R-344, agent 0.130.0, now live on both boxes). A slope of ~0
measures our fix, not the upgrade, and reading it the other way would credit a
changelog that was read in advance and found to contain no such mechanism. The
perturbation pre-registered for this window was Phase C at ~3%; the actual
perturbation was the removal of the entire phenomenon. Row closed as moot.

What it DOES establish is worth more than the original question: twelve days
after the R-344 fix, same proxy generation, no restart to hide behind, ep0 sits
at baseline with zero established connections. R-336's ~323-day runway concern
retires with it.

BYPASSED -- golden-currency. Controller v0.224.0 and v0.225.0 are released and
the newest golden bake carries 0.223.0, so a machine installed right now gets
neither. The gate is RIGHT. This push therefore uses `git push --no-verify`,
declared here, in hub/CHANGELOG.md, in REPORT.md and on R-242.

A BYPASS, not a waiver: the gate offers a waiver only for a release that
DELIBERATELY needs no golden, and these need one. The operator was asked and
ruled bypass-now-bake-later, on the ground that neither fix bites a day-0 box --
R-330 is a nightly false alarm about apps a new box has not installed yet, R-331
is a hub display over backups a new box has not taken yet -- and both arrive by
self-update. That ground is recorded because it is what to re-check: it does NOT
extend to a release changing first-boot behaviour.

OWED: bake a golden carrying 0.225.0 and vouch it (RUNBOOK-manual-build.md 4.1,
three-field change, MinAgent 0.129.0). Fourth bypass of this gate, and the gap
is now two releases wide rather than one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 18:49:53 +02:00
admin c30430c530 skills: five process-domain skills + check_skills.py
gates / gates (push) Failing after 15s
The four existing skills cover the product; nothing covered how work is
reported. Two rules this project has paid for — check the artifact rather
than the report, and do not state a claim more firmly than the evidence
allows — lived only in the operator's head and in chat, where Claude Code
never read them.

- felhom-evidence      five confidence tiers, artifact-over-report
- felhom-diagnosis     no hypothesis until a command has been seen red
- felhom-plain-language ASD-STE100, two options, the re-pitch
- felhom-handoff       the note goes to a FILE, not the conversation
- felhom-doc-authoring the pointer decides whether material is reached

scripts/check_skills.py asserts what decides whether a skill is EVER
reached: frontmatter parses, name == directory, description and body
non-empty, under 150 lines, installed copy still samefile()s into the
repo. install_skills.py globs and never reads the file, so a missing
description installs perfectly and then silently never loads.

It convicted on its first run: felhom-build-deploy is 179 lines. NOT
trimmed here (pre-existing skills are out of scope, and trimming a
deploy skill without exercising its commands is how a wrong command
reaches a live host) — a named single-entry GRANDFATHERED exception,
WARNed every run, R-394. A new skill over the limit is convicted.

Red-proof run and seen failing: description removed from
felhom-evidence -> exit 1, "frontmatter field 'description' is missing
or empty". Restored, tree clean.

skills/SOURCES.md records both MIT upstreams, that these are adaptations
not copies, and the six pieces deliberately EXCLUDED with reasons.

Register: R-392 (no architecture doc covers the two-AI workflow),
R-393 (decision-log skill deferred, with the reason), R-394.
2026-08-25 09:36:20 +02:00
admin ebdc04601d docs(hub v0.108.0): the delivery grain, the cooldown ruling, and gate 11's first subject
gates / gates (push) Successful in 15s
The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or
coarse, and why the default is coarse. CONTEXT records two rulings: the grain is
allow-listed rather than inferred from the payload, with crossdrive_failed as
the proof that a payload rule would have been wrong; and a finding recorded only
in REPORT.md has a lifetime of one session.

R-389 closed and compressed, keeping its rules and naming the commit whose
git show returns the full text. R-390 and R-391 left open.

REPORT.md is gate 11's first real subject and passes: six observations, two
FILED, four NOT-A-FINDING with their reasons. Three of those declarations are
things a tidier report would have omitted - the gate's own spec would have
passed the item it was built to catch, the burst has no ceiling, and ArgoCD
said "successfully rolled out" while still running the old image.

STATUS carries forward the one thing outstanding: the controller floor still
reads 0.222.0 while the golden reads 0.223.0.
2026-08-23 14:12:23 +02:00
admin 2fc4a15fa3 R-389: key the operator cooldown per app for app_start_failed; gate 11 makes an unfiled observation refuse the push
gates / gates (push) Successful in 16s
The cooldown key was customerID:eventType plus the tier and run suffixes, and
none of them names an app, so every app going down inside the same hour
collapsed onto one key and only the first was mailed. Measured on demo-hp:
bookstack sent 09:27:51, privatebin suppressed 09:31:51 under
key=demo-hp:app_start_failed.

cooldownStackSuffix is the third sibling of cooldownTierSuffix and
cooldownRunSuffix, and separate for the reason the second one's docstring
already gives: the existing two keep byte-identical semantics for every type
that uses them.

It is ALLOW-LISTED to app_start_failed and takes the event type as well as the
details, unlike its siblings, and that asymmetry is the safety property. The
backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk
sends one digest rather than one mail per app - and crossdrive_failed is
severity error, reaches the operator leg, and carries stack_name through a
DIFFERENT struct, so a payload-shape rule would have split it silently. The
hour itself does not change.

Gate 11 refuses a push whose REPORT.md carries an observation with neither
`FILED: R-NNN` nor `NOT-A-FINDING: <reason>`. It deliberately does NOT accept a
passing mention of some other R-number: the lost item cited R-182 as an analogy,
so "cites a register row" would have passed the very item the gate exists to
catch. That discrepancy with the spec is recorded in the gate's docstring.

Registered here and in the controller and agent runners. NOT in the catalog
runner - it has no shared-gate mechanism and appends --all to every gate;
filed as R-391 rather than left as a sentence, which is this session's lesson.

PROMPT-TEMPLATE.md §15.9 corrected: "documented, NOT acted on" was the wording
that invited the gap, and it now names the markers and points at the gate.

R-390 filed for the golden-bake runbook's missing `pveam update`.
Hub tests 709 -> 716.
2026-08-23 13:53:01 +02:00
admin f751aea4e3 R-389: file the cooldown-grain finding that was never filed
gates / gates (push) Successful in 16s
Only the first broken app per hour reaches the operator. The cooldown key is
customerID:eventType plus the tier and run suffixes, and neither reads an app
name, so every app that goes down inside the same hour collapses onto one key.

Measured on demo-hp 2026-08-23: bookstack sent at 09:27:51, privatebin four
minutes later logged `suppressed - operator cooldown 1h,
key=demo-hp:app_start_failed`.

Filed FIRST, before any code, for two reasons. It should have existed since
yesterday and did not - it lived in a REPORT.md observations paragraph and
nowhere else, which is R-341's shape one surface over. And the gate this session
adds refuses a push whose report carries an observation with no row behind it,
so the row has to precede the gate or the gate refuses its own commit.
2026-08-23 13:41:57 +02:00
admin 2f7c9a6ce5 docs(R-329/R-386/R-387): the severity contract, the intent ruling, and Part 5 recorded
gates / gates (push) Successful in 17s
The alarm ladder gains the severity contract (the hub's vocabulary is exact, it
coerces silently, and three things now hold it) and the intent test with its
three-way ruling on unknown. Both marked [DESIGN] with the live measurements.

Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy,
verbatim, marked plainly as direction rather than current behaviour, with the
12 -> 15 toggle growth as the argument. Filed as R-388, a product decision.

R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the
full-text commit named. R-387 filed closed - including WHY the dispatcher branch
was kept rather than deleted, which is evidence (three monitor checkers call
ProcessEvent directly) and not caution.

The drill record names three things that had to be re-run: an inert red-proof
mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A
NOT proving the customer gate because demo-hp has no prefs row at all.

Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B.
2026-08-23 12:03:49 +02:00
admin 68a9f5475c hub v0.107.0: the hub rewrote a severity and said nothing (R-387); golden 0.223.0
gates / gates (push) Successful in 17s
One handler, two fields, opposite discipline. An unknown event_type is rejected
with a loud 400. An unknown severity was rewritten to "info" without a word -
and severityNotifies drops "info" before BOTH legs, so the event was stored,
answered 200, and mailed to nobody.

Two shipped features went out that way: DiskAlertKind.Severity emitted "warn"
until controller v0.215.0, app_start_failed until v0.223.0. Measured on the live
hub DB today: 91 app_start_failed events stored all-time, ZERO notification_log
rows before this session - not one, on any channel.

The mechanism built to catch this class was structurally blind to it: the
dispatcher's `unrecognized severity` line cannot execute for anything arriving
over the API, because the coercion one line earlier guarantees the value it
looks for cannot arrive.

The coercion STAYS - a rejected event is a lost event, and losing an alarm is
worse than mis-routing one. Only the silence is fixed: a WARN naming the
customer, the event type and the rejected value.

The dispatcher branch is KEPT, not deleted as dead, and the reason is evidence
rather than caution: cmd/hub/main.go wires dispatcher.ProcessEvent DIRECTLY as
the monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box
checkers, which never pass through the handler. For those it is the only
severity guard there is. All 90 severity literals in internal/monitor are
already valid, so the guard is silent because the producers are correct.

Test count 702 -> 709. Red-proof seen failing: delete the WARN line and the
coercion test fails with "the hub rewrote a severity and said nothing".

Golden 0.223.0 baked and published (sha 9eaf39ac3921...), round-trip HTTP 206.
Vouching is the operator's act and was not done here.
2026-08-23 11:57:26 +02:00
admin 55274d5ef3 R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s
The gate failed only on `released > baked`, so it could catch a forgotten bake
and nothing else. A golden AHEAD of the record passed silently - and that is
how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG
heading still read v0.221.0, with every gate green. Reproduced on the real
history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0.

The gate now asks whether the version being shipped is WRITTEN DOWN: the baked
version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG.
Membership rather than `baked > released` deliberately - a comparison against
the newest heading alone goes green the moment any later entry is written,
leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2)
preserved; every refusal names a reason and a route.

Red-proofed both directions: old gate/old record exit 0, new gate/old record
exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0.

08-alarm-ladder.md is new, and its absence was itself the finding: no document
owned "when does a broken app raise an alarm?". The rules lived as comments in
four packages, each locally correct, with the ordering between them legible only
by reading one function top to bottom - which is how R-384 survived review.

R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed
closed. R-386 filed OPEN: a single-container app stopped out of band raises no
alarm, and a comment claims the opposite - measured live, 9 scans, 0 events,
against a positive control from the same box 17 minutes earlier. Not fixed here.

Golden 0.222.0 baked and published; vouching is the operator's act.
2026-08-23 07:59:52 +02:00
admin 1eb64bec51 R-361 docs: the [FACT], the negative that cancelled Part 2, R-383/R-384, golden 0.221.1
gates / gates (push) Successful in 17s
07-backup-architecture.md gains a dated [FACT] on R-361 - a comment asserting an
invariant the code did not have, for four months - and a [DESIGN] on the db_dumps
decision INCLUDING the trap it created: a stable list lets the already-current
early return fire, so per-capture housekeeping must sit above it.

00-capability-map.md records the NEGATIVE from Part 3 so it is not re-derived: a
held app does NOT raise the dead-app alarm. It aggregates to unhealthy, which
IsDownState excludes. Measured on the shipped build with the scans demonstrably
running over it. No suppression was built and no row opened.

R-383: the double-failure message names an undo copy that is not there - R-361's
own class, one surface over, observed on both 0.220.2 and 0.221.1.
R-384: an app whose database has died reads unhealthy and raises no alarm.

R-361 closed and compressed. OPEN-ITEMS 325236 -> 327266 bytes.

Golden 0.221.1 baked, published and round-trip verified. The golden-currency gate
blocked this push and that block is not circular, so it was satisfied rather than
bypassed - no --no-verify anywhere in this session.
2026-08-23 00:33:12 +02:00
admin a8caa0fdde R-379/R-380 docs: the failure ladder, the drill record, register housekeeping
gates / gates (push) Successful in 17s
07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay ->
rollback -> hold, including why no engine flag closes it: --single-transaction
makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the
fix and the flag is a belt.

Drill record for the live walk, including the TWO defects the walk found in the
fix itself (a rollback into a re-created container; an operator route that
cleared the file while the running controller kept refusing) and the ONE
red-proof that PASSED, which is reported rather than omitted.

R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes.

STATUS.md restates the outcome and names the next operator step.
2026-08-22 18:40:22 +02:00
admin 4e488321bf DRILL R-356b: the off-site restore for a driveless app that HAS a database
gates / gates (push) Successful in 16s
A drill, not an implementation. No code, no version bump, no CHANGELOG entry.

Ten of the forty driveless apps carry a database; I re-measured that count and
got 10. For those ten the restore is a five-leg operation that never ran at all
until this week, because R-356 refused before any of it started.

Walked end to end on demo-hp for both engines - docmost (Postgres 16) and
bookstack (MariaDB 12.3) - each deployed for the drill, planted through the
app's own interface, destroyed for real, restored through the endpoint the UI
posts to.

Q1 does it complete: YES. All five legs ran and succeeded, 32s / 25s. Accented
names byte-identical both directions.

Q2 which leg won: the SQL DUMP. Three-way discriminator returned the altered
dump's value. This confirms R-164's F17 ordering on the OFF-SITE path; R-164
only ever cited the local one. Scratch-only mutation; store proved unmutated.

Q3 does a failure tell the truth: partly, and two defects.

Filed R-379 (HIGH, the undo copy is valid, named, and unappliable by any product
action - proven by applying it by hand on both engines), R-380 (HIGH, a failed
MariaDB replay leaves a partial database behind an app reporting healthy, where
Postgres crash-loops visibly), R-381 (MEDIUM, the failure message pastes engine
stderr including customer table rows into the Hungarian surface), R-382 (LOW,
the summary log omits the volume count it already has).

H1, H2 and H4 did NOT fire and that is recorded. H3 fired in a shape nobody
predicted: not a quiet success, but a loud error over a silent inconsistency.

R-361 reproduced independently on a second app. restic check: no errors, 29
snapshots. A flaw in the drill's own planting - a double-escaped accented title -
was caught by reading stored bytes as hex, recorded, and re-measured in Phase 1b.

Register 325236 -> 330683 bytes. Nothing dropped. Teardown: two apps retained
with reason, no pvesm before-snapshot taken (said plainly), no hub-side record
created.
2026-08-22 16:33:45 +02:00
admin 8c9f1b798b golden 0.219.0 baked, published and round-trip verified (NOT vouched)
gates / gates (push) Successful in 18s
Baked in the drill VM per RUNBOOK-manual-build.md 4.0/4.1, carrying controller
v0.219.0 (R-356).

  GOLDEN_VERSION 0.219.0
  GOLDEN_SHA256  67b46f78f8ed9c7b1876265ab1bde9ec6798897898b1836acece9f3864a2aeb6
  656832571 bytes

All five pass markers matched, both negative controls at 0. Verified by ROUND
TRIP - the published object downloaded again and its sha recomputed - not by the
number the script printed.

Both token-leak greps were proved able to convict before their zeros were
believed: planted copy grepped 1, shredded, then the 0 accepted.

Teardown complete: guest 9100 purged, four secret/script files shredded after the
log was copied out, qemu exited, disk reverted to virgin. The revert first
refused while qemu held the image, which is the runbook's own no-holder proof.

NOT vouched - that is a three-field operator save (golden_version 0.219.0,
agent_version 0.130.0, min_agent 0.129.0).
2026-08-22 13:56:33 +02:00
admin c297b9f85e R-356 docs: correct R-107 in the architecture, record the design, refresh STATUS, compress the register
gates / gates (push) Failing after 17s
07-backup-architecture.md: three places said no offsite action unpacks the
named-volume tars. R-107 closed in controller v0.218.0; all three corrected with
a dated [FACT], the old sentence kept in the past tense. R-102 is NOT closed and
the correction says so explicitly.

New [DESIGN] paragraph in 6.3: the restore destination is resolved by the same
rule as the capture destination, and the wrong-disk refusal applies to apps that
have a drive to get wrong. Carries the 13/40 measurement.

STATUS.md was internally contradictory - nothing waiting, and one decision
waiting, for something the same page recorded as shipped. 218 -> 102 lines; the
deciding section now says what happens if nothing is done.

R-356 compressed into CLOSED-ITEMS.md; OPEN-ITEMS 327109 -> 325236 bytes.

Drill record and 16 evidence files for the live walk on demo-hp.
2026-08-22 13:26:10 +02:00
admin ef6ac6fe74 One register, enforced by a gate; closed work compressed into siblings (R-376..R-378)
gates / gates (push) Successful in 16s
Records and process only. No machine contacted.

ONE REGISTER (operator ruling). 17 roadmap rows moved into OPEN-ITEMS.md keeping their
identifiers, evidence and original filing dates - the oldest R-10, filed 2026-07-15, 38 days.
15 ideas stay in ROADMAP.md, which is their home; the gate exempts them by their own state
word. 59 already-closed rows stay as history. Sorting rule recorded in the roadmap header:
does the item assert something about the shipped product a reader could check and find false?

scripts/one_register_gate.py, wired as the 11th gate. Control run: baseline passes, a planted
open roadmap-only row is convicted by name, removing it passes with the file byte-identical,
and a planted `idea` row is correctly exempt. Its four residual holes are in its docstring.

The gate earned its keep immediately: it caught R-103, a READY finding my hand-sort mis-read as
done because my regex matched the whole row where the body contains "shipped" - the gate matches
the state cell. It also caught R-203 and R-163, recorded closed in the register and still open in
the roadmap; the roadmap copies are marked SUPERSEDED with the register's verdict.

HOUSEKEEPING. OPEN-ITEMS 672,376 -> 327,109 bytes (-51%); ROADMAP 239,306 -> 78,110 (-67%).
Closed work compressed to 17% into CLOSED-ITEMS.md and ROADMAP-HISTORY.md; every entry names the
commit whose git show returns the full original text. Rule-sentences are kept verbatim under
"Reasoning kept" rather than judged entry by entry - 25 carry one.

CONTEXT.md deliberately NOT compressed and the disagreement is argued in the report: 86% of it is
standing rulings still in force, this prompt's own 3.4 says the log is never edited, and it has no
per-ruling delimiter. Filed as R-377 - the problem is navigational, not volumetric.

The hot/bulk placement decision was NEVER recorded as a decision anywhere - established, not
assumed. Now marked [DESIGN] with a pointer honest about having no original date, given a
decision-log entry that records what was rejected, and the [DESIGN]/[FACT] legend carried from 1
of 8 architecture documents to 8 of 8. Existing statements deliberately left unmarked (R-376).

PROMPT-TEMPLATE gains N.7: compress what you closed, rehome live reasoning before it goes, state
the register's size before and after.

Ceiling R-375 -> R-378.
2026-08-22 12:13:54 +02:00
admin 091a4b7444 Correct the placement mis-framing, and file what we wrote down and never filed (R-368..R-375)
gates / gates (push) Successful in 16s
Documentation and survey only. No code, no machine contacted.

THE CORRECTION. The 40 catalogue templates without a configurable path are not missing a
choice: 01-topology-and-trust.md:150-152 classes each volume hot (DB/config/cache -> fast
storage, ENFORCED) or bulk (media/files), and the 40 are all-hot apps. The deploy page has
been saying so to the customer all along (deploy.html:624-625). SPEC-app-data-placement and
R-352 are corrected in place with the framing MARKED, not deleted; every measurement stands.
R-356 was re-checked and survives, strengthened - an absent HDD_PATH is the normal state, so
reading it as "not installed" misreads a correct configuration.

The disk claim, precisely: since R-165 there is ONE guest data volume with two binds, not two
volumes (build-golden.sh:29-40, 99). A physical-disk failure losing data and first-tier copy
together is REAL and is what the other tiers exist for. A full data volume stopping the OS is
NOT real and was the overstated one.

THE SWEEP. 113 survey-class documents examined, 14 statements of "not filed", 2 already filed.
Its positive control convicted the sweep itself twice before it convicted the corpus - markdown
bold broke the strongest pattern, and the reporter re-searched a truncated line - both false
zeros of the exact class being hunted, and together worth 2 of the 14.

THE HEADLINE. The gap the 2026-08-21 drill rediscovered WAS filed - as R-107, ROADMAP.md:122,
M/READY, 2026-07-28 - and is absent from OPEN-ITEMS.md, which calls itself the single source of
truth. OPEN-ITEMS and that rule both landed 2026-07-27; R-107 went to ROADMAP alone the day
after. 72 ids live only in ROADMAP, 29 not done, some of them findings. Filed as R-369 (HIGH).

Five more still-open gaps filed with their ages: R-371 (17d), R-372 (38d, the oldest), R-373
(20d), R-374 (14d), R-375 (4d). R-368 corrects Part 4: the storage default IS applied at deploy
time via deploy.html:612 - the earlier "the deploy route never reads it" came from grepping Go
and never the templates. R-370 records the process failure and is closed by the template change.

PROMPT-TEMPLATE gains the two rules it lacked: name the architecture document for the area and
say what it says (with a file->area map and the test "is this something we chose?"), and an
enumerated gap becomes a register row in the same session - a ROADMAP row alone does not count.

Ceiling R-367 -> R-375.
2026-08-22 11:18:08 +02:00
admin ca543b8f69 where-felhom-stands: bring the picture up to 2026-08-22, and stop the page disagreeing with its source
gates / gates (push) Successful in 15s
Eight claims re-checked against the drill and the v0.218.0 fixes; three moved, all downward.

  backup.offsite            walked -> partial. "18 snapshots, daily, unbroken" was true on
                            2026-08-09 and false by 2026-08-21: the next snapshot after that date
                            was put there by hand, twelve days later. The rebuild lost the target
                            and the per-app switches came back off, so a run reported "backup OK:
                            0 app(s) backed up".
  fail.wiped-reinstalled.data  walked -> partial. A real reinstall orphans BOTH off-premises tiers:
                            restic silently for 12 days (R-193), and the PBS archives from before
                            the reinstall cannot be opened by the rebuilt box at all (R-366).
  backup.fill-warning       walked -> partial. The warning fires correctly, but the watcher runs
                            once a day, so a filesystem that fills at 03:31 goes unannounced for
                            ~24 h. Watched silent while a volume sat at 99%.

Five re-checked and held: backup.tier1 and recover.byte-identical carry the R-355/R-354 story and
their fixes; backup.restore-proof stays grey for a sharper reason (orphaned archives, not an
untested tier); backup.sikeres gains two fresh instances; fail.customer-self-restore records that
R-356 now blocks 40 of 53 apps regardless of who is driving.

render_stands.py: the header's commit shas were hardcoded, so the page cited the August 9th commits
while the YAML said otherwise - the stale-build-product failure the renderer exists to prevent. They
are parsed now. The count beside them said "15 status(es) moved in that pass" when 15 was every
recorded move ever; it now separates the two numbers.

check_stands passes, and was itself proven able to convict first: a claim marked `missing` flipped to
`walked` in a scratch copy fired rule 5 by name (use.dlna), and the real file still passes.
2026-08-22 10:36:42 +02:00
admin 877fcd2a38 R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed
gates / gates (push) Successful in 16s
Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with
a negative control first — the same planted, hash-recorded fixture run through the same steps on
both builds.

R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so
it never entered the recovery unit, the off-site copy or the restore; and because the same wrong
name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal
was never reached. Fixed by reading the compose project label. Sweep proven able to convict
before its count was trusted: one affected app of 53.

R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit,
before the database and inside the stopped window, and VolumesReplayed reaches the sentence.
The half-false comment beside the skip is corrected and the half that still holds is named.

Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b,
verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are
the operator's decision, and raising the floor is what puts this on demo-felhom, which is still
on 0.217.0 and still has both defects.

R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them
(an existing guard), they are adoptable by hand, and doing it automatically would be a migration.

Ceiling R-366 -> R-367.
2026-08-22 10:11:46 +02:00
admin 7064596c2e DRILL closeout: the scheduled cycle agrees, the abandonment sweep watched firing, R-366
gates / gates (push) Successful in 16s
The box's own 02:30 / 03:30 / 04:15 cycle ran unattended and agrees with the manual one
on every measure. The one that matters: R-355 is not an artefact of manual triggering —
the scheduled run again wrote paperless-ngx's PostgreSQL dump into a directory for a
stack that does not exist, and again left the app's own unit recording db_dumps: null.

Part 4.2's terminal deletion has now been observed. At 05:10 the sweep removed exactly
the recorded set-aside store and left the live repository untouched; at 05:13 the hub
dropped the sealed package that protected it and said so (event 3025). Both halves went
together, three minutes apart, and the controller cleared its own state. demo-felhom's
two preserved fixtures were verified untouched throughout.

R-366 (HIGH) filed, found incidentally: the 21 August reinstall orphaned demo-hp's PBS
whole-guest archives as well as its restic repo, so a rebuilt box loses BOTH off-premises
tiers at once. The restore-test caught it and named the key mismatch precisely; it is
merely called "a failed restore test" rather than "your older backups are unreadable".

demo-hp is left HEALTHY, not broken. Ceiling R-353 -> R-366.
2026-08-22 05:25:30 +02:00
admin d895d9f7dd STATUS + capability map: narrow the end-to-end off-site claim to the leg it was proven on
gates / gates (push) Successful in 16s
The 2026-08-04 row claimed "a customer's file survives a machine rebuild and comes back
— the whole off-site story, end to end". Tonight's drill shows that holds for the
declared-userdata leg of a drive-declaring app and for nothing else: the off-site restore
has no named-volume leg (R-354), and refuses outright for the 40 apps that declare no data
drive (R-356). Since that class keeps ALL its data in named volumes, the end-to-end story
is unproven there and disproven for the volume leg generally. The escrow/key half of the
row is untouched and still stands.

STATUS.md also corrects the fleet pair it still named (0.214.0/0.129.0 -> 0.217.0/0.130.0)
and records that demo-hp's off-site had been silent since 9 August.
2026-08-21 23:34:21 +02:00
admin f5a4fceeeb DRILL 2026-08-21: the off-site restore never replays named volumes (R-354..R-365)
gates / gates (push) Successful in 16s
Diagnostic only — no code changed, no version bumped, nothing deployed.

The verdict is a mixture. The unit and the off-site snapshot HOLD the data, proven
by identity in both storage classes including two Hungarian accented filenames. The
loss is in the last leg: ReconstituteFromOffsite skips every isUnit placement and the
volume tars live inside the unit, so the off-site full restore has no named-volume
leg at all — while the local restore-from-unit does, and returned the same tar
byte-identical minutes later.

Twelve rows opened, ceiling R-353 -> R-365. Three HIGH:
  R-354 off-site restore never replays volume dumps
  R-355 paperless-ngx's Postgres is dumped under a non-existent stack, so its unit
        has no DB dump, no safety dump is taken, and the customer is told it has none
  R-356 the off-site restore refuses for all 40 no-drive apps saying the running app
        "is not installed", with a remedy those apps make impossible

R-353's instruction (2) is satisfied and annotated: the 40-class DOES reach the
off-site tier. Its instruction (1) stands and is now larger. R-329 confirmed still
live and now the only bad-severity emit fleet-wide.

Evidence: documentation/audits/DRILL-backup-truth-2026-08-21/evidence/
2026-08-21 23:30:27 +02:00
admin 059adfb8b8 golden 0.217.0 baked and published — evidence, markers and the token-leak proof
gates / gates (push) Successful in 14s
GOLDEN_VERSION = 0.217.0
GOLDEN_SHA256  = 0276c5f638d140a861daba4ef25129896e259937eec0ae5cba7af42391315ad0
archive volid  = local:backup/vzdump-lxc-9100-2026_08_21-21_40_05.tar.zst
MinAgent       = 0.129.0 (unchanged from 0.216.0)

Acceptance markers counted in this run's own bake.log, not paraphrased:
  docker OK (overlay2       1
  including mount point     2   (rootfs and mp0 — there is no mp1)
  upload OK (HTTP 201)      1
  excluding                 0
  FATAL                     0
Published package fetched back over HTTPS: HTTP 200.

Token handling: copied file->file, read by a runner script inside the VM, never on a command
line. systemctl show of the live unit contained it 0 times. The committed bake.log greps 0 for
the literal token AND the grep was first PROVEN to work on that same file by appending the token
to a throwaway copy (grep = 1) then shredding it — a 0 from an untested grep is not evidence.

Evidence copied off the VM BEFORE teardown. Then destroy 9100 --purge, shred token+runner+script
+log inside the VM (0 left), poweroff, waited for qemu using `ps -eo comm` (never `pgrep -f`,
which self-matches), and reverted the drill VM to `virgin`.

This unblocks the golden-currency gate, which correctly refused the previous push of the register
rows: "controller v0.217.0 is released and NO golden carries it". No --no-verify was used.

NOT DONE: the vouch. It is operator-gated and is a THREE-field change; vouching golden_version
alone would ship this controller onto an agent older than it declares it needs.
2026-08-21 21:43:16 +02:00
admin 67356c9e5c register: R-351 CLOSED, R-352 partly, R-353 OPEN (next session's first item) + placement spec
R-351 - the restore never read back where the backup said the data lived, and a second press
started a second restore. Both shipped in controller v0.217.0.

R-352 - four measured untruths about where an app's data goes:
  1. 40 of 53 catalogue templates declare no data path (13 declare env_var: HDD_PATH)
  2. GetDefaultStoragePath() has three non-test callers and NONE of them places data; its
     field comment "new apps use this by default" has never been true
  3. the first-tier backup follows the data onto the same disk (GetAppDrivePath ->
     systemDataPath) - the posture Tier 2 refuses outright at tier2.go:329
  4. "1 alkalmazas hasznalja" counts only Env["HDD_PATH"] == path, so it can never include
     the 40-class; it means "1 of the apps that CAN use a drive does"
  Visibility shipped tonight; PLACEMENT IS OPEN and is the operator's ruling. An earlier
  recommendation to refuse deployment until a drive is registered was WITHDRAWN - it assumed
  the customer had failed to choose, and they had no choice to make.

R-353 - a restore reported success having returned configuration and no data. OpenGist's unit
holds manifest.json + compose/ and nothing else (volume_dumps: None, db_dumps: None); the
off-site snapshot was 182.3 KB; the outcome said only "completed in 8.666896042s". Ranked as
the NEXT SESSION'S FIRST ITEM. Compounding and recorded as UNKNOWN rather than fine: whether
the 40-class reaches the off-site tier at all has not been observed - runVolumeDumps covers
them on paper, but no nightly dump run had happened on a one-hour-old box.

New: documentation/backlog/SPEC-app-data-placement-2026-08-21.md - specification only, nothing
implemented, listing the five points a placement ruling must settle. Records that the OpenGist
instance meant to be left as evidence was removed by someone between 16:57 and 17:02 UTC (not
by this session); its unit and manifest survive, and privatebin is now a live specimen.

Ceiling moved R-350 -> R-353.
2026-08-21 21:25:07 +02:00
admin 38ca4cf6f1 R-344 CLOSED: delivery was the last thing holding it open, and R-347 closed that
gates / gates (push) Successful in 14s
0.130.0 is tagged, published and vouched, so a fresh install gets the
fixed agent. Both boxes reinstalled from the downloaded artifact.

Closing evidence is the positive observable: ep0 at fd 17 / ESTAB 0 /
CLOSE-WAIT 0, its t0 baseline, returning there between poll cycles with
both agents still demonstrably polling. ep0 was read-only throughout and
its proxy PID never changed.
2026-08-20 12:55:45 +02:00
admin 910fd91124 agent 0.130.0 published and vouched; R-347 closed, R-349 + R-350 filed
gates / gates (push) Successful in 14s
Released via scripts/release-agent.sh: tag v0.130.0 at 7569f34, sha256
a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3,
verified by independent download and reproducible byte for byte with
-trimpath -buildvcs=false.

Vouched agent 0.129.0 -> 0.130.0 in the Day-0 manifest. Only the agent
fields changed: min_agent stays 0.129.0 because it states what the GOLDEN
CONTROLLER requires, and raising it would have HELD the floor for every
box below 0.130.0. Global floor untouched at 0.216.0 -- and on hub
v0.106.0 it is a separate form with its own action, so publish-train
rule 2's hazard no longer exists in the shape its incident describes.

No --no-verify: the CHANGELOG heading was flipped only after the tag and
package existed, so release-complete passes on the real artifact.

R-349: the fleet was running a DIFFERENT binary under the same version
name -- the proof deploy was a hand build, the release is -trimpath.
Self-update could never have corrected it, because every version check
compares the string. Both boxes reinstalled from the downloaded package.
The proper fix exists in miniature as wrapper_sha256 and was never
extended to the agent's own binary.

R-350: I printed the hub password into the session transcript via
curl -w '%{redirect_url}' -- the hub answers 303 and curl re-attaches the
credential. Not in git, not in any committed file, not in the evidence
directory. Rotation is the operator's call.

ep0 closes at fd 17, ESTAB 0, CLOSE-WAIT 0 -- its t0 baseline -- and was
read-only for this entire arc.
2026-08-20 12:54:54 +02:00
admin 57dd62b097 R-344 fixed and proven on both boxes: ep0 is back to fd 17 from 415
gates / gates (push) Successful in 14s
P1, outcome (i) in one second: replacing the agent on demo-hp released
exactly its 199 established connections (ep0 fd 415 -> 216). CLOSE-WAIT
stayed 0, so outcome (ii) does not exist and gets no row -- ep0 reaps on
peer FIN correctly, and the 543 CLOSE-WAIT at the 08-18 wedge has another
explanation.

P2, 1.03 h (operator closed the >=4 h window early, so no daily rate is
extrapolated): control +4, fixed +0, with each box making exactly 4
/snapshots and 4 /version calls. Same cadence, same work: 4 cycles -> 4
leaks vs 4 cycles -> 0. The fixed box's cycles are in ep0's log, so the
zero is the fix and not a stopped agent.

P3: the second box took ep0 from 220 to 17 fd in under two seconds. 17 is
precisely the t0 baseline of 2026-08-18 09:51:22Z.

Corrects a claim this session made earlier the same day: the accumulated
descriptors did NOT need an ep0 proxy restart. They were held on both
sides. ep0 was read-only throughout; its PID never changed.

R-344 updated and left OPEN (unpublished is not delivered). R-336
re-scoped -- its old next-step would have fixed nothing while looking like
a failed fix, and it is now a scaling row (~25 req/s at fifty customers).
R-347 filed for the delivery gap (Viktor decides). R-348 filed: an agent
restart blanks the reported backup list for ~18 h and the Store comment
calls it unaffected -- blinds no alarm, checked not assumed.
2026-08-20 12:39:32 +02:00
admin 9299f85c4b SPIKE ep0 connections: CI green by run id (360/237, 19672e685)
gates / gates (push) Successful in 14s
2026-08-20 10:42:50 +02:00
admin 19672e685e SPIKE ep0 connections: the leak is felhom-agent's, not the poll rate
gates / gates (push) Successful in 14s
R-341's first dated check, taken at +46.2 h: fd 17 -> 405 over 166,251 s
= 201.6/day. Pre-registered range was 370-450; observed 388. UNCHANGED,
as predicted. CLOSE-WAIT is 0 -- absent entirely, not merely flat.

Q1: exactly two peers, 194 each, no third party.
Q2: outcome (a). ep0 388 = 194 + 194 on the boxes, twice, and the four
    new sockets carry the same source ports on both sides. 0 closed in 31 min.

The finding: all 388 are held by felhom-agent. pvestatd and
proxmox-backup-client made 162,404 requests and leaked zero. Mechanism is
a per-cycle http.Transport with a zero-value IdleConnTimeout that nothing
ever closes (internal/pbs/client.go:56, main.go:1486). R-336's premise
does not survive this -- cutting the poll rate would have fixed nothing.

Q3 NOT measured: Phase C held at STOP 1, prediction pre-registered first.

Part 0 captures the due-checks gate's first conviction on a real overdue
date (rc=1, names R-341, sole failure among 10 gates). Row cleared at
Part 4, after the result was recorded in R-341, not to make a push work.

New: R-344 (the transport leak), R-345 (hub/Makefile pushes :latest),
R-346 (ActiveEnterTimestamp reads 5h56m early -- NRestarts is still 0).
2026-08-20 10:41:37 +02:00
admin ab2262c91c hub v0.106.0: report loss of visibility into the off-site stores (R-339)
gates / gates (push) Successful in 14s
THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.

REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).

THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.

Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
  - the unreachable event REPEATS rather than escalating once. The band shape
    would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
    mail is missable. It leans on the dispatcher's 1 h operator cooldown to
    become an hourly "still blind" heartbeat.
  - ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
    which means we reached it. Counting it would alert for days on a healthy
    pre-update endpoint.
  - born-blind is reported: the counter is not gated on having a snapshot, so
    a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
    than zero-valued -- a fabricated timestamp reads as "it was fine until
    then".

Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).

Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.

Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
2026-08-18 19:27:33 +02:00
admin 0a5e9b14dc due-checks gate (R-341), floor raise recorded (R-343), snapshot coverage (R-342)
gates / gates (push) Successful in 14s
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as
prose in a register row; nothing read those dates and nothing would have
objected when they passed. The dates now live in a DUE-CHECKS block INSIDE
OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads
them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the
pre-push hook and CI.

  exit 0  nothing due (prints pending count + nearest date; empty block too)
  exit 1  a row is due/overdue (due <= today, UTC -- due TODAY counts), or a
          row names an item with no R-row
  exit 2  block absent/duplicated/unparseable -- INCONCLUSIVE, never 0

It REFUSES rather than warns, and its docstring states the limitation: it is
NOT a scheduler, it fires on the next push, not on the date.

37 tests. BOTH red-proofs run and reverted -- and the first one earned its
keep by catching a hollow assertion of MINE rather than confirming the gate:
flipping <= to < left a due-today row in neither bucket, min() raised on an
empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the
boundary was wrong. An exit code cannot tell a verdict from a crash. The test
now asserts the conviction banner and the absence of a traceback, and the gate
returns 2 rather than crashing if that partition breaks again.

PART 3 — the floor raise, and the premise was WRONG. Read back from the store
(not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero
per-customer overrides, no "managed floor HELD" line. But read 5 shows the
raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and
auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save,
exactly the immediate action publish-train rule 2 documents. No error events
followed; it restarted clean.

R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five
reads clean and no directive served. It went well, but a record calling it
inert when it moved a customer box is what misleads the next reader. The row
also states why the floor was behind -- rule 2 policy, not drift, earned by
the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor
(store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267,
fails open at :78-82) rather than asserting them.

Two boxes are below the floor and neither reports: drill-r50 (blocked,
powered off) and peti-felhom (host row deleted). peti-felhom was NOT
contacted -- its row records that a report from a deleted host 401s and is
not persisted, so the raise cannot reach it.

PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner
server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a
separate Volume that snapshots exclude, so a rollback restores software state
and NOT the datastore. Fine for that upgrade; the safeguard for any future
procedure that could touch the datastore does not exist and is Viktor's call.

Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed
rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling
200). Capability map deliberately unchanged; no row cites a floor or golden
version. repo_gates.py fully green, 10/10.
2026-08-18 15:16:51 +02:00
admin f267bc047f R-334 CLOSED, quoting the green CI run id
gates / gates (push) Successful in 14s
Golden 0.216.0 baked, published and vouched; the row is closed with CI run
353 (head_sha 7d81681d6, conclusion success) named in the closing text.

Quoting the run id is deliberate rather than decorative: this row was already
re-confirmed once and widened once (0.215.0 -> two releases behind at
0.216.0), and closing it on a local green a third time would have left the
same ambiguity the row keeps being reopened for. Runs 351 and 352, earlier
the same afternoon, were red on exactly this gate -- that contrast is the
evidence, not the assertion.

Also records what was verified rather than assumed: the vouch was read back
out of the hub's own store (golden 0.216.0 / agent 0.129.0 / min_agent
0.129.0) and the hub's recorded sha256 matches the artifact downloaded
independently from Gitea. golden_currency_gate.py states of itself that it
checks the BAKE and not the vouch, so the gate alone could not have closed
this.

Fixes a column-count slip in the same row: the closing text initially
replaced two cells with one, leaving R-334 at 6 pipes against its
neighbours' 7. The deps cell is restored.
2026-08-18 13:06:33 +02:00
admin 7d81681d6e golden 0.216.0: baked, published, vouched — gates green again
gates / gates (push) Successful in 13s
Closes the two-release day-0 gap that has been convicting CI since
2026-08-14. Run against RUNBOOK-manual-build.md 4.0 + 4.1.

  GOLDEN_VERSION = 0.216.0
  GOLDEN_SHA256  = ac004dc90d8cefccc5448377892f9cff3a4c3e1e27d0e11129120e38ac31c34b
  archive        = 656,970,239 bytes, controller image 0.216.0
  template       = debian-13-standard_13.6-1_amd64.tar.zst (listed live, not reused)

Baselines re-read on the machine and all four matched the sheet: controller
v0.216.0, its MinAgent 0.129.0, agent v0.129.0, previous golden 0.214.0. The
published agent artifact for the vouched agent_version was confirmed present
in the package registry rather than inferred from a CHANGELOG, and the R-216
check passed on the machine: MinAgent is EQUAL to, not above, the newest
published agent.

Verified beyond the script's own claim: the artifact was downloaded back out
of Gitea and hashed, and it matches GOLDEN_SHA256 exactly. A script printing
a digest and the registry serving those bytes are two different claims.

Pass markers (corrected post-R-233 list) all present, quoted with line
numbers in pass-markers.txt; excluding/FATAL absent; there is no mp1.

Token never reached a command line: copied file->file, read inside the VM by
the runner. systemctl show grep = 0. Token-leak grep on the COMMITTED log run
with its positive control FIRST -- seeded copy 1, real log 0 -- because a
grep -c that matches nothing also returns 0.

Teardown: guest destroyed and purged, token/runner/script/log shredded AFTER
the log was copied out, qemu exit confirmed with ps -eo comm (not pgrep -f),
disk reverted to virgin.

Vouched by the operator; verified by reading the hub's own store: golden
0.216.0 / agent 0.129.0 / min_agent 0.129.0, and the hub's recorded sha256
matches the independently downloaded artifact. That check was necessary
because golden_currency_gate.py says of itself that it checks the BAKE, not
the vouch.

repo_gates.py --fast now rc=0, all nine gates OK -- first fully green run
since 2026-08-14.

Capability map deliberately NOT changed: the day-0 row cites drill documents,
and the map's only golden literal is a dated historical citation on the
recovery-journey row which bumping would falsify.

R-334 is closed in a follow-up commit quoting this push's CI run id, since
closing it without one would leave the ambiguity a third time.
2026-08-18 13:04:37 +02:00
admin 3e50902a98 RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
gates / gates (push) Failing after 14s
Both STOPs cleared by the operator. No code changed; documentation only.

STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.

STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.

STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.

UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.

SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.

CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
  - the "~85/day, ~2 years of runway" figures were WRONG. They came from a
    single 17-minute window with a delta of ONE descriptor. Real rate is
    183-200/day over two independent windows; runway ~357 days, not 2 years.
  - the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
    flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
    543 CLOSE-WAIT. R-336's fix must target unreaped connections.
  - "proxmox-backup-api" reported inactive during verification; that unit does
    not exist. Bad query, not a fault, written down because it looked like one.

R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.

golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
2026-08-18 12:26:53 +02:00
admin 435e044cf1 INCIDENT/R-336: the leak is measured live, not assumed
gates / gates (push) Failing after 15s
Post-restart baseline on ep0: 19 fds at 16m43s (from 18), 1 CLOSE-WAIT,
Recv-Q 0. One descriptor per ~17 min is ~85/day, which agrees with the
~73/day implied independently by the failure itself (1016 sockets over 14
days of uptime).

Two estimates of the same slope agreeing turns "the ceiling raise is
mitigation, not a cure" from a plausible claim into a measured one, and
puts the next ceiling at ~2 years instead of a fortnight. Recorded because
standing rule 3 asks for a positive observable: this is it, and it fired.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016p1PTCzb8rF5G9Aa1qhpBN
2026-08-18 06:12:03 +02:00