07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay ->
rollback -> hold, including why no engine flag closes it: --single-transaction
makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the
fix and the flag is a belt.
Drill record for the live walk, including the TWO defects the walk found in the
fix itself (a rollback into a re-created container; an operator route that
cleared the file while the running controller kept refusing) and the ONE
red-proof that PASSED, which is reported rather than omitted.
R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes.
STATUS.md restates the outcome and names the next operator step.
A drill, not an implementation. No code, no version bump, no CHANGELOG entry.
Ten of the forty driveless apps carry a database; I re-measured that count and
got 10. For those ten the restore is a five-leg operation that never ran at all
until this week, because R-356 refused before any of it started.
Walked end to end on demo-hp for both engines - docmost (Postgres 16) and
bookstack (MariaDB 12.3) - each deployed for the drill, planted through the
app's own interface, destroyed for real, restored through the endpoint the UI
posts to.
Q1 does it complete: YES. All five legs ran and succeeded, 32s / 25s. Accented
names byte-identical both directions.
Q2 which leg won: the SQL DUMP. Three-way discriminator returned the altered
dump's value. This confirms R-164's F17 ordering on the OFF-SITE path; R-164
only ever cited the local one. Scratch-only mutation; store proved unmutated.
Q3 does a failure tell the truth: partly, and two defects.
Filed R-379 (HIGH, the undo copy is valid, named, and unappliable by any product
action - proven by applying it by hand on both engines), R-380 (HIGH, a failed
MariaDB replay leaves a partial database behind an app reporting healthy, where
Postgres crash-loops visibly), R-381 (MEDIUM, the failure message pastes engine
stderr including customer table rows into the Hungarian surface), R-382 (LOW,
the summary log omits the volume count it already has).
H1, H2 and H4 did NOT fire and that is recorded. H3 fired in a shape nobody
predicted: not a quiet success, but a loud error over a silent inconsistency.
R-361 reproduced independently on a second app. restic check: no errors, 29
snapshots. A flaw in the drill's own planting - a double-escaped accented title -
was caught by reading stored bytes as hex, recorded, and re-measured in Phase 1b.
Register 325236 -> 330683 bytes. Nothing dropped. Teardown: two apps retained
with reason, no pvesm before-snapshot taken (said plainly), no hub-side record
created.
07-backup-architecture.md: three places said no offsite action unpacks the
named-volume tars. R-107 closed in controller v0.218.0; all three corrected with
a dated [FACT], the old sentence kept in the past tense. R-102 is NOT closed and
the correction says so explicitly.
New [DESIGN] paragraph in 6.3: the restore destination is resolved by the same
rule as the capture destination, and the wrong-disk refusal applies to apps that
have a drive to get wrong. Carries the 13/40 measurement.
STATUS.md was internally contradictory - nothing waiting, and one decision
waiting, for something the same page recorded as shipped. 218 -> 102 lines; the
deciding section now says what happens if nothing is done.
R-356 compressed into CLOSED-ITEMS.md; OPEN-ITEMS 327109 -> 325236 bytes.
Drill record and 16 evidence files for the live walk on demo-hp.
Records and process only. No machine contacted.
ONE REGISTER (operator ruling). 17 roadmap rows moved into OPEN-ITEMS.md keeping their
identifiers, evidence and original filing dates - the oldest R-10, filed 2026-07-15, 38 days.
15 ideas stay in ROADMAP.md, which is their home; the gate exempts them by their own state
word. 59 already-closed rows stay as history. Sorting rule recorded in the roadmap header:
does the item assert something about the shipped product a reader could check and find false?
scripts/one_register_gate.py, wired as the 11th gate. Control run: baseline passes, a planted
open roadmap-only row is convicted by name, removing it passes with the file byte-identical,
and a planted `idea` row is correctly exempt. Its four residual holes are in its docstring.
The gate earned its keep immediately: it caught R-103, a READY finding my hand-sort mis-read as
done because my regex matched the whole row where the body contains "shipped" - the gate matches
the state cell. It also caught R-203 and R-163, recorded closed in the register and still open in
the roadmap; the roadmap copies are marked SUPERSEDED with the register's verdict.
HOUSEKEEPING. OPEN-ITEMS 672,376 -> 327,109 bytes (-51%); ROADMAP 239,306 -> 78,110 (-67%).
Closed work compressed to 17% into CLOSED-ITEMS.md and ROADMAP-HISTORY.md; every entry names the
commit whose git show returns the full original text. Rule-sentences are kept verbatim under
"Reasoning kept" rather than judged entry by entry - 25 carry one.
CONTEXT.md deliberately NOT compressed and the disagreement is argued in the report: 86% of it is
standing rulings still in force, this prompt's own 3.4 says the log is never edited, and it has no
per-ruling delimiter. Filed as R-377 - the problem is navigational, not volumetric.
The hot/bulk placement decision was NEVER recorded as a decision anywhere - established, not
assumed. Now marked [DESIGN] with a pointer honest about having no original date, given a
decision-log entry that records what was rejected, and the [DESIGN]/[FACT] legend carried from 1
of 8 architecture documents to 8 of 8. Existing statements deliberately left unmarked (R-376).
PROMPT-TEMPLATE gains N.7: compress what you closed, rehome live reasoning before it goes, state
the register's size before and after.
Ceiling R-375 -> R-378.
Documentation and survey only. No code, no machine contacted.
THE CORRECTION. The 40 catalogue templates without a configurable path are not missing a
choice: 01-topology-and-trust.md:150-152 classes each volume hot (DB/config/cache -> fast
storage, ENFORCED) or bulk (media/files), and the 40 are all-hot apps. The deploy page has
been saying so to the customer all along (deploy.html:624-625). SPEC-app-data-placement and
R-352 are corrected in place with the framing MARKED, not deleted; every measurement stands.
R-356 was re-checked and survives, strengthened - an absent HDD_PATH is the normal state, so
reading it as "not installed" misreads a correct configuration.
The disk claim, precisely: since R-165 there is ONE guest data volume with two binds, not two
volumes (build-golden.sh:29-40, 99). A physical-disk failure losing data and first-tier copy
together is REAL and is what the other tiers exist for. A full data volume stopping the OS is
NOT real and was the overstated one.
THE SWEEP. 113 survey-class documents examined, 14 statements of "not filed", 2 already filed.
Its positive control convicted the sweep itself twice before it convicted the corpus - markdown
bold broke the strongest pattern, and the reporter re-searched a truncated line - both false
zeros of the exact class being hunted, and together worth 2 of the 14.
THE HEADLINE. The gap the 2026-08-21 drill rediscovered WAS filed - as R-107, ROADMAP.md:122,
M/READY, 2026-07-28 - and is absent from OPEN-ITEMS.md, which calls itself the single source of
truth. OPEN-ITEMS and that rule both landed 2026-07-27; R-107 went to ROADMAP alone the day
after. 72 ids live only in ROADMAP, 29 not done, some of them findings. Filed as R-369 (HIGH).
Five more still-open gaps filed with their ages: R-371 (17d), R-372 (38d, the oldest), R-373
(20d), R-374 (14d), R-375 (4d). R-368 corrects Part 4: the storage default IS applied at deploy
time via deploy.html:612 - the earlier "the deploy route never reads it" came from grepping Go
and never the templates. R-370 records the process failure and is closed by the template change.
PROMPT-TEMPLATE gains the two rules it lacked: name the architecture document for the area and
say what it says (with a file->area map and the test "is this something we chose?"), and an
enumerated gap becomes a register row in the same session - a ROADMAP row alone does not count.
Ceiling R-367 -> R-375.
NOT WALKED moved 32 -> 35 of 55, which the repo checklist asks be stated. Also records the
hardcoded-sha defect found in render_stands.py and the positive control run on check_stands.py.
Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with
a negative control first — the same planted, hash-recorded fixture run through the same steps on
both builds.
R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so
it never entered the recovery unit, the off-site copy or the restore; and because the same wrong
name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal
was never reached. Fixed by reading the compose project label. Sweep proven able to convict
before its count was trusted: one affected app of 53.
R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit,
before the database and inside the stopped window, and VolumesReplayed reaches the sentence.
The half-false comment beside the skip is corrected and the half that still holds is named.
Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b,
verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are
the operator's decision, and raising the floor is what puts this on demo-felhom, which is still
on 0.217.0 and still has both defects.
R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them
(an existing guard), they are adoptable by hand, and doing it automatically would be a migration.
Ceiling R-366 -> R-367.
The box's own 02:30 / 03:30 / 04:15 cycle ran unattended and agrees with the manual one
on every measure. The one that matters: R-355 is not an artefact of manual triggering —
the scheduled run again wrote paperless-ngx's PostgreSQL dump into a directory for a
stack that does not exist, and again left the app's own unit recording db_dumps: null.
Part 4.2's terminal deletion has now been observed. At 05:10 the sweep removed exactly
the recorded set-aside store and left the live repository untouched; at 05:13 the hub
dropped the sealed package that protected it and said so (event 3025). Both halves went
together, three minutes apart, and the controller cleared its own state. demo-felhom's
two preserved fixtures were verified untouched throughout.
R-366 (HIGH) filed, found incidentally: the 21 August reinstall orphaned demo-hp's PBS
whole-guest archives as well as its restic repo, so a rebuilt box loses BOTH off-premises
tiers at once. The restore-test caught it and named the key mismatch precisely; it is
merely called "a failed restore test" rather than "your older backups are unreadable".
demo-hp is left HEALTHY, not broken. Ceiling R-353 -> R-366.
The 2026-08-04 row claimed "a customer's file survives a machine rebuild and comes back
— the whole off-site story, end to end". Tonight's drill shows that holds for the
declared-userdata leg of a drive-declaring app and for nothing else: the off-site restore
has no named-volume leg (R-354), and refuses outright for the 40 apps that declare no data
drive (R-356). Since that class keeps ALL its data in named volumes, the end-to-end story
is unproven there and disproven for the volume leg generally. The escrow/key half of the
row is untouched and still stands.
STATUS.md also corrects the fleet pair it still named (0.214.0/0.129.0 -> 0.217.0/0.130.0)
and records that demo-hp's off-site had been silent since 9 August.
Diagnostic only — no code changed, no version bumped, nothing deployed.
The verdict is a mixture. The unit and the off-site snapshot HOLD the data, proven
by identity in both storage classes including two Hungarian accented filenames. The
loss is in the last leg: ReconstituteFromOffsite skips every isUnit placement and the
volume tars live inside the unit, so the off-site full restore has no named-volume
leg at all — while the local restore-from-unit does, and returned the same tar
byte-identical minutes later.
Twelve rows opened, ceiling R-353 -> R-365. Three HIGH:
R-354 off-site restore never replays volume dumps
R-355 paperless-ngx's Postgres is dumped under a non-existent stack, so its unit
has no DB dump, no safety dump is taken, and the customer is told it has none
R-356 the off-site restore refuses for all 40 no-drive apps saying the running app
"is not installed", with a remedy those apps make impossible
R-353's instruction (2) is satisfied and annotated: the 40-class DOES reach the
off-site tier. Its instruction (1) stands and is now larger. R-329 confirmed still
live and now the only bad-severity emit fleet-wide.
Evidence: documentation/audits/DRILL-backup-truth-2026-08-21/evidence/
Both STOPs cleared by the operator. No code changed; documentation only.
STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.
STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.
STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.
UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.
SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.
CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
- the "~85/day, ~2 years of runway" figures were WRONG. They came from a
single 17-minute window with a delta of ONE descriptor. Real rate is
183-200/day over two independent windows; runway ~357 days, not 2 years.
- the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
543 CLOSE-WAIT. R-336's fix must target unreaped connections.
- "proxmox-backup-api" reported inactive during verification; that unit does
not exist. Bad query, not a fault, written down because it looked like one.
R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.
golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
Checklist item "confirm your own push's CI run by run ID". Run 348,
head_sha ebfd0967c, conclusion failure, elapsed 13 s -- the inherited
golden-currency conviction (R-334), not a new fault: CI's only step is the
same repo_gates.py --fast entry point, and 13 s is the honest-failure band
rather than R-265's reap band.
Stated plainly that the run LOG could not be read (runs/348/logs and
tasks/348/logs both 404 authenticated as admin, runs/348/jobs empty, web
endpoint 302), so naming the gate is an inference from the local run plus
gates.yml -- not CI's own words. Standing rule 2: a "no access" claim names
what was tried.
Also notes the [felhom CI] gates FAILED mail run 348 will send, so it is not
read as a second incident alongside this morning's backup alerts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016p1PTCzb8rF5G9Aa1qhpBN
Two whole_guest_backup_failed alerts at 04:30 and 04:32 CEST were one
incident, and not on either customer box: ep0's proxmox-backup-proxy was
active, holding its listening socket, and accepting nothing.
Root cause: accept() returning EMFILE. The process held exactly 1024 fds
-- its systemd-default soft RLIMIT_NOFILE -- of which 1016 were sockets
and 547 connections sat in CLOSE-WAIT. The 1024-deep accept backlog had
overflowed (Recv-Q 1025), so every client timed out. It was wedged from
its own loopback too, which is what moved this from a network problem to
a process problem.
Fed by ~85k requests/day (a flat 3,538/hour) against an endpoint written
to weekly, that leak reached the ceiling in 14 days of uptime.
Fix: LimitNOFILE=65536 drop-ins for both PBS units, restart, verified from
both boxes (200 in ~0.1s, felhom-pbs active), then re-drove the missed
backups through the product path -- POST /backup?target=felhom-pbs on each
agent's local API, not a hand-run vzdump.
demo-felhom ct/9201/2026-08-18T03:57:43Z 4.10 GB 36.4s
demo-hp ct/9201/2026-08-18T03:58:43Z 4.29 GB 41.5s
Both host reports now carry felhom-pbs success=true, so the hub is green on
the evidence rather than on a restart having been performed. No data lost,
no backup skipped: the daily local tier was never affected and the PBS tier
is weekly, so the window cost exactly one attempt.
Evidence copied off ep0 BEFORE the restart, per standing rule 5.
Filed: R-336 (the ~1 req/s poll rate is the real defect; the raised ceiling
is mitigation, not a cure), R-337 (a status endpoint that trailed its own
artifact by minutes then caught up -- WATCHING, downgraded from the defect
I first wrote, because it self-corrected), R-338 (demo-hp is not on the
R-50 island at all and nodes.md says it is; its local API is bound to the
customer LAN).
R-334 updated: still open, now one version wider (controller 0.216.0 vs
golden 0.214.0). golden-currency is the only failing gate and is inherited
-- it reads files this session did not touch -- so this push used
--no-verify, stated per .claude/rules/gates.md.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016p1PTCzb8rF5G9Aa1qhpBN
Three rules carried forward. A name must separate on the STEM, not the noun — naming
this secret after the act it is used in would have recreated the trap, because the
other factor on the same page is the „Párosító kód". A guard is worth what its positive
control is worth: this one's selftest convicted its own step-3 case and found a defect
in the guard itself. And a suppression must rest on the machine's own declaration, then
be checked for the SECOND door — recording the disabled state rather than deleting it
is what let the deadline check skip it too.
Yesterday's report is preserved to audits/ because it carries the only record of the
self-heal verdict (Part C was dropped, so that reasoning is in no register row) — the
rule written last night, applied to itself the first time it mattered.
A4's evidence contradicts the task's framing and the register now says so: peti-felhom
is a real machine with 482 reports and a real person behind it; david -> tester-1 is a
record that has never had a host, an escrow or a report. The risk is real; only the
word that named it was wrong.
CONTEXT gains three standing rulings: read the fact that carries the RISK (heals_last_hour,
not state) and decode the one that decides the question (heal_succeeded — R-260's lesson);
one secret in two situations keeps its name and changes its sentence, and the mail must
name the page the machine actually shows; and evidence dies in the INTERMEDIATE revert —
with the corollary that a durable citation may never point at a file whose contract is to
be overwritten, which is why the R-316 report was moved to audits/ before this one was
written.
The live test read as a FAILURE for twenty minutes because I stripped only double
quotes from a credentials value wrapped in SINGLE ones, sending a literal ' as part
of the recovery code. Correctly unquoted: the old code returns 422 with
opens_retained=true and the supersession date; a wrong code still returns 400.
The same bug produced the R-308 finding in the previous report. The dashboard
password is fine - HTTP 302 with a session cookie on the first try. Third time this
project has produced a wrong 'the credential is stale' verdict from that one trap.
ListSupersededEscrow had zero production callers for nineteen days. It is the only
reader of a retained identity_blob, so the retention shipped in v0.93.0 was
material the product could not reach - proven on the fixture 2026-08-12, where a
code that opens a retained package was answered as a code that opened nothing.
New GET /api/v1/hosts/<id>/escrow/retained: self-scoped exactly as the current-row
GET, same recovery-mode gate, same audit event written BEFORE the bytes leave,
capped at 16. Rows with a NULL identity_blob are WITHHELD and returned as
unopenable_count - they retain the PBS key, not the repository password, so they
can never open what the caller is asking about, and serving them would let the
screen claim an earlier package is openable on exactly the boxes the original
defect hurt. The count is returned because their existence is load-bearing and
underivable by the caller.
The trade, stated rather than waved through: the hub still cannot read any of it -
sealed bytes in, sealed bytes out, no decrypt path, no recovery code ever held.
What widens is volume, bounded by self-scope, the recovery-mode gate and the cap.
The response is a NAMED TYPE, not a map, so the wire-contract gate can resolve it;
the wire is declared as a fourth ROOT and the gate now checks 182 tags rather than
174. A positive control shows that check is name-presence, not decodability -
filed as R-315 rather than reported as coverage.
Also: golden 0.214.0 baked, published and round-trip verified; the countdown on
demo-felhom cancelled on the operator's ruling (R-307); the spike that halted
Part 3 recorded as R-312; the set-aside store found unrecoverable as R-313.
Six hub tests through the real endpoint; four red-proofs asserted applied.
Three verdicts, kept separate because collapsing them is how this assumption
survived a week.
(a) The material IS retained. host_escrow_superseded id 11 is the first retained
row in fleet history to carry identity_blob (572 B), byte-identical to the
pre-supersession row (sha256 a10032341c8584ed...).
(b) The retained material DOES open the old store. Unsealed with the old recovery
code it yielded a password byte-identical to the pre-change one, and restored
three planted files byte-identical from a store the box itself could no longer
open - including a Hungarian accented filename verified as raw bytes. Negative
control ran first and failed closed.
(c) The customer has NO route, and is misinformed. ListSupersededEscrow has zero
production callers; the recovery path selects FROM host_escrow. Asked with the
code that had just worked by hand, the product answered "the recovery code did
not open the sealed bundle". A valid code for retained history is reported as a
bad code - the R-224 class again. R-304, rank 1.
Both installer faults were watched happening first, so installer-v1.27.0 is now
published (tag + both webpage.yaml refs). Pre-fix: the box came up on controller
0.98.3 against a vouched 0.213.0, below the floor and below the version carrying
the recovery screen; and our own uninstall left dnsmasq on 0.0.0.0:53 so our own
next install refused. R-297 and R-300 CLOSED.
Also filed R-305 (the dnsmasq fix fires once per machine - the leftover returns
on the second reinstall, proven), R-306 (--preflight-only writes state it says it
does not), R-307 (a live abandon countdown on demo-felhom, firing 2026-08-24 -
operator decision), R-308 (stored controller password stale), R-309 (the day-0
runbook's publication claim has been false since R-110), R-310 (two edges).
Ceiling R-303 -> R-310. Capability map moved: the retention claim is now marked
operator-only. Phase A logs did not survive the intermediate revert; recorded.
CENSUS (read-only, hub store, tester's machine not contacted): no machine that is
not ours can be in the state that cost demo-felhom its history. The hub holds
escrow for three hosts; both demo boxes lost their pre-fix key in the same four
hours on 2026-08-04; peti-felhom and david have no host row and no escrow at all.
A control ran FIRST and had to pass -- the query returned "present (572 bytes)"
for a host known to have material and "absent (NULL)" for one known not to.
Corrected my own instrument on the way: a date-only comparison mislabelled both
losses as after the fix, so the in-force moment is now pinned from the hub's first
post-fix escrow row (11:11:37Z), which independently agrees with the register.
PART 1 ESTABLISHED. The prune is recorded inside R-267 -- the row about the
Configuration page being slow -- because pruning artifacts is what made that page
fast. Arithmetic checks (23+7=30, plus three versions that only surfaced after the
first thirty moved them onto page one = 33) and the PAGINATED listing shows both
generics at exactly ten. R-291's blocking condition is released: the operator was
being asked to establish something already written down. And my counter-argument
yesterday was wrong in exactly the way R-267 warns about -- "containers hold 19"
came from an unpaginated query; paginated they hold 270 and 169.
RECEIPTS: three restored (drives.enrol, backup.tier1, fail.lost-recovery-code),
each citing the document that walked it; the map already read PROVEN-LIVE for all
three, so this follows the map rather than raising a status in the view. NINE
HONEST GREYS. fault.selfheal's best hit argues against it -- an incident recording
self-heal's absence through a 1h15m outage.
THE DECAY RULE FIRED FOR THE FIRST TIME. backup.restore-proof has a receipt from
28 July and is superseded anyway: demo-hp's restore-test failed 5 August and the
box has since been rebuilt. A claim about a continuing behaviour cannot rest on an
old observation. The capability map still reads PROVEN-LIVE and is now the thing
out of step -- recorded, not silently rewritten.
PART 4 specified, not implemented. The orphan card promises restorability the box
rendering it cannot evaluate: the discriminator is on the hub and no wire field
carries it. A conditional promise the system cannot evaluate is the same defect as
an unconditional false one, so the copy stops promising, says what happens, and
names a route. Ships with the next controller change so one bake covers both.
R-273's owed guards are both built and closed. R-291 records what CI stopped
covering and why, so it can be widened deliberately rather than discovered.
R-292 is new and was found by a test failing for the wrong reason:
artifact_sha_invalid conflates "version missing", "registry unreachable" and
"bad sha" into one message. v0.102.0 works around it by ORDERING -- the probes
run first, so an unreachable registry is reported as unreachable -- but the
message itself is untouched.
CONTEXT gains the rule this session is about: a check and the policy it enforces
must read the same number from the same place, or they drift and the drift looks
like a defect in something else. Two corollaries, both of which cost something:
a bounded check must print what it stopped covering on every run, and an
unreadable policy is INCONCLUSIVE rather than unbounded.
Stated in the report rather than glossed: Part 4 (finding receipts for the twelve
downgraded claims) was NOT done and is a shortfall, not a decision -- splitting
it would have produced exactly the half-checked green the exercise exists to
prevent. Part 5 was droppable and dropped. The tag-push green is not re-proved
tonight and is not claimed; the evidence offered is runs 190 and 216.
55 claims verified. Twelve moved, all downwards: walked 32 -> 20, built 5 -> 17.
Register ceiling R-284 -> R-290.
THE RULE DID NOT FIRE THE WAY IT WAS EXPECTED TO. Not one downgrade came from
code moving under an old proof. All twelve came from step 1 of the same rule --
the cited evidence does not exist. Measured: of the 28 capability-map rows
behind the page's claims, 8 carry a tests/ or audits/ path and 20 carry prose
only. The green dots were drawn from rows that cite an argument, not a walk
(R-290). The map, not the dataset, is what needs fixing -- it still says
PROVEN-LIVE for all twelve.
And once it ran backwards: fault.operator-email looked contradicted by R-182,
but live source shows the backup_run_failures digest allowlisted, operator-only
and templated, with recovery_unit_capture_failed now record-only. The claim is
right and the REGISTER ROW is stale (R-289). The session went looking for stale
proofs and found a stale defect.
R-281 WITHDRAWN -- wrong in both directions, settled by the operator's mailbox.
The tripwire DID fire (escrow_blob_served 10:19:41Z = 12:19 CEST) and false
error-severity alarms fired too, for deliberate attended work (R-285). The
measurement's cause is ESTABLISHED: the P7 query copied hub.db without hub.db-wal,
and the signature is exact -- it reported "2 events all day, newest 00:30:07",
and the rows at or before 00:30:07 number exactly 2. Timezone and wrong-key were
tested and refuted. The control had been drawn from the same stale snapshot as
the measurement, which is why it agreed (R-286).
Part 4: NO WORKFLOW CHANGED, deliberately. The gate is not ref-sensitive -- it
enumerates from the Gitea tags API, and both previous tag pushes passed. The red
is TRUE: run 267 saw v0.120.0 downloadable, run 284 on the same commit saw 404.
Who deleted the package is NOT established and is not guessed (R-287).
The page is now generated from where-felhom-stands.yaml by scripts/render_stands.py:
static, zero script tags, every moved status carrying a visible "changed, was X"
chip. The React bundle -- whose content was gzip+base64 inside a JS module map --
is kept as a dated snapshot. scripts/check_stands.py gates the data and convicted
51 problems in my own first draft before the staged positive control ever ran.
Operator-present drill on demo-felhom guest 9201. Preflight refused with the key removed, went green
after one 60s tick with the SAME MainPID (1993397 both sides, so no restart), and all 45 config keys
came back identical. Positive control run before the change so the green afterwards is a measurement,
not an artefact of the probe. Marker sha unchanged throughout.
Four defects of one family, all shipped today: something the box already knows, thrown away or drawn
as its opposite. Agent v0.128.0, controller v0.210.0. NO HUB CODE, no hub bump, no ArgoCD sync.
R-265 (this repo). timeout-minutes: 5 on the gates job — every honest run in the observed session
finished in 18-34s, so this is ~9x the slowest and far under whatever reaped run 264 at 834s with no
log. The alarm mail now carries Elapsed (start stamp via $GITHUB_ENV; an absent stamp prints
"unknown (no start stamp)", never a bogus 1.7-billion-second figure) and its "names itself in the run
log" sentence is qualified so it cannot mislead when there is no log.
⚠ THE UNKNOWN IS NOT CLOSED. Whether the if: failure() alarm fires for a REAPED job is still
unverified. The timeout makes the reap unreachable in practice; it does not answer what happens in
one. Demonstrating it means deliberately hanging a run on main, which would leave the branch red for
a parallel session. Said in the workflow comment, the changelog, R-265 and the report — none of them
claiming it is answered.
GOLDEN 0.210.0 baked, published, round-trip verified, NOT VOUCHED. The currency gate went red the
moment the controller was bumped — correct — and is closed by the bake, never --no-verify. No
--no-verify anywhere this session.
⚠ THE AGENT WAS NOT PUBLISHED UNTIL THIS SESSION CHECKED, AND IT MATTERED. R-221's fix is in the
AGENT, and a fresh install takes its agent from the Day-0 manifest. The binary had been hand-deployed
to felhom-pve and never published, so agent_version 0.128.0 was not selectable and a fresh install
would have received 0.127.0 — the golden would have carried the controller fixes and NOT the one the
headline defect needed. Caught by checking each Day-0 value was FETCHABLE rather than assuming.
Published from the live-deployed bytes, sha-verified across the hop first.
Registers. R-221, R-259, R-258, R-265 CLOSED. R-266 MINTED (READY): the failed root statfs still
travels to the hub as a 0-of-0 disk; ranked LOW because it is the quiet direction — it can only miss
a true alarm, never raise a false one — and it is now a two-repo wire change governed by G-1's gate.
Highest ID moved R-265 -> R-266.
CONTEXT S-39 rules the convention this project was missing: "we do not know" is never drawn as
"fine", and the codebase has ONE way of saying it — an explicit ...Known bool companion checked in
the template. ROADMAP G-3 was explicitly blocked on that decision and is unblocked; what remains
there is a survey-and-convert of existing sites, not the gate.
Capability map row 93 CHECKED and it was NOT claiming something untrue — it is about the operator
notification path. But its narrative ("the page you open to ask whether ONE app is backed up")
invites the wrong reading, and the adjacent thing WAS false until v0.210.0, so the row now records
that the two halves disagreed and only the operator half was true.
Six red-proofs across the two code repos, each with the mutation asserted applied. The one that
matters: Part 1 Scenario A FAILED against today's tree, with the intended message.
Part 1's operator-present live validation is OWED and is the session's STOP.
repo_gates --fast: all 8 OK.
New shared scripts/instructions_gate.py, registered in controller_gates.py and
agent_gates.py, never copied into a sibling repo (the reuse_refs_check.py
precedent). 20 fixture tests, all asserting the effect: exit code AND that the
message names the file and the reason.
It is a consistency gate, not a budget gate, and the failure message says so. A
/context reading measured the instruction files at 15k tokens against 869k free in
a 1M window -- space is not the constraint, and a future reader must not re-derive
the wrong reason. The 200-line ceiling is adherence guidance; a file nobody can
hold in their head is where contradictions hide, and five were found here.
Checks run against effective text (HTML comments stripped, because they are
stripped before injection): the line ceiling; every .claude/rules/*.md declares
paths: or an explicit unconditional: true; no component version literal; no
TEMPORARY block carrying a past date; and the workspace-root CLAUDE.md is
byte-identical to its versioned copy -- the live file sits outside any git repo,
so that copy is its only version-controlled record.
Two traps recorded so they are not reintroduced: a bare \d+\.\d+\.\d+ matches the
first three octets of every IPv4 (the gate excludes dotted quads, or it fails on
192.168.0.180 in the agent's own file); and unconditional: true is NOT a Claude
Code feature but this project's own marker.
Workspace-root CLAUDE.md 208 -> 182 lines (142 effective), copy kept identical.
The nine-instance invariant table moved into the felhom-testing skill, which
triggers when writing or reviewing a test; all three directive bullets stayed in
the core. felhom.eu/CLAUDE.md got surgical corrections only and is knowingly still
over the ceiling at 227 effective lines -- closing it needs the restructure R-229
defers, said plainly rather than quietly absorbed.
CONTEXT.md gains standing ruling S-35. OPEN-ITEMS.md gains R-229.
Docs only -- no Go, no version bump, nothing built or deployed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JJc8sAGRWmavP3rMtdpkr2
- OPEN-ITEMS: R-196 CLOSED; R-204 items 1-3 CLOSED with item 4 named and
its dependency stated. Header restates that R-202, the 1.2 GB ciphertext
deletion and R-198's still-unit-proven retention all REMAIN OPEN.
- capability map: the recovery row keeps its 'with a person present'
qualifier, names which crutch remains, and cites the three now gone.
- 07-backup-architecture: new 7.0 - what a customer can and cannot do
ALONE, the four steps in a table with status. This is the section a
future reader will use to answer that question.
- CONTEXT: standing ruling S-32, superseding S-31 steps 2-5.
- STATUS: rewritten to one screen per its own header; removes a corrupted
half-overwritten section left from the drill session.
- ROADMAP: R-196 and R-204 collapsed.
Read-only recon of the escrow -> recovery chain, from a dead node to an open
repository. No production code, no build, no version bump.
Headline: the hub's superseded-escrow retention does NOT retain the offsite
repository password. host_escrow_superseded has no identity_blob column and
demoteCurrentEscrowTx copies only the K-escrow blob, so what survives a
supersession is the PBS datastore key, not the restic repo password. The next
escrow ceremony -- which the system tells a rebuilt box's customer to run --
destroys the last copy. Both demo boxes crossed that line on 2026-08-04.
Also established:
- the hub's blob-serving endpoints (re-enroll / restore-directive) have zero
callers anywhere: agent, hub UI, scripts, runbooks (R-199)
- POST /backup/offbox/inject-password is routed and handled but no template
contains the form (R-200)
- nothing in the recovery path has ever been exercised; the one live
round-trip proof (2026-06-10) predates the ResticRepoPassword field (R-201)
- a fail-closed mint refusal IS implementable: the report ACK already carries
escrow{identity_blob_present, restic_pw_sha256} and the controller discards
it whenever no offbox target exists
Corrections: yesterday's spike annotated (candidate (b) overturned in part --
unattended recovery is impossible, customer-present is not); capability-map
retention claim struck through and replaced with what the code does.
Deliverable: documentation/audits/RECON-offsite-dr-chain-2026-08-04.md
Register: new R-198..R-201; R-193 and R-192 updated; STATUS.md refreshed.
The 20-minute latch expired at 10:20:29 and the hub logged degraded -> ok
(agent_capability_recovered) at 10:30:40. Final state on both boxes: agent
0.124.1, two ACL rows on /storage/felhom-backup.
Part 0's gate PASSED — ep0 prunes both namespaces daily since 2026-07-27 (18
tasks, all OK) — but three of my own queries said the opposite and all three were
broken instruments. Acting on them would have disabled the only pruning attempt
while filing a finding that nothing prunes.
Also records that v0.124.0's transition record failed in production with a green
test suite, that two red-proofs did not fail on the first attempt (one could not
compile, one asserted a helper rather than the path), and that two hollow tests
were caught in one file.
Four SCHEDULED runs, none triggered by hand: demo-felhom host 83.8s / offsite
540.4s; demo-hp host 109.3s / offsite 300.1s. Each restored into a scratch guest,
booted, verified and destroyed itself; zero 990000 guests or volumes afterwards
and both local-lvm figures returned to their pre-run values.
Both boxes had BOTH tiers due at once, so R-86's ordering was observed live for
the first time: never-proven sorted first, each box took its HOST tier, deferred
the offsite one, and picked it up on the next evaluation six hours later. The
host-tier proofs reached the hub through R-189's merge — demo-felhom's report
carries two tiers, and the local one can only have come from disk.
The capability map's optimistic half is cashed, with its scope stated: these two
boxes, not the fleet.
Surfaced and filed rather than fixed:
- R-190: a storage ACL that demonstrably worked at 04:44 was gone by 09:24, with
a reinstall, any logged pveum activity and any cluster-log entry ruled out.
- R-191: every weekly offsite backup uploads successfully and then fails the job
on a prune the box is deliberately not allowed to do (R-89 moved it
server-side; both boxes still arm keep_last=2).
Two corrections to yesterday's record: the R-185 drift DID surface as 403s on the
write path (six, with the hub raising whole_guest_backup_failed at the first), and
my earlier "no restore_test_* events" was produced by grepping a 404 page.
- OPEN-ITEMS: R-86 CLOSED with the trap in its own wording recorded (the literal
reading is never true on a daily tier); R-87 re-ranked UP because R-86 built
most of what it waited for; R-185 (the agent cannot list demo-felhom's host
backup tier — a missing storage ACL, pre-existing), R-186 (a released binary's
sha is not reproducible from its tag), R-187 (R-115's publish leg had never
actually run) filed. R-184 was the highest ID in use.
- ROADMAP: R-86 collapsed, keeping the reasoning and correcting the shape the row
itself proposed — which would have been the never-fires version.
- 07-backup-architecture: new contract section — restore-testing is per ARCHIVE
GENERATION, with the trap and what did not change (S-1).
- 00-capability-map: the unattended restore-proof row upgraded to PROVEN-LIVE on
the 635 s due-triggered offsite run, with the restart and teardown evidence.
- CONTEXT: S-17 (the rule, the trap, the config key, the hub's derivation) and
S-18 (ep0 is Tier 2 — extends D-d's protected list to three machines).
Numbered 17/18 because S-14 and S-15 were already duplicated in the file.
- STATUS: rewritten for the operator, trimmed back to one screen.
R-182 CLOSED (controller v0.194.0 + hub v0.90.0/.1), proven live on demo-hp. The
hub's notification_log for the run reads: two per-app failures RECORDED, one
digest SENT naming both, and the customer channel SKIPPED with operator_only.
Against the measured previous behaviour — two failures, one email naming one
app, one leaving no trace anywhere.
Scenario D proved itself on an event I had not planned: disk_critical alarmed on
two filesystems, the second was collapsed by the cooldown, and that collapse is
now visible WITH ITS KEY. Yesterday it would have left nothing at all.
A gap the spec did not anticipate is recorded with its fix: the per-app event
also fires from the periodic sweep, outside any run, so making it record-only
would have created a NEW silence. The sweep emits a digest too, with no run_id,
so it stays under the ordinary hourly cooldown.
ep0: MEASURED on the box — 7757 MB (8 GB), 4 vCPU, and the 4 GiB swapfile
SURVIVED the resize and is active (checked, because a resize is a stop/start).
The 40 GB local disk is UNCHANGED, so no disk figure was touched anywhere.
Five documents corrected — three of which the task's list did not name, found by
searching. Two audit/evidence documents ANNOTATED, body untouched: they record
what was true when written and that is their value.
R-90 CLOSED. R-86 unblocked and re-ranked, stated honestly: 8 GB is comfortable,
not unbounded — the original OOM was a 14.46 GB restore — so the restore-test
cadence should still be paced, just not by fear of the endpoint.
target-selection.md's "D-d did not name ep0 either way" is deliberately left
standing. It is the operator's question, not CC's.
STATUS.md 127 -> 83 lines, items rather than sentences.
R-182's direction REVERSED by Part 0's measurement. Filed yesterday as "the
reserve re-alerts on every status refresh" — too many alerts, seen at the
sending end. Measured at the receiving end: 9 events received today, 2 operator
emails sent. When two apps are refused in the same second the operator is told
about ONE; the other is dropped before LogNotification, so it leaves no row on
any channel and cannot be audited. The operator cooldown key is
customerID:eventType(+tier) and the capture-failed event carries `app` but no
`tier`, so the key has no app identifier. Same failure mode as R-97a, in a
second event type that never opted into the narrow fix. Nothing changed —
Part 0 was investigation only.
Correction owed: yesterday's report said "one recovery_unit_capture_failed per
app, HTTP 200". True of what the CONTROLLER pushed; a reader would take it as
"the operator was told about each app", which is false.
R-110 CLOSED (installer v1.23.0). Both channels moved. The spec's mechanism for
channel 2 rested on a factual error — the run-time fetches are sixteen, not
nine, and come from felhom-agent, not this repo — so no tag here could cover
them; pinned to the agent version being installed instead, on the operator's
ruling. Channel 3 needed no change: the URL never carried a ref, so no hub
change and no hub bump.
R-115 CLOSED. release-agent.sh builds, tags, publishes and verifies by an
independent download; check-published-versions.py refuses a tag with no package;
CI now runs the full gate set so it actually runs.
R-183 NEW+CLOSED: a fresh install fetched the vouched agent binary and its
sixteen config files from two different refs, and nothing compared them.
R-184 NEW: nothing stops the hub vouching a version that was never released.
The R-115 gate cannot see it — measured, the hub manifest and Gitea's package
listing are both 401 anonymously.
capability map: new PROVEN-LIVE row for the published installer channel.
STATUS.md 138 -> 127 lines.
R-181 CLOSED (controller v0.193.0 + v0.193.1) and proven live on demo-hp for
BOTH reserve terms. The reserve is now a per-app, per-run ADMISSION decision
taken before the app's first write and covering all three write legs, and it
gained a size term. The refusal's wording was not weakened; the behaviour moved
so it became true, verified by sha256 tree fingerprint.
R-156 CLOSED — papra's template mounts the app's own data root. Precondition
re-measured rather than inherited (both boxes were wiped today).
Part 4, documentation only, nothing built:
- R-110 WAITING-ON-OPERATOR -> READY. Ruling: option (b), the installer's
publish channel moves to a TAG. Recorded with the condition that decides
whether it works at all — it must cover BOTH the /scripts/ git-sync and the
nine files the installer fetches from raw/branch/main.
- R-115 WAITING-ON-OPERATOR -> READY. Ruling: mechanism (b), a build-side gate
refusing to deploy or vouch an unpublished version. The third instance (agent
v0.120.0) would have silently downgraded both demo boxes while succeeding.
R-182 NEW: the periodic status refresh has no admission scope, so a refused app
re-alerts on every poll (measured: a second alert pair 13s after the run's).
Pre-existing in v0.192.0; deliberately not fixed in the R-181 task.
capability map: the local-backup row moves to PROVEN-LIVE in BOTH halves.
ROADMAP: R-165 collapses to CLOSED; R-181 collapsed into it.
07-backup-architecture.md: the reserve's contract stated as what the code
provides (S-1 — an architectural contract changed in the same session).
STATUS.md trimmed 150 -> 111 lines, "What's broken" no longer holds shipped
work, and the stale "After:" line (pointing at work that shipped on 2 August)
is fixed.