- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two
live-only defects, §6.4 part 1 SHIPPED.
- Capability map: a failed update is undone by the box - PROVEN-LIVE.
- Live evidence on 9202: three apps undone by the product with seeds before
the backup, after it and seconds before the press read back; cut-off copy
held honestly; power cut during the undo resumed; manual press after undo.
- Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643
ruled; R-646 opened. STATUS asks the floor question.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost
(PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came
back with data written before AND after the backup. The product's loader
cannot do it: over a migrated PG database it fails on the new tables'
foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads
rc 0 into an empty database. No-DB apps have no last-second copy.
- Part 2: one press jumps A -> C; the box's catalog clone is depth 1.
Ladder format recommended: update_ladder in .felhom.yml, not git history.
- Part 3: memory watch red-proof results (harness change in the catalog repo).
- Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643).
- Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT.
No product code. Live catalog untouched; 9202 back on it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Part C shipped the same day the pilot was read: fifty apps in three pushes, 1 031 of
1 032 strings. The English Apps list shows ZERO Hungarian app descriptions across all
53 apps — the only Hungarian left on it is the "Naprakész" badge (R-589) and the
language picker naming itself, which is correct.
The Hungarian Apps list is identical to the pre-slice capture once the per-session
CSRF token AND Docker's own "Up N hours" container string are normalised. Both
normalisations are stated in the evidence rather than applied quietly — the second
one moved because two hours of wall clock passed between captures, not because any
copy changed.
Fleet floor raised to 0.257.0 with the declared MinAgent 0.131.0, above the vouched
golden so the declaration carries it. demo-felhom went 0.255.0 -> 0.257.0 by itself
in under 12 seconds and THEN rendered the English tagline: the floor delivered the
feature, not a version string.
Rows: R-593 (papra describes a session-signing key as "the app's subdomain" — the one
string left untranslated) and R-594 (the catalog gate can convict a retrieval promise
but has no way to REGISTER a true one, which the shared vocabulary's design calls
for). R-560 closed. 281 -> 287 rows.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
10-localisation.md §7 goes from [DESIGN, proposed] to [FACT], with the numbers it was
specced against corrected — and one correction chose an instrument rather than a
footnote. The catalog has 1 032 copy strings, not 835. 832 carry a Hungarian letter,
which was right. But the ASCII-ONLY Hungarian is ~120 strings, not three: „Aldomain"
appears 53 times and „A szerver domain neve" 53 times, and the three the plan named
(„Igen"/„Nem"/„Nincs") do not occur in this catalog at all. An accent-only gate passes
every one of them inside an English block — R-565's blind spot arriving again in a
different repo.
New §10.6 records what was proven live rather than reasoned about: the English pages
show English; the seven Hungarian pages are byte-identical before and after the push
apart from the per-session CSRF token; and a 0.255.0 box with the block synced onto it
renders identically and logs no warning, in a 93-line window that contains the sync's
own lines, so the absence is evidence and not a dead log.
Rows R-589 (the update badge is Hungarian on an English page), R-590 (the data-folder
card's backup promise, likewise, and it is a promise about the customer's files),
R-591 (Stack.Copy() deep-copies five Meta fields and not the new I18n map — safe
today, which is precisely why it is a row), R-592 (three defects inside the new
catalog gate, closed the same session, each found by its own decoy). R-560 updated:
Parts A and B done, Part C waiting on the operator's read of the pilot.
STATUS asks for that read, and for the floor to 0.257.0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Design written after the spike ran (controller v0.247.0 live on demo-hp): mechanism,
flow, fallback, gates per language, catalog model, sliced plan with costs, operator
rulings 1-4 recorded, CC decisions 5-6, open decisions 1b and 7.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Architecture: 08 records ruling A (45 m, the round-9 arithmetic, the cost) and
the two new event types with their audiences; 03 records the slow counter as
built; 07 records the restore-record persistence as a REVERSED design for the
restore record only; CONTEXT.md carries the day's rulings.
Guide: the recovery code moves after the first apps and waits for the yellow bar.
Register: R-549 and R-550 closed PROVEN-LIVE; R-546 closed on red-proofed tests
with its live walk owed by R-551 (no Tier-0 box is paused AND agent-connected).
R-552 filed: an interrupted-restore notice for a removed app never clears -
found in my own v0.246.0 after the release was built.
Evidence: Part A (hub prints 45m/1h30m), Part C delivery on HP and N100, B.4(a)
live proof and its teardown.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The tier-3 pause is the zero-knowledge escrow design and is untouched. What was
missing was the ASK, while the backup page promised the copy that had never run.
- VOLUNTEER-first-hour.md: a new step 6, right after the dashboard password and
before the first app - what the code is, where, write it on PAPER, and that
Felhom cannot get it back for them. Sections 6..12 renumbered to 7..13.
- day0-install.md A.2b: the operator step for a REBUILT box, which was missing.
Acknowledged delete -> the hub re-issues by itself; otherwise ONE press of
"Re-issue PBS credentials" (F-14 ruling 2026-07-13, hub/internal/web/pbsdr.go).
This is the correction to last night's "zero presses" note.
- 07-backup-architecture.md: 6.1 records tier-3's paused state as a DESIGN, and
2 records that the household is asked from first login.
- capability map: the first-hour row's last gap closed, with what it still does
not claim (no volunteer has walked the ask from the written guide).
- register: R-543 CLOSED with the live measurements; R-545 filed (nothing
un-configures an off-site target). R-511 was already closed yesterday.
- STATUS: the answered publish question removed (1.28.0 is live), readiness yes.
- evidence: red-proofs, the two-box live validation, teardown on three layers,
and both of my own mistakes in this session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Off-site is ON by default for a new customer — shared, 100 GB soft quota prefilled,
the checkbox kept so an operator can opt a customer out. The reason is this repo's
own [FACT]: the whole-guest tiers do not carry the data drive and a Tier-1 unit has
no file leg, so with this unticked a one-drive box keeps NO copy of the household's
own files. Measured on a fresh box the same day.
The quota is prefilled because the fill warning only fires when quota_gb > 0.
Also registers controller v0.244.0's app_deploy_started / app_deploy_failed in both
allowedEventTypes and customerMessages, per the rule that the two move together.
Red-proofed: dropping the default fails the new render test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Phase 0: the public ISO never auto-installs by construction (no answer.toml,
G1); the operator re-affirmed the interactive installer 2026-09-14.
- felhom-bootstrap.sh: mask pvebanner.service, write a Hungarian /etc/issue
(no :8006 admin URL); pairing banner names the Tulajdonosi jelmondat and
paints through the CONSOLE_DEV seam (R-496). Harness: 8 checks, red first;
fake hub now sends a pairing code (the banner was never tested, R-502).
- hub: created flash + Credentials block tell the operator to hand the phrase
over; the self-bind mail names the operator (R-497). Tests red first.
- iso-release-gate G14-G16; domain ruling in 01-topology + CONTEXT; R-494
narrowed to P3; R-502..R-504 filed; volunteer guide and day-0 A.2 aligned.
ISO_VERSION 1.27.0 (not built, not published).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Hub v0.112.0 serves a floor above the golden with a declared MinAgent;
controller v0.239.0 reached both demo boxes by that floor in 14 s and 15 s
and updates on any backup tier. 09 §3 decisions 7 and 8, §6/§6.1; 07 §6
line; capability map row; STATUS items 15/16 done and the cadence line
corrected; CONTEXT; register: R-470/R-472/R-475 compressed to CLOSED-ITEMS
(full text at 2f5d3af), R-477..R-480 opened, R-474 reproduced a third
time. OPEN-ITEMS 431689 -> 432156 bytes, CLOSED-ITEMS 118051 -> 120598.
Evidence: documentation/audits/rulings-r472-r475-2026-09-13/.
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the
proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure.
Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/).
Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched
golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the
runbook, STATUS, CONTEXT, R-468 and the gate docstring.
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Operator rulings 2026-09-13, both shipped the same day:
- MariaDB finishes its own conversion (catalog eec1228/bd32830/3525e35). Harness E3/E3b `proven`
with engine_state_after "already upgraded to 12.3.3-MariaDB [exit=1]", the skip line gone, C3
still `failed`; landed on demo-hp through the real 15-min cycle, nothing recreated, one deliberate
restart logged "MariaDB upgrade not required" with the app serving. Evidence:
documentation/audits/r459-close-2026-09-13/. The engine-major rule + gate keep every engine
inside its major until Slice 4 (R-448) — removal tracked as R-469.
- Goldens on a cadence, not per release. golden_currency_gate.py reads a dated waiver
(documentation/tests/golden-waiver.yml, <= 14 days, row-bound): valid + BEHIND -> loud advisory,
exit 0; expired -> red again naming the date; UNRECORDED (R-385) never covered; malformed -> 2,
never 0. Tests cases 5-15 incl. the R-421 decoy; red-proof old-vs-new on the real behind tree.
R-242's vouch half stays open. Cadence in RUNBOOK-manual-build.md §4.2 + the checklist.
- Golden 0.236.0 baked, round-tripped, vouched, floor raised 0.232.0 -> 0.236.0
(documentation/tests/golden-0.236.0-2026-09-13/) — the last per-release bake; the waiver was
issued AFTER it landed. No --no-verify anywhere in this session.
Rows: R-459 CLOSED, R-467 CLOSED, R-242 narrowed; R-468/R-469/R-470/R-471 opened. 09 §3 gains
decisions 5 and 6; STATUS items 11 and 12 closed; CONTEXT records the cadence ruling.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Opens with Part 1's answer because everything reads differently after it: a customer's own account can
reach the snapshot DOOR and is REFUSED writes to it, but sees the tree EMPTY.
The write-refusal is the load-bearing sentence of the whole R-95 re-scope and it is now PROVEN rather
than cited - the control write to the account home succeeded and was cleaned up, the write into
/.zfs/snapshot returned `dest open ...: Failure`, and nothing was left behind. Identical on both boxes.
storage-box-pool-1 IS u629488, so the emptiness is per-sub-account filtering rather than absence -
which means recovery is an operator act in a browser today (R-432), and that decides whether R-95's
remedy can ever be product-driven.
STATUS carries two items for Viktor in plain words: read one snapshot name off the panel (two
minutes, and it may make recovery product-reachable), and IGNORE the alarm mail he received today -
the live firing was required to prove delivery and nothing was deleted.
Five of my own mistakes are named, including the one that matters most: my first escalation-only test
was HOLLOW and its red-proof PASSED. It re-swept the same report, so the baseline had already moved
and the latch was never consulted. That is why red-proofs are run.
THE CLASS, now a row: an instrument that matches a LABEL rather than the fact it names. Five
instances - R-410, R-400, R-378, R-419, R-94 - and EVERY ONE was found by accident, by someone
looking at something else. The gates enforce every other rule in this project, including the rule
that findings must be written down rather than left in prose. Nothing had ever checked the gates.
METHOD, and it is the transferable part: for each gate, construct the label WITHOUT the fact - a
directory with the right name and no bake log, a handler case that exists only in a comment, a note
whose prose mentions the marker it lacks - run the gate, record what it says. No verdict was reached
by reading. Reading is how all five hid.
RESULT: 29 distinct scripts (35 registrations; three are shared across three runners). 19 sound, 4
holes left OPEN with rows, 6 that no plausible decoy could be built for and are named UNTESTED rather
than called sound. A gate nobody tried to fool is UNKNOWN.
SCOPE IS A FACT TOO - the largest single cause, and mundane. Eight gates decided what to look at with
os.listdir, one level. Every one was green AND CORRECT today, and every one would have gone blind the
moment anyone added a subdirectory. mojibake and docker-v already used os.walk, caught the identical
planted file, and are the control that proves the cause was the listing and not the decoy.
IN THIS REPO: hub-confirm and manifest-bearer now walk. observations_gate (R-419, CLOSED) requires a
marker at a line start or after a sentence boundary and strips inline code spans - a note SAYING it
carries no marker no longer satisfies the marker test. closed-register now CONVICTS on a row it
cannot parse instead of warning: FOUR rows were in that state, TWO of them written by the session
that closed them the day before, and every one was exempt from the only check that reads that file.
The rows were repaired first and the conviction added second - registering a failing gate refuses
every push.
THE META-GATE: decoy_coverage_gate.py refuses a gate registered without a decoy or a named exemption.
It convicted ITSELF the moment it was registered, which is how it came to have one. Coverage is a
DECLARATION the gate AST-parses, never a grep - searching a test file for a gate's name would be the
very shape this sweep exists to find. The 20 uncovered gates are listed by name (R-426).
NOT FIXED, each with a row and a decoy asserting TODAY's behaviour so the fix must be deliberate:
R-422 reuse-refs (only 7 extensions; a rotted .md citation is invisible), R-423 site (PAGES is a
hardcoded list of 7), R-424 one-register (a defect parked as `idea`), R-425 offbox-rename (fixed
FILES list). R-427: closed_register_gate checks ONE direction - twelve open rows carry a closed
verdict and were NOT moved, because telling finished from partly-finished is a judgement and R-378
is the record of a machine getting it wrong.
FIVE DECOYS WITHDRAWN AS ILLEGITIMATE, mine, named in the audit. A decoy nobody would write proves
nothing, and manufacturing a finding to fill a row is worse than an honest NO.
No product code. No version bump. No image. No golden owed. All four runners green.
Register: OPEN 172 -> 178, CLOSED 160 -> 161.
THE RULING WAS NEITHER OPTION AS FRAMED. Both offered answers - narrow the gate, or leave it and
write waivers - argued about the gate, and the gate was never the problem.
DIAGNOSIS, from live source: golden_currency_gate.py never looks at the push. It compares the
controller's newest CHANGELOG heading against this repo's bake evidence and returns the same
verdict whatever you are pushing - correct for a standing invariant, wrong as a push gate. And
controller_gates.py had NO golden-currency entry at all. So the repo where a release happens never
checked, and the repo that cannot create the debt was refused on every push. 18 of the last 24
pushes here touched no code - measured, and the new classifier agrees EXACTLY - most of them by
construction, because the controller's code is in one repo and its register lives in this one. SIX
of those 18 were bake records, so the push that PAYS the debt is itself documents-only: the gate
was blocking its own cure.
Not the waiver its docstring prescribes: that clause was written for a release nobody wants a
golden for. R-417 was a release we DID want a golden for, on a night the runbook forbade baking. A
waiver would have recorded a lie.
RULING: block the push that can create the debt, notify the push that cannot.
The gate's logic, exit codes and wording are BYTE-IDENTICAL. Only the consequence changed, for one
gate, on one kind of push, with a loud ADVISORY block so nothing goes quiet.
R-242's vouch half is amended in place to say it is UNTOUCHED and still open - a baked-but-unvouched
golden still passes both the gate and the new notice. Do not read R-404's closure as closing it.
FILED: R-418 - this runner's docstring listed ELEVEN gates while THIRTEEN were registered;
one-register and closed-register ran undocumented since 2026-08-24. Enumeration fixed here, the
correspondence is still unenforced. R-419 - observations_gate.py accepts an item whose body merely
CONTAINS "NOT-A-FINDING", even in prose disclaiming it; found by accident when a planted test
observation passed and my live validation proved nothing. R-420 - controller_gates.py could not
express a non-blocking gate at all before today.
Register: OPEN 171 -> 172, CLOSED 158 -> 160.
The four existing skills cover the product; nothing covered how work is
reported. Two rules this project has paid for — check the artifact rather
than the report, and do not state a claim more firmly than the evidence
allows — lived only in the operator's head and in chat, where Claude Code
never read them.
- felhom-evidence five confidence tiers, artifact-over-report
- felhom-diagnosis no hypothesis until a command has been seen red
- felhom-plain-language ASD-STE100, two options, the re-pitch
- felhom-handoff the note goes to a FILE, not the conversation
- felhom-doc-authoring the pointer decides whether material is reached
scripts/check_skills.py asserts what decides whether a skill is EVER
reached: frontmatter parses, name == directory, description and body
non-empty, under 150 lines, installed copy still samefile()s into the
repo. install_skills.py globs and never reads the file, so a missing
description installs perfectly and then silently never loads.
It convicted on its first run: felhom-build-deploy is 179 lines. NOT
trimmed here (pre-existing skills are out of scope, and trimming a
deploy skill without exercising its commands is how a wrong command
reaches a live host) — a named single-entry GRANDFATHERED exception,
WARNed every run, R-394. A new skill over the limit is convicted.
Red-proof run and seen failing: description removed from
felhom-evidence -> exit 1, "frontmatter field 'description' is missing
or empty". Restored, tree clean.
skills/SOURCES.md records both MIT upstreams, that these are adaptations
not copies, and the six pieces deliberately EXCLUDED with reasons.
Register: R-392 (no architecture doc covers the two-AI workflow),
R-393 (decision-log skill deferred, with the reason), R-394.
The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or
coarse, and why the default is coarse. CONTEXT records two rulings: the grain is
allow-listed rather than inferred from the payload, with crossdrive_failed as
the proof that a payload rule would have been wrong; and a finding recorded only
in REPORT.md has a lifetime of one session.
R-389 closed and compressed, keeping its rules and naming the commit whose
git show returns the full text. R-390 and R-391 left open.
REPORT.md is gate 11's first real subject and passes: six observations, two
FILED, four NOT-A-FINDING with their reasons. Three of those declarations are
things a tidier report would have omitted - the gate's own spec would have
passed the item it was built to catch, the burst has no ceiling, and ArgoCD
said "successfully rolled out" while still running the old image.
STATUS carries forward the one thing outstanding: the controller floor still
reads 0.222.0 while the golden reads 0.223.0.
The alarm ladder gains the severity contract (the hub's vocabulary is exact, it
coerces silently, and three things now hold it) and the intent test with its
three-way ruling on unknown. Both marked [DESIGN] with the live measurements.
Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy,
verbatim, marked plainly as direction rather than current behaviour, with the
12 -> 15 toggle growth as the argument. Filed as R-388, a product decision.
R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the
full-text commit named. R-387 filed closed - including WHY the dispatcher branch
was kept rather than deleted, which is evidence (three monitor checkers call
ProcessEvent directly) and not caution.
The drill record names three things that had to be re-run: an inert red-proof
mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A
NOT proving the customer gate because demo-hp has no prefs row at all.
Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B.
Records and process only. No machine contacted.
ONE REGISTER (operator ruling). 17 roadmap rows moved into OPEN-ITEMS.md keeping their
identifiers, evidence and original filing dates - the oldest R-10, filed 2026-07-15, 38 days.
15 ideas stay in ROADMAP.md, which is their home; the gate exempts them by their own state
word. 59 already-closed rows stay as history. Sorting rule recorded in the roadmap header:
does the item assert something about the shipped product a reader could check and find false?
scripts/one_register_gate.py, wired as the 11th gate. Control run: baseline passes, a planted
open roadmap-only row is convicted by name, removing it passes with the file byte-identical,
and a planted `idea` row is correctly exempt. Its four residual holes are in its docstring.
The gate earned its keep immediately: it caught R-103, a READY finding my hand-sort mis-read as
done because my regex matched the whole row where the body contains "shipped" - the gate matches
the state cell. It also caught R-203 and R-163, recorded closed in the register and still open in
the roadmap; the roadmap copies are marked SUPERSEDED with the register's verdict.
HOUSEKEEPING. OPEN-ITEMS 672,376 -> 327,109 bytes (-51%); ROADMAP 239,306 -> 78,110 (-67%).
Closed work compressed to 17% into CLOSED-ITEMS.md and ROADMAP-HISTORY.md; every entry names the
commit whose git show returns the full original text. Rule-sentences are kept verbatim under
"Reasoning kept" rather than judged entry by entry - 25 carry one.
CONTEXT.md deliberately NOT compressed and the disagreement is argued in the report: 86% of it is
standing rulings still in force, this prompt's own 3.4 says the log is never edited, and it has no
per-ruling delimiter. Filed as R-377 - the problem is navigational, not volumetric.
The hot/bulk placement decision was NEVER recorded as a decision anywhere - established, not
assumed. Now marked [DESIGN] with a pointer honest about having no original date, given a
decision-log entry that records what was rejected, and the [DESIGN]/[FACT] legend carried from 1
of 8 architecture documents to 8 of 8. Existing statements deliberately left unmarked (R-376).
PROMPT-TEMPLATE gains N.7: compress what you closed, rehome live reasoning before it goes, state
the register's size before and after.
Ceiling R-375 -> R-378.
Documentation and survey only. No code, no machine contacted.
THE CORRECTION. The 40 catalogue templates without a configurable path are not missing a
choice: 01-topology-and-trust.md:150-152 classes each volume hot (DB/config/cache -> fast
storage, ENFORCED) or bulk (media/files), and the 40 are all-hot apps. The deploy page has
been saying so to the customer all along (deploy.html:624-625). SPEC-app-data-placement and
R-352 are corrected in place with the framing MARKED, not deleted; every measurement stands.
R-356 was re-checked and survives, strengthened - an absent HDD_PATH is the normal state, so
reading it as "not installed" misreads a correct configuration.
The disk claim, precisely: since R-165 there is ONE guest data volume with two binds, not two
volumes (build-golden.sh:29-40, 99). A physical-disk failure losing data and first-tier copy
together is REAL and is what the other tiers exist for. A full data volume stopping the OS is
NOT real and was the overstated one.
THE SWEEP. 113 survey-class documents examined, 14 statements of "not filed", 2 already filed.
Its positive control convicted the sweep itself twice before it convicted the corpus - markdown
bold broke the strongest pattern, and the reporter re-searched a truncated line - both false
zeros of the exact class being hunted, and together worth 2 of the 14.
THE HEADLINE. The gap the 2026-08-21 drill rediscovered WAS filed - as R-107, ROADMAP.md:122,
M/READY, 2026-07-28 - and is absent from OPEN-ITEMS.md, which calls itself the single source of
truth. OPEN-ITEMS and that rule both landed 2026-07-27; R-107 went to ROADMAP alone the day
after. 72 ids live only in ROADMAP, 29 not done, some of them findings. Filed as R-369 (HIGH).
Five more still-open gaps filed with their ages: R-371 (17d), R-372 (38d, the oldest), R-373
(20d), R-374 (14d), R-375 (4d). R-368 corrects Part 4: the storage default IS applied at deploy
time via deploy.html:612 - the earlier "the deploy route never reads it" came from grepping Go
and never the templates. R-370 records the process failure and is closed by the template change.
PROMPT-TEMPLATE gains the two rules it lacked: name the architecture document for the area and
say what it says (with a file->area map and the test "is this something we chose?"), and an
enumerated gap becomes a register row in the same session - a ROADMAP row alone does not count.
Ceiling R-367 -> R-375.
THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.
REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).
THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.
Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
- the unreachable event REPEATS rather than escalating once. The band shape
would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
mail is missable. It leans on the dispatcher's 1 h operator cooldown to
become an hourly "still blind" heartbeat.
- ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
which means we reached it. Counting it would alert for days on a healthy
pre-update endpoint.
- born-blind is reported: the counter is not gated on having a snapshot, so
a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
than zero-valued -- a fabricated timestamp reads as "it was fine until
then".
Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).
Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.
Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as
prose in a register row; nothing read those dates and nothing would have
objected when they passed. The dates now live in a DUE-CHECKS block INSIDE
OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads
them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the
pre-push hook and CI.
exit 0 nothing due (prints pending count + nearest date; empty block too)
exit 1 a row is due/overdue (due <= today, UTC -- due TODAY counts), or a
row names an item with no R-row
exit 2 block absent/duplicated/unparseable -- INCONCLUSIVE, never 0
It REFUSES rather than warns, and its docstring states the limitation: it is
NOT a scheduler, it fires on the next push, not on the date.
37 tests. BOTH red-proofs run and reverted -- and the first one earned its
keep by catching a hollow assertion of MINE rather than confirming the gate:
flipping <= to < left a due-today row in neither bucket, min() raised on an
empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the
boundary was wrong. An exit code cannot tell a verdict from a crash. The test
now asserts the conviction banner and the absence of a traceback, and the gate
returns 2 rather than crashing if that partition breaks again.
PART 3 — the floor raise, and the premise was WRONG. Read back from the store
(not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero
per-customer overrides, no "managed floor HELD" line. But read 5 shows the
raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and
auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save,
exactly the immediate action publish-train rule 2 documents. No error events
followed; it restarted clean.
R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five
reads clean and no directive served. It went well, but a record calling it
inert when it moved a customer box is what misleads the next reader. The row
also states why the floor was behind -- rule 2 policy, not drift, earned by
the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor
(store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267,
fails open at :78-82) rather than asserting them.
Two boxes are below the floor and neither reports: drill-r50 (blocked,
powered off) and peti-felhom (host row deleted). peti-felhom was NOT
contacted -- its row records that a report from a deleted host 401s and is
not persisted, so the raise cannot reach it.
PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner
server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a
separate Volume that snapshots exclude, so a rollback restores software state
and NOT the datastore. Fine for that upgrade; the safeguard for any future
procedure that could touch the datastore does not exist and is Viktor's call.
Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed
rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling
200). Capability map deliberately unchanged; no row cites a floor or golden
version. repo_gates.py fully green, 10/10.
Three rules carried forward. A name must separate on the STEM, not the noun — naming
this secret after the act it is used in would have recreated the trap, because the
other factor on the same page is the „Párosító kód". A guard is worth what its positive
control is worth: this one's selftest convicted its own step-3 case and found a defect
in the guard itself. And a suppression must rest on the machine's own declaration, then
be checked for the SECOND door — recording the disabled state rather than deleting it
is what let the deadline check skip it too.
Yesterday's report is preserved to audits/ because it carries the only record of the
self-heal verdict (Part C was dropped, so that reasoning is in no register row) — the
rule written last night, applied to itself the first time it mattered.