Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
28 KiB
OPEN-ITEMS — the narrative sections, as they stood on 2026-10-03
History, not a register. These sections sat between the register tables of
documentation/backlog/OPEN-ITEMS.mduntil the 2026-10-03 triage. They are moved here word for word (git show 9e2786c:documentation/backlog/OPEN-ITEMS.mdholds them in place). Every row they discuss is inOPEN-ITEMS.md(open) orCLOSED-ITEMS.md(finished). A status word below is the status ON THE DATE OF ITS SECTION and may be stale — the registers are the authority. The ranking paragraphs ("Why the TOP READY rows rank this way", "The 2026-08-02 intake, ranked") are superseded bydocumentation/audits/backlog-triage-2026-10-03/RECOMMENDATION.md.
Operator rulings — 2026-08-04
Recorded here because a ruling that lives only in a conversation binds nobody (the R-96 standing rule).
- Run the recovery drill, after R-198. R-198 shipped in hub v0.93.0; the drill is the next
session. Design:
audits/RECON-offsite-dr-chain-2026-08-04.md§10 — demo-hp, ~3–4 h, a recovery code created and KEPT, a sentinel file, wipe, reinstall, recover, and pass = a byte-identical sha256, not "the repository opened". → R-201 - Delete the orphaned ciphertext (~1.2 GB across the two demo boxes, in set-aside restic stores nothing prunes). STILL OWED — not done in v0.93.0. It is a destructive act on a protected endpoint and belongs to a session that is scoped for it, not to a release that ships a schema change. → R-193
- Accept the risk on R-193 candidate (c) — no repository password retained on the Proxmox host. This is what makes R-198 load-bearing rather than tidy: with no host-retained copy, the customer-present recovery path is the ONLY way back from a rebuild, and that path runs entirely through the retained identity blob. → R-193, R-199, R-200, R-201
Still open and untouched by v0.93.0: R-199, R-200, R-201. UPDATED 2026-08-04 (evening), after
hub v0.94.0 + agent v0.125.0 + controller v0.195.0:
- R-199 — CLOSED, proven on hardware. Chain links 6–8 are assembled and walked. The offsite
repository password came back out of the sealed bundle byte-identical to the one on disk
(
c60c8bc737a6…from three independent sources: the box's file, the recovered bundle, and the hash the hub already stored). - R-200 — the plumbing half shipped, the customer-facing form did not, deliberately.
- R-201 — PASSED 2026-08-04 (night run). A customer's file survived a machine rebuild and came back byte-identical, through the customer's own restore flow. It passed only because a person was there: four manual interventions stood between the recovered key and the restored file, none of them in any design document → R-204.
- R-202 — untouched. The orphan card still promises recoverability unconditionally.
- The orphaned ciphertext deletion (~1.2 GB) is STILL OWED — ruling 2, above.
UPDATED 2026-08-05, after controller v0.198.0 + hub v0.95.0 (R-204):
- R-204 items 1–3 — CLOSED. The reset code works without a restart; a re-issue no longer marks a healthy escrow stale (→ R-196 CLOSED); a unit restore states what it did NOT restore.
- R-204 item 4 — OPEN and unstarted: a rebuilt box cannot obtain an off-site credential unaided. It needs an operator ruling on the one-shot credential design → R-193.
- Still open and untouched by this session, stated so nothing is presumed closed by association: R-202 (the orphan card's unconditional promise), the ~1.2 GB orphaned-ciphertext deletion (ruling 2 — still owed, still needs its own scoped session), and R-198's retention, which remains UNIT-PROVEN ONLY — nothing has superseded a key in production, and proving it needs a SECOND deliberate wipe. That retention drill is the next item, and it is not this session's.
v0.93.0 made the key survive. v0.94.0/v0.125.0/v0.195.0 make it come back. v0.197.0 got the file into the snapshot and the drill got it out again. v0.198.0/v0.95.0 remove three of the four crutches the drill needed — the fourth is R-193, and until it goes the recovery is still operator-assisted.
R-201 — THE RE-WALK, 2026-08-06 (attended)
The question was asked a second time, on the fixed build, on a brand-new appliance built from the published ISO. The answer is still no — but it is a nearer no.
| half | verdict |
|---|---|
| the data | PASS — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename whose NAME BYTES are also byte-identical (verified as hex, not as rendered text). Restored in 23 s out of the pre-destruction snapshot a7bc23bd, through the customer's own two-step full-restore flow |
| the journey | FAIL — two dead ends, against Phase 1's four. One needed a command line inside the guest, one a Proxmox-host action |
RTO: the unaided figure is STILL UNDEFINED, because the unaided journey still does not complete. Attended: login 11:42:22 → key placed +45 s → tier up +24 m 12 s (after intervention 1) → all three sentinels restored and verified +30 m 13 s (after intervention 2). The 30 m figure must not be quoted as the customer number. The only segment that reflects the product working alone is 23 seconds to pull 12.8 MB back once everything was in place.
The two dead ends: R-218's consume half (row corrected above) and R-220 (drives unenrollable after a rebuild — reproduced and red-proved again; without it no app can be redeployed, and without a redeployed app the restore page is empty, which is R-213's territory and follows from R-220 rather than being separate).
What PASSED and is worth keeping: the recovery screen appeared without being sought
(/ → /launcher → /recovery); it answered all three of its questions and its seal date matched the
hub's created_at exactly; the emailed reset code worked first try; the unlock took 1.528 s —
a real unseal — and placed the key; and R-225's fix was seen working in the wild (the store read
„a pillanatképek száma még ismeretlen" rather than a false zero).
R-216 part 4 reproduced live: the reinstall downgraded the agent 0.126.0 → 0.125.0, back to the vouched version — an operator's hand-fix undone by the very event that makes recovery necessary.
⚠ THE DELIVERY GAP, and it is owed. A fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes. They were installed by hand. Fleet delivery needs a golden carrying 0.202.0 and a vouched agent 0.126.0. Nothing was vouched; that is the operator's act. This re-walk proves the JOURNEY on the fixed build; it does NOT prove a customer would receive that build.
Evidence: tests/rewalk-r201-2026-08-06/journal.md.
CAMPAIGN 11 — the recovery journey, 2026-08-05
The whole journey was walked end to end for the first time, on a throwaway appliance built from the
published ISO. The data came back byte-identical; the journey did not exist. Ten findings from
Phases 1 and 3, R-214 … R-223 (seven fixed in controller v0.201.0 + hub v0.97.0/0.97.1;
three deliberately still open, each blocking a real flow), plus five from Phase 2's injected
faults, R-224 … R-228. Evidence: tests/campaign11-evidence-2026-08-05/journal.md (Phases 0/1/3)
and journal-phase24.md (Phases 2/4). Campaign document:
audits/CAMPAIGN-11-recovery-journey-2026-08-05.md.
Phase 2's verdict in one line. The cryptography, the retention and the transport all work and are now proven live. What fails is being told the truth: a mistyped code, a hub outage, a stopped agent and a correct code for a retained earlier package all produce one message, and three of the four are wrong.
Phase 2 — the injected faults, 2026-08-05/06 (unattended)
ALL FIVE CLOSED 2026-08-06 in controller v0.202.0 + agent v0.126.0. The rule they now enforce, stated so it outlives them: on the unlock path the customer is blamed only after a real attempt REFUSED their code; every other outcome, including an unclassifiable one, says something else.
STILL OPEN AND DELIBERATELY UNTOUCHED BY THAT WORK — said explicitly rather than left to inference: R-214 (the console never stops showing a stale pairing code), R-220 (a rebuilt box's drives cannot be re-enrolled — currently worked around BY HAND on the campaign venue, which is the only reason an app could be deployed there at all), R-221 (a rebuilt box cannot run the escrow ceremony), R-213 (putting files back), R-202 (the orphan card's unconditional promise — now the last place on that surface still promising recoverability, two doors from where R-228 removed the same promise).
Eleven faults, each judged on the message, not the outcome, each with a positive control proving
the fault was real. Full observables: tests/campaign11-evidence-2026-08-05/journal-phase24.md.
Instruction files — deferred half, 2026-08-06
Recorded against existing rows by Phase 2:
- R-216 — §4.1 is now MEASURED, not deduced. The previous session could only offer two absences.
The box's own
/settingsrenders „Minimális verzió (üzemeltető) 0.200.0" (GetFloor(), whose only writer is the report-ACK handler; both hold branches serveFloor="", pinned bymanaged_floor_test.go:94), and a cold-started controller logssettle-gate: GO — at/above floor 0.200.0 (we are 0.201.0)against the same line readingfloor still unknown after 1m30swhile the hold was in force. The hub's HELD lines ran every 15 min to 22:12:06 and stopped, with a liveness control proving the hub kept logging. The floor is served. - ⚠ A correction to how that positive was to be taken.
SetFloor's line isu.dbg(...), gated oncfg.Logging.Level == "debug"and written to the logger — it can never reach the logx debug ring, so it cannot appear in/api/debug/logsat any level. A controller restart alone would not have produced it. Confirmed with a level census on the ring first (1196 DEBUG / 2802 INFO / 2 WARN), so the absence was known to be structural rather than evidential. - R-218 — the live half is STILL NOT MEASURED, deliberately. The fix is present and correct
(
needsOffsiteCredentialnow retires on the target, not the key), but the venue has a target, so the box correctly does not declare; declaring here would be the bug. The state that exercises it is shape (a), which the venue no longer holds. Recorded as not measured rather than inferred from the unit test. - R-217 — its fix HELD under exactly its fault (F5): with the store blocked after a successful unlock, the page rendered M3 and no listing block at all; the three false-claim strings are absent, verified in UTF-8 with accented positive controls present.
- R-215 — its fix is present and the
GET /recoverygate consults the same predicate as the POST sibling. - R-199's back-pointer in
architecture/00-capability-map.mdwas already added — the brief lists it as owed; it is present on the escrow-recovery row, explicitly labelled as the omitted back-pointer. No action taken; the brief's assumption was stale.
Untouched by this session, stated so nothing is presumed closed by association: R-213 (putting files back — the half the recovery screen deliberately does not do) and R-202 (the orphan card's unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing cost of).
CAMPAIGN 12 — the class sweep, 2026-08-08 (unattended)
Eight rows, grouped by class so the classes are visible as classes. Full method, controls, blind
spots and the Part-4 gating ranking: audits/CAMPAIGN-12-class-sweep-2026-08-08.md. Gating candidates
are in ROADMAP.md, not here. C1 produced no new instance and has no row, deliberately.
⚠ A correction the campaign owed to its own brief: the task described C5's escrow_stale instance
as "closed individually". It is not — it is R-247, READY. The live repo is the source.
CAMPAIGN 12 follow-through — G-1 built, R-260 and R-247 closed, 2026-08-08
The gate was built BEFORE the fixes and was seen failing on 40 fields, captured verbatim in
documentation/tests/wire-contract-gate-2026-08-08/BEFORE.md. That order was the method, not
bureaucracy: the night before, deadcode was rejected for class C6 precisely because it was made to
prove itself first and found neither of the two defects it was meant for.
⚠ A COUNT THIS SESSION'S OWN PROMPT GOT WRONG, corrected against the repo rather than quoted. The prompt said "465 emitted tags, eight unreachable". R-260's wording was "at least eight DECISION-BEARING facts", never eight tags in total, and its own census already listed more. Measured on the three declared wires: 210 tags checked, 51 skipped, 40 convicted. Two prompt claims were wrong this week and both were caught the same way.
Two things the gate's CONTROL caught before it was trusted, each a defect in the instrument:
- A substring false negative.
grep -F healed_atalso matchesprivsep_healed_at, so a genuinely dropped field read as received. R-260 namedhealed_at, so its absence from the output was the tell. Now a whole-token regex. dr_recipeis not wholly opaque. The hub stores each half asjson.RawMessageand re-emits nested shapes verbatim — but the TOP-LEVEL section keys are decoded byhostHalfShape/appHalfShape, and those are allow-lists: a section an emitter adds is silently dropped until named in both. That already costoffsite_restic(R-122). The gate is therefore opaque BELOW depth 1, so the sections are checked; treating the whole subtree as opaque would have put R-122's shape back outside its reach.
THE FORTY, BY DISPOSITION. Full per-field reasons live in the gate's own ALLOWLIST, where each
entry is a claim someone can re-check.
| # | field(s) | direction | decision | what changed |
|---|---|---|---|---|
| 1 | oob.operator_key_configured |
agent → hub | RECEIVE AND ACT ON IT | decoded as a pointer; oobDegraded now fails when the key is absent, and the alert NAMES it. Hub v0.99.0 |
| 2 | oob.wg_handshake_age_s, oob.healed_at |
agent → hub | RECEIVE, for the message only | decoded into HostOOBRow and put in the event payload; deliberately NOT in the predicate — widening the check beyond the fact that is now arriving is how a check stops being read |
| 3 | escrow_stale |
hub → controller | RECEIVE AND ACT ON IT | report.EscrowStatus.Stale; the box tells a withheld hash from a hash-less one. Controller v0.209.0. This is R-247 |
| 4 | host.{cpu_temp_c,loadavg,memory_total_bytes,memory_used_bytes,uptime_seconds}, system.{load_avg_1,5,15, memory_total_mb, memory_used_mb, temperature_celsius, uptime_seconds} |
agent/controller → hub | NO CONSUMER WANTED — redundant | allowlisted: the hub decodes cpu_percent / memory_percent / disk_percent from the same stanzas and every threshold is expressed on those |
| 5 | guests.spec.{disk_bytes,memory_bytes} |
agent → hub | NO CONSUMER WANTED — redundant | allowlisted: guest sizing is hub-owned INTENT (the manifest), not mirrored reality |
| 6 | storage_targets.smart.model_name |
agent → hub | NO CONSUMER WANTED — redundant | allowlisted: a display label with no threshold on it; smart.health and every counter the hub bands on ARE decoded |
| 7 | wireguard.last_handshake_age_s |
agent → hub | NO CONSUMER WANTED — redundant | allowlisted: wgsync reconciles from its own state, and the OOB path's own handshake age is now decoded |
| 8 | guest_net + its 7 children, selfupdate_pending(_version), mgmt_plane.healed_recently, pbs_dr.applied_at, restore_tests.mount_{parity,inventory}, config_hash, reporting_disabled, stacks, storage.migrated_to, backup.last_db_dump, backup.last_integrity_check |
agent/controller → hub | NO CONSUMER TODAY, AND ONE IS ARGUABLY OWED | allowlisted against R-264, which stays OPEN. Allowlisting is not deciding, and the entries say so |
A C7 instance found while doing it, and corrected. HostReport.SelfUpdatePending's own comment
claimed "The hub reads an absent field as pending=false, the correct default." The hub has no field
for it and reads nothing either way. Comment corrected in the agent (no version bump — comment only).
The hub's OOB test fixture was part of why this survived. oobReport() omitted
operator_key_configured entirely, so every pre-existing scenario ran against a report shape no
released agent produces. Fixed, and the new tests drive the raw JSON decode boundary — a test that
builds the receiving struct by hand cannot see a field that never decodes, which is the whole class.
Explicitly still open, untouched by this session: R-246 (the wrong stale flag on demo-hp —
clearing it is an operator act hub-side), R-255, R-256, R-257, R-258, R-259, R-261, R-262, R-263, and
C7's test-comment half, which Campaign 12 recorded as owed, not done (60 of 2652 production
invariant comments sampled; none of the 1440 test comments).
The seed that never ran twice, and three pictures that were not true — 2026-08-08
Four defects of one family: something the box already knows, either thrown away or drawn as its
opposite. Agent v0.128.0, controller v0.210.0, gates.yml (no hub change, no hub bump).
R-221's writer, ESTABLISHED at file:line rather than assumed — the prompt asked for this and it
was owed. step_agent_config (felhom.eu/scripts/felhom-host-install.sh:2396) renders agent.json
from base = {} unless an explicit --preserve-from is passed (flag :1246, defaulting empty at
:256), and writes it with O_TRUNC (:2579). The render never writes an escrow section at
all — grep over the whole heredoc returns zero hits. The pbsdr marker lives host-side
(<agent-state>/pbsdr/marker.json) and survives. So a rebuild keeps the marker and takes the key:
same descriptor, same hash, early return, seed never re-runs. The attribution in R-221 was
correct. A rebuild is nonetheless only the case that was measured — the same hole opens for a
hand-edited or restored config, which is the honest reason the fix is at the seam and not in the
installer.
The idempotent early return was KEPT, and that is load-bearing: it stops a converged box
re-running Proxmox operations every 60 s.
TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls asserts zero recorded runner calls on
that tick, so a "fix" that simply deleted the return fails the test. Verified by mutation.
The §7.3 truth table as implemented (R-258):
| this app's own most recent dump result | restore point | verdict |
| any of its databases failed | yes | error — cross |
| all clean | yes | ok — tick |
| none recorded (no database / no run yet) | yes | no icon, time only, title „Erről a mentésről nincs eredményünk." |
| any | no | no tier-1 row at all, unchanged |
An existing test was asserting the defect and was corrected, not deleted.
TestBuildAppBackupRows_Tier1FromRestorePoints expected "ok" for a FullBackupStatus with no
LastDBDump at all — a green tick derived from nothing but a file's existence, i.e. Scenario G.
Its real subject, the Tier1LastRun time, is unchanged.
The convention is now ruled (§7.2, CONTEXT.md S-39): a …Known bool companion beside the
figures. ROADMAP.md G-3 was blocked on that decision and is unblocked.
Six red-proofs, every one demonstrated failing and restored, each with the mutation asserted applied. The one that matters: Scenario A fails against today's tree with the intended message — so the test tests the defect.
Explicitly still open, untouched by this session: R-246, R-255, R-256, R-257, R-261, R-262, R-263, R-264 (the twenty-one undecided facts — a design session of its own), R-240, R-243, R-202, R-213, R-244, R-214/R-235, and C7's test-comment half, which Campaign 12 recorded as owed, not done. G-8's other half (a hub-side check that notices a vouch has been forgotten) was deliberately not built: it is hub work whose payoff is a daily email, and this session already ends with a bake-and-vouch cycle in front of the operator.
Why the TOP READY rows rank this way
This covers the next few only — it is deliberately not a full ordering of the table above, so that there is one ranking to maintain rather than two.
- R-95 — the largest data exposure: the tier holding the customer's documents and photos is the
one whose credential can delete. THIS PARAGRAPH WAS STALE UNTIL 2026-09-01 AND ITS OLD TEXT IS
NAMED SO THE CORRECTION IS NOT MISTAKEN FOR A RE-RANK. It said the snapshot mitigation was
"armed (daily 00:00, keep 7), but it has taken zero snapshots so far". Both halves were wrong:
seven daily snapshots do exist (R-429), and the word "armed" was withdrawn by the spike that same
day. What is true now: the snapshots exist and the box cannot write into them, but no account
we hold can read anything out of one (R-433, measured — 777,600 names, zero hits, controlled), so
they do not yet bound this exposure. The root cause is untouched either way — the box can still
forget --pruneits own repo, from two call sites. The ORDER of this list is unchanged and is Viktor's; only the facts under item 1 were corrected. - R-94 — de-ranked 2026-07-29. The prior rationale ("until it moves every hub-driven install gets the pre-R-82 default") was false: the constant selects no script and every install already fetches 1.22.0. What remains is a wrong label plus two pieces of dead safety equipment — a drift gate nobody runs and a test that compares a constant to itself. Cheap and worth doing; not high-consequence, and it blocks nothing.
R-86— CLOSED 2026-08-03, agent v0.121.0 + hub v0.91.0, proven live on demo-felhom.- R-87 — SPIKED 2026-08-31 and now a DECISION, not work. It also spent 2026-08-22..31 in
CLOSED-ITEMS.mdby mistake while this paragraph ranked it fourth and pointed at nothing (R-405). The spike measured it rather than designing it: a scratch restore of all 8 apps costs 25 s against the 40.3 s weekly check, but it would have caught ONE of the five drill-found restore defects. Recommendation: build the NARROW version — prove the snapshot still CONTAINS a recoverable unit — or close the row. Viktor's call; seeaudits/SPIKE-restic-restore-test-2026-08-31.md. The 2026-08-03 rationale, kept: R-86 built most of what it was waiting for (per-archive due-ness, a proof that names its archive, and a staleness window that learns a tier's rhythm). What is left is restic-specific — there is no scratch-guest analogue — so it still needs its own design, but it is no longer waiting on a scheduling model that did not exist. R-185— CLOSED 2026-08-03, agent v0.123.0 + installer 1.24.0, proven live on both demo boxes. The silence was fixed as well as the grant: the box now asks whether it may READ each tier it depends on, because an empty listing cannot distinguish forbidden from newborn.R-189— CLOSED 2026-08-03 with R-188 and R-186, agent v0.122.0. The three reporting/release signals that misreported their own work are fixed; R-185 is the one that remains open from that group and is untouched by this — it is a missing storage ACL on demo-felhom, not a reporting defect.- R-110 — last because it is not a READY row: the ruling is the operator's, not CC's, and there is nothing for CC to build until it lands. Ranked here rather than omitted because it is the only item on this page about the publish channel of the most privileged artifact Felhom ships, and today's exposure is zero — which makes now the cheapest moment it will ever be to decide.
The 2026-08-02 intake (R-156 … R-164), ranked
Filed in one pass from Campaign 10, its two spikes, and the 53-template catalog persistence sweep. R-156 and R-157 had lived only in audit documents — the identical "minted in a spike doc and never carried across" failure the register already records for R-153/R-154/R-155, caught by the sweep's own §8.0 while it was happening. R-158 was minted by a second session on the same day for an unrelated finding, which is why the sweep's proposals were renumbered to R-159…R-162 at filing time.
- R-157 — highest: a
deployed: trueapp can stay down indefinitely after a power cut or hard reset, and in mechanism B nothing reports it on any channel (0 currently down). It is the only row here where the customer loses service and has no signal at all. - R-156 — the class is now detectable and two of three apps are fixed; what remains is papra's referral, one app, well understood. (Promoted 2026-08-02: R-161 was ranked here because nothing ran the gate; it now has a mandated entry point, so R-156's residue is the larger remaining item.)
- R-163 — a real ceiling that silently caps local backup once an app outgrows
mp1, and it gates Tier-2 and Tier-3 as well. Ranked below the above only because overflow itself is safe today — it refuses per app and preserves the last good unit byte-identical. RE-FRAMED 2026-08-02: no longer waiting on a ratio — decision D-a mergesmp1away, so the row is now the record of the constraint and the work moves to R-165 (with R-167 shipping in the same step). R-165 inherits this rank; it is the highest-ranked item that must land before any external install. - R-158 — the gap that makes R-163 dangerous: cross the size line and one page tells you. On its own it is a notification gap, not a silent failure, which is why it sits here and not higher.
- R-164 — blocked on a predicate, no customer impact today; it only becomes urgent if the unit size in R-163 is judged unacceptable, since the tar-drop is the cheapest way to halve it.
- R-161 — de-ranked 2026-08-02, ruled and shipped at reduced scope. The gate now has one
mandated entry point (
catalog_gates.py), which is the shape that actually gets run here. What is left is the automatic half, and that is sufficient while one person touches templates — so it ranks low by design, not by neglect. Revisit when a second does. - R-162 —
WATCHINGonly. A limitation that fails closed; revisit if a non-overlay driver ships.
R-159 and R-160 are SHIPPED and are not ranked; they are filed to record the class, and R-159's
class (an image VOLUME at an unmounted path) is still live — immich-server has one today.