R-97 collapsed to its shipped one-liner in ROADMAP and closed in OPEN-ITEMS.
PROMPT-TEMPLATE N.5 now names FOUR coupled artifacts instead of two: the
capability map, ROADMAP, the owning architecture doc (ruled as S-1 in CONTEXT.md
but never reflected in the template CC actually reads, so it bound nobody), and
OPEN-ITEMS.md. Tasks must now report which register rows they opened, closed or
re-ranked.
Ops: R-90 swap done (interim; CX33 still blocked), R-95 mitigation armed but zero
snapshots taken so it moves to WATCHING rather than closed, R-91 gate still not
satisfied. CONTEXT.md datastore path corrected to /mnt/pbs-datastore.
R-88 split: Part 1 (the failure breaker) SHIPPED in controller v0.176.0 and live
on both boxes; Part 2 (unknown != never) stays OPEN and is agent-side.
Phase 0 established the root cause at source: newestArchiveOn's (time.Time, bool)
signature cannot represent 'unknown', so a storage read ERROR collapses into a
positive 'no successful backup recorded yet'. The errored and genuine-never paths
are byte-identical on the wire, which is why Part 2 cannot be done controller-side.
R-97: the whole-guest backup tier has no failure signal to the hub at all —
internal/quiesce never imports internal/notify, so three failed backups and three
app-stack outages produced zero backup_failed events. Its only trace was a
customer-tier Hungarian app_start_failed for an app the backup itself had stopped.
Read-only triage found work that was agreed or discovered but never given an id:
R-95 restic offsite credential CAN delete — answers the parallel question R-89
raised and left open. Per-customer subaccounts report readonly=False, the
controller runs forget --prune from the box, and the sftp: backend cannot
express append-only. Storage Box snapshots (snapshot_limit=10, plan=null,
0 used) are server-side and SFTP cannot delete them — an unused zero-code
mitigation.
R-94 hub pins hostInstallVersion 1.19.0 while host-install ships 1.20.0, so a
hub-driven install still gets the pre-R-82 backup default.
R-90 ep0 has no swap at all and OOM'd today; gates R-86.
R-91 the pre-migration 13 GB datastore copy still occupies ep0's root disk.
R-92 PBS-DR gauge granularity. R-93 drill-r50 fixture tension.
R-96 two standing rules agreed in chat and never committed (the third, N.5's
third leg, IS committed at CONTEXT.md:8).
Open state was spread across ROADMAP, CONTEXT.md, four audits/, three runbooks,
per-session REPORT.md files and a chat log. This is the one page to read first:
every row has a state (BLOCKED/READY/WAITING-ON-OPERATOR/WATCHING) and an owner,
and the READY rows are ranked with reasoning.
R-88 is the recommended next task — quiesce's nil-age fail-open stops every app
stack every 5 minutes with no backoff and bypasses the maintenance window, and
its trigger (a PBS read failure) is live given ep0's demonstrated OOM.
CONTEXT.md now records that OPEN-ITEMS.md is authoritative and that REPORT.md is
overwritten per session.
Supervised runbook execution. No code, no version bump.
The felhom-pbs tier had reported `job errors` on EVERY demo-hp backup
since the tier was created on 07-26, while the data landed correctly
every time: `DatastoreBackup` grants Datastore.Backup but not
Datastore.Prune, so the box's keep_last=2 prune was denied.
Operator ruling: retention is a COMMERCIAL attribute owned by the hub;
ep0 executes. Box tokens therefore stay write-only - a compromised box
must not be able to delete its own offsite backups. No grant was widened
and felhom-tenantsync.sh is unchanged (the ruling makes it correct).
Increment 1:
- boxes stop attempting prune. allowPBSPrune is DERIVED
(`!t.Primary && t.KeepLast > 0`), so keep_last: 0 on the PBS tier
disables both the --prune-backups value and the gate in one config
edit, and the tier stays armed. Verified prune_pbs_allowed=false on
both boxes with no tier REJECTED line.
- per-namespace prune jobs on ep0, keep-last 2, daily 03:30 UTC
(05:30 CEST), dry-run gated. demo-hp 3->2, demo-felhom untouched,
chunk count unchanged (prune removes indexes, not chunks).
Write proof CLOSED: 08:25:47 job errors -> 09:37:29 TASK OK, snapshot
2026-07-27T09:37:29Z, chunks 9787->9813, prune step absent entirely.
Driven through POST /api/guest-backup/trigger (the UI path), not
--selftest and not raw vzdump. Hub gauge evidence explicitly NOT
satisfied - the delta is below its 0.1 GB display granularity.
GC scheduled sun 04:30 UTC and deliberately NOT run: every chunk still
carries a fresh atime from the migration copy, so a run today would
reclaim nothing. verify-new enabled per operator ruling, turning an
inert hub alarm live.
Legacy demo-felhom-01 namespace deleted with its two ACL entries and its
token (operator ruling, confirmed twice) so nothing dangles.
R-89 records the target architecture and carries the unanswered parallel
question: does the restic key on storage-box-pool-1 have DELETE rights?
If so the daily app-data tier has the identical exposure and append-only
is the equivalent answer.
ep0 is Etc/UTC, not CEST - corrected in the record.
Corrects two wrong severity readings with evidence from the box and the code.
The PBS outage was ~15 min (07:00-07:18 UTC), caused by a global OOM at 06:58:12:
proxmox-backup-proxy peaked at 3.2G on a 3.8G box and a concurrent 1.9G rsync
tipped it over. Root SSH to that box works from DooPlex via the public IP, not
from felhom-pve via the tunnel IP — the documented path I failed to try first.
R-88: internal/quiesce has NO failure limiter, backoff or breaker; the loop
stopped after three cycles only because PBS recovered. Verified additionally that
scheduledRunAllowed (quiesce.go:476-478) returns true whenever lastAgeSecs is nil,
so the same missing value that makes every poll due also bypasses the time-of-day
gate — the cycles ran outside the [04:30,08:30) window. Fixing the due-verdict
without fixing the nil-age bypass would leave the hole open.
Measured on demo-felhom while the offsite PBS service was down: the controller
re-polls /backup/due every ~5 min, still gets 'due' (storage unreachable + cold
store), and runs the FULL quiesce cycle each time — all four customer app stacks
stopped and restarted for a backup that cannot succeed. ~19 s of app downtime per
cycle, unbounded. The first entry called this bounded and event-only; it is an
availability fault.
Observed live on demo-felhom 2026-07-27 07:02:57 UTC: an agent restart while the
offsite PBS service was down produced a doomed vzdump at that tier. R-84's
read-error fallback to the in-memory record is correct alone but empty after a
restart, so "cannot read the storage" resolved to "no backup has ever been
taken" = due. Same class R-81 fixed in the hub, one layer down in the agent:
unreachable must be UNKNOWN, not resolved.
The unattended offsite restore-test on demo-felhom passed: 14.46 GB archive,
duration_s=635.07 (10m35s), then it rotated to the local tier. Persisted state
confirms the credit: {"felhom-pbs": "2026-07-27T06:14:42Z"}.
CORRECTION: I estimated ~2 hours for this restore. It took 10m35s. I derived
the estimate from a download rate measured during the FAILED attempt, which was
running under contention; the real link does ~1.4 GB/min. I then used that wrong
figure to raise a design concern — that the heavy-op gate would block backups
for hours on this box — which at 10 minutes largely evaporates. An estimate
extrapolated from a degraded measurement is not a measurement.
The SPEC's closing risk note is corrected in place, with the original left
visible for the lesson.
R-86 (NEXT, operator ruling 2026-07-27): backup-ALIGNED restore-test scheduling
— test a tier ~1 day after ITS OWN backup. R-85 schedules on a free-running
interval, which cannot express 'the day after the PBS backup': any fixed offset
drifts, so alignment would be luck. Shape: trigger from the tier's own last
successful backup rather than a clock. Interim in force: 302400s (3.5d), which
lands each tier ~weekly — the cadence half of the ruling, not the alignment half.
R-87: the restic app-data offsite tier is NEVER restore-tested. R-85 covers
whole-guest vzdump tiers only; the agent has no restic surface. That is arguably
the tier that matters most — the only one that survives losing the box AND
carries the customer's app data, since the whole-guest snapshot excludes the
bind-mounted drives. Exactly the state PBS was in before R-85.
REPORT.md: the full R-80 -> R-85 arc, including a section on the seven mistakes
I made and the two recurring shapes behind them (inferring behaviour from an
artifact instead of the code that consumes it; reading a result without its exit
code). Records demo-felhom's restore-test as IN FLIGHT at close, with the manual
recovery step if the deferred restart watcher does not complete.
Hub gate green (17 packages, rc=0).
- ROADMAP: R-85 row. Code SHIPPED; rotation NOT YET OBSERVED LIVE, stated as
such rather than written as done.
- Capability map: a new row for UNATTENDED restore-proof, IMPLEMENTED not
PROVEN-LIVE, kept distinct from the R-82 row that a MANUAL selftest earned.
That distinction is the same one the activation-vs-arrival split made.
- 03-host-agent §8: the scheduler covers every tier, oldest-proven first; the
spec is per-run; a restore-test joins the one-heavy-op gate. The safety
properties that must not be re-derived are listed.
- 07: restore-proof recorded as a per-tier property. Doc still NOT ratified.
- 06: corrects S4.1's 'the offsite restore-test now runs unattended' — it
silently stopped being true when local_backup_target was retargeted to 'local',
the SECOND time in that doc that a correct mechanism was broken by its input
changing underneath it.
- CONTEXT + REUSE.
Hub gate green (17 packages, rc=0).
Written against verified state, not assumption. Records three gaps the original
Phase 5 ordering does not cover:
1. hub v0.77.0 is COMMITTED BUT NOT DEPLOYED — manifest pins 0.76.0 and the pod
runs 0.76.0, so the R-85 signal exists only in git. The original Phase 5 only
mentions the agent.
2. R-85 has no ROADMAP row.
3. The agent CHANGELOG says v0.104.0-dev; an ldflags version disagreeing with
the CHANGELOG is the reconciliation problem hub 0.73.2 already caused.
One ordering correction: THE HUB GOES FIRST. Agent v0.104.0 makes the offsite
tier testable; hub v0.77.0 makes a failure audible. Agent-first means rotation
begins with nothing listening — two tiers able to fail silently instead of one,
which is the fault R-85 exists to end. Also drops the retired drill box, so the
rollout is demo-hp -> demo-felhom.
Names the phase's most likely SILENT failure: the new rotation state lands at
/var/lib/felhom-agent/restore-test-state.json and the agent is non-root. If that
is not writable, RecordSuccess warns and continues — a quiet return to one tier
being starved, not a crash.
Flags for operator judgement: demo-felhom's 14.46 GB offsite archive makes its
unattended restore-test a ~2h operation every other day, holding the heavy-op
gate throughout. Ruled when the only measured restore was demo-hp's 4 minutes.
R-84 resolved by asking the STORAGE rather than persisting the store: ground
truth, so a pruned archive correctly stops counting where a persisted record
would keep claiming a backup that no longer exists. Proven live on both boxes
with the in-memory store cold.
demo-hp's FIRST EVER offsite backup landed (4.25 GB) — the R-82 finding closed
on the box where it was worst. Controller v0.175.0 deployed to both boxes.
Slice D.1 — host-install 1.20.0: a FRESH box defaults to local-daily +
offsite-weekly (felhom-pbs, 604800s, keep_last=2). setdefault semantics proven
both ways: fresh gets the tier, an UPGRADE preserves the existing backup block
verbatim — so an in-place upgrade can never silently start writing to an
offsite datastore. Existing boxes are migrated explicitly.
Slice E:
- 07-backup-architecture.md: honest status header per CONTEXT ruling S-2, with
an explicit STALE-outside-the-PBS-tier verdict (the controller tiers were last
verified 41 controller versions ago). The PBS row claimed 'PBS on DooPlex'
(the retired spike store) with no cadence; it now names felhom-pbs ->
felhom-offsite on ep0 over wg-felhom, weekly, keep_last=2. NOT marked
ratified — that is Viktor's review of the section 10 list. Discharges R-83.
- 06-offsite-connectivity.md: the target-split remaining-work note collapsed
(shipped), and records HOW S4.1's tier-aware timeout silently regressed — the
mechanism was never removed, its INPUT changed when local_backup_target was
retargeted to 'local'. Also notes S4.1 already diagnosed the teardown 403 as a
phantom (a timeout consequence, not an ACL gap).
- capability map: new row for recurring offsite backups actually LANDING, as
distinct from the existing row proving ACTIVATION. IMPLEMENTED, not
PROVEN-LIVE — the restore round-trip has not completed under the fixed code.
- ROADMAP: R-82 SHIPPED with its remaining gate named, R-83 DISCHARGED, R-84
left open.
- CONTEXT + REPORT: the arc, including the mid-arc correction I had to make.
Third instance of one class (hub v0.12.0, v0.73.0, this), fixed as a class.
On 2026-07-26 03:00 UTC expected_backup_missed fired on demo-felhom, demo-hp
and drill-r50 at once; the demo-felhom one reached the CUSTOMER channel
claiming "newest backup is 176h0m0s old". Nothing was wrong — three vzdump
archives were on disk. Cause: the agent backup store is in-memory, so the
R-50 fleet restart emptied `backups` until the next run, and the hub read
empty as "no backup exists".
- assessBackupFreshness returns OK/UNKNOWN/MISSED instead of `missed bool`;
absence is UNKNOWN until it outlives an anchored window. Still pure.
- store.GetHostReportsSince + monitor.newestBackupEvidence read the hubs own
retained history (bounded 7-day lookback, early-exit on fresh evidence) —
"when did I last SEE evidence of a backup?" The anchor was free: the hub
already retains 90 days. No agent change, no new persisted state.
- store.GetFirstHostReportAt anchors absence at first contact, reusing the
existing 26h threshold as the grace (no new knob, the v0.73.0 shape).
- Deferrals logged + counted; reason strings kept distinct.
- backupStaleAfter untouched; landmine recorded (a weekly PBS snapshot would
alarm six days in seven) and owned by R-82.
Tests 493->508. Red-proofs A/B/C observed and restored; A reproduces the live
message verbatim. Replayed the real 03:00 reports (600/417/77 rows): all
three now silent.
Source: documentation/audits/DIAG-backup-missed-2026-07-26.md
The allowlist entry is REQUIRED, not cosmetic: handleEvent 400s an unknown
event_type, so controller v0.173.0's new drift alert would be silently inert
without it. Shipped with the controller that emits it.
Docs:
- RUNBOOK-local-api-endpoint-drift.md — how to repair a drift, including the
step everyone will want to skip (establish which value is CORRECT from what
the agent is actually bound to, rather than assuming bootstrap.json wins) and
what success looks like (SILENCE, not a "recovered" line, because a fresh
controller's healthy first observation is not logged). Records both
2026-07-26 repairs.
- ROADMAP: R-77 shipped; R-78 the local_api authority ruling, with the
clobber-a-working-channel risk spelled out in BOTH directions so it is not
resolved opportunistically; R-79 the whole-surface English-strings sweep;
R-80 expected_backup_missed, flagged as likely outranking R-77 because 7.3
days of stale backup materially exceeds the ~1.5-day channel outage, so the
causal link the DIAG hedged on cannot be the whole story.
- Capability map: note against the drive-wizard row (every agent-backed
capability rides this channel) that a silent drift class is now detected.
NO row status flips — detection is not prevention.
New documentation/controller/import-and-data-paths.md: the canonical import root
(and why it is NOT a registered StoragePath), the three data_paths roles, the
Fork-3 validation asymmetry, the class-driven copy rule, and the seven
invariants a future change must not break.
Capability map "File access via browser" — status DELIBERATELY UNCHANGED. The
drop-zone now has its own FileBrowser source and the app page carries a deep
link, both verified live, but nothing drove the FileBrowser HTTP UI (no browser
on DooPlex), so the row's standing "browse is exercised in no doc" caveat still
holds and PROVEN-LIVE remains unearned.
R-75 collapsed to its shipped one-liner. R-76 left open — this task does not fix
it, and nothing built here assumes an import/* directory stays 2775.
R-75 (spiked, GO) names the capability-map row it would flip: "File access via
browser" (00-capability-map.md line 96), currently IMPLEMENTED with the caveat
that browse/download through FileBrowser is exercised in no doc. Carries the
mandatory determinism constraint from P6 (sort + red-proof, or FileBrowser
force-recreates on every sync pass), the zero-removals invariant for
`documents`, the url.PathEscape-not-QueryEscape trap, and the four design forks
with evidence + recommendation, all awaiting operator ruling.
R-76 is minted for the two PRE-EXISTING defects the spike surfaced and
deliberately did not fix: FileBrowser Quantum creating 0644/0755 without
propagating setgid (breaking the shared-group chain one level below any
customer-created folder -- latent only because every userdata-touching app runs
uid 1000), and import/calibre living at 755 on demo-felhom where its same-app
sibling media/books is 2775.
Source: audits/SPIKE-catalog-data-paths-2026-07-26.md
- B2 demo-hp + B3 demo-felhom migrated to the island (agent 0.96.0), apps
served throughout (0 container restarts), island /storage 200, LAN DNS pinned
to the LAN IP, hub reports 0.96.0. No rollback.
- capability-map 'site/network change' row PARTIAL -> PROVEN-LIVE
- ROADMAP R-50 -> SHIPPED (fleet-migrated); add R-74 (island on Peti's cluster)
- nodes.md: both boxes island-bound, agent 0.96.0
Provisioned nested-PVE drill 'drill-r50' (qm300 on demo-hp) via the v1.25.0
nested-vm ISO through the real day-0, then ran the R-50 empirical spike:
- vmbr9 portless island bridge + guest island NIC hot-add (LAN undisturbed)
- F1 replay money shot: LAN move survives on the island; LAN-literal bind
reproduces the 2026-07-20 daemon-exit bug verbatim
- dnsmasq trap confirmed live + lan_resolver.host_ip fix proven
- pin address-independent (leaf SHA-256 unchanged, HTTP 200 over island)
- survival matrix: agent/guest/host-cold-reboot all return on the island
Docs: SPIKE verdict BLOCKED->GO, ROADMAP R-50 SPIKED->GO, nodes.md drill VM,
REPORT overwrite.
New capability-map row (IMPLEMENTED; §13 endpoint-level live on 9201). Records the
ruling: member accounts are superseded by the capability-URL guest share for
launcher sharing; per-member tile visibility parked under the SSO/members arc (R-15).
Updated the launcher row's member-coupling note and R-15 accordingly.
Virgin-ISO nested drill closed the train: dead-NIC install baked the
fallback (incl. the dead default gateway), the R-59 screen painted
(capture committed beside the spike doc), the cable move healed +
registered at the hub in 23s unaided, and the build's rootpw file
matched the installed box's shadow hash. R-59 SHIPPED with the recorded
deviation (first-boot gate; installer-initrd abort out of scope by
operator ack). R-60 SHIPPED (spike + drill cited; F-P9 route-flush fix
included). R-61 slice 1 SHIPPED. New R-62 row (hub delete-dialog
cosmetics, XS). Capability map: new PROVEN-LIVE row (nested != metal,
said so). Cleanup verified: felhom-pve interfaces byte-identical,
bridge/VMs/ISO removed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UuFPHmHNrCJj1VhY6QdDMU
Per the 2026-07-21 refresh brief: R-39 interim blocks (B4/E1) and the R-36
manual-Save block (C4) deleted — both shipped and proven live; freemail.hu
gate proven (R-4 COMPLETE); golden/floor-lift note now cites two shapes
(rehearsal + virgin HP t740 day-0 lift 0.153.0->0.156.0); A3 loader table
per operations/nodes.md (N100=mkimage/SB-off per record, HP t740=shim/SB
ENABLED); B2 multi-NIC cabled-port gotcha (R-59/R-60 pending); new A5 gate
(agent >=0.93.0 deployed box-side before the first escrow ceremony); D
offboarding pointer to §G (R-25b). DRAFT status and the C7 graduation gate
unchanged. ROADMAP R-25b pointer follows the rename.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UuFPHmHNrCJj1VhY6QdDMU
Found validating v0.69.0 against the live hub. demo-vm-felhom was deleted
on 07-18 and was still on the Customers list AND still raising offsite_stale
(10 events, latest 07-21 17:34, operator email at 19:34) — because
GetCustomers() is report-derived and no lifecycle tier ever deleted a report.
New leg 3 (residue), before the record purge: reports, app_telemetry,
app_log_tails, log_tail_requests, customer_notifications, plus the
credential-bearing appliance_registrations and selfbind_tokens. Audit
(events, notification_log) and F-14 provenance still survive.
Ghost customers are now deletable: 404 means "nothing here", not "no config
row". With no config row the offsite descriptor is unknowable, so the Hetzner
and descriptor legs record skipped_no_config rather than a bare "skipped".
Two more red-proofs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J55BQE1gE2V4ffud5jweGS
POST /configs/{id}/delete now runs hosts -> RESET -> purge behind three
acknowledgements, a typed customer-id, a stale-preview check and the
ONLINE-host refusal (every gate before any write, so a refusal has zero
side effects). The shallow handleConfigDelete is gone.
Two invariants are asserted, not just commented: ruling 3 is preserved by
construction (leg 2 never sees a host row) and retained escrow custody is
purged exactly once, in leg 3 (leg 2 runs with purgeEscrow=false).
handleCustomerReset's committed half was extracted as commitCustomerReset;
the standalone RESET path is byte-identical to v0.68.1 and its suite is
untouched. Five red-proofs run.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J55BQE1gE2V4ffud5jweGS
R-59 no-DHCP install must hard-abort (it baked 192.168.100.2 static and
completed - a box that can never call home). R-60 first-boot NIC sweep
self-heal. R-61 the baked root password must be knowable; a fixed well-known
password is explicitly rejected.
Positive evidence same-session: R-21 slice C PROVEN on a SECOND, virgin board
(HP t740) - and the shim loader booted with Secure Boot ENABLED, retiring the
assumption that Felhom installs need SB off. Fresh-box floor lift
0.153.0 -> 0.156.0 during day-0 cited on the publish-train row.
ISO README gains the t740 five-NIC trap: the 4-port igb card gets no lease,
the onboard r8169 port does.
R-58 records the operator ruling (2026-07-21) with the argument verbatim: the
installer should list available storage devices, excluding the install media,
and let one be selected. Third ISO mode alongside unattended-serial and
match-nothing-safety; unattended stays the appliance/factory mode. Slice 1 is
the abort-screen candidate table, same enumeration code, and it collapses the
two-boot dance on its own. Matters most for BYO/reinstall, where the serial is
unknown and a wrong guess is destructive.
ISO README gains the HP section: shim proven on this board by the safety boot,
the uncommitted-armed-profile pattern, verify-from-inside-the-ISO, and a
pointer to the prior-LVM abort that is the one likely failure on a
second-hand disk.
R-55's reboot leg ran operator-present on 9201: immich UI-stopped -> stayed
stopped across pct reboot, calibre-web recreated, zero alerts, ~15s.
R-57 records the lifecycle mechanism with the operator's abandoned-app
requirements verbatim and plant-it as the motivating case, including why the
retired/ directory move was wrong and the v0.158.1 pointer-receiver defect.
PROMPT-TEMPLATE: standard 'For the operator' plain-language section, mandatory
for M+ tasks and anything with a STOP.
ROADMAP rulings (operator, 2026-07-21): R-25b full-teardown cascade with three
acks + typed name (M-sized, spec to follow, no longer blocks R-3); R-11 channel
= direct Messenger, doc is the architect's; R-42 option (a), sidecars follow the
app; R-17 delete the archive - spike-lite found NO tooling verb targets it, so
it is an operator console action; R-4 complete (freemail.hu verified).
R-55 + R-41 slice 1 marked shipped; new R-56 (app difficulty classification -
the constructive half of the glance ruling).
scripts/build-hub.sh v1.23.0: the hub build script was outside any repo. Adopted
verbatim + versioned; the build-dir path is now a symlink to it.
felhom-testing skill: the ~1/5 recovery-code 'known flake' is retired - it was a
real defect the test was correctly detecting.
Dead primary: degraded in 13 s, exactly one app_start_failed, banner rendered and
self-cleared. Boot orphan: recovered in one attempt with zero alerts. Dead dhclient:
detected in 57 s on process liveness while the lease was still live, healed 120 s after the
kill — the tunnel never dropped, so the outage was prevented rather than observed.
P1 answered as a by-product: bookstack StartedAt == the moment bootrecon StartStack
returned, so unless-stopped did NOT resurrect it. F5 hypothesis confirmed.
New R-55, surfaced by the leg designed to prove the opposite: the boot bind gate recreates
and STARTS every deployed drive-backed app unconditionally, so a customer Stop does not
survive a reboot for those apps. Predates R-52 and does not implicate it, but it narrows
R-52's practical scope and needs a ruling.
R-51's roadmap diagnosis is corrected at the source: aggregation returned StateRunning
("partial") for a running/stopped mix, so the stack read RUNNING and IsDownState was never
consulted about at all — the constraint that row protects was never in tension
with the fix.
New R-54 row closes the INCIDENT-guest-dhclient-killed-2026-07-20 §5 OPEN RISK, and records
the design fact that makes it work: liveness of the DHCP client is itself a probe, because
the damage is timed and the address outlives its cause by 1-2 hours. The static-guest leg is
deliberately deferred to R-50.
New capability-map row is IMPLEMENTED, not PROVEN-LIVE: one leg is live (the watchdog's
healthy cycle on felhom-pve), the three that matter are destructive and operator-present and
have not run.
PROMPT-TEMPLATE §10 gains the seam-discipline row, including that a strings.Contains source
assertion is NOT sufficient — a commented-out call still contains the string.
The operator pressed Re-issue PBS credentials and the chain closed in 13 seconds. The
identical click on 2026-07-18 did nothing at all.
hub 08:39:31Z fresh mint, generation 0 -> 1; descriptor gains secret_generation: 1
(token_id + fingerprint BYTE-IDENTICAL — the invisible re-key shape)
agent 10:39:34 felhom-pbs-apply read felhom-pbs (leg b: the impossible read)
agent 10:39:38 ERROR REJECTED ... applied and DEAD, previous_state=applied
(leg c: the R-39 state, loud)
hub 08:39:45Z consumed_at stamped
agent 10:39:45 one-time token secret consumed (leg a: NO short-circuit)
agent 10:39:45 reconcile (set-only, no --server)
agent 10:39:47 pbsdr: converged state=applied
Corroboration: marker hash moved to afbb3b41… (it was byte-identical to the pre-reissue
marker in the failure); secret mtime 2026-07-18 -> 2026-07-21 10:39:45; new credential
probes 200; three consecutive reports trace applied -> auth_failed -> applied; ZERO
self-heal escalations, one mint, one consume, no consumed-failed.json — the box healed
through the descriptor path before the damper was ever needed.
Recorded for future runbooks: the operator first pressed the OFFSITE re-issue (two
distinct Re-issue actions exist). Harmless to PBS-DR, but it rotated the restic password
and correctly marked the escrow STALE, so the ceremony had to be re-run. Name the surface
explicitly next time.
R-39's three legs are closed and deployed: the hub stamps a monotonic secret_generation
so a re-key finally moves the descriptor hash; the wrapper gains a narrow read verb so
the non-root agent can read the credential it writes; and ProbeAuth turns a 401 into a
loud auth_failed the existing damper escalates to a fresh mint. Plus a consumed_at
honesty gauge for the applied-but-never-consumed disagreement.
Recorded in the R-39 row, because both are the kind of thing a future reader needs:
- A load-bearing fact the spec did not flag, checked rather than trusted: Apply bails out
if the storage status probe ERRORS and adopt converges without consuming when the
storage reads active, so the fix depended on PVE's 401 behaviour. PVE's storage_info
wraps activation in eval{} and leaves active=0, so a 401 returns HTTP 200 with
active:0 — never an API error. The chain is sound by proof, not inference.
- A defect I shipped and caught: v0.91.0 built the probe seam and main.go never wired it,
so the leg was inert while every test passed. Same class as controller v0.154.0 the day
before. Fixed in v0.91.1 (artifact superseded, not overwritten); v0.91.2 made a healthy
probe observable so "no auth_failed" can never again be confused with "never probed".
The DR-tier capability row is deliberately NOT upgraded to PROVEN-LIVE: the decisive
evidence is STOP-2, the operator pressing Re-issue and the box converging where the
identical click did nothing on 2026-07-18.
R-50b(a) shipped — wrapper sha256 in the manifest + agent reporting + host drift surface,
with unknown-on-either-side reading as quiet rather than drift. (b)/(c) remain open: the
wrapper is still fetched unversioned from raw/branch/main.
The operator moved the global floor to a version the box did NOT run (0.153.0 ->
v0.154.0) and the managed self-update fired exactly once:
06:57:13Z UpdateState pending, initiated_by=auto-floor
06:57:17Z agent: controller swap requested 0.153.0 -> 0.154.0
06:57:21Z container restarted
06:57:29Z agent: new controller healthy (16 s save -> healthy)
Over a 39-minute window: swap requests 1, agent-driven bootstrap restarts 1,
rollbacks 0, container RestartCount 0. VerifyStartup confirmed on the next boot;
the following periodic check logged "Current version 0.154.0 is up to date" —
the at/above-floor branch correctly doing nothing.
The 2026-07-20 attempt proved nothing because it targeted an already-running
version; that was the whole reason this leg stayed open.
Disclosed in both rows: a hand-deploy of v0.155.0 at 07:17:10 falls inside the
observation window and is what StartedAt shows afterwards. It never goes through
SwapController, so the swap-count assertions hold across the full window — and it
incidentally re-confirmed the at/above-floor branch (0.155.0 running against a
0.154.0 floor -> updater did nothing).
R-23(b) (cosmetic Waiter "recovered" log timing) remains open.
R-48 — the offsite restore controls collapse to one „Visszaállítás…" entry per app plus
a per-app wizard with three described intent cards. Shipped in controller v0.154.0
(3a9d744). Live click-through still pending the operator's floor save.
R-39 — the planned v0.90.1 artifact publish was CANCELLED as a false signal (operator
ruling 2026-07-21). 9596d5a changes zero non-test Go files; its own message says "the Go
binary is unchanged". The fix is the felhom-pbs-apply wrapper, which felhom-pve has
carried since 2026-07-18 and which every new install fetches from raw/branch/main
regardless of binary version. Publishing would have delivered no behaviour change and
advertised a versioned fix the artifact channel never carried.
R-50b (new) — that stop surfaced the real defect: a root-owned privileged host artifact
is delivered unversioned from main, absent from the Day-0 manifest, so the fleet has no
way to answer which wrapper a given host is running.
Capability map — the destroy-then-recover drill ran through the customer UI:
photos deleted, TRASH EMPTIED, full files+database restore. 40 files placed
against 6 in the earlier non-destructive run, 1 DB dump replayed rc-0, 11
assets active, no drift, timeline confirmed. That is the proof the 6D
downgrade asked for, so the offsite-restore row earns PROVEN-LIVE. The
customer-restore row records the honest residual: an operator ran it, so the
row's literal 'a customer, not the operator' wording still owes one pass.
ROADMAP R-23(a) — the STOP-2 floor save released the held wait in the SAME
SECOND (hub 18:56:27 CEST = controller 16:56:27Z), out-of-cycle report 2s
later, generation advanced 0 -> 1. Still open: the self-restart single-fire
leg, since the floor was set to a version the box already ran.
Trap recorded: the wake is logx.Debugf, so it is invisible in docker logs at
INFO and lives only in the debug ring.