09 §3 decisions 172-173 (kernel lane yes, option A + C; R-32 leftovers left, P4) recorded before the work
gates / gates (push) Successful in 3m10s
gates / gates (push) Successful in 3m10s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -915,6 +915,19 @@ its length, and both fixes cost something the household would notice — operato
|
||||
127. **The agent's three by-design abilities (`03` §3.1) stay for now**; revisited before the first paying customer.
|
||||
*Operator ruling 2026-10-05.* (R-861)
|
||||
|
||||
### 2026-10-07 (14:43) — operator rulings on the kernel lane and R-32 (recorded before the work)
|
||||
|
||||
172. **The kernel lane: yes — design option A (the one-shot flag on the ESP), with option C on the one-shot entry**
|
||||
(`audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`). A box may restart at night for a kernel update. **The
|
||||
household is told by mail the day before**, with what to do if the box is not back by morning (unplug it, plug it
|
||||
back in). Each kernel set is approved by the operator after ring 0 ran it, as the Docker and Proxmox sets. **A
|
||||
frozen new kernel needing a person's power cycle is accepted** for the first customers, because the household is
|
||||
told what to do. *Operator ruling 2026-10-07 14:43.* (R-836)
|
||||
173. **R-32's old off-site leftovers (about 2.7 GB): leave them for now.** They were computed (box total minus live
|
||||
customers), never listed; deleting them needs the pool box's main login, which reaches every customer's data.
|
||||
R-32 stays open only for the leftovers, re-ranked P4, blocked on „a safe main-account path". This supersedes the
|
||||
„delete the old RESET leftovers once" half of decision 167. *Operator ruling 2026-10-07 14:43.* (R-32)
|
||||
|
||||
### 2026-10-07 (07:58) — operator rulings on the night designs (recorded before the work)
|
||||
|
||||
162. **Hub releases at night: NO.** A hub image build or hub deploy happens only in a session the operator attends.
|
||||
|
||||
@@ -230,3 +230,70 @@
|
||||
12:19:45 302
|
||||
12:19:50 302
|
||||
12:19:55 302
|
||||
12:20:01 302
|
||||
12:20:06 302
|
||||
12:20:11 302
|
||||
12:20:16 302
|
||||
12:20:21 302
|
||||
12:20:26 302
|
||||
12:20:31 302
|
||||
12:20:36 302
|
||||
12:20:41 302
|
||||
12:20:46 302
|
||||
12:20:51 302
|
||||
12:20:56 302
|
||||
12:21:01 302
|
||||
12:21:06 302
|
||||
12:21:11 302
|
||||
12:21:16 302
|
||||
12:21:21 302
|
||||
12:21:26 302
|
||||
12:21:31 302
|
||||
12:21:36 302
|
||||
12:21:41 302
|
||||
12:21:46 302
|
||||
12:21:51 302
|
||||
12:21:56 302
|
||||
12:22:01 302
|
||||
12:22:06 302
|
||||
12:22:11 302
|
||||
12:22:16 302
|
||||
12:22:21 302
|
||||
12:22:26 302
|
||||
12:22:31 302
|
||||
12:22:36 302
|
||||
12:22:41 302
|
||||
12:22:46 302
|
||||
12:22:52 302
|
||||
12:22:57 302
|
||||
12:23:02 302
|
||||
12:23:07 302
|
||||
12:23:12 302
|
||||
12:23:17 302
|
||||
12:23:22 302
|
||||
12:23:27 302
|
||||
12:23:32 302
|
||||
12:23:37 302
|
||||
12:23:42 302
|
||||
12:23:47 302
|
||||
12:23:52 302
|
||||
12:23:57 302
|
||||
12:24:02 302
|
||||
12:24:07 302
|
||||
12:24:12 302
|
||||
12:24:17 302
|
||||
12:24:22 302
|
||||
12:24:27 302
|
||||
12:24:32 302
|
||||
12:24:37 302
|
||||
12:24:42 302
|
||||
12:24:47 302
|
||||
12:24:52 302
|
||||
12:24:57 302
|
||||
12:25:02 302
|
||||
12:25:07 302
|
||||
12:25:12 302
|
||||
12:25:17 302
|
||||
12:25:22 302
|
||||
12:25:27 302
|
||||
12:25:32 302
|
||||
|
||||
@@ -149,7 +149,6 @@ stopping line that lies.
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **R-32** | Backup & restore | P2 | **[P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible.** The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's `"hetzner":"ok"` leg destroys the sub-account, but **a Hetzner sub-account is an access-control object, not a data object** — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-32.md`) — pick A: purge the repo + set-aside copies through the sub-account's own password login BEFORE deleting it (no main-account credential); the „repo data destroyed" comment at `hub/internal/offsite/offsite.go:276-280` is false for the shared tier; part (3) waits on a read-only measurement (`du` on port 23). Operator: agree to route A. **2026-10-07 07:58: `09` §3 decision 167 — option A yes, and delete the old RESET leftovers once.** **2026-10-07 (Part D step 2, read only):** the hub shows bx11 holding 3.3 GB of data while the four live customers account for ~0.6 GB → **~2.7 GB belongs to no live customer**; no live customer shows a set-aside folder, so the leftovers are in homes whose sub-account is gone, reachable only by the pool box's MAIN account. **The one-time clean-up is STOPPED:** no main-account login for bx11 is held on DooPlex (the API token in the credentials file sees no box). Needs: the main-account login, or the operator lists and deletes in the Hetzner console (`audits/day-2026-10-07/D/D2-listing-readonly.txt`). **2026-10-07: option A DELIVERED (hub 0.142.0):** RESET purges the repo and every `.orphaned-*` through the sub-account's own login before deleting it, and keeps the sub-account if the purge fails (red-proved, `audits/day-2026-10-07/D/`). Not yet seen live (no scratch RESET today). **The one-time clean-up of the ~2.7 GB of leftovers is STOPPED:** it needs the pool box's main-account login (see above). | — | **Ruling from the run (three parts, deliberately separate):** (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable **BY DESIGN** → RESET gains a **main-account purge of the customer base dir** (the existing operator ack already covers it); (2) the **move-aside guard STAYS** for reinstall-*without*-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator **Restic tab shows per-customer directory bytes vs attributed snapshot bytes**, so dead data cannot hide. Measured on the pool box that night: **49 M attributed** (2 snapshots, 48.717 MiB) against **1.4 G + 3.0 M unattributed** across TWO `.orphaned-*` dirs. Evidence `restic-and-pool.txt` | CC |
|
||||
| **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-05 — owner Viktor.** (b) partly: the hub database now leaves DooPlex nightly, encrypted, to ep0 (R-173); everything else in DooPlex's backup still stays on the box. (a) partly: the hub copy alarms through Prometheus (`HubDBBackupStale`); `notify_failure` is still a no-op for the rest. (c)–(h) unchanged. **READY** for the rest | — | — | operator |
|
||||
| **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY (L) — NEW 2026-08-12, RANK 1** | R-198, R-199, R-224, R-241 | Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. **Until one of those, the honest position is that retention is an operator-only capability.** At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked | operator + CC |
|
||||
| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. **2026-10-06 (night), from Part C:** demo-hp's off-site tier was NOT overdue — its last copy is 2026-10-01 20:15Z (ep0's listing, verify ok), so with the 7-day cadence it is due ~2026-10-08; the night of 2026-10-06→07 is most likely local-only on both demo boxes (demo-felhom's off-site landed 2026-10-06 04:21Z). The two-tier night under the new rule is then ~2026-10-08 on demo-hp. **2026-10-07 (morning): the local-tier night and one press READ BACK** (`audits/readback-2026-10-07/RESULT-B-D.md`): the night stop on demo-hp (9 apps) was ~91 s (was 5 min 47 s), demo-felhom (1 app) ~11 s; one press on demo-hp: 80 s from press to the last app (per app 39–79 s), the copy finished 4 min later with the apps running, only the local tier ran; the page's „kb. 1–1,5 perc" holds. Two channels each (controller log + agent journal / container StartedAt + a 5-s HTTP sampler). **Left:** the first night with both tiers due on demo-hp (~2026-10-08). | — | Read back the ~2026-10-08 night (both tiers on demo-hp); then close | CC |
|
||||
@@ -179,6 +178,7 @@ stopping line that lies.
|
||||
| **R-816** | Backup & restore | P4 | **No off-site failure class has ever been seen live.** F-DIAG (controller v0.182.0, 2026-07-28) split off-site failures into six causes — quota, orphaned, no_repo, no_units, transport, unknown — each with its own Hungarian message. None of the six has been exercised by a real failure on a box; the recovery inventory records it only as a known limit (`documentation/architecture/_recovery-inventory-2026-07-28.md:955`). Filed 2026-10-03 from the F-DIAG row's residue when that row moved to `CLOSED-ITEMS.md`. | **READY — filed 2026-10-03 (triage); owner: CC.** Exercise each class once on a scratch guest (a full quota, a missing repository, a blocked transport) and read the message the household sees. | — | — | CC |
|
||||
| **R-832** | Backup & restore | P4 | **ep0's copy in a place outside both Hetzner and the operator's home (roadmap).** Today DooPlex (the operator's home) holds it (decision 71). A Hetzner Storage Box would share a provider with ep0 and with every household's file backups, and cannot run PBS, so the copy could not be verified or restored from directly. | **DEFERRED — later, if the product grows** | — | — | operator |
|
||||
| **R-878** | Backup & restore | P4 | **A catch-up (R-871) runs the database-dump leg in the DAY, and that leg stops an app with a volume for its copy — the household may notice the stop, and a large volume makes it longer.** MEASURED 2026-10-05 on demo-felhom: the catch-up at 08:25:02 stopped opengist, copied 182.5 KB, started it again — about 1 s, then a few seconds of `health: starting`; the night does exactly the same, unseen. Nothing measured for a large volume. Fix direction (if it matters): skip the volume copy of a running app in a DAYTIME catch-up and leave it to the next night, or warn. `audits/catchup-2026-10-05/partA/live-demo-felhom.txt` | **READY — owner: CC** **2026-10-05 (burn-down night): NEEDS A LIVE MEASUREMENT** — a timed catch-up on 9202 for an app with a multi-GB volume, then the skip-or-warn choice. | — | measure a large volume first | CC |
|
||||
| **R-32** | Backup & restore | P4 | **[P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible.** The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's `"hetzner":"ok"` leg destroys the sub-account, but **a Hetzner sub-account is an access-control object, not a data object** — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-32.md`) — pick A: purge the repo + set-aside copies through the sub-account's own password login BEFORE deleting it (no main-account credential); the „repo data destroyed" comment at `hub/internal/offsite/offsite.go:276-280` is false for the shared tier; part (3) waits on a read-only measurement (`du` on port 23). Operator: agree to route A. **2026-10-07 07:58: `09` §3 decision 167 — option A yes, and delete the old RESET leftovers once.** **2026-10-07 (Part D step 2, read only):** the hub shows bx11 holding 3.3 GB of data while the four live customers account for ~0.6 GB → **~2.7 GB belongs to no live customer**; no live customer shows a set-aside folder, so the leftovers are in homes whose sub-account is gone, reachable only by the pool box's MAIN account. **The one-time clean-up is STOPPED:** no main-account login for bx11 is held on DooPlex (the API token in the credentials file sees no box). Needs: the main-account login, or the operator lists and deletes in the Hetzner console (`audits/day-2026-10-07/D/D2-listing-readonly.txt`). **2026-10-07: option A DELIVERED (hub 0.142.0):** RESET purges the repo and every `.orphaned-*` through the sub-account's own login before deleting it, and keeps the sub-account if the purge fails (red-proved, `audits/day-2026-10-07/D/`). Not yet seen live (no scratch RESET today). **The one-time clean-up of the ~2.7 GB of leftovers is STOPPED:** it needs the pool box's main-account login (see above). **2026-10-07 14:43 — `09` §3 decision 173: leave the ~2.7 GB of old leftovers for now; re-ranked P2 → P4.** Option A itself is delivered (hub 0.142.0). What stays open is ONLY the leftovers: they were computed (box total minus live customers), never listed, and deleting them needs the pool box's MAIN login, which reaches every customer's data. **Blocker: a safe main-account path.** | — | **Ruling from the run (three parts, deliberately separate):** (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable **BY DESIGN** → RESET gains a **main-account purge of the customer base dir** (the existing operator ack already covers it); (2) the **move-aside guard STAYS** for reinstall-*without*-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator **Restic tab shows per-customer directory bytes vs attributed snapshot bytes**, so dead data cannot hide. Measured on the pool box that night: **49 M attributed** (2 snapshots, 48.717 MiB) against **1.4 G + 3.0 M unattributed** across TWO `.orphaned-*` dirs. Evidence `restic-and-pool.txt` | CC |
|
||||
|
||||
## Storage & devices — 6 rows (P3 4, P4 2)
|
||||
|
||||
@@ -214,7 +214,7 @@ stopping line that lies.
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. | — | — | CC |
|
||||
| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. **2026-10-07 14:43 — `09` §3 decision 172: the kernel lane is YES — option A (the ESP one-shot flag) with option C (lockup → panic) on the one-shot entry; a box may restart at night; the household is mailed the day before (what to do if it is not back: unplug, plug back in); each kernel set is operator-approved after ring 0; a frozen new kernel needing a power cycle is ACCEPTED for the first customers. BUILD STARTED 2026-10-07 (the kernel-lane arc).** | — | — | CC |
|
||||
| **R-862** | Box system & updates | P3 | **Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its `felhom-os-apply` (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it.** FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; `felhom-op` cannot become root). The steps: `runbooks/config-bundle.md` "Tester 2". Then CC signs the bundle and reads it back. | **WAITING-ON-OPERATOR** — the operator said (2026-10-04 ~19:05) he will try through his tunnel **2026-10-07 07:58 (`09` §3 decision 170): the operator believes Tester 2 never set up the off-site backup (no escrow); the laptop may stay off for days; the tester plans to add an HDD.** **2026-10-07 (Part H, read only):** in Tester 2's last 10 notifications (2026-10-05 09:16 → 2026-10-07 06:58) the household got 0 mails, the operator 10 (host_down ×9, expected_dbdump_missed ×1); control: the household channel shows on the other three customers. So the household is not mailed daily while the box is off; no row (`audits/day-2026-10-07/H/`). | the operator's WireGuard tunnel | the operator runs the three bootstrap commands; CC sends the bundle | operator |
|
||||
| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC |
|
||||
| **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |
|
||||
|
||||
Reference in New Issue
Block a user