rulings 87-89 recorded before the work (live-restore ON; crash restart with a limit; versions visible in the hub); R-852 filed (versions shown nowhere)
gates / gates (push) Successful in 30s
gates / gates (push) Successful in 30s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -16,6 +16,10 @@
|
|||||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||||
|
|
||||||
|
|
||||||
|
> **Rulings 2026-10-04 (~15:17) — recorded before the work (System page + Docker lane + crash restart brief).** `09` §3
|
||||||
|
> decisions **87** (Docker `live-restore` ON fleet-wide, never off by a plain restart), **88** (a crashed host restarts by
|
||||||
|
> itself, with a limit — R-851) and **89** (the operator sees the boxes' OS / Proxmox versions in the hub — R-852).
|
||||||
|
|
||||||
> **2026-10-04 (late afternoon) — OS updates: host fast lane + fleet view + alarms BUILT (`11` §8.2–§8.3); the tunnel
|
> **2026-10-04 (late afternoon) — OS updates: host fast lane + fleet view + alarms BUILT (`11` §8.2–§8.3); the tunnel
|
||||||
> status is true (R-841).** Agent v0.141.0 → v0.141.1, hub v0.131.0 → v0.131.1, controller v0.292.0 (cloudflared
|
> status is true (R-841).** Agent v0.141.0 → v0.141.1, hub v0.131.0 → v0.131.1, controller v0.292.0 (cloudflared
|
||||||
> readiness health check). **Decided by CC unattended — operator may reverse:** `09` §3 decisions **84** (appliance proof
|
> readiness health check). **Decided by CC unattended — operator may reverse:** `09` §3 decisions **84** (appliance proof
|
||||||
|
|||||||
@@ -778,6 +778,16 @@ its length, and both fixes cost something the household would notice — operato
|
|||||||
83. **Next: the tunnel status (R-841), the host fast lane (`11` §8 step 3), and other OS-update improvements.**
|
83. **Next: the tunnel status (R-841), the host fast lane (`11` §8 step 3), and other OS-update improvements.**
|
||||||
*Operator ruling 2026-10-04 ~12:20.*
|
*Operator ruling 2026-10-04 ~12:20.*
|
||||||
|
|
||||||
|
### 2026-10-04 (~15:17) — three operator rulings (recorded before the work)
|
||||||
|
|
||||||
|
87. **Docker `live-restore` is ON for every box** (`11` §5.8, option A). It is turned on once — by the golden and by a
|
||||||
|
one-time step on each installed box — and never turned off by a plain restart (R-835). *Operator ruling 2026-10-04
|
||||||
|
~15:17.*
|
||||||
|
88. **A crashed host restarts by itself, with a limit** (R-851): *"Yes, but maybe not indefinitely."* After a limit of
|
||||||
|
crashes in a short time the box stays off and the operator is told. *Operator ruling 2026-10-04 ~15:17.*
|
||||||
|
89. **The operator must see the boxes' OS and Proxmox versions in the hub** (*"Where should I be able to see the
|
||||||
|
OS/Proxmox versions of the boxes?"* — today: nowhere; R-852). *Operator ruling 2026-10-04 ~15:17.*
|
||||||
|
|
||||||
### 2026-10-04 (afternoon) — decided by CC unattended — operator may reverse (host fast lane brief)
|
### 2026-10-04 (afternoon) — decided by CC unattended — operator may reverse (host fast lane brief)
|
||||||
|
|
||||||
84. **Which record proves a box is an appliance, for the host fast lane?** Options: (a) `agent.json`
|
84. **Which record proves a box is an appliance, for the host fast lane?** Options: (a) `agent.json`
|
||||||
|
|||||||
@@ -358,7 +358,7 @@ stopping line that lies.
|
|||||||
| **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC |
|
| **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC |
|
||||||
| **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
|
| **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
|
||||||
|
|
||||||
## Hub & operator — 24 rows (P2 1, P3 7, P4 16)
|
## Hub & operator — 25 rows (P2 1, P3 8, P4 16)
|
||||||
|
|
||||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||||
|---|---|---|---|---|---|---|---|
|
|---|---|---|---|---|---|---|---|
|
||||||
@@ -386,6 +386,7 @@ stopping line that lies.
|
|||||||
| **R-844** | Hub & operator | P4 | **The household's OS-update line exists only on the hub's customer timeline.** 2026-10-04: the box itself has no event surface for agent results (the controller UI shows no timeline), so `os_update_applied` is a hub customer event (info: recorded, never mailed). Its stored text is the hub's English sentence; the hu/en bundle text (`mail.event.os_update_applied`) is used only if it is ever mailed. Fix direction: a controller-side line (the controller already polls the agent's local API) when the box gets a household timeline. `audits/os-guest-lane-2026-10-04/partG/hub-customer-timeline-demo-hp.txt` | **READY — owner: CC** | — | — | CC |
|
| **R-844** | Hub & operator | P4 | **The household's OS-update line exists only on the hub's customer timeline.** 2026-10-04: the box itself has no event surface for agent results (the controller UI shows no timeline), so `os_update_applied` is a hub customer event (info: recorded, never mailed). Its stored text is the hub's English sentence; the hu/en bundle text (`mail.event.os_update_applied`) is used only if it is ever mailed. Fix direction: a controller-side line (the controller already polls the agent's local API) when the box gets a household timeline. `audits/os-guest-lane-2026-10-04/partG/hub-customer-timeline-demo-hp.txt` | **READY — owner: CC** | — | — | CC |
|
||||||
| **R-848** | Hub & operator | P4 | **A held host package is invisible to the hub.** MEASURED 2026-10-04 on demo-hp (undo runbook proof): with `tzdata` held after a by-hand undo, the wrapper's `pending` stayed 78 — apt's simulation leaves held packages out, so the fleet view shows nothing and no alarm can see a hold that was forgotten. The hold lives only in the incident's register row (`runbooks/os-updates-host-undo.md`). Fix direction: the wrapper reports `apt-mark showhold` and the fleet line shows it. | **READY — owner: CC** | — | — | CC |
|
| **R-848** | Hub & operator | P4 | **A held host package is invisible to the hub.** MEASURED 2026-10-04 on demo-hp (undo runbook proof): with `tzdata` held after a by-hand undo, the wrapper's `pending` stayed 78 — apt's simulation leaves held packages out, so the fleet view shows nothing and no alarm can see a hold that was forgotten. The hold lives only in the incident's register row (`runbooks/os-updates-host-undo.md`). Fix direction: the wrapper reports `apt-mark showhold` and the fleet line shows it. | **READY — owner: CC** | — | — | CC |
|
||||||
| **R-849** | Hub & operator | P4 | **The fleet view's GUEST "reboot needed since" never clears.** 2026-10-04: the guest is scanned only after an install (R-845, one `pct exec`), so a guest restart is never seen; the guest line keeps the date of the last install that said "needed". No alarm reads the guest line (the reboot alarm is host-only), so it is display only. Fix direction: scan the guest on every pass too (one `pct exec`, ~1 s) or hide the guest date. `audits/os-host-lane-2026-10-04/partC/live/fleet-during-ring1-test.json` | **READY — owner: CC** | — | — | CC |
|
| **R-849** | Hub & operator | P4 | **The fleet view's GUEST "reboot needed since" never clears.** 2026-10-04: the guest is scanned only after an install (R-845, one `pct exec`), so a guest restart is never seen; the guest line keeps the date of the last install that said "needed". No alarm reads the guest line (the reboot alarm is host-only), so it is display only. Fix direction: scan the guest on every pass too (one `pct exec`, ~1 s) or hide the guest date. `audits/os-host-lane-2026-10-04/partC/live/fleet-during-ring1-test.json` | **READY — owner: CC** | — | — | CC |
|
||||||
|
| **R-852** | Hub & operator | P3 | **The operator cannot see any box's Debian, Proxmox, kernel or Docker version anywhere.** FOUND 2026-10-04 (operator question: *"Where should I be able to see the OS/Proxmox versions of the boxes?"*): no box reports them — the agent's host report carries no Debian version, no `pveversion`, no running or next-boot kernel, no guest Debian or Docker engine version (`felhom-agent/internal/proxmox/types.go:35` has a `PVEVersion` field that nothing reads); the OS fleet view is `GET /os/fleet` JSON only, with no page, no tab and no buttons for ring, switch or "approve now"; its "release" is Felhom's OS release id, not a Debian or Proxmox version. `09` §3 decision 89. | **READY — owner: CC (this session, Part A)** | — | — | CC |
|
||||||
|
|
||||||
## Business & legal — 7 rows (P2 4, P4 3)
|
## Business & legal — 7 rows (P2 4, P4 3)
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user