DRILL new household on golden 0.282.0: 0 interventions; ready for a first real tester on a new record
gates / gates (push) Successful in 28s
gates / gates (push) Successful in 28s
Audit doc, STATUS one sentence, capability-map first-hour row (walk re-proven, day-one off-site sentence narrowed: R-720/R-726/R-727), teardown in four layers, R-600 measured again. Secret scan over all audits: 0 hits, positive control 1. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -757,7 +757,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-597** | **[P2-MEDIUM] The setup code is three Hungarian words, inside an otherwise fully English e-mail, sent to a household the hub knows is English.** FOUND 2026-09-20 by the slice-6 drill. The mail is English end to end (slice 3 working); the code it carries was **`képző-szkítia-ásatás`** — 20 characters, 3 words, **5 of them outside ASCII** (ő, í, á×2, é). An English speaker must copy three words they cannot read, spell or say aloud, and type them into a box on a keyboard that has no ő. They can paste — until the day they read the code to someone over the telephone, which is precisely what a three-word code is FOR. **The same generator feeds the recovery code (10 words) and the owner passphrase (5 words)**, so the fault is one wordlist wide, not one mail wide: this walk saw the passphrase too and it is Hungarian. **Fix shape:** an English wordlist chosen per `customer.language`, with the same word count and the same entropy, and a test that pins BOTH lists' entropy and that no word in either needs a character outside the reader's keyboard. **Not a rename of the existing words** — a second list. **CLOSED 2026-09-21, hub v0.119.0.** **One third of the row was wrong: the recovery code was never Hungarian.** `felhom-agent` mints it (`internal/escrow`) from the **EFF large wordlist** and always has — ten English words, ≈129 bits. The hub does not own that secret and no row was opened for it: a second definition here is the drift `backupTargetAbsentText` already demonstrates across two repos. The two the hub DOES mint now follow the household: setup code 3 hu words (44.6 bits) → **4 en words (51.7)**, owner passphrase 5 hu (74.3) → **6 en (77.5)**, list and count chosen together by `RandomPassphraseFor(lang, use)` so a caller cannot pair an English list with a Hungarian count. **The floor is computed from the embedded lists at test time, not compared with a constant** — red-proofed at 3 English words (38.77 vs 44.56). Hungarian is byte-unchanged, and the list length is pinned so a swap cannot move it quietly. **The task's proposed "read it over the phone" filter was MEASURED and NOT adopted** — it removes 5270 of 7772 words (68%, 12.92 → 11.29 bits/word) and would make this list stricter than the one the product already uses for the code a household writes on paper during a disaster; what it reached for is kept as an assertion (`TestEnglishListIsTranscribable`: 3-9 lower-case ASCII letters, no digit, no separator). Decision recorded in source, **operator may reverse**. Also: **no claim mail ever stated a word count** — the only count wording was the bind page's passphrase hint, whose English half is now count-free. | **CLOSED 2026-09-21 — hub v0.119.0** |
|
||||
| **R-598** | **[P2-MEDIUM] The Backup page's two protection warnings — the ones that say whether the household's files are safe — are Hungarian on an English dashboard.** FOUND 2026-09-20 by the slice-6 drill on a fresh box, and confirmed on the demo box. Of 73 lines on `/backups` exactly four are Hungarian: **„Csak egy másolat készül (nincs második meghajtó) — a 3-2-1 mentéshez csatlakoztasson egy második meghajtót vagy offsite tárolót"**, **„A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok ellen véd, lemezhiba ellen nem"**, and the two backup-target names **„Helyi tároló (local)"** and **„Biztonsági szerver – külön hardver (PBS)"**. They come from `internal/web/backup_handlers.go` (12 Hungarian literals) and `internal/web/backup_target_offer.go` (10) — again composed sentences handed to the page, the R-573/R-596 shape. **It matters more than its line count:** those two warnings are the only place the product tells a household that one copy on one disk is not protection, and the volunteer guide's §9 sends every tester to exactly this page to read exactly these two sentences. **Fix shape:** keys + args for both files, with the retrieval-promise gate run over the English (these sentences are about what a backup does and does not protect). **CLOSED 2026-09-21, controller v0.259.0.** The row's count of `backup_handlers.go` was 12; **nine are code and three are Hungarian inside COMMENTS**. The offer file's ten is right. `degradedMessageFor` now returns a **KEY** — the decision stays language-free and in one place, the words are chosen by the caller that knows the reader — and `buildTierViews` / `backupTargetLabel` / `loadGuestBackup` take the language the way `buildDataPathCards` already did. **The English is asserted to carry the same NEGATION the Hungarian does** ("protects against corrupted files, **but not** against a disk failure"); an English sentence that promised disk-failure protection would be worse than leaving it Hungarian. **Proven LIVE on guest 9201** for the two tier names ("Local storage (felhom-backup)", "Backup server – separate hardware (PBS)"); **the two warnings themselves were NOT walked live** — that box is healthy and a healthy box renders nothing by design, and producing the state would mean un-assigning a live backup target. They are covered by render tests through the real handler in both states. **An apostrophe cost a render:** the first English absent-drive sentence never matched because `html/template` escapes `'` to `'` — caught by the test, not by review. | **CLOSED 2026-09-21 — controller v0.259.0; the two warnings proven by render test, not live** |
|
||||
| **R-599** | **[P3-LOW] A drill's teardown is blocked for 30 minutes by design, and nothing says so.** FOUND 2026-09-20 tearing the slice-6 drill down. VM destroyed at 17:12Z; the hub then refused **both** `POST /configs/<id>/delete` (409, *host … is ONLINE*) and the host delete (`deletable:false`) — correctly, because an online host would receive permanent 401s. But "online" is not a liveness probe: it is a **report-staleness window**, and the window is **45 minutes** — `manifests/hub.yaml` sets `alerting.stale_threshold: "45m"`, which `hostStatus()` reads (`ok` under it, `stale` over, `down` at 2x). A machine that no longer exists therefore reads ONLINE for three quarters of an hour. **Measured the boring way, and worth recording:** this row first said 30 minutes, because `monitor/host_staleness.go`'s literal default says 30m — the DEPLOYED value is in the manifest, and the 409s kept coming after the half hour was up. Reading a default and calling it the live value is the same mistake in a smaller coat. **The consequence is not theoretical:** a session that destroys its VM and then tears down the hub side walks away believing the delete failed, or leaves the customer behind — and the 2026-09-14 drill's teardown had the same shape without recording this. **Fix shape (smallest first):** the 409 body says *how long* it will refuse ("the last report was N minutes ago; deletion opens at HH:MM"), and `runbooks/target-selection.md`'s drill section names the wait. A force flag is NOT proposed — the refusal is right, only silent about its own clock. | **READY — rank P3-LOW; owner: CC (hub)** |
|
||||
| **R-600** | **[P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0.** FOUND 2026-09-20 by the slice-6 drill's teardown, **measured on ep0 rather than inferred from the hub**. The customer delete cascade finished at 19:41:54 with `customer DELETE cascade COMPLETE for drill-en-0920 (journal #18) — full teardown`, and every hub-side row was gone (0 configs, 0 hosts, 17 residue rows purged, PBS tenancy deprovisioned, escrow demoted). **Three minutes later `wg show wg0 allowed-ips` on ep0 still listed `10.77.0.5/32`** — the drill box's peer — because `wgsync` pushes on its own cycle. **Watched to the end rather than assumed: the peer was gone by 17:47:46Z — it outlived the *full teardown* line by about 6 minutes.** (My first estimate said ~35, read off the gap between two log lines; the sync runs oftener and only LOGS when something changes. That is the second time in this session that a period inferred from two log lines was wrong — the other was the delete's own staleness window. **A period read off two log lines is not a measurement.**) **The 2026-09-14 drill's findings say "the teardown removes it through the host delete"; measured, the host delete removes the hub's RECORD and the peer goes on the next push.** The mechanism is not broken — it is asynchronous, and the log line claims a completeness it does not yet have. **Why it is P2 rather than P3:** a session that tears down, reads *full teardown*, and leaves is the normal case; the peer outlives it by minutes, and the third teardown layer is the one the workspace rules single out as the one that gets forgotten. Six minutes is short — but the session that reads *full teardown* and leaves has no way to know it is six and not six hours. **Fix shape (smallest first):** the cascade triggers a wgsync push before it logs COMPLETE, or the log line says what is still pending and when ("wg peer removal queued; next push in N min"). A session should not have to read ep0 to know whether a teardown finished. **2026-09-25 (Peti's retirement):** nothing to remove for `peti-felhom-86d37d` — its host was deleted 2026-07-15 and ep0's live `wg show` carries no peer beyond the demo boxes', drill-r50 and the operator OOB (`audits/retire-peti-2026-09-25/A1-ep0-before.txt`); whether that July delete removed a peer, or none existed, is not recorded. **-- 2026-09-28: the Day-0 test install's peer (`drill-g0276`, key `Ly0yjK…`, 10.77.0.5) was checked on ep0 the day after its customer DELETE: gone from `wg show`, absent from `/etc/wireguard/*.conf` (control: a live key found), no namespace or file names the customer — the hub's sync removed it; nothing was removed by hand.** `audits/logins-nvme-2026-09-28/E/`. | **READY - rank P2-MEDIUM; owner: CC (hub)** |
|
||||
| **R-600** | **[P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0.** FOUND 2026-09-20 by the slice-6 drill's teardown, **measured on ep0 rather than inferred from the hub**. The customer delete cascade finished at 19:41:54 with `customer DELETE cascade COMPLETE for drill-en-0920 (journal #18) — full teardown`, and every hub-side row was gone (0 configs, 0 hosts, 17 residue rows purged, PBS tenancy deprovisioned, escrow demoted). **Three minutes later `wg show wg0 allowed-ips` on ep0 still listed `10.77.0.5/32`** — the drill box's peer — because `wgsync` pushes on its own cycle. **Watched to the end rather than assumed: the peer was gone by 17:47:46Z — it outlived the *full teardown* line by about 6 minutes.** (My first estimate said ~35, read off the gap between two log lines; the sync runs oftener and only LOGS when something changes. That is the second time in this session that a period inferred from two log lines was wrong — the other was the delete's own staleness window. **A period read off two log lines is not a measurement.**) **The 2026-09-14 drill's findings say "the teardown removes it through the host delete"; measured, the host delete removes the hub's RECORD and the peer goes on the next push.** The mechanism is not broken — it is asynchronous, and the log line claims a completeness it does not yet have. **Why it is P2 rather than P3:** a session that tears down, reads *full teardown*, and leaves is the normal case; the peer outlives it by minutes, and the third teardown layer is the one the workspace rules single out as the one that gets forgotten. Six minutes is short — but the session that reads *full teardown* and leaves has no way to know it is six and not six hours. **Fix shape (smallest first):** the cascade triggers a wgsync push before it logs COMPLETE, or the log line says what is still pending and when ("wg peer removal queued; next push in N min"). A session should not have to read ep0 to know whether a teardown finished. **2026-09-25 (Peti's retirement):** nothing to remove for `peti-felhom-86d37d` — its host was deleted 2026-07-15 and ep0's live `wg show` carries no peer beyond the demo boxes', drill-r50 and the operator OOB (`audits/retire-peti-2026-09-25/A1-ep0-before.txt`); whether that July delete removed a peer, or none existed, is not recorded. **-- 2026-09-28: the Day-0 test install's peer (`drill-g0276`, key `Ly0yjK…`, 10.77.0.5) was checked on ep0 the day after its customer DELETE: gone from `wg show`, absent from `/etc/wireguard/*.conf` (control: a live key found), no namespace or file names the customer — the hub's sync removed it; nothing was removed by hand.** `audits/logins-nvme-2026-09-28/E/`. **MEASURED AGAIN 2026-09-30 (a HOST delete, not a customer delete; new-household drill):** the host was deleted 07:23:03Z and `wgsync: pushed 4 peers` at 07:23:51Z — `10.77.0.5/32` gone from ep0 48 s later, the other four peers unchanged (`audits/evidence-drill-new-household-2026-09-30/teardown/layer4-ep0.txt`). | **READY - rank P2-MEDIUM; owner: CC (hub)** |
|
||||
| **R-601** | **[P2-MEDIUM] ~~demo-hp is unreachable~~ — WRONG, WITHDRAWN THE SAME DAY. The box was never down; MY ROUTES WERE.** Filed 2026-09-21 morning after `ssh demo-hp`, the hub-vaulted break-glass over the tailnet, `demo-hp-lan`, a ping and `ip neigh` on felhom-pve all failed, and `tailscale status` said *`demo-hp … offline, last seen 30d ago`*. **The operator looked at the hub and said it was ONLINE. It was**: it had reported 13 minutes earlier, and it has been up **4 weeks 2 days**. **Two stale facts, each enough on its own:** (1) `~/.ssh/config` sends `demo-hp` to the tailnet address `100.76.96.79`, and **tailscale is not installed on that box at all** (checked on it: no `tailscaled`, no `tailscale` binary) — so that entry is a dead peer from an earlier build and can never answer; (2) `demo-hp-lan` and `nodes.md` both say `192.168.0.87`, and the box is **statically** on **`192.168.0.104/24`**, bridge-port `nic0` (nodes.md says `enp2s0f0`). **The hub knew the right address the whole time** — every host report carries `addresses: [{iface: vmbr0, cidr: 192.168.0.104/24}, …]`. **What I actually did wrong, and it is the part worth keeping:** I ran `ip neigh` on felhom-pve, and `192.168.0.104 … STALE` was *in that output*, four lines above the `192.168.0.87 … FAILED` I quoted. I searched the output for the address I expected instead of reading it for the address that was there. **The standing rule says a "no access" claim must list what was tried; it does not say the list makes the claim true.** Six failed routes to a stale address are six failures of one assumption, not six pieces of evidence. **FIXED:** both `~/.ssh/config` entries repointed to `.104` (each carrying a comment saying why, including that there is no tailscale on this box), both verified live; `nodes.md` corrected. | **CLOSED 2026-09-21 — withdrawn, the claim was false; the routes are fixed** |
|
||||
| **R-604** | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** |
|
||||
| **R-605** | **[P3-LOW] A catalog gate that REFUSED TO RUN and a gate that ran and could not decide print the same word, so a reader cannot tell which happened.** FOUND 2026-09-21 while answering why the chaos night's update round could not run. On 2026-09-17 `check-image-resolvable` and `check-volume-persistence` both returned INCONCLUSIVE and the drawn `update` action was replaced with `use` (`audits/DRILL-chaos-night-2026-09-17.md:692-695`). **Neither script is defective — they behaved exactly as designed**, and both headers say why: a detector that cannot prove itself must refuse to report rather than guess (`check-image-resolvable.py` cites the 2026-07-21 incident where a Docker Hub throttle read as 24 of 65 pins falsely dead). **What is missing is the DISTINCTION.** `check-volume-persistence.py`'s `self_test` refuses to evaluate ANY app when it cannot build its canary image — a harness-level refusal — while `classify()` returns a per-app UNDETERMINED for an app that wrote nothing; `check-image-resolvable.py` likewise separates a harness-level canary failure (exit 2 at `check()` L180-183) from a per-pin throttle (L121-128). **`catalog_gates.py`'s VERDICT map collapses all of them into one `INCONCLUSIVE` label**, so the operator-facing summary cannot say whether the gate ran at all. **The cost is real and already paid:** no raw stdout of the 2026-09-17 run survives in either evidence directory, so the exact triggering path is INFERRED from the code plus the documented throttle precedent, not observed — a distinct summary line would have recorded it for free. **Fix shape:** have each gate's exit distinguish "the harness refused" from "the result is undetermined" (a third exit code, or a marker line the runner matches), and have `catalog_gates.py` print the two differently. **Ships with a decoy each way (R-421): a run whose canary fails must NOT read as a per-app undetermined, and vice versa.** Small. | **READY — rank P3-LOW; owner: CC (catalog)** |
|
||||
|
||||
Reference in New Issue
Block a user