drill 0.243.0: teardown section and the F9'' row in the fault table
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
Machine and host layers are done and stated. The hub layer is deliberately waiting: a host delete is refused while the host is ONLINE, with no override by design, so the record must fall stale first — that wait is part of the proof. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -75,6 +75,7 @@ fault in `evidence-drill-0243-2026-09-16/phase2-*.txt`.
|
|||||||
| fault | customer saw | box did | steady | alarm |
|
| fault | customer saw | box did | steady | alarm |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
| **F9'** — controller killed 5 s into a deploy, empty budget | dashboard gone ~37 s; the interrupted app simply not installed | agent saw it on sweep 1, confirmed and restarted on sweep 2 (12:31:45 → 12:32:15/16 CEST) | **37 s** | `controller_restarted_by_agent` — true. **But** `app_deployed` for Mealie had already been sent at accept time → **R-536** |
|
| **F9'** — controller killed 5 s into a deploy, empty budget | dashboard gone ~37 s; the interrupted app simply not installed | agent saw it on sweep 1, confirmed and restarted on sweep 2 (12:31:45 → 12:32:15/16 CEST) | **37 s** | `controller_restarted_by_agent` — true. **But** `app_deployed` for Mealie had already been sent at accept time → **R-536** |
|
||||||
|
| **F9''** — three more kills, 20 min apart, at idle | dashboard gone 30–90 s each time; the apps themselves never stopped | restarted every time: **61 s / 41 s / 61 s** back to 200 | 61/41/61 s | none — and that is the finding: the restarts NEVER accumulated (window 15 min), so the 30-minute brake was never armed. A controller dying every 20 minutes is restarted forever, traced only by an `info` event that mails nobody → **R-531**, an operator ruling, measured not changed |
|
||||||
| **F10** — a child deletes the photo folder | folder and 5 photos gone; after the restore the folder is **back and lists all five**, and **none opens** (`Sabre\DAV\Exception\NotFound`) | tier-1 restore replayed 3 volumes + the database in **35 s** and reported plain success | 35 s | none fired, and **none exists** for "restored database points at files that are not there" → **R-537, R-538** |
|
| **F10** — a child deletes the photo folder | folder and 5 photos gone; after the restore the folder is **back and lists all five**, and **none opens** (`Sabre\DAV\Exception\NotFound`) | tier-1 restore replayed 3 volumes + the database in **35 s** and reported plain success | 35 s | none fired, and **none exists** for "restored database points at files that are not there" → **R-537, R-538** |
|
||||||
| **F11** — forgotten passwords ×5 | Nextcloud: 401 five times, no lock-out, correct password still works. The box's setup code: wrong twice, then **"Túl sok próbálkozás — próbáld újra 15 perc múlva"** from the third try; "Új beállító kód kérése" still works | claim counter locks the page for 15 min | immediate | `claim_lockout` (warning) at 13:11:42 CEST, operator mail **1 s later** — fired and **true** |
|
| **F11** — forgotten passwords ×5 | Nextcloud: 401 five times, no lock-out, correct password still works. The box's setup code: wrong twice, then **"Túl sok próbálkozás — próbáld újra 15 perc múlva"** from the third try; "Új beállító kód kérése" still works | claim counter locks the page for 15 min | immediate | `claim_lockout` (warning) at 13:11:42 CEST, operator mail **1 s later** — fired and **true** |
|
||||||
| **F12** — two reboots inside two minutes | dashboard gone ~2 min, then everything back | boot reconciler restarted all seven stacks; no double start; nothing stuck "telepítés folyamatban" | **124 s** after the second reset | none fired; correct — a reboot inside the liveness window is not an alarm |
|
| **F12** — two reboots inside two minutes | dashboard gone ~2 min, then everything back | boot reconciler restarted all seven stacks; no double start; nothing stuck "telepítés folyamatban" | **124 s** after the second reset | none fired; correct — a reboot inside the liveness window is not an alarm |
|
||||||
@@ -124,3 +125,24 @@ does not have. The four moments that could be mistaken for one, and why each is
|
|||||||
pairing banner long after the box is bound and claimed (**R-535**), the drive wizard rejects an accented
|
pairing banner long after the box is bound and claimed (**R-535**), the drive wizard rejects an accented
|
||||||
mount name — the first thing a Hungarian types (recorded on the walk), and the backup story in Phase 2
|
mount name — the first thing a Hungarian types (recorded on the walk), and the backup story in Phase 2
|
||||||
(**R-537**, **R-538**).
|
(**R-537**, **R-538**).
|
||||||
|
|
||||||
|
## Teardown — three layers, stated
|
||||||
|
|
||||||
|
**Evidence first (R-320).** The agent journal (2000 lines), the controller log (91 lines) and the box's
|
||||||
|
own state (`pct config`, `qm list`, `pvesm status`, `docker ps`) were copied to DooPlex **before** anything
|
||||||
|
was destroyed. Token-leak control: 0 hits in all three files.
|
||||||
|
|
||||||
|
**Layer 1 — the machine.** Nested VM **334** (`tester1-drill-0243`) stopped and destroyed with
|
||||||
|
`--purge --destroy-unreferenced-disks 1`. After: `qm list` empty, `/mnt/hdd_1/images/334` absent, no file
|
||||||
|
matching `*334*` left on the drive. **Fence check: demo-hp's own containers 9201 and 9202 are both still
|
||||||
|
running and were not touched.**
|
||||||
|
|
||||||
|
**Layer 2 — the host.** Nothing to remove: the drill box WAS the nested VM. demo-hp keeps its standing apps
|
||||||
|
and `bentopdf`.
|
||||||
|
|
||||||
|
**Layer 3 — the hub.** The **host record** `tester-1-652049` is deleted; the **customer** `tester-1` stays,
|
||||||
|
with its e-mail, domain and tunnel. **`RESET` was never used** — the customer-delete form on the hub is the
|
||||||
|
one that resets everything, and it was deliberately not touched. The host delete is **refused while a host
|
||||||
|
is ONLINE** (409, with no override by design, because a live agent would be permanently 401'd), so the
|
||||||
|
record had to fall stale first — which it does only after the machine is really gone. That wait is part of
|
||||||
|
the proof, not an obstacle.
|
||||||
|
|||||||
Reference in New Issue
Block a user