drill 0.243.0: Phase 2 faults F10/F11/F12 measured, R-537 and R-538 filed
gates / gates (push) Successful in 20s
gates / gates (push) Successful in 20s
F10 (a child deletes the photo folder) is the finding: on a one-drive box with no off-site tier the household's own files are in NO backup — the whole-guest tiers exclude mp8 by design and the app's file leg lives at tier 2/3. The app page still labels tier 1 „DB + Konfig + Adatok" (R-537), and the restore reports success while leaving Nextcloud listing five photos it cannot open, after wiping the app's own trash which still held every byte (R-538). F11: the claim page locks out after the SECOND wrong code (15 minutes), the alarm fires and is true; Nextcloud does not lock out. F12: two reboots 60 s apart, all seven stacks back in 124 s, and the supervisor did not count the boots. Evidence: audits/evidence-drill-0243-2026-09-16/phase2-f10.txt, -f11.txt, -f12.txt. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -720,6 +720,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
| **R-534** | **[P1-HIGH] The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks `Datastore.Modify`, so every re-issue fails.** MEASURED 2026-09-16 on the drill box (`tester-1-652049`, fresh install from published ISO 1.27.1): the WG-registration hook refused as R-511 describes („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action"); the operator pressed exactly that, hub v0.114.0's ADOPT path ran, and the endpoint answered **`Process exited with status 255 (stderr: Error: permission check failed - missing Datastore.Modify on /datastore/felhom-offsite)`** → HTTP 502, no descriptor written, no secret stored (fail-closed, correct). So the code fix of 2026-09-15 is sound and INERT: a rebuilt box has no whole-guest off-site tier, and the customer's „Távoli rendszermentés" stays absent. **Not run by hand on ep0** (fenced). **Fix shape (operator):** grant the hub's tenantsync user `Datastore.Modify` on `/datastore/felhom-offsite` (it already holds the create/delete grants the provision path uses), or give the script a token-only re-key op that needs no Modify. Evidence: `audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt`. | **WAITING-ON-OPERATOR — rank P1-HIGH; owner: operator (ep0 grant) · CC (verify after)** |
|
||||
| **R-535** | **[P2-MEDIUM] The box's console still says „a doboz készen áll, és a párosításra vár" long after the box is bound, claimed and running apps — and it promises that the screen refreshes itself.** MEASURED 2026-09-16 on the drill box (fresh install, ISO 1.27.1, controller 0.243.0): bind succeeded 10:01:14Z, claim 10:06:18Z, four apps deploying by 10:22Z — and at 10:23Z the console still showed the pairing banner with the code `37S-NFE` and the line „Ez a képernyő magától frissül — nincs teendő a doboznál" (`audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png`). A volunteer watching the monitor has no way to tell the box is finished; worse, the screen says it updates itself, so waiting longer does not help. **Fix shape:** the first-boot banner unit re-renders on claim/bind state (the controller already knows: it reports `controller_started` and the hub holds `claimed`), showing „A doboz össze van kötve — a vezérlőpult a https://felhom.<domain> címen érhető el"; or at minimum stop printing the pairing code once the appliance is claimed. | **READY — rank P2-MEDIUM; owner: CC (ISO payload `felhom-bootstrap.sh` — ships with the next ISO)** |
|
||||
| **R-536** | **[P2-MEDIUM] The hub is told „Alkalmazás telepítve" the moment a deploy is ACCEPTED, so an install that never finishes is recorded as a completed one.** MEASURED 2026-09-16 on the drill box: the deploy of `mealie` was accepted at 12:31:36 CEST and the hub logged `Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie` in the SAME second; the controller was then killed 5 s in (F9'), and after the agent restarted it the stack read `not_deployed / deployed=false / deploying=false` — i.e. the app was never installed, and nothing corrected the event. Source confirms the ordering: `internal/api/router.go` writes the 202 „Telepítés elindítva" and then calls `NotifyAppDeployed` immediately, while the comment right above it says the deploy „runs asynchronously (compose pull/up + health happen after this returns)". The event is `info`, so nobody is mailed — but the customer timeline and the hub's app history record a completed install that did not happen (the „presence is not success" class). **Also measured, same shape:** an interrupted deploy leaves `/opt/docker/stacks/<app>/app.yaml` behind (written at accept) while the stack reads not-deployed — second instance after 2026-09-15's homebox; moved aside on the box. **Fix shape:** emit `app_deployed` from the async path when the stack reaches running/healthy (or emit `app_deploy_started` at accept and `app_deployed` at completion), and remove the accept-time `app.yaml` on a failed deploy. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. | **READY — rank P1-HIGH; owner: CC (controller)** |
|
||||
| **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. | **READY — rank P1-HIGH; owner: CC (controller)** |
|
||||
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |
|
||||
| **R-526** | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** |
|
||||
| **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** FOUND 2026-09-15: `stacks/metadata.go` parses it; `grep -rn LockedAfterDeploy` finds no reader; `deploy.html` renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in `02-controller-module-map.md`; the flag is a seam never wired. **Fix shape:** remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
Reference in New Issue
Block a user