drill 0.243.0: Phase 2 faults F10/F11/F12 measured, R-537 and R-538 filed
gates / gates (push) Successful in 20s

F10 (a child deletes the photo folder) is the finding: on a one-drive box with no
off-site tier the household's own files are in NO backup — the whole-guest tiers
exclude mp8 by design and the app's file leg lives at tier 2/3. The app page still
labels tier 1 „DB + Konfig + Adatok" (R-537), and the restore reports success while
leaving Nextcloud listing five photos it cannot open, after wiping the app's own
trash which still held every byte (R-538).

F11: the claim page locks out after the SECOND wrong code (15 minutes), the alarm
fires and is true; Nextcloud does not lock out. F12: two reboots 60 s apart, all
seven stacks back in 124 s, and the supervisor did not count the boots.

Evidence: audits/evidence-drill-0243-2026-09-16/phase2-f10.txt, -f11.txt, -f12.txt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 13:15:41 +02:00
parent bca45aaed1
commit 725a81a66a
13 changed files with 507 additions and 0 deletions
@@ -65,3 +65,26 @@ Vouched as a three-field change — golden 0.243.0 + agent 0.131.0 + min agent 0
**New finding on the walk: R-535 (P2)** — 25 minutes after a successful bind AND claim, with apps deploying, the box's console still showed „a doboz készen áll, és a párosításra vár" with the pairing code, and still promised „Ez a képernyő magától frissül".
## Phase 2 — the faults
Five faults, each recorded the same way: what the customer saw · what the box did · time to steady ·
did an alarm fire and was it true · did an alarm that should have fired stay silent. Evidence per
fault in `evidence-drill-0243-2026-09-16/phase2-*.txt`.
| fault | customer saw | box did | steady | alarm |
|---|---|---|---|---|
| **F9'** — controller killed 5 s into a deploy, empty budget | dashboard gone ~37 s; the interrupted app simply not installed | agent saw it on sweep 1, confirmed and restarted on sweep 2 (12:31:45 → 12:32:15/16 CEST) | **37 s** | `controller_restarted_by_agent` — true. **But** `app_deployed` for Mealie had already been sent at accept time → **R-536** |
| **F10** — a child deletes the photo folder | folder and 5 photos gone; after the restore the folder is **back and lists all five**, and **none opens** (`Sabre\DAV\Exception\NotFound`) | tier-1 restore replayed 3 volumes + the database in **35 s** and reported plain success | 35 s | none fired, and **none exists** for "restored database points at files that are not there" → **R-537, R-538** |
| **F11** — forgotten passwords ×5 | Nextcloud: 401 five times, no lock-out, correct password still works. The box's setup code: wrong twice, then **"Túl sok próbálkozás — próbáld újra 15 perc múlva"** from the third try; "Új beállító kód kérése" still works | claim counter locks the page for 15 min | immediate | `claim_lockout` (warning) at 13:11:42 CEST, operator mail **1 s later** — fired and **true** |
| **F12** — two reboots inside two minutes | dashboard gone ~2 min, then everything back | boot reconciler restarted all seven stacks; no double start; nothing stuck "telepítés folyamatban" | **124 s** after the second reset | none fired; correct — a reboot inside the liveness window is not an alarm |
| **M1** — memory pressure (R-528) | app restarted 9–10 times | all three OOM signals **silent**: `OOMKilled=false`, `docker events oom` empty, the cgroup invisible inside the guest, `dmesg` unreadable | — | no OOM alarm is possible on this box, same as on scratch 9202 |
**F10 is the finding of this drill.** On a one-drive box with no off-site tier — the state every fresh
install starts in — the household's own files are in **no backup at all**: the whole-guest tiers exclude
the data drive by design (`07-backup-architecture.md`, "[FACT] What the whole-guest tiers do NOT carry",
confirmed live: `excluding bind mount point mp8 … (not a volume)`), and the app's file leg lives at
tier 2 / tier 3, both unset. The design is not the defect. The defects are that the page calls tier 1
„DB + Konfig + Adatok" and prints the drive size beside it (**R-537**), and that a restore reports
success while leaving the app listing files it cannot open — and wipes the app's own trash, which still
held every byte (**R-538**).
@@ -12,3 +12,13 @@
2026/09/16 12:21:58 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Paperless-ngx
2026/09/16 12:22:18 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Nextcloud
page after: Rendszermentés (teljes mentés) A teljes szerver — alkalmazások, beállítások és adatbázisok együtt — időszakos mentése, amelyből az egész készülék visszaállítható. Ezt a host-ügynök készíti és kezeli. ✓ Utolsó teljes mentés 2026-09-16 12:08 (20 perce) 624.3 MB Helyi tároló (local) Naprakész Következő mentés 0 órája — a mentési ablakon belül – Visszaállítás ellenőrizve Még nem futott Helyi tároló (local) ✓ Utolsó sikeres mentés: 2026-09-16 12:08 (20 perce) · 624.3 MB Naprakész Biztonsági szerver – külön hardver (PBS) nincs beállítva Mentés folyamatban… (fázis: snapshotted ) Mentés most A mentés
## 2026-09-16T10:31:39Z manual backup finished: guest lock cleared; archive vzdump-lxc-9201-2026_09_16-12_27_11.tar.zst = 2 175 571 124 B (the scheduled one at 12:08 was 654 665 901 B, before the four apps)
## hub events/mails for tester-1 in the last 25 min:
2026/09/16 12:08:25 [INFO] Event from tester-1: backup_tier_skipped (warning) — Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it.
2026/09/16 12:08:25 [INFO] Operator email sent for tester-1/backup_tier_skipped
2026/09/16 12:21:17 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: PrivateBin
2026/09/16 12:21:38 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Vaultwarden
2026/09/16 12:21:58 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Paperless-ngx
2026/09/16 12:22:18 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Nextcloud
2026/09/16 12:31:36 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie
## 2026-09-16T10:50:29Z manual backup finished (guest lock cleared)
@@ -0,0 +1,85 @@
## F10 - a child deletes the photo folder (Nextcloud) - 2026-09-16T11:01:38Z
seed: folder 'Fotok' created through Nextcloud WebDAV (the app's own file interface), MKCOL 201,
5 files nyaralas-1..5.jpg, 200k/400k/600k/800k/1000k = 3 000 000 B total, every PUT 201.
on-disk: 68204151 /mnt/felhom-drives/adatlemez/appdata/nextcloud
on-disk: 5
on-disk: 25337 /mnt/felhom-drives/adatlemez/backups
2026-09-16T11:01:57Z pressing 'Mentes most' on the app-backup page (POST /api/backup/run)
trigger: {"ok":true,"message":"Mentés elindítva"}
http=200
2026-09-16T11:02:58Z backup finished; status: {"ok":true,"data":{"db_dump":{"count":2,"duration":"47.936765144s","last_run":"2026-09-16T11:02:45.477832211Z","success":true},"enabled":true,"running":false}}
restore page after the app backup:
TEXT: Pillanatkép: — Válasszon alkalmazást — Még nincs mentés felhasználói adattal. A visszaállítás felülírja az alkalmazás jelenlegi adatait a kiválasztott mentés állapotával. Az alkalmazás a folyamat során automatikusan leáll és újraindul. Megértettem, visszaállítás indítása. Visszaállítás indítása Importálás mentett csomagból (.fab) Hordozható mentéscsomag (.fab) Hordozható pillanatfelvétel — bárhol tárolhatod, és bármikor visszatöltheted egy meghajtóról. A folyamatos védelmet az 1–3. szintű mentés adja. Jelszavas titkosítás (opcionális) Üresen hagyva a csomag titkosítás nélkül készül. Nextcloud Letöltés (.fab) Paperless-ngx Letöltés (.fab) PrivateBin Letöltés (.fab) Vaultwarden Letöltés (.fab)
2026-09-16T11:05:06Z the child deletes the folder: DELETE /Fotok http=204
folder after the delete: PROPFIND http=404 (404 = gone)
## MEASURED before the restore attempt:
## 1) The app backup ("1. mentes", tier 1) was made at 11:02:45Z by the customer-visible "Mentes most"
## button (POST /api/backup/run -> {"ok":true,"message":"Mentes elindItva"}; the backup folder grew
## 25 337 B -> 978 MB, so the button does more than its status field, which reports db_dump only).
## 2) Its manifest lists db-dumps + THREE docker volume dumps (html, db_data, redis). It does NOT list
## the drive-side app data (/mnt/felhom-drives/adatlemez/appdata/nextcloud), where the photos live.
## 3) Proof by listing the html tar (29 346 entries):
## positive control "version.php" = 3 hits (the listing works)
## "data/" entries = 1284, but "./data/" is the EMPTY bind-mount point
## "Fotok" = 0 hits
## "nyaralas"= 0 hits
## And: find over the whole backups tree for *appdata*/*hdd*/*Fotok* = nothing.
## 4) The whole-guest backup does not cover them either - vzdump log of the 12:27 CEST run:
## "including mount point rootfs ('/')", "including mount point mp0 ('/var/lib/felhom')",
## "excluding bind mount point mp8 ('/mnt/felhom-drives') from backup (not a volume)".
## 5) The app-backup page nevertheless labels tier 1 "DB + Konfig + Adatok" and prints "Nextcloud
## Adatlemez 65.1 MB" - a size measured on exactly the data it does not copy.
## 6) Tier 2 ("off-drive masolat") and tier 3 (offsite) both read "Nincs beallitva" on this box.
2026-09-16T11:06:56Z pressing the restore button: POST /backup/restore stack_name=nextcloud snapshot_id=helyi
restore POST http=302
2026-09-16T11:07:42Z restore status: {"ok":true,"data":{"running":false,"op":"restore","stack":"nextcloud","started_at":"2026-09-16T11:06:56.710474757Z","last":{"op":"restore","stack":"nextcloud","ok":true,"message":"A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult.","finished_at":"2026-09-16T11:07:31.194473339Z"},"last_recent":true}}
after the restore, through Nextcloud's own interface:
PROPFIND /Fotok http=207
file:
file: nyaralas-1.jpg
file: nyaralas-2.jpg
file: nyaralas-3.jpg
file: nyaralas-4.jpg
file: nyaralas-5.jpg
GET /Fotok/nyaralas-1.jpg http=503 bytes=276
trash PROPFIND http=207
trash:
RE-MEASURE at 2026-09-16T11:08:38Z, app healthy (nextcloud Up >2 min):
app status.php http=200 <- positive control, app is serving
PUT control kontroll.txt http=201 <- positive control, WebDAV write works
GET control kontroll.txt http=200 bytes=6 <- positive control, WebDAV read works
GET /Fotok/nyaralas-1.jpg http=404 bytes=249
GET /Fotok/nyaralas-2.jpg http=503 bytes=276
GET /Fotok/nyaralas-3.jpg http=503 bytes=276
GET /Fotok/nyaralas-4.jpg http=503 bytes=276
GET /Fotok/nyaralas-5.jpg http=503 bytes=276
body of the failed download (first 200 chars):
<?xml version="1.0" encoding="utf-8"?><d:error xmlns:d="DAV:" xmlns:s="http://sabredav.org/ns"> <s:exception>Sabre\DAV\Exception\NotFound</s:exception> <s:message>File with name /Fotok/nyaralas-1
## RESULT - F10 (a child deletes the photo folder, Nextcloud, fresh box, one drive, no tier 2/3)
## customer saw: the folder and the 5 photos vanish from Nextcloud (DELETE 204, PROPFIND 404).
## After the restore the folder is BACK and lists all 5 photos - but NONE of them opens:
## GET nyaralas-1..5 = 404 / 503 x4, body "Sabre\DAV\Exception\NotFound".
## Positive controls at the same moment: status.php 200, WebDAV PUT kontroll.txt 201, GET 200 (6 B).
## So the failure is real, not a starting app and not a broken login.
## box did: POST /backup/restore (stack_name=nextcloud, snapshot_id=helyi) -> 302, finished in 35 s,
## "A(z) nextcloud: 3 adatkotet es az adatbazis visszaallitva - az alkalmazas ujraindult."
## It replayed the 3 docker volumes + the MariaDB dump. The DB dump (11:01:45Z) KNOWS the 5 photos,
## because they were uploaded at ~11:00Z. The BYTES were never in any backup (see MEASURED above).
## time to steady: 35 s restore, app healthy ~50 s later.
## how much was lost: all 5 files, 3 000 000 B - every photo the folder ever held.
## does the page say so? NO. The result line says the volumes and the database were restored and the
## app restarted. Nothing says the app's files on the data drive are not part of this backup, and
## nothing says the restored database now points at files that do not exist.
## where the bytes actually are: Nextcloud's OWN trash on the data drive -
## appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg (all 5 present).
## But the restored database no longer lists them: the trash PROPFIND returns an EMPTY list, so the
## customer cannot press "restore from trash" either. The restore made the trash unreachable.
## alarm fired and true? none fired.
## alarm that should have and did not: none is defined for "restored DB references missing files".
## DESIGN CHECK (a design is not a defect, R-370): 07-backup-architecture.md is explicit -
## "[FACT] What the whole-guest tiers do NOT carry ... mp8 /mnt/felhom-drives ... out of vzdump scope",
## and the file leg for nextcloud exists at TIER 2 and TIER 3 only (section 6.2 table: Tier-3
## mandatory = calibre-web, immich, nextcloud, paperless-ngx). Tier 1 carrying no drive-side files is
## therefore BY DESIGN. What is NOT by design is the page calling tier 1 "DB + Konfig + Adatok" and
## printing "Adatlemez 65.1 MB" next to it, and a restore reporting plain success while leaving the
## app inconsistent. Those two are the rows.
@@ -0,0 +1,61 @@
## F11 - forgotten passwords, part 1: Nextcloud, 5 wrong tries 2026-09-16T11:09:35Z
try 1: http=401 (took 868 ms)
try 2: http=401 (took 726 ms)
try 3: http=401 (took 702 ms)
try 4: http=401 (took 725 ms)
try 5: http=401 (took 687 ms)
now the CORRECT password again (positive control - is the customer locked out?):
http=200 (took 610 ms) 207 = still allowed in
web login page http=200
## F11 part 2 - wrong CLAIM code x5, measured on this box (already claimed at 10:06:18Z)
the claim page itself: http=200
page says: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új beállító kód kérése Felhom — Otthoni szerver kezelés felhom.eu
wrong claim try 1: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m
wrong claim try 2: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m
wrong claim try 3: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m
wrong claim try 4: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m
wrong claim try 5: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m
## F11 part 2 - CORRECTED run (my first run was refused by CSRF - I measured my own mistake, not the product)
wrong code try 1: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új
wrong code try 2: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új
wrong code try 3: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új
wrong code try 4: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új
wrong code try 5: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új
NOTE: this box was already claimed at 10:06:18Z. The /claim URL still answers 200 and now renders
'Jelszo visszaallitasa' (password reset) with a 'Uj beallito kod kerese' button - so the claim
page is NOT one-shot-dead after a successful claim; it becomes the reset surface.
## F11 part 2 - FINAL run (correct CSRF: cookie felhom_claim_csrf + _csrf field + X-CSRF-Token header)
wrong code try 1: http=200 (198 ms) page says: Hibás vagy lejárt kód Add meg az e-mailben kapott beállító kódot, majd válassz új
wrong code try 2: http=200 (187 ms) page says: Hibás vagy lejárt kód Add meg az e-mailben kapott beállító kódot, majd válassz új
wrong code try 3: http=200 (177 ms) page says: Túl sok próbálkozás — próbáld újra 15 perc múlva. Add meg az e-mailben kapott beá
wrong code try 4: http=200 (209 ms) page says: Túl sok próbálkozás — próbáld újra 15 perc múlva. Add meg az e-mailben kapott beá
wrong code try 5: http=200 (172 ms) page says: Túl sok próbálkozás — próbáld újra 15 perc múlva. Add meg az e-mailben kapott beá
CORRECTION of my own summary line: there IS a lock-out, and it arrives on the THIRD try.
Tries 1-2: "Hibas vagy lejart kod". Tries 3-5: "Tul sok probalkozas - probald ujra 15 perc mulva."
No captcha and no growing delay (172-209 ms throughout) - it is a counter, not a slow-down.
the household's escape route while locked out - pressing 'Uj beallito kod kerese':
POST /claim/request-new-code http=200
page says: Ha az e-mail cím regisztrálva van, elküldtük a kódot. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jel
## RESULT - F11 (forgotten passwords x5, both kinds)
## A) Nextcloud app password, 5 wrong tries:
## customer saw: 401 five times, each ~700 ms, no lock-out, no captcha; the CORRECT password then
## worked immediately (200). Positive controls in the same run: status.php 200, WebDAV PUT 201/GET 200.
## box did: nothing - no event, no mail, no restart.
## time to steady: none needed.
## alarm fired and true? none. alarm that should have and did not: none is expected for an app login.
## B) The box's own setup code (the claim page), 5 wrong tries:
## customer saw: tries 1-2 "Hibas vagy lejart kod"; tries 3-5 "Tul sok probalkozas - probald ujra
## 15 perc mulva." So the lock-out starts on the THIRD attempt and lasts 15 minutes. Timing was
## flat (172-209 ms), so it is a counter, not a slow-down.
## the escape route WORKS while locked out: POST /claim/request-new-code -> 200 and
## "Ha az e-mail cim regisztralva van, elkuldtuk a kodot." (deliberately neutral wording).
## one-shot behaviour, MEASURED not assumed: this box was claimed at 10:06:18Z, and /claim still
## answers 200 afterwards - it becomes the "Jelszo visszaallitasa" (password reset) surface with a
## "Uj beallito kod kerese" button. The URL is NOT dead after a successful claim.
## NOT WALKED: the same five wrong codes on a SECOND, never-claimed box. That needs a second fresh
## install (~35 min) and the drill box was needed for F9''. So the "first-ever claim" variant of
## the counter is untested; what is proven here is the reset path on a claimed box.
## MY OWN MISTAKE, recorded: my first two runs of B were refused with "Ervenytelen urlap" because I
## paired the claim page's pre-auth token wrongly (the page sets cookie felhom_claim_csrf and carries
## TWO _csrf fields). I measured my own error, said so, and re-ran with cookie + _csrf field +
## X-CSRF-Token header, which the product accepted.
@@ -0,0 +1,70 @@
## F12 — two reboots inside two minutes (box: nested VM 334 on demo-hp)
2026-09-16T10:50:10Z before: dashboard=200 apps=200
reset 1 at 2026-09-16T10:50:11Z
reset 2 at 2026-09-16T10:51:12Z (60 s after the first)
dashboard 200 again at 2026-09-16T10:53:16Z — 124 s after the second reset
apps: privatebin=200 vault=200 paperless=302 cloud=200
Sep 16 12:53:06 tester1 felhom-agent[1128]: time=2026-09-16T12:53:06.326+02:00 level=WARN msg="controller-supervisor: guest list unavailable — skipping sweep (ownership unproven)" err="pro
Sep 16 12:53:07 tester1 felhom-agent[1128]: time=2026-09-16T12:53:07.077+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1
Sep 16 12:53:36 tester1 felhom-agent[1128]: time=2026-09-16T12:53:36.356+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1
Sep 16 12:53:36 tester1 felhom-agent[1128]: time=2026-09-16T12:53:36.385+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1
Sep 16 12:53:36 tester1 felhom-agent[1128]: time=2026-09-16T12:53:36.388+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1
Sep 16 12:54:06 tester1 felhom-agent[1128]: time=2026-09-16T12:54:06.388+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1
paperless-webserver Up 34 seconds (healthy)
paperless-redis Up 44 seconds (healthy)
paperless-postgres Up 44 seconds (healthy)
nextcloud Up 56 seconds (healthy)
nextcloud-db Up About a minute (healthy)
nextcloud-redis Up About a minute (healthy)
felhom-controller Up About a minute (healthy)
vaultwarden Up About a minute (healthy)
privatebin Up About a minute (healthy)
filebrowser Up About a minute (healthy)
cloudflared Up About a minute
traefik Up About a minute
## app states after the two resets (2026-09-16T10:57:17Z):
## supervisor block in the host report:
Traceback (most recent call last):
File "<string>", line 1, in <module>
import json;print(json.load(open("/var/lib/felhom-agent/bootstrap.json"))["local_api_token"])
~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
FileNotFoundError: [Errno 2] No such file or directory: '/var/lib/felhom-agent/bootstrap.json'
Traceback (most recent call last):
File "<string>", line 1, in <module>
import sys,json;r=json.load(sys.stdin);print(json.dumps(r.get("controller_supervisor"),indent=1))
~~~~~~~~~^^^^^^^^^^^
File "/usr/lib/python3.13/json/__init__.py", line 293, in load
return loads(fp.read(),
cls=cls, object_hook=object_hook,
parse_float=parse_float, parse_int=parse_int,
parse_constant=parse_constant, object_pairs_hook=object_pairs_hook, **kw)
File "/usr/lib/python3.13/json/__init__.py", line 346, in loads
return _default_decoder.decode(s)
~~~~~~~~~~~~~~~~~~~~~~~^^^
File "/usr/lib/python3.13/json/decoder.py", line 345, in decode
obj, end = self.raw_decode(s, idx=_w(s, 0).end())
~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/lib/python3.13/json/decoder.py", line 363, in raw_decode
raise JSONDecodeError("Expecting value", s, err.value) from None
json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
## uptime + boot count:
up 5 minutes
reboot system boot 7.0.2-6-pve Wed Sep 16 12:51 - still running
reboot system boot 7.0.2-6-pve Wed Sep 16 12:50 - crash
reboot system boot 7.0.2-6-pve Wed Sep 16 11:58 - crash
## RESULT — F12 (two `qm reset`s 60 s apart on VM 334)
## customer saw: dashboard unreachable ~2 min; back at 10:53:16Z, 124 s after the SECOND reset.
## box did: boot reconciler brought all seven stacks up by itself (privatebin 200, vaultwarden 200,
## paperless 302 = its login redirect, nextcloud status.php 200); NO double start observed;
## mealie stayed not_deployed (the F9'-interrupted deploy), NOT stuck in "telepites folyamatban".
## time to steady: 124 s to the dashboard, ~3 min to every app healthy.
## supervisor did NOT count the boots: whole-journal "RESTARTED the controller" = 1 (the F9' one at
## 12:32, POSITIVE control that the grep string matches when it happens); since 12:49 = 0,
## while 4 controller-supervisor lines in the same window prove it was running and sweeping.
## During the boot it logged "guest list unavailable - skipping sweep (ownership unproven)" -
## the unprovisioned/ownership guard doing its job.
## journal is PERSISTENT on this box (/var/log/journal exists), so the pre-reboot lines are real.
## alarm fired and true? none fired. alarm that should have and did not: none - a reboot inside the
## node-liveness dedupe window produces no host_* mail, which matches the design.
@@ -0,0 +1,13 @@
## F9'' - three more controller kills, 20 minutes apart, at IDLE (no deploy running)
## Question to MEASURE (not to change): does a restart after 20 minutes of healthy uptime
## count against the 3-restarts-per-15-minutes budget, or does the window start fresh?
## Baseline: the last restart before this run was F9' at 12:32:16 CEST (1 of 3 at that time).
## kill 1 at 2026-09-16T11:12:41Z
felhom-controller
dashboard 200 again after 61 s
Sep 16 13:12:37 tester1 felhom-agent[1128]: time=2026-09-16T13:12:37.326+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=40 guests_evaluated=1 controllers_not_running=0
Sep 16 13:13:07 tester1 felhom-agent[1128]: time=2026-09-16T13:13:07.323+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status
Sep 16 13:13:37 tester1 felhom-agent[1128]: time=2026-09-16T13:13:37.415+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit
Sep 16 13:13:38 tester1 felhom-agent[1128]: time=2026-09-16T13:13:38.655+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
RESTARTED lines in the whole journal so far: 2
waiting 20 minutes before the next kill
@@ -0,0 +1,19 @@
## F9' — controller killed 5 s into a deploy, on a FRESH crash-loop budget (agent started 12:02, no restarts yet)
2026-09-16T10:31:36Z starting the mealie deploy
deploy http=202
{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
killed at 2026-09-16T10:31:41Z — 5 s into the deploy (the "(37901 s …)" text my script printed here was a timestamp-parsing bug of mine; the real interval is the scripted sleep of 5 s between the 202 and the kill)
dashboard health 200 again at 2026-09-16T10:32:18Z — 37 s after the kill
mealie after: {'state': 'not_deployed', 'deployed': False, 'deploying': False}
Sep 16 12:31:45 tester1 felhom-agent[3602]: time=2026-09-16T12:31:45.160+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status
Sep 16 12:32:15 tester1 felhom-agent[3602]: time=2026-09-16T12:32:15.231+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit
Sep 16 12:32:16 tester1 felhom-agent[3602]: time=2026-09-16T12:32:16.606+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
Sep 16 12:32:16 tester1 felhom-agent[3602]: time=2026-09-16T12:32:16.606+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=60 guests_evaluated=1 controllers_not_running=1
## RESULT (R-531's missing timing, now pinned):
## kill 5 s into a deploy, on an EMPTY crash-loop budget → dashboard health 200 again 37 s after the kill
## (agent saw it at 12:31:45, confirmed and restarted at 12:32:15/16 — two sweeps, as designed).
## The interrupted app ended `not_deployed / deployed=false / deploying=false` — NOT stuck in "telepítés folyamatban".
## Budget after this single restart: 1 of 3 in the 15-minute window.
2026-09-16 10:31:36.748187371 +0000 253 /opt/docker/stacks/mealie/app.yaml
deployed deployed_at env locked_fields desired_state
moved-aside
@@ -0,0 +1,65 @@
are supported and installed on your system.
## 2026-09-16T10:35:16Z M1 — OOM signals on THIS box (R-528). before:
limit=1342177280 oom=false restarts=0 status=running
81: memory: 128M
134: memory: 128M
## 2026-09-16T10:37:30Z after the 128M cap: false restarting 9
cgroup: not visible
docker oom events since 2026-09-16T10:35:29Z:
kernel OOM lines: 0
dmesg not readable in guest
2
## 2026-09-16T10:40:13Z cap restored: limit=1342177280 health=unhealthy restarts=10
are supported and installed on your system.
## 2026-09-16T10:40:44Z memory lines now:
81: memory: 1280M
112: memory: 256M
134: memory: 1280M
81: memory: 1280M
112: memory: 256M
134: memory: 128M
## 2026-09-16T10:47:37Z paperless after fix: ws=unhealthy limit=1342177280 redis=134217728
## M1 RESULT (R-528), on THIS box (nested VM, customer LXC 9201, Docker in the guest):
## cap 128M → paperless-webserver RESTARTING with RestartCount 9–10, and:
## - docker inspect .State.OOMKilled = FALSE
## - docker events --filter event=oom (since the cap change) = EMPTY
## - the container's cgroup is NOT visible from inside the guest (find under /sys/fs/cgroup → nothing),
## so memory.events / oom_kill could not be read there
## - dmesg is not readable in the guest
## So all three signals are silent here, exactly as on scratch 9202 (2026-09-15). BIGNIGHT's VM 333 did read
## oomkilled=true for the worker-killed-inside-a-running-container shape; a MAIN-process kill + restart reports
## nothing. The controller v0.243.0 OOM tag therefore cannot fire for this (restart) shape on a real box.
## MISTAKE, stated: my cap edit used a blunt sed, so the restore rewrote the REDIS limit too (128M → 1280M) —
## the same mistake as 2026-09-15 on 9202. Repaired in the next step (redis back to 128M, webserver 1280M).
paperless-redis Up 7 minutes (healthy)
paperless-webserver Restarting (1) 57 seconds ago
paperless-postgres Up 12 minutes (healthy)
---
File "/usr/local/lib/python3.12/site-packages/django/db/backends/base/base.py", line 256, in connect
self.connection = self.get_new_connection(conn_params)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/django/utils/asyncio.py", line 26, in inner
return func(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/django/db/backends/postgresql/base.py", line 332, in get_new_connection
connection = self.Database.connect(**conn_params)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/psycopg/connection.py", line 120, in connect
raise last_ex.with_traceback(None)
django.db.utils.OperationalError: connection failed: connection to server at "172.20.0.3", port 5432 failed: FATAL: password authentication failed for user "paperless"
s6-rc: warning: unable to start service init-migrations: command exited 1
/run/s6/basedir/scripts/rc.init: warning: s6-rc failed to properly bring all the services up! Check your logs (in /run/uncaught-logs/current if you have in-container logging) for more information.
/run/s6/basedir/scripts/rc.init: fatal: stopping the container.
## 2026-09-16T10:48:27Z repairing paperless through the CONTROLLER's own path (my manual compose up broke its env):
stop: {"ok":true,"message":"Stack paperless-ngx stop completed"}
http=200
start: {"ok":true,"message":"Stack paperless-ngx start completed"}
http=200
paperless state after the controller-driven restart: running at 2026-09-16T10:49:35Z
## MISTAKE, stated (2026-09-16): the M1 measurement recreated the paperless webserver with a bare
## `docker compose up -d` INSIDE the guest. The controller injects the app's environment (DB password etc.) from
## `app.yaml` at deploy time, so a hand-run compose brings the container up WITHOUT it: the webserver then failed
## `password authentication failed for user "paperless"` against its own postgres and crash-looped.
## Nothing about the product; the fault was the manual path. Repaired by restarting the stack through the
## CONTROLLER's own stop/start API, which re-renders the environment. Lesson for the next measurement: drive app
## containers through the controller, or set the cap via `docker update --memory`, never by re-running compose.
@@ -0,0 +1,48 @@
## Alarm truth table — operator mails actually DELIVERED during the drill (read from the mailbox, not from the hub log)
## (to admin@felhom.eu, via the Gmail connector, 2026-09-16)
1. 12:04:01 CEST [Felhom] ✅ tester-1: node_recovered (severity INFO) — "Reports resumed (was down for 0m)".
TRUE but odd: the box had never been down; it was newly enrolled. Also: severity INFO was MAILED, while the
project's rule says severityNotifies DROPS info before both legs — to be checked against the dispatcher (below).
2. 12:08:26 CEST [Felhom] ⚠️ tester-1: backup_tier_skipped (severity warning) — "Whole-guest backup tier felhom-pbs
skipped: its storage does not exist on the host (never provisioned or removed)". TRUE: R-534 blocked the tier, so
the tier's storage is genuinely absent. This is R-518's cheap half firing on a real box, with its operator mail.
No other operator mail arrived in the window (the third thread in the mailbox is a DMARC report, unrelated).
## Faults injected so far: none (Phase 2 starts with F9'). Alarms that SHOULD have fired and did not: none so far.
## Why the INFO mail is correct (checked in code, not assumed):
## `node_recovered` is emitted with severity "info" (monitor/staleness.go emitTransition) and handed to
## dispatcher.ProcessEvent, which returns before both legs for "info" — BUT the recovery branch
## (`recoveredPairedDownTypes`, dispatcher.go:124) runs BEFORE that gate, deliberately, so the operator hears the
## all-clear for a down it was told about. The customer leg stays pairing-gated.
## In context the mail is also TRUE: customer `tester-1` had been reporting nothing since the BIGNIGHT teardown, so the
## new box's first report IS a recovery, and "(was down for 0m)" is the age of the gap the checker could see.
## No row filed.
## ALARM TRUTH TABLE - second pass, 2026-09-16 (hub 0.115.0), built from the hub's own log lines
## (pod hub-85478f77d7-mpxkr, times CEST) paired with the customer timeline on /customers/tester-1.
##
## event when severity operator mail? true in context?
## selfbind_link_sent 11:59:55 info yes (self-bind e-mail) yes - the operator pressed it
## claim_reissued_reenroll 12:01:56 info yes (claim code to customer) yes - 3rd generation after reinstall
## controller_started 12:03:30 info no correct - info, not paired
## node_recovered 12:04:00 info YES - mail sent BY DESIGN: the recovery branch runs
## before the severity gate; true here
## backup_tier_skipped 12:08:25 warning yes yes - the off-site tier has no storage
## backup_tier_skipped 12:37:17 warning SUPPRESSED - cooldown correct - same key within the hour
## (key=tester-1:backup_tier_skipped:felhom-pbs)
## app_deployed x5 12:21-12:31 info no NOT always true -> R-536 (sent at accept)
## app_start_failed (paperless) 12:48:47 warning yes TRUE - and it was MY damage (the
## hand-run compose, recorded in phase2-m1)
## controller_restarted_by_agent 12:48 info no (info) yes - F9' restart #1, timestamps match
## backup_tier_skipped 12:58:13 warning SUPPRESSED - cooldown correct
## claim_lockout 13:11:42 warning YES, 1 s later TRUE - F11 part 2, third wrong code
## claim reset code (gen 4) 13:12:05 info yes (to the customer) yes - "Uj beallito kod kerese"
##
## What this pass adds to the first one:
## 1) The 5-minute node/host dedupe (R-529) was NOT exercised today: no node_* or host_* alarm fired
## after 12:04, because the box stayed up and reporting. The two reboots of F12 were far too short
## to reach the 30-minute staleness threshold. So the widened bypass is SHIPPED but UNPROVEN-LIVE
## on this box; the proof from 2026-09-15 on the other box still stands.
## 2) The one-hour operator cooldown is PROVEN here twice, by its own log line, on backup_tier_skipped.
## 3) Every warning that fired today was true. No warning that should have fired stayed silent, EXCEPT
## the class R-538 names: there is no alarm at all for "a restore finished and the app is now
## inconsistent", so nothing could have fired.
@@ -0,0 +1,39 @@
## 2026-09-16T11:14:14Z ep0 READ-ONLY listing of tester-1's namespace (nothing removed, nothing written)
+================+====================+=========+
| name | path | comment |
+================+====================+=========+
| felhom-offsite | /mnt/pbs-datastore | |
+================+====================+=========+
--- groups in namespace tester-1:
Error: error building client for repository felhom-offsite - no password input mechanism available
## 2026-09-16T11:14:24Z ep0 listing by filesystem (read-only; the client needs a password we deliberately do not hold here)
--- namespaces:
demo-felhom
demo-hp
tester-1
--- tester-1 groups:
total 12
drwxr-xr-x 3 backup backup 4096 Sep 14 16:04 .
drwxr-xr-x 5 backup backup 4096 Sep 14 15:48 ..
drwxr-xr-x 2 backup backup 4096 Sep 15 08:33 ct
--- ct groups and their snapshots:
## 2026-09-16T11:14:42Z ep0 deeper listing (read-only)
--- tester-1/ct contents:
total 8
drwxr-xr-x 2 backup backup 4096 Sep 15 08:33 .
drwxr-xr-x 3 backup backup 4096 Sep 14 16:04 ..
--- snapshots per group:
group: /mnt/pbs-datastore/ns/tester-1/ct/*/
--- other namespaces, for scale:
9201
## CONCLUSION (nothing was removed by this session; ep0 was read only)
## Before the drill (09:12:21Z): ns tester-1 existed with an EMPTY ct group dir - the old box's
## snapshots had been removed the previous night, with the operator's explicit yes.
## After the whole drill (11:14:42Z): ns tester-1/ct is STILL EMPTY.
## Why: this box never had an off-site tier at all. The re-issue failed on the endpoint token's missing
## Datastore.Modify grant (R-534), so no descriptor, no secret, no upload path. The box therefore
## wrote ZERO snapshots off-site during the entire drill.
## For scale, the neighbouring namespace demo-hp/ct holds group 9201 - proof that this listing method
## DOES show groups when they exist (positive control).
## Consequence for Phase 3: "restore one DB-backed app from the off-site tier onto scratch 9202"
## CANNOT BE WALKED on this box. Not skipped for time - there is no off-site copy to restore from.
@@ -0,0 +1,52 @@
## tester-1 event timeline as the hub shows it (customer page), read 2026-09-16T11:13:18Z
Sep 16 11:11 | warning | claim_lockout | Túl sok hibás beállító kód — a beállító oldal 15 percre zárolva | controller
Sep 16 10:58 | warning | backup_tier_skipped | Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. | controller
Sep 16 10:53 | info | controller_started | Controller elindult (0.243.0) | controller
Sep 16 10:48 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller
Sep 16 10:48 | info | controller_restarted_by_agent | Host tester-1-652049 guest 9201: the agent restarted the controller (controller container exited on 2 consecutive sweeps) — restart #1 since the agent started | hub
Sep 16 10:37 | warning | backup_tier_skipped | Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. | controller
Sep 16 10:32 | info | controller_started | Controller elindult (0.243.0) | controller
Sep 16 10:31 | info | app_deployed | Alkalmazás telepítve: Mealie | controller
Sep 16 10:27 | info | recovery_credential_revealed | Az üzemeltető lekérte a géped konzolos hozzáférési jelszavát (távoli hibaelhárítás). | hub
Sep 16 10:22 | info | app_deployed | Alkalmazás telepítve: Nextcloud | controller
Sep 16 10:21 | info | app_deployed | Alkalmazás telepítve: Paperless-ngx | controller
Sep 16 10:21 | info | app_deployed | Alkalmazás telepítve: Vaultwarden | controller
Sep 16 10:21 | info | app_deployed | Alkalmazás telepítve: PrivateBin | controller
Sep 16 10:08 | warning | backup_tier_skipped | Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. | controller
Sep 16 10:04 | info | node_recovered | Reports resumed (was down for 0m) | hub
Sep 16 10:03 | info | controller_started | Controller elindult (0.243.0) | controller
Sep 16 10:01 | info | claim_reissued_reenroll | Új beállító kódot küldtünk a szerver újratelepítése után (3. generáció) az ügyfél címére. | hub
Sep 16 10:01 | info | appliance_credential_delivered | Új eszköz (bare-metal telepítés) megkapta a hozzáférést és megkezdi a beállítást. | hub
Sep 16 10:01 | info | appliance_bound | Az ügyfél saját maga kötötte össze az új eszközt (bare-metal telepítés); a hozzáférést a doboz a következő lekérdezéskor megkapja. | customer_selfbind
Sep 16 09:59 | info | selfbind_link_sent | Self-bind link e-mailed (operator button) | hub
Sep 14 23:25 | error | node_down | No report received for 1h | hub
Sep 14 22:55 | warning | node_stale | No report received for 30m | hub
Sep 14 22:54 | warning | host_stale | Host tester-1-a61396: no report for 30m | hub
Sep 14 22:10 | info | node_recovered | Reports resumed (was stale for 5m) | hub
Sep 14 22:10 | info | controller_started | Controller elindult (0.242.0) | controller
Sep 14 22:04 | warning | node_stale | No report received for 30m | hub
Sep 14 21:34 | info | app_deployed | Alkalmazás telepítve: Homebox | controller
Sep 14 21:26 | info | node_recovered | Reports resumed (was stale for 1m) | hub
Sep 14 21:25 | warning | node_stale | No report received for 30m | hub
Sep 14 21:04 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: warn) | controller
Sep 14 20:54 | warning | health_degraded | Rendszer állapot romlott (volt: ok) | controller
Sep 14 20:44 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: warn) | controller
Sep 14 20:43 | info | storage_reconnected | Meghajtó újra csatlakoztatva: Adatlemez | controller
Sep 14 20:39 | warning | health_degraded | Rendszer állapot romlott (volt: ok) | controller
Sep 14 20:38 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller
Sep 14 20:38 | warning | app_start_failed | Telepített alkalmazás nem fut: Immich | controller
Sep 14 20:38 | warning | app_start_failed | Telepített alkalmazás nem fut: Jellyfin | controller
Sep 14 20:38 | warning | app_start_failed | Telepített alkalmazás nem fut: Nextcloud | controller
Sep 14 20:38 | error | storage_disconnected | Meghajtó váratlanul leválasztva: Adatlemez | controller
Sep 14 20:29 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: warn) | controller
Sep 14 20:28 | info | storage_reconnected | Meghajtó újra csatlakoztatva: Adatlemez | controller
Sep 14 19:58 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller
Sep 14 19:58 | warning | app_start_failed | Telepített alkalmazás nem fut: Jellyfin | controller
Sep 14 19:58 | warning | app_start_failed | Telepített alkalmazás nem fut: Immich | controller
Sep 14 19:58 | warning | app_start_failed | Telepített alkalmazás nem fut: Nextcloud | controller
Sep 14 19:58 | error | storage_disconnected | Meghajtó váratlanul leválasztva: Adatlemez | controller
Sep 14 19:54 | info | controller_started | Controller elindult (0.242.0) | controller
Sep 14 19:42 | info | controller_started | Controller elindult (0.242.0) | controller
Sep 14 19:42 | error | backup_failed | an app-data backup (volume dump) was interrupted by a controller restart — 1 app(s) were left stopped and have been restarted | controller
Sep 14 19:31 | info | controller_started | Controller elindult (0.242.0) | controller
## rows printed: 50
@@ -0,0 +1,19 @@
## 2026-09-16T11:15:11Z Phase 3 - morning-after sweep (in the quiet gap between F9'' kills)
public front doors:
https://paste.enkicsifelhom.hu/ -> 200
https://vault.enkicsifelhom.hu/ -> 200
https://paperless.enkicsifelhom.hu/ -> 302
https://cloud.enkicsifelhom.hu/status.php -> 200
https://fajlok.enkicsifelhom.hu/ -> 404
stack states through the dashboard API:
version labels:
controller (dashboard footer): golden on the hub for this host:
(the earlier empty read was my expired session, not a broken box - re-measured below)
cloudflared running deployed=False subdomain=
filebrowser running deployed=False subdomain=
nextcloud running deployed=True subdomain=cloud
paperless-ngx running deployed=True subdomain=paperless
privatebin running deployed=True subdomain=paste
traefik running deployed=False subdomain=
vaultwarden running deployed=True subdomain=vault
controller version on the page: 0.243.0
+3
View File
@@ -720,6 +720,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-534** | **[P1-HIGH] The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks `Datastore.Modify`, so every re-issue fails.** MEASURED 2026-09-16 on the drill box (`tester-1-652049`, fresh install from published ISO 1.27.1): the WG-registration hook refused as R-511 describes („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action"); the operator pressed exactly that, hub v0.114.0's ADOPT path ran, and the endpoint answered **`Process exited with status 255 (stderr: Error: permission check failed - missing Datastore.Modify on /datastore/felhom-offsite)`** → HTTP 502, no descriptor written, no secret stored (fail-closed, correct). So the code fix of 2026-09-15 is sound and INERT: a rebuilt box has no whole-guest off-site tier, and the customer's „Távoli rendszermentés" stays absent. **Not run by hand on ep0** (fenced). **Fix shape (operator):** grant the hub's tenantsync user `Datastore.Modify` on `/datastore/felhom-offsite` (it already holds the create/delete grants the provision path uses), or give the script a token-only re-key op that needs no Modify. Evidence: `audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt`. | **WAITING-ON-OPERATOR — rank P1-HIGH; owner: operator (ep0 grant) · CC (verify after)** |
| **R-535** | **[P2-MEDIUM] The box's console still says „a doboz készen áll, és a párosításra vár" long after the box is bound, claimed and running apps — and it promises that the screen refreshes itself.** MEASURED 2026-09-16 on the drill box (fresh install, ISO 1.27.1, controller 0.243.0): bind succeeded 10:01:14Z, claim 10:06:18Z, four apps deploying by 10:22Z — and at 10:23Z the console still showed the pairing banner with the code `37S-NFE` and the line „Ez a képernyő magától frissül — nincs teendő a doboznál" (`audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png`). A volunteer watching the monitor has no way to tell the box is finished; worse, the screen says it updates itself, so waiting longer does not help. **Fix shape:** the first-boot banner unit re-renders on claim/bind state (the controller already knows: it reports `controller_started` and the hub holds `claimed`), showing „A doboz össze van kötve — a vezérlőpult a https://felhom.<domain> címen érhető el"; or at minimum stop printing the pairing code once the appliance is claimed. | **READY — rank P2-MEDIUM; owner: CC (ISO payload `felhom-bootstrap.sh` — ships with the next ISO)** |
| **R-536** | **[P2-MEDIUM] The hub is told „Alkalmazás telepítve" the moment a deploy is ACCEPTED, so an install that never finishes is recorded as a completed one.** MEASURED 2026-09-16 on the drill box: the deploy of `mealie` was accepted at 12:31:36 CEST and the hub logged `Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie` in the SAME second; the controller was then killed 5 s in (F9'), and after the agent restarted it the stack read `not_deployed / deployed=false / deploying=false` — i.e. the app was never installed, and nothing corrected the event. Source confirms the ordering: `internal/api/router.go` writes the 202 „Telepítés elindítva" and then calls `NotifyAppDeployed` immediately, while the comment right above it says the deploy „runs asynchronously (compose pull/up + health happen after this returns)". The event is `info`, so nobody is mailed — but the customer timeline and the hub's app history record a completed install that did not happen (the „presence is not success" class). **Also measured, same shape:** an interrupted deploy leaves `/opt/docker/stacks/<app>/app.yaml` behind (written at accept) while the stack reads not-deployed — second instance after 2026-09-15's homebox; moved aside on the box. **Fix shape:** emit `app_deployed` from the async path when the stack reaches running/healthy (or emit `app_deploy_started` at accept and `app_deployed` at completion), and remove the accept-time `app.yaml` on a failed deploy. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. | **READY — rank P1-HIGH; owner: CC (controller)** |
| **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. | **READY — rank P1-HIGH; owner: CC (controller)** |
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |
| **R-526** | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** |
| **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** FOUND 2026-09-15: `stacks/metadata.go` parses it; `grep -rn LockedAfterDeploy` finds no reader; `deploy.html` renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in `02-controller-module-map.md`; the flag is a seam never wired. **Fix shape:** remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | **READY — rank P3-LOW; owner: CC** |