diff --git a/documentation/audits/DRILL-prove-fixes-0243-2026-09-16.md b/documentation/audits/DRILL-prove-fixes-0243-2026-09-16.md index 57455505..35b2beac 100644 --- a/documentation/audits/DRILL-prove-fixes-0243-2026-09-16.md +++ b/documentation/audits/DRILL-prove-fixes-0243-2026-09-16.md @@ -65,3 +65,26 @@ Vouched as a three-field change — golden 0.243.0 + agent 0.131.0 + min agent 0 **New finding on the walk: R-535 (P2)** — 25 minutes after a successful bind AND claim, with apps deploying, the box's console still showed „a doboz készen áll, és a párosításra vár" with the pairing code, and still promised „Ez a képernyő magától frissül". + +## Phase 2 — the faults + +Five faults, each recorded the same way: what the customer saw · what the box did · time to steady · +did an alarm fire and was it true · did an alarm that should have fired stay silent. Evidence per +fault in `evidence-drill-0243-2026-09-16/phase2-*.txt`. + +| fault | customer saw | box did | steady | alarm | +|---|---|---|---|---| +| **F9'** — controller killed 5 s into a deploy, empty budget | dashboard gone ~37 s; the interrupted app simply not installed | agent saw it on sweep 1, confirmed and restarted on sweep 2 (12:31:45 → 12:32:15/16 CEST) | **37 s** | `controller_restarted_by_agent` — true. **But** `app_deployed` for Mealie had already been sent at accept time → **R-536** | +| **F10** — a child deletes the photo folder | folder and 5 photos gone; after the restore the folder is **back and lists all five**, and **none opens** (`Sabre\DAV\Exception\NotFound`) | tier-1 restore replayed 3 volumes + the database in **35 s** and reported plain success | 35 s | none fired, and **none exists** for "restored database points at files that are not there" → **R-537, R-538** | +| **F11** — forgotten passwords ×5 | Nextcloud: 401 five times, no lock-out, correct password still works. The box's setup code: wrong twice, then **"Túl sok próbálkozás — próbáld újra 15 perc múlva"** from the third try; "Új beállító kód kérése" still works | claim counter locks the page for 15 min | immediate | `claim_lockout` (warning) at 13:11:42 CEST, operator mail **1 s later** — fired and **true** | +| **F12** — two reboots inside two minutes | dashboard gone ~2 min, then everything back | boot reconciler restarted all seven stacks; no double start; nothing stuck "telepítés folyamatban" | **124 s** after the second reset | none fired; correct — a reboot inside the liveness window is not an alarm | +| **M1** — memory pressure (R-528) | app restarted 9–10 times | all three OOM signals **silent**: `OOMKilled=false`, `docker events oom` empty, the cgroup invisible inside the guest, `dmesg` unreadable | — | no OOM alarm is possible on this box, same as on scratch 9202 | + +**F10 is the finding of this drill.** On a one-drive box with no off-site tier — the state every fresh +install starts in — the household's own files are in **no backup at all**: the whole-guest tiers exclude +the data drive by design (`07-backup-architecture.md`, "[FACT] What the whole-guest tiers do NOT carry", +confirmed live: `excluding bind mount point mp8 … (not a volume)`), and the app's file leg lives at +tier 2 / tier 3, both unset. The design is not the defect. The defects are that the page calls tier 1 +„DB + Konfig + Adatok" and prints the drive size beside it (**R-537**), and that a restore reports +success while leaving the app listing files it cannot open — and wipes the app's own trash, which still +held every byte (**R-538**). diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase1-manualbackup.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase1-manualbackup.txt index fc1d8e6e..9c0bca3a 100644 --- a/documentation/audits/evidence-drill-0243-2026-09-16/phase1-manualbackup.txt +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase1-manualbackup.txt @@ -12,3 +12,13 @@ 2026/09/16 12:21:58 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Paperless-ngx 2026/09/16 12:22:18 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Nextcloud page after: Rendszermentés (teljes mentés) A teljes szerver — alkalmazások, beállítások és adatbázisok együtt — időszakos mentése, amelyből az egész készülék visszaállítható. Ezt a host-ügynök készíti és kezeli. ✓ Utolsó teljes mentés 2026-09-16 12:08 (20 perce) 624.3 MB Helyi tároló (local) Naprakész Következő mentés 0 órája — a mentési ablakon belül – Visszaállítás ellenőrizve Még nem futott Helyi tároló (local) ✓ Utolsó sikeres mentés: 2026-09-16 12:08 (20 perce) · 624.3 MB Naprakész Biztonsági szerver – külön hardver (PBS) nincs beállítva Mentés folyamatban… (fázis: snapshotted ) Mentés most A mentés +## 2026-09-16T10:31:39Z manual backup finished: guest lock cleared; archive vzdump-lxc-9201-2026_09_16-12_27_11.tar.zst = 2 175 571 124 B (the scheduled one at 12:08 was 654 665 901 B, before the four apps) +## hub events/mails for tester-1 in the last 25 min: +2026/09/16 12:08:25 [INFO] Event from tester-1: backup_tier_skipped (warning) — Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. +2026/09/16 12:08:25 [INFO] Operator email sent for tester-1/backup_tier_skipped +2026/09/16 12:21:17 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: PrivateBin +2026/09/16 12:21:38 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Vaultwarden +2026/09/16 12:21:58 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Paperless-ngx +2026/09/16 12:22:18 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Nextcloud +2026/09/16 12:31:36 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie +## 2026-09-16T10:50:29Z manual backup finished (guest lock cleared) diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f10.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f10.txt new file mode 100644 index 00000000..46eb0ef8 --- /dev/null +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f10.txt @@ -0,0 +1,85 @@ +## F10 - a child deletes the photo folder (Nextcloud) - 2026-09-16T11:01:38Z + seed: folder 'Fotok' created through Nextcloud WebDAV (the app's own file interface), MKCOL 201, + 5 files nyaralas-1..5.jpg, 200k/400k/600k/800k/1000k = 3 000 000 B total, every PUT 201. + on-disk: 68204151 /mnt/felhom-drives/adatlemez/appdata/nextcloud + on-disk: 5 + on-disk: 25337 /mnt/felhom-drives/adatlemez/backups + 2026-09-16T11:01:57Z pressing 'Mentes most' on the app-backup page (POST /api/backup/run) + trigger: {"ok":true,"message":"Mentés elindítva"} + http=200 + 2026-09-16T11:02:58Z backup finished; status: {"ok":true,"data":{"db_dump":{"count":2,"duration":"47.936765144s","last_run":"2026-09-16T11:02:45.477832211Z","success":true},"enabled":true,"running":false}} + restore page after the app backup: + TEXT: Pillanatkép: — Válasszon alkalmazást — Még nincs mentés felhasználói adattal. A visszaállítás felülírja az alkalmazás jelenlegi adatait a kiválasztott mentés állapotával. Az alkalmazás a folyamat során automatikusan leáll és újraindul. Megértettem, visszaállítás indítása. Visszaállítás indítása Importálás mentett csomagból (.fab) Hordozható mentéscsomag (.fab) Hordozható pillanatfelvétel — bárhol tárolhatod, és bármikor visszatöltheted egy meghajtóról. A folyamatos védelmet az 1–3. szintű mentés adja. Jelszavas titkosítás (opcionális) Üresen hagyva a csomag titkosítás nélkül készül. Nextcloud Letöltés (.fab) Paperless-ngx Letöltés (.fab) PrivateBin Letöltés (.fab) Vaultwarden Letöltés (.fab) + 2026-09-16T11:05:06Z the child deletes the folder: DELETE /Fotok http=204 + folder after the delete: PROPFIND http=404 (404 = gone) +## MEASURED before the restore attempt: +## 1) The app backup ("1. mentes", tier 1) was made at 11:02:45Z by the customer-visible "Mentes most" +## button (POST /api/backup/run -> {"ok":true,"message":"Mentes elindItva"}; the backup folder grew +## 25 337 B -> 978 MB, so the button does more than its status field, which reports db_dump only). +## 2) Its manifest lists db-dumps + THREE docker volume dumps (html, db_data, redis). It does NOT list +## the drive-side app data (/mnt/felhom-drives/adatlemez/appdata/nextcloud), where the photos live. +## 3) Proof by listing the html tar (29 346 entries): +## positive control "version.php" = 3 hits (the listing works) +## "data/" entries = 1284, but "./data/" is the EMPTY bind-mount point +## "Fotok" = 0 hits +## "nyaralas"= 0 hits +## And: find over the whole backups tree for *appdata*/*hdd*/*Fotok* = nothing. +## 4) The whole-guest backup does not cover them either - vzdump log of the 12:27 CEST run: +## "including mount point rootfs ('/')", "including mount point mp0 ('/var/lib/felhom')", +## "excluding bind mount point mp8 ('/mnt/felhom-drives') from backup (not a volume)". +## 5) The app-backup page nevertheless labels tier 1 "DB + Konfig + Adatok" and prints "Nextcloud +## Adatlemez 65.1 MB" - a size measured on exactly the data it does not copy. +## 6) Tier 2 ("off-drive masolat") and tier 3 (offsite) both read "Nincs beallitva" on this box. + 2026-09-16T11:06:56Z pressing the restore button: POST /backup/restore stack_name=nextcloud snapshot_id=helyi + restore POST http=302 + 2026-09-16T11:07:42Z restore status: {"ok":true,"data":{"running":false,"op":"restore","stack":"nextcloud","started_at":"2026-09-16T11:06:56.710474757Z","last":{"op":"restore","stack":"nextcloud","ok":true,"message":"A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult.","finished_at":"2026-09-16T11:07:31.194473339Z"},"last_recent":true}} + after the restore, through Nextcloud's own interface: + PROPFIND /Fotok http=207 + file: + file: nyaralas-1.jpg + file: nyaralas-2.jpg + file: nyaralas-3.jpg + file: nyaralas-4.jpg + file: nyaralas-5.jpg + GET /Fotok/nyaralas-1.jpg http=503 bytes=276 + trash PROPFIND http=207 + trash: + RE-MEASURE at 2026-09-16T11:08:38Z, app healthy (nextcloud Up >2 min): + app status.php http=200 <- positive control, app is serving + PUT control kontroll.txt http=201 <- positive control, WebDAV write works + GET control kontroll.txt http=200 bytes=6 <- positive control, WebDAV read works + GET /Fotok/nyaralas-1.jpg http=404 bytes=249 + GET /Fotok/nyaralas-2.jpg http=503 bytes=276 + GET /Fotok/nyaralas-3.jpg http=503 bytes=276 + GET /Fotok/nyaralas-4.jpg http=503 bytes=276 + GET /Fotok/nyaralas-5.jpg http=503 bytes=276 + body of the failed download (first 200 chars): + Sabre\DAV\Exception\NotFound File with name /Fotok/nyaralas-1 +## RESULT - F10 (a child deletes the photo folder, Nextcloud, fresh box, one drive, no tier 2/3) +## customer saw: the folder and the 5 photos vanish from Nextcloud (DELETE 204, PROPFIND 404). +## After the restore the folder is BACK and lists all 5 photos - but NONE of them opens: +## GET nyaralas-1..5 = 404 / 503 x4, body "Sabre\DAV\Exception\NotFound". +## Positive controls at the same moment: status.php 200, WebDAV PUT kontroll.txt 201, GET 200 (6 B). +## So the failure is real, not a starting app and not a broken login. +## box did: POST /backup/restore (stack_name=nextcloud, snapshot_id=helyi) -> 302, finished in 35 s, +## "A(z) nextcloud: 3 adatkotet es az adatbazis visszaallitva - az alkalmazas ujraindult." +## It replayed the 3 docker volumes + the MariaDB dump. The DB dump (11:01:45Z) KNOWS the 5 photos, +## because they were uploaded at ~11:00Z. The BYTES were never in any backup (see MEASURED above). +## time to steady: 35 s restore, app healthy ~50 s later. +## how much was lost: all 5 files, 3 000 000 B - every photo the folder ever held. +## does the page say so? NO. The result line says the volumes and the database were restored and the +## app restarted. Nothing says the app's files on the data drive are not part of this backup, and +## nothing says the restored database now points at files that do not exist. +## where the bytes actually are: Nextcloud's OWN trash on the data drive - +## appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg (all 5 present). +## But the restored database no longer lists them: the trash PROPFIND returns an EMPTY list, so the +## customer cannot press "restore from trash" either. The restore made the trash unreachable. +## alarm fired and true? none fired. +## alarm that should have and did not: none is defined for "restored DB references missing files". +## DESIGN CHECK (a design is not a defect, R-370): 07-backup-architecture.md is explicit - +## "[FACT] What the whole-guest tiers do NOT carry ... mp8 /mnt/felhom-drives ... out of vzdump scope", +## and the file leg for nextcloud exists at TIER 2 and TIER 3 only (section 6.2 table: Tier-3 +## mandatory = calibre-web, immich, nextcloud, paperless-ngx). Tier 1 carrying no drive-side files is +## therefore BY DESIGN. What is NOT by design is the page calling tier 1 "DB + Konfig + Adatok" and +## printing "Adatlemez 65.1 MB" next to it, and a restore reporting plain success while leaving the +## app inconsistent. Those two are the rows. diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f11.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f11.txt new file mode 100644 index 00000000..55f665b2 --- /dev/null +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f11.txt @@ -0,0 +1,61 @@ +## F11 - forgotten passwords, part 1: Nextcloud, 5 wrong tries 2026-09-16T11:09:35Z + try 1: http=401 (took 868 ms) + try 2: http=401 (took 726 ms) + try 3: http=401 (took 702 ms) + try 4: http=401 (took 725 ms) + try 5: http=401 (took 687 ms) + now the CORRECT password again (positive control - is the customer locked out?): + http=200 (took 610 ms) 207 = still allowed in + web login page http=200 +## F11 part 2 - wrong CLAIM code x5, measured on this box (already claimed at 10:06:18Z) + the claim page itself: http=200 + page says: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új beállító kód kérése Felhom — Otthoni szerver kezelés felhom.eu + wrong claim try 1: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m + wrong claim try 2: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m + wrong claim try 3: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m + wrong claim try 4: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m + wrong claim try 5: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m +## F11 part 2 - CORRECTED run (my first run was refused by CSRF - I measured my own mistake, not the product) + wrong code try 1: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új + wrong code try 2: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új + wrong code try 3: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új + wrong code try 4: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új + wrong code try 5: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új + NOTE: this box was already claimed at 10:06:18Z. The /claim URL still answers 200 and now renders + 'Jelszo visszaallitasa' (password reset) with a 'Uj beallito kod kerese' button - so the claim + page is NOT one-shot-dead after a successful claim; it becomes the reset surface. +## F11 part 2 - FINAL run (correct CSRF: cookie felhom_claim_csrf + _csrf field + X-CSRF-Token header) + wrong code try 1: http=200 (198 ms) page says: Hibás vagy lejárt kód Add meg az e-mailben kapott beállító kódot, majd válassz új + wrong code try 2: http=200 (187 ms) page says: Hibás vagy lejárt kód Add meg az e-mailben kapott beállító kódot, majd válassz új + wrong code try 3: http=200 (177 ms) page says: Túl sok próbálkozás — próbáld újra 15 perc múlva. Add meg az e-mailben kapott beá + wrong code try 4: http=200 (209 ms) page says: Túl sok próbálkozás — próbáld újra 15 perc múlva. Add meg az e-mailben kapott beá + wrong code try 5: http=200 (172 ms) page says: Túl sok próbálkozás — próbáld újra 15 perc múlva. Add meg az e-mailben kapott beá + CORRECTION of my own summary line: there IS a lock-out, and it arrives on the THIRD try. + Tries 1-2: "Hibas vagy lejart kod". Tries 3-5: "Tul sok probalkozas - probald ujra 15 perc mulva." + No captcha and no growing delay (172-209 ms throughout) - it is a counter, not a slow-down. + the household's escape route while locked out - pressing 'Uj beallito kod kerese': + POST /claim/request-new-code http=200 + page says: Ha az e-mail cím regisztrálva van, elküldtük a kódot. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jel +## RESULT - F11 (forgotten passwords x5, both kinds) +## A) Nextcloud app password, 5 wrong tries: +## customer saw: 401 five times, each ~700 ms, no lock-out, no captcha; the CORRECT password then +## worked immediately (200). Positive controls in the same run: status.php 200, WebDAV PUT 201/GET 200. +## box did: nothing - no event, no mail, no restart. +## time to steady: none needed. +## alarm fired and true? none. alarm that should have and did not: none is expected for an app login. +## B) The box's own setup code (the claim page), 5 wrong tries: +## customer saw: tries 1-2 "Hibas vagy lejart kod"; tries 3-5 "Tul sok probalkozas - probald ujra +## 15 perc mulva." So the lock-out starts on the THIRD attempt and lasts 15 minutes. Timing was +## flat (172-209 ms), so it is a counter, not a slow-down. +## the escape route WORKS while locked out: POST /claim/request-new-code -> 200 and +## "Ha az e-mail cim regisztralva van, elkuldtuk a kodot." (deliberately neutral wording). +## one-shot behaviour, MEASURED not assumed: this box was claimed at 10:06:18Z, and /claim still +## answers 200 afterwards - it becomes the "Jelszo visszaallitasa" (password reset) surface with a +## "Uj beallito kod kerese" button. The URL is NOT dead after a successful claim. +## NOT WALKED: the same five wrong codes on a SECOND, never-claimed box. That needs a second fresh +## install (~35 min) and the drill box was needed for F9''. So the "first-ever claim" variant of +## the counter is untested; what is proven here is the reset path on a claimed box. +## MY OWN MISTAKE, recorded: my first two runs of B were refused with "Ervenytelen urlap" because I +## paired the claim page's pre-auth token wrongly (the page sets cookie felhom_claim_csrf and carries +## TWO _csrf fields). I measured my own error, said so, and re-ran with cookie + _csrf field + +## X-CSRF-Token header, which the product accepted. diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f12.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f12.txt new file mode 100644 index 00000000..f18afcdf --- /dev/null +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f12.txt @@ -0,0 +1,70 @@ +## F12 — two reboots inside two minutes (box: nested VM 334 on demo-hp) + 2026-09-16T10:50:10Z before: dashboard=200 apps=200 + reset 1 at 2026-09-16T10:50:11Z + reset 2 at 2026-09-16T10:51:12Z (60 s after the first) + dashboard 200 again at 2026-09-16T10:53:16Z — 124 s after the second reset + apps: privatebin=200 vault=200 paperless=302 cloud=200 +Sep 16 12:53:06 tester1 felhom-agent[1128]: time=2026-09-16T12:53:06.326+02:00 level=WARN msg="controller-supervisor: guest list unavailable — skipping sweep (ownership unproven)" err="pro +Sep 16 12:53:07 tester1 felhom-agent[1128]: time=2026-09-16T12:53:07.077+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1 +Sep 16 12:53:36 tester1 felhom-agent[1128]: time=2026-09-16T12:53:36.356+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1 +Sep 16 12:53:36 tester1 felhom-agent[1128]: time=2026-09-16T12:53:36.385+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1 +Sep 16 12:53:36 tester1 felhom-agent[1128]: time=2026-09-16T12:53:36.388+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1 +Sep 16 12:54:06 tester1 felhom-agent[1128]: time=2026-09-16T12:54:06.388+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1 +paperless-webserver Up 34 seconds (healthy) +paperless-redis Up 44 seconds (healthy) +paperless-postgres Up 44 seconds (healthy) +nextcloud Up 56 seconds (healthy) +nextcloud-db Up About a minute (healthy) +nextcloud-redis Up About a minute (healthy) +felhom-controller Up About a minute (healthy) +vaultwarden Up About a minute (healthy) +privatebin Up About a minute (healthy) +filebrowser Up About a minute (healthy) +cloudflared Up About a minute +traefik Up About a minute +## app states after the two resets (2026-09-16T10:57:17Z): + +## supervisor block in the host report: +Traceback (most recent call last): + File "", line 1, in + import json;print(json.load(open("/var/lib/felhom-agent/bootstrap.json"))["local_api_token"]) + ~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +FileNotFoundError: [Errno 2] No such file or directory: '/var/lib/felhom-agent/bootstrap.json' +Traceback (most recent call last): + File "", line 1, in + import sys,json;r=json.load(sys.stdin);print(json.dumps(r.get("controller_supervisor"),indent=1)) + ~~~~~~~~~^^^^^^^^^^^ + File "/usr/lib/python3.13/json/__init__.py", line 293, in load + return loads(fp.read(), + cls=cls, object_hook=object_hook, + parse_float=parse_float, parse_int=parse_int, + parse_constant=parse_constant, object_pairs_hook=object_pairs_hook, **kw) + File "/usr/lib/python3.13/json/__init__.py", line 346, in loads + return _default_decoder.decode(s) + ~~~~~~~~~~~~~~~~~~~~~~~^^^ + File "/usr/lib/python3.13/json/decoder.py", line 345, in decode + obj, end = self.raw_decode(s, idx=_w(s, 0).end()) + ~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^ + File "/usr/lib/python3.13/json/decoder.py", line 363, in raw_decode + raise JSONDecodeError("Expecting value", s, err.value) from None +json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0) +## uptime + boot count: +up 5 minutes +reboot system boot 7.0.2-6-pve Wed Sep 16 12:51 - still running +reboot system boot 7.0.2-6-pve Wed Sep 16 12:50 - crash +reboot system boot 7.0.2-6-pve Wed Sep 16 11:58 - crash + +## RESULT — F12 (two `qm reset`s 60 s apart on VM 334) +## customer saw: dashboard unreachable ~2 min; back at 10:53:16Z, 124 s after the SECOND reset. +## box did: boot reconciler brought all seven stacks up by itself (privatebin 200, vaultwarden 200, +## paperless 302 = its login redirect, nextcloud status.php 200); NO double start observed; +## mealie stayed not_deployed (the F9'-interrupted deploy), NOT stuck in "telepites folyamatban". +## time to steady: 124 s to the dashboard, ~3 min to every app healthy. +## supervisor did NOT count the boots: whole-journal "RESTARTED the controller" = 1 (the F9' one at +## 12:32, POSITIVE control that the grep string matches when it happens); since 12:49 = 0, +## while 4 controller-supervisor lines in the same window prove it was running and sweeping. +## During the boot it logged "guest list unavailable - skipping sweep (ownership unproven)" - +## the unprovisioned/ownership guard doing its job. +## journal is PERSISTENT on this box (/var/log/journal exists), so the pre-reboot lines are real. +## alarm fired and true? none fired. alarm that should have and did not: none - a reboot inside the +## node-liveness dedupe window produces no host_* mail, which matches the design. diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f9dprime.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f9dprime.txt new file mode 100644 index 00000000..58ec7b71 --- /dev/null +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f9dprime.txt @@ -0,0 +1,13 @@ +## F9'' - three more controller kills, 20 minutes apart, at IDLE (no deploy running) +## Question to MEASURE (not to change): does a restart after 20 minutes of healthy uptime +## count against the 3-restarts-per-15-minutes budget, or does the window start fresh? +## Baseline: the last restart before this run was F9' at 12:32:16 CEST (1 of 3 at that time). +## kill 1 at 2026-09-16T11:12:41Z +felhom-controller + dashboard 200 again after 61 s +Sep 16 13:12:37 tester1 felhom-agent[1128]: time=2026-09-16T13:12:37.326+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=40 guests_evaluated=1 controllers_not_running=0 +Sep 16 13:13:07 tester1 felhom-agent[1128]: time=2026-09-16T13:13:07.323+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status +Sep 16 13:13:37 tester1 felhom-agent[1128]: time=2026-09-16T13:13:37.415+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit +Sep 16 13:13:38 tester1 felhom-agent[1128]: time=2026-09-16T13:13:38.655+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse + RESTARTED lines in the whole journal so far: 2 + waiting 20 minutes before the next kill diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt new file mode 100644 index 00000000..148d70e9 --- /dev/null +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt @@ -0,0 +1,19 @@ +## F9' — controller killed 5 s into a deploy, on a FRESH crash-loop budget (agent started 12:02, no restarts yet) + 2026-09-16T10:31:36Z starting the mealie deploy + deploy http=202 +{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"} + killed at 2026-09-16T10:31:41Z — 5 s into the deploy (the "(37901 s …)" text my script printed here was a timestamp-parsing bug of mine; the real interval is the scripted sleep of 5 s between the 202 and the kill) + dashboard health 200 again at 2026-09-16T10:32:18Z — 37 s after the kill + mealie after: {'state': 'not_deployed', 'deployed': False, 'deploying': False} +Sep 16 12:31:45 tester1 felhom-agent[3602]: time=2026-09-16T12:31:45.160+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status +Sep 16 12:32:15 tester1 felhom-agent[3602]: time=2026-09-16T12:32:15.231+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit +Sep 16 12:32:16 tester1 felhom-agent[3602]: time=2026-09-16T12:32:16.606+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse +Sep 16 12:32:16 tester1 felhom-agent[3602]: time=2026-09-16T12:32:16.606+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=60 guests_evaluated=1 controllers_not_running=1 +## RESULT (R-531's missing timing, now pinned): +## kill 5 s into a deploy, on an EMPTY crash-loop budget → dashboard health 200 again 37 s after the kill +## (agent saw it at 12:31:45, confirmed and restarted at 12:32:15/16 — two sweeps, as designed). +## The interrupted app ended `not_deployed / deployed=false / deploying=false` — NOT stuck in "telepítés folyamatban". +## Budget after this single restart: 1 of 3 in the 15-minute window. +2026-09-16 10:31:36.748187371 +0000 253 /opt/docker/stacks/mealie/app.yaml +deployed deployed_at env locked_fields desired_state +moved-aside diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt new file mode 100644 index 00000000..8c245dc1 --- /dev/null +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt @@ -0,0 +1,65 @@ + are supported and installed on your system. +## 2026-09-16T10:35:16Z M1 — OOM signals on THIS box (R-528). before: +limit=1342177280 oom=false restarts=0 status=running +81: memory: 128M +134: memory: 128M +## 2026-09-16T10:37:30Z after the 128M cap: false restarting 9 + cgroup: not visible + docker oom events since 2026-09-16T10:35:29Z: + kernel OOM lines: 0 +dmesg not readable in guest +2 +## 2026-09-16T10:40:13Z cap restored: limit=1342177280 health=unhealthy restarts=10 + are supported and installed on your system. +## 2026-09-16T10:40:44Z memory lines now: +81: memory: 1280M +112: memory: 256M +134: memory: 1280M +81: memory: 1280M +112: memory: 256M +134: memory: 128M +## 2026-09-16T10:47:37Z paperless after fix: ws=unhealthy limit=1342177280 redis=134217728 +## M1 RESULT (R-528), on THIS box (nested VM, customer LXC 9201, Docker in the guest): +## cap 128M → paperless-webserver RESTARTING with RestartCount 9–10, and: +## - docker inspect .State.OOMKilled = FALSE +## - docker events --filter event=oom (since the cap change) = EMPTY +## - the container's cgroup is NOT visible from inside the guest (find under /sys/fs/cgroup → nothing), +## so memory.events / oom_kill could not be read there +## - dmesg is not readable in the guest +## So all three signals are silent here, exactly as on scratch 9202 (2026-09-15). BIGNIGHT's VM 333 did read +## oomkilled=true for the worker-killed-inside-a-running-container shape; a MAIN-process kill + restart reports +## nothing. The controller v0.243.0 OOM tag therefore cannot fire for this (restart) shape on a real box. +## MISTAKE, stated: my cap edit used a blunt sed, so the restore rewrote the REDIS limit too (128M → 1280M) — +## the same mistake as 2026-09-15 on 9202. Repaired in the next step (redis back to 128M, webserver 1280M). +paperless-redis Up 7 minutes (healthy) +paperless-webserver Restarting (1) 57 seconds ago +paperless-postgres Up 12 minutes (healthy) +--- + File "/usr/local/lib/python3.12/site-packages/django/db/backends/base/base.py", line 256, in connect + self.connection = self.get_new_connection(conn_params) + ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + File "/usr/local/lib/python3.12/site-packages/django/utils/asyncio.py", line 26, in inner + return func(*args, **kwargs) + ^^^^^^^^^^^^^^^^^^^^^ + File "/usr/local/lib/python3.12/site-packages/django/db/backends/postgresql/base.py", line 332, in get_new_connection + connection = self.Database.connect(**conn_params) + ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + File "/usr/local/lib/python3.12/site-packages/psycopg/connection.py", line 120, in connect + raise last_ex.with_traceback(None) +django.db.utils.OperationalError: connection failed: connection to server at "172.20.0.3", port 5432 failed: FATAL: password authentication failed for user "paperless" +s6-rc: warning: unable to start service init-migrations: command exited 1 +/run/s6/basedir/scripts/rc.init: warning: s6-rc failed to properly bring all the services up! Check your logs (in /run/uncaught-logs/current if you have in-container logging) for more information. +/run/s6/basedir/scripts/rc.init: fatal: stopping the container. +## 2026-09-16T10:48:27Z repairing paperless through the CONTROLLER's own path (my manual compose up broke its env): + stop: {"ok":true,"message":"Stack paperless-ngx stop completed"} + http=200 + start: {"ok":true,"message":"Stack paperless-ngx start completed"} + http=200 + paperless state after the controller-driven restart: running at 2026-09-16T10:49:35Z +## MISTAKE, stated (2026-09-16): the M1 measurement recreated the paperless webserver with a bare +## `docker compose up -d` INSIDE the guest. The controller injects the app's environment (DB password etc.) from +## `app.yaml` at deploy time, so a hand-run compose brings the container up WITHOUT it: the webserver then failed +## `password authentication failed for user "paperless"` against its own postgres and crash-looped. +## Nothing about the product; the fault was the manual path. Repaired by restarting the stack through the +## CONTROLLER's own stop/start API, which re-renders the environment. Lesson for the next measurement: drive app +## containers through the controller, or set the cap via `docker update --memory`, never by re-running compose. diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase3-alarm-truth-table.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-alarm-truth-table.txt new file mode 100644 index 00000000..090d56fc --- /dev/null +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-alarm-truth-table.txt @@ -0,0 +1,48 @@ +## Alarm truth table — operator mails actually DELIVERED during the drill (read from the mailbox, not from the hub log) +## (to admin@felhom.eu, via the Gmail connector, 2026-09-16) + 1. 12:04:01 CEST [Felhom] ✅ tester-1: node_recovered (severity INFO) — "Reports resumed (was down for 0m)". + TRUE but odd: the box had never been down; it was newly enrolled. Also: severity INFO was MAILED, while the + project's rule says severityNotifies DROPS info before both legs — to be checked against the dispatcher (below). + 2. 12:08:26 CEST [Felhom] ⚠️ tester-1: backup_tier_skipped (severity warning) — "Whole-guest backup tier felhom-pbs + skipped: its storage does not exist on the host (never provisioned or removed)". TRUE: R-534 blocked the tier, so + the tier's storage is genuinely absent. This is R-518's cheap half firing on a real box, with its operator mail. + No other operator mail arrived in the window (the third thread in the mailbox is a DMARC report, unrelated). +## Faults injected so far: none (Phase 2 starts with F9'). Alarms that SHOULD have fired and did not: none so far. +## Why the INFO mail is correct (checked in code, not assumed): +## `node_recovered` is emitted with severity "info" (monitor/staleness.go emitTransition) and handed to +## dispatcher.ProcessEvent, which returns before both legs for "info" — BUT the recovery branch +## (`recoveredPairedDownTypes`, dispatcher.go:124) runs BEFORE that gate, deliberately, so the operator hears the +## all-clear for a down it was told about. The customer leg stays pairing-gated. +## In context the mail is also TRUE: customer `tester-1` had been reporting nothing since the BIGNIGHT teardown, so the +## new box's first report IS a recovery, and "(was down for 0m)" is the age of the gap the checker could see. +## No row filed. + +## ALARM TRUTH TABLE - second pass, 2026-09-16 (hub 0.115.0), built from the hub's own log lines +## (pod hub-85478f77d7-mpxkr, times CEST) paired with the customer timeline on /customers/tester-1. +## +## event when severity operator mail? true in context? +## selfbind_link_sent 11:59:55 info yes (self-bind e-mail) yes - the operator pressed it +## claim_reissued_reenroll 12:01:56 info yes (claim code to customer) yes - 3rd generation after reinstall +## controller_started 12:03:30 info no correct - info, not paired +## node_recovered 12:04:00 info YES - mail sent BY DESIGN: the recovery branch runs +## before the severity gate; true here +## backup_tier_skipped 12:08:25 warning yes yes - the off-site tier has no storage +## backup_tier_skipped 12:37:17 warning SUPPRESSED - cooldown correct - same key within the hour +## (key=tester-1:backup_tier_skipped:felhom-pbs) +## app_deployed x5 12:21-12:31 info no NOT always true -> R-536 (sent at accept) +## app_start_failed (paperless) 12:48:47 warning yes TRUE - and it was MY damage (the +## hand-run compose, recorded in phase2-m1) +## controller_restarted_by_agent 12:48 info no (info) yes - F9' restart #1, timestamps match +## backup_tier_skipped 12:58:13 warning SUPPRESSED - cooldown correct +## claim_lockout 13:11:42 warning YES, 1 s later TRUE - F11 part 2, third wrong code +## claim reset code (gen 4) 13:12:05 info yes (to the customer) yes - "Uj beallito kod kerese" +## +## What this pass adds to the first one: +## 1) The 5-minute node/host dedupe (R-529) was NOT exercised today: no node_* or host_* alarm fired +## after 12:04, because the box stayed up and reporting. The two reboots of F12 were far too short +## to reach the 30-minute staleness threshold. So the widened bypass is SHIPPED but UNPROVEN-LIVE +## on this box; the proof from 2026-09-15 on the other box still stands. +## 2) The one-hour operator cooldown is PROVEN here twice, by its own log line, on backup_tier_skipped. +## 3) Every warning that fired today was true. No warning that should have fired stayed silent, EXCEPT +## the class R-538 names: there is no alarm at all for "a restore finished and the app is now +## inconsistent", so nothing could have fired. diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase3-ep0-after.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-ep0-after.txt new file mode 100644 index 00000000..eb0f1779 --- /dev/null +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-ep0-after.txt @@ -0,0 +1,39 @@ +## 2026-09-16T11:14:14Z ep0 READ-ONLY listing of tester-1's namespace (nothing removed, nothing written) ++================+====================+=========+ +| name | path | comment | ++================+====================+=========+ +| felhom-offsite | /mnt/pbs-datastore | | ++================+====================+=========+ +--- groups in namespace tester-1: +Error: error building client for repository felhom-offsite - no password input mechanism available +## 2026-09-16T11:14:24Z ep0 listing by filesystem (read-only; the client needs a password we deliberately do not hold here) +--- namespaces: +demo-felhom +demo-hp +tester-1 +--- tester-1 groups: +total 12 +drwxr-xr-x 3 backup backup 4096 Sep 14 16:04 . +drwxr-xr-x 5 backup backup 4096 Sep 14 15:48 .. +drwxr-xr-x 2 backup backup 4096 Sep 15 08:33 ct +--- ct groups and their snapshots: +## 2026-09-16T11:14:42Z ep0 deeper listing (read-only) +--- tester-1/ct contents: +total 8 +drwxr-xr-x 2 backup backup 4096 Sep 15 08:33 . +drwxr-xr-x 3 backup backup 4096 Sep 14 16:04 .. +--- snapshots per group: + group: /mnt/pbs-datastore/ns/tester-1/ct/*/ +--- other namespaces, for scale: +9201 +## CONCLUSION (nothing was removed by this session; ep0 was read only) +## Before the drill (09:12:21Z): ns tester-1 existed with an EMPTY ct group dir - the old box's +## snapshots had been removed the previous night, with the operator's explicit yes. +## After the whole drill (11:14:42Z): ns tester-1/ct is STILL EMPTY. +## Why: this box never had an off-site tier at all. The re-issue failed on the endpoint token's missing +## Datastore.Modify grant (R-534), so no descriptor, no secret, no upload path. The box therefore +## wrote ZERO snapshots off-site during the entire drill. +## For scale, the neighbouring namespace demo-hp/ct holds group 9201 - proof that this listing method +## DOES show groups when they exist (positive control). +## Consequence for Phase 3: "restore one DB-backed app from the off-site tier onto scratch 9202" +## CANNOT BE WALKED on this box. Not skipped for time - there is no off-site copy to restore from. diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase3-events-timeline.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-events-timeline.txt new file mode 100644 index 00000000..3e73f1aa --- /dev/null +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-events-timeline.txt @@ -0,0 +1,52 @@ +## tester-1 event timeline as the hub shows it (customer page), read 2026-09-16T11:13:18Z + Sep 16 11:11 | warning | claim_lockout | Túl sok hibás beállító kód — a beállító oldal 15 percre zárolva | controller + Sep 16 10:58 | warning | backup_tier_skipped | Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. | controller + Sep 16 10:53 | info | controller_started | Controller elindult (0.243.0) | controller + Sep 16 10:48 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller + Sep 16 10:48 | info | controller_restarted_by_agent | Host tester-1-652049 guest 9201: the agent restarted the controller (controller container exited on 2 consecutive sweeps) — restart #1 since the agent started | hub + Sep 16 10:37 | warning | backup_tier_skipped | Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. | controller + Sep 16 10:32 | info | controller_started | Controller elindult (0.243.0) | controller + Sep 16 10:31 | info | app_deployed | Alkalmazás telepítve: Mealie | controller + Sep 16 10:27 | info | recovery_credential_revealed | Az üzemeltető lekérte a géped konzolos hozzáférési jelszavát (távoli hibaelhárítás). | hub + Sep 16 10:22 | info | app_deployed | Alkalmazás telepítve: Nextcloud | controller + Sep 16 10:21 | info | app_deployed | Alkalmazás telepítve: Paperless-ngx | controller + Sep 16 10:21 | info | app_deployed | Alkalmazás telepítve: Vaultwarden | controller + Sep 16 10:21 | info | app_deployed | Alkalmazás telepítve: PrivateBin | controller + Sep 16 10:08 | warning | backup_tier_skipped | Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. | controller + Sep 16 10:04 | info | node_recovered | Reports resumed (was down for 0m) | hub + Sep 16 10:03 | info | controller_started | Controller elindult (0.243.0) | controller + Sep 16 10:01 | info | claim_reissued_reenroll | Új beállító kódot küldtünk a szerver újratelepítése után (3. generáció) az ügyfél címére. | hub + Sep 16 10:01 | info | appliance_credential_delivered | Új eszköz (bare-metal telepítés) megkapta a hozzáférést és megkezdi a beállítást. | hub + Sep 16 10:01 | info | appliance_bound | Az ügyfél saját maga kötötte össze az új eszközt (bare-metal telepítés); a hozzáférést a doboz a következő lekérdezéskor megkapja. | customer_selfbind + Sep 16 09:59 | info | selfbind_link_sent | Self-bind link e-mailed (operator button) | hub + Sep 14 23:25 | error | node_down | No report received for 1h | hub + Sep 14 22:55 | warning | node_stale | No report received for 30m | hub + Sep 14 22:54 | warning | host_stale | Host tester-1-a61396: no report for 30m | hub + Sep 14 22:10 | info | node_recovered | Reports resumed (was stale for 5m) | hub + Sep 14 22:10 | info | controller_started | Controller elindult (0.242.0) | controller + Sep 14 22:04 | warning | node_stale | No report received for 30m | hub + Sep 14 21:34 | info | app_deployed | Alkalmazás telepítve: Homebox | controller + Sep 14 21:26 | info | node_recovered | Reports resumed (was stale for 1m) | hub + Sep 14 21:25 | warning | node_stale | No report received for 30m | hub + Sep 14 21:04 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: warn) | controller + Sep 14 20:54 | warning | health_degraded | Rendszer állapot romlott (volt: ok) | controller + Sep 14 20:44 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: warn) | controller + Sep 14 20:43 | info | storage_reconnected | Meghajtó újra csatlakoztatva: Adatlemez | controller + Sep 14 20:39 | warning | health_degraded | Rendszer állapot romlott (volt: ok) | controller + Sep 14 20:38 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller + Sep 14 20:38 | warning | app_start_failed | Telepített alkalmazás nem fut: Immich | controller + Sep 14 20:38 | warning | app_start_failed | Telepített alkalmazás nem fut: Jellyfin | controller + Sep 14 20:38 | warning | app_start_failed | Telepített alkalmazás nem fut: Nextcloud | controller + Sep 14 20:38 | error | storage_disconnected | Meghajtó váratlanul leválasztva: Adatlemez | controller + Sep 14 20:29 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: warn) | controller + Sep 14 20:28 | info | storage_reconnected | Meghajtó újra csatlakoztatva: Adatlemez | controller + Sep 14 19:58 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller + Sep 14 19:58 | warning | app_start_failed | Telepített alkalmazás nem fut: Jellyfin | controller + Sep 14 19:58 | warning | app_start_failed | Telepített alkalmazás nem fut: Immich | controller + Sep 14 19:58 | warning | app_start_failed | Telepített alkalmazás nem fut: Nextcloud | controller + Sep 14 19:58 | error | storage_disconnected | Meghajtó váratlanul leválasztva: Adatlemez | controller + Sep 14 19:54 | info | controller_started | Controller elindult (0.242.0) | controller + Sep 14 19:42 | info | controller_started | Controller elindult (0.242.0) | controller + Sep 14 19:42 | error | backup_failed | an app-data backup (volume dump) was interrupted by a controller restart — 1 app(s) were left stopped and have been restarted | controller + Sep 14 19:31 | info | controller_started | Controller elindult (0.242.0) | controller +## rows printed: 50 diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase3-morning-after.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-morning-after.txt new file mode 100644 index 00000000..923aab46 --- /dev/null +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-morning-after.txt @@ -0,0 +1,19 @@ +## 2026-09-16T11:15:11Z Phase 3 - morning-after sweep (in the quiet gap between F9'' kills) + public front doors: + https://paste.enkicsifelhom.hu/ -> 200 + https://vault.enkicsifelhom.hu/ -> 200 + https://paperless.enkicsifelhom.hu/ -> 302 + https://cloud.enkicsifelhom.hu/status.php -> 200 + https://fajlok.enkicsifelhom.hu/ -> 404 + stack states through the dashboard API: + version labels: + controller (dashboard footer): golden on the hub for this host: + (the earlier empty read was my expired session, not a broken box - re-measured below) + cloudflared running deployed=False subdomain= + filebrowser running deployed=False subdomain= + nextcloud running deployed=True subdomain=cloud + paperless-ngx running deployed=True subdomain=paperless + privatebin running deployed=True subdomain=paste + traefik running deployed=False subdomain= + vaultwarden running deployed=True subdomain=vault + controller version on the page: 0.243.0 diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index d34ec994..7ec060e1 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -720,6 +720,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **R-534** | **[P1-HIGH] The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks `Datastore.Modify`, so every re-issue fails.** MEASURED 2026-09-16 on the drill box (`tester-1-652049`, fresh install from published ISO 1.27.1): the WG-registration hook refused as R-511 describes („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action"); the operator pressed exactly that, hub v0.114.0's ADOPT path ran, and the endpoint answered **`Process exited with status 255 (stderr: Error: permission check failed - missing Datastore.Modify on /datastore/felhom-offsite)`** → HTTP 502, no descriptor written, no secret stored (fail-closed, correct). So the code fix of 2026-09-15 is sound and INERT: a rebuilt box has no whole-guest off-site tier, and the customer's „Távoli rendszermentés" stays absent. **Not run by hand on ep0** (fenced). **Fix shape (operator):** grant the hub's tenantsync user `Datastore.Modify` on `/datastore/felhom-offsite` (it already holds the create/delete grants the provision path uses), or give the script a token-only re-key op that needs no Modify. Evidence: `audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt`. | **WAITING-ON-OPERATOR — rank P1-HIGH; owner: operator (ep0 grant) · CC (verify after)** | | **R-535** | **[P2-MEDIUM] The box's console still says „a doboz készen áll, és a párosításra vár" long after the box is bound, claimed and running apps — and it promises that the screen refreshes itself.** MEASURED 2026-09-16 on the drill box (fresh install, ISO 1.27.1, controller 0.243.0): bind succeeded 10:01:14Z, claim 10:06:18Z, four apps deploying by 10:22Z — and at 10:23Z the console still showed the pairing banner with the code `37S-NFE` and the line „Ez a képernyő magától frissül — nincs teendő a doboznál" (`audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png`). A volunteer watching the monitor has no way to tell the box is finished; worse, the screen says it updates itself, so waiting longer does not help. **Fix shape:** the first-boot banner unit re-renders on claim/bind state (the controller already knows: it reports `controller_started` and the hub holds `claimed`), showing „A doboz össze van kötve — a vezérlőpult a https://felhom. címen érhető el"; or at minimum stop printing the pairing code once the appliance is claimed. | **READY — rank P2-MEDIUM; owner: CC (ISO payload `felhom-bootstrap.sh` — ships with the next ISO)** | +| **R-536** | **[P2-MEDIUM] The hub is told „Alkalmazás telepítve" the moment a deploy is ACCEPTED, so an install that never finishes is recorded as a completed one.** MEASURED 2026-09-16 on the drill box: the deploy of `mealie` was accepted at 12:31:36 CEST and the hub logged `Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie` in the SAME second; the controller was then killed 5 s in (F9'), and after the agent restarted it the stack read `not_deployed / deployed=false / deploying=false` — i.e. the app was never installed, and nothing corrected the event. Source confirms the ordering: `internal/api/router.go` writes the 202 „Telepítés elindítva" and then calls `NotifyAppDeployed` immediately, while the comment right above it says the deploy „runs asynchronously (compose pull/up + health happen after this returns)". The event is `info`, so nobody is mailed — but the customer timeline and the hub's app history record a completed install that did not happen (the „presence is not success" class). **Also measured, same shape:** an interrupted deploy leaves `/opt/docker/stacks//app.yaml` behind (written at accept) while the stack reads not-deployed — second instance after 2026-09-15's homebox; moved aside on the box. **Fix shape:** emit `app_deployed` from the async path when the stack reaches running/healthy (or emit `app_deploy_started` at accept and `app_deployed` at completion), and remove the accept-time `app.yaml` on a failed deploy. | **READY — rank P2-MEDIUM; owner: CC (controller)** | +| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. | **READY — rank P1-HIGH; owner: CC (controller)** | +| **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. | **READY — rank P1-HIGH; owner: CC (controller)** | | **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** | | **R-526** | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** | | **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** FOUND 2026-09-15: `stacks/metadata.go` parses it; `grep -rn LockedAfterDeploy` finds no reader; `deploy.html` renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in `02-controller-module-map.md`; the flag is a seam never wired. **Fix shape:** remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | **READY — rank P3-LOW; owner: CC** |