drill 0.243.0: Phase 2 faults F10/F11/F12 measured, R-537 and R-538 filed
gates / gates (push) Successful in 20s
gates / gates (push) Successful in 20s
F10 (a child deletes the photo folder) is the finding: on a one-drive box with no off-site tier the household's own files are in NO backup — the whole-guest tiers exclude mp8 by design and the app's file leg lives at tier 2/3. The app page still labels tier 1 „DB + Konfig + Adatok" (R-537), and the restore reports success while leaving Nextcloud listing five photos it cannot open, after wiping the app's own trash which still held every byte (R-538). F11: the claim page locks out after the SECOND wrong code (15 minutes), the alarm fires and is true; Nextcloud does not lock out. F12: two reboots 60 s apart, all seven stacks back in 124 s, and the supervisor did not count the boots. Evidence: audits/evidence-drill-0243-2026-09-16/phase2-f10.txt, -f11.txt, -f12.txt. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -65,3 +65,26 @@ Vouched as a three-field change — golden 0.243.0 + agent 0.131.0 + min agent 0
|
||||
|
||||
**New finding on the walk: R-535 (P2)** — 25 minutes after a successful bind AND claim, with apps deploying, the box's console still showed „a doboz készen áll, és a párosításra vár" with the pairing code, and still promised „Ez a képernyő magától frissül".
|
||||
|
||||
|
||||
## Phase 2 — the faults
|
||||
|
||||
Five faults, each recorded the same way: what the customer saw · what the box did · time to steady ·
|
||||
did an alarm fire and was it true · did an alarm that should have fired stay silent. Evidence per
|
||||
fault in `evidence-drill-0243-2026-09-16/phase2-*.txt`.
|
||||
|
||||
| fault | customer saw | box did | steady | alarm |
|
||||
|---|---|---|---|---|
|
||||
| **F9'** — controller killed 5 s into a deploy, empty budget | dashboard gone ~37 s; the interrupted app simply not installed | agent saw it on sweep 1, confirmed and restarted on sweep 2 (12:31:45 → 12:32:15/16 CEST) | **37 s** | `controller_restarted_by_agent` — true. **But** `app_deployed` for Mealie had already been sent at accept time → **R-536** |
|
||||
| **F10** — a child deletes the photo folder | folder and 5 photos gone; after the restore the folder is **back and lists all five**, and **none opens** (`Sabre\DAV\Exception\NotFound`) | tier-1 restore replayed 3 volumes + the database in **35 s** and reported plain success | 35 s | none fired, and **none exists** for "restored database points at files that are not there" → **R-537, R-538** |
|
||||
| **F11** — forgotten passwords ×5 | Nextcloud: 401 five times, no lock-out, correct password still works. The box's setup code: wrong twice, then **"Túl sok próbálkozás — próbáld újra 15 perc múlva"** from the third try; "Új beállító kód kérése" still works | claim counter locks the page for 15 min | immediate | `claim_lockout` (warning) at 13:11:42 CEST, operator mail **1 s later** — fired and **true** |
|
||||
| **F12** — two reboots inside two minutes | dashboard gone ~2 min, then everything back | boot reconciler restarted all seven stacks; no double start; nothing stuck "telepítés folyamatban" | **124 s** after the second reset | none fired; correct — a reboot inside the liveness window is not an alarm |
|
||||
| **M1** — memory pressure (R-528) | app restarted 9–10 times | all three OOM signals **silent**: `OOMKilled=false`, `docker events oom` empty, the cgroup invisible inside the guest, `dmesg` unreadable | — | no OOM alarm is possible on this box, same as on scratch 9202 |
|
||||
|
||||
**F10 is the finding of this drill.** On a one-drive box with no off-site tier — the state every fresh
|
||||
install starts in — the household's own files are in **no backup at all**: the whole-guest tiers exclude
|
||||
the data drive by design (`07-backup-architecture.md`, "[FACT] What the whole-guest tiers do NOT carry",
|
||||
confirmed live: `excluding bind mount point mp8 … (not a volume)`), and the app's file leg lives at
|
||||
tier 2 / tier 3, both unset. The design is not the defect. The defects are that the page calls tier 1
|
||||
„DB + Konfig + Adatok" and prints the drive size beside it (**R-537**), and that a restore reports
|
||||
success while leaving the app listing files it cannot open — and wipes the app's own trash, which still
|
||||
held every byte (**R-538**).
|
||||
|
||||
@@ -12,3 +12,13 @@
|
||||
2026/09/16 12:21:58 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Paperless-ngx
|
||||
2026/09/16 12:22:18 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Nextcloud
|
||||
page after: Rendszermentés (teljes mentés) A teljes szerver — alkalmazások, beállítások és adatbázisok együtt — időszakos mentése, amelyből az egész készülék visszaállítható. Ezt a host-ügynök készíti és kezeli. ✓ Utolsó teljes mentés 2026-09-16 12:08 (20 perce) 624.3 MB Helyi tároló (local) Naprakész Következő mentés 0 órája — a mentési ablakon belül – Visszaállítás ellenőrizve Még nem futott Helyi tároló (local) ✓ Utolsó sikeres mentés: 2026-09-16 12:08 (20 perce) · 624.3 MB Naprakész Biztonsági szerver – külön hardver (PBS) nincs beállítva Mentés folyamatban… (fázis: snapshotted ) Mentés most A mentés
|
||||
## 2026-09-16T10:31:39Z manual backup finished: guest lock cleared; archive vzdump-lxc-9201-2026_09_16-12_27_11.tar.zst = 2 175 571 124 B (the scheduled one at 12:08 was 654 665 901 B, before the four apps)
|
||||
## hub events/mails for tester-1 in the last 25 min:
|
||||
2026/09/16 12:08:25 [INFO] Event from tester-1: backup_tier_skipped (warning) — Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it.
|
||||
2026/09/16 12:08:25 [INFO] Operator email sent for tester-1/backup_tier_skipped
|
||||
2026/09/16 12:21:17 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: PrivateBin
|
||||
2026/09/16 12:21:38 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Vaultwarden
|
||||
2026/09/16 12:21:58 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Paperless-ngx
|
||||
2026/09/16 12:22:18 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Nextcloud
|
||||
2026/09/16 12:31:36 [INFO] Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie
|
||||
## 2026-09-16T10:50:29Z manual backup finished (guest lock cleared)
|
||||
|
||||
@@ -0,0 +1,85 @@
|
||||
## F10 - a child deletes the photo folder (Nextcloud) - 2026-09-16T11:01:38Z
|
||||
seed: folder 'Fotok' created through Nextcloud WebDAV (the app's own file interface), MKCOL 201,
|
||||
5 files nyaralas-1..5.jpg, 200k/400k/600k/800k/1000k = 3 000 000 B total, every PUT 201.
|
||||
on-disk: 68204151 /mnt/felhom-drives/adatlemez/appdata/nextcloud
|
||||
on-disk: 5
|
||||
on-disk: 25337 /mnt/felhom-drives/adatlemez/backups
|
||||
2026-09-16T11:01:57Z pressing 'Mentes most' on the app-backup page (POST /api/backup/run)
|
||||
trigger: {"ok":true,"message":"Mentés elindítva"}
|
||||
http=200
|
||||
2026-09-16T11:02:58Z backup finished; status: {"ok":true,"data":{"db_dump":{"count":2,"duration":"47.936765144s","last_run":"2026-09-16T11:02:45.477832211Z","success":true},"enabled":true,"running":false}}
|
||||
restore page after the app backup:
|
||||
TEXT: Pillanatkép: — Válasszon alkalmazást — Még nincs mentés felhasználói adattal. A visszaállítás felülírja az alkalmazás jelenlegi adatait a kiválasztott mentés állapotával. Az alkalmazás a folyamat során automatikusan leáll és újraindul. Megértettem, visszaállítás indítása. Visszaállítás indítása Importálás mentett csomagból (.fab) Hordozható mentéscsomag (.fab) Hordozható pillanatfelvétel — bárhol tárolhatod, és bármikor visszatöltheted egy meghajtóról. A folyamatos védelmet az 1–3. szintű mentés adja. Jelszavas titkosítás (opcionális) Üresen hagyva a csomag titkosítás nélkül készül. Nextcloud Letöltés (.fab) Paperless-ngx Letöltés (.fab) PrivateBin Letöltés (.fab) Vaultwarden Letöltés (.fab)
|
||||
2026-09-16T11:05:06Z the child deletes the folder: DELETE /Fotok http=204
|
||||
folder after the delete: PROPFIND http=404 (404 = gone)
|
||||
## MEASURED before the restore attempt:
|
||||
## 1) The app backup ("1. mentes", tier 1) was made at 11:02:45Z by the customer-visible "Mentes most"
|
||||
## button (POST /api/backup/run -> {"ok":true,"message":"Mentes elindItva"}; the backup folder grew
|
||||
## 25 337 B -> 978 MB, so the button does more than its status field, which reports db_dump only).
|
||||
## 2) Its manifest lists db-dumps + THREE docker volume dumps (html, db_data, redis). It does NOT list
|
||||
## the drive-side app data (/mnt/felhom-drives/adatlemez/appdata/nextcloud), where the photos live.
|
||||
## 3) Proof by listing the html tar (29 346 entries):
|
||||
## positive control "version.php" = 3 hits (the listing works)
|
||||
## "data/" entries = 1284, but "./data/" is the EMPTY bind-mount point
|
||||
## "Fotok" = 0 hits
|
||||
## "nyaralas"= 0 hits
|
||||
## And: find over the whole backups tree for *appdata*/*hdd*/*Fotok* = nothing.
|
||||
## 4) The whole-guest backup does not cover them either - vzdump log of the 12:27 CEST run:
|
||||
## "including mount point rootfs ('/')", "including mount point mp0 ('/var/lib/felhom')",
|
||||
## "excluding bind mount point mp8 ('/mnt/felhom-drives') from backup (not a volume)".
|
||||
## 5) The app-backup page nevertheless labels tier 1 "DB + Konfig + Adatok" and prints "Nextcloud
|
||||
## Adatlemez 65.1 MB" - a size measured on exactly the data it does not copy.
|
||||
## 6) Tier 2 ("off-drive masolat") and tier 3 (offsite) both read "Nincs beallitva" on this box.
|
||||
2026-09-16T11:06:56Z pressing the restore button: POST /backup/restore stack_name=nextcloud snapshot_id=helyi
|
||||
restore POST http=302
|
||||
2026-09-16T11:07:42Z restore status: {"ok":true,"data":{"running":false,"op":"restore","stack":"nextcloud","started_at":"2026-09-16T11:06:56.710474757Z","last":{"op":"restore","stack":"nextcloud","ok":true,"message":"A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult.","finished_at":"2026-09-16T11:07:31.194473339Z"},"last_recent":true}}
|
||||
after the restore, through Nextcloud's own interface:
|
||||
PROPFIND /Fotok http=207
|
||||
file:
|
||||
file: nyaralas-1.jpg
|
||||
file: nyaralas-2.jpg
|
||||
file: nyaralas-3.jpg
|
||||
file: nyaralas-4.jpg
|
||||
file: nyaralas-5.jpg
|
||||
GET /Fotok/nyaralas-1.jpg http=503 bytes=276
|
||||
trash PROPFIND http=207
|
||||
trash:
|
||||
RE-MEASURE at 2026-09-16T11:08:38Z, app healthy (nextcloud Up >2 min):
|
||||
app status.php http=200 <- positive control, app is serving
|
||||
PUT control kontroll.txt http=201 <- positive control, WebDAV write works
|
||||
GET control kontroll.txt http=200 bytes=6 <- positive control, WebDAV read works
|
||||
GET /Fotok/nyaralas-1.jpg http=404 bytes=249
|
||||
GET /Fotok/nyaralas-2.jpg http=503 bytes=276
|
||||
GET /Fotok/nyaralas-3.jpg http=503 bytes=276
|
||||
GET /Fotok/nyaralas-4.jpg http=503 bytes=276
|
||||
GET /Fotok/nyaralas-5.jpg http=503 bytes=276
|
||||
body of the failed download (first 200 chars):
|
||||
<?xml version="1.0" encoding="utf-8"?><d:error xmlns:d="DAV:" xmlns:s="http://sabredav.org/ns"> <s:exception>Sabre\DAV\Exception\NotFound</s:exception> <s:message>File with name /Fotok/nyaralas-1
|
||||
## RESULT - F10 (a child deletes the photo folder, Nextcloud, fresh box, one drive, no tier 2/3)
|
||||
## customer saw: the folder and the 5 photos vanish from Nextcloud (DELETE 204, PROPFIND 404).
|
||||
## After the restore the folder is BACK and lists all 5 photos - but NONE of them opens:
|
||||
## GET nyaralas-1..5 = 404 / 503 x4, body "Sabre\DAV\Exception\NotFound".
|
||||
## Positive controls at the same moment: status.php 200, WebDAV PUT kontroll.txt 201, GET 200 (6 B).
|
||||
## So the failure is real, not a starting app and not a broken login.
|
||||
## box did: POST /backup/restore (stack_name=nextcloud, snapshot_id=helyi) -> 302, finished in 35 s,
|
||||
## "A(z) nextcloud: 3 adatkotet es az adatbazis visszaallitva - az alkalmazas ujraindult."
|
||||
## It replayed the 3 docker volumes + the MariaDB dump. The DB dump (11:01:45Z) KNOWS the 5 photos,
|
||||
## because they were uploaded at ~11:00Z. The BYTES were never in any backup (see MEASURED above).
|
||||
## time to steady: 35 s restore, app healthy ~50 s later.
|
||||
## how much was lost: all 5 files, 3 000 000 B - every photo the folder ever held.
|
||||
## does the page say so? NO. The result line says the volumes and the database were restored and the
|
||||
## app restarted. Nothing says the app's files on the data drive are not part of this backup, and
|
||||
## nothing says the restored database now points at files that do not exist.
|
||||
## where the bytes actually are: Nextcloud's OWN trash on the data drive -
|
||||
## appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg (all 5 present).
|
||||
## But the restored database no longer lists them: the trash PROPFIND returns an EMPTY list, so the
|
||||
## customer cannot press "restore from trash" either. The restore made the trash unreachable.
|
||||
## alarm fired and true? none fired.
|
||||
## alarm that should have and did not: none is defined for "restored DB references missing files".
|
||||
## DESIGN CHECK (a design is not a defect, R-370): 07-backup-architecture.md is explicit -
|
||||
## "[FACT] What the whole-guest tiers do NOT carry ... mp8 /mnt/felhom-drives ... out of vzdump scope",
|
||||
## and the file leg for nextcloud exists at TIER 2 and TIER 3 only (section 6.2 table: Tier-3
|
||||
## mandatory = calibre-web, immich, nextcloud, paperless-ngx). Tier 1 carrying no drive-side files is
|
||||
## therefore BY DESIGN. What is NOT by design is the page calling tier 1 "DB + Konfig + Adatok" and
|
||||
## printing "Adatlemez 65.1 MB" next to it, and a restore reporting plain success while leaving the
|
||||
## app inconsistent. Those two are the rows.
|
||||
@@ -0,0 +1,61 @@
|
||||
## F11 - forgotten passwords, part 1: Nextcloud, 5 wrong tries 2026-09-16T11:09:35Z
|
||||
try 1: http=401 (took 868 ms)
|
||||
try 2: http=401 (took 726 ms)
|
||||
try 3: http=401 (took 702 ms)
|
||||
try 4: http=401 (took 725 ms)
|
||||
try 5: http=401 (took 687 ms)
|
||||
now the CORRECT password again (positive control - is the customer locked out?):
|
||||
http=200 (took 610 ms) 207 = still allowed in
|
||||
web login page http=200
|
||||
## F11 part 2 - wrong CLAIM code x5, measured on this box (already claimed at 10:06:18Z)
|
||||
the claim page itself: http=200
|
||||
page says: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új beállító kód kérése Felhom — Otthoni szerver kezelés felhom.eu
|
||||
wrong claim try 1: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m
|
||||
wrong claim try 2: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m
|
||||
wrong claim try 3: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m
|
||||
wrong claim try 4: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m
|
||||
wrong claim try 5: http=200 page: Jelszó visszaállítása — Felhom Jelszó visszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (m
|
||||
## F11 part 2 - CORRECTED run (my first run was refused by CSRF - I measured my own mistake, not the product)
|
||||
wrong code try 1: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új
|
||||
wrong code try 2: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új
|
||||
wrong code try 3: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új
|
||||
wrong code try 4: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új
|
||||
wrong code try 5: http=200 page: isszaállítása Tester 1 Érvénytelen űrlap — töltsd újra az oldalt. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jelszó megerősítése Jelszó beállítása Új
|
||||
NOTE: this box was already claimed at 10:06:18Z. The /claim URL still answers 200 and now renders
|
||||
'Jelszo visszaallitasa' (password reset) with a 'Uj beallito kod kerese' button - so the claim
|
||||
page is NOT one-shot-dead after a successful claim; it becomes the reset surface.
|
||||
## F11 part 2 - FINAL run (correct CSRF: cookie felhom_claim_csrf + _csrf field + X-CSRF-Token header)
|
||||
wrong code try 1: http=200 (198 ms) page says: Hibás vagy lejárt kód Add meg az e-mailben kapott beállító kódot, majd válassz új
|
||||
wrong code try 2: http=200 (187 ms) page says: Hibás vagy lejárt kód Add meg az e-mailben kapott beállító kódot, majd válassz új
|
||||
wrong code try 3: http=200 (177 ms) page says: Túl sok próbálkozás — próbáld újra 15 perc múlva. Add meg az e-mailben kapott beá
|
||||
wrong code try 4: http=200 (209 ms) page says: Túl sok próbálkozás — próbáld újra 15 perc múlva. Add meg az e-mailben kapott beá
|
||||
wrong code try 5: http=200 (172 ms) page says: Túl sok próbálkozás — próbáld újra 15 perc múlva. Add meg az e-mailben kapott beá
|
||||
CORRECTION of my own summary line: there IS a lock-out, and it arrives on the THIRD try.
|
||||
Tries 1-2: "Hibas vagy lejart kod". Tries 3-5: "Tul sok probalkozas - probald ujra 15 perc mulva."
|
||||
No captcha and no growing delay (172-209 ms throughout) - it is a counter, not a slow-down.
|
||||
the household's escape route while locked out - pressing 'Uj beallito kod kerese':
|
||||
POST /claim/request-new-code http=200
|
||||
page says: Ha az e-mail cím regisztrálva van, elküldtük a kódot. Add meg az e-mailben kapott beállító kódot, majd válassz új jelszót. Beállító kód Új jelszó (min. 12 karakter) Új jel
|
||||
## RESULT - F11 (forgotten passwords x5, both kinds)
|
||||
## A) Nextcloud app password, 5 wrong tries:
|
||||
## customer saw: 401 five times, each ~700 ms, no lock-out, no captcha; the CORRECT password then
|
||||
## worked immediately (200). Positive controls in the same run: status.php 200, WebDAV PUT 201/GET 200.
|
||||
## box did: nothing - no event, no mail, no restart.
|
||||
## time to steady: none needed.
|
||||
## alarm fired and true? none. alarm that should have and did not: none is expected for an app login.
|
||||
## B) The box's own setup code (the claim page), 5 wrong tries:
|
||||
## customer saw: tries 1-2 "Hibas vagy lejart kod"; tries 3-5 "Tul sok probalkozas - probald ujra
|
||||
## 15 perc mulva." So the lock-out starts on the THIRD attempt and lasts 15 minutes. Timing was
|
||||
## flat (172-209 ms), so it is a counter, not a slow-down.
|
||||
## the escape route WORKS while locked out: POST /claim/request-new-code -> 200 and
|
||||
## "Ha az e-mail cim regisztralva van, elkuldtuk a kodot." (deliberately neutral wording).
|
||||
## one-shot behaviour, MEASURED not assumed: this box was claimed at 10:06:18Z, and /claim still
|
||||
## answers 200 afterwards - it becomes the "Jelszo visszaallitasa" (password reset) surface with a
|
||||
## "Uj beallito kod kerese" button. The URL is NOT dead after a successful claim.
|
||||
## NOT WALKED: the same five wrong codes on a SECOND, never-claimed box. That needs a second fresh
|
||||
## install (~35 min) and the drill box was needed for F9''. So the "first-ever claim" variant of
|
||||
## the counter is untested; what is proven here is the reset path on a claimed box.
|
||||
## MY OWN MISTAKE, recorded: my first two runs of B were refused with "Ervenytelen urlap" because I
|
||||
## paired the claim page's pre-auth token wrongly (the page sets cookie felhom_claim_csrf and carries
|
||||
## TWO _csrf fields). I measured my own error, said so, and re-ran with cookie + _csrf field +
|
||||
## X-CSRF-Token header, which the product accepted.
|
||||
@@ -0,0 +1,70 @@
|
||||
## F12 — two reboots inside two minutes (box: nested VM 334 on demo-hp)
|
||||
2026-09-16T10:50:10Z before: dashboard=200 apps=200
|
||||
reset 1 at 2026-09-16T10:50:11Z
|
||||
reset 2 at 2026-09-16T10:51:12Z (60 s after the first)
|
||||
dashboard 200 again at 2026-09-16T10:53:16Z — 124 s after the second reset
|
||||
apps: privatebin=200 vault=200 paperless=302 cloud=200
|
||||
Sep 16 12:53:06 tester1 felhom-agent[1128]: time=2026-09-16T12:53:06.326+02:00 level=WARN msg="controller-supervisor: guest list unavailable — skipping sweep (ownership unproven)" err="pro
|
||||
Sep 16 12:53:07 tester1 felhom-agent[1128]: time=2026-09-16T12:53:07.077+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1
|
||||
Sep 16 12:53:36 tester1 felhom-agent[1128]: time=2026-09-16T12:53:36.356+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1
|
||||
Sep 16 12:53:36 tester1 felhom-agent[1128]: time=2026-09-16T12:53:36.385+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1
|
||||
Sep 16 12:53:36 tester1 felhom-agent[1128]: time=2026-09-16T12:53:36.388+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1
|
||||
Sep 16 12:54:06 tester1 felhom-agent[1128]: time=2026-09-16T12:54:06.388+02:00 level=INFO msg="stale-lock: scanning pool guests" pool=felhom listed=1 scanned=1
|
||||
paperless-webserver Up 34 seconds (healthy)
|
||||
paperless-redis Up 44 seconds (healthy)
|
||||
paperless-postgres Up 44 seconds (healthy)
|
||||
nextcloud Up 56 seconds (healthy)
|
||||
nextcloud-db Up About a minute (healthy)
|
||||
nextcloud-redis Up About a minute (healthy)
|
||||
felhom-controller Up About a minute (healthy)
|
||||
vaultwarden Up About a minute (healthy)
|
||||
privatebin Up About a minute (healthy)
|
||||
filebrowser Up About a minute (healthy)
|
||||
cloudflared Up About a minute
|
||||
traefik Up About a minute
|
||||
## app states after the two resets (2026-09-16T10:57:17Z):
|
||||
|
||||
## supervisor block in the host report:
|
||||
Traceback (most recent call last):
|
||||
File "<string>", line 1, in <module>
|
||||
import json;print(json.load(open("/var/lib/felhom-agent/bootstrap.json"))["local_api_token"])
|
||||
~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
FileNotFoundError: [Errno 2] No such file or directory: '/var/lib/felhom-agent/bootstrap.json'
|
||||
Traceback (most recent call last):
|
||||
File "<string>", line 1, in <module>
|
||||
import sys,json;r=json.load(sys.stdin);print(json.dumps(r.get("controller_supervisor"),indent=1))
|
||||
~~~~~~~~~^^^^^^^^^^^
|
||||
File "/usr/lib/python3.13/json/__init__.py", line 293, in load
|
||||
return loads(fp.read(),
|
||||
cls=cls, object_hook=object_hook,
|
||||
parse_float=parse_float, parse_int=parse_int,
|
||||
parse_constant=parse_constant, object_pairs_hook=object_pairs_hook, **kw)
|
||||
File "/usr/lib/python3.13/json/__init__.py", line 346, in loads
|
||||
return _default_decoder.decode(s)
|
||||
~~~~~~~~~~~~~~~~~~~~~~~^^^
|
||||
File "/usr/lib/python3.13/json/decoder.py", line 345, in decode
|
||||
obj, end = self.raw_decode(s, idx=_w(s, 0).end())
|
||||
~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^
|
||||
File "/usr/lib/python3.13/json/decoder.py", line 363, in raw_decode
|
||||
raise JSONDecodeError("Expecting value", s, err.value) from None
|
||||
json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
|
||||
## uptime + boot count:
|
||||
up 5 minutes
|
||||
reboot system boot 7.0.2-6-pve Wed Sep 16 12:51 - still running
|
||||
reboot system boot 7.0.2-6-pve Wed Sep 16 12:50 - crash
|
||||
reboot system boot 7.0.2-6-pve Wed Sep 16 11:58 - crash
|
||||
|
||||
## RESULT — F12 (two `qm reset`s 60 s apart on VM 334)
|
||||
## customer saw: dashboard unreachable ~2 min; back at 10:53:16Z, 124 s after the SECOND reset.
|
||||
## box did: boot reconciler brought all seven stacks up by itself (privatebin 200, vaultwarden 200,
|
||||
## paperless 302 = its login redirect, nextcloud status.php 200); NO double start observed;
|
||||
## mealie stayed not_deployed (the F9'-interrupted deploy), NOT stuck in "telepites folyamatban".
|
||||
## time to steady: 124 s to the dashboard, ~3 min to every app healthy.
|
||||
## supervisor did NOT count the boots: whole-journal "RESTARTED the controller" = 1 (the F9' one at
|
||||
## 12:32, POSITIVE control that the grep string matches when it happens); since 12:49 = 0,
|
||||
## while 4 controller-supervisor lines in the same window prove it was running and sweeping.
|
||||
## During the boot it logged "guest list unavailable - skipping sweep (ownership unproven)" -
|
||||
## the unprovisioned/ownership guard doing its job.
|
||||
## journal is PERSISTENT on this box (/var/log/journal exists), so the pre-reboot lines are real.
|
||||
## alarm fired and true? none fired. alarm that should have and did not: none - a reboot inside the
|
||||
## node-liveness dedupe window produces no host_* mail, which matches the design.
|
||||
@@ -0,0 +1,13 @@
|
||||
## F9'' - three more controller kills, 20 minutes apart, at IDLE (no deploy running)
|
||||
## Question to MEASURE (not to change): does a restart after 20 minutes of healthy uptime
|
||||
## count against the 3-restarts-per-15-minutes budget, or does the window start fresh?
|
||||
## Baseline: the last restart before this run was F9' at 12:32:16 CEST (1 of 3 at that time).
|
||||
## kill 1 at 2026-09-16T11:12:41Z
|
||||
felhom-controller
|
||||
dashboard 200 again after 61 s
|
||||
Sep 16 13:12:37 tester1 felhom-agent[1128]: time=2026-09-16T13:12:37.326+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=40 guests_evaluated=1 controllers_not_running=0
|
||||
Sep 16 13:13:07 tester1 felhom-agent[1128]: time=2026-09-16T13:13:07.323+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status
|
||||
Sep 16 13:13:37 tester1 felhom-agent[1128]: time=2026-09-16T13:13:37.415+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit
|
||||
Sep 16 13:13:38 tester1 felhom-agent[1128]: time=2026-09-16T13:13:38.655+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
|
||||
RESTARTED lines in the whole journal so far: 2
|
||||
waiting 20 minutes before the next kill
|
||||
@@ -0,0 +1,19 @@
|
||||
## F9' — controller killed 5 s into a deploy, on a FRESH crash-loop budget (agent started 12:02, no restarts yet)
|
||||
2026-09-16T10:31:36Z starting the mealie deploy
|
||||
deploy http=202
|
||||
{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
|
||||
killed at 2026-09-16T10:31:41Z — 5 s into the deploy (the "(37901 s …)" text my script printed here was a timestamp-parsing bug of mine; the real interval is the scripted sleep of 5 s between the 202 and the kill)
|
||||
dashboard health 200 again at 2026-09-16T10:32:18Z — 37 s after the kill
|
||||
mealie after: {'state': 'not_deployed', 'deployed': False, 'deploying': False}
|
||||
Sep 16 12:31:45 tester1 felhom-agent[3602]: time=2026-09-16T12:31:45.160+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status
|
||||
Sep 16 12:32:15 tester1 felhom-agent[3602]: time=2026-09-16T12:32:15.231+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exit
|
||||
Sep 16 12:32:16 tester1 felhom-agent[3602]: time=2026-09-16T12:32:16.606+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 conse
|
||||
Sep 16 12:32:16 tester1 felhom-agent[3602]: time=2026-09-16T12:32:16.606+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=60 guests_evaluated=1 controllers_not_running=1
|
||||
## RESULT (R-531's missing timing, now pinned):
|
||||
## kill 5 s into a deploy, on an EMPTY crash-loop budget → dashboard health 200 again 37 s after the kill
|
||||
## (agent saw it at 12:31:45, confirmed and restarted at 12:32:15/16 — two sweeps, as designed).
|
||||
## The interrupted app ended `not_deployed / deployed=false / deploying=false` — NOT stuck in "telepítés folyamatban".
|
||||
## Budget after this single restart: 1 of 3 in the 15-minute window.
|
||||
2026-09-16 10:31:36.748187371 +0000 253 /opt/docker/stacks/mealie/app.yaml
|
||||
deployed deployed_at env locked_fields desired_state
|
||||
moved-aside
|
||||
@@ -0,0 +1,65 @@
|
||||
are supported and installed on your system.
|
||||
## 2026-09-16T10:35:16Z M1 — OOM signals on THIS box (R-528). before:
|
||||
limit=1342177280 oom=false restarts=0 status=running
|
||||
81: memory: 128M
|
||||
134: memory: 128M
|
||||
## 2026-09-16T10:37:30Z after the 128M cap: false restarting 9
|
||||
cgroup: not visible
|
||||
docker oom events since 2026-09-16T10:35:29Z:
|
||||
kernel OOM lines: 0
|
||||
dmesg not readable in guest
|
||||
2
|
||||
## 2026-09-16T10:40:13Z cap restored: limit=1342177280 health=unhealthy restarts=10
|
||||
are supported and installed on your system.
|
||||
## 2026-09-16T10:40:44Z memory lines now:
|
||||
81: memory: 1280M
|
||||
112: memory: 256M
|
||||
134: memory: 1280M
|
||||
81: memory: 1280M
|
||||
112: memory: 256M
|
||||
134: memory: 128M
|
||||
## 2026-09-16T10:47:37Z paperless after fix: ws=unhealthy limit=1342177280 redis=134217728
|
||||
## M1 RESULT (R-528), on THIS box (nested VM, customer LXC 9201, Docker in the guest):
|
||||
## cap 128M → paperless-webserver RESTARTING with RestartCount 9–10, and:
|
||||
## - docker inspect .State.OOMKilled = FALSE
|
||||
## - docker events --filter event=oom (since the cap change) = EMPTY
|
||||
## - the container's cgroup is NOT visible from inside the guest (find under /sys/fs/cgroup → nothing),
|
||||
## so memory.events / oom_kill could not be read there
|
||||
## - dmesg is not readable in the guest
|
||||
## So all three signals are silent here, exactly as on scratch 9202 (2026-09-15). BIGNIGHT's VM 333 did read
|
||||
## oomkilled=true for the worker-killed-inside-a-running-container shape; a MAIN-process kill + restart reports
|
||||
## nothing. The controller v0.243.0 OOM tag therefore cannot fire for this (restart) shape on a real box.
|
||||
## MISTAKE, stated: my cap edit used a blunt sed, so the restore rewrote the REDIS limit too (128M → 1280M) —
|
||||
## the same mistake as 2026-09-15 on 9202. Repaired in the next step (redis back to 128M, webserver 1280M).
|
||||
paperless-redis Up 7 minutes (healthy)
|
||||
paperless-webserver Restarting (1) 57 seconds ago
|
||||
paperless-postgres Up 12 minutes (healthy)
|
||||
---
|
||||
File "/usr/local/lib/python3.12/site-packages/django/db/backends/base/base.py", line 256, in connect
|
||||
self.connection = self.get_new_connection(conn_params)
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
File "/usr/local/lib/python3.12/site-packages/django/utils/asyncio.py", line 26, in inner
|
||||
return func(*args, **kwargs)
|
||||
^^^^^^^^^^^^^^^^^^^^^
|
||||
File "/usr/local/lib/python3.12/site-packages/django/db/backends/postgresql/base.py", line 332, in get_new_connection
|
||||
connection = self.Database.connect(**conn_params)
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
File "/usr/local/lib/python3.12/site-packages/psycopg/connection.py", line 120, in connect
|
||||
raise last_ex.with_traceback(None)
|
||||
django.db.utils.OperationalError: connection failed: connection to server at "172.20.0.3", port 5432 failed: FATAL: password authentication failed for user "paperless"
|
||||
s6-rc: warning: unable to start service init-migrations: command exited 1
|
||||
/run/s6/basedir/scripts/rc.init: warning: s6-rc failed to properly bring all the services up! Check your logs (in /run/uncaught-logs/current if you have in-container logging) for more information.
|
||||
/run/s6/basedir/scripts/rc.init: fatal: stopping the container.
|
||||
## 2026-09-16T10:48:27Z repairing paperless through the CONTROLLER's own path (my manual compose up broke its env):
|
||||
stop: {"ok":true,"message":"Stack paperless-ngx stop completed"}
|
||||
http=200
|
||||
start: {"ok":true,"message":"Stack paperless-ngx start completed"}
|
||||
http=200
|
||||
paperless state after the controller-driven restart: running at 2026-09-16T10:49:35Z
|
||||
## MISTAKE, stated (2026-09-16): the M1 measurement recreated the paperless webserver with a bare
|
||||
## `docker compose up -d` INSIDE the guest. The controller injects the app's environment (DB password etc.) from
|
||||
## `app.yaml` at deploy time, so a hand-run compose brings the container up WITHOUT it: the webserver then failed
|
||||
## `password authentication failed for user "paperless"` against its own postgres and crash-looped.
|
||||
## Nothing about the product; the fault was the manual path. Repaired by restarting the stack through the
|
||||
## CONTROLLER's own stop/start API, which re-renders the environment. Lesson for the next measurement: drive app
|
||||
## containers through the controller, or set the cap via `docker update --memory`, never by re-running compose.
|
||||
@@ -0,0 +1,48 @@
|
||||
## Alarm truth table — operator mails actually DELIVERED during the drill (read from the mailbox, not from the hub log)
|
||||
## (to admin@felhom.eu, via the Gmail connector, 2026-09-16)
|
||||
1. 12:04:01 CEST [Felhom] ✅ tester-1: node_recovered (severity INFO) — "Reports resumed (was down for 0m)".
|
||||
TRUE but odd: the box had never been down; it was newly enrolled. Also: severity INFO was MAILED, while the
|
||||
project's rule says severityNotifies DROPS info before both legs — to be checked against the dispatcher (below).
|
||||
2. 12:08:26 CEST [Felhom] ⚠️ tester-1: backup_tier_skipped (severity warning) — "Whole-guest backup tier felhom-pbs
|
||||
skipped: its storage does not exist on the host (never provisioned or removed)". TRUE: R-534 blocked the tier, so
|
||||
the tier's storage is genuinely absent. This is R-518's cheap half firing on a real box, with its operator mail.
|
||||
No other operator mail arrived in the window (the third thread in the mailbox is a DMARC report, unrelated).
|
||||
## Faults injected so far: none (Phase 2 starts with F9'). Alarms that SHOULD have fired and did not: none so far.
|
||||
## Why the INFO mail is correct (checked in code, not assumed):
|
||||
## `node_recovered` is emitted with severity "info" (monitor/staleness.go emitTransition) and handed to
|
||||
## dispatcher.ProcessEvent, which returns before both legs for "info" — BUT the recovery branch
|
||||
## (`recoveredPairedDownTypes`, dispatcher.go:124) runs BEFORE that gate, deliberately, so the operator hears the
|
||||
## all-clear for a down it was told about. The customer leg stays pairing-gated.
|
||||
## In context the mail is also TRUE: customer `tester-1` had been reporting nothing since the BIGNIGHT teardown, so the
|
||||
## new box's first report IS a recovery, and "(was down for 0m)" is the age of the gap the checker could see.
|
||||
## No row filed.
|
||||
|
||||
## ALARM TRUTH TABLE - second pass, 2026-09-16 (hub 0.115.0), built from the hub's own log lines
|
||||
## (pod hub-85478f77d7-mpxkr, times CEST) paired with the customer timeline on /customers/tester-1.
|
||||
##
|
||||
## event when severity operator mail? true in context?
|
||||
## selfbind_link_sent 11:59:55 info yes (self-bind e-mail) yes - the operator pressed it
|
||||
## claim_reissued_reenroll 12:01:56 info yes (claim code to customer) yes - 3rd generation after reinstall
|
||||
## controller_started 12:03:30 info no correct - info, not paired
|
||||
## node_recovered 12:04:00 info YES - mail sent BY DESIGN: the recovery branch runs
|
||||
## before the severity gate; true here
|
||||
## backup_tier_skipped 12:08:25 warning yes yes - the off-site tier has no storage
|
||||
## backup_tier_skipped 12:37:17 warning SUPPRESSED - cooldown correct - same key within the hour
|
||||
## (key=tester-1:backup_tier_skipped:felhom-pbs)
|
||||
## app_deployed x5 12:21-12:31 info no NOT always true -> R-536 (sent at accept)
|
||||
## app_start_failed (paperless) 12:48:47 warning yes TRUE - and it was MY damage (the
|
||||
## hand-run compose, recorded in phase2-m1)
|
||||
## controller_restarted_by_agent 12:48 info no (info) yes - F9' restart #1, timestamps match
|
||||
## backup_tier_skipped 12:58:13 warning SUPPRESSED - cooldown correct
|
||||
## claim_lockout 13:11:42 warning YES, 1 s later TRUE - F11 part 2, third wrong code
|
||||
## claim reset code (gen 4) 13:12:05 info yes (to the customer) yes - "Uj beallito kod kerese"
|
||||
##
|
||||
## What this pass adds to the first one:
|
||||
## 1) The 5-minute node/host dedupe (R-529) was NOT exercised today: no node_* or host_* alarm fired
|
||||
## after 12:04, because the box stayed up and reporting. The two reboots of F12 were far too short
|
||||
## to reach the 30-minute staleness threshold. So the widened bypass is SHIPPED but UNPROVEN-LIVE
|
||||
## on this box; the proof from 2026-09-15 on the other box still stands.
|
||||
## 2) The one-hour operator cooldown is PROVEN here twice, by its own log line, on backup_tier_skipped.
|
||||
## 3) Every warning that fired today was true. No warning that should have fired stayed silent, EXCEPT
|
||||
## the class R-538 names: there is no alarm at all for "a restore finished and the app is now
|
||||
## inconsistent", so nothing could have fired.
|
||||
@@ -0,0 +1,39 @@
|
||||
## 2026-09-16T11:14:14Z ep0 READ-ONLY listing of tester-1's namespace (nothing removed, nothing written)
|
||||
+================+====================+=========+
|
||||
| name | path | comment |
|
||||
+================+====================+=========+
|
||||
| felhom-offsite | /mnt/pbs-datastore | |
|
||||
+================+====================+=========+
|
||||
--- groups in namespace tester-1:
|
||||
Error: error building client for repository felhom-offsite - no password input mechanism available
|
||||
## 2026-09-16T11:14:24Z ep0 listing by filesystem (read-only; the client needs a password we deliberately do not hold here)
|
||||
--- namespaces:
|
||||
demo-felhom
|
||||
demo-hp
|
||||
tester-1
|
||||
--- tester-1 groups:
|
||||
total 12
|
||||
drwxr-xr-x 3 backup backup 4096 Sep 14 16:04 .
|
||||
drwxr-xr-x 5 backup backup 4096 Sep 14 15:48 ..
|
||||
drwxr-xr-x 2 backup backup 4096 Sep 15 08:33 ct
|
||||
--- ct groups and their snapshots:
|
||||
## 2026-09-16T11:14:42Z ep0 deeper listing (read-only)
|
||||
--- tester-1/ct contents:
|
||||
total 8
|
||||
drwxr-xr-x 2 backup backup 4096 Sep 15 08:33 .
|
||||
drwxr-xr-x 3 backup backup 4096 Sep 14 16:04 ..
|
||||
--- snapshots per group:
|
||||
group: /mnt/pbs-datastore/ns/tester-1/ct/*/
|
||||
--- other namespaces, for scale:
|
||||
9201
|
||||
## CONCLUSION (nothing was removed by this session; ep0 was read only)
|
||||
## Before the drill (09:12:21Z): ns tester-1 existed with an EMPTY ct group dir - the old box's
|
||||
## snapshots had been removed the previous night, with the operator's explicit yes.
|
||||
## After the whole drill (11:14:42Z): ns tester-1/ct is STILL EMPTY.
|
||||
## Why: this box never had an off-site tier at all. The re-issue failed on the endpoint token's missing
|
||||
## Datastore.Modify grant (R-534), so no descriptor, no secret, no upload path. The box therefore
|
||||
## wrote ZERO snapshots off-site during the entire drill.
|
||||
## For scale, the neighbouring namespace demo-hp/ct holds group 9201 - proof that this listing method
|
||||
## DOES show groups when they exist (positive control).
|
||||
## Consequence for Phase 3: "restore one DB-backed app from the off-site tier onto scratch 9202"
|
||||
## CANNOT BE WALKED on this box. Not skipped for time - there is no off-site copy to restore from.
|
||||
@@ -0,0 +1,52 @@
|
||||
## tester-1 event timeline as the hub shows it (customer page), read 2026-09-16T11:13:18Z
|
||||
Sep 16 11:11 | warning | claim_lockout | Túl sok hibás beállító kód — a beállító oldal 15 percre zárolva | controller
|
||||
Sep 16 10:58 | warning | backup_tier_skipped | Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. | controller
|
||||
Sep 16 10:53 | info | controller_started | Controller elindult (0.243.0) | controller
|
||||
Sep 16 10:48 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller
|
||||
Sep 16 10:48 | info | controller_restarted_by_agent | Host tester-1-652049 guest 9201: the agent restarted the controller (controller container exited on 2 consecutive sweeps) — restart #1 since the agent started | hub
|
||||
Sep 16 10:37 | warning | backup_tier_skipped | Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. | controller
|
||||
Sep 16 10:32 | info | controller_started | Controller elindult (0.243.0) | controller
|
||||
Sep 16 10:31 | info | app_deployed | Alkalmazás telepítve: Mealie | controller
|
||||
Sep 16 10:27 | info | recovery_credential_revealed | Az üzemeltető lekérte a géped konzolos hozzáférési jelszavát (távoli hibaelhárítás). | hub
|
||||
Sep 16 10:22 | info | app_deployed | Alkalmazás telepítve: Nextcloud | controller
|
||||
Sep 16 10:21 | info | app_deployed | Alkalmazás telepítve: Paperless-ngx | controller
|
||||
Sep 16 10:21 | info | app_deployed | Alkalmazás telepítve: Vaultwarden | controller
|
||||
Sep 16 10:21 | info | app_deployed | Alkalmazás telepítve: PrivateBin | controller
|
||||
Sep 16 10:08 | warning | backup_tier_skipped | Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. | controller
|
||||
Sep 16 10:04 | info | node_recovered | Reports resumed (was down for 0m) | hub
|
||||
Sep 16 10:03 | info | controller_started | Controller elindult (0.243.0) | controller
|
||||
Sep 16 10:01 | info | claim_reissued_reenroll | Új beállító kódot küldtünk a szerver újratelepítése után (3. generáció) az ügyfél címére. | hub
|
||||
Sep 16 10:01 | info | appliance_credential_delivered | Új eszköz (bare-metal telepítés) megkapta a hozzáférést és megkezdi a beállítást. | hub
|
||||
Sep 16 10:01 | info | appliance_bound | Az ügyfél saját maga kötötte össze az új eszközt (bare-metal telepítés); a hozzáférést a doboz a következő lekérdezéskor megkapja. | customer_selfbind
|
||||
Sep 16 09:59 | info | selfbind_link_sent | Self-bind link e-mailed (operator button) | hub
|
||||
Sep 14 23:25 | error | node_down | No report received for 1h | hub
|
||||
Sep 14 22:55 | warning | node_stale | No report received for 30m | hub
|
||||
Sep 14 22:54 | warning | host_stale | Host tester-1-a61396: no report for 30m | hub
|
||||
Sep 14 22:10 | info | node_recovered | Reports resumed (was stale for 5m) | hub
|
||||
Sep 14 22:10 | info | controller_started | Controller elindult (0.242.0) | controller
|
||||
Sep 14 22:04 | warning | node_stale | No report received for 30m | hub
|
||||
Sep 14 21:34 | info | app_deployed | Alkalmazás telepítve: Homebox | controller
|
||||
Sep 14 21:26 | info | node_recovered | Reports resumed (was stale for 1m) | hub
|
||||
Sep 14 21:25 | warning | node_stale | No report received for 30m | hub
|
||||
Sep 14 21:04 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: warn) | controller
|
||||
Sep 14 20:54 | warning | health_degraded | Rendszer állapot romlott (volt: ok) | controller
|
||||
Sep 14 20:44 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: warn) | controller
|
||||
Sep 14 20:43 | info | storage_reconnected | Meghajtó újra csatlakoztatva: Adatlemez | controller
|
||||
Sep 14 20:39 | warning | health_degraded | Rendszer állapot romlott (volt: ok) | controller
|
||||
Sep 14 20:38 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller
|
||||
Sep 14 20:38 | warning | app_start_failed | Telepített alkalmazás nem fut: Immich | controller
|
||||
Sep 14 20:38 | warning | app_start_failed | Telepített alkalmazás nem fut: Jellyfin | controller
|
||||
Sep 14 20:38 | warning | app_start_failed | Telepített alkalmazás nem fut: Nextcloud | controller
|
||||
Sep 14 20:38 | error | storage_disconnected | Meghajtó váratlanul leválasztva: Adatlemez | controller
|
||||
Sep 14 20:29 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: warn) | controller
|
||||
Sep 14 20:28 | info | storage_reconnected | Meghajtó újra csatlakoztatva: Adatlemez | controller
|
||||
Sep 14 19:58 | warning | app_start_failed | Telepített alkalmazás nem fut: Paperless-ngx | controller
|
||||
Sep 14 19:58 | warning | app_start_failed | Telepített alkalmazás nem fut: Jellyfin | controller
|
||||
Sep 14 19:58 | warning | app_start_failed | Telepített alkalmazás nem fut: Immich | controller
|
||||
Sep 14 19:58 | warning | app_start_failed | Telepített alkalmazás nem fut: Nextcloud | controller
|
||||
Sep 14 19:58 | error | storage_disconnected | Meghajtó váratlanul leválasztva: Adatlemez | controller
|
||||
Sep 14 19:54 | info | controller_started | Controller elindult (0.242.0) | controller
|
||||
Sep 14 19:42 | info | controller_started | Controller elindult (0.242.0) | controller
|
||||
Sep 14 19:42 | error | backup_failed | an app-data backup (volume dump) was interrupted by a controller restart — 1 app(s) were left stopped and have been restarted | controller
|
||||
Sep 14 19:31 | info | controller_started | Controller elindult (0.242.0) | controller
|
||||
## rows printed: 50
|
||||
@@ -0,0 +1,19 @@
|
||||
## 2026-09-16T11:15:11Z Phase 3 - morning-after sweep (in the quiet gap between F9'' kills)
|
||||
public front doors:
|
||||
https://paste.enkicsifelhom.hu/ -> 200
|
||||
https://vault.enkicsifelhom.hu/ -> 200
|
||||
https://paperless.enkicsifelhom.hu/ -> 302
|
||||
https://cloud.enkicsifelhom.hu/status.php -> 200
|
||||
https://fajlok.enkicsifelhom.hu/ -> 404
|
||||
stack states through the dashboard API:
|
||||
version labels:
|
||||
controller (dashboard footer): golden on the hub for this host:
|
||||
(the earlier empty read was my expired session, not a broken box - re-measured below)
|
||||
cloudflared running deployed=False subdomain=
|
||||
filebrowser running deployed=False subdomain=
|
||||
nextcloud running deployed=True subdomain=cloud
|
||||
paperless-ngx running deployed=True subdomain=paperless
|
||||
privatebin running deployed=True subdomain=paste
|
||||
traefik running deployed=False subdomain=
|
||||
vaultwarden running deployed=True subdomain=vault
|
||||
controller version on the page: 0.243.0
|
||||
@@ -720,6 +720,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
| **R-534** | **[P1-HIGH] The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks `Datastore.Modify`, so every re-issue fails.** MEASURED 2026-09-16 on the drill box (`tester-1-652049`, fresh install from published ISO 1.27.1): the WG-registration hook refused as R-511 describes („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action"); the operator pressed exactly that, hub v0.114.0's ADOPT path ran, and the endpoint answered **`Process exited with status 255 (stderr: Error: permission check failed - missing Datastore.Modify on /datastore/felhom-offsite)`** → HTTP 502, no descriptor written, no secret stored (fail-closed, correct). So the code fix of 2026-09-15 is sound and INERT: a rebuilt box has no whole-guest off-site tier, and the customer's „Távoli rendszermentés" stays absent. **Not run by hand on ep0** (fenced). **Fix shape (operator):** grant the hub's tenantsync user `Datastore.Modify` on `/datastore/felhom-offsite` (it already holds the create/delete grants the provision path uses), or give the script a token-only re-key op that needs no Modify. Evidence: `audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt`. | **WAITING-ON-OPERATOR — rank P1-HIGH; owner: operator (ep0 grant) · CC (verify after)** |
|
||||
| **R-535** | **[P2-MEDIUM] The box's console still says „a doboz készen áll, és a párosításra vár" long after the box is bound, claimed and running apps — and it promises that the screen refreshes itself.** MEASURED 2026-09-16 on the drill box (fresh install, ISO 1.27.1, controller 0.243.0): bind succeeded 10:01:14Z, claim 10:06:18Z, four apps deploying by 10:22Z — and at 10:23Z the console still showed the pairing banner with the code `37S-NFE` and the line „Ez a képernyő magától frissül — nincs teendő a doboznál" (`audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png`). A volunteer watching the monitor has no way to tell the box is finished; worse, the screen says it updates itself, so waiting longer does not help. **Fix shape:** the first-boot banner unit re-renders on claim/bind state (the controller already knows: it reports `controller_started` and the hub holds `claimed`), showing „A doboz össze van kötve — a vezérlőpult a https://felhom.<domain> címen érhető el"; or at minimum stop printing the pairing code once the appliance is claimed. | **READY — rank P2-MEDIUM; owner: CC (ISO payload `felhom-bootstrap.sh` — ships with the next ISO)** |
|
||||
| **R-536** | **[P2-MEDIUM] The hub is told „Alkalmazás telepítve" the moment a deploy is ACCEPTED, so an install that never finishes is recorded as a completed one.** MEASURED 2026-09-16 on the drill box: the deploy of `mealie` was accepted at 12:31:36 CEST and the hub logged `Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie` in the SAME second; the controller was then killed 5 s in (F9'), and after the agent restarted it the stack read `not_deployed / deployed=false / deploying=false` — i.e. the app was never installed, and nothing corrected the event. Source confirms the ordering: `internal/api/router.go` writes the 202 „Telepítés elindítva" and then calls `NotifyAppDeployed` immediately, while the comment right above it says the deploy „runs asynchronously (compose pull/up + health happen after this returns)". The event is `info`, so nobody is mailed — but the customer timeline and the hub's app history record a completed install that did not happen (the „presence is not success" class). **Also measured, same shape:** an interrupted deploy leaves `/opt/docker/stacks/<app>/app.yaml` behind (written at accept) while the stack reads not-deployed — second instance after 2026-09-15's homebox; moved aside on the box. **Fix shape:** emit `app_deployed` from the async path when the stack reaches running/healthy (or emit `app_deploy_started` at accept and `app_deployed` at completion), and remove the accept-time `app.yaml` on a failed deploy. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. | **READY — rank P1-HIGH; owner: CC (controller)** |
|
||||
| **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. | **READY — rank P1-HIGH; owner: CC (controller)** |
|
||||
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |
|
||||
| **R-526** | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** |
|
||||
| **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** FOUND 2026-09-15: `stacks/metadata.go` parses it; `grep -rn LockedAfterDeploy` finds no reader; `deploy.html` renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in `02-controller-module-map.md`; the flag is a seam never wired. **Fix shape:** remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
Reference in New Issue
Block a user