BIGNIGHT F7 disk full: box holds, English banner, operator alarm silenced by cooldown; R-516/R-521 amended
gates / gates (push) Successful in 19s
gates / gates (push) Successful in 19s
This commit is contained in:
@@ -719,12 +719,12 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-513** | **[P1-HIGH — SECURITY] Every box's file manager (FileBrowser, `files.<domain>`, a launcher tile) accepts the login `admin` / `admin`, and on demo-hp that login page is on the public internet.** MEASURED 2026-09-14 (BIGNIGHT): `POST /api/auth/login?username=admin` with `X-Password: admin` → **200 + a session token** on VM 333 (fresh ISO 1.27.1 install, controller 0.242.0) and on demo-hp guests **9201 and 9202** (loopback, `Host: files.enkisfelhom.hu`); negative control `admin` / wrong → **401** on all three. On VM 333 that token lists both sources — „Adatlemez" (the data drive's `userdata`: documents, media, photos…) and „Beolvasás" (with `paperless`) — `GET /api/users?id=self` 200. **Public exposure, measured by GET only:** `https://files.enkisfelhom.hu/` from DooPlex through Cloudflare → 200, FileBrowser Quantum, `passwordAvailable:true, noAuth:false` (`phase3/filebrowser-public-reachability-demo-hp.txt`); no login was attempted over the internet. The generated `config.yaml` sets no admin credential (FileBrowser's own default applies); no screen shows the customer any FileBrowser login. Geo-restriction narrows who can reach it, it does not authenticate. **Not changed tonight** (9201's standing state is fenced; no product code). **Fix shape:** the controller sets a generated admin password (or proxy auth behind the dashboard session) at stack creation and on every existing box, and shows it where the customer finds app credentials. | **READY — rank P1-HIGH; owner: CC (controller) · operator (rotate on live boxes first)** |
|
||||
| **R-514** | **[P2-MEDIUM] Paperless-ngx dies silently when a family uploads 20 documents at once: its worker is OOM-killed inside the catalog's 768 MB cap, 11 uploads fail, 8 wait forever, and the app still reads „Fut".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, catalog paperless-ngx 2.20.15, `paperless-webserver` limit 805 306 368 B): 20 small 3-page PDFs posted through `/api/documents/post_document/` at 18:29:57Z. At 18:31:22Z the VM kernel logged `Memory cgroup out of memory: Killed process … (gs)` ×2, `([celeryd: celer)`, `([celery beat] -)` — constraint MEMCG of that container; `docker inspect` → `oomkilled=true restarts=0`. Paperless's own task list: **11 FAILURE (`WorkerLostError`), 1 STARTED, 8 PENDING**, unchanged 13 minutes later; **0 documents**. The controller shows the app running and healthy; no event, no alarm (`phase3/paperless-tasks.txt`, `paperless-oom-check.txt`). A household scanning a drawer of bills sees nothing arrive and no reason. **Fix shape:** raise the cap or set `PAPERLESS_TASK_WORKERS=1` / `PAPERLESS_THREADS_PER_WORKER=1` in the template, and let the dead-app/health check see an OOM-killed worker. | **READY — rank P2-MEDIUM; owner: CC (catalog)** |
|
||||
| **R-515** | **[P3-LOW] The Paperless-ngx app page tells the customer to log in with `admin / admin`, and that login does not exist: the deploy form generates the admin password.** MEASURED 2026-09-14 (BIGNIGHT, VM 333): `/apps/paperless-ngx` „Első lépések — Jelentkezz be: admin / admin" and „Alapértelmezett belépés admin / admin"; the catalog template generates `PAPERLESS_ADMIN_PASSWORD` (`password:16`) and the token call with the generated value succeeded (`phase3/seed-paperless.txt`). A household following the page is refused at its first login. Same card also sends documents to „FileBrowser … import/paperless" — the FileBrowser login is R-513. **Fix:** the card points at „Beállítások → Automatikusan generált értékek" as gokapi's does. | **READY — rank P3-LOW; owner: CC (catalog)** |
|
||||
| **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. | **READY — rank P3-LOW; owner: CC (controller + catalog)** |
|
||||
| **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. | **READY — rank P3-LOW; owner: CC (controller + catalog)** |
|
||||
| **R-517** | **[P1-HIGH] After a failed off-site whole-system backup, „Biztonsági mentés" tells the customer the full backup is current and that a remote copy on separate hardware exists — neither is true.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, controller 0.242.0, customer `tester-1` with the DR tier ticked but never provisioned — R-511): „Mentés most" at 19:03:23Z; the local tier succeeded (8 877 619 753 B, 362 s); the `felhom-pbs` tier then failed — agent: `could not activate storage 'felhom-pbs': storage 'felhom-pbs' does not exist`; `pvesm status` lists only `local` and `local-lvm`. At 19:15:36Z `/backups` read: „✗ · **Utolsó teljes mentés 2026-09-14 21:09 (5 perce) · 0 B · Biztonsági szerver – külön hardver (PBS) · Naprakész**" and „✓ **Távoli rendszermentés — külön hardveren (PBS)**" (`phase4/backup-pages-after.txt`). The successful 8.9 GB local backup is no longer shown; a 0-byte failed attempt is labelled up to date; the remote tier is ticked as present. The CLAUDE.md rule „presence is not success" in page form: an attempt's timestamp stands in for a result. The hub did raise a true `whole_guest_backup_failed (error)` to the operator; the customer's page says the opposite. **Fix shape:** the tile shows the newest SUCCESSFUL backup per tier, a failed tier as failed, and a tier whose storage does not exist as „nincs beállítva". **Also measured after F2's reboot (19:45:41Z):** the whole-system tile read „– · Utolsó teljes mentés – · Méret / cél · **Naprakész**" — no backup listed at all, still labelled up to date (the 21:03 CEST local success no longer shown). | **READY — rank P1-HIGH; owner: CC (controller)** |
|
||||
| **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
| **R-519** | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. | **READY — rank P3-LOW; owner: CC (drill)** |
|
||||
| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |
|
||||
| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user