New apps 2026-10-10: Grocy and LubeLogger on the website; Monica stopped
gates / gates (push) Failing after 11m41s
gates / gates (push) Failing after 11m41s
The catalogue side is app-catalog-felhom.eu b7f0f7c. Here: the evidence, the website, the register and the operator's view. Website - Two cards in the Otthon & Eletmod / Home & Lifestyle section of BOTH apps pages, with assets: grocy-logo.svg is grocy's own icon with its single fill made white like the other logos, lubelogger-logo.png is the app's own icon with its dark background dropped and the mark made white (the rule SparkyFitness's PNG follows), and six screenshots of each app's own UI with a household's own data, taken headless on the bench from the published template. - The app count moved 56 -> 58 in 15 places per language set, both languages, and the open-source tile 49 -> 51. Checked by asking the same patterns for the new number afterwards. marketing/facebook/COPY.md still says 56 and is NOT changed: the post it carries is already scheduled. Evidence - documentation/audits/new-apps-2026-10-10/ — FIT.md (checklist group 0 for all three, with the Hungarian-UI column), the bench and box transcripts, the memory samples, the screenshots and the gate runs. Register: 137 -> 139 rows, 2 opened, 0 closed. - R-926 after a remove-keeping-backups and restore, nobody has checked what the app page shows for an after_install app's generated password. Measured here: LubeLogger is safe (its login is derived from the environment at every start, the new password signs in, the data is back); grocy's install password still signs in from the restored database but the deployed environment no longer carries ADMIN_PASSWORD at all. Seven apps are in the class. - R-927 Monica, stopped at checklist 0.2 with the measurements, waiting on the operator. Gate scripts: the same console trap in nineteen of them and in repo_gates.py itself, where it ABORTED THE WHOLE RUNNER at the first gate — printing a non-ASCII character on this workstation's cp1250 console raised UnicodeEncodeError before the gate had decided anything, and reuse_refs_check.py died while printing a NOTE. All now reconfigure their own streams. Red-proof that the remaining script-test failures are not mine: test_due_checks_gate.py fails the same 5 of 42 with the change reverted.
This commit is contained in:
@@ -122,10 +122,11 @@ stopping line that lies.
|
||||
| **R-494** | Install & onboarding | P4 | **NARROWED 2026-09-14 by operator ruling → [P3-LOW] the hub COULD create the tunnel at customer creation, for a domain already on Cloudflare. Not blocking: every customer has their own domain and the operator creates the tunnel per day-0 A.1 (`architecture/01-topology-and-trust.md`).** *Original finding, kept:* **[P1-HIGH] A new customer's dashboard has NO reachable address unless the operator hand-makes a Cloudflare tunnel — the link in the setup-code mail is dead.** MEASURED 2026-09-14 on a fresh install from the public ISO (drill intervention **I1**): the claim mail points at `https://felhom.drill0242.felhom.eu`; that name has **no A and no AAAA** record (`dig @1.1.1.1`, control `felhom.enkisfelhom.hu` resolves); the hub has **no tunnel- or DNS-creation code** (`hub/internal/cloudflare/` holds only geo-rule removal; `cf_tunnel_token` is a pasted, optional form field, `configs.go:1478`) — day-0 runbook A.1 makes it a manual Cloudflare-dashboard step that nothing on the customer-create page asks for; the box's own split-horizon resolver on the appliance LAN IP answered `google.com` but not the dashboard name at 13:27:39Z; the agent applied the record at **13:27:44Z** (`lanresolver: applied split-horizon record … ip=192.168.0.158`, 3 m 46 s after the controller started), so the box CAN answer the name — **but only to a device that uses the box as its DNS server, and no document, screen or mail tells a household to do that**; the router and the installer-offered DNS answer nothing. The page was reachable only at the guest's LAN address with the name forced (`curl --resolve …:443:192.168.0.158`). **A volunteer could not have done that.** **What it needs:** an operator ruling — the hub creates the tunnel and DNS at customer creation, or the product gives a household a LAN address that works with no DNS change. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: operator-side automation; the operator creates the tunnel by hand per the day-0 runbook.** | — | — | CC |
|
||||
| **R-504** | Install & onboarding | P4 | **[P3-LOW] `iso.felhom.eu` cannot show an index page on its own — its root returns 404, and the download page lives on the website instead.** MEASURED 2026-09-14: `https://iso.felhom.eu/` and `/index.html` → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded `index.html` at `/` was **not measured** (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at `felhom.eu/letoltes` (published with the ISO, after the operator's yes). **Remaining:** a redirect from `iso.felhom.eu/` to that page needs a Cloudflare rule the session has no credential for. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule)** **Re-ranked 2026-10-03: P3→P4: households are sent to the website's download page; the bare address is cosmetic.** | — | — | operator |
|
||||
|
||||
## Apps & catalog — 11 rows (P3 3, P4 8)
|
||||
## Apps & catalog — 12 rows (P3 3, P4 9)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **R-927** | Apps & catalog | P4 | **Monica was checked against the new-app checklist and STOPPED at 0.2 — the operator decides whether to keep it out.** Measured 2026-10-10 (`audits/new-apps-2026-10-10/FIT.md`): no release of any kind since 2025-04-21 and that one a prerelease; no STABLE release since v4.1.2 on 2024-05-04, 29 months; the `4.x` branch the README calls "the stable and current version" has not moved since 2024-05-04; `main` says it is the beta; the commit history has a 13-month hole, and the two commits that end it add a login-page notice reading that the hosted instance and all its data will be deleted at the end of December 2026 "ahead of the new Monica version". 790 issues open. Nothing in the catalog covers a family address book, so keeping it out costs nothing we have; letting it in means publishing an app whose next version is a rewrite with no migration story. Radicale already carries contacts over CardDAV. | **OPEN — needs the operator's word** | the operator | Keep it out and look again when the rewrite ships, or say to build it anyway | operator |
|
||||
| **R-562** | Apps & catalog | P3 | **[P3-LOW] Dates and sizes are not formatted for any locale — and the Hungarian pages disagree with themselves.** FOUND 2026-09-17 by the i18n inventory §2.8: the two template date layouts differ (`2006. 01. 02. 15:04` Hungarian vs `2006-01-02 15:04` ISO); 10 layout literals in `internal/web` Go and 25 elsewhere pick formats ad hoc; sizes print a decimal POINT (`%.1f GB`, 4 helpers) where Hungarian uses a comma; `timeAgo`/`nextRunLabel`/`pruneLabel` produce Hungarian words outside the three converted pages. Not changed by v0.247.0 (Hungarian bytes are frozen by the parity rule). **Fix shape:** one date and one size formatter per language in `internal/i18n`, the Hungarian output deliberately changed in ONE reviewed release with the parity fixtures re-captured for that release only and the change named in its CHANGELOG. Needs an operator word on the Hungarian format (comma, date style). | **READY - rank P3-LOW; owner: CC** **2026-10-06 night: not started — the row needs the operator's word on the Hungarian format (decimal comma, date style) before any code.** | — | — | CC |
|
||||
| **R-676** | Apps & catalog | P3 | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. **-- 2026-09-30: the first-start restarts are explained.** immich's first-start geodata import OOM-kills its database at 512M on a guest with no swap (R-732, measured: 61–104 kills); the 2026-09-17 chaos-night case (DB connection dropped during the import on a 6 GB guest) fits it. Fixed in the catalog (`56c4888`, 768M). The watch itself (decision 28 on a DEPLOY's first start) is unchanged. | **OPEN — P3; owner: CC (watch)** | — | — | CC |
|
||||
| **R-76** | Apps & catalog | P4 | **FileBrowser-created folders break the setgid chain, and a drop-zone's mode is not stable** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size S, roadmap state `idea (surfaced by the R-75 spike, 2026-07-26)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-06 night: not a catalog fix** — the setgid chain is set by the controller's FileBrowser setup; a controller row. | — | Two related findings from `audits/SPIKE-catalog-data-paths-2026-07-26.md` P3/P5, both **pre-existing** and deliberately left alone by that spike. **(a)** FileBrowser Quantum 1.3.3 creates files `0644` and folders `0755` and does **not** propagate the setgid bit — even though the entrypoint wrapper's `umask 002` really is in effect (`/proc/1/status` `Umask: 0002`). Group inheritance itself works (a file uploaded into a 2775 group-100 dir landed group 100, not the process gid 1000), so the convention's *group* half holds and only its *mode* half is lost. The consequence is proven with a control: inside a UI-created `0755` folder a gid-1000 process's file landed group **1000**, while the identical write into the 2775 parent landed group **100**. So **any folder a customer creates through FileBrowser breaks the shared-group chain one level down.** Latent today — every userdata-touching catalog app that declares an identity declares uid/gid **1000**, the same uid FileBrowser runs as, so owner permissions mask it; it bites the day a content app runs as a different non-root uid with gid 1000. The comment at `infra/infra.go:156` is right that the image ignores `-e UMASK` but does not say t | CC |
|
||||
@@ -147,10 +148,11 @@ stopping line that lies.
|
||||
| **R-683** | App updates | P3 | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | — | — | CC |
|
||||
| **R-785** | App updates | P3 | **[P3-LOW] SparkyFitness is pinned 11 releases and a major behind upstream (v0.17.3; upstream v1.7.3, v1.6.0 dated 2026-07-24).** READ 2026-10-01 (`audits/visitors-2026-10-01/C/bench/C1-previous-tag.txt`). **Needs:** an update walk 0.17 → 1.x through the ladder (bench + box), after R-784 is decided. | **OPEN — rank P3-LOW; owner: CC (after R-784)** | — | — | CC |
|
||||
|
||||
## Backup & restore — 30 rows (P2 4, P3 12, P4 14)
|
||||
## Backup & restore — 31 rows (P2 4, P3 13, P4 14)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **R-926** | Backup & restore | P3 | **After a household removes an `after_install` app keeping its backups and then restores it, nobody has checked whether the app page still shows a password that works.** A `type: password` deploy field is NOT in `stacks.PortableSecretEnvVars` (that list is `type: secret` only, `deploy.go:1199`), so it does not travel in the backup unit — the same shape as R-765, where Radicale's restored login was replaced by a value no page showed and every phone lost its calendar. MEASURED on 9202 2026-10-10 for grocy: remove keeping the backups, restore, and **the household's own password still signs in and the seeded data reads back** — so the DATABASE half is right. What was NOT settled is the PAGE half: `app.yaml` holds the value encrypted at rest, so reading it proves nothing (a 76-character string that of course did not sign in), and `POST /apps/grocy/initial-credentials/reveal` answered 404 for this app. Seven apps are in the same class: bookstack, calibre-web, claper, dawarich, grocy, mealie, wger. | **OPEN — the question is unanswered, not the answer bad** | — | Ask the controller what the app page renders for a generated `type: password` after a restore (one endpoint read, one app), and if it differs from the one that works, decide: carry the value in the unit, or re-run `after_install` after a restore | CC |
|
||||
| **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-08 — owner Viktor.** (a) DONE with the operator's yes in chat: `notify_failure` now mails admin@felhom.eu through Resend; proven by one test mail that reached the inbox (`audits/day-2026-10-08/r232/`; no backup was started). (b) partly: the hub database leaves DooPlex nightly to ep0 (R-173); everything else stays on the box. (c)–(h) unchanged. **2026-10-09 (operator yes in chat): (b) and (h) DONE for Gitea and the secrets.** Nightly 00:20 `felhom-dooplex-offsite` pushes Gitea (614 MB of repositories, the `gitea` database dump taken first, `app.ini`) and the nightly GPG secrets export, encrypted with a NEW key, to ep0 `operator` (`host/dooplex-gitea`) on the hub-DB write-only token — no change on ep0; first push 57 s; restore test weekly (manifest, `git fsck` every repo, `pg_restore --list`), run once: 27 805 files, 10 repos OK; `DooplexGiteaOffsiteStale`/`DooplexGiteaRestoreTestStale` in force (`absent()` included, promtool + 2 red-proofs); every backup unit now mails on failure (`OnFailure=`, dry run reached the inbox). **First real Gitea restore:** into a throwaway on the bench, 10/10 repos, product `main` = live, a file byte for byte, a login (`runbooks/gitea-restore.md`). The hub came back too (R-173, closed). **The registry is left out on purpose:** its images rebuild from the code. **Correction to the recon §7:** the API mirror holds ONE repo (`homelab-manifests`), not all — Gitea had a single copy, on DooPlex. `audits/dooplex-survival-2026-10-09/` **LEFT:** (c) append-only for the on-box tree, (d)–(g) as written; the paper copy of the new key (operator). **READY** for the rest | — | — | operator |
|
||||
| **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY — the operator mail is BUILT on main 2026-10-08 (decision 183; controller `a40729a` + hub, ships with the next releases).** The honesty fix is on main too (agent `91b9405`, controller `75b3b39`). Left: the runbook „open a retained package for a household" (design option C, second slice). Design `audits/day-2026-10-08/design-R-304.md`. | R-198, R-199, R-224, R-241 | Release controller + hub; write the retained-package runbook from the 2026-08-12 drill §4; then close | CC |
|
||||
| **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-893.md` — re-verified; and the screen's „your data is back as it was" (`err.backup.db_restore_failed_rolled_back`) is false for files, other volumes and the app version. Question D8 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D8 (`09` §3 decision 192):** yes, both, in that order — first the app stays stopped for support (option C), next „put back exactly as it was" (option A). | **NARROWED — 2026-10-09: the first half DELIVERED in controller 0.304.0. Slice 0 (the live measurement) NOT run: 9202 has no off-site, and a dump cannot be cut by hand inside an encrypted off-site copy on Tester 1; it needs a scratch off-site repo — a plan, not a finding.** **NARROWED 2026-10-08 — the first half (D8 option C, the app held stopped) is built on controller main and ships tomorrow; what stays open is the second half, „put back exactly as it was” (option A: a pre-restore copy of the app's volumes and placed files, a fit check, disk for one copy; the slice-0 measurement on 9202 first).** **OPEN — filed 2026-10-06** **2026-10-06 night: verified in source, no code** — `offbox_reconstitute.go:758` (undo dump from the live DB), `:778-860` (files and volumes from the snapshot), `:804` (the snapshot's definition is written when its version differs), `:898` (the rollback loads the newer dump over the older volume; nothing writes the live definition back). Not a reorder fix: it needs R-638 option B (a rebuilding loader) or a pre-restore volume copy (disk cost; R-685's class). Which state a household gets after a failed off-site replay is the operator's call. Next: the 9202 measurement, then the design. | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC |
|
||||
|
||||
Reference in New Issue
Block a user