Compare commits
49 Commits
8e873ee4a0
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
| 4372b46b44 | |||
| 6f78bd9f0e | |||
| 8925dcc633 | |||
| 0a460bcdea | |||
| a5b4d29af4 | |||
| 68df84bcaa | |||
| bf1798479d | |||
| 7ee00d59a8 | |||
| 900e6d86e2 | |||
| daeed175ca | |||
| 91e4ace9b8 | |||
| e221476ff5 | |||
| 5f060e3d1e | |||
| 86def579ed | |||
| f790734ec2 | |||
| b43705ac7b | |||
| c11d4fbf2e | |||
| 7f3944eff9 | |||
| efe76093ce | |||
| 17af3b65f6 | |||
| 85de3f9b87 | |||
| c8da8e0439 | |||
| 30650cad6e | |||
| d75ad0fdf3 | |||
| 81d04a6ae2 | |||
| 557629d2bf | |||
| e0d884565f | |||
| 301fe4532a | |||
| e8c56c440a | |||
| 1122b5c65b | |||
| 40d34c5576 | |||
| 8012ce4640 | |||
| b26d292a64 | |||
| f5a0aeb0b8 | |||
| 58ce696e16 | |||
| ab2b304987 | |||
| b018ca9092 | |||
| 53d8131b20 | |||
| afba622fb6 | |||
| cea8502f0b | |||
| 5d8922e619 | |||
| d4be9f6ff1 | |||
| 0826e41b31 | |||
| 2b30733b0d | |||
| 3d7a2761fc | |||
| 9bb45eaaa2 | |||
| ff25f1076d | |||
| 1b0678fa05 | |||
| 6da3a7d866 |
@@ -13,7 +13,8 @@ unconditional: true
|
||||
- A register row **you or another CC session filed**, with owner CC, at P3 or a bounded P2, that
|
||||
needs **no operator decision**, touches **no customer data by design**, and introduces **no
|
||||
mechanism nobody has measured**. Smallest first.
|
||||
- A defect you find while exercising the product, filed as a row **before** you fix it.
|
||||
- A defect you find while exercising the product, filed as a row **before** you fix it — **unless it is small**:
|
||||
a small finding is fixed in the session and never filed (the size rule, `OPEN-ITEMS.md` „How a row is filed").
|
||||
- Hygiene: register compression, stale citations, rows with no owner, documents that contradict
|
||||
live source.
|
||||
|
||||
@@ -43,7 +44,8 @@ it **first** in the morning note. A decision you cannot write in that shape is o
|
||||
6. **One release per repo per session**, with a CHANGELOG entry (controller: with its `MinAgent`
|
||||
line), REPORT overwritten, floor raised to deliver it. **No golden unless a drill or fresh install
|
||||
needs one** (the waiver, R-468). **No `--no-verify`.**
|
||||
7. **An enumerated gap becomes a row in the same session.** Prose is not a record.
|
||||
7. **An enumerated gap becomes a row in the same session — or, if it is small, is fixed in it** (the size rule).
|
||||
Prose is not a record.
|
||||
8. **Hungarian text is searched with ASCII fragments**, with a positive and a negative control.
|
||||
9. **Never leave a half-state.** If time runs out, revert to clean and say what was reverted.
|
||||
10. **Teardown, three layers, stated** — machine, host, hub — or "provisioned nothing".
|
||||
|
||||
+54
@@ -16,6 +16,60 @@
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
|
||||
|
||||
> **2026-10-05 (late night) — burn-down round 2 (releases).** Register 292 → 199 (1 opened: R-888; 94 closed: 43
|
||||
> accepted by the operator 18:23, 51 fixed). Releases: hub v0.137.0 (`557629d`, deployed), agent v0.147.0 (tag, sha
|
||||
> `642c4d19…`, bundle `326527d0…`, signed jobs to 3 boxes), controller v0.297.0 (`1453cfc` + `6f1ba1f`), golden 0.297.0
|
||||
> (`8cebc42e…`, vouched with agent 0.147.0 / min_agent 0.131.0; floors 0.297.0 for demo-hp, demo-felhom, tester-1),
|
||||
> catalog `4828dc7`. R-124: recipe root namespace = `""` (+ runbook). R-887 mechanism from Gitea's log: a FetchTask the
|
||||
> runner abandons after assignment → zombie stop after ~10 min; load = an outside crawler + the session's own 15-page CI
|
||||
> polling (now one `runs?head_sha=` call per minute). New gates: `stands` (felhom.eu), `gofmt` (controller; NOT
|
||||
> CHECKED out loud on the Go-less runner). R-469 not done: the permission check refused the catalog CLAUDE.md edit.
|
||||
> Report: `REPORT-burndown2-2026-10-05.md`.
|
||||
|
||||
> **2026-10-05 (night) — the burn-down (no release; DooPlex/ep0 untouched).** Register 336 → 292 (1 opened — R-887 CI runner fault — 45
|
||||
> closed): 24 fixed by later work + 2 duplicates (each re-checked; `audits/burndown-2026-10-05/partA-table.md` holds all
|
||||
> 317 P3/P4 verdicts), 19 small fixes with tests/red-proofs (catalog `29ac711`, agent `d833163`, controller `114ff27`,
|
||||
> felhom.eu `ab2b304`…). New gate `script-tests` (every `scripts/**/test_*.py` per push, R-885); `closed_register_gate`
|
||||
> RULE 4; `site_gates` page walk (R-423, exemption dropped); hub build no longer pushes `:latest`. **The size rule**
|
||||
> (`OPEN-ITEMS.md` „How a row is filed", the four `unprompted-work.md` copies, `PROMPT-TEMPLATE.md` §9.2/N.7.3): a
|
||||
> small finding is fixed in-session, not filed; reports state four numbers. Controller/agent/hub carry an
|
||||
> `## unreleased` CHANGELOG head for comment/test-only changes — the next release folds it in. Part C: 45 rows await
|
||||
> the operator's „close as accepted" (STATUS). Report: `REPORT-burndown-2026-10-05.md`.
|
||||
|
||||
> **2026-10-05 (evening) — the hub database off DooPlex (hub v0.136.0; operator rulings `09` 125–127).** R-173 option A
|
||||
> IN FORCE: hub `internal/dbsnap` writes `VACUUM INTO /data/snapshots/hub-<UTC>.db` at 02:00 Budapest (keep 2, `ErrBusy`
|
||||
> on overlap, start-up catch-up when >24 h; `05` §16.3); DooPlex `scripts/hub-db-backup/` (installed by `install.sh` to
|
||||
> `/usr/local/sbin`, units + timers in `/etc/systemd/system`, config `/etc/felhom-hub-backup/{env,token-push,
|
||||
> token-restore,enc.key}` all root 0600) pushes at 02:30 to ep0 `felhom-offsite` ns `operator` as `dooplex-hub@pbs!push`
|
||||
> (`DatastoreBackup`), restore-tests Sun 04:30 as `!restore` (`DatastoreReader`); ep0 prune job `prune-operator-hubdb`
|
||||
> (03:45, keep-daily 14, keep-weekly 8, max-depth 0). Success-only textfile metrics → `HubDBBackupStale` (26 h, critical)
|
||||
> and `HubDBRestoreTestStale` (8 d) in homelab-manifests `backup-freshness`. `hub/cmd/hubdb-check` opens a restored copy
|
||||
> with the seal key (runbook §3). Hub PVC 2 Gi + label `enabled`. Collateral: Longhorn instance-manager restart (R-882),
|
||||
> zipline pinned 4.7.0 (R-883). R-519 proven live on 9202 (now controller 0.296.0) and CLOSED. Alarm proven end to end (fired 14:22Z, Alertmanager email sent, 0 failed; cleared by a real push); Alertmanager nflog/silences unwritable since the restart (R-886). Register 332 → 336.
|
||||
> Report: `REPORT-hub-db-offsite-2026-10-05.md`.
|
||||
|
||||
> **2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller
|
||||
> v0.296.0, agent v0.146.1 + bundle `42333e96…`, golden 0.296.0 vouched with agent 0.146.1, min_agent 0.131.0).** CC decisions
|
||||
> 119–124, *operator may reverse*. Hub: R-135 a cookie-less state change needs Basic + `X-Felhom-Operator` (`05` §16.1); R-133
|
||||
> console passwords sealed with the off-site seal/key (`05` §16.2; 4 rows sealed live); R-604/R-530 the System page's "Version
|
||||
> floors" table + Agent cell, `agent_behind` (7 d) and `floor_raise_skipped` (`05` §5, `08` §6.3); R-508 no-e-mail banner.
|
||||
> Controller: R-519 `run_record.go` (a cut run said until a complete one; live cut refused by the permission check — operator
|
||||
> asked), R-518 copy with today's measurement (5 min 47 s, demo-hp). Agent: R-861 narrowed (`03` §3.1 — exact sudo regexes,
|
||||
> `felhom-priv-apply`, fixed hook/parent files, signed update verified as root, escrow paths pinned; three residuals named);
|
||||
> **R-880: a bundle that adds a path needs a STEP bundle** (`scripts/build-step-bundle.py`, `runbooks/config-bundle.md`).
|
||||
> R-173 measured (backed up only on DooPlex, by label drift) + `runbooks/RUNBOOK-hub-db-offsite-backup.md`; decision in
|
||||
> STATUS. Closed R-133, R-135, R-508, R-509, R-530, R-604, R-880; opened R-879, R-881; register 336 → 332.
|
||||
|
||||
> **2026-10-05 (afternoon) — a box that is not always on (controller v0.295.0, agent v0.145.0, hub v0.134.0, golden 0.295.0
|
||||
> vouched with agent 0.145.0, min_agent 0.131.0).** Rulings 109–111 (`09` §3: R-871 option A — a missed night runs once when the
|
||||
> box comes back; the household's banner; Tester 2 read only). CC decisions 112–118, *operator may reverse*. Design `07`
|
||||
> §6.1.1 (new): `internal/nightchain` ledger (`night-ledger.json`, an ATTEMPT record) + `CatchUp` (start + resume triggers,
|
||||
> 15 min, backup legs only, `catchUpLegs`) + `scheduler.DailyLateLimit` (60 min) + `quiesce.SetCatchUpFn`; the banner
|
||||
> (`nightchain.ComputeBanner`, 26 h, metrics-record suggestion, durable dismissal). Hub: R-872 down boxes judged at 48 h/72 h
|
||||
> (open, dated check 2026-10-06), R-873 household liveness mail weekly, `backup_catchup_done` allowed. Agent: R-876 dpkg journal
|
||||
> repair (live, second crash), R-874 first restore-test check 30 min after start (live), R-875. Closed R-871, R-873..R-877;
|
||||
> opened R-877 (closed), R-878. Floors: demo-hp, demo-felhom, tester-1 at 0.295.0. Report: `REPORT-catchup-2026-10-05.md`.
|
||||
|
||||
> **2026-10-05 (day) — the night's fixes (controller v0.294.0, agent v0.144.0 + v0.144.1, golden 0.294.0 vouched with agent
|
||||
> 0.144.1, min_agent 0.131.0).** Rulings 100–103 (`09` §3: Tester 1's CF tokens NOT rotated → R-870; A1 by day on demo-hp
|
||||
> only, with the operator's go; Tester 2 is a laptop off at night; re-sign Tester 2 only if online — it never was).
|
||||
|
||||
@@ -0,0 +1,122 @@
|
||||
# REPORT — the burn-down: the open-items list gets shorter — 2026-10-05 (night)
|
||||
|
||||
| Part | Result |
|
||||
|---|---|
|
||||
| **A** — stale sweep, every P4 then every P3 row, oldest first | **done** — 317 rows checked against `main`; 24 closed as fixed by later work, 2 closed as duplicates (facts merged); table `documentation/audits/burndown-2026-10-05/partA-table.md` |
|
||||
| **B** — small fixes, batched | **done** — 19 rows fixed and closed in four repos, each with a test (and a red-proof where a check changed); **no release** (see below) |
|
||||
| **C** — the „not worth doing" list | **done** — 45 rows in `STATUS.md`, one line each with my pick; none closed; 2 more (leaked tokens) listed as actions for the operator |
|
||||
| **D** — stop the growth | **done** — the size rule and the four-numbers rule, in the register, the rules file (4 copies) and the report template |
|
||||
|
||||
| Rows before | Rows after | Opened | Closed |
|
||||
|---|---|---|---|
|
||||
| **336** | **292** | **1** | **45** |
|
||||
|
||||
Counted by `register_shape_gate.py`'s method (`| **R-n** |` lines in `OPEN-ITEMS.md`). Target was ≥ 40 fewer: 44 net (45 closed, 1 opened — R-887, below).
|
||||
|
||||
## Baselines (re-verified at the start)
|
||||
|
||||
felhom.eu `53d8131b20` (hub v0.136.0) · controller `7690c27f86` (v0.296.0) · agent `e06ed97fa8` (v0.146.1) · catalog
|
||||
`917a779cca`. Register 336: P2 19, P3 139, P4 178.
|
||||
|
||||
## Why no release
|
||||
|
||||
The brief allowed one release per repo, but also said „DooPlex: no change". The hub runs on DooPlex, so a hub release
|
||||
could not be deployed; delivering a controller or agent release needs a floor raise or signed jobs through the hub. So
|
||||
Part B fixed only what needs no release: documents, comments, tests, gates, catalog tooling. Controller, agent and hub
|
||||
each carry an `## unreleased` head in `CHANGELOG.md` for these; the next release folds it into its own entry. Code
|
||||
rows that need a release stay open (they are in the Part A table as STILL-TRUE-SMALL with their fix described).
|
||||
|
||||
## Part A — the sweep
|
||||
|
||||
Eight read-only checker agents took 40 rows each (P4 then P3, oldest id first). No machine was reached. Groups over all
|
||||
317: FIXED-BY-LATER-WORK 29, DUPLICATE 2, STILL-TRUE-SMALL 91, STILL-TRUE-NOT-SMALL 129, NOT-WORTH-IT 43, UNCHECKED 23.
|
||||
|
||||
**Every FIXED and DUPLICATE verdict was re-checked before closing**: 22 cited proof lines re-grepped (all present; one
|
||||
first missed by my own shell quoting). Of the 29 „fixed": **24 closed**; **R-274 kept** (only one of its two halves was
|
||||
checked); **R-700, R-704, R-706, R-723 moved to Part C** — their code fix and tests exist, but each row waits for a live
|
||||
observation, and closing them would silently drop that. R-766 was additionally checked in the live hub image (read
|
||||
only): the new app logos are in `/usr/share/felhom/assets-seed/`.
|
||||
|
||||
Closed as fixed by later work: R-184, R-207, R-287, R-289, R-373, R-390, R-427, R-437, R-464, R-501, R-602, R-617, R-705,
|
||||
R-766, R-50b, R-121, R-200, R-450, R-489, R-573, R-622, R-635, R-235, R-282. Duplicates: R-755 → R-762, R-446 → R-440
|
||||
(the unique fact of each moved into the survivor). Each closed row's evidence is in `CLOSED-ITEMS.md`.
|
||||
|
||||
## Part B — the fixes, by repo
|
||||
|
||||
| Repo (commit) | Row | Fix | Test / red-proof |
|
||||
|---|---|---|---|
|
||||
| felhom.eu `ab2b304` + `58ce696` | R-885 | gate `script-tests` runs every `scripts/**/test_*.py` per push (13 → 15 suites, ~20 s); a Python-sqlite3 stand-in when CI has no `sqlite3` | 5 decoys; 2 red-proofs convict |
|
||||
| felhom.eu `f5a0aeb` | R-376 | marker legend in `08`, `09`, `11` (all 11 numbered docs now) | — (docs) |
|
||||
| felhom.eu `f5a0aeb` | R-817 | dated clarification under `09` decision 56 — same image, seen before and after a swap (`controllerswap.go:236-240`, `:289`) | — (docs) |
|
||||
| felhom.eu `f5a0aeb`, controller `e563733` | R-818 | correction notes under hub v0.109.0, controller v0.224.0/v0.225.0 — those findings never had rows | — (docs) |
|
||||
| catalog `29ac711` | R-799 | MeTube `POST /add` sends `download_type` | test + red-proof |
|
||||
| catalog `29ac711` | R-761 | logo name `.svg` then `.png` in template comment, REUSE, checklist | — (comment) |
|
||||
| catalog `29ac711` | R-391 | CLAUDE.md: no observations section by convention | — (docs) |
|
||||
| agent `d833163` | R-291 | retention record names its source (R-267 prune, R-287); dead reader dropped | reader still reads 10 |
|
||||
| agent `d833163` | R-348 | restart comment says what a restart blanks | — (comment) |
|
||||
| controller `114ff27` | R-263 | „only writer that GRANTS"; AST scan of internal/ + cmd/ | 2 red-proofs convict |
|
||||
| controller `114ff27` | R-368 | `IsDefault` comment names the form as the one that applies it | — (comment) |
|
||||
| felhom.eu `b26d292` | R-418 | gate list in the docstring = `GATES` | test + red-proof |
|
||||
| felhom.eu `b26d292` | R-345 | no `:latest` in `hub/Makefile` **and in the real release script `build-hub.sh`** (found during the fix; nothing pulls it) | test; 2 red-proofs |
|
||||
| felhom.eu `b26d292` | R-416 | `closed_register_gate` RULE 4 — duplicate id in CLOSED-ITEMS | decoy + red-proof |
|
||||
| felhom.eu `b26d292` | R-261 | `CountSelfBindTokens` documented as a test accessor | — (comment) |
|
||||
| felhom.eu `b26d292` | R-262 | hostRestoreTest is a deliberate subset; the test found a **third** unmodelled agent field, `skipped` (by design) | test; 2 red-proofs |
|
||||
| felhom.eu `b26d292` | R-286 | standing rule 3: a control from a DIFFERENT channel (both CLAUDE.md copies) | instructions gate |
|
||||
| felhom.eu `b26d292` | R-588 | one home for ISO release records; 1.28.0 pointer | — (docs) |
|
||||
| felhom.eu (final commit) | R-423 | `site_gates` fails on a page `PAGES` does not list; exemption dropped | decoy + red-proof |
|
||||
|
||||
Red-proof records: `documentation/audits/burndown-2026-10-05/` (`r885-`, `r263-`, `r262-`, `r423-red-proof.txt`). The
|
||||
red-proofs of R-418, R-345, R-416 and R-799 were run in the session and each convicted, but their output was **not
|
||||
saved to a file** — re-run them by mutating as described in each CHANGELOG entry. Suites: hub `go test ./...` green; controller
|
||||
`go build/vet/test ./...` green; agent build + vet green; all four repos' gates green at each push.
|
||||
|
||||
**Fixed without a row** (the new rule, used once): the `:latest` push in `build-hub.sh` — folded into R-345's fix.
|
||||
|
||||
**CI was red for five pushes, said plainly:** the new `script-tests` gate ran the hub-DB script tests on the CI runner
|
||||
(Alpine, BusyBox) for the first time, and 9 of 15 failed — BusyBox `date` cannot read `…T…Z`. Locally (GNU tools) and in
|
||||
the pre-push hook everything was green, so the push went through and CI caught it: jobs 1351, 1352, 1353, 1358, 1359 =
|
||||
failure, each mailed to the operator (`RESEND-ACCEPTED` in the log). Fixed in `40d34c5` (a form both `date`s read;
|
||||
red-proved with a BusyBox-only PATH: the old form gives the same 9 failures) → **job 1360 = success**. The gate did what
|
||||
it was built for, on its first day.
|
||||
|
||||
**A second CI fault, not mine to fix (R-887, opened):** felhom.eu job 1361 (`1122b5c`) and controller job 1357 (`114ff27`) were
|
||||
never picked up by the runner (no `task` line in its log; task id = job id + 1), and Gitea failed them after ~10–13 min with
|
||||
no log and no alarm. Re-run through the API: **controller 1357 → success (32 s)**; felhom.eu 1361 → never picked up again,
|
||||
failure. Suspected stale runner registration; my token cannot list runners. Final CI per repo: agent `d833163` → job 1356
|
||||
success; catalog `29ac711` → job 1355 success; controller `114ff27` → job 1357 (re-run) success; felhom.eu `40d34c5` → job
|
||||
1360 success (the later docs-only commits: see the final check below). The copy installed on DooPlex is the previous revision (same behaviour on GNU
|
||||
`date`); not reinstalled because DooPlex was out of scope.
|
||||
|
||||
**One slip, said plainly:** R-885's gate commit (`ab2b304`) went out WITHOUT closing the row — my closing file had a
|
||||
JSON error. It was closed one commit later (`58ce696`). „Close in the same commit" was broken once.
|
||||
|
||||
## Part C — in `STATUS.md`
|
||||
|
||||
45 rows, one line each: what it is, what fixing costs, what happens if never, my pick (close for all 45). Two more —
|
||||
R-831 and R-870, leaked tokens — are listed as actions for the operator (pick: keep until rotated).
|
||||
|
||||
## Part D — the rule, where it lives, quoted
|
||||
|
||||
`documentation/backlog/OPEN-ITEMS.md`, „How a row is filed":
|
||||
|
||||
> **Fix small, do not file (the size rule — operator brief 2026-10-05, the burn-down).** The register grew because every
|
||||
> session closed a few rows and filed a few small new ones. So: **a finding that is cosmetic or small — fixable in the
|
||||
> session in about 30 minutes, in a repo the session may change — is FIXED in that session, with a test where it changes
|
||||
> behaviour, and NOT filed.** It is recorded in that repo's `CHANGELOG.md` and in the session report under „fixed without
|
||||
> a row". **Only a finding that needs a decision, a design, a larger build, or a change the session may not make** (a
|
||||
> protected machine, a repo out of scope, a release budget already spent) **becomes a row.** This narrows — it does not
|
||||
> repeal — „an enumerated gap becomes a row": a small gap leaves the session fixed, which is a record too.
|
||||
>
|
||||
> **Every session report states four numbers:** rows before, rows after, rows opened, rows closed — counted the way
|
||||
> `register_shape_gate.py` counts (`| **R-n** |` lines in this file).
|
||||
|
||||
Carried, so no instruction contradicts it: `.claude/rules/unprompted-work.md` §1 and §3.7 (all four identical copies:
|
||||
workspace root, felhom.eu, controller, catalog — commits `b018ca9`, `df85a07`, `7a19491`) and
|
||||
`documentation/PROMPT-TEMPLATE.md` §9.2 (the exception) and N.7.3 (four numbers). The morning-note line in the rules file
|
||||
already asked for opened/closed/before/after. No gate was weakened or bypassed.
|
||||
|
||||
## Teardown
|
||||
|
||||
Provisioned nothing. No machine changed: the only reads were the hub pod's asset directory (R-766) and none on any box.
|
||||
Checker agents were read-only. Scratch: the batch files and results in the session scratchpad; the results are copied to
|
||||
the audit folder.
|
||||
@@ -0,0 +1,115 @@
|
||||
# REPORT — burn-down round 2: the operator's answer recorded, R-887 re-diagnosed, small rows fixed WITH releases — 2026-10-05 (late night)
|
||||
|
||||
| Part | Result |
|
||||
|---|---|
|
||||
| **A** — rulings, then R-887 | **done** — rulings commit `301fe45` (count after: **249**); R-887 re-diagnosed from the logs (the restart idea refuted; the mechanism then SEEN in Gitea's own log), dated check 2026-10-12 |
|
||||
| **B** — R-124, the small rows, the 23 unchecked | **done** — R-124 fixed (agent v0.147.0 + runbook); 50 more rows fixed and closed with tests and red-proofs; the 23 checked from source (1 duplicate closed, facts added to 8 rows, the rest left as they need a live box or a decision) |
|
||||
| **B.4** — releases, delivered the normal way | **done** — hub v0.137.0 deployed; agent v0.147.0 released + signed jobs (binary and bundle) to demo-hp, demo-felhom, Tester 1; controller v0.297.0 + golden 0.297.0 baked, vouched, floors raised, all three boxes on 0.297.0; catalog pushed |
|
||||
| **C** — numbers and record | **done** — STATUS shows 199 and asks nothing about the closed list |
|
||||
|
||||
| Rows before | Rows after | Opened | Closed |
|
||||
|---|---|---|---|
|
||||
| **292** | **199** | **1** (R-888) | **94** (43 accepted by the operator + 51 fixed/merged) |
|
||||
|
||||
Counted by `register_shape_gate.py`'s method. Target ≤ 220: met.
|
||||
|
||||
## Baselines (re-verified at the start)
|
||||
|
||||
felhom.eu `e8c56c440a` (hub v0.136.0) · controller `114ff2761a` (v0.296.0) · agent `d83316326e` (v0.146.1) · catalog
|
||||
`29ac711d26` · golden 0.296.0 · register 292. The agent clone had a stray `scripts/__pycache__/` from round 1 — removed.
|
||||
|
||||
## Part A — the rulings commit and R-887
|
||||
|
||||
- `301fe45`: 43 rows closed as „accepted by the operator, 2026-10-05", each with its one-line reason from the list;
|
||||
R-124 and R-698 kept (R-698 owner → operator); R-831/R-870 carry the not-rotated rulings; R-887 records the screenshot
|
||||
(one runner, ID 2, online). STATUS: the list and the rotate/runners requests removed. **Count after: 249.** CI run
|
||||
1363 success.
|
||||
- **R-887, from the logs:** the runner's last restart was 13:24:42Z; the lost attempts started 15:05–15:46Z — **not a
|
||||
restart**. Four lost attempts (not two): each without a runner `task` line, each failed at a :38-second mark 10–13 min
|
||||
after assignment. Gitea's log for that hour had rotated. **Then it happened again at 17:15Z with the log intact:**
|
||||
`slow POST …/RunnerService/FetchTask for 10.42.0.42, elapsed 3192ms` → `context canceled` → 17:28:39
|
||||
`clear_tasks.go … stopTasks() … task 1371` — the runner abandoned its fetch after Gitea assigned the task; Gitea's
|
||||
zombie stop failed it. Load at that minute: an outside crawler on public commit pages, and this session's CI waiter
|
||||
(15-page job listings at 13–31 s each). The waiter now makes ONE `runs?head_sha=` call a minute. A lost run re-runs
|
||||
with `POST …/actions/runs/<id>/rerun` (used twice: controller run 1357 → success; catalog run 1368 → success).
|
||||
**Dated check 2026-10-12** in DUE-CHECKS. The fix on DooPlex (runner fetch timeout, crawler) is the operator's.
|
||||
|
||||
## Part B — fixes by repo
|
||||
|
||||
**agent v0.147.0** (`f1b9b41`, CI 1365; tag `v0.147.0`; binary sha256 `642c4d19…`, bundle `326527d0…`, verified by
|
||||
download; CHANGELOG `208fac8`, CI 1367): R-124, R-118, R-269, R-317 — red-proofs `audits/burndown2-2026-10-05/r124-red-proof.txt`,
|
||||
`agent-red-proofs.txt`. **Delivery:** vouched (agent 0.147.0, golden 0.296.0 first), signed `agent_update` ×3, then
|
||||
`agent_config_update` ×3 (felhom-op-1, ttl 45 m); hub System page: demo-hp, demo-felhom, Tester 1 — agent 0.147.0,
|
||||
root files 0.147.0 (`delivery/`). Tester 2 offline — nothing sent.
|
||||
|
||||
**hub v0.137.0** (`557629d`, CI 1369; manifest `81d04a6`; CI 1370): R-277, R-581, R-600, R-544, R-855, R-134, R-92,
|
||||
R-292, R-599, R-725, R-728, R-208 (hub half) — red-proofs `felhom-eu-red-proofs.txt`. **Deployed:** ArgoCD Synced/Healthy
|
||||
at `d75ad0f`, image `felhom-hub:0.137.0`, `felhom-hub 0.137.0 starting`, healthz 200; R-855's new line seen live
|
||||
(„after 2 healthy ring-0 night(s)"). (The build ran while a helper was still appending to an audit text file outside
|
||||
`hub/` — the image is the committed `hub/` tree; said here because the clean-tree gate is literal.)
|
||||
|
||||
**felhom.eu gates/tools/docs** (same commits): R-819 (`stands` gate), R-857, R-555, R-364 (`hu_grep.py` + REUSE.md),
|
||||
R-587, R-571, R-129 (demo-hp authenticates with DooPlex's own key — corrected everywhere it said „no key"), R-124 runbook.
|
||||
New script tests pass under a BusyBox + bash + python3 + git PATH (the CI runner's tools): 19/19.
|
||||
|
||||
**controller v0.297.0** (`1453cfc`; CI run 1371 **FAILED** — the new gofmt gate was INCONCLUSIVE on the Go-less runner;
|
||||
fixed in `6f1ba1f`, CI 1372 success): R-591, R-568, R-567, R-363, R-547, R-10, R-552, R-251, R-104, R-619, R-362, R-675,
|
||||
R-256, R-257, R-240, R-365, R-425, R-565, R-564, R-603, R-454, R-208, R-457 (swept, nothing left) + two twins found and
|
||||
fixed on the way (the top-bar countdown at 0 days; nine more shared references in `deepCopyStack`). Red-proofs
|
||||
`controller-red-proofs.txt` (two first attempts that did not convict are marked, with valid re-runs). **MinAgent 0.131.0.**
|
||||
**Image** `felhom-controller:0.297.0`. **Golden 0.297.0** baked per RUNBOOK §4.0–4.1 (`documentation/tests/golden-0.297.0-2026-10-05/`:
|
||||
all pass markers, round trip sha `8cebc42e…`, token leak 0 with a working control, teardown to `virgin`). **Vouched**
|
||||
(agent 0.147.0, golden 0.297.0, min_agent 0.131.0) and **floors** 0.297.0 for demo-hp, demo-felhom, tester-1.
|
||||
**Delivered:** demo-hp and demo-felhom `felhom-controller:0.297.0 … (healthy)`; Tester 1 reports Controller 0.297.0
|
||||
(„Controller frissítve: 0.296.0 → 0.297.0").
|
||||
|
||||
**catalog** (`4828dc7`; CI run 1368 lost by R-887, re-run success): R-593, R-760, R-594, R-605, R-781, R-806 (scheme half;
|
||||
row narrowed), plus a stale runner test (expected 11 gates, 12 exist) and a test that never ran (outside its class) —
|
||||
fixed, not filed.
|
||||
|
||||
**Not done, and why:** R-469 and R-605's exit-code line in the catalog's `CLAUDE.md` — **the permission check refused
|
||||
the instruction-file edit**; the operator is asked (rule 5). R-126 needs an operator choice. R-325 needs a same-step
|
||||
felhom.eu gate change (left). R-377 (CONTEXT headings) not attempted. Installer rows (R-179, R-180, R-275, R-276, R-306,
|
||||
R-130, R-310, R-881), R-136 (logs every operator out), R-502 (Docker in CI), R-798 (a live app definition) and the
|
||||
larger controller rows (R-492, R-569, R-575, R-615, R-616, R-498, R-718) were left on purpose.
|
||||
|
||||
**Opened:** R-888 — two report fields the hub never reads (a decision). **Seen, not a row:** Tester 1's crash guard reads
|
||||
TRIPPED since 07:57Z — the morning's two deliberate test crashes; it re-arms by itself after 24 h (`runbooks/crash-guard.md`).
|
||||
|
||||
## The 23 rows the first burn-down could not check
|
||||
|
||||
Checked from source by a read-only agent (`audits/burndown2-2026-10-05/unchecked-results.jsonl`). Closed: R-350 (duplicate
|
||||
of R-132, facts merged). Facts added to the open rows R-607, R-883, R-886, R-884, R-756, R-91, R-338, R-488. The three
|
||||
„not worth it" ones are on STATUS for the operator. The rest need a live box reading (the settle command is in the table).
|
||||
|
||||
| Row | Group | Evidence / how to settle (abridged) |
|
||||
|---|---|---|
|
||||
| R-76 | UNCHECKABLE-FROM-SOURCE | Image changed since the 1.3.3 finding: felhom-controller@114ff27 controller/internal/infra/infra.go:27 FileBrowserImage = "gtstef/filebrowser:1.5.6-stable". The comment infra.go:207-208 still asserts folders come out '2775 with the parent's setgid' -- the exact claim R-76 measured false on 1.3.3; no test pins it (git log --gre |
|
||||
| R-91 | UNCHECKABLE-FROM-SOURCE | Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of delet |
|
||||
| R-209a | UNCHECKABLE-FROM-SOURCE | Pure live state on DooPlex (whether a reboot has happened and the post-boot check passed). No source claim to test. — settle: uptime -s; cat /var/log/felhom-store-postboot-check.log; ls -d /var/lib/containerd.pre-move-2026-08-05; df -h / |
|
||||
| R-337 | NOT-WORTH-IT | The row's first question ('establish the intended refresh path') is answered by source: GET /backup/status reads only the agent's in-memory store (felhom-agent@d833163 internal/localapi/server.go:1258 -> pickLatestBackup :1304-1318), and the ONLY writer is the job goroutine after the whole runner returns: server.go:885 b, err : |
|
||||
| R-375 | NOT-WORTH-IT | The signal (audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:168-171) is pvesm status showing felhom-pbs Total/Used/Avail = 0. Nothing in the product consumes those numbers for a PBS target: felhom-agent@d833163 internal/backup/runner.go:265 if st == nil // st.Type == "pbs" // st.Avail <= 0 { return true, "" } (space preflight sk |
|
||||
| R-488 | STILL-TRUE-SMALL | The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded interval := 5 * time.Second and time.Sleep(3 * time.Second) // initial settling time, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go: |
|
||||
| R-504 | UNCHECKABLE-FROM-SOURCE | Live HTTP behaviour of iso.felhom.eu; curl to hosts is outside this checker. Source side: documentation/runbooks/VOLUNTEER-first-hour.md:14 still says the root has no index (R-504); the download page exists at website/letoltes.html. — settle: curl -sI https://iso.felhom.eu/ / head -1 |
|
||||
| R-644 | UNCHECKABLE-FROM-SOURCE | Live scratch-box state. Source context: app-catalog-felhom.eu templates/gokapi/docker-compose.yml:26-29 seeds config.json with an EMPTY Password only when config.json is absent, then runs --deployment-password; a config.json that exists with an empty/plain password (e.g. a restored volume or an interrupted first boot) matches |
|
||||
| R-814 | UNCHECKABLE-FROM-SOURCE | Hetzner account state; nothing in source records a deletion. — settle: Hetzner Storage Box API (read-only): GET https://api.hetzner.com/v1/storage_boxes/611421 with the operator's API token (stored out-of-band) -> 404 = deleted, else read .storage_box.status |
|
||||
| R-815 | UNCHECKABLE-FROM-SOURCE | PBS server-side state on ep0; no GC completion record in the docs (grep). — settle: ssh root@ep0 'proxmox-backup-manager garbage-collection status felhom-offsite; proxmox-backup-manager task list --all --limit 20 / grep -i garbage' |
|
||||
| R-884 | UNCHECKABLE-FROM-SOURCE | Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml prom/prometheus:v3.14.0 -> v3.15.0 (monitoring.yaml:419), and the monitoring Application has no automated syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an un |
|
||||
| R-132 | UNCHECKABLE-FROM-SOURCE | Whether HUB_PW was rotated is not in source. The hub stores a UI-set password in hub_settings with an updated_at column: felhom.eu hub/internal/store/store.go:2182 key operator_password_hash, setSetting :2200-2207 writes updated_at = datetime('now'). No commit records a rotation (git log --grep rotate/HUB_PW since 2026-07-31 |
|
||||
| R-298 | UNCHECKABLE-FROM-SOURCE | Template gate unchanged: felhom-controller@114ff27 controller/internal/web/templates/storage.html:364 if(d.role==='user-data'){ else :368 protected, no actions. Dependency R-280 is CLOSED (CLOSED-ITEMS.md:600, v0.211.0). Whether the bug bites depends on the role the agent gives the drive: felhom-agent@d833163 internal/storage/ |
|
||||
| R-338 | UNCHECKABLE-FROM-SOURCE | nodes.md:86-88 still claims demo-hp is on the R-50 island (local_api on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent con |
|
||||
| R-350 | DUPLICATE | of R-132 — Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence. |
|
||||
| R-542 | NOT-WORTH-IT | Still true in source, and by design: felhom-agent@d833163 internal/localapi/disks.go:438 initialize = append(initialize, c) // every unclaimed disk can be initialized; a disk mounted under /mnt/felhom-drives counts as UNCLAIMED on purpose (internal/storage/claim.go:88-103, R-220, so drives survive a guest rebuild), and the mkf |
|
||||
| R-607 | STILL-TRUE-SMALL | Diagnosed from source (both questions the row asks). felhom-controller@114ff27 controller/internal/sync/sync.go:236-241 rescans ONLY if len(newApps) > 0 // len(updated) > 0, and :257-258 says 'nincs változás' when both are empty. updated counts stack-dir copies only (copyTemplates, :447 updated = append(updated, appName) a |
|
||||
| R-683 | UNCHECKABLE-FROM-SOURCE | Evidence in repo: audits/night-2026-09-24/E/round-03-controller.log is the POST-cut log only (first lines 12:05:09Z: update.go:1335 'interrupted in verifying (started 2026-09-24T12:04:21Z)', :589 undo copies '.pre-update-20260924T120426Z'); the pre-cut log that would show a backing-up phase is lost (row says so). No later power- |
|
||||
| R-756 | UNCHECKABLE-FROM-SOURCE | Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 if !m.DriveLive(hddPath) -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 return m.isMountPoint(hddPath) -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH |
|
||||
| R-862 | UNCHECKABLE-FROM-SOURCE | Waits on the operator's by-hand bootstrap on Tester 2; no commit records it (felhom.eu log since 2026-10-04). runbooks/config-bundle.md:77 'CC sends the bundle by the signed job and reads it back on the System page'; a box behind the vouched bundle for 7 days raises os_config_bundle_behind (:83). — settle: Ask the operator wheth |
|
||||
| R-882 | UNCHECKABLE-FROM-SOURCE | Longhorn instance-manager runtime state on DooPlex; nothing in homelab-manifests addresses it (no commit since 87dfc29 names it). — settle: sudo kubectl -n longhorn-system get pods -l longhorn.io/component=instance-manager -o custom-columns=NAME:.metadata.name,START:.status.startTime ; systemctl show k3s containerd iscsid -p Act |
|
||||
| R-883 | STILL-TRUE-SMALL | homelab-manifests@87dfc29 still has moving tags (grep image lines without a numeric tag): admin-system/toolbox.yaml:12 nicolaka/netshoot:latest (a bare Pod); calibre-system/cwa.yaml:826 calibre-web-automated:dev; outline-system/outline.yaml:270 minio/minio:latest; tandoor-system/recipe-importer.yaml:26 gitea.dooplex.hu/admin/rec |
|
||||
| R-886 | STILL-TRUE-SMALL | homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root nobody image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts |
|
||||
|
||||
## Teardown
|
||||
|
||||
Machines: drill VM — build guest destroyed, token/scripts/log shredded, powered off, disk back on `virgin`. Boxes: only the
|
||||
normal deliveries above. Hub: only the deploy, the vouch and the floors. Scratch secrets (hub password file, hub key file,
|
||||
signed envelopes) are shredded at the end of the session.
|
||||
@@ -0,0 +1,78 @@
|
||||
# REPORT — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (2026-10-05, afternoon)
|
||||
|
||||
Brief: "a box that was off at night catches up when it comes back (decision A) …" (operator, 2026-10-05). Evidence:
|
||||
`documentation/audits/catchup-2026-10-05/`. Architecture read before the claims: `07` §6.1, `08` (cool-downs, §6.2–6.3),
|
||||
`09` §3 decision 11, `11` §5.4.1 and §8; the Part F spike (`audits/night-fixes-2026-10-05/partF/FINDINGS.md`).
|
||||
|
||||
## 1. The Part table
|
||||
|
||||
| Part | State | Note |
|
||||
|---|---|---|
|
||||
| §1 rulings recorded first | **done** | `09` decisions 109–111 |
|
||||
| A.0 — the design | **done** | new `07` §6.1.1 (`[DESIGN — ruled 2026-10-05]` for 109–110; CC decisions 112–118 "operator may reverse"), `08` §6.4 |
|
||||
| A — the catch-up (R-871) | **done** | spike on 9202 first (off across 09:05 → `db-dump scheduled for 2026-10-06 09:05`, nothing ran); built controller v0.295.0; 13 red-proofs; live on 9202 (twice — the second after a host crash interrupted the wait) and on demo-felhom |
|
||||
| B — the banner (decision 110) | **done, changed** | rule + real-page tests + red-proofs; served live on 9202 and closed by its real route. **No screenshot: DooPlex has no browser or page renderer** (checked: no chromium/firefox/wkhtmltoimage/playwright); the page HTML as the box served it is the evidence. 9202's ledger was set by hand to a 3-day-old dump to make it show (scratch box, stated). |
|
||||
| C — R-872, R-873, R-874, R-875 | **done; R-872 live pending** | hub v0.134.0 (R-872, R-873), agent v0.145.0 (R-874, R-875); each red-proofed. R-874 live on demo-felhom. R-872's first live 05:00 run: dated check 2026-10-06. R-873: tests only (no live occurrence — Tester 2 stayed off). |
|
||||
| D — R-876 self-repair | **done** | agent v0.145.0; 3 red-proofs; live with the operator's go: crash mid-unpack, the next pass repaired dpkg by itself and finished; bundle signed to all three boxes |
|
||||
| E — release, golden, records | **done** | controller v0.295.0, agent v0.145.0, hub v0.134.0 (one each); golden 0.295.0 baked + vouched (gate OK); floors for demo-hp, demo-felhom, tester-1 |
|
||||
|
||||
## 2. Claims in the brief that turned out wrong (or only partly true)
|
||||
|
||||
1. **"A controller start is the right trigger"** — not alone. A host **resume** is needed too: Go's timers run on
|
||||
CLOCK_MONOTONIC, which does not count suspended time, so on a suspended laptop the 02:30 timer fires hours late — and
|
||||
would start the app-update leg at noon. Built: a late daily timer (> 60 min) is skipped, and a resume watch triggers
|
||||
the catch-up. *Reasoned and unit-tested; NOT measured — no suspend was allowed on a demo box.*
|
||||
2. **"The whole-guest backup cannot collide with the catch-up"** — it CAN: its 48 h safety valve fires on the first
|
||||
5-minute poll after a start, and the catch-up's database dump needs the apps up. Built: each waits for the other.
|
||||
3. **"The box keeps a record of when it was on"** — TRUE, and it was not designed as one: the controller's system-metrics
|
||||
table (one sample a minute, kept 30 days). Used for "off at 02:30" and the suggestion.
|
||||
4. **"The repair step can check the journal without losing R-845's speed"** — TRUE: `--audit` and the journal are read
|
||||
in ONE `sh -c` call; a clean pass still costs one call (pinned by a test).
|
||||
5. **"Apps keep running during a catch-up"** (my morning STATUS said "to be measured") — the dump leg stops an app with a
|
||||
volume for its copy: **measured 1 s** for opengist in the day (R-878).
|
||||
6. R-873's brief offered "only when the box was off for more than 24 h" — rejected (a real outage would reach the
|
||||
household a day late); the weekly rule was chosen (decision 116).
|
||||
|
||||
## 3. What was proven, with numbers
|
||||
|
||||
- **Part A (live).** 9202: W 09:35, off 07:33–07:38 UTC, start 07:38:09 → `missed [db-dump] … ONE catch-up in 15m0s` →
|
||||
07:53:09 dump, done in 21 s. Host crash at 07:56 interrupted the next wait → new start 07:58:35 → ONE catch-up of all
|
||||
three legs at 08:13:35, done in 22 s. demo-felhom: W 10:07, controller parked + stopped 08:05–08:10 (apps running) →
|
||||
catch-up 08:25:02, dump in 1 s; hub event `backup_catchup_done` "Kimaradt mentés pótolva: a doboz ki volt kapcsolva
|
||||
10:07-kor, a mentés most elkészült." — no mail. Window set back to 02:30.
|
||||
- **Part B.** `nightchain.ComputeBanner` tests (shown / fresh / upgrade day / dismissed then back after a new miss / gone
|
||||
after a success / off-site counts / no pattern / usually on); web tests through ServeHTTP and the real dismiss route.
|
||||
- **Part C.** R-872 Tester 2 shape → both alarms; a dump 20 h ago → quiet; a new box → quiet. R-873: night 1 household
|
||||
2 mails, night 2 household 0 / operator 2, after 8 days household again. R-874 live: demo-felhom agent start
|
||||
07:38:46 → first check 08:08:46 → restore-test passed in 29 s.
|
||||
- **Part D (live, operator's go).** 13 packages rolled back; crash 07:56:03.308 UTC during
|
||||
`dpkg --force-confold … --unpack`; new boot 07:56:41 (38 s), guard armed, 1 unclean boot in window; at boot
|
||||
`--audit` clean, journal 1 file, 1 package new / 12 old; next pass (nobody touched dpkg): `REPAIR configured=0
|
||||
journal=1` → `PLAN upgrade=12` → `DONE rc=0 upgraded=12`, healthy; package list identical (279 lines). No mail.
|
||||
|
||||
## 4. Rows
|
||||
|
||||
Register before **340**, after **336**. Closed (6): R-871, R-873, R-874, R-875, R-876, R-877. Narrowed: R-872 (fixed,
|
||||
dated check 2026-10-06). Opened (2): **R-877** (the Tester 1 VM had no start-on-boot; the morning crash left it off
|
||||
1 h 17 min — filed and closed), **R-878** (P4, a daytime catch-up stops a volume app for its dump). STATUS updated.
|
||||
|
||||
## 5. Slips of mine, said plainly
|
||||
|
||||
- **This morning's report said every box was healthy; the Tester 1 VM had been off since the 06:14 crash** (R-877). Found
|
||||
at 07:31 from its stale report time.
|
||||
- **The hub image 0.134.0 was first built from a commit that was not yet pushed** (my unstaged doc edits blocked the
|
||||
`git pull`; the push was then refused by the golden gate). I pushed after the golden bake and REBUILT the image from
|
||||
the pushed commit `1b0678fa`; the deployed image is the rebuilt one (`sha256:f823b10e…`).
|
||||
- Two of my red-proofs did not convict at first (a masked mutation, a test that matched another call); both tests were
|
||||
strengthened and the red-proofs re-run — recorded in the red-proof files.
|
||||
|
||||
## 6. Teardown, three layers
|
||||
|
||||
- **Machines:** 9202: controller 0.295.0 (by hand, allowed there), window back to 02:30, its ledger was set by hand for
|
||||
the banner and has since been overwritten by real catch-up runs. demo-felhom: window 02:30 again, controller unparked.
|
||||
demo-hp: every package current, list identical. Tester 1: running, start-on-boot set. Bake VM: CT 9100 destroyed,
|
||||
token shredded, `virgin`, qemu gone.
|
||||
- **Hosts:** the park file on felhom-pve removed; nothing provisioned.
|
||||
- **Hub:** floors 0.295.0 (MinAgent 0.131.0) for demo-hp, demo-felhom, tester-1; artifacts vouched (agent 0.145.0,
|
||||
golden 0.295.0); 6 signed jobs (agent_update and agent_config_update for each of the three boxes); hub 0.134.0 deployed
|
||||
(ArgoCD Synced/Healthy). Tester 2: read only, offline all session, nothing sent.
|
||||
@@ -0,0 +1,129 @@
|
||||
# REPORT — the hub database off DooPlex (R-173 option A) and the torn-backup live test (R-519) — 2026-10-05 (evening)
|
||||
|
||||
| Part | Result | Releases / changes |
|
||||
|---|---|---|
|
||||
| **A** — PVC label + nightly `VACUUM INTO` snapshot | **done** — label `enabled`, volume 2 Gi, snapshot 353 MiB in 44 s, keep 2, mutex; 6 red-proofs | hub **v0.136.0** (one release) |
|
||||
| **B** — ep0: namespace, two tokens, prune job | **done** — `operator` ns, `dooplex-hub@pbs` with `!push` (DatastoreBackup) and `!restore` (DatastoreReader), `prune-operator-hubdb` (ns `operator` only); household jobs unchanged | ep0 config only |
|
||||
| **C** — DooPlex push / restore-test units, key, alarm | **done** — versioned scripts + units, 15 tests, 9 red-proofs; `enc.key` made, operator saved the paper key; two alarm rules, `promtool`-proven, 2 red-proofs | `scripts/hub-db-backup/`, homelab-manifests rules |
|
||||
| **D** — live proof | **done** — push, ep0 listing, restore test, token limits, paper-key restore, alarm fired + mailed + cleared; R-519 live on 9202; runbook §3 tested (steps 1–3) | `hub/cmd/hubdb-check` (tool, not in the image) |
|
||||
|
||||
**Side events (all with the operator's word):** the Longhorn instance-manager on DooPlex was restarted (77/77 volumes
|
||||
healthy in 110 s); my failed offline-grow attempt kept the hub down ~9.5 min (12:53–13:02Z); zipline came back on a
|
||||
newer release and was pinned to 4.7.0.
|
||||
|
||||
## Wrong claims (in the brief and the runbook)
|
||||
|
||||
1. **Runbook §1: "the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any
|
||||
sync."** The Volume kept `enabled` through every sync since February while the PVC said `disabled` — Longhorn did not
|
||||
copy it down (`partA/step1-labels.txt`, before-state). Which one Longhorn reads was not measured; both say `enabled` now.
|
||||
2. **Brief + runbook: "keep 2 snapshots" on the 1 Gi volume.** Two 353 MiB snapshots plus the 370 MB live database did
|
||||
not fit (594 MB free). The operator chose to grow to 2 Gi.
|
||||
3. **Runbook Step 3 implied the snapshot is quick (0.63 s measured on a scratch copy).** On the live Longhorn volume: 44 s.
|
||||
4. **Runbook Step 2: `--schedule 'daily 03:45'`** — not a PBS calendar event (`unable to parse … 'daily'`); `03:45` is.
|
||||
5. **Runbook Step 2: `proxmox-backup-client namespace create … root@pam`** needs a password CC does not have;
|
||||
`proxmox-backup-debug api create …/namespace` works (the CLI then crashes printing the result — the namespace exists).
|
||||
**`user generate-token … --output-format json`** is refused; the API path is `…/access/users/<u>/token/<name>`.
|
||||
6. **Brief: the push token "DatastoreBackup only (no prune, no read)".** No prune: TRUE (`missing Datastore.Modify|
|
||||
Datastore.Prune`). **No read: FALSE** — the push token restored its own copy (`restore complete … 352.676 MiB`): PBS
|
||||
lets a backup's owner read it. Scope: ns `operator` only, and the content is encrypted with a key ep0 never sees.
|
||||
7. **Runbook Step 2: a token's ACL alone grants it** — PBS cuts a token's rights down by its user's, so the user needs
|
||||
both roles on the path too (done; the user has no password).
|
||||
8. **The 1 Gi → 2 Gi growth was assumed routine** (`allowVolumeExpansion=true`). It failed: the instance-manager called a
|
||||
vanished host PID (`nsenter: cannot open /host/proc/196610/ns/mnt`), and offline growth was blocked by the expansion's
|
||||
own attachment ticket (`partA/step1-expansion-failure.txt`, `step1-offline-expansion.txt`).
|
||||
9. **Runbook §3 as proposed: "a working reveal proves the key matches"** — it needs a running hub on the copy. A check
|
||||
that needs no hub now exists (`hubdb-check`), and the procedure says how to rebuild the key from the paper `data` field.
|
||||
10. **My own: "the hub is down ~2 minutes" for the offline grow** — it was ~9.5 minutes, and it did not work.
|
||||
|
||||
## Part A — evidence (`documentation/audits/hub-db-offsite-2026-10-05/partA/`)
|
||||
|
||||
- Architecture read first: `05-hub-architecture.md` §16 (new §16.3), `07-backup-architecture.md`, the runbook.
|
||||
- Hub baseline `cea8502f`-era main at v0.135.0; commits `d4be9f6` (code, docs, manifest), `5d8922e` (image tag 0.136.0).
|
||||
- `internal/store/snapshot.go` (`SnapshotInto` = `VACUUM INTO ?`), `internal/dbsnap` (`.tmp` → rename, 0600, keep 2,
|
||||
`ErrBusy`, `NeedsCatchUp`), `cmd/hub/main.go` (02:00 Budapest + start-up catch-up).
|
||||
- Tests: `dbsnap_test.go` — integrity_check `ok` and equal row counts in EVERY table vs. the live DB with rows still in
|
||||
the WAL (precondition asserted: a plain file copy held 0 of 150 hosts); keep 2; never two at once; a failed write
|
||||
leaves no file; catch-up age. `cmd/hub/r173_wiring_test.go` (AST). Red-proofs R1–R6 convict (`red-proof.txt`,
|
||||
`red-proof-run1.txt`; R2's first run convicted by HANGING — the test now fails in 2 s).
|
||||
- Live: `db snapshot written: hub-20261005T123037Z.db (369807360 bytes, 44.276s)`, mode `-rw-------`.
|
||||
- PVC/Volume labels both `enabled`; PVC capacity 2Gi; `/data` 2,028,392 KB, 37 % used (`step1-im-restart.txt`).
|
||||
- Full hub suite green (`go build/vet/test ./...`).
|
||||
|
||||
## Part B — evidence (`partB/`)
|
||||
|
||||
- Before (`ep0-before.txt`): prune jobs `prune-demo-hp`, `prune-demo-felhom` (ns each, keep-last 2, 03:30); GC
|
||||
`sun 04:30`; users `root@pam`, `felhom@pbs`; no `operator` ns.
|
||||
- After (`ep0-after.txt`): the same two jobs unchanged + `prune-operator-hubdb` (ns `operator`, max-depth 0, 03:45,
|
||||
keep-daily 14, keep-weekly 8); GC unchanged; user `dooplex-hub@pbs`; tokens `!push`, `!restore`; ACLs only on
|
||||
`/datastore/felhom-offsite/operator`.
|
||||
- Token secrets: ep0 → pipe → root 0600 files on DooPlex (36 bytes each), never on a command line or printed. No
|
||||
household namespace was listed or read (the `operator` check read only whether that one name exists).
|
||||
|
||||
## Part C — evidence (`partC/`)
|
||||
|
||||
- Tools present: `sqlite3` 3.46.1, `proxmox-backup-client` 4.2.3, `kubectl` (k3s) as root, textfile dir
|
||||
`/var/lib/node_exporter/textfile_collector` (already scraped), tunnel `felhom-ep0-pbs-tunnel` active on 127.0.0.1:18007.
|
||||
- `scripts/hub-db-backup/` (commit `cea8502f`), installed with `install.sh`; config
|
||||
`/etc/felhom-hub-backup/{env,token-push,token-restore,enc.key}` all root 0600. Timers enabled after the manual runs:
|
||||
next push Tue 02:31, next restore test Sun 2026-10-11 04:31.
|
||||
- `enc.key` created (`kdf none`, fingerprint `b2:19:bf:36:…`); **STOPPED for the paper key; the operator saved the `data`
|
||||
field** before the first push.
|
||||
- Tests: 15 (`test_hub_db_backup.py`, fakes for `kubectl` and `proxmox-backup-client`). Red-proofs P1–P9 all convict
|
||||
(`red-proof.txt`; P4 first did NOT convict — masked by integrity_check — and P5 errored; both tests were strengthened).
|
||||
**Not in CI** (R-885).
|
||||
- Alarm: `homelab-manifests` `ebc14b0`; `promtool check rules` (4 rules) and `promtool test rules` green; red: threshold
|
||||
260 h → FAILED; `absent()` removed → FAILED (the first attempt at that mutation did not apply — recorded)
|
||||
(`alarm-rule-test.txt`, `bf_test.yml`). Only `ConfigMap/prometheus-rules` was synced; `POST /-/reload` → 200; rules
|
||||
listed in `/api/v1/rules`.
|
||||
|
||||
## Part D — evidence (`partD/`)
|
||||
|
||||
- **Push** (`push-1.txt`): unit `Result=success`; `checked: 369807360 bytes, integrity ok, 4 host(s)`; `Encryption key
|
||||
fingerprint: b2:19:bf:36:…`; 352.676 MiB (17.242 MiB compressed) in 7.19 s; success timestamp written.
|
||||
- **ep0 listing with the read-only token:** `host/dooplex-hub/2026-10-05T13:50:18Z 352.677 MiB`.
|
||||
- **Restore test** (`restore-test-1.txt`): `checked: integrity ok, 4 host(s), 4 sealed console password(s), 0 readable`;
|
||||
success timestamp written; no restored file left.
|
||||
- **Token limits + paper key** (`token-limits-and-paperkey.txt`): forget refused for both tokens; restore token cannot
|
||||
back up; neither can list the datastore root; push token CAN restore its own copy (claim 6); a key file rebuilt from
|
||||
the `data` field only → restore `rc=0`, integrity `ok`, 4 hosts; without a key → `missing key`.
|
||||
- **Alarm** (`alarm-drill.txt`): success file hidden 13:51:46Z → pending 13:52:06Z → **firing 14:22Z** → Alertmanager:
|
||||
1 active alert, receiver `email-notifications`, `notifications_total{email}` 5 and `failed_total` 0 → a real push
|
||||
14:23 (2 s) → inactive 14:23:42Z, email counter 6. **Both mails arrived** (operator's inbox screenshot, 2026-10-05):
|
||||
„[FIRING] HubDBBackupStale" 16:22 and „[RESOLVED] HubDBBackupStale" 16:27 CEST.
|
||||
- **R-519 on 9202** (`r519/`): 9202 raised to controller 0.296.0 by hand (`00-upgrade-9202.txt`); bookstack installed;
|
||||
run 1 complete (33 s); run 2 cut by `docker restart felhom-controller` at 14:05:59.67Z, 0.8 s after `Volume dump:
|
||||
bookstack/bookstack_bookstack_config` and before its database volume. After: the run record turned `interrupted`;
|
||||
the controller restarted bookstack itself; **/backups and /backups/apps both carry `data-interrupted-run`** ("A
|
||||
legutóbbi mentés (2026-10-05 16:05) megszakadt …"; negative control 0); **the restore point reads 14:04:44Z** — the
|
||||
run-1 database volume, its oldest part (config 14:05:58, SQL 14:05:54). Run 3 complete → notice gone from both pages,
|
||||
record clear, the torn `.tar.tmp` replaced. Bookstack removed through the product: 0 containers, 0 volumes, no
|
||||
folder, no backups.
|
||||
- **Runbook §3** (`restore-procedure/drill.txt`): restore from ep0 → `hubdb-check` with the seal key from the Secret
|
||||
(file → file) → `hosts=4 console_passwords_opened=4 failed=0`; a random key → `opened=0 failed=4` FAILED. Steps 4–5
|
||||
(into a live PVC) not run. `hubdb-check` red-proof R7 convicts.
|
||||
|
||||
## Rows and register
|
||||
|
||||
- **Closed:** R-519 (live-proven) → `CLOSED-ITEMS.md` in the same commit.
|
||||
- **Narrowed:** R-173 (option A in force; left: runbook §3 steps 4–5), R-232 ((b) and (a) partly, for the hub DB).
|
||||
- **R-231 not touched** (it is about `/opt/backup/scripts/`); the new units are versioned from day one.
|
||||
- **Opened:** R-882 (Longhorn stale host PID), R-883 (8 workloads on a moving tag; zipline pinned), R-884 (`monitoring`
|
||||
Prometheus Deployment OutOfSync), R-885 (scripts' Python tests not in CI), R-886 (Alertmanager cannot write
|
||||
nflog/silences since the restart).
|
||||
- **Register: 332 → 336.** `STATUS.md` updated (Tonight section; the two old "needs you" items marked decided/done).
|
||||
|
||||
## Teardown — three layers
|
||||
|
||||
- **Machines:** 9202 — bookstack removed through the product; its controller stays 0.296.0 (was 0.295.0); helper files
|
||||
in the guest removed. DooPlex — scratch restores shredded; the `hubdb-check` binary shredded; the units, timers,
|
||||
config and keys stay (they are the deliverable).
|
||||
- **Hosts:** ep0 — the new user, tokens, namespace, prune job stay (the deliverable); nothing else changed. demo-hp — none.
|
||||
- **Hub:** none provisioned. The hub runs v0.136.0.
|
||||
- **Secrets:** a scan of 7,132 committed/working files against the 8 real secret values of this session → 0 hits. Scratch
|
||||
copies (demo password, bookstack deploy values, the zipline pre-change dump, rule files, used signed envelopes)
|
||||
shredded.
|
||||
|
||||
## Unproven / say-so
|
||||
|
||||
- `unproven.py --summary`: walked 20, partial 17, built 14, missing 4 — NOT WALKED 35 of 55; no number moved (the hub-DB capability is a new row in `00` §G, not one of the 55 walked claims).
|
||||
- Runbook §3 steps 4–5.
|
||||
@@ -0,0 +1,131 @@
|
||||
# REPORT — the hub's own safety, boxes left behind, honest backup wording, two onboarding rows, and the agent's admin permissions (R-135, R-133, R-173/R-232, R-604, R-530, R-518, R-519, R-861, R-508, R-509) — 2026-10-05, late afternoon
|
||||
|
||||
Brief: "an open-items batch — the hub's own safety (CSRF, the console credential at rest, the hub database in
|
||||
backups), a fleet view that shows boxes left behind, honest backup wording, two stale onboarding rows; and the agent
|
||||
permission fix (R-861) as its own Part". Evidence: `documentation/audits/hub-safety-2026-10-05/part{A..H}/`, the golden
|
||||
`documentation/tests/golden-0.296.0-2026-10-05/`. Architecture read before the claims: `05-hub-architecture.md`,
|
||||
`_hub-review.md`, `04-control-plane-authorization.md`, `03-host-agent.md` §3/§11, `07` §6.1, `08` §6.3,
|
||||
`runbooks/target-selection.md`, `runbooks/secrets.md`, `runbooks/ep0-datastore-copy.md`,
|
||||
`audits/RECON-dooplex-backup-2026-08-06.md`.
|
||||
|
||||
Baselines (re-verified at the start): felhom.eu `9bb45eaaa2`, felhom-agent `61345790ed` (v0.145.0), felhom-controller
|
||||
`477e2548db` (v0.295.0); register 336 rows, highest R-878.
|
||||
|
||||
## 1. The Part table
|
||||
|
||||
| Part | State | Note |
|
||||
|---|---|---|
|
||||
| A — CSRF (R-135) | **done** | hub v0.135.0: no session → Basic credentials + `X-Felhom-Operator` (decision 120); every state-changing route in one table (`partA/route-table.md`, 38 routes + an unknown path); red-proof: the old shape lets 39 of 39 through; live: 403 / pass / 401 |
|
||||
| B — console password at rest (R-133) | **done** | the off-site seal and key reused (decision 121); 4 legacy rows sealed live, 0 left plain; reveal still opens demo-hp's Proxmox; wrong key fails closed; 2 red-proofs. What a DB backup still holds readable → R-879 |
|
||||
| C — hub DB in backups (R-173, R-232) | **done (read only) — decision with you** | it IS backed up, only on DooPlex, by a label drift; no failure alarm; steps in `runbooks/RUNBOOK-hub-db-offsite-backup.md`; decision in STATUS |
|
||||
| D — boxes left behind (R-604, R-530) | **done** | System page "Version floors" + Agent cell (live), `agent_behind` 7 d + `floor_raise_skipped` (tests, 3 red-proofs). The mail was not exercised live (needs a global raise) |
|
||||
| E — honest backup wording (R-518, R-519) | **done, changed** | R-518: the copy was already honest; today's measurement added (5 min 47 s). R-519: dating was already fixed (v0.275.0); the page notice + the synthesised status fixed (v0.296.0, 4 red-proofs). **The live cut on 9202 was refused by the permission check — asked** |
|
||||
| F — the agent's admin permissions (R-861) | **done, changed** | nine root paths, not four; `03` §3.1 written AFTER the build (not "design first"); agent v0.146.1 (after a review found three holes in v0.146.0); delivered by a two-step bundle (R-880); live on both demo boxes: sudo 93/93, capability 67/67; three residuals named, row stays open narrowed |
|
||||
| G — onboarding rows (R-508, R-509) | **done** | R-509 closed by three matched real mails; R-508 closed (e-mail set since 09-14 + a new page warning, red-proof) |
|
||||
| H — release and records | **done, changed** | hub 0.135.0, controller 0.296.0, agent 0.146.1 (+ 0.146.0 never delivered); golden 0.296.0 baked + vouched; floors + signed jobs for demo-hp, demo-felhom, tester-1; docs `00`, `03`, `05`, `07`, `08`, `09`, `11`-runbook |
|
||||
|
||||
## 2. Claims in the brief that turned out wrong
|
||||
|
||||
1. **"The hub's own database is in no backup."** It is in one — only on DooPlex. Longhorn's `backup-daily` /
|
||||
`backup-weekly` copy `hub-data` every night (last 2026-10-05 02:06 UTC, Completed, 713 MB) to DooPlex's own `sda1`.
|
||||
R-173's "excluded" is the PVC label (`recurring-job-group.longhorn.io/default: disabled`, set 2026-02-16 with no
|
||||
reason); the live Longhorn Volume carries `enabled` — a hand-set drift that keeps the backup alive and can be undone
|
||||
by any sync. Nothing leaves DooPlex, and nothing alarms if it fails (R-232 stands).
|
||||
2. **"The backup page promises 'a few seconds'."** Not since controller v0.243.0 / v0.267.0: the text already said
|
||||
"several minutes (about 8 minutes on a 12-app box)". Measured today on demo-hp (9 apps): 5 min 47 s, local tier only.
|
||||
v0.296.0 adds today's figure and "minutes, not seconds".
|
||||
3. **"R-519: fix the dating."** The dating was already fixed in controller v0.275.0 (R-696): a restore point carries the
|
||||
time of its OLDEST part. What was still missing was the sentence on the page, and the page's synthesised "last
|
||||
database backup … OK" after a restart — both fixed in v0.296.0.
|
||||
4. **"R-861: four admin-command groups."** It was nine ways to root, not four: besides the four named (guest hook,
|
||||
intermediary script/unit, escrow, self-update), the mount units, the dnsmasq drop-ins, the WireGuard config and the
|
||||
OOB sshd config were each installed from agent-written files, and almost every `*` in the arguments matched spaces
|
||||
(measured with real sudo 1.9.16: 23 of 29 attack lines allowed).
|
||||
5. **"Deliver the agent fix by the signed bundle."** Not possible in one step: an installed `felhom-os-apply` refuses
|
||||
a bundle naming a path it does not know (R16), and v0.146.1's bundle adds four. Delivered by a step bundle (R-880).
|
||||
6. **"Design first" (Part F).** I built first and wrote the `03` §3.1 section after the code, in the same session — the
|
||||
section records what was built, group by group.
|
||||
|
||||
## 3. Per Part — tests, red-proofs, live proof
|
||||
|
||||
**A.** `hub/internal/web/r135_csrf_test.go` (5 tests). Red-proof `partA/red-proof.txt` (39 of 39 convicted). Live
|
||||
`partA/live.txt` (hub 0.135.0, ClusterIP): Basic, no header, `Origin: evil` → **403**; unknown path, no header → **403**;
|
||||
with `X-Felhom-Operator: cli` → **404** (passed the gate); header without credentials → **401**; GET → 200. The skill and the
|
||||
memory note now carry the header.
|
||||
|
||||
**B.** `store/r133_recovery_seal_test.go` (4), `web/r133_reveal_wrongkey_test.go`, `cmd/hub/r133_wiring_test.go`. Red-proofs
|
||||
`partB/red-proof.txt` (plaintext save — the first attempt did not compile, re-run with a compiling mutation; the wiring).
|
||||
Live `partB/live-db.txt`: hub start `console passwords sealed at rest (4 legacy plaintext row(s) sealed now)`; the live DB
|
||||
copy (scratch, shredded) shows 4 rows `enc:v1:`, 0 not sealed. `partB/live-reveal.txt`: reveal on demo-hp → 200, a
|
||||
32-char password that minted a PVE ticket (200; a wrong one 401); the timeline event recorded.
|
||||
**What a hub DB backup now holds:** the console and off-site passwords sealed (useless without `OFFSITE_SECRET_KEY`, which
|
||||
exists only on DooPlex); still readable: box API keys, owner passphrases + customer API keys, PBS-DR token values (R-879).
|
||||
|
||||
**C.** Readings `partC/readings.txt` (read only). Steps `runbooks/RUNBOOK-hub-db-offsite-backup.md`: keys off the box first;
|
||||
fix the PVC label in git; a write-only namespace on ep0's PBS; a hub `VACUUM INTO` nightly snapshot (a later hub release);
|
||||
the encrypted push via the existing tunnel; a weekly restore test (`PRAGMA integrity_check`, row counts, every console
|
||||
password still sealed); two Prometheus alarms through the existing mail receiver (`absent()` included); a proof run.
|
||||
|
||||
**D.** `osupdates/r530_agent_alarm_test.go` (3), `web/r604_floor_held_back_test.go` (4), `cmd/hub` wiring. 3 red-proofs
|
||||
(`partD/red-proof.txt`). Live `partD/live-system-page.txt`: global floor 0.292.0, three per-customer floors 0.295.0 (age
|
||||
"unknown" — set before v0.135.0); Tester 2 `0.142.0 → 0.145.0 (since 2026-10-05)`, the demo boxes "current".
|
||||
|
||||
**E.** `internal/backup/run_record_test.go` (3), `cmd/controller/run_record_wiring_test.go` (2), `TestR518_*`, parity
|
||||
cases. 4 red-proofs (`partE/red-proof.txt`). Measurement `partE/r518-measure.txt`. The 9202 reproduction: a throwaway
|
||||
bookstack installed (09:35:56Z) and a complete baseline run (09:37, 35 s); the cut was refused by the permission check;
|
||||
bookstack removed through the product (`partE/teardown-9202.txt`: no container, volume, folder or backup left).
|
||||
|
||||
**F.** Design `03` §3.1. `configs/test_felhom_priv_apply.py` (32), `AgentUpdate` (8), `SelfupdateWrapperConfinement`,
|
||||
`StepBundle` (3), Go contract tests (4 packages), `TestSudoersRefusesTheR861Injections`, `TestManifestCoveredBySudoers`.
|
||||
Red-proofs F1–F9 + S1–S3 (`partF/red-proof.txt`; F1 masked on its first run — strengthened; F3 errored rather than
|
||||
failed — clean assertion added). Real sudo, container (`partF/sudo-container-proof.txt`): old 23/29 attacks allowed, new
|
||||
0/29, 64/64 commands allowed. Pre-flight on both boxes' live files: all OK. **Live after the bundle:** `sudo -l` 93/93 on
|
||||
demo-hp and demo-felhom (`partF/live-sudo-after-*.txt`; before, on demo-felhom: 23 attacks allowed —
|
||||
`live-sudo-before-demo-felhom.txt`); the checker run as the agent user → SAME on every real file (on demo-felhom the drive unit has no staged copy — an older path wrote it — so that one read `[P1] no staged file`; that box's `/mnt/hdd_1` is the whole-system backup storage, not a household drive, so no bind under `/mnt/felhom-drives` is expected); a staged unit over
|
||||
`/etc/sudoers.d` refused `[U3]`, nothing installed; the old `install` route → `a password is required`.
|
||||
**Capability check after the bundle: demo-hp 67/67, demo-felhom 67/67, Tester 1 all ok (hub page), nothing degraded.**
|
||||
A gap seen on demo-felhom: between the new agent (~10:45 UTC) and the bundle (11:09) its OOB-sshd reconcile logged "install
|
||||
failed" every minute (the expected gap); 0 errors after the bundle.
|
||||
|
||||
**G.** `partG/r509-real-mails.txt` (three host-delete sends matched to mailbox arrivals within 1 s), `web/r508_no_email_banner_test.go`
|
||||
(3 branches) + red-proof.
|
||||
|
||||
## 4. Release and delivery
|
||||
|
||||
- **hub v0.135.0** (`3d7a2761` code, `2b30733b` manifest) — built from the pushed commit, ArgoCD Synced/Healthy, image tag
|
||||
0.135.0. (A later comment-only change in `server.go` is in this session's docs commit; the image is unchanged by it.)
|
||||
- **controller v0.296.0** (`ff69074`), MinAgent 0.131.0. Floors 0.296.0 for demo-hp, demo-felhom, tester-1 → both demo
|
||||
boxes ran 0.296.0 (healthy) within ~30 min; 9202 (scratch) stays 0.295.0.
|
||||
- **agent v0.146.0** (`6ab1e7c`, released, **never vouched or delivered**) → **v0.146.1** (`fdd8717`) after the review.
|
||||
Step bundle `0.146.1-step1` (sha `8482851e…`, built from the 0.145.0 bundle, only `felhom-os-apply` replaced).
|
||||
- **golden 0.296.0** baked and vouched with agent 0.146.1 / min_agent 0.131.0 (`documentation/tests/golden-0.296.0-2026-10-05/`).
|
||||
- **Per box, signed with felhom-op-1:** `agent_update` 0.146.1 → `agent_config_update` 0.146.1-step1 (`written=1 same=20`,
|
||||
self-check ok) → `agent_config_update` 0.146.1 (`written=3 same=22`, self-check ok). demo-hp, demo-felhom, Tester 1 all
|
||||
report agent 0.146.1 and root files 0.146.1 (`partH/fleet-after.txt`). **Tester 2: DOWN all session, nothing sent.**
|
||||
|
||||
## 5. Rows
|
||||
|
||||
Register **336 → 332**. Closed (7): R-133, R-135, R-508, R-509, R-530, R-604, and R-880 (opened and closed today).
|
||||
Narrowed: R-861 (three residuals), R-173 (measured; waiting on you), R-518 (copy; per-tier quiesce left), R-519 (live cut
|
||||
left). Opened (2): R-879 (hub.db still holds readable secrets), R-881 (installer uninstall misses `felhom-priv-apply`).
|
||||
The section counts in `OPEN-ITEMS.md` were recomputed (several were already out of date).
|
||||
|
||||
## 6. Slips of mine, said plainly
|
||||
|
||||
- **Two agent releases** (0.146.0, 0.146.1) against "one per repo". 0.146.0 had three security holes a background review
|
||||
found after I pushed it; it was never vouched or sent.
|
||||
- **I did not design Part F first** as the brief asked; the `03` section was written after the build.
|
||||
- **My first version waiter read the wrong page cell** and reported the boxes as not updated; I re-read the right cell.
|
||||
- **Two red-proofs did not convict on the first run** (F1 masked, the R-133 plaintext mutation did not compile); both
|
||||
were fixed and re-run.
|
||||
|
||||
## 7. Teardown, three layers
|
||||
|
||||
- **Machines:** 9202 — the throwaway bookstack removed through the product, nothing left; its controller stays 0.295.0.
|
||||
The demo boxes keep their real files (the checker reported SAME; the one staged attack file was deleted).
|
||||
Bake VM: CT 9100 destroyed, token/script/log shredded, qemu stopped, disk back to `virgin`.
|
||||
- **Hosts:** nothing provisioned. The pre-flight copies of the checker (`/tmp/felhom-priv-apply-check`) and the case
|
||||
files were removed from both hosts.
|
||||
- **Hub:** hub 0.135.0 deployed; artifacts vouched (agent 0.146.1, golden 0.296.0); floors 0.296.0 for three customers;
|
||||
9 signed jobs (3 × agent_update, 6 × agent_config_update), all consumed. The step package `felhom-agent/0.146.1-step1`
|
||||
stays published on purpose (Tester 2 will need it). No customer or appliance record created.
|
||||
@@ -74,6 +74,7 @@
|
||||
| `offsitekeys.Registrar` (`Install` / `Confirm` / `Audit` / `OpenWindow` / `CloseWindow` / `MoveAside`) | hub/internal/offsitekeys/offsitekeys.go | `(ctx, Target, password, …)` | EVERY write to a sub-account's `.ssh/authorized_keys` and every repo move-aside | **The only writer of that file, and the only deleter on a sub-account (`DeleteSetAside`: `<repo>.orphaned-*` only, decision 74).** `read()` is read-only (R-827). Uses the provider's port-23 restricted shell (`dd of=` takes stdin, `mv` overwrites, `test` does NOT exist — measured); an unpinned line is a deletion route and is dropped on every install; the window line goes FIRST (first match wins). Never `rm`. |
|
||||
| `offsitekeys.Service` (`RegisterKey`, `ConfirmKey`, `AuditAll`, `OpenWindowFor`, `CloseWindowFor`, `SweepExpiredWindows`) | hub/internal/offsitekeys/service.go | — | Binding the registrar to the store, descriptor and operator events | The box-facing API (`/api/v1/offsite/register-key…`) answers with NO credential — pinned by `TestOffsiteKeyEndpoints_AuthAndNoPasswordInAnyResponse`. |
|
||||
| `(*Store).SaveOneTimeSecret` / `OffsitePassword` / `SealLegacyOffsiteSecrets` | hub/internal/store/offsite_seal.go | — | Storing / reading the sub-account password | **Sealed AES-256-GCM; no key → refused (fail-closed).** Under `go test` every store gets a fixed key (`testing.Testing()`); production needs `OFFSITE_SECRET_KEY`. Never serve the value to a box. |
|
||||
| `(*Store).sealAtRest` / `openAtRest` / `sealAPIKey` / `apiKeyHash` / `SealLegacyBoxSecrets` | hub/internal/store/r879_box_seal.go | — | Any box-facing secret column (`hosts.api_key`, `customer_configs.api_key` / `retrieval_password`, `host_pbs_secrets.value`) | **R-879: same seal; a looked-up key is matched on `api_key_hash`, never opened** — box auth needs no sealing key. A sealed value that does not open sets `SecretsUnreadable` (field ""): serve/compare paths 500 on it and `SaveCustomerConfig` / `UpsertHost` refuse it (a load-modify-save would blank the secret). A new writer of an API key must write the hash in the same statement (`sealAPIKey`). Pinned by `store/r879_box_seal_test.go`, `api/r879_box_seal_test.go`. |
|
||||
|
||||
### Host views & lifecycle / offsite endpoints (v0.47.0, hub/internal/web + store)
|
||||
|
||||
@@ -148,6 +149,7 @@
|
||||
| `store.GuestID` | hub/internal/store/store.go (~L1268) | `(hostID string, vmid int) string` | Canonical guest primary key | Never hand-concatenate host+vmid. |
|
||||
| `(*Store).GetHostReportsSince` + `GetFirstHostReportAt` + `monitor.newestBackupEvidence` | hub/internal/store/store.go, hub/internal/monitor/deadline.go | `(customerID, since) ([]HostReportRow, error)`; `(customerID) (time.Time, error)`; `(rows, now) (time.Time, bool)` | **Asking "when did the hub last SEE evidence of X?" instead of "what does the latest report say?"** — the R-81 anchor. The agent's reporters are point-in-time and forget across a restart; the hub retains ~90 d of host-reports and does not. | The three go together: window scan + first-contact anchor + a bounded lookback (`backupEvidenceLookback`). **Never judge a report-derived absence on the LATEST report alone** — that is the bug class R-81 fixed for the third time. The scan early-exits on sufficiently-fresh evidence, so don't reorder rows away from newest-first. |
|
||||
| `scheduleDaily` | hub/cmd/hub/main.go (~L449) | `(ctx, name, "HH:MM", fn, logger)` | Daily jobs in Europe/Budapest (prune etc.) | Blocking — run as goroutine. `parseHM` returns 0,0 (midnight) on bad input. |
|
||||
| `hu_grep.py` | scripts/hu_grep.py | `PATTERN --anchor ASCII [--negative TEXT] PATH…` | ANY search for Hungarian (accented) text in files — prints file:line hits, and refuses to report a zero unless an ASCII anchor hits and a negative control misses (R-364) | Reads bytes in Python, never via a shell; refuses a pattern that arrived transformed (octal escapes, U+FFFD). Exit 0 found · 1 tested zero · 2 refused |
|
||||
|
||||
## 2. Canonical patterns (copy structure from THE named file)
|
||||
|
||||
@@ -197,6 +199,7 @@
|
||||
| `api.ConfigTemplateProvider` | hub/internal/api/handler.go (~L24) | `web.TemplateFetcher` (Gitea-pulled controller.yaml template) | stub providers in api tests |
|
||||
| `api.LatestVersionProvider` | hub/internal/api/handler.go (~L31) | `web.VersionChecker` (registry poll) | hub/internal/api/config_version_ack_test.go |
|
||||
| `mailRateLimiter.now` (func seam) | hub/internal/api/mail.go (~L27) | `time.Now` | hub/internal/api/mail_test.go clock injection |
|
||||
| Installer behaviour tests (function lift + recording PATH stubs) | scripts/test_hostinstall.py | functions copied word for word out of `scripts/felhom-host-install.sh` | the same file — use it for any new installer function (BusyBox-safe; the `script-tests` gate runs it) |
|
||||
| `web.tenancyProvisioner` | hub/internal/web/pbsdr.go | `*tenantsync.Client` (pinned SSH to ep0's felhom-tenantsync) | `fakeTenancy` in hub/internal/web/pbsdr_test.go; in-process SSH server in hub/internal/tenantsync/client_test.go |
|
||||
| Cross-repo: ep0 tenancy surface | `scripts/felhom-tenantsync.sh` (JSON stdin/stdout forced command) | installed on ep0 per runbook offsite-endpoint.md §10 | provision/reissue/fingerprint ops; token secret rides stdout ONLY; the peersync script/key are untouched |
|
||||
| Cross-repo: controller → hub | `POST /api/v1/report` (frozen) + `POST /api/v1/event` | felhom-controller repo | new event types MUST enter `allowedEventTypes` (hub/internal/api/handler.go ~L1063) or the controller gets 400 |
|
||||
|
||||
@@ -1,13 +1,191 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Ready for the first real tester (Tester-2): yes. Tester 2 is a laptop that is switched off at night (your word,
|
||||
2026-10-05) — it was offline all session; nothing was sent to it.**
|
||||
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was
|
||||
sent to it.**
|
||||
|
||||
**Updated 2026-10-05 (day, the night's fixes): every box of ours healthy. Fixed and proven live: the off-site clean-up
|
||||
now deletes old copies (both demo boxes), a new box's first app install, the update's disk-space check, a killed
|
||||
update's lost report. The power cut in the middle of an update was tested on demo-hp with your go: the box came back by
|
||||
itself in 37 s, but the next update failed until I ran one command by hand — filed (R-876), fix next session.
|
||||
One decision for you below (a box that is off at night). Report: `REPORT-night-fixes-2026-10-05.md`.**
|
||||
**Updated 2026-10-05 (late night, burn-down round 2): every box of ours healthy. The open-items list is at 199 (was 292
|
||||
at the start of this round, 336 this morning). Report: `REPORT-burndown2-2026-10-05.md`.**
|
||||
|
||||
## Tonight, last (2026-10-05): the list at 199
|
||||
|
||||
**What happened:**
|
||||
- **Your answer is recorded** (below): 43 rows closed as accepted.
|
||||
- **51 more rows fixed and closed**, with one release per repository, delivered the normal way: hub **0.137.0**
|
||||
(live), agent **0.147.0** (on demo-hp, demo-felhom and Tester 1, root files too), controller **0.297.0** (on all
|
||||
three), new-install image **0.297.0** (baked, checked, approved), app catalog updated.
|
||||
- **R-124 is fixed** (your ruling): the recovery recipe now writes the backup-server namespace the way the server reads it.
|
||||
- **The CI fault (R-887) is understood:** when Gitea is busy, the CI runner's request for work can time out after Gitea
|
||||
already gave it the job; the job is then never run and is failed 10–13 minutes later. Gitea was busy because of an
|
||||
outside web crawler and because of MY CI checks, which asked too much — mine now ask once a minute in one small request.
|
||||
- **Two of my own CI misses tonight:** a new check needed Go, which the CI machine does not have (fixed); a red run was
|
||||
the fault above (re-run passed).
|
||||
|
||||
**The numbers:** 292 before → **199 after**; 1 opened; 94 closed.
|
||||
|
||||
**Needs you (none urgent; if you do nothing, each stays open as it is):**
|
||||
1. **R-469** — a one-paragraph rewording in the app catalog's instruction file; my permission check refused editing
|
||||
instruction files. Say "go" and it is done in a minute.
|
||||
2. **R-126** — should a network share be offered as an export destination? Two options in the row; pick one.
|
||||
3. **R-888** — two facts the boxes report that the hub never shows (installed-app list, retired drives). Needed or not?
|
||||
4. **R-887** — CI: raise the runner's fetch timeout and/or slow the crawler on Gitea's public pages. If nothing: now and
|
||||
then a CI run fails without running; it can be re-run.
|
||||
5. Three more rows look „not worth doing" (R-337, R-375, R-542 — the check found each is by design or harmless). Close
|
||||
them as accepted? If you say nothing they stay.
|
||||
|
||||
## Your answer to the burn-down list (2026-10-05 18:23), recorded
|
||||
|
||||
- **43 rows closed as accepted by you**, each with its reason from the list.
|
||||
- **R-124 stays open and is fixed in this session** (a recovery step that fails during a real recovery).
|
||||
- **R-698 stays open as a known risk** (a backup keeps an app's image name, not the image). It is yours; nobody works on it now.
|
||||
- **R-831 and R-870** (the printed tokens) stay as they are, by your earlier rulings; the rows keep their steps.
|
||||
- **R-887** (the CI jobs that never ran): your screenshot shows ONE runner, online. My „old copy of the runner" guess was
|
||||
wrong; the row says so.
|
||||
|
||||
## Tonight, later (2026-10-05, night): the list got shorter
|
||||
|
||||
**Decisions:** none of mine.
|
||||
|
||||
**What happened:**
|
||||
- **Every lower-priority row was checked against today's code** (317 rows). 24 described problems a later change had
|
||||
already fixed; 2 were duplicates. Those are closed, each with the change that fixed it.
|
||||
- **19 small rows were fixed and closed** in four repositories — wrong comments and documents, missing tests, gates that
|
||||
checked less than they claimed. No release was needed: nothing that runs on a box or on the hub changed.
|
||||
- **One real fix found on the way:** the hub's build script pushed a `latest` image tag on every release. Nothing used
|
||||
it; it is gone, and a test keeps it gone.
|
||||
- **You got 5 „CI failed" mails tonight — my fault, fixed.** The new test gate ran one of today's test suites on the CI
|
||||
machine for the first time; it uses a different `date` tool, and 9 tests failed. Fixed; CI is green again (job 1360).
|
||||
- **A separate CI fault on DooPlex (new row R-887):** two CI jobs were never run and were marked failed after ~10 minutes,
|
||||
with no log and no mail. (The „old runner copy" guess was wrong — see above.)
|
||||
- **A new rule, so the list stops growing:** a small problem found during work (about 30 minutes) is fixed in that
|
||||
session and never added to the list. Every report now states four numbers: rows before, after, opened, closed.
|
||||
|
||||
**The numbers:** 336 before → **292 after**; 1 opened (the CI fault below); 45 closed.
|
||||
|
||||
## Tonight (2026-10-05, evening): the hub database off DooPlex; the cut-backup check on the scratch box
|
||||
|
||||
**Decisions:** none of mine. Yours (`09` 125–127): ep0 for the copy; the scratch-box restart allowed; the agent's three
|
||||
by-design abilities stay. Today you also chose: grow the hub's disk to 2 GiB, and restart Longhorn's disk manager.
|
||||
|
||||
**What works now (proven live):**
|
||||
- **The hub makes a clean copy of its database every night at 02:00** (hub 0.136.0). The first one: 353 MB, 44 s.
|
||||
- **DooPlex checks it, locks it with a key ep0 never sees, and sends it to ep0 at 02:30.** First send: 7 s.
|
||||
- **Every Sunday at 04:30 DooPlex takes the copy back from ep0 and checks it**: it opens, it is whole, every console
|
||||
password in it is still locked. Done once by hand today: 4 boxes, 4 locked passwords, 0 readable.
|
||||
- **The two ep0 accounts can only do their one job:** the sending one cannot delete, the checking one cannot write,
|
||||
neither can see the households' backups. (The sending one can read back its own locked copies — that is how ep0 works.)
|
||||
- **Your two saved keys work:** a key rebuilt from the paper copy you saved opened the copy, and your saved lock key
|
||||
opened all 4 console passwords in it (a wrong key opened none).
|
||||
- **The alarm:** proven by Prometheus' own rule test; the real alarm mail is below.
|
||||
- **A backup cut off by a restart is now said on the backup pages** (scratch box): the page said so, the restore point
|
||||
kept the older time of the part that was not redone, and the next full backup cleared the message.
|
||||
|
||||
**What broke, and what I did:**
|
||||
- **The hub's disk would not grow:** Longhorn's disk manager on DooPlex was stuck. My first try (stopping the hub so the
|
||||
disk could grow offline) did not work and kept the hub **down about 9.5 minutes**. You approved restarting the disk
|
||||
manager: all 77 disks were back in under 2 minutes, and the hub's disk is 2 GiB now.
|
||||
- **Zipline did not come back after that restart:** it is set to "always the newest", so it pulled a new release that
|
||||
refused its database. I pinned it to the previous release; it runs and its database is updated. 7 more apps on
|
||||
DooPlex use "always the newest" (new row).
|
||||
|
||||
- **The alarm mail:** I hid the "copy sent" signal on purpose; after 30 minutes the alarm fired (14:22) and the mail
|
||||
system sent it without error; a real send cleared it a minute later. **You confirmed both mails arrived:**
|
||||
"[FIRING] HubDBBackupStale" 16:22 and "[RESOLVED] HubDBBackupStale" 16:27.
|
||||
- **The mail system (Alertmanager) cannot save its own notes since the Longhorn restart.** Mail still goes out; but a
|
||||
silence you set would be lost at its next restart (new row).
|
||||
|
||||
**Register:** 332 → 336 rows (1 closed: the cut-backup check; 5 opened: the Longhorn fault, the "always newest" apps, a
|
||||
monitoring sync drift, script tests not in CI, the mail system's notes).
|
||||
|
||||
**Needs you:**
|
||||
1. **Nothing urgent.** If you do nothing, the copy runs every night and you get a mail only if it stops.
|
||||
2. **When convenient:** pin the 7 other DooPlex apps that use "always the newest" (or tell me to list them for you). If
|
||||
you do nothing, any restart may upgrade one of them by surprise, as it did zipline.
|
||||
3. **Tester 2's one-time step** is unchanged (below).
|
||||
|
||||
## Today (2026-10-05, late afternoon): the hub's own safety; boxes left behind; the agent's admin rights
|
||||
|
||||
**Decisions I took myself (you may reverse each — `09` decisions 119–124):**
|
||||
- A box behind the approved agent for 7 days raises an alarm to you; a global version raise that cannot move a box
|
||||
sends you one mail naming it.
|
||||
- Scripts that post to the hub with the password must now send one extra header; a web page on another site cannot.
|
||||
- The console passwords are locked with the same key as the off-site passwords (one key to keep safe, not two).
|
||||
- The agent's rights are narrowed with exact rules and one checking helper, not one helper per command.
|
||||
- A cut-off backup is shown on the backup page until a backup runs all the way through.
|
||||
- A new agent whose root files add a file is delivered in two signed steps (the old box would refuse it in one).
|
||||
|
||||
**What works now (proven live):**
|
||||
- **Form protection:** a password post without the header is refused (403); a browser on another site cannot add it.
|
||||
- **Console passwords locked:** all 4 were sealed at the hub's start; the demo-hp one still opens its Proxmox (checked).
|
||||
- **Boxes left behind:** the System page lists the three per-box version floors and shows Tester 2's agent 4 releases
|
||||
behind. The alarm and the mail are proven by tests only (they need 7 days / a global raise).
|
||||
- **The agent cannot make itself root any more:** before, the real sudo let 23 of 29 attack commands through; now 0, on
|
||||
demo-hp and demo-felhom, and every agent feature still passes its check (67 of 67) on all three boxes.
|
||||
- Agent 0.146.1, controller 0.296.0 and hub 0.135.0 on demo-hp, demo-felhom and Tester 1; new-install image 0.296.0.
|
||||
|
||||
**Found today:**
|
||||
- **The hub database is backed up — but only inside DooPlex**, and only because a hand-set label says so; nothing tells
|
||||
anyone if that backup fails. (Your decision below.)
|
||||
- **A new agent's root files could not reach any box in one step** (an older box refuses files it does not know). Fixed
|
||||
with a two-step delivery; written down for next time.
|
||||
- **A security review of my own agent change found three holes** before it went to any box; fixed in a second agent
|
||||
release (0.146.1). Two agent releases today, against "one per repo" — the first was never sent anywhere.
|
||||
- **The backup page already said "about 8 minutes"**, not "a few seconds". Measured today on demo-hp (9 apps): about 6
|
||||
minutes. Both figures are on the page now.
|
||||
|
||||
**Needs you:**
|
||||
1. **(DECIDED 2026-10-05 evening: A, done — see Tonight)** **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings):
|
||||
- **A — my pick: ep0's backup server**, encrypted on DooPlex before it leaves, with a weekly restore test and an
|
||||
alarm mail. Costs one small change on ep0 (a write-only account) and keeping two keys in your password manager.
|
||||
- **B: a separate Hetzner Storage Box account** with restic. More new parts to look after than A.
|
||||
- **If you decide nothing:** the database stays only on DooPlex. A fire or theft there loses every box's console
|
||||
password, the escrow records and the customer settings; each box would need re-pairing by hand. Steps:
|
||||
`documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md`.
|
||||
- **Either way, first:** put the hub's lock key (`OFFSITE_SECRET_KEY`) in your password manager — without it a copy
|
||||
of the database cannot open the console passwords.
|
||||
2. **(DONE 2026-10-05 evening — see Tonight)** **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the
|
||||
controller in the middle of a backup. Say "go" and the next session does it once on 9202; if not, the fix stays
|
||||
proven by tests only.
|
||||
3. **Three things the agent can still do, by design** (each written in `03` §3.1): pick which controller image its own
|
||||
guest runs; install the operator SSH key for the limited `felhom-op` user; see the box's backup key during the
|
||||
recovery-code ceremony. If you do nothing, they stay as they are until before the first paying customer.
|
||||
4. **Tester 2's one-time step** is unchanged (below).
|
||||
|
||||
## Earlier today (2026-10-05, afternoon): a box that is not always on; the self-repair after a power cut
|
||||
|
||||
**Decisions I took myself (you may reverse each — `09` decisions 112–118):**
|
||||
- The make-up run starts 15 minutes after the box comes back; a backup due within 30 minutes is left to its normal time.
|
||||
- A laptop that sleeps through the night: a nightly job that wakes up more than an hour late is skipped (otherwise app
|
||||
updates would start at noon); the backups are made up instead. Tested, not measured (I may not suspend a box).
|
||||
- The banner appears when the last backup is over 26 hours old; it suggests the latest evening hour the box is usually
|
||||
on (5 of the last 7 days), and nothing when the box is usually on at its backup time.
|
||||
- A box that is off at the 05:00 check now raises the missed-backup alarm after 2 nights without a database backup (3
|
||||
without a whole-box backup) — not after one, so a box that broke last night gives only its "offline" alarm.
|
||||
- The household hears "your server cannot be reached" at most once a week; you still hear every one.
|
||||
- The restore-test's first check is 30 minutes after the agent starts (a box on for short times now gets tested).
|
||||
- The update's repair step now also looks at dpkg's journal — the place the power cut left its mark.
|
||||
|
||||
**What works now (proven live):**
|
||||
- **A missed night is made up once** (your choice A): the scratch box and demo-felhom were off across their backup time;
|
||||
15 minutes after they came back, the missed backups ran by themselves (seconds). The household's timeline got one line:
|
||||
"Kimaradt mentés pótolva: a doboz ki volt kapcsolva 10:07-kor, a mentés most elkészült." No mail.
|
||||
- **The banner** (your idea) appeared on the scratch box, with the "change the backup time" button; "Close" kept it
|
||||
closed. *No screenshot: there is no browser on DooPlex; I captured the page as the box served it.*
|
||||
- **The power cut, again** (demo-hp, your go): back by itself in 38 seconds; **the next update repaired dpkg by itself
|
||||
and finished — nobody typed anything.** No mail.
|
||||
- **A restore-test 30 minutes after an agent start** ran and passed on demo-felhom.
|
||||
- Controller 0.295.0, agent 0.145.0 (+ its root files) on demo-hp, demo-felhom and Tester 1; hub 0.134.0; new-install
|
||||
image 0.295.0 baked and approved.
|
||||
|
||||
**Found today:**
|
||||
- **My slip from this morning:** the Tester 1 test machine did not restart after the morning crash and stayed off for
|
||||
1 h 17 min; my morning report said every box was healthy. It now starts by itself after a crash (proven by the second
|
||||
crash).
|
||||
- **What the household may notice from a make-up run:** the database backup stops an app with stored files for its copy —
|
||||
1 second for opengist. The night does the same unseen; a big app may take longer, in the day (filed, small).
|
||||
|
||||
**Needs you:** nothing urgent. **Tester 2's one-time step** is unchanged (below). The new missed-backup alarm will be
|
||||
checked at tomorrow's 05:00 run; if Tester 2 is still off, you will get its first real "backup missed" mail — that is
|
||||
the fix working, not a new fault.
|
||||
|
||||
## Today (2026-10-05, day): the night's fixes, the power cut by day, a box that is off at night
|
||||
|
||||
@@ -38,7 +216,7 @@ One decision for you below (a box that is off at night). Report: `REPORT-night-f
|
||||
|
||||
**Needs you:**
|
||||
|
||||
1. **A box that is off every night (Tester 2) — what does the product promise?** Today such a box never gets its
|
||||
1. **[DECIDED 2026-10-05 08:42 — option A, built the same afternoon; see above]** **A box that is off every night (Tester 2) — what does the product promise?** Today such a box never gets its
|
||||
nightly database backups, second copy or off-site copy; the whole-box backup runs only about every 2 days; no
|
||||
alarm says so; the household is mailed "your server cannot be reached" every night. No design document covers it
|
||||
(R-871).
|
||||
|
||||
@@ -257,7 +257,9 @@ Then: [exact refusal — HTTP status, error, and the proven non-effect, e.g. "m
|
||||
typecheck; do not accumulate compile errors.
|
||||
2. **Minimal changes:** build only what's listed. No "while I'm here" refactors. Note anything worth
|
||||
fixing under "Observations" (§15) — don't act on it, but **do file it**: §15.9's marker rule means
|
||||
"not acted on" never means "not recorded".
|
||||
"not acted on" never means "not recorded". **Exception — the size rule (2026-10-05):** a finding that is cosmetic
|
||||
or small (about 30 minutes, in a repo this session may change) is FIXED in the session with a test and listed under
|
||||
„fixed without a row" — not filed (`OPEN-ITEMS.md` „How a row is filed").
|
||||
3. **No silent failures:** never swallow a parse/exec error — log it. Check a subprocess's **own** exit
|
||||
code; never pipe in a way that hides a 127. (The silent `.felhom.yml` quoting bug + the spike's
|
||||
exit-swallow lesson.)
|
||||
@@ -423,8 +425,9 @@ inside an entry about something else, and a finding rediscovered because nobody
|
||||
shape `[DESIGN]` in the architecture document and point it at the log entry** (§4's map).
|
||||
**Losing a reason is how a deliberate design becomes a bug in someone's eyes** — that cost four
|
||||
mis-filed defect reports in August 2026 (R-370, R-376).
|
||||
3. **State the register's size in the report, before and after.** A number every session is what makes
|
||||
growth visible; prose about tidiness is not a mechanism.
|
||||
3. **State the register's size in the report as FOUR numbers: rows before, rows after, rows opened, rows closed**
|
||||
(2026-10-05). A number every session is what makes growth visible; prose about tidiness is not a mechanism.
|
||||
Before/after alone hid that sessions closed some and opened as many.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -196,10 +196,11 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| Multiple household users / per-person accounts | — | **MISSING** | — | Single dashboard password; acceptable for alpha → R-15 |
|
||||
| A second login step for the dashboard (a TOTP code or a passkey) | — | **MISSING** | — | One password, one bcrypt hash (`controller/internal/web/auth.go:37-44`) → R-811 (added 2026-10-03) |
|
||||
| The household can leave Felhom, or outlive it — the box runs without the hub, the household owns its domain, tunnel and off-site account, and can export everything | — | **MISSING** (as a written answer) | the LOST-hub half only: `_recovery-inventory-2026-07-28.md` §D2.4, `07` §8 row 11b | Leaving and hand-over are answered nowhere → R-810 (spike, added 2026-10-03) |
|
||||
| **The host agent cannot reach root without the operator key: exact sudo patterns, no agent-written file installed where root reads it without a content check, the agent binary only by an operator-signed update checked as root** | agent **v0.146.1** (R-861; delivered by a step bundle, R-880) | **PROVEN-LIVE on both demo boxes (2026-10-05) — with three named residuals** | `audits/hub-safety-2026-10-05/partF/` (real sudo: 23 of 29 attacks allowed before, 0 after; 64 capability commands allowed; `sudo -l` 93/93 on demo-hp and demo-felhom after the bundle; capability probe 67/67; a staged unit over `/etc/sudoers.d` refused live) | `03` §3.1: the controller-swap image ref (guest-scoped), the felhom-op SSH key (hub-delivered, unsigned; felhom-op's sudo is scoped), the escrow ceremony relays R → R-861 |
|
||||
| WireGuard base infra always-on; OOB operator access (felhom-sshd, /32 peer) | agent v0.72, hub v0.35 | **IMPLEMENTED** | `SPIKE-oob-wg-operator-peer-2026-07-05`, `SPIKE-felhom-sshd-2026-07-05` | Mutual-repair desired-state arc not built → R-13. **CHECKED 2026-08-08 (R-260) and this row was NOT claiming something untrue** — it claims the capability is implemented, never that it is monitored, so no correction was owed. What WAS untrue is narrower and sat one layer down: **the hub's own OOB health check could not see whether the operator's key was installed.** `HostOOBRow` mirrored five of the agent's eight OOB fields, so `operator_key_configured` — emitted every heartbeat since agent v0.72.0, i.e. from this row's own vintage — was discarded by `encoding/json` on arrival, and `oobDegraded` returned `ok` for a box with felhom-sshd active, reachable, a valid config, a configured peer and **no operator key at all**. `operator_peer_configured`, which it did read, only says the peer IP is in desired-state — that OOB is MEANT to work, not that entry is possible. Fixed hub v0.99.0; the missing key now degrades and the alert NAMES it; a stanza too old to carry the field is reported distinctly and is never a silent ok. Pinned end-to-end from raw report JSON by `TestHostOOB_MissingOperatorKey_EndToEnd` and `TestHostOOB_NoKeyField_IsNotSilentlyOK_EndToEnd` |
|
||||
| The operator can see WHERE a managed host is — its LAN address and its WireGuard address, on the host page | agent **v0.119.0**, hub **v0.85.0** | **PROVEN-LIVE** (2026-07-31) | `audits/host-addresses-visible-2026-07-31.md` | Before this the LAN IP was **not reportable at all** — `HostMetrics` carried no address of any kind — and the WG IP existed only in `/offsite`'s peer table keyed by pubkey (peer→host, never host→peer). New wire field `addresses[]`, one row per (interface, address); `IsGlobalUnicast()` is the whole filter, chosen by MEASURING both demo boxes, and it needs no veth/fwbr denylist because that plumbing carries no IP. Rendered live on both 0.119.0 hosts matching their `ip addr` ground truth exactly. **Two honesty properties carry the risk and are both red-proofed:** WireGuard shows the hub ALLOCATION and whether the box CONFIRMS holding it (allocation alone cannot tell a live tunnel from a peer never applied), and an agent below 0.119.0 renders **UNKNOWN, never "no addresses"** — proven live on `drill-r50-0a4f9a` (0.113.0). **Not covered:** a two-LAN-bridge box and a real WG drift, neither of which exists to observe |
|
||||
| The operator can see whether a managed host's **guests still have working networking** — and **how often the watchdog had to repair them** | agent **v0.92.0** (emitter, 2026-07-21), hub **v0.104.0** (reader, 2026-08-13) | **IMPLEMENTED** | `backlog/OPEN-ITEMS.md` R-319; `hub/internal/web/hosts_guestnet_test.go` (7 tests, fixtures copied verbatim from `demo-felhom-8363b5`'s live `host_reports` row) | The agent emitted `guest_net` on every heartbeat for **twenty-three days** while the string occurred **nowhere** in `felhom.eu/hub/` — stored as raw text in `report_json`, read by nothing (R-260/R-264, the first of that census's readers to be built). **The fact that carries the risk is `heals_last_hour`, not `state`:** a guest the watchdog keeps repairing is healthy at every instant anyone looks, so rendering the state alone would give it a green tick — the failed-disk-drawn-as-a-healthy-empty-disk shape. `heal_succeeded` is decoded beside it, because six FAILED repairs is a guest that is down while six successful ones is a nuisance. **Unknown is never drawn as healthy:** three absences, three sentences (agent < 0.92.0; a capable agent that sent nothing; a guest whose own state the watchdog did not assert), and a malformed stanza degrades to unknown without a 500. **Three red-proofs, each mutation asserted applied by grep before its run**, including the one that matters — removing the unknown branches and watching a silent machine render as healthy. **Positive control that it is WIRED and not merely written: the wire-contract gate's checked-tag count rose 182 → 190** as the eight `guest_net` allowlist entries were deleted (an allowlisted tag is skipped, so leaving them would have meant these fields were never checked) | **IMPLEMENTED, not PROVEN-LIVE, and the distinction is the honest half.** Every scenario is proven against the real wire in tests, and the healthy case renders correctly for the live fleet — but **no machine has ever been observed with a climbing repair count on this card**, because neither demo box has needed a repair since the watchdog shipped. The signal this card exists for has therefore never been seen firing on hardware. It moves to PROVEN-LIVE the first time a real repair count is watched appearing. **No alarm was added, deliberately** (R-319): the incident behind this was about nobody being able to SEE the condition, and a new email on a fleet of two demo machines is untested noise — revisit when a third machine exists or when a count is seen climbing |
|
||||
| Break-glass management-plane recovery | agent v0.71, hub v0.84 | **IMPLEMENTED** | `runbooks/break-glass.md` | hub v0.84.0 adds an **operator-SESSION** retrieval path (host page → Console access → Reveal; `POST /hosts/{id}/reveal-recovery-credential`, CSRF-gated, writes a customer-visible `recovery_credential_revealed` event) beside the pre-existing **global-key** one (`GET /api/v1/admin/hosts/{id}/recovery-credential`), which is untouched and stays the route for when the hub UI itself is down. **The credential half is now PROVEN (2026-07-31):** the vaulted `demo-hp-bb76ea` password was verified against the box's own `/etc/shadow` hash AND minted a real PVE ticket — `POST /api2/json/access/ticket` → **HTTP 200, `root@pam`, 367-char ticket**, the exact API the login form submits to. Still IMPLEMENTED rather than PROVEN-LIVE overall, because the path has not been exercised on a REAL lockout (SSH was available throughout). Discovered during that check: the card's Copy button shipped `disabled` until a Reveal and silently no-opped, leaving ANOTHER host's password in the clipboard — fixed in hub v0.86.0. The vaulted secret is plaintext at rest → **R-133** |
|
||||
| Break-glass management-plane recovery | agent v0.71, hub v0.84 | **IMPLEMENTED** | `runbooks/break-glass.md` | hub v0.84.0 adds an **operator-SESSION** retrieval path (host page → Console access → Reveal; `POST /hosts/{id}/reveal-recovery-credential`, CSRF-gated, writes a customer-visible `recovery_credential_revealed` event) beside the pre-existing **global-key** one (`GET /api/v1/admin/hosts/{id}/recovery-credential`), which is untouched and stays the route for when the hub UI itself is down. **The credential half is now PROVEN (2026-07-31):** the vaulted `demo-hp-bb76ea` password was verified against the box's own `/etc/shadow` hash AND minted a real PVE ticket — `POST /api2/json/access/ticket` → **HTTP 200, `root@pam`, 367-char ticket**, the exact API the login form submits to. Still IMPLEMENTED rather than PROVEN-LIVE overall, because the path has not been exercised on a REAL lockout (SSH was available throughout). Discovered during that check: the card's Copy button shipped `disabled` until a Reveal and silently no-opped, leaving ANOTHER host's password in the clipboard — fixed in hub v0.86.0. **Sealed at rest since hub v0.135.0 (R-133 CLOSED):** the off-site seal and key; live 2026-10-05: 4 legacy rows sealed at start-up, 0 left plain, and the demo-hp reveal still minted a PVE ticket (HTTP 200; a wrong password 401) — `audits/hub-safety-2026-10-05/partB/`. A database backup now needs `OFFSITE_SECRET_KEY` too → R-173 |
|
||||
|
||||
## F. Notifications & monitoring
|
||||
|
||||
@@ -214,6 +215,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| Always-on debug rings + on-demand log-bundle pulls with TTL/custody | controller v0.116, agent v0.83, hub v0.46 | **PROVEN-LIVE** | debug rings live-exercised `CAMPAIGN-3` fix-6 (1000-cap ring, ~55min horizon under load) | The **log-bundle-pull TTL/custody** half is changelog-only (no dedicated observability audit doc); ring persistence across restart is a known gap |
|
||||
| Operator alerting (Healthchecks → monitoring@felhom.eu) | k3s, Resend | **IMPLEMENTED** | operator infra, stated in production since 02-04; no corpus validation doc |
|
||||
| Backup-deadline alerting (`expected_backup_missed`) is ANCHORED — absence of signal is UNKNOWN, not failure | hub v0.75.0 | **IMPLEMENTED** | `audits/DIAG-backup-missed-2026-07-26.md` + red-proofs A/B/C + replay of the real 2026-07-26 03:00 reports (all three silent) | **No row status flips** — this signal had a FALSE-POSITIVE class (three instances: hub v0.12.0, v0.73.0, R-81), now anchored at first contact and read across retained host-report history. Still unit-proven only, not live-fired at a real deadline. The *underlying* PBS/offsite-DR tier gap it exposed is → R-82. | Per the status enum, no citation → not PROVEN-LIVE. Demoted pending an operator-cited live alert (candidate re-upgrade — see REPORT) |
|
||||
| A box that is not always on (off at its backup time) makes the missed night up once when it comes back; the household sees a banner; the operator gets the missed-backup alarm, down or not | controller v0.295.0, hub v0.134.0, agent v0.145.0 | **PARTIAL — the catch-up and the banner PROVEN-LIVE (9202 + demo-felhom; banner on 9202 by a hand-set ledger); R-872's alarm proven by tests, live at the next 05:00 (dated check); a host SUSPEND reasoned + unit-tested, not measured** | `audits/catchup-2026-10-05/partA/`, `partB/`; design `07` §6.1.1 | R-872 live; a suspend never measured |
|
||||
|
||||
## G. Fleet & operator (hub)
|
||||
|
||||
@@ -232,7 +234,10 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
|
||||
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
||||
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/` |
|
||||
| **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live |
|
||||
| **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs | `audits/hub-db-offsite-2026-10-05/`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | Runbook §3 steps 4–5 (into a live PVC) not exercised (R-173) |
|
||||
| **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` |
|
||||
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
|
||||
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
|
||||
|
||||
|
||||
@@ -370,6 +370,23 @@ own; every caller that is not the customer must decide for itself whether the ap
|
||||
| `storage_handlers.go` (1600 L) | **DELETE (→agent)** | Format/attach/mount/disconnect/migrate-drive/decommission disk UI. Any survivor is a **thin client calling the agent API** (e.g. per-volume placement requests). | hazard |
|
||||
| `templates/` (HTML, non-Go) | **PORT** | Remove disk-wizard + DR pages; keep app/deploy/backup/settings pages. | needs-rework |
|
||||
|
||||
#### Alert placement — inline under the storage bars, or the top banner (R-571)
|
||||
|
||||
**[FACT, read from source 2026-10-05, felhom-controller `114ff27`, `controller/internal/web/alerts.go`]** Every
|
||||
dashboard alert is an `Alert` with two placement fields: `PageOnly` (the pages it may appear on; empty = every
|
||||
page) and `Inline` (rendered by the page template in place, not by the layout's banner). `GetAlerts` returns
|
||||
the list (endpoint-drift, agent-channel and dead-app alerts first, then the rest sorted error > warning > info,
|
||||
capped at five plus an overflow line), and `layout.html` paints only those that are not `Inline` and match the
|
||||
page; `GetInlineAlerts(page)` hands the dashboard and monitoring pages their inline ones.
|
||||
|
||||
**Exactly one warning is inline today:** the „storage is not on a separate drive" health warning. It is
|
||||
`PageOnly: dashboard, monitoring` and `Inline: true`, so it sits quietly under the storage bars; **every other
|
||||
warning, including every off-site failure (`07` §6.7), renders in the top banner on every page.** The choice
|
||||
is made by the warning's KIND (`monitor.WarnKindStorageNotSeparate`, read with `report.WarningKindAt`), never
|
||||
by its words — the earlier Hungarian-substring test would have moved the warning to the red banner on every
|
||||
page the day the sentence was translated (R-553, pinned by `TestR553_DiskWarningPlacementSurvivesWordingChange`).
|
||||
A new inline warning needs its own kind, not a text match.
|
||||
|
||||
### `scripts/`
|
||||
| File | Class | Reason | Risk |
|
||||
|---|---|---|---|
|
||||
|
||||
@@ -72,6 +72,59 @@ Explicitly does **not**:
|
||||
- **Root-minimized (boundary settled — Phase 3 B3).** The agent runs as a **non-root** service user with the scoped `FelhomAgent` token for all API-covered work + a **narrow `sudoers` allowlist** for true host ops. Per Phase 3 (B3) the boundary is settled: the entire per-customer guest lifecycle — provision (by restore, §9), config, start/stop, snapshot, backup, **restore**, destroy — is token-covered. Genuine OS-root is confined to: (1) building/refreshing the **golden base image** (`keyctl` create is `root@pam`-only — one-time at enrollment + a maintenance cadence, §9); (2) **host mounts** (USB mount-by-UUID, systemd mount units / fstab); (3) **SMART / hardware sensors**. Root therefore never sits on the per-customer path. See `proxmox-platform.md` §3.6 for the role + boundary table.
|
||||
- ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:<pin>`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent does not manage it; it READS its health (R-841, agent v0.141.0). Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins").
|
||||
|
||||
### 3.1 The agent's admin commands, group by group — and how each is narrowed (R-861, agent v0.146.1) `[DESIGN — 2026-10-05, CC; operator may reverse]`
|
||||
|
||||
**The question.** Can a compromised agent PROCESS (running as the `felhom-agent` user) become root on its host without
|
||||
the operator's key? Read on 2026-10-04 (R-861) and measured 2026-10-05 with the real sudo 1.9.16 in a throwaway
|
||||
container: **with the v0.145.0 sudoers, yes — 23 of 29 attack command lines were allowed.** Two shapes did it:
|
||||
|
||||
1. **A glob in the arguments.** Sudo's `*` in arguments also matches spaces, so one grant smuggled extra options:
|
||||
`pct set [0-9]* -onboot 1` allowed `pct set 100 --dev0 /dev/sda -onboot 1` (a raw host disk for the guest);
|
||||
`mount --bind /mnt/*/felhom-data /mnt/felhom-drives/*` allowed a `..` path onto `/etc/sudoers.d`; `nft add element …
|
||||
*` allowed `; flush ruleset`.
|
||||
2. **A file the agent wrote, installed where root reads it.** A `.mount` unit (bind any directory over `/etc`), a
|
||||
dnsmasq drop-in (`dhcp-script=` runs as root), the wg-quick config (`PostUp=` runs as root), the OOB sshd config
|
||||
(`AuthorizedKeysFile` + `StrictModes no`), the guest pre-start hook (Proxmox runs it as root), the shared-parent boot
|
||||
script, and the agent BINARY itself (the escrow ceremony and the guest hook run it as root; `apply` took a sha the
|
||||
agent passed).
|
||||
|
||||
**The rule after v0.146.1** (the R-861 fix direction: each becomes a root-owned wrapper that checks its own input, or a
|
||||
fixed file, delivered by the signed config bundle):
|
||||
|
||||
- Every varying argument list is a **sudo regular expression** (`^…$`): one value per slot, a fixed character set, no
|
||||
`..`, no extra argument. Literal lines stay literal.
|
||||
- **No file the agent wrote is installed where root reads it.** Either the content is FIXED and comes with the signed
|
||||
bundle (the hook, the shared parent), or a root wrapper checks the CONTENT against the agent's own renderers before
|
||||
installing it (`felhom-priv-apply`), or the operator's signature is checked as root (`felhom-os-apply agent_update`).
|
||||
- The pins: `TestManifestCoveredBySudoers` (every command the agent runs is still allowed), `TestSudoersRefusesTheR861
|
||||
Injections` (the 29 attacks are not), the real-sudo run of both (`audits/hub-safety-2026-10-05/partF/
|
||||
sudo-container-proof.txt`), and `sudo -l -U felhom-agent` on both demo boxes after the bundle.
|
||||
|
||||
| Group | What it is for | How it is narrowed (v0.146.1) | Left open |
|
||||
|---|---|---|---|
|
||||
| `FELHOM_MOUNT` | fs-UUID mount units for enrolled drives | install only via `felhom-priv-apply unit <name>`: `[Unit]` only Description + `After=local-fs-pre.target`, `Where=` `/mnt/<name>` or `/mnt/felhom-drives/<name>` and equal to the unit name, `What=` a UUID or a network source, no `bind`/`suid`/`dev`, no continuation lines; systemctl verbs on `mnt-…\.mount` only | — |
|
||||
| `FELHOM_NETMOUNT` | NAS automount pairs, re-arm, clean-up | same checker; a network share must carry `nosuid,nodev` (the agent now renders them); `rm`/`rmdir`/`reset-failed` one exact name | — |
|
||||
| `FELHOM_DISK` | SMART, thin-pool, PV and pool reads | exact device / LV patterns (no extra options such as `smartctl -s off`, `lvs --config`) | read-only |
|
||||
| `FELHOM_PROVISION` | bootstrap config mount, autostart | exact `mpN` spec (`…/guests/<vmid>/bootstrap,mp=/…[,ro=1]`) — no smuggled `--dev0` | — |
|
||||
| `FELHOM_FORMAT` | data-bearing probe + guarded mkfs | one device path, no space; `mkfs` only through `felhom-mkfs-guarded` (its own root checks) with `ext4`/`xfs` | blkid/lsblk read any `/dev` path (read-only) |
|
||||
| `FELHOM_DNSMASQ` | the LAN split-horizon resolver | drop-ins via `felhom-priv-apply dnsmasq` (only `bind-interfaces`, `no-resolv`, `listen-address`, `server`, `local`, `address`); `rm` one exact name; exact `pct exec` reads | — |
|
||||
| `FELHOM_GUESTHOOK` | the pre-start self-heal hook | the hook is a FIXED bundle file; the agent only checks it (`SnippetReady`) and registers it; exact vmid/slot | — |
|
||||
| `FELHOM_INTERMEDIARY` | the shared drive parent + live drive binds | boot script + unit are FIXED bundle files (the agent only enables the unit); one-segment drive names (no leading dot, no `..`) | — |
|
||||
| `FELHOM_CONTROLLERSWAP` | the managed controller update | exact vmid; image ref pinned to `gitea.dooplex.hu/admin/felhom-controller:X.Y.Z` for the image check; the inspect template stays free text | **guest-scoped by design**: a compromised agent can still `tee` a chosen (pinned-registry) image ref and restart the guest's bootstrap — the household's data, not host root |
|
||||
| `FELHOM_STALELOCK` / `FELHOM_SCRATCH_TEARDOWN` | stale-lock clear; failed restore-test scratch | exact vmid; the scratch band `99000[0-9]` was already exact | — |
|
||||
| `FELHOM_WG` | the off-site tunnel | conf via `felhom-priv-apply wg` (only the keys `renderConf` writes; no `PostUp`/`PreUp`/`DNS`/`Table`; `/32` only) | — |
|
||||
| `FELHOM_SELFUPDATE` | commit / rollback of the A/B flip | **`apply` removed**: the flip runs only inside `felhom-os-apply agent_update`, after the operator signature, host, window and nonce are checked as root and the staged bytes are hashed ONCE and copied to a root-owned dir (`/var/lib/felhom-os-apply/agent-update/`); the wrapper accepts only that dir | — |
|
||||
| `FELHOM_SSHD` | the out-of-band operator sshd | config via `felhom-priv-apply sshd-config` (the ONE template, only the Port varies, never 22); the felhom-op key via `sshd-key` (one plain key, no `command=`/`from=` options) | felhom-op's key itself is hub-delivered, not signed: a compromised agent can install its own key for **felhom-op** — whose sudo is scoped (`felhom-op.sudoers`), not root |
|
||||
| `FELHOM_OOB` | the OOB firewall sets | `add element` takes exactly `{ <ip>[/n] }` or `{ <port> }` — no chained command | — |
|
||||
| `FELHOM_PBSDR` / `FELHOM_BACKUPTARGET` | PBS DR entry; whole-system backup target | unchanged: the arguments stay coarse, and the root wrappers (`felhom-pbs-apply`, `felhom-backup-target-apply`) are the gate (fixed verbs, own validation) | coarse argv into a checking wrapper |
|
||||
| `FELHOM_ESCROW` | the recovery-code ceremony (runs the agent binary as root) | the binary is only ever an operator-signed one (`FELHOM_SELFUPDATE`); as root it pins the PVE secret dir and the WG state dir, refuses a storage id that is a path, and reads its two staged files by walking the path with `openat(O_NOFOLLOW)` (no symlink anywhere) | **by design the agent relays R**, so a compromised agent can still learn this box's PBS key through the ceremony — not root, but the backup key |
|
||||
| `FELHOM_SELFHEAL` / `FELHOM_GUESTNET` / `FELHOM_OSAPPLY` | networking restart; guest DHCP watchdog; OS updates | exact; `felhom-os-apply --plan …` stays the glob line on purpose — the bundle's own self-check reads that exact text, and the wrapper refuses any other plan path (R1) | — |
|
||||
|
||||
**What this does not change.** The operator key (`/etc/felhom/operator-signers`, root-owned, never a bundle path) stays
|
||||
the one trust root; the agent's API token is untouched; a box gets the new rule only through the signed
|
||||
`agent_config_update` (the order: signed `agent_update` first, then the bundle — after the bundle, an agent below
|
||||
0.146.0 cannot update itself on that box).
|
||||
|
||||
## 4. Control model — reconcile + signed destructive ops
|
||||
|
||||
Two channels, split by **reversibility**, not by transport.
|
||||
@@ -655,7 +708,8 @@ buildable until then; recorded here so the front-half built in slice 7 lands rea
|
||||
runs as root at guest start (and `pct reboot` is granted); `FELHOM_INTERMEDIARY` installs a script and a systemd unit
|
||||
that run as root at boot; `FELHOM_ESCROW` runs the agent binary as root, and `FELHOM_SELFUPDATE apply` accepts a sha
|
||||
the agent itself passes. So a compromised agent PROCESS is root on its host; the root-owned trust files (decision 93,
|
||||
the bundle's R17) are defence in depth, not a boundary, until R-861 narrows these grants.
|
||||
the bundle's R17) are defence in depth, not a boundary, until R-861 narrows these grants. **Narrowed in agent
|
||||
v0.146.1 — §3.1 lists every group, the rule now, and what stays open.**
|
||||
- **Controller (the easy case — it's a guest).** The agent owns the controller's lifecycle,
|
||||
so the **agent updates the controller**: snapshot-before-update (free rollback, because the
|
||||
controller *is* a snapshottable guest) → pull new image → redeploy → health-check → rollback
|
||||
|
||||
@@ -130,6 +130,17 @@ R-216. Either way the box's reported agent must meet the chosen MinAgent, else t
|
||||
the Hosts page and logged once per change as `managed floor SERVED`. The rules the operator follows:
|
||||
`runbooks/publish-train-rules.md` rule 1.
|
||||
|
||||
**Boxes left behind (hub v0.135.0, R-604 + R-530).** A per-customer floor wins over the global one, so a global
|
||||
raise does not move a box whose OWN floor is lower — and `managed floor SERVED` is logged once per change, so that box
|
||||
was silent (demo-hp missed four raises, 2026-09-21). Now the raise logs one line per such customer and sends ONE
|
||||
operator mail naming them (`floor_raise_skipped`); a per-customer floor records when it was set
|
||||
(`customer_configs.min_controller_set_at`; "unknown" for one set before v0.135.0), and the System page's "Version
|
||||
floors" table lists the global floor, every per-customer floor with its age, and which ones the global cannot move.
|
||||
Agents are a separate train (they update only by a per-box signed job, R-530's ruling): the System page shows each
|
||||
box's agent against the vouched one ("0.142.0 → 0.145.0 (since …)", amber, red after the wait), and a box behind the
|
||||
vouched agent for 7 days raises `agent_behind` (warning, operator-only; `OS_ALARM_AGENT_BEHIND_AFTER`; the clock starts
|
||||
when the hub first sees the box behind). `[DESIGN — CC 2026-10-05, operator may reverse; `09` decision 119]`
|
||||
|
||||
## 6. Authorization — signed-op queue + editing flow
|
||||
|
||||
Implements Part 4's gate on the hub side. The hub holds **no signing key**.
|
||||
@@ -403,3 +414,58 @@ count appeared was the bind page's passphrase hint, and its English half is now
|
||||
phrase you received from your operator during setup") — because "five words" stops being true for an
|
||||
English household, and was already wrong for one whose passphrase predates this release. The
|
||||
Hungarian „öt szó" is correct and unchanged.
|
||||
|
||||
## 16. The operator surface's own safety [DESIGN — hub v0.135.0, CC 2026-10-05, operator may reverse]
|
||||
|
||||
### 16.1 Form protection (CSRF) on both login paths (R-135)
|
||||
|
||||
The operator logs in two ways: a browser session (`__Host-hub_session` cookie since v0.138.0, R-136 — always Secure, `Path=/`, no Domain, so a sibling
|
||||
subdomain cannot plant one; + a per-session token on every form) and HTTP
|
||||
Basic for scripts. Until v0.135.0 a state-changing request with NO cookie skipped the token check — on the reasoning
|
||||
that it must be a script. It need not be: a browser caches Basic credentials per origin and resends them on a
|
||||
cross-site form POST (SameSite does not govern the Authorization header). Now a request without a session passes
|
||||
only with Basic credentials AND the header `X-Felhom-Operator` (any value; scripts send `cli`). A page on another site
|
||||
cannot add a custom header without a CORS preflight, which the hub never answers. The gate sits in `ServeHTTP` before
|
||||
the route switch, so it covers every route at once (38 state-changing routes + an unknown path in
|
||||
`r135_csrf_test.go`); `/login` and the public `/bind/<token>` stay exempt (no operator session to ride; the bind token
|
||||
is the capability). The other choice — dropping browser-usable Basic auth entirely — was not taken: CC's headless runs
|
||||
and the runbooks drive the hub with Basic auth, and the header costs them one flag.
|
||||
|
||||
### 16.2 Secrets at rest in `hub.db` (R-133, R-821)
|
||||
|
||||
Two columns are SEALED (AES-256-GCM, `enc:v1:` + nonce, one key: `OFFSITE_SECRET_KEY` from `Secret/offsite-secret-key`,
|
||||
never in the database or git): the off-site sub-account passwords (`one_time_secrets.value`, v0.127.0) and, since
|
||||
v0.135.0, each box's break-glass console password (`host_recovery.secret`) — the same helpers, not a second scheme.
|
||||
Legacy rows are sealed in place at start-up (`SealLegacyRecoverySecrets`, measured live: 4 rows). No key → a save is
|
||||
refused; a wrong key → a reveal is a 500 with nothing in the body or the log. Both retrieval paths (the operator page and
|
||||
the global-key API) open through `GetHostRecoveryCredential`, so the break-glass route still works with the UI down.
|
||||
|
||||
**Since v0.138.0 (R-879) four more columns are sealed the same way:** each box's hub API key (`hosts.api_key`), each
|
||||
household's controller API key and owner passphrase (`customer_configs.api_key`, `retrieval_password`) and the PBS-DR
|
||||
token values (`host_pbs_secrets.value`). Hashing or deleting was not possible — every one of these values is served again
|
||||
(re-enroll returns the key, the config ships the controller key, the operator page shows the passphrase, a PBS token can
|
||||
be re-staged). The two API keys also carry an UNKEYED SHA-256 twin (`api_key_hash`, backfilled at every start), and a box
|
||||
is looked up by that hash — **so a box authenticates even when the sealing key is missing or wrong** (`TestR879_BoxAuthSurvivesFailedSealing`).
|
||||
Legacy rows are sealed at start-up (`SealLegacyBoxSecrets`, idempotent, non-fatal). A value that does not open marks the
|
||||
record unreadable: its serve paths answer 500, a save refuses it (never blanks it), a PBS token is not consumed. **Roll-back
|
||||
to a hub before v0.138.0** needs the columns opened first: `felhom-hub -unseal-box-secrets` (all or nothing; exit 1 =
|
||||
nothing changed) — run in the running pod, then put the old image back. `guests.api_key` is an unused, always-empty
|
||||
column and is not sealed. *Decided by CC unattended — operator may reverse (`09` §3 decision 132).*
|
||||
**And the key is the other half:** a backup of the database restores a hub that can open the sealed columns only
|
||||
with the same `OFFSITE_SECRET_KEY`; today that key exists only on DooPlex (the k8s Secret, and the GPG secrets export
|
||||
on the same machine). The off-site plan for the database and its key: `runbooks/RUNBOOK-hub-db-offsite-backup.md`
|
||||
(R-173 — option A decided 2026-10-05, `09` decision 125; §16.3).
|
||||
|
||||
### 16.3 The nightly database snapshot (R-173, hub v0.136.0)
|
||||
|
||||
Every night at **02:00 Budapest** the hub writes `<data>/snapshots/hub-<UTC stamp>.db` with SQLite's `VACUUM INTO`:
|
||||
one statement, one point in time, every committed write included — the rows still in `hub.db-wal` too, which a file copy
|
||||
of `hub.db` loses (measured in `internal/dbsnap` tests: a plain copy held 0 of 150 fresh rows). Written as `.tmp`, mode
|
||||
`0600`, then renamed, so a reader never sees half a file. The newest **2** stay (the volume grew from 1 GiB to 2 GiB
|
||||
for them, operator choice 2026-10-05; one snapshot measured 370 MB). Two runs never overlap (a second gets `ErrBusy`).
|
||||
Each run logs `db snapshot written: <name> (<bytes>, <duration>)`. At start-up the hub runs one when the newest is older
|
||||
than 24 h. **The hub ships nothing itself and holds no ep0 credential:** DooPlex's `felhom-hub-db-backup` unit picks the
|
||||
newest snapshot up at 02:30, checks it (`PRAGMA integrity_check`), encrypts it and pushes it to ep0's `operator`
|
||||
namespace (`runbooks/RUNBOOK-hub-db-offsite-backup.md`). Pinned by `internal/dbsnap/dbsnap_test.go` and
|
||||
`cmd/hub/r173_wiring_test.go`.
|
||||
|
||||
|
||||
@@ -133,6 +133,17 @@ startup a record still marked running becomes a failed, interrupted result („A
|
||||
restore, and raised once as `restore_interrupted`. **Notification cooldowns stay in memory** — the
|
||||
precedent is kept for what it was written for.
|
||||
|
||||
**A backup RUN cut off by a power cut or a restart is said (R-519, controller v0.296.0, `09` decision 123).** The
|
||||
same shape for the app-data run: `appdata-run.json` beside the restore record, written at both ends of a run. A start
|
||||
that finds it still running turns it into a notice on /backups and /backups/apps („A legutóbbi mentés (…) megszakadt,
|
||||
mert a doboz vagy a vezérlő újraindult…"), kept until a run ends with every step OK. Measured BIGNIGHT F2 (2026-09-14):
|
||||
before this, both pages said nothing and the synthesised „Utolsó adatbázis mentés … OK" was read off the fresh `.sql`
|
||||
the cut run left beside last night's tars; that line now reads failed after a cut. Each restore point's time was
|
||||
already its OLDEST part (the data block, v0.275.0) — so a torn unit is dated by its stale tars, never by its new dump.
|
||||
*Live: PROVEN 2026-10-05 on 9202 (operator ruling 126): a run cut by a controller restart between bookstack's config
|
||||
dump and its database volume → both pages carry the notice, the restore point reads the older volume's time, the next
|
||||
complete run clears it (`audits/hub-db-offsite-2026-10-05/partD/r519/`; R-519 closed).*
|
||||
|
||||
### Lane 2 — the operator: guest and host recovery
|
||||
|
||||
Rebuilding an LXC guest, rebuilding a host, and re-establishing a box's identity are **operator
|
||||
@@ -238,7 +249,8 @@ the identity bundle's shape is `{tunnel_token, pbs_token, wg_private_key, restic
|
||||
parts a host-loss recovery would read (INV Part D2.3): `hosts.dr_record_json` is `{}` on all three
|
||||
hosts; `host_escrow.directive_json` is `{}` on both escrowed hosts; `dr_recipe.host_half.drives` is
|
||||
`[]` on every customer including two with enrolled data drives; and `dr_recipe.host_half.pbs.namespace`
|
||||
reads `"root"` while the real namespaces are `demo-felhom` / `demo-hp`. → **R-105**, **R-106**.
|
||||
reads `"root"` while the real namespaces are `demo-felhom` / `demo-hp`. → **R-105**, **R-106**. *Since agent v0.147.0 (R-124) a genuine root namespace is
|
||||
recorded as `""` (PBS's own spelling) beside `namespace_state: resolved`; the word `root` is no longer written.*
|
||||
|
||||
---
|
||||
|
||||
@@ -396,6 +408,80 @@ warning that should precede such a failure is R-685.
|
||||
`pct config 9201`, both hosts). So the whole-guest tiers carry the guest and **none of the customer's
|
||||
data drives** — 916 GB on demo-felhom, 938 GB on demo-hp.
|
||||
|
||||
### 6.1.1 A box that is not always on — the catch-up, the banner, the alarms (R-871..R-874)
|
||||
|
||||
> **Why this section exists (R-871):** until 2026-10-05 no architecture document covered a box that is OFF at its
|
||||
> window W. Tester 2 is a laptop switched off at night (operator, 2026-10-05). Measured (`audits/night-fixes-2026-10-05/
|
||||
> partF/FINDINGS.md`, and live on 9202 `audits/catchup-2026-10-05/partA-spike/`): the controller's daily jobs always
|
||||
> schedule the NEXT future time, so a missed 02:30/03:30/04:15 waited for the next night for ever; no alarm fired
|
||||
> while the box was down at 05:00; the household was mailed "your server cannot be reached" every night.
|
||||
|
||||
**[DESIGN — ruled 2026-10-05, `09` decision 109] A missed night runs ONCE when the box comes back.** Built controller
|
||||
v0.295.0 (`internal/nightchain`):
|
||||
|
||||
- **What a box does when it was off at W.** On a controller START and on a host RESUME, the controller asks its
|
||||
night ledger (`<data>/night-ledger.json`) which backup legs missed their last scheduled time. The ledger records when
|
||||
each leg last RAN TO ITS END — an attempt record, used only to decide "missed", never as evidence a backup exists
|
||||
(`R-100`'s rule). A leg that ran and failed was NOT missed: failures have their own alarms.
|
||||
- **What the catch-up runs:** the three BACKUP legs, in the night's order — database dump, second copy, off-site copy
|
||||
— through the SAME wrapped leg bodies the scheduled jobs run (`main.go` `withLeg`, `catchUpLegs`).
|
||||
- **What it does not run:** the app-update leg and every Docker step (they restart apps; they wait for a real night).
|
||||
The OS fast lane is the agent's and follows a whole-guest backup, not the catch-up (below).
|
||||
- **When:** 15 minutes after the trigger (apps settle; a box switched on and off again at once does nothing). A leg
|
||||
whose own next scheduled time is under 30 minutes away is left to its normal run.
|
||||
- **Only once:** several missed nights = one catch-up (the question is about the LAST scheduled time). A normal night
|
||||
followed by a daytime restart = none. A power cut in the middle of the chain = only the legs that did not end.
|
||||
- **Never two at once:** every leg, scheduled or catch-up, holds one lock; a second trigger while one is pending does
|
||||
nothing.
|
||||
- **The whole-guest backup.** It keeps its own agent-side behaviour (inside [W+2h, W+6h), or the 48 h safety valve —
|
||||
about one every 2 days for an evening-only box). **It CAN collide with a catch-up** (the valve fires on the first
|
||||
5-minute poll after a start), so each waits for the other: a scheduled quiesce defers while a catch-up runs
|
||||
(`quiesce` `SetCatchUpFn`), and a catch-up waits up to 2 h while a quiesce holds the apps.
|
||||
- **A suspended host** (a laptop lid): Go's timers run on CLOCK_MONOTONIC, which does not count suspended time, so
|
||||
the 02:30 timer would fire hours late — and would start the app-update leg at noon. Two rules: a daily job whose
|
||||
timer fires more than 60 minutes after its wall-clock time is SKIPPED (`scheduler.DailyLateLimit`), and a resume
|
||||
watch (wall clock vs monotonic clock, checked every minute) triggers the catch-up. *Reasoned from the Go and Linux
|
||||
clock semantics and unit-tested with injected clocks; not measured live (no suspend was allowed on a demo box).*
|
||||
- **The first start on this release** seeds the ledger: nothing before it counts as missed (no surprise catch-up on
|
||||
every box at the upgrade).
|
||||
- **What the household sees:** one timeline line — "Kimaradt mentés pótolva: a doboz ki volt kapcsolva 02:30-kor, a
|
||||
mentés most elkészült." / "Missed backup made now: the box was off at 02:30, so the backup ran when it came back on."
|
||||
(event `backup_catchup_done`, info: recorded, never mailed).
|
||||
|
||||
**[DESIGN — ruled 2026-10-05, `09` decision 110] The banner (the operator's idea).** On every page of a logged-in
|
||||
household: when the last daily backup (the database dump, and the off-site copy when configured) is over **26 h** old —
|
||||
when the last one was, "the box was off at backup time (02:30)" when the box's own record says so, and a suggested
|
||||
time. The record is the controller's system-metrics table (one sample a minute, kept 30 days): "on" at W = a sample
|
||||
within 5 minutes of it; "usually on" in an hour = on in it on at least 5 of the last 7 days. It **never changes the
|
||||
time** — a button opens the backup-time setting. The household can close it: it stays closed until the NEXT missed
|
||||
backup time (durable, in the ledger) and disappears by itself after a successful night.
|
||||
|
||||
*Decided by CC — operator may reverse* (`09` decisions 112–116): the 15-minute delay and 30-minute leave-to-normal
|
||||
line; the 60-minute late-fire limit; 26 h as the banner's line (one night + the chain's two hours, = the hub's
|
||||
`backupStaleAfter`); the suggestion = the LATEST hour H with H, H+1, H+2 usually on (a later evening disturbs least),
|
||||
none when the current window is already usually on; the catch-up's line on the household's timeline but no mail.
|
||||
|
||||
**The operator's side (hub v0.134.0, `08` §6.4).** R-872: a box DOWN at the 05:00 deadline is no longer skipped — it
|
||||
is judged on longer lines (48 h without a dump, 72 h without a whole-guest backup), so a box that died last night
|
||||
raises only its staleness alarm, and a box off at every deadline raises `expected_dbdump_missed` /
|
||||
`expected_backup_missed`. R-873: the household hears "your server cannot be reached" at most once per 7 days (the
|
||||
operator still gets every edge). R-874 (agent v0.145.0): the restore-test's first due-check runs 30 minutes after the
|
||||
agent starts, so a box with short power-on sessions is still restore-tested.
|
||||
|
||||
**[FACT, 2026-10-05] Proven live** (`audits/catchup-2026-10-05/partA/`): 9202 off across 09:35, on at 07:38 UTC → the
|
||||
catch-up at 07:53:09 made the dump (21 s); a crash of its host mid-wait → at the next start ONE new catch-up made all
|
||||
three legs (22 s); demo-felhom's controller off across 10:07 → the dump made at 08:25:03, 15 min after the start, and
|
||||
the household's timeline line reached the hub. **What the household may notice:** the dump leg stops an app with a
|
||||
volume for its copy — measured 1 s for opengist (182.5 KB), the same as at night; a large volume takes longer, in the
|
||||
day (R-878).
|
||||
|
||||
**Edge cases, stated.** A box on only in the day with W at night: the catch-up makes the backups every day it is
|
||||
switched on (after 15 min); the banner suggests an evening time once a pattern exists. A box switched on during the
|
||||
chain: legs already past are made up, legs still ahead run normally. A box whose controller restarts in the day after
|
||||
a normal night: nothing. A box off for weeks: one catch-up when it returns; the hub's staleness alarm has been
|
||||
running all along. An upgrade from an older release: the ledger is seeded, the first missed night after it is the
|
||||
first one made up.
|
||||
|
||||
### 6.2 Coverage per app class — and an unresolved count
|
||||
|
||||
**[FACT]** Of 53 catalog templates, **52 keep data in Docker named volumes**; exactly **13** carry a
|
||||
@@ -587,6 +673,11 @@ successes only. After an agent restart the success is read back from the tier's
|
||||
**What a run may do** (R-518, cheap half). A tier the agent reports `storage: absent` is dropped before
|
||||
anything is stopped, logged, and reported once as `backup_tier_skipped`; `unknown` is never skipped.
|
||||
**Still open:** quiescing per tier, so a slow second tier does not keep every app down.
|
||||
**Measured 2026-10-05 on demo-hp (9 apps, controller v0.295.0):** „Mentés most" stopped the apps at 09:19:08Z, the
|
||||
local tier ran 09:19:29–09:24:09, the PBS tier was busy (the controller logged a retry in 15 min; no second stop was seen in the next 55 min), the last app was back at 09:24:55Z — the
|
||||
longest stop **5 min 47 s**, for the local tier alone. The button text and its confirm (v0.296.0) give both
|
||||
measurements (≈6 min / 9 apps, ≈8 min / 12 apps), say "minutes, not seconds", and that the off-site copy in the same
|
||||
run makes it longer.
|
||||
|
||||
### 6.5 Kept data — what a removed app leaves on the drive (controller v0.274.0, `09` §3 decision 36)
|
||||
|
||||
@@ -705,6 +796,32 @@ digests all resolve today (`audits/version-travel-2026-09-26/A7/`). Options are
|
||||
it; older images of the app are deleted. A restore that needs an older version re-pulls it — as before; the limit
|
||||
above is unchanged, and kept data (decision 40) is not touched by the image clean-up.
|
||||
|
||||
### 6.7 Why an off-site run failed — the failure classifier (R-571)
|
||||
|
||||
**[FACT, read from source 2026-10-05, felhom-controller `114ff27`]** When an off-site (restic) run fails, the
|
||||
controller names ONE cause before it writes the note the customer's page shows days later
|
||||
(`ClassifyOffsiteFailure` in `controller/internal/backup/offbox.go`; the head line is the bundle key
|
||||
`note.offsite.fail_<class>`, followed by the run time and the sanitised error). The classes, in the order
|
||||
they are tested:
|
||||
|
||||
| Class | Decided by | What it means for the household |
|
||||
|---|---|---|
|
||||
| `orphaned` | our sentinel `ErrOffboxOrphaned` | the remote store was made with a key this box no longer has; nothing new reaches it until the operator acts |
|
||||
| `quota` | our sentinel `ErrOffsiteQuota` (the pre-run soft-quota gate, R-553) | the backup did not fit the remote space; the run was refused before upload |
|
||||
| `locked` | our sentinel `ErrOffsiteLocked`, or restic's lock text (R-104) | an interrupted earlier run left the store locked and both self-heal layers failed |
|
||||
| `no_units` | text: „produced no snapshots" | there was nothing to send — no chosen app had a backup on any drive |
|
||||
| `no_repo` | restic's text: „unable to open config file" / „is there a repository…" | nothing exists at the remote location |
|
||||
| `transport` | ssh/restic/rclone text: connection refused/reset, timeout, permission denied, host key, handshake, DNS, unreachable | the remote store could not be reached (network or sign-in) |
|
||||
| `unknown` | everything else | the cause is not known — the page says so instead of guessing |
|
||||
|
||||
**Two kinds of signal, and the difference matters.** The first three are OUR sentinels: they survive a
|
||||
translation of our own text, which is why R-553 replaced the Hungarian-word match for `quota`. The text
|
||||
signatures are **restic's, ssh's and rclone's own English output** — external strings we neither write nor
|
||||
translate. A new restic or OpenSSH version that rewords an error moves that failure to `unknown`; it never
|
||||
moves it to a wrong class. The order is deliberate: a cause that cannot be told apart returns `unknown`
|
||||
rather than being folded into a neighbour. Where the warning is SHOWN on the dashboard is a separate rule —
|
||||
`02-controller-module-map.md`, „Alert placement".
|
||||
|
||||
## 7. The recovery chain (D3) — the reason this document exists
|
||||
|
||||
**[DESIGN] 3-2-1 describes copies. It does not describe recovery.**
|
||||
|
||||
@@ -1,5 +1,14 @@
|
||||
# 08 — The app-down alarm ladder
|
||||
|
||||
> **How to read this document.** Where a statement is marked, it is marked like this — the same wording as
|
||||
> `07-backup-architecture.md:11-17`, carried here on 2026-10-05 (R-376, the three documents written after the
|
||||
> 2026-08-22 pass):
|
||||
>
|
||||
> - **[DESIGN]** — a decision taken. Not derived from code; the code may not implement it yet.
|
||||
> - **[FACT]** — an observed property, carrying a `file:line`, a live command output or a citation.
|
||||
>
|
||||
> **An unmarked statement means "not yet classified", never "observed"** (R-376).
|
||||
|
||||
**Written 2026-08-23, with controller v0.222.0 (R-384).**
|
||||
|
||||
**The absence is the finding.** Until this file existed, no document owned the question *"when does a
|
||||
@@ -327,6 +336,8 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
|
||||
| `host_crash_restart` | warning | the box's crash guard reports a NEW unclean boot (a crash, a power cut or a hard reset; hub v0.132.0) | — (one per boot) | `api/crash_test.go` |
|
||||
| `host_crash_guard_tripped` | error | the guard tripped: the next crash leaves the box OFF | the re-arm → `host_crash_guard_rearmed` (info) | `api/crash_test.go` |
|
||||
| `host_kernel_oops` | warning | a kernel oops this boot (taint D) — the box keeps running | — (once per boot) | `api/crash_test.go` |
|
||||
| `agent_behind` | warning | the box has run an agent OLDER than the vouched one for **7 days** (from when the hub first saw it behind; an unreadable version never counts; nothing vouched → nothing behind) — agents update only by a per-box signed job (R-530), so this is the "nobody signed for this box" alarm (hub v0.135.0) | the box reports the vouched agent (or newer) | `osupdates/r530_agent_alarm_test.go` |
|
||||
| `floor_raise_skipped` | warning | a GLOBAL controller floor was raised and one or more boxes keep their own LOWER per-customer floor, so the raise does not move them — ONE mail naming them all (R-604, hub v0.135.0) | — (one per raise) | `web/r604_floor_held_back_test.go` |
|
||||
|
||||
- **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report
|
||||
breaks a `not_running` run.
|
||||
@@ -334,9 +345,29 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
|
||||
protected-container check recreates it within 5 minutes, and the host reports every 15. So `tunnel_down` catches
|
||||
what the box cannot heal — a running container with no connection (wrong token, blocked network).
|
||||
- The OS alarms are checked **hourly**, re-sent at most **once a week** while true, and forgotten when false, so the
|
||||
next occurrence is announced again. The four numbers are configuration (`OS_ALARM_STALE_AFTER`,
|
||||
`OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`) — *decided by CC unattended,
|
||||
operator may reverse* (`11` §8.3).
|
||||
next occurrence is announced again. The numbers are configuration (`OS_ALARM_STALE_AFTER`,
|
||||
`OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`, `OS_ALARM_BUNDLE_BEHIND_AFTER`,
|
||||
`OS_ALARM_AGENT_BEHIND_AFTER`) — *decided by CC unattended, operator may reverse* (`11` §8.3; `09` decision 119).
|
||||
|
||||
---
|
||||
|
||||
## 6.4 A box that is not always on: the missed-backup deadline and the household's outage mail [DESIGN, hub v0.134.0, 2026-10-05]
|
||||
|
||||
Design home: `07` §6.1.1 (`09` decisions 109–110; CC decisions 115–116, *operator may reverse*).
|
||||
|
||||
- **R-872 — a box DOWN at the 05:00 deadline is judged, not skipped.** The check used to skip every customer whose node
|
||||
is `down` ("they already have staleness events"), so a box down at EVERY deadline — a laptop off at night — was
|
||||
never judged (measured 2026-10-05: `1 skipped (down)`). Now a down box is judged on longer lines:
|
||||
`expected_dbdump_missed` after **48 h** without a `db_dump_completed`, `expected_backup_missed` after **72 h**
|
||||
without a whole-guest backup in any retained host report — never for a box first seen less than 48 h ago. A box that
|
||||
died last night still raises only its staleness alarm. A `disabled` box is still skipped (R-321). Pinned by
|
||||
`monitor/r872_down_box_test.go`.
|
||||
- **R-873 — "your server cannot be reached" reaches the HOUSEHOLD at most once per 7 days** (`node_stale`,
|
||||
`node_down`, `host_stale`, `host_down`, read from the persisted notification log, so a hub restart does not reset
|
||||
it). The operator still gets every edge. The recovery mail stays paired with a down mail the household actually
|
||||
received (§6.2's pairing), so a held-back down mail also holds back its recovery. Pinned by
|
||||
`notify/r873_liveness_weekly_test.go`.
|
||||
- `backup_catchup_done` (info) is the box's "missed backup made now" line: recorded, never mailed.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -1,5 +1,14 @@
|
||||
# 09 — How an app update works, and what it is becoming
|
||||
|
||||
> **How to read this document.** Where a statement is marked, it is marked like this — the same wording as
|
||||
> `07-backup-architecture.md:11-17`, carried here on 2026-10-05 (R-376, the three documents written after the
|
||||
> 2026-08-22 pass):
|
||||
>
|
||||
> - **[DESIGN]** — a decision taken. Not derived from code; the code may not implement it yet.
|
||||
> - **[FACT]** — an observed property, carrying a `file:line`, a live command output or a citation.
|
||||
>
|
||||
> **An unmarked statement means "not yet classified", never "observed"** (R-376).
|
||||
|
||||
> **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.**
|
||||
> Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen
|
||||
> deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who
|
||||
@@ -620,6 +629,11 @@ R-636's louder repeated alarm.
|
||||
controller images are deleted by the same in-use rule as decision 53 — *operator ruling 2026-10-01 (R-745, option 3A).*
|
||||
**Why:** about 50 old controller versions sat on each demo box (R-745); a release is ~400 MB unpacked and several ship
|
||||
a day. Registry tags are never deleted by this.
|
||||
*Clarified 2026-10-05 (R-817), the ruling unchanged:* a swap records the image RUNNING when it starts
|
||||
(`felhom-agent internal/localapi/controllerswap.go:236-240`, `st.Previous`) and a failed swap writes exactly that
|
||||
image back (`:289`). After a good swap that image is "the one before" the running one — so „the one before it" here
|
||||
and R-745's „rolls back to the RUNNING image" name the same image, seen before and after the swap. The controller
|
||||
never hands the agent an older roll-back target.
|
||||
|
||||
### 2026-10-01 — decided by CC unattended, operator may reverse
|
||||
|
||||
@@ -815,6 +829,116 @@ its length, and both fixes cost something the household would notice — operato
|
||||
1.30.0, agent 0.142.0 and the re-made golden; it lacks only agent 0.142.1's wrapper fix (R-858). *Operator ruling
|
||||
2026-10-04 ~18:49.*
|
||||
|
||||
### 2026-10-05 (08:42) — three operator rulings (recorded before the work; the catch-up brief)
|
||||
|
||||
109. **R-871, option A: a missed night runs once when the box comes back.** A box that was off during its backup time
|
||||
makes the missed backups soon after it is switched on; updates that restart apps still wait for a real night.
|
||||
**Rejected:** B — "the box must stay on at night", with an alarm only. *Operator ruling 2026-10-05.*
|
||||
110. **The household is told on its dashboard** (the operator's idea): a banner the household can close, shown when a
|
||||
daily backup was missed — when the last backup was, that the box was off during the backup time, and a suggested
|
||||
different time. It never changes the time by itself. *Operator ruling 2026-10-05.*
|
||||
111. **Tester 2: read only.** If it is online, CC re-signs its agent update (decision 97, act 1) and reads it back.
|
||||
Nothing else. *Operator ruling 2026-10-05.*
|
||||
|
||||
### 2026-10-05 (day, catch-up brief) — decided by CC — operator may reverse
|
||||
|
||||
112. **When the catch-up runs (R-871).** Options: (a) at once after the box comes back — a box switched on and off
|
||||
again starts a dump each time; (b) 15 minutes later, re-checking first. **Chosen (b)**, as the brief proposed;
|
||||
a leg due within 30 minutes is left to its normal run. `07` §6.1.1.
|
||||
113. **What a host suspend does to the night (R-871).** Options: (a) only a controller start triggers the catch-up —
|
||||
a suspended laptop's 02:30 timer then fires hours late (Go timers run on CLOCK_MONOTONIC) and starts the
|
||||
app-update leg at noon; (b) skip any daily job that fires over 60 minutes late and trigger the catch-up from a
|
||||
resume watch. **Chosen (b).** Reasoned and unit-tested with injected clocks; not measured live.
|
||||
114. **The banner's line and its suggestion (decision 110).** Options for the line: 24 h (a slow night flickers it),
|
||||
26 h (one night + the chain's two hours, = the hub's `backupStaleAfter`), 48 h (two missed nights before a word).
|
||||
**Chosen 26 h.** Suggestion: the latest hour H with H, H+1, H+2 usually on (on ≥ 5 of the last 7 days); none
|
||||
when the current window is already usually on (a new time would not help).
|
||||
115. **R-872's lines for a box that is down at the deadline.** Options: (a) judge a down box like an up one — a box
|
||||
that died last night raises a backup alarm on top of its staleness alarm; (b) 48 h without a dump / 72 h
|
||||
without a whole-guest backup (the agent's own 48 h valve + a day). **Chosen (b).** hub v0.134.0, `08` §6.4.
|
||||
116. **R-873's rule.** Options: (a) household liveness mail only after 24 h down — a real outage reaches the household
|
||||
a day late; (b) at most once per 7 days to the household, the operator every edge. **Chosen (b)** (the brief's
|
||||
first example): the first outage of a week still reaches the household at once. hub v0.134.0, `08` §6.4.
|
||||
117. **R-874's first restore-test check after start.** Options: at start (a crash loop hammers a failing tier — the
|
||||
earned restraint), 30 minutes after start, or keep one interval (6 h — a short-session box never reaches it).
|
||||
**Chosen 30 minutes.** agent v0.145.0.
|
||||
118. **R-876's repair trigger.** Options: (a) always run `dpkg --configure -a` (R-845's speed lost on every clean
|
||||
pass); (b) read `--audit` AND the update journal in ONE `sh -c` call, repair when either shows something, and as
|
||||
a belt repair + retry once when apt itself says "dpkg was interrupted". **Chosen (b)**: a clean pass still costs
|
||||
one call (pinned by a test). agent v0.145.0.
|
||||
|
||||
### 2026-10-05 (afternoon) — decided by CC — operator may reverse (the hub-safety / R-861 brief)
|
||||
|
||||
119. **Boxes left behind (R-604, R-530).** Options for the agent alarm's wait: 3 days (a box off for a long weekend
|
||||
alarms), 7 days (one week, the same wait as the bundle-behind alarm, R-840), 14 days. **Chosen 7 days**
|
||||
(`OS_ALARM_AGENT_BEHIND_AFTER`), counted from when the hub first sees the box behind. A global floor raise names,
|
||||
in one operator mail, every box whose own LOWER floor it cannot move. hub v0.135.0, `05` §5.
|
||||
120. **How the Basic-auth operator path is protected from cross-site POSTs (R-135).** Options: (a) drop browser-usable
|
||||
Basic auth (CC's headless runs and every runbook POST break); (b) require a custom header on a cookie-less
|
||||
state change (a browser cannot add one cross-site without a CORS preflight the hub never answers; scripts add one
|
||||
flag). **Chosen (b)**, header `X-Felhom-Operator`. hub v0.135.0, `05` §16.1.
|
||||
121. **The console password's seal (R-133).** Options: (a) a second key and scheme for `host_recovery`; (b) the off-site
|
||||
seal and key already in force (R-821). **Chosen (b)** — the brief asked for the existing pattern, and one key is
|
||||
one custody question. Consequence named: a database backup needs this key off DooPlex too (R-173). `05` §16.2.
|
||||
122. **How the agent's admin commands are narrowed (R-861).** Options per group: (a) a root wrapper per group with its
|
||||
own argv; (b) exact sudo regex patterns for every varying argument + ONE content checker for every agent-written
|
||||
file root reads + fixed bundle files where the content never varies + the signed update verified as root by the
|
||||
existing `felhom-os-apply`. **Chosen (b)** — fewest new root programs, and the checker is pinned to the agent's
|
||||
own renderers by contract tests. Delivery order: signed `agent_update` first, then the bundle. agent v0.146.1,
|
||||
`03` §3.1.
|
||||
123. **How long a cut-off backup run is said on the page (R-519).** Options: until the next run of any kind (a failed
|
||||
run would clear the warning), until the next run that ends with every step OK, or until dismissed. **Chosen: until
|
||||
a run ends with every step OK.** controller v0.296.0.
|
||||
124. **How a bundle that adds a path reaches a box (R-880).** An installed `felhom-os-apply` refuses any path not in its
|
||||
own table (R16). Options: (a) copy the new wrapper onto each box by hand as root (does not scale, leaves the signed
|
||||
route); (b) a STEP bundle — the box's current bundle with only `felhom-os-apply` replaced, published as
|
||||
`<ver>-step1` — then the release's bundle, both by signed jobs. **Chosen (b)**, `felhom-agent/scripts/build-step-bundle.py`;
|
||||
delivered to demo-hp, demo-felhom and Tester 1 on 2026-10-05. `11` §5.4.2 rule unchanged.
|
||||
|
||||
### 2026-10-05 (14:05) — three operator rulings (recorded before the work; the hub-DB off-site brief)
|
||||
|
||||
125. **The hub database's off-site copy goes to ep0's backup server** (option A of the hub-safety STATUS decision):
|
||||
encrypted on DooPlex, pushed to a write-only namespace on ep0, restore-tested weekly, alarmed. Rejected: B, a separate
|
||||
Hetzner Storage Box account (more new parts to maintain). *Operator ruling 2026-10-05.* (R-173)
|
||||
126. **R-519's live test is approved:** one controller restart on scratch 9202 in the middle of a backup. *Operator ruling
|
||||
2026-10-05.*
|
||||
127. **The agent's three by-design abilities (`03` §3.1) stay for now**; revisited before the first paying customer.
|
||||
*Operator ruling 2026-10-05.* (R-861)
|
||||
|
||||
### 2026-10-05 (night) — decided by CC unattended — operator may reverse (the burn-down night)
|
||||
|
||||
131. **What the installer's 120 GiB local-lvm check is (R-130).** Options: (a) rename it a recommendation and say in the
|
||||
warning that the install continues — matches what `runbooks/day0-install.md` already documents („proceed only if you
|
||||
sized the grows deliberately"), costs nothing; (b) make it refuse — blocks small boxes that install and run today
|
||||
(a 75 GiB drill box) and changes a behaviour operators rely on, with no measured floor behind 120. **Chosen (a)**;
|
||||
reversible by renaming back. Whether a REAL floor exists is unmeasured. Installer 1.32.0 (`installer-v1.32.0`, not cut).
|
||||
|
||||
132. **A hub with no sealing key, asked to write a NEW box secret (R-879).** Options: (a) refuse — enrolment, customer
|
||||
creation, passphrase regeneration and a PBS mint fail until the key is back, existing boxes unaffected; (b) write it
|
||||
in plaintext — silently re-creates the readable secret the row removes. **Chosen (a)**: `05` §16.2 already says
|
||||
„no key → a save is refused" for every sealed column, and the production hub runs with the key. The box lookup is
|
||||
an UNKEYED hash so that authentication never depends on the key. Hub v0.138.0.
|
||||
133. **What removing a household's own off-site target does to its repository password (R-729, R-545).** Options: (a)
|
||||
shred it always — a hub-held package then hits the R-241 mint refusal, and history on the NAS loses its on-box key;
|
||||
(b) refuse the whole press while the hub holds a package — refused on practically every box, R-729 unsolved; (c)
|
||||
clear the target, SSH key and known-host line always; keep the password whenever anything could depend on it (a
|
||||
hub package, escrowed, a successful run, snapshots), delete it otherwise. **Chosen (c)**: R-241's rule (never drop
|
||||
a key a package protects) and „never the repository"; keeping a key is reversible, deleting it is not. Cost: on
|
||||
most boxes the 0600 password file stays with no target. `07`. Controller (next release).
|
||||
|
||||
### 2026-10-05 (~21:00) — three rulings, the reviewer's picks given to the operator (recorded before the work; the burn-down night)
|
||||
|
||||
128. **R-126 — a `.fab` export onto a network drive.** **Refuse an export WITHOUT a password to any network drive; with
|
||||
a password it is allowed.** Rejected: the row's other option, filtering network drives out of the export
|
||||
destination list. *Ruling 2026-10-05 ~21:00, the burn-down night brief §1.*
|
||||
129. **R-856 — a crash restart reached the household twice.** **After a crash boot, the controller's app mails wait the
|
||||
same grace period as after a normal start.** The hub's "restarted after an unexpected stop" line stays the one
|
||||
message about the crash. *Ruling 2026-10-05 ~21:00, the burn-down night brief §1.* `08` §5.
|
||||
130. **R-888, R-337, R-375 — closed.** R-888: the hub needs neither `stacks.deployed` nor `storage.decommissioned` today
|
||||
(no hub page reads them; the fields stay on the wire and stay allow-listed in `scripts/wire_contract_gate.py`).
|
||||
R-337: resolved by itself and never seen again. R-375: a note with no defect behind it (nothing in the product
|
||||
reads a PBS target's size). *Ruling 2026-10-05 ~21:00, the burn-down night brief §1.*
|
||||
|
||||
### 2026-10-05 (06:49) — four operator rulings (recorded before the work; the night-fixes brief)
|
||||
|
||||
100. **Tester 1's Cloudflare tokens, shown in the 2026-10-04 night session's output, are NOT rotated** (option B) —
|
||||
|
||||
@@ -1,5 +1,14 @@
|
||||
# 11 — Operating-system updates: the host, the guest and the Docker engine
|
||||
|
||||
> **How to read this document.** Where a statement is marked, it is marked like this — the same wording as
|
||||
> `07-backup-architecture.md:11-17`, carried here on 2026-10-05 (R-376, the three documents written after the
|
||||
> 2026-08-22 pass):
|
||||
>
|
||||
> - **[DESIGN]** — a decision taken. Not derived from code; the code may not implement it yet.
|
||||
> - **[FACT]** — an observed property, carrying a `file:line`, a live command output or a citation.
|
||||
>
|
||||
> **An unmarked statement means "not yet classified", never "observed"** (R-376).
|
||||
|
||||
> | | |
|
||||
> |---|---|
|
||||
> | **Status** | **NOT RATIFIED — a PROPOSAL with operator rulings, corrected by the 2026-10-04 spike (§7.1, C1–C12); §8 step 2 BUILT 2026-10-04 (§8.1).** Ratification is Viktor's review, not an editor's. |
|
||||
@@ -639,6 +648,19 @@ Agent v0.144.0 + v0.144.1 (wrapper and agent; `09` decisions 106–108). Evidenc
|
||||
(runbook `crash-guard.md`), then the next pass installed the 12. Operator mail `os_update_failed` (true); no
|
||||
household mail; the household's timeline showed "Controller elindult".
|
||||
|
||||
### 8.5 The self-repair after a power cut as BUILT (2026-10-05, agent v0.145.0) `[FACT]`
|
||||
|
||||
- **R-876 fixed.** The wrapper reads `dpkg --audit` AND dpkg's update journal (`/var/lib/dpkg/updates/`) in ONE call
|
||||
and repairs when either shows something; as a belt, when apt itself says "dpkg was interrupted", it repairs and
|
||||
retries once (`09` decision 118). A clean pass still costs one call (R-845's speed, pinned by a test).
|
||||
- **Proven live** (operator's go, demo-hp, `audits/catchup-2026-10-05/partD/`): 13 guest packages rolled back, crash
|
||||
at 07:56:03 UTC during dpkg's unpack, back by itself (new boot 07:56:41, guard armed, 1 unclean boot in the window);
|
||||
at boot `--audit` clean, journal 1 file, the same shape as §8.4. **The next pass, with nobody touching the box:**
|
||||
`REPAIR configured=0 journal=1`, then `DONE rc=0 upgraded=12`, healthy; the guest's package list equals the one
|
||||
before the rollback. No mail (the pass did not fail); the household's timeline: "Controller elindult".
|
||||
- **R-874.** The agent's restore-test due-check runs 30 minutes after start (then every interval), so a box with short
|
||||
power-on sessions is restore-tested (`09` decision 117).
|
||||
|
||||
## 9. Where the rest lives
|
||||
|
||||
- The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`).
|
||||
|
||||
@@ -0,0 +1,17 @@
|
||||
### CI fix red-proof: the previous ISO date form, BusyBox-only PATH (the Alpine runner's tools)
|
||||
AssertionError: 'integrity_check' not found in "date: invalid date '2026-10-05T15:18:27Z'\nfelhom-hub-db-backup: FAILED: cannot parse snapshot time 2026-10-05T15:18:27Z\n"
|
||||
felhom-hub-db-backup: FAILED: cannot parse snapshot time 2026-10-05T15:18:27Z
|
||||
AssertionError: 'stopped snapshotting' not found in "date: invalid date '2026-10-04T12:28:28Z'\nfelhom-hub-db-backup: FAILED: cannot parse snapshot time 2026-10-04T12:28:28Z\n"
|
||||
AssertionError: "bytes, the pod's file is" not found in "date: invalid date '2026-10-05T15:18:28Z'\nfelhom-hub-db-backup: FAILED: cannot parse snapshot time 2026-10-05T15:18:28Z\n"
|
||||
felhom-hub-db-backup: FAILED: cannot parse snapshot time 2026-10-05T15:18:28Z
|
||||
felhom-hub-db-backup: FAILED: cannot parse snapshot time 2026-10-05T15:18:28Z
|
||||
felhom-hub-db-backup: FAILED: cannot parse snapshot time 2026-10-05T15:18:28Z
|
||||
felhom-hub-db-backup: FAILED: cannot parse snapshot time 2026-10-05T15:18:28Z
|
||||
Ran 15 tests in 0.781s
|
||||
FAILED (failures=9)
|
||||
### the fix, BusyBox-only PATH
|
||||
Ran 15 tests in 3.328s
|
||||
OK
|
||||
### the fix, normal GNU PATH
|
||||
Ran 15 tests in 2.555s
|
||||
OK
|
||||
@@ -0,0 +1,317 @@
|
||||
{"id": "R-10", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/appbackup/dbdump.go:364 `if err := tmpFile.Sync(); err != nil {` then :390 `if err := os.Rename(tmpPath, finalPath); err != nil {` with no directory Sync after; the twin at controller/internal/backup/backup.go:948 `_ = dir.Sync()` does sync the dir. Origin: audits/CAMPAIGN-6E-2026-07-15.md:129.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/appbackup/dbdump.go"], "change": "After the os.Rename in DumpOne, open filepath.Dir(finalPath) and call a best-effort dir.Sync() (log at DEBUG on error), mirroring atomicPromoteTar in backup.go:948.", "test": "Unit test that DumpOne still produces the final file and leaves no .tmp; fsync itself is not observable in a unit test, so add a small syncDir seam and assert it is called with the dump directory (red-proof by removing the call).", "minutes": 30}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-25", "sev": "P4", "category": "Storage & devices", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/web/storage_handlers.go:153 `uuid := resolveEnrollUUID(ctx, agent, device)` still resolves by device PATH after format, then AssignDisk(uuid) at the next step; FormatResult (controller/internal/agentapi/client.go:384-395) carries DurableID only for the confirmation path, not the new fs UUID. Binding resolve+assign to the format's durable-id needs the agent to return the new fs identity (two repos).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-76", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKED", "evidence": "Behaviour is FileBrowser-image runtime behaviour (mode/setgid of UI-created folders), only observable on a live box. The image has changed since the finding: controller/internal/infra/infra.go:27 `FileBrowserImage = \"gtstef/filebrowser:1.5.6-stable\"` (finding was on 1.3.3). The comment at infra.go:207-208 still asserts `umask 002 so folders the customer creates here come out group-writable (2775 with the parent's setgid)`, which the row says was false on 1.3.3 — needs a re-measure on 1.5.6 before deciding. Note: the register row itself is truncated mid-sentence (\"does not say t\").", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-89", "sev": "P4", "category": "Business & legal", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No retention policy object in hub: `grep -rln -i 'retentionpolicy|retention_policy' felhom.eu/hub` returns nothing (felhom.eu@53d8131b). Commercial per-customer policy = money/product decision + new reconciler.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-91", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKED", "evidence": "Whether /srv/pbs-felhom still exists on ep0 is live-only (ep0 is protected; not touched). Source-side: CONTEXT.md:3656 still reads \"`/srv/pbs-felhom` is 13 G of dead weight on `/` awaiting R-91's go-ahead\". Extra fact found: documentation/runbooks/offsite-endpoint.md:24 still says the datastore `felhom-offsite` is at `/srv/pbs-felhom` and :119 `proxmox-backup-manager datastore create felhom-offsite /srv/pbs-felhom`, contradicting RUNBOOK-ep0-datastore-volume-2026-07-27.md:8 (moved to /mnt/pbs-datastore). The row's CONTEXT.md:1018 citation is stale (now :3656).", "dup_of": null, "unique_facts": "offsite-endpoint.md:24 and :119 still name /srv/pbs-felhom as the live datastore path (stale since the 2026-07-27 move to /mnt/pbs-datastore) — a doc fix independent of the deletion; CONTEXT citation moved from :1018 to :3656.", "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-92", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu@53d8131b hub/internal/web/pbsdr_box.go:57 and :64 `view.UsedStr = fmtBytesGB(snap.UsedBytes)`; hub/internal/web/offsite_box.go:54 `return fmt.Sprintf(\"%.1f GB\", float64(b)/float64(int64(1)<<30))` — still 0.1 GB granular. Note the row's own trigger (\"when retention becomes customer-visible\") has not fired.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["hub/internal/web/pbsdr_box.go", "hub/internal/web/offsite_box.go", "hub/internal/web/templates/offsite.html"], "change": "Add an exact-bytes value to the PBS DR view (e.g. UsedBytesExact rendered as a title= tooltip or a MB-precision string below 10 GB) without changing fmtBytesGB for other callers.", "test": "Table test on the view builder: two snapshots 50 MB apart render different exact strings; render test of offsite.html shows the exact value.", "minutes": 30}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-93", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "Premise gone per the row itself (R-461, CLOSED-ITEMS.md:636): drill-r50 VM no longer exists on either demo box. target-selection.md:111 keeps the fence with the note \"the VM does not exist anywhere, so the fence currently protects nothing\". Only remaining references are comments/tests (hub/internal/monitor/deadline_anchor_test.go:16, deadline_tiers.go:59).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "A row about choosing between two fixtures, neither of which exists any more.", "cost": "Building a new synthetic drift fixture is a design task (M), not a fix of this row.", "if_never": "Nothing breaks; there is no drift fixture either way. If one is wanted, it is a new row.", "pick": "close-as-accepted (operator word needed: close, or reopen as 'build a drift fixture'); also drop the dead drill-r50 fence in target-selection.md:111 at close"}, "minutes_spent": 4}
|
||||
{"id": "R-99", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No phantom-snapshot cleanup in felhom.eu/hub or felhom-agent (grep -i phantom finds only agent runner/test detection code; no removal path). Deletion on a customer datastore is a separate operator ruling per the row — customer data.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-104", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/backup/offbox.go:193-222 ClassifyOffsiteFailure has cases NoUnits/NoRepo/Transport and `default: return OffsiteFailUnknown` — no lock case, although offbox.go:826 already defines `var offboxLockRe = regexp.MustCompile(`repository is already locked`)`. The self-heal half is built (offbox.go:839-858 unlock --remove-all + retry once), as the row's 2026-08-22 note says.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/backup/offbox.go", "controller/internal/i18n/locales/hu.json", "controller/internal/i18n/locales/en.json", "controller/internal/backup/*_test.go"], "change": "Add an OffsiteFailLocked class matched by offboxLockRe in ClassifyOffsiteFailure (before transport) and a cause line in OffsiteFailureMessage telling the operator the repository is locked by an interrupted run and how it clears.", "test": "Table test: a restic 'repository is already locked' error classifies as Locked (red-proof: fails today as Unknown); i18n parity gate for the new key.", "minutes": 50}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-124", "sev": "P4", "category": "Backup & restore", "group": "NOT-WORTH-IT", "evidence": "Still true: felhom-agent@e06ed97 internal/hub/dr_recipe.go:61 `const PBSRootNamespace = \"root\"` and :276 `c.Namespace = PBSRootNamespace`; agent internal/pbs/report.go:24 `ns = \"root\"`. The comment at dr_recipe.go:57-58 already documents that PBS spells it \"\".", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "The disaster-recovery recipe writes the PBS root namespace as the word 'root', but PBS itself uses an empty name, so a pasted '--ns root' fails.", "cost": "Changing it alters a wire field read by the hub (cross-repo wire contract + recipe producers), for a case no customer has: every box writes a per-customer namespace.", "if_never": "An operator restoring a box with NO namespace line would get one failed command and have to drop --ns; no data risk. The constant's comment already warns.", "pick": "close-as-accepted"}, "minutes_spent": 4}
|
||||
{"id": "R-129", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "Docs still say no key: felhom.eu@53d8131b documentation/operations/nodes.md:110 `### Access — there is no baked SSH key` and :112 \"no operator public key is on this box\"; target-selection.md:111 still flags R-129 unresolved; MEMORY.md:34 says `ssh demo-hp`, NO KEY→G1. Whether the key works today is a live fact (not checked — no ssh in this pass).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["documentation/operations/nodes.md", "documentation/runbooks/target-selection.md"], "change": "After one read-only `ssh -o BatchMode=yes demo-hp true` (and reading root's authorized_keys comment to name the key), rewrite nodes.md 'Access' section to the measured truth and drop the R-129 caveat in target-selection.md:111-112 (also update the memory index line).", "test": "Positive control: the BatchMode ssh succeeds/fails as the doc now states; repo_gates.py doc gates pass.", "minutes": 30}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-134", "sev": "P4", "category": "Security & access", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu@53d8131b hub/internal/cloudflare/unblock.go:117 `for _, name := range []string{domain, parentDomain(domain)} {` and :136-141 parentDomain strips exactly one label (`strings.SplitN(domain, \".\", 2)`); controller strips progressively (controller/internal/cloudflare/zone.go per row).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["hub/internal/cloudflare/unblock.go", "hub/internal/cloudflare/unblock_test.go (new)"], "change": "Extract a pure zoneCandidates(domain) []string that yields the name and every parent down to two labels, and loop resolveZone over it (same order: most specific first).", "test": "Table test on zoneCandidates: 'a.b.felhom.eu' yields [a.b.felhom.eu b.felhom.eu felhom.eu] (red-proof: one-label version yields only two); no HTTP needed because apiBase is a const.", "minutes": 40}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-161", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "Automatic half exists: app-catalog-felhom.eu@917a779 .gitea/workflows/gates.yml:40 `run: cd ws/app-catalog-felhom.eu && python3 scripts/catalog_gates.py --fast`; .githooks/pre-push:84 runs the same. Only the runtime volume-persistence gate stays a manual periodic run, by operator ruling 2026-08-02.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "The runtime check that app data lands on a volume is run by hand, not on every push.", "cost": "Automating it means CI pulling and starting ~53+ app images per push — slow, and the row itself says such CI gets disabled.", "if_never": "A template that writes data outside its volume can ship until the next periodic run catches it; the static gates and pre-push still run.", "pick": "close-as-accepted (residual is a deliberate ruling; owner operator)"}, "minutes_spent": 3}
|
||||
{"id": "R-162", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "Gate is app-catalog-felhom.eu scripts/check-volume-persistence.py (`docker diff` at :41, :72); behaviour on a non-overlay driver is not reachable from source and the row says it fails closed. Status WATCHING, no defect.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "If Docker ever ran on a storage driver where `docker diff` does not work, the persistence gate would refuse to report and blame the prober instead of the driver.", "cost": "A driver probe + reworded message in the catalog script, ~1 h, for a driver nobody runs.", "if_never": "Nothing, unless a non-overlay driver ships; even then the gate fails closed (no false green).", "pick": "close-as-accepted"}, "minutes_spent": 3}
|
||||
{"id": "R-164", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Predicate still absent: felhom-controller controller/internal/appbackup/dbdump.go:544 still only WARNs `its accounts table has NO rows`; restore still replays dump + tar (internal/backup/restore_unit.go:114-118 hasReplayableDump). Blocked on a design (live-vs-dump per-table counts).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-169", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "Working-style ruling owed by operator; nothing in source to fix. Current nets per row: pre-push hooks + Gitea runner alarm (e.g. app-catalog .gitea/workflows/gates.yml:40).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "CI only reports after a push lands, because every repo pushes straight to main with no pull request.", "cost": "Making CI blocking needs branch protection plus a PR workflow for every change — a slower way of working for a one-operator project.", "if_never": "A `--no-verify` push can land broken code until the operator reads the CI alarm e-mail.", "pick": "close-as-accepted (row itself says decide only if the window ever costs something)"}, "minutes_spent": 2}
|
||||
{"id": "R-177", "sev": "P4", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller@7690c27 controller/cmd/controller/main.go:1546 `sched.Daily(\"fill-watch\", \"03:30\", func(ctx context.Context) error { return fillWatcher.Check() })`; internal/scheduler/scheduler.go:269 has GetJobs but grep finds no RunNow/Trigger method and no run-job route in internal/web. Needs a new operator-gated trigger endpoint (auth surface) — a new mechanism, solve together with R-279.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-184", "sev": "P4", "category": "Box system & updates", "group": "FIXED-BY-LATER-WORK", "evidence": "Fixed by felhom.eu b55fc17d \"hub v0.102.0 — refuse to vouch a version that cannot be installed (R-273)\" — exactly shape (b), validate at vouch time in the hub. felhom.eu/hub/internal/web/configs.go:1358 `res := s.gitea.PackageDownloadable(ctx, t.pkg, t.version, t.file)` and :1365 `s.logger.Printf(\"[WARN] artifact vouch REFUSED: %s package %s is NOT downloadable (R-287)\", ...)`; tag leg at :1343 TagServesFile; unreachable registry also refuses.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-194", "sev": "P4", "category": "Box system & updates", "group": "NOT-WORTH-IT", "evidence": "PVE behaviour, not our code; row states the self-repair already tolerates it (fires on the next probe after the cache clears). No source change to check.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "Proxmox caches permissions, so a removed storage grant can still read as present for seconds to minutes; our self-repair notices only after the cache expires.", "cost": "Adding a second signal (storage content listing) to the agent's grant probe is a new mechanism, needs live measurement on a box.", "if_never": "A lost grant is noticed up to ~16 min late; the repair still happens on its own.", "pick": "close-as-accepted"}, "minutes_spent": 2}
|
||||
{"id": "R-206", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "homelab-manifests@87dfc29 (/home/kisfenyo/git/homelab-manifests): no daemon.json template in homelab-ansible (grep finds only a comment at roles/node_housekeeping/templates/node-housekeeping.sh.j2:17 and homelab-ansible/CLAUDE.md:54). Part (b) was superseded by fc9fbb8 (\"correct the expired Docker rationale\"): the script now says at :13-20 do NOT add docker calls, the GC policy in daemon.json is the control point. DooPlex work — not unprompted.", "dup_of": null, "unique_facts": "Part (b) (prune in the role) is superseded by fc9fbb8 — the role now deliberately forbids docker calls and names daemon.json GC policy as the control; remaining scope is (a) template daemon.json in Ansible + (c) restart-and-verify.", "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-207", "sev": "P4", "category": "Process & tooling", "group": "FIXED-BY-LATER-WORK", "evidence": "Fixed by homelab-manifests fc9fbb8 \"node_housekeeping: guard DRY_RUN, correct the expired Docker rationale, pin container log rotation\". /home/kisfenyo/git/homelab-manifests/homelab-ansible/roles/node_housekeeping/templates/node-housekeeping.sh.j2:137 `if [[ \"${DRY_RUN}\" == \"1\" ]]; then` inside write_metrics, :138 logs \"file left untouched\".", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-208", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller@7690c27 controller/Dockerfile:12 `ARG VERSION=dev` and :13 `ARG GIT_COMMIT=unknown` sit above :19 `RUN go mod download || true`; felhom.eu@53d8131b hub/Dockerfile:3 `ARG VERSION=dev`, :4 `ARG BUILD_TIME=unknown` above :9 `RUN go mod download || true`.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller (and the identical one-line move in felhom.eu/hub — two repos, each trivial)", "files": ["controller/Dockerfile", "felhom.eu: hub/Dockerfile"], "change": "Move the ARG VERSION/GIT_COMMIT (controller) and ARG VERSION/BUILD_TIME (hub) declarations down to just above the final `go build` RUN.", "test": "Build twice with different --build-arg VERSION on a clean tree; the second build must show `RUN go mod download` CACHED; `--version` of the built binary still shows the passed version.", "minutes": 30}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-209a", "sev": "P4", "category": "Process & tooling", "group": "UNCHECKED", "evidence": "Live-only: whether DooPlex has rebooted and /var/log/felhom-store-postboot-check.log says PASS. Not read (DooPlex is Tier 2, operator ruled no reboot; this pass touches no machine). No source claim to check.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-210", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "Operator ruling owed; row records CC's view 'not worth doing for the space' (~27 GB reclaimable vs 199 GB free). Workspace CLAUDE.md also forbids `docker image prune -a` on DooPlex.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "193 old controller/hub images exist only on DooPlex and cannot be re-pulled; the question is whether to delete them.", "cost": "An operator decision plus a careful targeted delete on the production host; returns ~27 GB.", "if_never": "~27 GB stays used on a disk with ~199 GB free; old images remain as clutter (and as the only copies of very old builds).", "pick": "close-as-accepted"}, "minutes_spent": 2}
|
||||
{"id": "R-213", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Row is a not-started design (live-vs-backup comparison, then put-back flow), operator-owned; nothing in source to verify against.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-230", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Owed rulings, not code: (a) bulk-correction ruling on MEMORY.md staleness (MEMORY.md index still carries version literals, e.g. 'ctrl 0.224.0', 'hub 0.109.0'); (c) spec-as-failing-test pilot not started. (b) closed. Operator decision required.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-246", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu@53d8131b hub/internal/store/store.go:3248 `func (s *Store) MarkEscrowStale(hostID string) error {` still has no production caller (grep: only definition + comments at offsite.go:208,216); stale_at still read (store.go:3182 clears it). Ruling owed by operator: evidential setter or retire the column (folds R-248).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-256", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/i18n/locales/hu.json:1406 `\"flash.offbox.mgr_unavailable\": \"A mentéskezelő nem elérhető.\",` used at controller/internal/web/offbox_handlers.go:54 and :197; sibling :1407 mgr_unreachable used at offbox_handlers.go:582; en.json:1415-1416 same shape.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/i18n/locales/hu.json", "controller/internal/i18n/locales/en.json"], "change": "Rewrite flash.offbox.mgr_unavailable / mgr_unreachable in both languages to say the backup service is not running yet and give a route (try again in a few minutes; if it persists, contact support).", "test": "i18n parity/accent gates (controller_gates.py) pass; a handler test with nil backupMgr asserts the redirect carries the key (exists or add one). Owner is operator (copy) — needs a nod on the wording.", "minutes": 20}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-261", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu@53d8131b hub/internal/store/selfbind.go:111 `func (s *Store) CountSelfBindTokens(customerID string) (int, error) {`; only callers are hub/internal/web/customer_delete_test.go:511 and selfbind_automint_test.go:29 — no production caller.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["hub/internal/store/selfbind.go"], "change": "Reword the doc comment (selfbind.go:106-110) to say it is a test accessor and name the two tests that pin the auto-mint invariant (selfbind_automint_test.go, customer_delete_test.go) — or, if the operator prefers, add one post-mint production check that logs [WARN] when count != 1.", "test": "Comment-only option: existing tests stay green; production-check option: unit test that a pre-seeded extra token produces the WARN (red-proof).", "minutes": 20}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-263", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/settings/settings.go:1655 `// from every other. This is the ONLY writer of StoragePath.BackupTarget — registration must never set` while :1699 `s.StoragePaths[i].BackupTarget = false` (ClearBackupTarget) also writes it; :1684 is SetBackupTarget's write.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/settings/settings.go", "controller/internal/settings/backup_target_role_test.go"], "change": "Change the comment to 'the only writer that GRANTS the role' and add a source-scanning test that finds every `.BackupTarget =` assignment in non-test settings code and fails if any other than SetBackupTarget can assign a non-false value.", "test": "The new test passes today; red-proof by planting a temporary `BackupTarget = true` in another function and seeing it fail.", "minutes": 45}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-264", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu@53d8131b scripts/wire_contract_gate.py still allowlists the six with _R264: :242 selfupdate_pending, :246 selfupdate_pending_version, :255 restore_tests.mount_parity, :258 restore_tests.mount_inventory, :281 backup.last_db_dump, :282 backup.last_integrity_check. Each reader is a design per the row.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-266", "sev": "P4", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/report/builder.go:94 `{Mount: \"/\", Label: \"SSD\", TotalGB: sysInfo.DiskTotalGB, UsedGB: sysInfo.DiskUsedGB, Percent: sysInfo.DiskPercent},` — no disk_known on the storage entry; hub has no disk_known (grep empty). Two-repo wire change gated by wire_contract_gate.py.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-279", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No operator/hub path to start an off-site run: grep for offbox run triggers in felhom.eu/hub/internal finds nothing; the only run entry is the customer dashboard handler (felhom-controller controller/internal/web/offbox_handlers.go:270 `if !s.backupMgr.OffboxRunnable() {`). Needs a new operator-authenticated trigger — sibling of R-177, not a duplicate (different job).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-284", "sev": "P4", "category": "Apps & catalog", "group": "NOT-WORTH-IT", "evidence": "Not a defect: felhom-controller@7690c27 controller/internal/web/templates/deploy.html:622 `<div id=\"storage-space-warn\" class=\"form-hint\" style=\"color:var(--warn);display:none\">` (hidden by default) and :808 `warn.style.display = freePct < 20 ? 'block' : 'none';` — correct direction. `git log -S \"freePct < 20\"` shows it unchanged since 69698a8 (v0.10.0), and deploy.html at c732fe1 (main as of 2026-08-09) already had display:none at :589. The 2026-08-09 report read raw HTML without running scripts.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "A reported 'almost full' warning on an empty disk; the code shows the warning only below 20% free and hides it by default, so the report was a reading of unrendered HTML.", "cost": "Nothing to fix; a JS render test would need a browser harness the project does not have.", "if_never": "Nothing — the warning never showed on a 93%-free disk.", "pick": "close-as-accepted (close as not-a-defect)"}, "minutes_spent": 5}
|
||||
{"id": "R-285", "sev": "P4", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No maintenance/expected-downtime concept in hub: `grep -rln -i 'maintenance|expected_downtime|quiet_until|snooze' felhom.eu/hub/internal` returns nothing. New mechanism (M).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-286", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "Lesson (a) not written anywhere: grep -i 'different channel|same channel|independent channel' over documentation/runbooks/workspace-CLAUDE.md, felhom.eu/skills/*, .claude/rules/* returns nothing; felhom.eu/skills/felhom-evidence/SKILL.md:52 has the positive-control rule ('Plant the thing, find it...') but not the different-channel requirement. Part (b): RUNBOOK-hub-db-offsite-backup.md:101 copies a snapshot ($SNAP) not bare hub.db, so no bare-`cat hub.db` runbook found in runbooks/.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["skills/felhom-evidence/SKILL.md (or documentation/runbooks/workspace-CLAUDE.md standing rule 3)"], "change": "Add one paragraph: a positive control must come from a different channel than the measurement (different query path, snapshot, API or clock); give the 2026-08-09 stale-snapshot case as the example. Put it in ONE home (pointer elsewhere).", "test": "python3 scripts/check_skills.py and repo_gates.py pass; re-read the skill to confirm it loads.", "minutes": 25}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-287", "sev": "P4", "category": "Process & tooling", "group": "FIXED-BY-LATER-WORK", "evidence": "Deleter established 2026-08-10 (R-267 newest-10 prune, recorded in the row itself); CI fixed by felhom-agent 53d047a \"Two guards, one number: bound the published check to the retention it must live with\" (R-291). felhom-agent@e06ed97 scripts/check-published-versions.py:101 `RETENTION_FILE = os.path.join(os.path.dirname(os.path.abspath(__file__)), \"retention-policy.json\")`, :213 `keep = retention_kept()`; scripts/retention-policy.json:37 `\"generic_versions_kept\": 10,`. Follow-up row R-291 is open (OPEN-ITEMS.md:439).", "dup_of": null, "unique_facts": "scripts/retention-policy.json _comment lines ~15-17 still say no register row records a package prune and 'Container packages currently hold 19 each' — both withdrawn by R-287 (prune is in R-267; 19 was an unpaginated count, real 270/169). Move this stale-comment fix into R-291.", "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-288", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu@53d8131b documentation/architecture/00-capability-map.md is now 210 125 bytes / 30 937 words / 253 lines (`wc`), larger than the 134 642 bytes measured in the row; :38 still reads `*Verified 2026-07-16 against evidence corpus @ felhom.eu tip `4b18cc5``. Restructure is an M doc surgery, operator-owned.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-289", "sev": "P4", "category": "Process & tooling", "group": "FIXED-BY-LATER-WORK", "evidence": "R-182 was closed by felhom.eu ef6ac6fe (2026-08-22, register compression): documentation/backlog/CLOSED-ITEMS.md:474 `| **R-182** | ... | **CLOSED — SHIPPED** (controller v0.194.0 + hub v0.90.0/.1, 2026-08-03) |`. The residue (digest never seen delivering) was since observed: documentation/audits/DRILL-chaos-night-2026-09-17.md:181 `backup_run_failures` „1 of 12 apps failed to back up in this nightly run: nextcloud\" listed as an alarm that fired and was true.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-290", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Gate exists (felhom.eu scripts/check_stands.py) but the map itself still carries the claims: documentation/architecture/00-capability-map.md has 95 'PROVEN-LIVE' occurrences (grep -c); demoting 12 rows or writing walk documents is M and blocked on R-288 per the row.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-291", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom-agent/scripts/retention-policy.json still says the 10 is 'NOT a ruling anyone has been able to locate' and recorded_by: \"CC, from the registry's observed state; NOT from a located operator ruling\" — but R-287 (OPEN-ITEMS.md:424) records it IS the operator's newest-10 rule executed under R-267. The file also lists a reader 'documentation/runbooks/registry-retention.md (felhom.eu)' that does not exist (find under felhom.eu/documentation returns no such file). check-published-versions.py:101-104 reads the number. The deeper min_agent-floor bound still needs hub network (not small, recorded only).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-agent", "files": ["scripts/retention-policy.json"], "change": "Rewrite the _comment/recorded_by to cite the operator's newest-10 rule (R-267/R-287) instead of 'observed, not a ruling', and drop or correct the non-existent registry-retention.md reader. Keep the min_agent-floor note as the recorded better bound; then close R-291.", "test": "python3 scripts/agent_gates.py --fast (check-published-versions reads the file; JSON must still parse and generic_versions_kept stay 10)", "minutes": 20}, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-292", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "hub/internal/web/templates/configuration.html:55 still reads 'the Gitea sha lookup failed (version missing / Gitea unreachable) or the manually-entered sha is invalid'; configs.go:1375 and :1395 both redirect to flash=artifact_sha_invalid; resolveArtifactSHA returns only (sha, ok bool) so the cause is lost.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu (hub)", "files": ["hub/internal/web/configs.go", "hub/internal/web/templates/configuration.html"], "change": "Make resolveArtifactSHA return a reason (not-found / unreachable / bad manual sha) and redirect to three distinct flashes (reuse artifact_unverifiable for unreachable, add artifact_version_missing, keep artifact_sha_invalid for a bad typed sha incl. the wrapper sha at :1395).", "test": "Handler test per cause asserting the redirect flash, with a fake gitea returning 404 / network error and a malformed manual sha; red-proof by running against old code.", "minutes": 50}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-310", "sev": "P4", "category": "Install & onboarding", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/scripts/felhom-host-install.sh:3060 sets GOLDEN_CHECK_WHY=\"it is controller $ver, but the vouched golden is $ART_GOLDEN_VER\" and :3080-3081 die \"...${GOLDEN_CHECK_WHY}.\\n The vouched golden is ${ART_GOLDEN_VER:-<unknown>}.\" — duplicate stands. :984 still reads the vmid confirm from /dev/tty; no runbook (day0-install.md, RUNBOOK-byo/appliance-deployment.md mention --uninstall) names the pty requirement.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/felhom-host-install.sh", "documentation/runbooks/day0-install.md"], "change": "Drop the second 'The vouched golden is' sentence when GOLDEN_CHECK_WHY already names it (or drop the version from :3060); add one runbook line: --uninstall needs an interactive terminal; --force does not bypass the typed vmid confirm.", "test": "bash -n on the script + python3 scripts/repo_gates.py --fast (hostinstall gate); grep the die text renders the version once.", "minutes": 20}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-315", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu/scripts/wire_contract_gate.py:30-48 still documents the test as a repo-wide literal-tag search ('IT PROVES REACHABILITY OF A NAME'); ROOTS at :88 includes the R-311 escrow/retained root. No receiver-type field-by-field comparison exists. Fix requires resolving receiver mirror types — a new mechanism (M).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-325", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller/controller/scripts/retrieval_promise_gate.py:54 still has its own literal STEMS = [\"visszaállíthat\", ...]; felhom.eu/scripts/hub_copy_gate.py:193-205 still reads it as a drift check. controller_gates.py:48-53 already imports shared scripts from the felhom.eu sibling, so the pattern exists.", "dup_of": null, "unique_facts": "Two-repo sequencing: hub_copy_gate.py:203-205 returns 'drift' when it cannot find a STEMS list, so the felhom.eu drift check must be removed in step with (or tolerate) the controller change.", "small_fix": {"repo": "felhom-controller", "files": ["controller/scripts/retrieval_promise_gate.py"], "change": "Import RETRIEVAL_STEMS from ../felhom.eu/scripts/customer_copy_vocab.py (same sibling-path pattern as controller_gates.py:48) and delete the STEMS literal; absent sibling = INCONCLUSIVE exit 2. Follow-up (felhom.eu, separate commit): remove hub_copy_gate.py's drift check, which would then fail to find STEMS.", "test": "python3 controller/scripts/controller_gates.py --fast; red-proof: remove a stem from the shared list and see the controller gate's decoy convict.", "minutes": 40}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-327", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu/documentation/architecture/where-felhom-stands.yaml:126-130 still: id claim.code-naming, title \"The same word is used for two different secrets across three surfaces; the email points at a page a rebuilt machine does not show\", status: partial. Needs the operator's capability-map ruling first (dataset may not be raised on its own).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-331", "sev": "P4", "category": "Storage & devices", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller/controller/internal/agentapi/diskverdict.go:34 uncorrectableFailCount = 64; :33 comment still defers 'growth-rate detection once the box keeps history'. No growth-rate rule found. New mechanism (M).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-336", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No source change reduces the ep0 poll rate (pvestatd interval is Proxmox-side, not in our repos). Design question (does the hub need a 15-min fill reading) remains; acceptance needs an ep0 access-log measurement. Scaling item, not small.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-337", "sev": "P4", "category": "Monitoring & notifications", "group": "UNCHECKED", "evidence": "Live-only behaviour (WATCHING). From source: felhom-agent/internal/localapi/server.go:518 serves GET /backup/status and :1258 answers from s.pickLatestBackup (the in-memory store), which suggests collection cadence, but the refresh path after an out-of-schedule run was not established within the time box.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-345", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/hub/Makefile:21 'docker tag $(IMAGE):$(VERSION) $(IMAGE):latest' and :22 'docker push $(IMAGE):latest' still present; only commit touching the Makefile is 77b5a4ce (initial).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu (hub)", "files": ["hub/Makefile"], "change": "Delete lines 21-22 (or move them behind an explicitly named opt-in target with a comment). Whether a stale :latest already sits on the registry is a separate live check for a session allowed to query it.", "test": "make -n docker-push | grep -c latest == 0", "minutes": 10}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-346", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "grep for ActiveEnterTimestamp|ExecMainStartTimestamp|InactiveExitTimestamp across felhom.eu/scripts, hub, felhom-agent, felhom-controller/controller, homelab-manifests (*.sh/*.go/*.py) returns ZERO hits; the R-341 row (now in CLOSED-ITEMS.md:211) carries the lstart reasoning. The row's only remaining action (check for other systemd-timestamp anchors) is answered: none exist.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "A warning that a future reader might anchor an uptime slope on systemd's ActiveEnterTimestamp instead of the process start time.", "cost": "Nothing left to fix; the audit the row asked for finds no such use anywhere.", "if_never": "A future script could make the mistake; there is no current instance.", "pick": "close-as-accepted (audit done, zero instances)"}, "minutes_spent": 4}
|
||||
{"id": "R-348", "sev": "P4", "category": "Monitoring & notifications", "group": "STILL-TRUE-SMALL", "evidence": "felhom-agent/internal/backup/store.go:28 still reads '// Backups are unaffected — their freshness has a ground truth on the storage (R-84).' The hub-side pin the row asks for ALREADY EXISTS: felhom.eu/hub/internal/monitor/deadline_anchor_test.go:52 TestBackupFreshness_AgentRestartBlindWindow_NoAlarm, :295 TestNewestBackupEvidence_ReachesPastEmptyReports, :366 TestCheckBackupDeadlines_RestartBlindWindow_NoEvent; constant at hub/internal/monitor/deadline.go:36 backupEvidenceLookback.", "dup_of": null, "unique_facts": "Hub pin test already exists (deadline_anchor_test.go:52/:295/:366) — only the agent comment remains.", "small_fix": {"repo": "felhom-agent", "files": ["internal/backup/store.go"], "change": "Reword the comment: the backups list IS lost on restart and refills only when a backup runs; the hub's freshness VERDICT is unaffected because it looks back 7 days (hub monitor/deadline.go backupEvidenceLookback, pinned by deadline_anchor_test.go TestCheckBackupDeadlines_RestartBlindWindow_NoEvent).", "test": "Comment-only; go vet ./internal/backup; the hub tests named above already pin the consequence.", "minutes": 10}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-352", "sev": "P4", "category": "Storage & devices", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Placement half is an open operator ruling (SPEC-app-data-placement-2026-08-21.md, 'Viktor rules'); deploy route still has no server-side default: GetDefaultStoragePath has no caller in internal/stacks (grep returns only internal/api/router.go:1137 systemInfo). Point (2) is carried by R-368.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-364", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "No helper exists: ls felhom.eu/scripts shows nothing grep/accent-related, and no script mentions '0x80' or 'negative control'. The proposal is a self-contained tool.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/hu_grep.py (new)", "scripts/test_hu_grep.py (new)"], "change": "A helper that, for a pattern containing a byte >= 0x80, also runs an ASCII anchor (must hit) and a negative control (must miss) and refuses to print a zero unless both behave; documented in the felhom-evidence or ui-hungarian rule as the way to search Hungarian text.", "test": "Unit test with a temp file: accented string present -> count; transformed (e.g. octal-escaped) input -> REFUSED not zero; red-proof by disabling the anchor check.", "minutes": 60}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-365", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller/controller/internal/i18n/locales/hu.json:368 'A kérésed szerint a korábbi távoli mentéseidet <strong>{{.AbandonDate}}</strong> napján véglegesen töröljük (még ... nap)'; handlers.go:1164-1167 sets AbandonDate/DaysLeft with no overdue branch; backups_remote.html:189 renders it unconditionally.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/web/handlers.go", "controller/internal/web/templates/backups_remote.html", "controller/internal/i18n/locales/hu.json", "controller/internal/i18n/locales/en.json"], "change": "Set AbandonOverdue when DueAt is past and render a new key ('a törlés esedékes, a következő napi karbantartáskor lefut' / English twin) instead of the future-tense sentence.", "test": "Render test with a clock past DueAt asserting the overdue key and not 'töröljük'; i18n parity + copy gates via controller_gates.py --fast; search with ASCII fragments.", "minutes": 45}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-367", "sev": "P4", "category": "Backup & restore", "group": "NOT-WORTH-IT", "evidence": "Prune guard confirmed still present: felhom-controller/controller/internal/backup/backup.go:1438 'continue // GUARD: an undeployed app's last backup is still its restore point'. The stranded file itself is on demo-hp (Tier 0) and was not checked live.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "One old 312 KB paperless database dump on the demo-hp test box sits under the app's old folder name; nothing reads or deletes it.", "cost": "An operator ruling plus a live hand-move or delete on one demo box, and an explanation that the adopted dump would not match the volume copies.", "if_never": "A small file stays on a disposable demo box; no customer is affected and the fix for new dumps already shipped (R-355).", "pick": "close-as-accepted (or delete it the next time the box is reprovisioned)"}, "minutes_spent": 4}
|
||||
{"id": "R-368", "sev": "P4", "category": "Storage & devices", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller/controller/internal/settings/settings.go:572 'IsDefault bool `json:\"is_default,omitempty\"` // new apps use this by default' (line moved from 453); the default is honoured only in templates/deploy.html:614 '{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}'; no internal/stacks caller of GetDefaultStoragePath/IsDefault.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/settings/settings.go"], "change": "Reword the field comment: the deploy FORM pre-selects this path (templates/deploy.html:614); the deploy API applies no default when HDD_PATH is omitted. (The alternative — a server-side default — is a behaviour change and not small.)", "test": "Comment-only; go vet ./internal/settings. Optionally a template render test asserting the default path is 'selected'.", "minutes": 15}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-371", "sev": "P4", "category": "Monitoring & notifications", "group": "NOT-WORTH-IT", "evidence": "Controller notifier events (internal/notify/notifier.go pushEventMsg list) include db_dump_completed, crossdrive_completed (:933) but no off-site success event; hub has no offsite *_completed type except offbox_abandon_completed.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "The weekly off-site backup sends no 'done' event, while the two local tiers do. Failures and an 8-day staleness deadline are already alarmed.", "cost": "A new event type in the controller plus the hub allowlist and severity mapping (two repos), or a written decision that silence-on-success is intended.", "if_never": "The operator sees off-site success only through its absence of alarms and the backup card, as today.", "pick": "close-as-accepted with one line in 07-backup-architecture saying success is silent by design because failure and staleness are alarmed"}, "minutes_spent": 6}
|
||||
{"id": "R-372", "sev": "P4", "category": "Hub & operator", "group": "NOT-WORTH-IT", "evidence": "No hub or controller surface for 'never produced a copy' (grep 'never produced|NeverProduced' finds only an unrelated comment at felhom-controller/controller/internal/backup/backup.go:724). The underlying F-6E-1 was judged demo-data churn and the controller warns loudly.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "An optional idea from July: show 'this second-drive copy was never made because its source is missing' separately from 'last copy failed' in the operator screen.", "cost": "A design call on the operator view plus a new state carried controller->hub->UI.", "if_never": "The operator keeps seeing the existing loud warning; no data risk.", "pick": "close-as-accepted"}, "minutes_spent": 5}
|
||||
{"id": "R-373", "sev": "P4", "category": "Box system & updates", "group": "FIXED-BY-LATER-WORK", "evidence": "Premise (20G/50G two-volume mismatch, 'nothing sets SysDataGrowGB') was retired by agent v0.120.0 one-data-volume work, commit cd6e267 'v0.120.0 — one data volume (R-165...)'. felhom-agent/internal/reconcile/bringup.go:191 '// SysDataGrowGB is a COMPATIBILITY INPUT since agent v0.120.0 (R-165). There is no longer a second' and :437 'growGB := spec.DataVolGrowGB + spec.SysDataGrowGB'; installer passes it (felhom-host-install.sh:3140 '-sysdata-grow \"$SYSDATA_GROW\"') and records the sizing at :2041-2058. Note: cd6e267 predates the row's filing date; the row quoted an older audit.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-374", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "felhom.eu/documentation/audits/CAMPAIGN-12-class-sweep-2026-08-08.md:121-123 still says 'three of the 19 were called borderline and left unfiled' without naming them; the doc's only commit is b7fb2117 and no evidence file in audits/ lists the 19.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "A July audit says three borderline cases were left unfiled but never named them.", "cost": "Re-running the whole C1 refusal sweep to re-find 19 cases and guess which three were meant — hours, and the guess cannot be checked.", "if_never": "Nobody can re-judge those three; the refusal class has had later sweeps and gates.", "pick": "close-as-accepted, with one line in the audit saying the three are not recoverable"}, "minutes_spent": 5}
|
||||
{"id": "R-375", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKED", "evidence": "Requires a read-only check on ep0 (token's datastore audit permission); not verifiable from source and ssh is out of scope for this checker.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-376", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "Legend present ('not yet classified') in 00..06 and 10 of documentation/architecture/, but MISSING in the newer 08-alarm-ladder.md, 09-update-architecture.md and 11-os-updates.md (grep -ci 'not yet classified' = 0 each); 09 has zero [DESIGN]/[FACT] marks.", "dup_of": null, "unique_facts": "Three architecture docs added after the row (08, 09, 11) lack the legend.", "small_fix": {"repo": "felhom.eu", "files": ["documentation/architecture/08-alarm-ladder.md", "documentation/architecture/09-update-architecture.md", "documentation/architecture/11-os-updates.md"], "change": "Carry the same marker legend paragraph into the three documents written after the 2026-08-22 pass; then close the row, since 'mark as sessions touch them' is a standing practice already in the template, not a defect.", "test": "grep -ci 'not yet classified' returns >=1 in every numbered architecture doc; python3 scripts/repo_gates.py --fast.", "minutes": 15}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-377", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/CONTEXT.md:1537 '## Standing rulings' runs to EOF: 189,685 bytes, 0 '###' sub-headings, 153 bullets, 39 distinct S- ids; each ruling starts as a bold paragraph e.g. '**S-39 — \"WE DO NOT KNOW\" IS NEVER DRAWN AS \"FINE\"...'.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["CONTEXT.md"], "change": "Turn each ruling's opening bold line '**S-NN — TITLE (date ...).**' into a '### S-NN — TITLE (date)' heading, changing no other byte; no compression, no reordering.", "test": "Scripted transform; assert 39 '### S-' headings and that git diff --word-diff touches only those opening lines (body bytes identical); repo_gates --fast.", "minutes": 45}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-390", "sev": "P4", "category": "Process & tooling", "group": "FIXED-BY-LATER-WORK", "evidence": "Commit 2344589a ('... runbook pveam note'); felhom.eu/documentation/runbooks/RUNBOOK-manual-build.md:154 '2. Run **`pveam update` first** — the `virgin` snapshot's template INDEX is stale too, and a stale index fails as a bogus'.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-391", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "app-catalog-felhom.eu/scripts/catalog_gates.py has no SHARED_ / observations entry (grep 'SHARED_|observations' returns nothing); app-catalog-felhom.eu/CLAUDE.md has no 'observation' mention.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "app-catalog-felhom.eu", "files": ["CLAUDE.md"], "change": "Take the row's second option: state in CLAUDE.md that the catalog REPORT.md carries no observations section by convention (findings go straight to the register), so gate 11 is not needed here. The runner refactor (first option) is the bigger alternative.", "test": "python3 scripts/catalog_gates.py --fast (instructions gate checks CLAUDE.md length) and python3 ../felhom.eu/scripts/check_skills.py unaffected.", "minutes": 15}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-392", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "ls felhom.eu/documentation/architecture shows no agent-tooling/workflow document (00-11 are all product; plus _design-review, _hub-review, _recovery-inventory). Writing a new architecture document is more than an hour and needs the operator's view of the split.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-393", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "No decision-log skill exists (felhom.eu/skills has 9 skills, none mention 'decision log'). Since then .claude/rules/unprompted-work.md §2 requires every unattended decision be recorded in CONTEXT.md + the owning architecture doc and put FIRST in the morning note (§4).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "A proposed skill plus helper script to log every decision an unattended run makes.", "cost": "A storage convention, a helper, a skill and rotation rules — a small project with its own acceptance tests.", "if_never": "Unattended decisions are still recorded under the unprompted-work rule (CONTEXT.md + morning note); only minor in-run choices go unlogged.", "pick": "close-as-accepted (superseded in practice by unprompted-work.md §2/§4)"}, "minutes_spent": 4}
|
||||
{"id": "R-394", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "wc -l felhom.eu/skills/felhom-build-deploy/SKILL.md = 186 (was 179 at filing — grew); scripts/check_skills.py:45 GRANDFATHERED still holds the exemption. Trim needs a session that can verify the build/deploy commands it keeps.", "dup_of": null, "unique_facts": "The skill has grown from 179 to 186 lines since filing.", "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-402", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No hub Go/template reads last_integrity_ok/_depth (grep in hub/internal returns nothing); still allowlisted at felhom.eu/scripts/wire_contract_gate.py:163 and :220. Needs the operator's decision on what the screen says.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-416", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "Partly covered: scripts/register_shape_gate.py:120 'RULE 3 — duplicate: {rid} already has a row at line ...' (added in 462ab4a5, R-627) now refuses a duplicate id WITHIN OPEN-ITEMS.md only (REG path at :84 = OPEN-ITEMS.md). CLOSED-ITEMS.md has no within-file duplicate rule; closed_register_gate.py:53 still lists '4. A duplicate id WITHIN one register escapes.'", "dup_of": null, "unique_facts": "OPEN-ITEMS half is already fixed by register_shape_gate RULE 3 (462ab4a5); only CLOSED-ITEMS remains.", "small_fix": {"repo": "felhom.eu", "files": ["scripts/register_shape_gate.py or scripts/closed_register_gate.py", "scripts/test_gate_decoys.py"], "change": "Apply the existing RULE 3 duplicate check to CLOSED-ITEMS.md too (suffixed ids like R-88a/R-88b stay distinct), and update the closed_register_gate.py:53 hole list to point at it.", "test": "Planted-duplicate decoy in a temp CLOSED file must convict; suffixed-id control must pass; run against the live CLOSED-ITEMS.md first to see it is clean.", "minutes": 35}, "not_worth": null, "minutes_spent": 7}
|
||||
{"id": "R-418", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "It drifted AGAIN: felhom.eu/scripts/repo_gates.py docstring lists 1-14 (+9b) = 15 gates while GATES has 17 — 'script-tests' and 'decoy-coverage' are registered but not listed (python import: len(GATES)=17). No test compares the two.", "dup_of": null, "unique_facts": "Live drift right now: script-tests and decoy-coverage missing from the docstring.", "small_fix": {"repo": "felhom.eu", "files": ["scripts/repo_gates.py", "scripts/test_repo_gates.py"], "change": "Add the two missing gates to the docstring list and a test that parses the docstring's gate labels and asserts they equal [g[0] for g in GATES].", "test": "python3 -m unittest scripts/test_repo_gates.py; red-proof: the test fails on today's docstring before the two lines are added.", "minutes": 30}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-420", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "felhom.eu/scripts/repo_gates.py:106 tuple is (label, path, args, fast, exemptible) — still no 'blocking' field, as the row says; no gate there needs one.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "The felhom.eu gate runner cannot mark a gate as advisory-only; the controller runner can.", "cost": "Add a field to the runner when a need appears.", "if_never": "Nothing today; no advisory gate is wanted in that repo.", "pick": "close-as-accepted (add it with the first gate that needs it)"}, "minutes_spent": 3}
|
||||
{"id": "R-421", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Deliberate class row ('stays open as the place the next instance is recorded'); its open instances R-422..R-426 are still open in this batch. Not a fixable item by itself.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-422", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu/scripts/reuse_refs_check.py:41 PATH_RE still ends '\\.(?:go|py|html|css|yml|yaml|sh)\\b' — no .md. Widening it needs a false-positive walk across all four repos' REUSE.md/CLAUDE.md citations (the row says that pass is the work).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-423", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/scripts/site_gates.py:22-27 hardcoded PAGES list; website/ today holds exactly those 9 files, so nothing is missed today, but a new page is not scanned.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/site_gates.py", "scripts/test_gate_decoys.py", "scripts/decoy_coverage_gate.py"], "change": "Discover website/**/*.html and FAIL on any page not in PAGES (or glob and keep PAGES only as exceptions); flip the decoy test to expect conviction and drop the 'site' EXEMPT entry.", "test": "Decoy: a temp website/decoy-page.html must now convict; repo_gates --fast green on the real tree.", "minutes": 45}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-424", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "felhom.eu/scripts/one_register_gate.py:23 '...defect written under `idea` looks exactly like a proposal to this gate.' — hole declared in the gate itself.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "The roadmap gate cannot tell a real defect filed as an 'idea' from a genuine idea.", "cost": "No mechanical fix exists; telling a defect from a proposal is human judgement.", "if_never": "A mis-filed defect could hide on the roadmap until a person reads it.", "pick": "close-as-accepted (hole stays declared in the gate's docstring)"}, "minutes_spent": 3}
|
||||
{"id": "R-425", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller/controller/scripts/offbox_rename_gate.py:16-20 FILES = backups.html, offbox_handlers.go, offbox.go only; templates/backups_remote.html exists and is not scanned, and customer copy now lives in internal/i18n/locales/hu.json (20 'NAS' hits) which is not scanned either.", "dup_of": null, "unique_facts": "Since localisation, the customer copy the gate guards moved to locales/hu.json, which FILES does not include — the gate may now be largely hollow, not only narrow.", "small_fix": {"repo": "felhom-controller", "files": ["controller/scripts/offbox_rename_gate.py", "controller/scripts tests (decoy)"], "change": "Scan by pattern (templates/backups*.html, *offbox*.go) plus internal/i18n/locales/hu.json, or assert FILES against a discovered set so an unclassified file fails; drop its decoy-coverage exemption.", "test": "Decoy: 'NAS-mentés' in a temp backups_offbox_extra.html and in a hu.json value must convict; real tree passes (check the 20 existing NAS strings are allowed forms).", "minutes": 50}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-426", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu/scripts/decoy_coverage_gate.py EXEMPT now has 19 entries (loaded via python): hub-copy is gone, felhom-agent 'release-complete' is new; group (d) gates (hostinstall, wire-contract, due-checks, published, image-resolvable, volume-persistence) all still exempt.", "dup_of": null, "unique_facts": "Count is 19 now, not 20: hub-copy left the list; felhom-agent release-complete joined it.", "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-427", "sev": "P4", "category": "Process & tooling", "group": "FIXED-BY-LATER-WORK", "evidence": "Commit 71b8c8c6 (Backlog triage Part B: '... closed_register_gate RULE 3 refuses a finished row in OPEN-ITEMS'); felhom.eu/scripts/closed_register_gate.py:10 'RULE 3 — (2026-10-03) no row in `OPEN-ITEMS.md` may carry a CLOSED-family word (CLOSED, SHIPPED,'. Of the 12 named rows, R-385/387/341/378/405/88a/88b/123 are now only in CLOSED-ITEMS.md; R-190 and R-352 remain open (partly-closed, as the row predicted).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-437", "sev": "P4", "category": "Process & tooling", "group": "FIXED-BY-LATER-WORK", "evidence": "felhom.eu 71b8c8c6 'Backlog triage Part B: 125 finished rows + 20 id-less rows moved to CLOSED-ITEMS ... closed_register_gate RULE 3 refuses a finished row in OPEN-ITEMS (decoys, seen red) ... register 444 -> 325'. felhom.eu/scripts/closed_register_gate.py:10 'RULE 3 — (2026-10-03) no row in `OPEN-ITEMS.md` may carry a CLOSED-family word'. Gate run today: 'closed-register gate OK — no open work filed as closed, no id in both registers.' (456 closed / 336 open, 0 convicted).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-445", "sev": "P4", "category": "Hub & operator", "group": "NOT-WORTH-IT", "evidence": "Still true in mechanism: hub/internal/web/apps.go:102-106 computes SuggestedLimit from fleet P95 with no deployment-count/live check. But hub/internal/web/apps.go:67 'since := parsePeriod(period, 7*24*time.Hour)' — the detail page defaults to a 7-day window, so a throwaway's samples leave the default view within a week on their own; the 2026-09-01 sample is long outside it.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "The hub's per-app memory suggestion can be built from samples of an app that no longer runs anywhere (e.g. a 15-minute test install).", "cost": "An operator ruling plus a hub change (age-out or exclude zero-deployment apps) and a test, ~1-2 h.", "if_never": "An operator glancing at the app page within 7 days of a throwaway install could see a suggestion based on it; after 7 days the default view drops it. Nothing customer-facing.", "pick": "close-as-accepted"}, "minutes_spent": 5}
|
||||
{"id": "R-451", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller/controller/internal/report/types.go ContainerDetailReport still carries only Name/State/CPUPercent/MemoryMB (no image field). Ruled (09 §3 decision 18), build deferred until fleet grows; needs controller payload + hub denormalisation + fleet page across two repos.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-454", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "The five files named are now clean (gofmt -l controller/internal/web/ prints nothing), but `gofmt -l controller` in felhom-controller lists 12 OTHER files today: cmd/controller/main.go, internal/agentapi/diskverdict.go, internal/api/update_reason_test.go, internal/appbackup/namespace_root_test.go, internal/appbackup/r381_undo_naming_test.go, internal/backup/r669_applied_meta_test.go, internal/family/family.go, internal/infra/infra.go, internal/notify/r636_oom_storm_test.go, internal/quiesce/tiers_test.go, internal/stacks/delete.go, internal/stacks/life_records.go. `grep -rn gofmt --include=*.py` in felhom-controller: no hit — controller/scripts/controller_gates.py still has no formatting gate. The row's 'the count can only grow' is proven.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/scripts/controller_gates.py", "the 12 files listed in evidence"], "change": "Add a gofmt -l gate (fails on any listed file, INCONCLUSIVE if gofmt missing) to controller_gates.py, see it red on today's tree, then one gofmt -w formatting commit for the 12 files.", "test": "Run controller_gates.py before (red, lists 12) and after (green); go build ./... and go vet unchanged.", "minutes": 30}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-457", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "No later commit references R-457 beyond the filing release (felhom-controller 38d28b5 v0.234.0 / 998aa31 REPORT). No faked-clock CI run exists (grep for faketime/FAKE_NOW in felhom-controller: no hit). The six candidate files are named but unread.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["internal/backup/offbox_test.go", "internal/web/handler_export_upload_test.go", "internal/web/r103_tier2_action_test.go", "internal/web/dashboard_backup_card_test.go", "internal/web/async_restore_test.go", "internal/stacks/installed_test.go"], "change": "Read the six candidates; for each date literal that feeds an assertion evaluated against time.Now(), derive it from now (as the R-457 fix did). The faked-future-date CI instrument is a separate, larger idea and should be split out or dropped.", "test": "Run the touched packages' unit tests that do not reach Docker (or rely on CI); record per file 'literal feeds real-clock assertion: yes/no'.", "minutes": 60}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-460", "sev": "P4", "category": "App updates", "group": "NOT-WORTH-IT", "evidence": "Fact about the BookStack app (API token mintable only via web UI; secure cookies over plain http give 419), not a code defect; re-measured 2026-09-21 per row. DB half proven by app-catalog-felhom.eu/scripts/upgrade-test.py.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "BookStack's uploaded files cannot be checked automatically after an upgrade; only its database can.", "cost": "An upstream headless token route, or a browser-driven step DooPlex cannot run.", "if_never": "BookStack upgrades stay half-proven automatically; file survival rests on volume persistence and manual checks.", "pick": "close-as-accepted"}, "minutes_spent": 2}
|
||||
{"id": "R-464", "sev": "P4", "category": "App updates", "group": "FIXED-BY-LATER-WORK", "evidence": "Lesson homed and harness uses the correct probe. app-catalog-felhom.eu b7ef0c4 'upgrade-test.py: record the engine's own view of its datadir'; app-catalog-felhom.eu/scripts/upgrade-test.py:278 '\"mariadb-upgrade --check-if-upgrade-is-needed --user=root \"'. felhom.eu d6837d98 (SPIKE R-459); felhom.eu/documentation/architecture/09-update-architecture.md:1647 '1. **Ask the engine, not the log.** MariaDB's entrypoint prints `MariaDB upgrade not required` on an' (cites R-464).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-488", "sev": "P4", "category": "Process & tooling", "group": "UNCHECKED", "evidence": "The claim is a measured suite runtime (5.5 min); confirming it needs running go test ./internal/backup, which I did not run (read-only; backup tests may reach real docker on DooPlex). No commit after filing (felhom-controller 24d7c54) mentions R-488 or a test-speed change in internal/backup (git log --grep on internal/backup since 2026-09-13: empty), so it is likely still true.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-492", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "cfg.Paths.HDDPath still defined and read: controller/internal/config/config.go:167 'HDDPath string `yaml:\"hdd_path\"`', :453 envStr(\"FELHOM_PATHS_HDD_PATH\", &cfg.Paths.HDDPath); readers internal/report/builder.go:69, internal/monitor/healthcheck.go:79, internal/api/router.go:1135, internal/web/server.go:988, cmd/controller/main.go:358 and :526. Only 2 test references (1 file).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/config/config.go", "controller/internal/report/builder.go", "controller/internal/monitor/healthcheck.go", "controller/internal/api/router.go", "controller/internal/web/server.go", "controller/cmd/controller/main.go (gitignored dir — use git add -f)", "settings.AutoDiscoverStoragePaths signature"], "change": "Remove Paths.HDDPath, its env binding and each reader's dead global branch, keeping the per-app/discovered fallbacks each reader already uses.", "test": "go build ./... proves no reader remains; existing tests of the six readers stay green; a config test that FELHOM_PATHS_HDD_PATH is no longer read.", "minutes": 60}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-494", "sev": "P4", "category": "Install & onboarding", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu/hub/internal/cloudflare/ holds only unblock.go; no tunnel/DNS creation code in hub (grep cfd_tunnel: none). Building it is a new Cloudflare-API mechanism on the hub; operator ruled it non-blocking.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-501", "sev": "P4", "category": "Process & tooling", "group": "FIXED-BY-LATER-WORK", "evidence": "felhom.eu a4993272 'CLAUDE.md: the CI-check recipe was wrong in two ways, both measured today'. felhom.eu/CLAUDE.md:176 'rows — a run can sit several pages earlier. **Scan every page** and match on `head_sha`; with a'; recipe at CLAUDE.md:166-168 loops every page.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-502", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/scripts/iso/test/bootstrap-modes.sh exists; `grep -rn 'bootstrap-modes|bootstrap_modes' scripts/*.py .gitea/workflows` in felhom.eu returns nothing — no gate or CI runs it.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/repo_gates.py", "scripts/iso/test/bootstrap-modes.sh"], "change": "Register bootstrap-modes.sh in repo_gates.py behind a docker-available check that reports INCONCLUSIVE (never pass) when docker/the felhom-iso-assistant image is absent, plus a decoy (a broken banner must turn it red).", "test": "Run repo_gates.py with docker present (green), with the harness sabotaged (red), and with docker hidden from PATH (INCONCLUSIVE).", "minutes": 60}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-503", "sev": "P4", "category": "Install & onboarding", "group": "NOT-WORTH-IT", "evidence": "Not built (by design); the 2026-07-31 ruling 'a person chooses the disk' stands per the row; no later commit references R-503 other than its filing context.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "An idea, offered and not chosen: the installer would pick the disk itself when there is exactly one.", "cost": "A spike with three measurements and then reversing two operator rulings.", "if_never": "Nothing changes: a person keeps choosing the disk, as ruled.", "pick": "close-as-accepted (as DECLINED by ruling; the three measurements stay in the closed row for any future reversal)"}, "minutes_spent": 2}
|
||||
{"id": "R-504", "sev": "P4", "category": "Install & onboarding", "group": "UNCHECKED", "evidence": "The claim (iso.felhom.eu/ returns 404) is live-only; I may not curl hosts. felhom.eu/documentation/runbooks/VOLUNTEER-first-hour.md:14 still says '`iso.felhom.eu/` itself still has no index — R-504'. Fix needs a Cloudflare rule the operator owns. Cosmetic; households use felhom.eu/letoltes (website/letoltes.html exists).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-507", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Needs measuring QEMU input-send-event or a VNC client against a live VM on felhom-pve — a live-machine spike, not a source change. No later commit references R-507.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-525", "sev": "P4", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Row itself states it is a new unmeasured mechanism (forwardAuth / Quantum proxy auth) needing a scratch-guest spike.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-526", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu/hub/internal/tenantsync/client.go:110 '// Deprovision DESTROYS the customer's PBS namespace, all its backup groups, and its token — the'; no token-only op exists. Needs a new op on protected ep0 and an operator yes/no.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-527", "sev": "P4", "category": "Apps & catalog", "group": "NOT-WORTH-IT", "evidence": "Row's wording is partly WRONG: the flag IS read — controller/internal/stacks/deploy.go:401 'if field.LockedAfterDeploy {', :790-791 and :1350-1374 copy it into appCfg.LockedFields, and deploy.go:710 UpdateStackConfig refuses a locked key. But UpdateStackConfig (deploy.go:681) has NO non-test caller (grep: only its own definition/logs), and deploy.html:508-518 renders every auto field `readonly`, so the effect the row describes (every field read-only after install regardless of the flag) is true.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "A catalog flag that marks some settings 'locked after install' changes nothing visible: the page makes every setting read-only anyway, and the edit path that would honour the flag is never called.", "cost": "Either delete the flag from the catalog and controller (catalog-wide edit + parity fixtures), or build an edit-after-install feature (a product change).", "if_never": "Nothing a household meets: all settings stay read-only, which is the safe direction.", "pick": "close-as-accepted (with the corrected facts: flag read into LockedFields, enforced only by uncalled UpdateStackConfig)"}, "minutes_spent": 8}
|
||||
{"id": "R-532", "sev": "P4", "category": "Apps & catalog", "group": "NOT-WORTH-IT", "evidence": "app-catalog-felhom.eu/templates/vaultwarden/docker-compose.yml:30 '- SIGNUPS_ALLOWED=${SIGNUPS_ALLOWED:-false}' (image vaultwarden/server:1.36.0-alpine, :22). The /api/config field is upstream Vaultwarden behaviour; the server refusal (400) is live-only and not re-measured here.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "Vaultwarden's web page shows a sign-up form even though sign-ups are off; the server then refuses it.", "cost": "An upstream fix, or a catalog note; nothing Felhom can change in the server's answer.", "if_never": "A stranger who finds the page sees a form that fails; an invited household is not affected.", "pick": "close-as-accepted"}, "minutes_spent": 3}
|
||||
{"id": "R-541", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Row: needs a new copy/re-key/release mechanism and a design; no later commit references R-541. Far off per re-rank (0.3% full, one pool box).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-544", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/hub/internal/web/hosts.go:917 's.logger.Printf(\"[INFO] host deleted: %s (escrow deleted: %v)\", hostID, deleteEscrow)'.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["hub/internal/web/hosts.go", "hub/internal/web/hosts_delete_test.go"], "change": "Change the log line to state the effect, e.g. 'host deleted: %s (escrow custody demoted to retained)' when escrow existed, and drop the boolean name from the text.", "test": "A test in hosts_delete_test.go capturing the logger output on an acknowledged delete asserts 'demoted to retained' and the absence of 'escrow deleted'; red-proof against today's line.", "minutes": 30}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-551", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "Row: behaviour proven by five red-proofed ServeHTTP tests; live proof needs a box in the 'paused AND agent-connected' state, which only a fresh install or a local-API token on scratch guest 9202 (a live-machine change) gives. No later commit references R-551.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "The escrow 'waiting for the agent' screens are proven by tests but never seen on a real box in that state.", "cost": "Provisioning a local-API token on the scratch guest, or walking it during the next fresh install.", "if_never": "Small risk the live page differs from the tested page; the state lasts ~17 minutes after a bind.", "pick": "close-as-accepted (add 'walk R-546 readiness branches' to the next fresh-install checklist instead)"}, "minutes_spent": 2}
|
||||
{"id": "R-555", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/scripts/wire_contract_gate.py:451 'def receiver_tokens(repo_root):' tokenises whole files: line 482 'toks.update(TOKEN_RE.findall(fh.read()))' — no comment stripping.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/wire_contract_gate.py"], "change": "Strip Go // and /* */ comments and template {{/* */}} before TOKEN_RE in receiver_tokens; add a decoy 'tag named only in a receiver comment must convict'. Newly surfacing tags each become a finding (allowlist with reason or a row).", "test": "Run the gate: decoy red, real tree result reviewed; the 'language' allowlist entry still justified.", "minutes": 60}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-564", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller/controller/scripts/retrieval_promise_gate.py:54 'STEMS = [\"visszaállíthat\", \"visszaszerezhet\", \"visszahozhat\", \"visszanyit\"]' — joined forms only, no split-verb pattern.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/scripts/retrieval_promise_gate.py"], "change": "Add split-form Hungarian patterns (állíthatók? vissza, (hoz|szerez|nyit)\\w* vissza), register the Hungarian occurrences found (the seven already reviewed in English), and add a planted split-verb decoy.", "test": "Gate red on the decoy, green on the tree after registration; search with ASCII fragments plus positive/negative controls per the Hungarian-search rule.", "minutes": 60}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-567", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-SMALL", "evidence": "controller/internal/web/templates/layout.html:83 '{{$storageOpen := or (eq .Page \"storage\") (eq .Page \"storage-network\")}}' — storage_init/storage_attach not included; storage_handlers.go:351 'data := s.baseData(tmpl, title)' passes the template name as Page.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/web/templates/layout.html", "controller/internal/web/testdata/i18n_parity (re-capture storage_init/attach fixtures)"], "change": "Add (eq .Page \"storage_init\") (eq .Page \"storage_attach\") to $storageOpen and mark the Meghajtók link active for them.", "test": "Render test: GET /storage/init and /storage/attach carry 'nav-group is-open' and the active storage link; red before, green after; re-capture parity fixtures.", "minutes": 40}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-568", "sev": "P4", "category": "Storage & devices", "group": "STILL-TRUE-SMALL", "evidence": "controller/internal/web/disk_health.go:124-152 diskHealthRows appends rows in resp.Disks order; no sort in the file (grep 'sort.' in disk_health.go: none).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/web/disk_health.go", "controller/internal/web/disk_health_test.go"], "change": "Sort rows by diskKey(d) (durable id, falling back to name) before returning.", "test": "Unit test feeding the fake agent two disks in both orders and asserting one identical row order; red-proof by removing the sort.", "minutes": 30}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-569", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "controller/internal/api/router.go:742 'if strings.Contains(err.Error(), \"protected\") {', also :745, :1018, :1021, :1024 ('not deployed'/'still running'), :1102, :1105, :1108 ('not orphaned').", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/api/router.go", "controller/internal/stacks (sentinel errors)", "router tests"], "change": "Add KindErrorf-style sentinels in internal/stacks (protected, not found, not deployed/still running, not orphaned), a statusFor helper per handler family replacing the three Contains blocks.", "test": "One table test per family passing a REWORDED error message and asserting the same status code; red against today's string matching.", "minutes": 60}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-570", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Fallback still present: controller/internal/web/handlers.go:1050 'offboxStaleWarningMarker = \"nincs mentésre jelölt alkalmazás\"'; producer internal/backup/offbox.go:1168. Closing depends on a fleet condition (every box one off-site run on >=0.251.0) — a watch, not a fix.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-571", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "grep ClassifyOffsiteFailure|PageOnly|Inline in felhom.eu/documentation/architecture/07-backup-architecture.md and 02-controller-module-map.md: no hit. Classifier lives at felhom-controller/controller/internal/backup/offbox.go:193 'func ClassifyOffsiteFailure(err error) OffsiteFailureClass {'.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["documentation/architecture/07-backup-architecture.md", "documentation/architecture/02-controller-module-map.md"], "change": "Add a short section to 07 listing the six failure classes, what each means for the customer, and that restic/ssh signatures are external; add an alert-placement paragraph (inline under storage bars vs top banner) to 02.", "test": "Doc-only: grep the two docs for ClassifyOffsiteFailure/PageOnly afterwards; the repo's doc gates stay green.", "minutes": 40}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-574", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "controller/internal/web/handler_debug.go still carries 40 lines with accented Hungarian string literals (grep -cP count); last touched by 0c702f8 v0.279.0, not converted. Labelling ~40 literals page-copy vs payload, adding en/hu keys and parity fixtures exceeds an hour.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-576", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller/controller/scripts/i18n_go_parity.py has no call-site argument-count or '+'-adjacency check (grep verb/argument: only VERB_RE/strip_verbs for text equality, lines 76-78, 196-199). Parsing multi-line Go call arguments reliably from Python with decoys is likely >1 h.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-577", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Waiting on an operator decision (what the share feature promises a stranger); felhom.eu/documentation/architecture/10-localisation.md table row 'the two guest share pages, the catch-all | a stranger / nobody | **no globe** | — (R-577, the operator's)'.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-579", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Gate deliberately deferred until R-554 deletes the first-boot wizard; R-554 is still OPEN (OPEN-ITEMS.md:131). Versionless links remain in controller/internal/setup/templates/setup_*.html:8 '<link rel=\"stylesheet\" href=\"/static/style.css\">' (8 files).", "dup_of": null, "unique_facts": "NEW: controller/internal/web/templates/monitoring.html:255 '<script src=\"/static/chart.min.js\"></script>' also has no ?v= — outside the wizard, so the future gate must cover or allowlist it. Could be folded into R-554's closing work.", "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-588", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/documentation/runbooks/iso-release-gate.md:319-322 'Result recording' says only 'in the release report' — names no home. documentation/tests/ holds iso-release-1.27.0, 1.27.1, 1.29.0 only; no record dir for 1.28.0 or later ISOs.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["documentation/runbooks/iso-release-gate.md"], "change": "Name the single home documentation/tests/iso-release-<ver>-<date>/ in the Result-recording section, and add a pointer dir/README for 1.28.0 to its audit record (evidence-backup-promise-2026-09-16/phaseD-iso-gate.txt). Optional existence check left out.", "test": "Doc-only; the repo doc gates stay green.", "minutes": 20}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-591", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "The copy is deepCopyStack in controller/internal/stacks/manager.go (row says Copy()); it deep-copies AppConfig (:1137-1148), DeployFields (:1162), OptionalConfig (:1174), Integrations (:1186) and has no I18n line (grep I18n in manager.go: none). Meta.I18n defined at internal/stacks/metadata.go:94.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/stacks/manager.go", "a stacks unit test"], "change": "Deep-copy Meta.I18n in deepCopyStack (map plus nested values).", "test": "Test mutates the copy's I18n overlay and asserts the original is unchanged; red without the copy.", "minutes": 30}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-594", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-SMALL", "evidence": "app-catalog-felhom.eu/scripts/check-copy-i18n.py: grep ALLOWLIST/allowlist — no hit; no way to register a true occurrence.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "app-catalog-felhom.eu", "files": ["scripts/check-copy-i18n.py"], "change": "Add ALLOWLIST_EN of (app, path, reason); registered occurrences pass, unregistered convict, and any entry matching nothing is itself a failure.", "test": "Decoys: a registered occurrence passes, an unregistered one convicts, a stale entry convicts. Mind catalog CI has no PyYAML (degraded mode must still run).", "minutes": 60}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-599", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/hub/internal/web/hosts.go:898 'http.Error(w, \"Host is ONLINE — deletion is refused (a live agent would receive 401s permanently).\", http.StatusConflict)' — no last-report age or opening time. target-selection.md: no mention of the stale-threshold wait (grep 45 min|stale_threshold: none).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["hub/internal/web/hosts.go", "hub/internal/web/configs.go (config delete 409)", "documentation/runbooks/target-selection.md"], "change": "Make both 409 bodies say when the last report arrived and when deletion opens (last report + configured stale threshold), and name the wait in target-selection.md's drill section.", "test": "Handler test with a host last reported N minutes ago asserts the 409 body carries the minutes and the opening time computed from the configured threshold (not the 30m literal).", "minutes": 60}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-602", "sev": "P4", "category": "Process & tooling", "group": "FIXED-BY-LATER-WORK", "evidence": "felhom.eu e02bc038 'hub v0.119.0 — ... R-596/R-598 closed' added the finding; felhom.eu/documentation/architecture/10-localisation.md:809-812 'the `felhom_lang` cookie and got the **Hungarian** page for `en`. ... The cookie is the right instrument for the anonymous claim page and the **wrong**' and :509 '`langFor`'s order is fixed: `?lang=` → **the household's setting when a session exists**'. Only the optional pointer from the workspace live-validation rules is absent (grep ?lang=/felhom_lang in CLAUDE.md files and .claude/rules: none) — a 5-minute add if wanted.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-603", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "No gate or helper: grep for html.EscapeString(want|'|R-603 in felhom-controller/controller scripts+internal: none. controller/internal/i18n/locales/en.json has 27 lines containing an apostrophe today, so the trap is live for any Contains assertion on those.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["a test helper in internal/web (e.g. assertPageContains)", "optionally controller/scripts gate"], "change": "Add a test helper that compares against html.EscapeString(want) and use it in the render tests that assert English copy; the bundle gate with a 27-entry allowlist is the larger alternative.", "test": "A test asserting an apostrophe-bearing en.json value through the helper passes; the same via raw strings.Contains fails (red-proof).", "minutes": 45}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-605", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-SMALL", "evidence": "app-catalog-felhom.eu/scripts/catalog_gates.py:122 'VERDICT = {0: \"OK\", 1: \"FAILED\", 2: \"INCONCLUSIVE\"}' — harness refusal and per-app undetermined both exit 2 and print the same word.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "app-catalog-felhom.eu", "files": ["scripts/catalog_gates.py", "scripts/check-volume-persistence.py", "scripts/check-image-resolvable.py"], "change": "Give harness-level refusal its own exit code (e.g. 3 = REFUSED) in both scripts and map it to a distinct label in catalog_gates.py.", "test": "Decoys each way: a forced canary failure reads REFUSED, a per-app undetermined reads INCONCLUSIVE.", "minutes": 60}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-610", "sev": "P4", "category": "App updates", "group": "NOT-WORTH-IT", "evidence": "felhom-controller@7690c27 controller/internal/stacks/update.go:1453 `case UpdatePhaseStarting, UpdatePhaseVerifying:` - starting and verifying share ONE recovery arm, the one measured three times. No fault injector exists (grep for faultinject/failpoint in controller/ returns nothing).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "A power cut landing inside the sub-second `starting` phase has never been measured; all three live cuts landed in `verifying`, which runs the same recovery code.", "cost": "An in-process fault injector in the controller (new mechanism) plus a live drill on a scratch guest.", "if_never": "Nothing new is learned: the code path is the same branch already proven three times; the residual risk is a difference between the two phases that the source does not show.", "pick": "close-as-accepted"}, "minutes_spent": 5}
|
||||
{"id": "R-617", "sev": "P4", "category": "Process & tooling", "group": "FIXED-BY-LATER-WORK", "evidence": "felhom.eu@462ab4a5 (2026-09-22) documentation/architecture/09-update-architecture.md:1969 \"`POST /api/v1/repos/migrate` is the route that works (the project's Gitea tokens carry\" - continues at :1970 \"`write:repository` but not `write:user`, so `POST /user/repos` answers 403\"; recipe at :1979. The one-line note the row asked for exists (in the architecture doc rather than operations/). The optional operator-scoped token is a separate wish, not the defect.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-618", "sev": "P4", "category": "App updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Remaining open work = the controller-side idea (let `verifying` accept docker's own `healthy`). grep for docker health status in felhom-controller@7690c27 controller/internal/stacks/update.go returns nothing; no R-618 reference in Go source. It is an undecided design question (operator), not a defect; the three probe fixes (app-catalog@793c4fb) and the gate scripts/check-probe-matches-compose.py are done.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-619", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/api/router.go:395 `meta, appCfg, err := r.stackMgr.GetDeployFields(name)` then :402 `\"metadata\": meta,` - metadata served verbatim, no derivation of Required for password; deploy.go:366 \"// We never silently auto-generate — the user needs to know their password.\" still refuses.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": "controller/internal/api/router.go (getDeployFields), new test in controller/internal/api/", "change": "In getDeployFields, copy meta.DeployFields and set Required=true for every field with Type==\"password\" before writing the response (copy, do not mutate the shared metadata; the web deploy page uses GetDeployFields separately and is untouched).", "test": "Go unit test: a stack whose .felhom.yml has a type: password field with required: false; GET /api/stacks/<n>/deploy-fields must answer required:true for it and leave a secret field's required as declared; red-proof by running it before the change.", "minutes": 45}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-621", "sev": "P4", "category": "App updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Capture is done: felhom-controller@7690c27 controller/internal/stacks/update.go:1054 `outDir := filepath.Join(dir, \"hold-logs\", ts)`. The open part (show it on the app page's hold panel / logs fallback) is not built: grep for hold-logs/holdLogs in controller templates and handlers returns only update.go:1050-1082 and undo.go:535,550 (writers). Surfacing needs a page change with HU/EN copy and a design for which hold's log to show.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-624", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Row's own latest update: remaining class is vaultwarden (closed sign-up by design) and code-server; the open decision is whether the harness may hold an app's admin secret (operator). Not verifiable further from source; needs a decision, not a fix.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-644", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKED", "evidence": "About the live state of gokapi on scratch guest 9202 (crash-loop, deployed:true). Only the box shows it; no ssh allowed.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-652", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "app-catalog@917a779 templates/romm/.felhom.yml:249 still carries `\"memory_peak_pct\": 80.9` with `\"memory_tight\": true` and no `memory_basis: anon` (contrast paperless-ngx/.felhom.yml:249 `\"memory_basis\": \"anon\"`). Needs a live re-measure of romm plus an undecided cache-thrash rule.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-654", "sev": "P4", "category": "Apps & catalog", "group": "NOT-WORTH-IT", "evidence": "app-catalog@917a779 templates/opengist/docker-compose.yml:11 `image: ghcr.io/thomiceli/opengist:1.15`; grep for \"/-/\" in templates/opengist/.felhom.yml returns nothing - app_info does not mention the moved pages.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "opengist 1.15 moved its pages under /-/; an old /login bookmark answers 404. The front page redirects correctly.", "cost": "A copy decision by the operator plus a HU/EN app_info line.", "if_never": "A household with an old deep bookmark sees a 404 once and re-bookmarks from the front page.", "pick": "close-as-accepted"}, "minutes_spent": 3}
|
||||
{"id": "R-687", "sev": "P4", "category": "App updates", "group": "NOT-WORTH-IT", "evidence": "Gaps (1)-(3) are unit-tested only (as the row says); item (4) proven live 2026-09-30. Cosmetic text still as described: felhom-controller@7690c27 controller/internal/quiesce/quiesce.go:846 `return true, fmt.Sprintf(\"the automatic update leg is running (it starts no step after %s)\", backupwindow.FmtHHMM(mod1440(startMin+stopMin)))` - derived from the window, not a manual leg's own deadline.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "Three live proofs a scratch box cannot give (a 3-hour leg, a failing off-site leg, a files_may_change step without a whole copy) plus one log text that names the window's deadline instead of a manually started leg's deadline.", "cost": "Each live gap needs a special venue (off-site target, long leg); the log text is ~30 min but only matters for the manual debug chain.", "if_never": "The unit tests remain the proof; an operator reading the manual-chain log sees a deadline 20 minutes off.", "pick": "close-as-accepted"}, "minutes_spent": 5}
|
||||
{"id": "R-688", "sev": "P4", "category": "Hub & operator", "group": "NOT-WORTH-IT", "evidence": "felhom.eu hub/internal/web/customer_delete.go:130 \"// R-688 (v0.125.0): NO leg of the cascade calls Cloudflare. The dialog used to promise\" and :133 `cfManual := cloudflareManualRemoval(cfg)` - the dialog lists the manual removal; no Cloudflare leg exists.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "Deleting a customer does not remove their Cloudflare tunnel and DNS records; the dialog now says so and lists what to remove by hand.", "cost": "A new Cloudflare leg in the reset cascade (API calls, ids, failure handling, tests) - a new mechanism touching an external service.", "if_never": "The operator removes a tunnel and a few DNS records by hand at each customer delete, guided by the preview.", "pick": "close-as-accepted"}, "minutes_spent": 4}
|
||||
{"id": "R-691", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Open work = the Tier 2 (second-drive) path of Use/Load is not live-proven; needs a two-drive Tier-0 box (9202 has one drive). Live-only gap, not a source defect; not checkable from source.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-693", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "grep for memory_scales_with_limit / scales_with_limit across app-catalog-felhom.eu, felhom-controller/controller and felhom.eu/scripts returns nothing - no basis that tells growth-to-fill from pressure exists. Needs a design (new harness signal).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-705", "sev": "P4", "category": "Process & tooling", "group": "FIXED-BY-LATER-WORK", "evidence": "The remaining half (manual whole-guest backup) EXISTS and predates the row: felhom-controller bbed5af (v0.47.0, 2026-06-12) 'backups page — whole-guest backup visibility + manual trigger'. Live source @7690c27: controller/internal/web/backup_handlers.go:324 `case r.URL.Path == \"/api/guest-backup/trigger\" && r.Method == http.MethodPost:`; :340 `if err := s.backupTrigger.TriggerNow(); err != nil {`; quiesce.go:427 \"manual backup requested — quiescing now\" (bypasses due-ness, all tiers); wired cmd/controller/main.go:2099 and the page button backups.html:229. Agent side: felhom-agent internal/localapi/server.go:514 `mux.HandleFunc(\"POST /backup\", ...)`. The controller half was built v0.279.0 (night-chain, handler_debug.go:79). The row's 'no manual trigger' claim was not true at writing; it is not chained into night-chain, which the row did not require.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-707", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Open work = live proof that the gate OPENS for seerr, outline and rallly (needs a media server / e-mail on a test box). Live-only proof gap; the gating itself is in source (catalog 6faf432 per row).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-718", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller@7690c27 controller/internal/web/templates/app_info.html:112 `<p>{{T \"app_info.close_signup_text\"}}</p>` - the close card has no restart line, while the window card has one only at :132-133 `{{- if .SignupNative}}` / `{{T \"app_info.signup_window_restart\"}}`. en.json:2505 close_signup_text says nothing about a restart.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": "controller/internal/web/templates/app_info.html, controller/internal/i18n/locales/hu.json + en.json, the app-info handler (if SignupNative is not set when CloseSignupOffered)", "change": "Add a key app_info.close_signup_restart (\"The app restarts once for this.\" / HU twin) and render it under the close card when the app has an after_setup env switch (the same SignupNative fact); add the same sentence to the gate-open confirmation where after_setup.env exists.", "test": "Render test per branch: app with after_setup.env shows the line on the close card, app without it does not (red first); run the design-v2 copy/i18n gates.", "minutes": 60}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-719", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Built: felhom.eu hub/internal/web/selfbind.go:160 \"// R-719 (v0.126.0): „Új linket kérek\" on an expired or used link.\" Open part is the operator's review of the changed shape and a live mint+send proof (unit-proven only) - an operator decision, not a CC fix.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-725", "sev": "P4", "category": "Install & onboarding", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu hub/internal/i18n/locales/hu.json:79 `\"bind.invalid.body\": \"A hivatkozás 7 napig érvényes. Ha lejárt, kérj újat az ügyfélszolgálattól, ...\"` still sits above hu.json:87 `\"bind.resend.button\": \"Új linket kérek\"`; en.json:79 'ask support for a new one'. The console 'V' is the ✔ glyph in scripts/iso/felhom-bootstrap.sh:99 `printf ' Felhom — a doboz össze van kötve. ✔\\n\\n'` (golden scripts/iso/test/golden/bound.hu.txt:2). The English JSON to phone apps is left on purpose.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": "hub/internal/i18n/locales/hu.json, hub/internal/i18n/locales/en.json (bind.invalid.body); optionally scripts/iso/felhom-bootstrap.sh:99 + scripts/iso/test/golden/bound.hu.txt", "change": "Reword bind.invalid.body to point at the button below (e.g. 'Ha lejárt, kérj újat lent.' / 'If it has expired, ask for a new one below.'), keeping the operator alternative; optionally drop the ✔ glyph the console font renders as 'V' and update the golden.", "test": "Hub i18n parity/copy tests and a render test of the expired bind page asserting the new sentence and absence of 'ügyfélszolgálattól'; ISO golden test if the glyph is changed. Ships with the next hub release.", "minutes": 30}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-731", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "The shape-switch control lives only in audit tools: felhom.eu/documentation/audits/catalog-currency-2026-09-30/00-currency.py, 04-analyse.py; no standing currency script in app-catalog-felhom.eu/scripts or felhom.eu/scripts (ls/grep for currency/shape returns only golden_currency_gate.py, which is unrelated). Making it standing means promoting a registry-reading tool with tests - more than an hour.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-734", "sev": "P4", "category": "App updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "grep for '.immich' / hash ignore list in app-catalog-felhom.eu/scripts/*.py and templates/immich/.felhom.yml returns nothing - no exclusion exists. The row says the rule change needs an operator word; calibre-web shows the mark is sometimes right, so the rule needs design.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-739", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "app-catalog@917a779 templates/wanderer/docker-compose.yml:117 `image: getmeili/meilisearch:v1.36.0`; grep MEILI_UPGRADE_DB in the compose returns nothing. Remaining: the template switch, a fixture (PocketBase create refused) and a measured step on the bench - live work.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-759", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Five checklist rows of wger need live measurement on 9202 (2.5, 3.7, 6.3, 8.2, 9.1); not verifiable from source.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-760", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-SMALL", "evidence": "app-catalog@917a779 templates/vikunja/docker-compose.yml: service vikunja (image: vikunja/vikunja:2.6.0, line 12) has no `healthcheck:` key and no comment explaining why (whole file read).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "app-catalog-felhom.eu", "files": "templates/vikunja/docker-compose.yml", "change": "Read the vikunja 2.6.0 image config for a HEALTHCHECK/shell; if none and the image has no shell, add a comment saying why there is no compose healthcheck (like adventurelog-frontend's R-655 comment); otherwise add a healthcheck of the family the image supports (REUSE.md §2). No image: line moves, so no catalog_since.", "test": "scripts/onboarding_gaps.py row 4.1 no longer lists vikunja (or lists it as explained); catalog_gates.py --fast green; if a healthcheck is added, a deploy on 9202 must read healthy.", "minutes": 45}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-761", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "app-catalog@917a779 templates/paperless-ngx/.felhom.yml:22 `# Logo: {assets.base_url}/assets/{slug}-logo.webp` vs felhom-controller controller/internal/config/config.go:511 `return fmt.Sprintf(\"/static/assets/%s-logo.svg\", slug)` and :516 `-logo.png`.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "app-catalog-felhom.eu", "files": "templates/paperless-ngx/.felhom.yml (comment block lines 20-24)", "change": "Change the comment to name `{slug}-logo.svg` (preferred) and `{slug}-logo.png` (fallback), matching config.go AppLogoURL/AppLogoPNGURL; comment-only.", "test": "grep shows no '-logo.webp' in the template comment; catalog_gates.py --fast green (copy-freeze gates ignore comments).", "minutes": 10}, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-764", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "grep smtp/mail in app-catalog templates/wger/.felhom.yml and docker-compose.yml finds only first_steps text (.felhom.yml:74 'Add meg az email címedet a beállításokban'); no smtp_mapping. A mapping needs a live boot proof with mail off (REUSE.md §2) - more than an hour; wger is hidden.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-766", "sev": "P4", "category": "Apps & catalog", "group": "FIXED-BY-LATER-WORK", "evidence": "Hub releases after the assets push (felhom.eu 40f07429, 2026-10-01): hub v0.131.0 (2026-10-04) .. v0.136.0 (d4be9f6f, 2026-10-05). The build copies website assets: felhom.eu scripts/build-hub.sh:98 `cp \"${WEBSITE_ASSETS_DIR}\"/*-logo.svg \"${BUILD_DIR}/assets/\" 2>/dev/null || true`, and the hub build workspace /mnt/5_hdd/felhom.eu/build/felhom-hub/workspace/assets/ holds radicale-logo.svg + 3 screenshots (also karakeep, dawarich) dated Oct 5 14:28; hub/Dockerfile:27 `COPY assets/ /usr/share/felhom/assets-seed/`. Not checked: what a live box shows (no machine access).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-768", "sev": "P4", "category": "Apps & catalog", "group": "NOT-WORTH-IT", "evidence": "Fit-check decline recorded in felhom.eu/documentation/audits/new-apps-2026-10-01/FIT.md (per row); nothing in source to fix.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "Grimoire is not built because upstream rules out public exposure and ships no image for v1.x; the row only watches for that to change.", "cost": "A periodic re-read of upstream at each catalog campaign.", "if_never": "Nothing; Karakeep covers bookmarks. The re-read can live in the next catalog campaign's checklist instead of an open row.", "pick": "close-as-accepted"}, "minutes_spent": 1}
|
||||
{"id": "R-769", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "New-app idea waiting on an operator decision (a fork as new upstream, after R-767). No source to check.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-770", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "New-app idea waiting on the operator's go/no-go (CC recommends not building). No source to check.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-771", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "New-app idea waiting on the operator's go/no-go. No source to check.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-779", "sev": "P4", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Live proof gap on the real Cloudflare tunnel needing the operator's phone off wifi; not checkable from source.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-781", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "app-catalog@917a779 scripts/test_gate_decoys.py:498 `sh([\"git\", \"clone\", \"-q\", \"file://\" + ROOT, cat], cwd=ROOT)` clones the REAL records while the stand-in sibling holds only :502-505 `proof.txt`; real records cite sibling evidence (grep -c 'felhom.eu/documentation': onboarding/karakeep.md 45, dawarich.md 43, radicale.md 40). Last change to the file (96829d0, 2026-10-02) added family-gate cases only. Test not run here (it can reach a registry).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "app-catalog-felhom.eu", "files": "scripts/test_gate_decoys.py (onboarding_cases)", "change": "After the clone, delete the real onboarding/<app>.md records (all but _TEMPLATE.md and the exempt wger.md the cases use) from the scratch clone, so genuine cases judge only the records they build; alternative: point the stand-in sibling at the real felhom.eu documentation tree read-only.", "test": "Run scripts/test_gate_decoys.py: the four genuine onboarding cases go from FAIL to ok, all decoys still refused; the real check-onboarding gate stays green.", "minutes": 30}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-786", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "app-catalog@917a779 onboarding/sparkyfitness.md has 8 '| open' rows, incl. :18 0.5, :20 0.7, :28 1.6, :29 1.7, :58 5.4, :68 8.3 (plus 0.1 licence = R-784, 2.1 = R-807). Most need live measurement (runtime internet, phone sign-in, second memory watch).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-793", "sev": "P4", "category": "Business & legal", "group": "NOT-WORTH-IT", "evidence": "Watch item from felhom.eu/documentation/audits/licences-2026-10-02/TABLE.md; nothing wrong as the catalog runs them (row's own reading).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "Four apps ship enterprise/BUSL code that is off as Felhom runs them; the row reminds us never to enable EE features or the -enterprise meilisearch image.", "cost": "Keeping a standing open row; the real guard would be a line in the licence table / REUSE.md read on each major.", "if_never": "Nothing changes unless someone turns on an EE feature; the rule survives in the licence audit.", "pick": "close-as-accepted"}, "minutes_spent": 2}
|
||||
{"id": "R-794", "sev": "P4", "category": "Business & legal", "group": "STILL-TRUE-NOT-SMALL", "evidence": "app-catalog@917a779 seven redis 7 images: dawarich/docker-compose.yml:164 `image: redis:7.4-alpine`, docmost:86, immich:123, outline:88, nextcloud:103, paperless-ngx:125, romm:136 `image: redis:7-alpine`. Moving each needs a harness-proven ladder step (7 apps).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-796", "sev": "P4", "category": "Apps & catalog", "group": "NOT-WORTH-IT", "evidence": "app-catalog@917a779 templates/metube/.felhom.yml:19 `family_gate: true` with no family_gate_except (by design, per row).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "MeTube's browser/phone 'send to MeTube' helpers cannot pass the family gate; households paste links in the page.", "cost": "A per-member token the gate accepts on /add only - a new design.", "if_never": "Households use the page; the helpers stay unusable. Reopen if a household asks.", "pick": "close-as-accepted"}, "minutes_spent": 2}
|
||||
{"id": "R-797", "sev": "P4", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "app-catalog@917a779 scripts/check-family-gate.py:126 `print(\"family-gate: rule 3 NOT CHECKED — no felhom.eu sibling with a baked golden at %s\" % sibling)` and :141 ' — rule 3 (golden >= 0.287.0) NOT CHECKED here' - the gap is stated, and the pre-push hook (with sibling) checks it.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "CI's single-repo clone cannot check rule 3 of the family-gate gate; it says NOT CHECKED, and the pre-push hook checks it.", "cost": "Giving catalog CI a felhom.eu sibling checkout (credentials, workflow change).", "if_never": "A push that bypasses the hook (--no-verify is forbidden) could skip rule 3; CI stays honest about it.", "pick": "close-as-accepted"}, "minutes_spent": 2}
|
||||
{"id": "R-798", "sev": "P4", "category": "Apps & catalog", "group": "STILL-TRUE-SMALL", "evidence": "app-catalog@917a779 templates/grimmory/docker-compose.yml:29 ` - SWAGGER_ENABLED=false` still present.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "app-catalog-felhom.eu", "files": "templates/grimmory/docker-compose.yml", "change": "Remove the dead SWAGGER_ENABLED line (or rename to API_DOCS_ENABLED=false, which v3.5.0 reads and defaults to false). No image: line moves.", "test": "catalog_gates.py --fast green; grep shows no SWAGGER_ENABLED. Note: a compose change syncs to boxes running grimmory and recreates the container, so the row's 'on the next Grimmory step' timing is reasonable.", "minutes": 10}, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-799", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "app-catalog@917a779 scripts/upgrade_fixtures_box.py:2183 `data=json.dumps({\"url\": self.URL, \"quality\": \"best\", \"format\": \"any\", \"auto_start\": True}), method=\"POST\")` - no download_type.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "app-catalog-felhom.eu", "files": "scripts/upgrade_fixtures_box.py (class MeTube.seed)", "change": "Add \"download_type\": \"video\" to the POST /add body.", "test": "Python syntax/import check of the fixture module; real proof rides the next MeTube bench/box walk (POST /add 200).", "minutes": 10}, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-804", "sev": "P4", "category": "Apps & catalog", "group": "NOT-WORTH-IT", "evidence": "app-catalog@917a779 templates/plant-it/docker-compose.yml:12 `image: msdeluise/plant-it:0.10.0`; templates/plant-it/.felhom.yml:23-26 \"The compose below is deliberately LEFT AS-IS (it pins `msdeluise/plant-it:0.10.0`, a repository ... It is not worth fixing\" and `lifecycle: abandoned`.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "plant-it's image repository does not exist; the template is already abandoned and not installable, and no box runs it.", "cost": "An operator choice to hide the template entirely (small catalog edit).", "if_never": "Nothing; the template is already documented as not installable on purpose.", "pick": "close-as-accepted"}, "minutes_spent": 2}
|
||||
{"id": "R-805", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "app-catalog-felhom.eu scripts/check-volume-persistence.py:342 'if m[\"class\"] == \"named-declared\" and m.get(\"files\", 0) == 0:' — only named volumes judged empty; binds not. Last change 917a779 (R-788). The rule change itself is small, but it flips Grimmory/komga/paperless-ngx/radarr/sonarr to UNDETERMINED and needs a live re-sweep to regenerate the verdict tables; the row also asks for a decision.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-806", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "app-catalog-felhom.eu scripts/check-volume-persistence.py:586 'code = _sh(a + [f\"http://{ip}:{port}{path}\"], timeout=40)' — plain http always; scheme reading exists only in scripts/upgrade_boxport.py:33 LB_SCHEME_RE and scripts/upgrade-test.py:407. The gramps-web no-answer half is not checkable from source.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "app-catalog-felhom.eu", "files": ["scripts/check-volume-persistence.py", "scripts/test_check_volume_persistence.py"], "change": "Make routed_ports return the traefik loadbalancer.server.scheme (reuse upgrade_boxport LB_SCHEME_RE) and have the GET exercise use that scheme with curl -k for https. The gramps-web :5000 non-answer stays a separate live look (narrow the row to it).", "test": "Unit test: a compose with scheme=https label yields an https:// URL in the built curl argv; decoy without the label yields http://.", "minutes": 45}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-807", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Per-app upload seeds for 13 apps (claper, crafty-controller, dawarich, docmost, gramps-web, immich, outline, sparkyfitness, tandoor, vikunja, wger, wishlist, zipline) + plex/wanderer; each needs a live fixture run. Gate rule at scripts/check-volume-persistence.py:342 still makes empty declared volumes UNDETERMINED (917a779).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-814", "sev": "P4", "category": "Hub & operator", "group": "UNCHECKED", "evidence": "Live Hetzner console state (box 611421 status); not visible in source. Operator action only.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-815", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKED", "evidence": "First GC completion on felhom-offsite is PBS server-side live state; grep of documentation found no GC completion record (DIAG-backup-missed-2026-07-26.md:43 'prune/GC history NOT COLLECTED').", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-816", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Needs a live exercise of six failure classes on a scratch guest; no source change can close it.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-817", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu documentation/architecture/09-update-architecture.md:619 '56. **A box keeps the controller image it runs and the one before it** (the self-update's roll-back target)'. Source disagrees: felhom-agent internal/localapi/controllerswap.go:240 'st.Previous, st.Current = prev, prev' (prev = image running when the swap began); felhom-controller controller/internal/stacks/controller_image_retention.go:21-22 '...the self-update's own roll-back needs only the running image... \"The previous\" is kept for a hand roll-back.' The decision text is the wrong one.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["documentation/architecture/09-update-architecture.md"], "change": "Add a dated correction under decision 56: the swap rolls back to the image running when the swap began (controllerswap.go Swap/rollback); the kept previous image is for a hand roll-back. Do not rewrite the ruling itself.", "test": "Doc only: grep the corrected line; repo_gates.py --fast passes.", "minutes": 15}, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-818", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu hub/CHANGELOG.md:759 '## v0.109.0 — the Backup card told every operator that every customer had no backups (2026-08-30, R-331)' and :766 '(R-330 is a nightly false alarm ...; R-331 is a hub display ...)'; OPEN-ITEMS.md:268/274 have R-330/R-331 as Disk health Phase 2/3. Line numbers moved from the row's 526-538 to 759-771. No correction line present.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["hub/CHANGELOG.md"], "change": "Add a dated correction note under the v0.109.0 entry (and the hub CHANGELOG mentions at :695/:701) naming the real closed rows from CLOSED-ITEMS.md. Also present in felhom-controller/CHANGELOG.md:3107 and :3234 (R-330/R-331 for v0.224.0/v0.225.0) — a second repo; either note it there too or narrow the row.", "test": "Doc only: grep the correction; gates pass.", "minutes": 25}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-819", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "Ran python3 scripts/check_stands.py (read-only): still convicts e.g. 'fail.stolen-machine: register id R-281 is not in OPEN-ITEMS.md', 'fail.customer-self-restore: register id R-356 ...'. grep 'stands' in scripts/repo_gates.py and .gitea/workflows/gates.yml: no hit. Last commit on the script 6088afcb.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/check_stands.py", "scripts/repo_gates.py", "where-felhom-stands.yaml (or its source)"], "change": "Let rule 3 accept an id found in CLOSED-ITEMS.md (and check the stand's status agrees), fix the R-273/R-356 dangling ids, register the gate in repo_gates.py with a decoy.", "test": "Run check_stands.py green; a decoy stand citing a non-existent id must fail; test_repo_gates.py passes.", "minutes": 60}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-832", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Deferred roadmap item: a third-location copy of ep0 is money + operator decision (decision 71). Nothing in source to change.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-844", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Needs a household timeline on the controller, which does not exist (row: 'when the box gets a household timeline'). A new surface, not a fix.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-855", "sev": "P4", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu hub/cmd/hub/main.go:450 'logger.Printf(\"[INFO] osupdates: the Docker engine set is approved only by the operator, after %d healthy ring-0 night(s)\", osSvc.DockerNights)' prints the raw value; internal/osupdates/service.go:169 'negative means none'.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["hub/cmd/hub/main.go", "hub/internal/osupdates/service.go"], "change": "Add a small helper (e.g. osSvc.DockerNightsEffective()) mapping negative→0 and 0→2, and print that in the start log.", "test": "Unit test of the helper for -1, 0, 3; red-proof by printing the raw value.", "minutes": 20}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-856", "sev": "P4", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Row is marked an operator design question (crash-restart suppression for app mails); not a defect yet.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-857", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu scripts/golden_currency_gate.py:146 'EVIDENCE_RE = re.compile(r\"^golden-(\\d+)\\.(\\d+)\\.(\\d+)-\\d{4}-\\d{2}-\\d{2}$\")' and :237-238 'found.sort() / return found[-1]' — two dirs with the same version tie-break by name, not by bake time; a '-rebake' suffix dir (documentation/tests/golden-0.292.0-2026-10-04-rebake/) does not match the regex at all, so the re-bake is never read.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/golden_currency_gate.py", "scripts/test_golden_currency_gate.py", "documentation/runbooks/RUNBOOK-manual-build.md (re-vouch line)"], "change": "Have newest_baked accept an optional suffix after the date and, for equal versions, prefer the newest bake-log timestamp (or refuse two dirs for one version); add 're-vouch at once after a same-version re-bake' to the runbook.", "test": "Test with two tmp dirs of one version holding different GOLDEN_SHA256 lines: the newer sha must be reported; red-proof against today's code.", "minutes": 45}, "not_worth": null, "minutes_spent": 7}
|
||||
{"id": "R-878", "sev": "P4", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Next action is 'measure a large volume first' — a live measurement; the fix direction is a behaviour change to the catch-up.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-881", "sev": "P4", "category": "Install & onboarding", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu scripts/felhom-host-install.sh: grep 'felhom-priv-apply' → no hit anywhere in the installer (uninstall does not remove it); :1679 '+ guest-hook snippet under /var/lib/vz/snippets/ (agent-installed at runtime)' still says runtime-installed.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/felhom-host-install.sh", "CHANGELOG"], "change": "Add an rm -f of /usr/local/sbin/felhom-priv-apply to the uninstall step (tolerate-absent, like felhom-pbs-apply at :1161) and fix the :1679 comment; ships at the next installer tag.", "test": "bash -n; --uninstall --dry-run output lists the removal; installer gates pass.", "minutes": 25}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-884", "sev": "P4", "category": "Monitoring & notifications", "group": "UNCHECKED", "evidence": "Live ArgoCD diff on DooPlex (forbidden to touch here). Related: homelab-manifests mon-system/monitoring.yaml:399-426 is the same prometheus Deployment R-211 concerns.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-885", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "On main (b018ca90) scripts/repo_gates.py runs no test_*.py suite. NOTE: the felhom.eu working tree holds UNCOMMITTED work for exactly this row by another session: '?? scripts/script_tests_gate.py' (docstring 'every Python test suite under scripts/ runs on every push (R-885)'), '?? scripts/test_script_tests_gate.py', ' M scripts/repo_gates.py' (registers 'script-tests'). Not fixed until pushed.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/script_tests_gate.py", "scripts/test_script_tests_gate.py", "scripts/repo_gates.py"], "change": "Finish and push the in-progress script_tests_gate.py (walks scripts/ for test_*.py, exit-code verdict, nesting guard) registered in repo_gates.py; coordinate with the session that owns the dirty tree.", "test": "test_script_tests_gate.py decoys (a failing suite, no suite found); one CI run green.", "minutes": 30}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-30", "sev": "P3", "category": "Hub & operator", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Design change (presence from the Dir-2 long-poll instead of the report clock), size M; no commit with R-30 after aa9c08f0 (filing).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-31", "sev": "P3", "category": "Hub & operator", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu hub/internal/web/configs.go:1594 'd, err := s.offsite.ProvisionOffsite(ctx, cfg.CustomerID, in)' still in-request; :1584 detaches from the request context (mid-cancel fixed) but no async/status card.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-35", "sev": "P3", "category": "Box system & updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller controller/internal/report/config_refresh.go:65 'config-refresh: applied config_version=%d — self-restarting to load it'; sessions are in-memory only: controller/internal/web/auth.go:259-260 's.sessions[token] = &session{'. Hot-apply or persisted sessions is a design change with security weight.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-49", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "app-catalog-felhom.eu templates/immich/.felhom.yml:28-34 backup block has no cache/volume exclusion; row itself says a capture-set exclusion needs its own ruling (data-loss-shaped).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-50b", "sev": "P3", "category": "Box system & updates", "group": "FIXED-BY-LATER-WORK", "evidence": "Claim 'fetched via fetch_raw from raw/branch/main — no tag, no pin' no longer true: bee68484 (installer v1.23.0, R-110/R-183) pinned fetch_raw to the vouched agent tag — felhom.eu scripts/felhom-host-install.sh:533 '\"$GITEA_BASE/$GITEA_OWNER/$AGENT_REPO/raw/tag/v$ART_AGENT_VER/$path\" \\' (leg b). Leg (c)-like signed delivery: felhom-agent c9fa2e7 (R-840 config bundle) and configs/test_felhom_config_bundle.py:264 covers /usr/local/sbin/felhom-pbs-apply. Residual worth one line if kept: the 0440 sudoers drift visibility note.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 7}
|
||||
{"id": "R-78", "sev": "P3", "category": "Box system & updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "An owed operator decision + spike (local_api authority); not a code defect.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-79", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller controller/internal/monitor/healthcheck.go:100 'fmt.Sprintf(\"SSD disk usage critical: %.0f%%\"', :213 'Protected container not running: %s'; rendered raw at controller/internal/web/alerts.go:241 'Message: issue, // ON THE WIRE ... not ours to translate; slice 3'. Whole-surface, seam needs a spike.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-118", "sev": "P3", "category": "Storage & devices", "group": "STILL-TRUE-SMALL", "evidence": "felhom-agent internal/localapi/disks.go:401 'if total, used, okc := statfsCapacity(d.MountPath); okc {' — no device-presence guard on the union path; devicePresent exists (disks.go:985) and is used just above (:381) for BoundUnderParent.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-agent", "files": ["internal/localapi/disks.go", "internal/localapi/disks_device_presence_test.go"], "change": "Only statfs the mount when s.devicePresent(d.MountPath) is true (else leave capacity zero/unknown); put statfsCapacity behind a seam var for the test.", "test": "Test: absent device + statfs seam returning root's numbers → row has no capacity; present device → capacity set; red-proof by removing the guard. Ships in the next agent release.", "minutes": 50}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-121", "sev": "P3", "category": "Box system & updates", "group": "FIXED-BY-LATER-WORK", "evidence": "3d7a2761 (hub v0.135.0, R-530/R-604 'boxes left behind listed and alarmed'): felhom.eu hub/internal/osupdates/service.go:79 'EventAgentBehind = \"agent_behind\" // warning, operator' with :180 'AgentBehindAfter: a box runs an agent older than the vouched one this long → an operator alarm' (7 d window, the staleness window the row asked for).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-126", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller controller/internal/web/handler_export.go:377-386 storageDriveList() appends every s.settings.GetStoragePaths() entry with no IsNetwork() filter; the predicate exists at controller/internal/settings/settings.go:618 'func (p StoragePath) IsNetwork() bool { return p.Kind == StorageKindNetwork }'.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/web/handler_export.go", "controller/internal/web/handler_export_test.go"], "change": "Skip IsNetwork() paths in the export-destination list and refuse them in isValidDrivePath for the export POST (keep scanning for import if wanted, via a separate list).", "test": "Handler test with one drive + one network path: destination list has only the drive; export POST to the NAS path is refused; red-proof without the filter. Needs a controller release.", "minutes": 50}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-127", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Leg (a) still true: grep 'data_key: true' in app-catalog-felhom.eu templates → only adventurelog, dawarich, homebox, papra, sparkyfitness; n8n N8N_ENCRYPTION_KEY, wanderer POCKETBASE_ENCRYPTION_KEY, calcom CALENDSO_ENCRYPTION_KEY, bookstack APP_KEY unflagged (templates/n8n/.felhom.yml:38 etc.). Leg (b) (regenerated DB password vs restored PGDATA) needs a design choice. Leg (a) alone is a ~45-min catalog change + label/flag agreement gate and could be split out.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-130", "sev": "P3", "category": "Install & onboarding", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu scripts/felhom-host-install.sh:348 'HARD_MIN_LVM_GIB=120 # a useful appliance won't fit below this on local-lvm' and :1760 '... || log_warn \"local-lvm free ~${free_gib} GiB < hard min ${HARD_MIN_LVM_GIB} GiB\"' — warns only.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/felhom-host-install.sh"], "change": "Take the cheap honest branch: rename to RECOMMENDED_MIN_LVM_GIB and reword the warning to 'below the recommended …' (making it refuse would change install behaviour and needs a ruling).", "test": "bash -n; installer gates; grep shows no 'hard min' left. Ships at the next installer tag.", "minutes": 20}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-132", "sev": "P3", "category": "Security & access", "group": "UNCHECKED", "evidence": "Whether HUB_PW was rotated is out-of-band operator state; not visible in source.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-136", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu hub/internal/web/server.go:857 'Name: \"hub_session\",' and readers at server.go:806, :896, :918 and apps.go:351.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["hub/internal/web/server.go", "hub/internal/web/apps.go", "a server test"], "change": "Introduce a const sessionCookieName = \"__Host-hub_session\" and use it at all five sites (Path=/, Secure, no Domain already hold). Every operator logs in once more; plain-HTTP browser access stops (Basic auth unaffected).", "test": "Test: login response sets __Host-hub_session with Secure+Path=/ and no Domain; a request with the old name is not a session.", "minutes": 30}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-137", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller controller/internal/cloudflare/waf.go:18 'globalRuleDesc = \"[felhom-geo] Global\"', :21 'appRuleDescPrefix = \"[felhom-geo] app:\"' — still not namespaced. Two-repo M change.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-138", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller controller/internal/infra/infra.go:157-159 writes CF_DNS_API_TOKEN when d.CFAPIToken != \"\"; no shared-zone guard in hub (grep shared.zone: none). Needs a policy decision first; no shared zone exists today.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-179", "sev": "P3", "category": "Install & onboarding", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu scripts/felhom-host-install.sh uninstall section :1121-1157 handles felhom-shared-parent and umounts under /mnt/felhom-drives, but grep 'x2ddrives|automount' → no hit: the per-share mnt-felhom\\x2ddrives-*.mount/.automount units are never stopped or removed.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/felhom-host-install.sh"], "change": "In uninstall step 4c, before the umount loop: stop+disable every mnt-felhom\\x2ddrives-*.automount/.mount unit, rm their files from /etc/systemd/system, daemon-reload (tolerate-absent; still never umount -l/-f).", "test": "bash -n; --uninstall --dry-run on a fixture listing; a demo-box uninstall run is the real check at the next installer tag.", "minutes": 50}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-180", "sev": "P3", "category": "Install & onboarding", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu scripts/felhom-host-install.sh:1800 'if pvesm status --storage \"$ARCHIVE_STORAGE\" ...' checks existence only; PVE_STORAGES=(local local-lvm felhom-pbs) at :322; no ARCHIVE_STORAGE ∈ PVE_STORAGES assertion.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/felhom-host-install.sh"], "change": "In the same pre-flight block, die (byo) / die (appliance) if ARCHIVE_STORAGE is not in PVE_STORAGES, with a message naming --acl-storages.", "test": "bash -n; a dry-run with --archive-storage felhom-backup must die in pre-flight; with local passes.", "minutes": 25}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-190", "sev": "P3", "category": "Box system & updates", "group": "NOT-WORTH-IT", "evidence": "Mitigation still in source: felhom-agent cmd/felhom-agent/main.go:656 'store-grant: GRANT WAS MISSING AND HAS BEEN SELF-REPAIRED — investigate the loss (R-190)'. Mechanism unexplained since 2026-08-03.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "A storage permission vanished once on demo-felhom in August and nobody knows why. Since agent 0.124.1 the agent puts it back by itself and mails the operator when it happens.", "cost": "Finding the cause needs live experiments with guest rebuilds and permission caches on a Proxmox host — hours, maybe days, with no guaranteed answer.", "if_never": "The box repairs the permission within one check cycle and the operator gets one e-mail each time; a recurrence would be visible and would reopen the question.", "pick": "close-as-accepted"}, "minutes_spent": 4}
|
||||
{"id": "R-200", "sev": "P3", "category": "Backup & restore", "group": "FIXED-BY-LATER-WORK", "evidence": "The remaining half (customer-facing recovery-code form: yell → R form → preview) shipped as the recovery screen: felhom-controller 636c51e 'R-193: the recovery screen — unlocking, and only unlocking (v0.200.0)'; controller/internal/web/templates/recovery.html:82 '<form id=\"unlock-form\" method=\"POST\" action=\"/recovery/unlock\" autocomplete=\"off\">', routed at internal/web/server.go:602. Plumbing half was 1b1366b (v0.196.0). The 64-hex inject-password route (server.go:783) stays a deliberate DR fallback with no form.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-211", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "homelab-manifests (/home/kisfenyo/git/homelab-manifests @87dfc29) mon-system/monitoring.yaml:420 'image: prom/prometheus:v3.15.0', :426 '--web.enable-lifecycle'; grep 'reload|checksum/config' → none. The manifest edit is small, but it rolls the production Prometheus on DooPlex (operator territory) and the same Deployment is OutOfSync per R-884 — do the two together.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-231", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Owner operator; DooPlex /opt/backup/scripts remains host state. Partially touched by cea8502f (scripts/hub-db-backup versioned in felhom.eu, cites R-231) but that covers only the hub-DB push, not /opt/backup/scripts or the same-disk/no-off-site facts.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-235", "sev": "P3", "category": "Install & onboarding", "group": "FIXED-BY-LATER-WORK", "evidence": "felhom.eu c033b3b6 'ISO 1.28.0 source: the console stops showing the pairing code once bound (R-535)'. scripts/iso/felhom-bootstrap.sh:538: `print_bound_banner # R-535: replace the pairing code on the console with the truth`. Same defect already CLOSED twice in CLOSED-ITEMS.md as R-535 (line 232) and R-214 (line 200, 'proven on a fresh install').", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-240", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller/controller/internal/backup/offbox.go:1168: `warns = append(warns, \"Sikeres — nincs mentésre jelölt alkalmazás\")`; web/handlers.go:1068-1069 returns lastWarning verbatim when toggledCount < 1, so the customer still sees 'Sikeres'. Run now also records warnKind=OffboxWarnNoAppsSelected (R-553), so the page no longer depends on the sentence for new runs.", "dup_of": null, "unique_facts": "handlers.go:1064-1066 comment: R-557 must not TRANSLATE the producer until R-570 closes (legacy kind=='' fallback matches the lowercase substring 'nincs mentésre jelölt alkalmazás'). A Hungarian rewording that keeps that lowercase substring stays safe for legacy boxes.", "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/backup/offbox.go", "controller/internal/backup/offbox_test.go", "controller/internal/backup/offsite_diag_test.go", "controller/internal/backup/r553_offsite_quota_test.go", "controller/internal/web/r553_stale_note_test.go"], "change": "Replace the producer string at offbox.go:1168 with wording that drops 'Sikeres' but keeps the lowercase marker substring, e.g. 'Ez a futás semmit nem mentett: nincs mentésre jelölt alkalmazás'; update the tests that pin the literal.", "test": "go test ./internal/backup ./internal/web -run 'Offbox|R553|Diag' (unit, no docker); red-proof: a test asserting the warning does not start with 'Sikeres' fails on old code.", "minutes": 45}, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-242", "sev": "P3", "category": "Box system & updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu/scripts/golden_currency_gate.py:23-24: 'It does **NOT** check that the golden was **VOUCHED**, because the vouched version lives ONLY in the hub's `hub_settings` table'; :40 'That vouch half is STILL open after 2026-09-13'. Last gate commits 5ef0f52b/ae59c31a did not add a vouch check.", "dup_of": null, "unique_facts": "Only the vouch half remains; it needs a hub-reading check (design: shape (c) hub-side checker) and cannot be a --fast gate. Operator ruling 2026-09-13 (goldens weekly + waiver file) reduces the urgency; candidate for NOT-WORTH-IT if the operator accepts.", "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-244", "sev": "P3", "category": "Hub & operator", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu/hub: no cascade leg touches app_log_issues — only writers are internal/store/telemetry.go:176-209 (upsert), :537 `DELETE FROM app_log_issues WHERE last_seen < ?`, :556/:575 operator deletes by app/id. cmd/hub/main.go:1004: `if n, err := s.PruneStaleIssues(time.Now().Add(-30 * 24 * time.Hour))` (since a757bee0, hub v0.4.0).", "dup_of": null, "unique_facts": "Not in the row: a 30-day stale-issue prune has existed since hub v0.4.0, so ORPHAN rows (only torn-down customers) age out on their own once not seen for 30 days — the 'accumulates one venue at a time' claim holds only for the SHARED rows, which keep a deleted customer's id in affected_customers while a live customer keeps reporting. The remaining fix is JSON de-referencing in the cascade + red-proof.", "small_fix": null, "not_worth": null, "minutes_spent": 7}
|
||||
{"id": "R-250", "sev": "P3", "category": "Install & onboarding", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu/hub/internal/offsite/offsite.go:71-72: `var defaultScanBackoff = []time.Duration{2 * time.Second, 4 * time.Second, 8 * time.Second, 16 * time.Second, 30 * time.Second}`; scanner.go:40 dials plain \"tcp\" (no A-record preference); no R-250 commit.", "dup_of": null, "unique_facts": "Not in the row: cmd/hub/main.go:617 `WriteTimeout: 60 * time.Second` — the ~60 s ladder already meets the server write deadline, so lengthening the ladder inside the POST needs a per-request deadline lift (internal/api/wait.go:58 shows the pattern) or async provisioning; that is why it is not a one-line change.", "small_fix": null, "not_worth": null, "minutes_spent": 7}
|
||||
{"id": "R-251", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller/controller/internal/backup/offbox_inventory.go:102-108: loops `for _, tag := range sn.Tags` and adds every non-empty tag to `newest` — no filter for 'felhom-offbox'. The marker filter exists only on the restore list (internal/web/offsite_restore_list.go:37 `const offboxMarkerTag = \"felhom-offbox\"`), not on the recovery inventory (web/recovery_handlers.go:175 `data[\"InvApps\"] = inv.Apps`).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/backup/offbox_inventory.go", "controller/internal/backup/offbox_inventory_test.go"], "change": "In offsiteNewestPerTag skip the 'felhom-offbox' marker tag (move the constant into package backup and reuse it from web); consider whether '_shares' should render as a named row or be skipped.", "test": "Extend inventoryFixture (offbox_inventory_test.go:22) so snapshots carry tags [felhom-offbox, <app>]; assert OffsiteInventoryList returns only the app rows and one size call per app; red-proof on current code (marker row appears).", "minutes": 40}, "not_worth": null, "minutes_spent": 9}
|
||||
{"id": "R-255", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller: still no page-wide runtime secret-sentinel test — `grep -l sentinel internal/web/*_test.go` hits only edge_safe_status_test.go, i18n_parity_test.go, recovery_test.go; controller/scripts/secret_in_markup_gate.py remains the only all-template net. No commit cites R-255.", "dup_of": null, "unique_facts": "Not in the row: the i18n parity test (internal/web/i18n_parity_test.go, TestI18nParity:516, 133 fixtures in testdata/i18n_parity) now builds per-page render data for most pages — the per-page fixture scaffolding this row priced as the main cost largely exists; a secret-sentinel pass could reuse it, making this cheaper than estimated (still > 1 h).", "small_fix": null, "not_worth": null, "minutes_spent": 7}
|
||||
{"id": "R-257", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller/controller/internal/i18n/locales/hu.json:1409: `\"flash.offbox.not_orphaned\": \"Az offsite tároló nincs elárvult állapotban.\"`, used at internal/web/offbox_handlers.go:317 (row cited :270); also hu.json:1216 err.backup.az_offsite_tarolo_nincs_elarvult_allapotban from backup/offbox.go:385.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/i18n/locales/hu.json", "controller/internal/i18n/locales/en.json"], "change": "Rewrite flash.offbox.not_orphaned (hu + en) to say what the customer tried, that it does not apply now, and where to look, without 'offsite'/'elárvult', e.g. 'A távoli mentés rendben van, nincs mit félretenni. Ha gondod van vele, írj nekünk.'; wording sign-off from the operator (owner Viktor).", "test": "Run the i18n parity/key tests (go test ./internal/i18n ./internal/web -run I18n) and an ASCII grep for 'elarvult' in the flash value as negative control.", "minutes": 25}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-262", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/hub/internal/api/handler.go:727-729: '// hostBackup / hostRestoreTest mirror the agent's hub.Backup / hub.RestoreTest wire // contract field-for-field'; struct hostRestoreTest (handler.go:746-761) has no mount_parity/mount_inventory, while felhom-agent/internal/hub/report.go:492-493 `MountParity string json:\"mount_parity,omitempty\"` / `MountInventory []string`. grep finds no mount_parity anywhere in hub Go code.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu (hub)", "files": ["hub/internal/api/handler.go", "hub/internal/api/host_test.go"], "change": "One-repo fix: narrow the comment to say hostBackup is field-for-field and hostRestoreTest is a deliberate SUBSET (lists mount_parity/mount_inventory as not modelled), and add a hub test that names the two agent fields as known-unmodelled so a future addition must edit it. Adding the fields + fixture is a two-repo change (byte-identical golden) and stays a separate choice.", "test": "go test ./internal/api -run HostReport (hub, no docker).", "minutes": 30}, "not_worth": null, "minutes_spent": 7}
|
||||
{"id": "R-269", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-SMALL", "evidence": "felhom-agent/internal/localapi/tokenstore.go:173-176: `if vmid, ok := s.byHash[want]; ok { if subtle.ConstantTimeCompare(...) == 1 { return vmid, true } }` — a superseded token's hash is still a direct hit; reload happens only on a miss (:180). Last change to the file f31a76f (v0.63.0, the reload-on-miss itself); no R-269 commit.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-agent", "files": ["internal/localapi/tokenstore.go", "internal/localapi/tokenstore_test.go"], "change": "On a map hit, also stat the store and reload when its size differs from loadedSize before answering (one cheap stat per auth), so a rotated-out token is rejected without depending on an unrelated miss.", "test": "Add TestTokenStore_RotatedOutTokenRejectedFirst: mint A, re-mint B from a second TokenStore instance on the same file, look up A FIRST on the long-lived instance -> want false. Red on current code (row says red-proved).", "minutes": 45}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-270", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller/controller/internal/bootstrap/bootstrap.go:255: `if cfg == nil || cfg.LocalAPI.Endpoint != \"\" {` (fill-only, never refreshes); DetectEndpointDrift (bootstrap.go:369-399) compares only the endpoint. No R-270 commit.", "dup_of": null, "unique_facts": "Blocked in substance on R-78 (OPEN, OPEN-ITEMS.md:323) — the operator's authority ruling (auto-reconcile vs detect-only) decides which fix is right. The rotation recipe itself is documented in memory 'local-API token rotation needs TWO files + TWO restarts'.", "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-271", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller/controller/internal/channelhealth/checker.go:152: `if prev != \"\" && prev != \"up\" {` (unseeded->up is silent); checker.go:87 alert text still says '(re-bootstrap)'. No R-271 commit.", "dup_of": null, "unique_facts": "Fix needs a persisted last state or a hub-seeded state (a new mechanism), so not an under-an-hour change.", "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-274", "sev": "P3", "category": "Box system & updates", "group": "FIXED-BY-LATER-WORK", "evidence": "felhom.eu eb600872 'R-297: installer compares a local golden against the manifest before using it'. scripts/felhom-host-install.sh:3074: `if golden_local_matches_manifest \"$GOLDEN_VOLID\"; then` (digest vs ART_GOLDEN_SHA, else baked marker vs ART_GOLDEN_VER; otherwise ignores the local golden and fetches, or dies if the operator named it, :3080). Line numbers differ from the triage note (2855-2905).", "dup_of": null, "unique_facts": "The BYO-disclosure half ('say what the install REUSES') was not checked; if wanted, it is a separate small wording row.", "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-275", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/scripts/felhom-host-install.sh:1059: `for _cfgbak in \"${agent_cfg}\".bak*; do [[ -e \"$_cfgbak\" ]] && run rm -f \"$_cfgbak\"; done` — glob still misses agent.json.campaign8-before, .campaign9-prev, .pre-e-target-move, .pre-prunegate.bak; :814 still claims 'config (+ its .bak backups)'. No R-275 commit.", "dup_of": null, "unique_facts": "The uid-reuse half (new service account inheriting uid 999) is not covered by the small fix; purging by directory makes it moot for /etc/felhom-agent.", "small_fix": {"repo": "felhom.eu", "files": ["scripts/felhom-host-install.sh", "scripts/hostinstall_gates.py"], "change": "Purge every sibling `\"${agent_cfg}\".*` (and then rmdir/rm the config dir) instead of `.bak*`; fix the WIPED line wording.", "test": "Add a hostinstall_gates.py check (or a bash harness under a temp dir with FELHOM paths) that creates agent.json.campaign8-before and agent.json.pre-prunegate.bak and asserts run_uninstall --dry-run lists both; red on current glob.", "minutes": 40}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-276", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/scripts/felhom-host-install.sh: no reference to wg-felhom/wg-quick anywhere (grep 'wg-quick\\|wg-felhom' empty); _uninstall_statement (:807-845) lists neither in WIPED nor KEPT. The tunnel is agent-managed: felhom-agent/internal/wgtunnel/manager.go:31 `confDest = \"/etc/wireguard/wg-felhom.conf\"`, :35 `unit = \"wg-quick@wg-felhom\"`.", "dup_of": null, "unique_facts": "Hub-side peer deregistration (hub/internal/store/wg_operator.go) is a second repo; the small fix covers the host side and names the hub peer under KEPT.", "small_fix": {"repo": "felhom.eu", "files": ["scripts/felhom-host-install.sh", "scripts/hostinstall_gates.py"], "change": "In run_uninstall (full scope): stop+disable wg-quick@wg-felhom and remove /etc/wireguard/wg-felhom.conf via run(); add it to WIPED, and add a KEPT line 'the hub-side WireGuard peer registration — remove it in the operator UI'.", "test": "Static gate in hostinstall_gates.py asserting the uninstall path names wg-quick@wg-felhom and the statement mentions it; --dry-run output check.", "minutes": 45}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-277", "sev": "P3", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "Part (a) FIXED by felhom.eu f5c9411e 'R-331 (hub half): the Backup card reads `offsite`, not the dead `backup` fields (v0.109.0)' (hub/internal/web/backup_card.go uses fmtBytesAuto :135-142). Part (b) still true: hub/internal/web/offsite_box.go:54 `return fmt.Sprintf(\"%.1f GB\", float64(b)/float64(int64(1)<<30))` — a 162 KB repo renders 0.0 GB.", "dup_of": null, "unique_facts": "Part (c) (stale offsite_delivery_stuck event read as current state) was NOT verified; if it is still wanted it should become its own row.", "small_fix": {"repo": "felhom.eu (hub)", "files": ["hub/internal/web/offsite_box.go", "hub/internal/web/offsite_box_test.go"], "change": "For UsageStr use fmtBytesAuto when the usage is below 1 GB (keep GB for the quota and the bar), so a non-empty repo never renders as 0.0 GB.", "test": "Unit test: 162*1024 bytes -> not '0.0 GB'; 0 bytes and >1 GB unchanged (go test ./internal/web -run Offsite).", "minutes": 25}, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-282", "sev": "P3", "category": "Install & onboarding", "group": "FIXED-BY-LATER-WORK", "evidence": "felhom.eu 4d6ec7c 'hub v0.104.0: ... the hub half of the naming (R-295)' + controller v0.211.0 (R-295 CLOSED, CLOSED-ITEMS.md:207) + R-323 hub v0.105.0. hub/internal/notify/templates.go:204: '// R-295, HUB HALF (2026-08-13). ONE NAME PER SECRET, and it is „Beállító kód\".'; hub/internal/claim/engine.go:51 `EmailReenroll EmailKind = \"reenroll\"` (mail names the setup page a rebuilt box shows).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-283", "sev": "P3", "category": "Install & onboarding", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu/hub/internal/claim/engine.go:7: '// engine: a rotation bumps the generation (single active code) and NEVER clears claimed_at.'; ReissueForReenroll (engine.go:195-212) rotates and mails but leaves the claim set — the hub still shows the customer as claimed after a guest rebuild. No R-283 commit.", "dup_of": null, "unique_facts": "The mail half is now right (EmailReenroll names the setup page, R-295 hub), so the R-282 consequence is gone; what remains is hub/box claim-state reconciliation (design).", "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-298", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKED", "evidence": "The template gate is still there: felhom-controller/controller/internal/web/templates/storage.html:364 `if(d.role==='user-data'){` else protected (:368). BUT the agent no longer reclassifies a backup-target drive: felhom-agent/internal/localapi/disks.go:1230-1236 ('WHY THIS IS NOT A ROLE RECLASSIFICATION ... the drive that now holds the whole-guest archives is ALSO the enrolled user-data drive') and disks.go:219/227 report Role from RoleForStorage plus a separate BackupTarget flag (958e54f, agent v0.112.0). storage/role.go:176-186 makes a local-dir with its own non-system device 'user-data'.", "dup_of": null, "unique_facts": "From source, a drive that is both user-data and the backup target should arrive as role 'user-data' and therefore be registrable; the row's 2026-08-10 observation contradicts that. Only a live /api/disks read on a box with that layout settles it.", "small_fix": null, "not_worth": null, "minutes_spent": 10}
|
||||
{"id": "R-306", "sev": "P3", "category": "Install & onboarding", "group": "STILL-TRUE-SMALL", "evidence": "felhom.eu/scripts/felhom-host-install.sh:418: `$DRY_RUN && return 0` (only DRY_RUN short-circuits _state_put); :1851/:1854 `_state_put dnsmasq_preexisting yes|no` run unguarded in preflight, while :226 says '--preflight-only: ... no state writes' and :1956-1961 guards only customer_id/mode. Blockers R-300 and R-305 are CLOSED.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/felhom-host-install.sh", "scripts/hostinstall_gates.py"], "change": "Make _state_put (and _state_mark) return 0 when PREFLIGHT_ONLY is true, so a preflight-only run writes nothing; the real run's preflight records ownership again.", "test": "Bash harness with STATE_DIR in a temp dir: run the preflight state block with PREFLIGHT_ONLY=true -> no state.json; red on current code. Or a static gate asserting _state_put checks PREFLIGHT_ONLY.", "minutes": 30}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-314", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller: StopAbandon is called only from cmd/controller/main.go:244 (CLI) — `grep -rn StopAbandon` shows internal/backup/offbox_abandon.go:374 and that CLI call, no web handler.", "dup_of": null, "unique_facts": "An 'operator-authenticated POST' on the controller needs an operator auth path the controller does not have today (or a hub-relayed command) — a new mechanism.", "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-317", "sev": "P3", "category": "Install & onboarding", "group": "STILL-TRUE-SMALL", "evidence": "felhom-agent/internal/lanresolver/lanresolver.go:107: `if _, err := os.Stat(\"/usr/sbin/dnsmasq\"); err != nil { // metadata read, no privilege needed` — still probes the dnsmasq-base file. No R-317 commit.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-agent", "files": ["internal/lanresolver/lanresolver.go", "internal/lanresolver/lanresolver_test.go"], "change": "Probe the dnsmasq unit (e.g. /usr/lib/systemd/system/dnsmasq.service or /lib/systemd/system/dnsmasq.service) behind a small stat seam instead of /usr/sbin/dnsmasq.", "test": "Seam test: binary present + unit absent -> install is attempted; unit present -> skipped. Red on current code.", "minutes": 40}, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-330", "sev": "P3", "category": "Storage & devices", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-agent/internal/hub/report.go:406-425 SmartSummary carries reallocated/pending/offline_uncorrectable + NVMe set only — no 187/188/199 fields. Wire change across agent + hub (+ controller), declared M.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-332", "sev": "P3", "category": "Storage & devices", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Closing condition is live-only (a real degrading disk or an injection through agent /disks -> controller -> hub). Last related commits ea16a21b/2fa1efc narrowed the restart half only; no commit records a live Hiba-from-counters verdict.", "dup_of": null, "unique_facts": "Needs an injection harness (new mechanism) or a real failing disk; cannot close from source.", "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-333", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "(b) felhom-agent/internal/storage/hostops.go:374: `out, stderr, err := h.runner.Run(ctx, h.bins.Smartctl, \"-a\", \"-j\", device)` — no -n standby. (a) still an operator decision (Viktor decides).", "dup_of": null, "unique_facts": "Not in the row: (b) is not a one-line agent change — configs/felhom-agent.sudoers:34 pins the exact argv `/usr/sbin/smartctl ^-a -j /dev/(...)$`, so adding `-n standby` also needs a sudoers change delivered by the signed bundle.", "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-338", "sev": "P3", "category": "Security & access", "group": "UNCHECKED", "evidence": "felhom.eu/documentation/operations/nodes.md:86-88 now says demo-hp 'Agent config shape (R-50 island): local_api on 169.254.253.1:8443/vmbr9, guest eth1 169.254.253.2/30', and :84 records demo-hp was reprovisioned (address 192.168.0.87 -> 192.168.0.104, read 2026-09-21). Whether the reprovisioned box is actually on the island is a live-box fact (agent.json, pct config) not provable from source.", "dup_of": null, "unique_facts": "A reprovision since the row was filed (fresh installs are born-on-island per memory) may have made nodes.md true; a read-only check of demo-hp's agent.json listen_addr settles it.", "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-340", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu/scripts/felhom-tenantsync.sh: no health op and no 8007 probe (grep '8007\\|health)' empty). Needs an ep0 (protected) script version bump + hub signal (M).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-349", "sev": "P3", "category": "Box system & updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-agent/internal/hub/report.go:282-292 reports only WrapperSHA256; no agent binary sha field in the report (grep AgentSHA256 finds only the hub's manifest entry, hub/internal/api/handler.go:2726). R-349 commits 40d857b/910fd911 are the manual correction only.", "dup_of": null, "unique_facts": "Fix spans agent (report own sha) + hub (compare to vouched agent_sha256).", "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-350", "sev": "P3", "category": "Security & access", "group": "UNCHECKED", "evidence": "Whether the hub operator password was rotated after 2026-08-20 lives only in the hub DB / operator; git log shows no rotation record (only 910fd911 filing it). Not determinable from source.", "dup_of": null, "unique_facts": "The reusable lesson is already in memory (curl-w-redirect-url-leaks-credentials). If the operator confirms no rotation and accepts the risk (value is in a local transcript on DooPlex only), this can close as accepted; rotation itself costs ~5 minutes via /configuration.", "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-362", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller/controller/internal/backup/offbox_restore.go:380, offbox.go:1929, shares_restore.go:116: `return fmt.Errorf(\"restore dir: %w\", err)` — no drive-state consultation on the restore path; IsDisconnected (settings.go:2001) is consulted only by backup legs/update guard (backup.go:609, tier2.go:457, …). No R-362 commit.", "dup_of": null, "unique_facts": "The settings IsDisconnected flag may lag a detach by seconds (the observed detach was 4 s into the restore), so the check should also test the drive mountpoint directly, not only the flag.", "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/backup/offbox_restore.go", "controller/internal/backup/offbox_restore_test.go"], "change": "When MkdirAll/write of the restore destination fails with EACCES/ENOENT, check whether the destination's drive is disconnected/decommissioned or no longer a mountpoint, and return a Hungarian error naming the drive instead of the raw permission error.", "test": "Unit test with a temp dir made read-only plus a settings stub marking the drive disconnected -> error names the drive, not 'permission denied'; red on current code.", "minutes": 60}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-363", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-SMALL", "evidence": "felhom-controller/controller/cmd/controller/main.go:1546: `sched.Daily(\"fill-watch\", \"03:30\", func(ctx context.Context) error { return fillWatcher.Check() })`. Watcher emits on escalation only with a persisted band (internal/fillwatch/fillwatch.go:133, :231), so a faster cadence does not spam.", "dup_of": null, "unique_facts": "Recommendation not followed: sharing the reserve's reading is a design change; an hourly sched.Every gives the same outcome cheaply because Check() is escalation-only and persisted.", "small_fix": {"repo": "felhom-controller", "files": ["controller/cmd/controller/main.go (gitignored dir — git add -f)", "a source-pin test"], "change": "Replace sched.Daily(\"fill-watch\", \"03:30\", …) with sched.Every(\"fill-watch\", time.Hour, …) (scheduler.go:104), keep the startup check.", "test": "fillwatch test: two consecutive Check() calls at the same band emit once (proves hourly is safe); a source-pin test that main.go registers fill-watch via Every.", "minutes": 30}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-388", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Product direction, operator's call; recorded as [DESIGN — DIRECTION] in documentation/architecture/08-alarm-ladder.md §8 per the row. Nothing in source to fix.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-401", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Event-triggered watch row: felhom-controller/controller/internal/backup/offbox_integrity.go:63 `const integrityCheckTimeout = 30 * time.Minute`, :78 `var integritySlowNoticeThreshold = 5 * time.Minute` — unchanged; the trigger (slow WARN on a large store) has not been recorded.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-409", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller/controller/internal/backup/recovery_unit.go:58 `Checksums map[string]string json:\"checksums\" // sha256 of captured compose/ files`; writers at :147, :152, :202 hash only compose/.felhom.yml/app.yaml — no db_dumps or volume_dumps hash.", "dup_of": null, "unique_facts": "Changes the recovery-unit manifest format (capture + verify side); not under an hour.", "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-412", "sev": "P3", "category": "Backup & restore", "group": "NOT-WORTH-IT", "evidence": "Leg 1 shipped: felhom-controller fcef8e0 / 62c6a8a (v0.232.0, R-412a) — hollow push now logs at WARN, pinned by TestR412a_EmptyPushDoesNotReadAsAPlainSuccess. Leg 2 (re-read the unit before push vs accept the race) has no code and is a decision.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "A narrow race: if a recovery unit is destroyed inside an off-site run after its own dump leg, the push ships the just-rebuilt hollow unit; the next run repairs it and the WARN line now says it carried no data.", "cost": "A re-read/re-validate step in the push path plus a red-proof drill; touches the capture/push boundary the architecture says not to guard (08 §8.2).", "if_never": "Rarely, one off-site snapshot of one app is hollow until the next nightly run; the R-403 mirror guard keeps the good secondary copy; the WARN makes it visible.", "pick": "close-as-accepted"}, "minutes_spent": 4}
|
||||
{"id": "R-433", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Hetzner answered (ticket per the row): file-level snapshot access is a MAIN-account capability. hub/internal/hetznerapi/hetznerapi.go still has no snapshot read method (only size_snapshots usage, :83-87). The owed proof (read a file from a snapshot with the main account; forced-command append-only key) is a live, credential-bound operator act.", "dup_of": null, "unique_facts": "The row's state cell 'BLOCKED-ON-PROVIDER' is stale — the provider has answered; the row should be restated as 'owed: one main-account measurement' (operator + CC).", "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-435", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom.eu/hub/internal/monitor/offsite.go:255: `snapshotDropFraction = 0.5 // more than half the history gone in one step` — no per-tag second signal exists.", "dup_of": null, "unique_facts": "Not in the row: STATUS.md no longer contains a 'noticed within a day' claim (grep -i 'within a day|deletion|half' empty), so that half of the row's reason for staying open is gone; what remains is the per-tag detector (new mechanism).", "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-440", "sev": "P3", "category": "App updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "app-catalog-felhom.eu @917a779: 15 templates still have no `update_ladder:` in .felhom.yml — bentopdf code-server glance gokapi gramps-web homebox homepage jellyfin onlyoffice plant-it plex recipe-importer seerr vaultwarden wanderer (calibre-web got its first step in 53a4a1d). For these, pins float with no recorded digest.", "dup_of": null, "unique_facts": "Closes per app as R-462 (widen the upgrade harness) proves a first ladder step; the pre-v0.269.0 installs heal on their next guarded update.", "small_fix": null, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-444", "sev": "P3", "category": "Box system & updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No fstrim anywhere in felhom-agent or felhom.eu/scripts (grep 'fstrim' over *.go/*.sh empty). Needs a new periodic host job (agent, privileged, sudoers/bundle) and possibly an operator surface.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-446", "sev": "P3", "category": "App updates", "group": "DUPLICATE", "evidence": "felhom-controller/controller/internal/stacks/updateorder.go:96: `if len(s.CatalogDigests) == 0 || s.CatalogTestedAt.IsZero() { return false }` — blind only for apps with no ladder entry, i.e. the same 15 templates R-440 lists (app-catalog has no update_ladder for them). Both rows close by the same act: each app's first proven ladder step (R-462).", "dup_of": "R-440", "unique_facts": "Move into R-440: the customer-visible consequence — the 'Naprakész' badge cannot go 'behind' for those 15 apps (updateorder.go:96); R-740 (same-tag re-test) is CLOSED (catalog 6a3ead9), so floating-tag drift for laddered apps is handled.", "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-450", "sev": "P3", "category": "App updates", "group": "FIXED-BY-LATER-WORK", "evidence": "The row's only remainder was 'the other ten PostgreSQL apps need two-venue proof'. R-463 CLOSED 2026-09-30 by felhom.eu 25cb3eb9 ('The last six PostgreSQL apps decided'): 8 of 11 moved by the box's own conversion, 3 (zipline, adventurelog, immich) stay by decision 42; CLOSED-ITEMS.md:223. Source proof: app-catalog templates/docmost/docker-compose.yml:62 'image: postgres:18-alpine' (also rallly:67, outline:64, paperless-ngx:101). The per-app engine gate stays as the permanent rule (catalog CLAUDE.md:115-122).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-458", "sev": "P3", "category": "App updates", "group": "NOT-WORTH-IT", "evidence": "Still true by design: Syncer.copyTemplates copies .felhom.yml verbatim. The row narrows the risk to type: api probes with expect; those exist (e.g. templates/dawarich/.felhom.yml:104 'expect:', adventurelog:93, immich:116, nextcloud:108). Quick catalog history scan (commits touching .felhom.yml health/expect lines AND an image: line) found only new-app commits (96829d0, 72247a3, 882ac14, 195129c, 4351d08) plus lifecycle/re-pin commits a325416/b3eabfd; b3eabfd's wanderer diff showed no healthcheck lines. So no evidence of a healthcheck changed together with a version move on an existing app.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "A frozen (pinned-behind) app can receive a newer health check from .felhom.yml. The only result is a false 'degraded'/dead-app alarm, never data loss, and only for type: api probes with an expect block.", "cost": "Freezing just the healthcheck key means the sync must parse and rebuild a metadata file on the path that touches every app every 15 minutes. That is new surface on a hot path.", "if_never": "Possibly a false alarm one day for an app pinned behind, if a catalog commit changes an expect-probe together with a version. History shows none so far, and the ladder now moves apps step by step.", "pick": "close-as-accepted"}, "minutes_spent": 12}
|
||||
{"id": "R-462", "sev": "P3", "category": "App updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Ongoing multi-session work (fixtures and ladders per app). Last progress: catalog e6f3ec2 (2026-09-30), audits/more-night-apps-2026-09-30/. 21 apps still have no ladder (per the row). Not checkable as done from source.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-468", "sev": "P3", "category": "Box system & updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "A standing pre-customer arrangement, not a defect. The mechanism is live: felhom.eu/scripts/golden_currency_gate.py:137 reads documentation/tests/golden-waiver.yml. The waiver file is currently ABSENT: deleted in felhom.eu 5efe6dae (2026-10-04, 'golden 0.292.0 vouched ... golden waiver deleted'). The row retires only at the first external install, so it stays as a watch.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-469", "sev": "P3", "category": "App updates", "group": "STILL-TRUE-SMALL", "evidence": "What stood between this row and its close was R-463, now CLOSED (felhom.eu 25cb3eb9; CLOSED-ITEMS.md:223). The per-app PostgreSQL gate (decision 35) is now the permanent rule. One thing still contradicts source: app-catalog-felhom.eu/CLAUDE.md:103 heading reads '- **A MariaDB major gets its OWN EDGE; PostgreSQL and MySQL may not cross a major at all.**', but the body at :115-122 lets PostgreSQL cross per app with engine_conversion + two-venue proof, and 8 apps already did (e.g. templates/docmost/docker-compose.yml:62 'image: postgres:18-alpine').", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "app-catalog-felhom.eu", "files": ["CLAUDE.md"], "change": "Reword the rule heading at CLAUDE.md:103 and the 'What is NOT lifted' paragraph (:112-114) to say that PostgreSQL crosses a major one app at a time with engine_conversion + both-venue proof (decision 35) and that MySQL stays refused. Then close R-469 citing R-463's closure. Do not delete the gate: it is now the per-app enforcement.", "test": "Run python3 scripts/catalog_gates.py (unchanged result expected). grep CLAUDE.md for 'may not cross a major at all' -> 0 hits.", "minutes": 20}, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-489", "sev": "P3", "category": "Apps & catalog", "group": "FIXED-BY-LATER-WORK", "evidence": "The residual (a unit-restore-recreated volume has no compose label, so the remove answered []) was fixed in felhom-controller 206b035 (v0.268.0, R-658). controller/internal/stacks/delete.go:1089-1090: '// appVolumeSet is every volume the removal accounts for: the ones carrying the project label AND the // ones the app's definition declares that Docker holds by name (R-658, v0.268.0).' delete.go:703 'resp.VolumesRemoved = removedVolumes(volsBefore, m.appVolumeSet(name, stackDir))'.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-498", "sev": "P3", "category": "Apps & catalog", "group": "STILL-TRUE-SMALL", "evidence": "Still literal: grep -l '\\.DOMAIN' over app-catalog templates/*/.felhom.yml -> 58 of 58; e.g. templates/bookstack/.felhom.yml:73 \"- 'Nyisd meg a wiki.DOMAIN címet a böngészőben'\". The controller renders first_steps as-is: controller/internal/web/templates/app_info.html:224 '{{range .AppInfo.FirstSteps}}<li>{{.}}</li>{{end}}' (also deploy.html:684). The precedent substitution exists for default creds only: controller/internal/web/known_login.go:67 'strings.ReplaceAll(meta.AppInfo.DefaultCreds, \"DOMAIN\", s.cfg.Customer.Domain)'.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/web/handlers.go (app info view model)", "controller/internal/web/known_login.go (pattern)", "new test in controller/internal/web/"], "change": "When building the app info view, rewrite each first_steps entry: replace '<word>.DOMAIN' with the stack's real address (installed SUBDOMAIN + customer domain when installed, else the template default subdomain + customer domain). Use the same approach as known_login.go:67.", "test": "A render test over every catalog template's first_steps (or a fixture of 3) asserting that no rendered string contains the literal 'DOMAIN'. A red-proof with the substitution removed must fail.", "minutes": 60}, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-516", "sev": "P3", "category": "Install & onboarding", "group": "STILL-TRUE-NOT-SMALL", "evidence": "By its own text it now waits for a Hungarian walk on a box with a second drive (items 4, 7, 8, 9, 10) and needs a separate row for item 11. That is live-box work, not source. No commit after 2026-09-20 names R-516 as closed.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-521", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No storage-disconnect suppression of app_start_failed: controller/cmd/controller/main.go:2796 'down := (stacks.IsDownState(st.State) || crashLooping) && !userStopped && !quiesced[st.Name]' (only user-stop/quiesce suppress), and controller/internal/notify/notifier.go:718 emits app_start_failed per newly-down app. It also needs the hub cooldown semantics changed (per-key cooldown outliving storage_reconnected) and an operator decision on customer mail policy. That is two repos plus a decision.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-522", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No tunnel-connection signal in the controller: grep for TunnelConnected/cloudflared connection state in controller/internal -> none. The tile comes from container metadata: controller/internal/web/inframeta.go:21-22 '\"cloudflared\": { DisplayName: \"Cloudflare Tunnel\"'. Fixing it needs a new state source (cloudflared metrics or the hub-push result) wired into the tile, plus a live internet-cut validation.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-531", "sev": "P3", "category": "Box system & updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Both measurements are done (audits/evidence-drill-0243-2026-09-16/). What remains is an operator design question about the crash-loop budget (a slow loop every 20 min is never paused). That is an operator decision, not code.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-540", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Still a single pool box: felhom.eu/hub/cmd/hub/main.go:367 'poolBoxID, _ := strconv.ParseInt(os.Getenv(\"HETZNER_POOL_BOX_ID\"), 10, 64)'. It needs a selection-rule design and eventually a second box (money). No risk today (0.3 % full per the row).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-542", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKED", "evidence": "The controller passes the agent's 'initialize' list through untouched: controller/internal/web/agent_disk_handlers.go:160-162 mergeAttachCandidates only appends to Attach. The agent puts every candidate under initialize: felhom-agent internal/localapi/disks.go:438 'initialize = append(initialize, c) // every unclaimed disk can be initialized'. Whether a REGISTERED in-guest drive still counts as 'unclaimed' depends on the agent's ListCandidateDisks claim state on a real box (internal/storage/candidates.go). Only a live box shows that. No commit names R-542.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-545", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Still no un-configure route: controller/internal/web/server.go:744-783 lists /backup/offbox/{config,toggle,enable-all,offer-dismiss,run,reset,status,restore,place,reconstitute,verify-copy/delete,confirm-escrow,inject-password}; nothing removes a target. The fix is a new destructive customer action (shred data/offbox/) with a hub-escrow refusal check (R-241 rule), Hungarian/English copy and a UI. That is more than an hour.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-547", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-SMALL", "evidence": "Still daily: controller/cmd/controller/main.go:1546 'sched.Daily(\"fill-watch\", \"03:30\", func(ctx context.Context) error { return fillWatcher.Check() })' plus one startup check (:1557-1566). The check is edge-triggered against PERSISTED state (comment at :1553-1555), so running it more often adds no repeat mails.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/cmd/controller/main.go", "felhom.eu/documentation/architecture/08-alarm-ladder.md (one sentence, optional)"], "change": "Also schedule the fill-watch on an interval (e.g. sched.Every(\"fill-watch-fast\", 15*time.Minute, ...) calling the same fillWatcher.Check) next to the daily run, and log the cadence. Alternative if the operator prefers: state in 08-alarm-ladder.md that a transient full disk is out of scope.", "test": "Unit test with a fake usage func: a target that crosses 95 % between two interval ticks yields exactly one notification, and a second tick at the same level yields none (pins edge-triggering under the faster cadence).", "minutes": 50}, "not_worth": null, "minutes_spent": 8}
|
||||
{"id": "R-548", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "The row's original fix shape SHIPPED: felhom-agent 0722b2c (2026-09-24, R-685) checks before start, internal/backup/runner.go:279 '... so a new one needs about %s (old archives are removed only after a successful backup)'. Shown in the UI by controller 44ae4de (v0.272.0). The 2026-09-30 addendum is still true and is the open part: the local tier is refused for ever (10 refusals on demo-hp) because the old archive is removed only after a success. Nothing gives the room back. That needs a design (prune before write, or move the tier), so it is not small. Consider re-scoping the row to that residual.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-552", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "The interrupted-restore notice is cleared only at controller/internal/backup/opstatus.go:61 'delete(m.opInterrupted, stack)' inside BeginRestoreOp. removeStack (controller/internal/api/router.go:955) clears only the update hold, at router.go:1054 'if cleared, err := r.sett.ClearUpdateHold(name); err != nil {'.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/backup/opstatus.go", "controller/internal/api/router.go", "tests in internal/backup and internal/api"], "change": "Add Manager.ClearInterruptedRestore(stack) (delete from opInterrupted + persistRestoreRecordLocked under m.mu). Call it in removeStack next to ClearUpdateHold, and log when it cleared something.", "test": "Unit test: record an interrupted restore for app X, call ClearInterruptedRestore(X), and assert that the list is empty and the persisted record no longer holds X. Wiring test: removeStack on a router with a fake backup manager calls the clear. Red-proof by removing the call.", "minutes": 45}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-554", "sev": "P3", "category": "Install & onboarding", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Still present: controller/internal/setup/ (setup.go, handlers.go, csrf.go, network.go, templates/), and controller/cmd/controller/main.go:328-330 'if setup.NeedsSetup(cfg) { ... runSetupMode(cfg, logger)'. The fix deletes a package and adds a new waiting page (HU+EN copy) with a red-proof. It also needs a check of drill/golden reliance on .needs-setup (controller/internal/web/handler_debug.go references it). More than an hour.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-562", "sev": "P3", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Needs an operator word on the Hungarian number/date format (the row says so) and a deliberate Hungarian-byte change release with parity re-capture. Not checked further.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-565", "sev": "P3", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "The English page test detects only accented letters: controller/internal/web/i18n_parity_test.go:563 'func huLetter(s string) bool {' used by TestI18nEnglishPages (:601). The ASCII Hungarian list exists only in the extractor: controller/scripts/i18n_extract.py:40 'ASCII_HU = re.compile(r\"\\b(Fut|Nincs|Igen|Nem|Hiba|Mentve|...'. No Go test uses it.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/web/i18n_parity_test.go"], "change": "Add an ASCII Hungarian word regex to TestI18nEnglishPages (seeded from i18n_extract.py ASCII_HU plus the words releases B/C found: mp, db, FIGYELEM, jelenlegi, majd a(z), Konfig, Megtartva, helyi, Befejezve, automatikus, kedd/szerda/szombat, szint), applied after the data mask.", "test": "Negative control: a plain English sentence passes. Decoy: plant 'mp' in an English value and the detector must flag it. Then run the existing English renders (go test ./internal/web -run TestI18nEnglishPages, no Docker).", "minutes": 55}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-573", "sev": "P3", "category": "Monitoring & notifications", "group": "FIXED-BY-LATER-WORK", "evidence": "felhom-controller 7c4a33b (v0.258.0, 'the last four Hungarian things an English household met ... R-573 the two channel banners'). controller/internal/web/alerts.go:113 'func (am *AlertManager) SetAgentChannelAlert(down bool, msgKey, msg string) {' and :136 'func (am *AlertManager) SetEndpointDriftAlert(drift bool, msgKey, msg string) {'. Both set MessageKey with msg only as a fail-open fallback.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-575", "sev": "P3", "category": "Apps & catalog", "group": "STILL-TRUE-SMALL", "evidence": "Still a plain string: controller/internal/stacks/deploy.go:1456 'func (m *Manager) memoryVerdict(newReqMB, newLimitMB, releasedReqMB, releasedLimitMB int) (refusal error, warning string) {'. Callers are deploy.go:302 and update.go:497. The consequence is stated at controller/internal/stacks/deploy_errors.go:41.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/stacks/deploy.go", "controller/internal/stacks/update.go", "the deploy response renderer in internal/web or internal/api"], "change": "Make memoryVerdict return (refusal error, warningKey string, warningArgs []any) instead of a msgHU string, and render the warning at the deploy answer with the request's language (the Alert/UpdateRefusal pattern).", "test": "Test that a deploy over the soft memory line returns the English warning for lang=en and the unchanged Hungarian bytes for hu (parity fixture). Red-proof: forcing hu rendering fails the en case.", "minutes": 60}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-578", "sev": "P3", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "The lock-reentrancy guard is still one test in one package: controller/internal/backup/offsite_diag_test.go:190 'TestNoteHelpersAreNotCalledUnderTheSettingsLock'. No other package has one, and there is no gate (grep for R-578 in controller -> none). The fix needs a cross-package AST gate that knows which methods read settings (or a re-entrant read path in Settings). That is a new mechanism, likely more than an hour with decoys.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-581", "sev": "P3", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "Still no tie-break: felhom.eu/hub/internal/store/store.go:1329 'SELECT customer_id, MAX(received_at) as max_time' joined on 'r.received_at = latest.max_time' (:1333). A same-second tie returns BOTH rows (a duplicate customer in GetCustomers), not just an arbitrary one. Other newest-report queries also order on received_at alone: store.go:1393 and :1446 'ORDER BY received_at DESC'. Compare :3574, which already has ', id DESC'.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu (hub)", "files": ["hub/internal/store/store.go", "hub/internal/store/*_test.go"], "change": "In GetCustomers, join on MAX(id) per customer_id instead of MAX(received_at), and add ', id DESC' to the ORDER BY at store.go:1393 and :1446.", "test": "Store test: write two reports for one customer in the same tick with different health/version, then assert that GetCustomers returns exactly one row for that customer and that it carries the second report. Red-proof against the current query (expect 2 rows or the wrong one).", "minutes": 40}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-584", "sev": "P3", "category": "Security & access", "group": "NOT-WORTH-IT", "evidence": "Still true as a process gap: the only written rule is felhom-controller/.claude/rules/ui-hungarian.md (credential-bearing helper cleanup). grep for 'shred' over the workspace .claude/rules and felhom.eu/skills -> no hits, so there is no 'last act of a phase' mechanism. The five files were already shredded (per the row).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "Helper scripts with the shared DEMO controller password inline were left in a demo guest's /tmp. The cleanup rule exists, but no mechanism enforces it.", "cost": "A real mechanism would be a push-helper wrapper that writes credentials to a 0600 file and shreds on exit, and every session would have to use it. Another rule line would be a wish, as the row itself says.", "if_never": "Demo-box (Tier 0) password litter may recur on throwaway guests. The password is a shared demo one, not a customer one, and rotating it is cheap.", "pick": "close-as-accepted"}, "minutes_spent": 5}
|
||||
{"id": "R-585", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Still finished Hungarian from callers: controller/internal/notify/notifier.go:408 'n.PushEvent(\"backup_failed\", \"error\", message, BackupDetails{Error: errMsg})', :487 offbox_enlarge_blocked, :497 db_dump_failed. None of the six types is in convertedProducers (controller/internal/notify/message_customer_test.go:39). The hub still has no customerMessages entry for offbox_enlarge_blocked (felhom.eu/hub/internal/notify/templates.go:94/153). The fix changes six producers' callers across packages: more than an hour.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-586", "sev": "P3", "category": "Process & tooling", "group": "NOT-WORTH-IT", "evidence": "The harness is still in no push-time gate or CI: grep for 'bootstrap-modes' over felhom.eu/scripts and .gitea -> none. It is required only in documentation/runbooks/iso-release-gate.md:285 'docker run --rm -v <repo>/scripts/iso:/work felhom-iso-assistant:trixie bash /work/test/bootstrap-modes.sh'. The FIFO fix shipped per the row.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "The ISO bootstrap harness runs only at each ISO release gate (G16), not on every push.", "cost": "CI needs a container-capable job (a new runner capability), or the push hook needs Docker. Both are new infrastructure for a harness that matters only when an ISO is released.", "if_never": "A harness break is caught at the next ISO release instead of at the push. ISO releases are infrequent, and that gate already blocks a broken one.", "pick": "close-as-accepted"}, "minutes_spent": 6}
|
||||
{"id": "R-587", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-SMALL", "evidence": "The files are gone (ls /mnt/5_hdd/felhom.eu/felhom-iso/out | grep -c rootpw -> 0). The guard is still not built: the publish still relies on the include pattern, felhom.eu/skills/felhom-build-deploy/SKILL.md:115 'rclone/rclone:latest copy /data R2:felhom-iso --include \"felhom-installer-<VER>*\"'. The build still emits the file for appliance mode: felhom.eu/scripts/iso/build-felhom-iso.sh:351 '... > \"$OUT_ISO.rootpw.txt\" )'.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu", "files": ["scripts/iso/build-felhom-iso.sh", "skills/felhom-build-deploy/SKILL.md (publish step)", "scripts/iso/test/ (a small test)"], "change": "At the start of build-felhom-iso.sh, refuse (non-zero exit, named reason) when any *.rootpw.txt exists in the output dir. In the SKILL publish block, add a pre-check line that aborts the rclone copy if a *.rootpw.txt is present in the publish source.", "test": "Shell test: a temp out dir with a decoy x.rootpw.txt makes the build entry refuse with exit != 0. The same dir without it proceeds past the check (stub the rest with an early-exit env flag).", "minutes": 45}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-593", "sev": "P3", "category": "Apps & catalog", "group": "STILL-TRUE-SMALL", "evidence": "Still wrong: app-catalog-felhom.eu/templates/papra/.felhom.yml:47 ' description: \"Az alkalmazás aldomainje\"' sits under AUTH_SECRET (:37). SUBDOMAIN (:30-35) has no description.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "app-catalog-felhom.eu", "files": ["templates/papra/.felhom.yml", "scripts/copy_freeze/hu.json"], "change": "Move 'Az alkalmazás aldomainje' to SUBDOMAIN, give AUTH_SECRET its own description (e.g. 'A munkamenetek aláírásához használt kulcs — ne generáld újra'), add the two English i18n.en descriptions, and re-capture papra's entries in copy_freeze/hu.json with the reason in the commit.", "test": "python3 scripts/catalog_gates.py green (copy-freeze and i18n coverage). Catalog English ceiling for papra reaches 14/14.", "minutes": 25}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-600", "sev": "P3", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "Log line unchanged: felhom.eu/hub/internal/web/customer_delete.go:310 's.logger.Printf(\"[INFO] customer DELETE cascade COMPLETE for %s (journal #%d) — full teardown\", customerID, journalID)'. The wgsync Trigger is called only from hub/internal/api/wg.go:86 'h.wgSyncer.Trigger()', not from the customer cascade. Measured lag is seconds to about 6 min (row).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu (hub)", "files": ["hub/internal/web/customer_delete.go", "hub/cmd/hub/main.go (wire the reconciler into the web server if not already)", "test in hub/internal/web"], "change": "Call the wgsync reconciler's Trigger() before the COMPLETE log line (nil-safe), and make the line say 'wg peer removal pushed' or 'queued for the next wgsync push' depending on whether a syncer is wired.", "test": "Web test with a fake reconciler: a completed delete cascade calls Trigger exactly once, and the log line contains the peer-removal state. Red-proof by removing the call.", "minutes": 50}, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-607", "sev": "P3", "category": "App updates", "group": "UNCHECKED", "evidence": "The row asks first for a live reproduction loop (push a tag, sync, read catalog_images on a timer) and has no diagnosis. Why the sync reports 'nincs változás' while the cache moved, and when CatalogImages refreshes, are live-box behaviour I could not settle from source in the time box. No commit after d19f07ea (filing) names R-607 as fixed.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-612", "sev": "P3", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "The memory half is fixed (catalog a5a729a, 'wishlist 512M (R-612)'). The open half, making a failed first-boot seed visible, needs a new detection mechanism (read the seed's exit/log, or a probe that checks the Role/Group rows). No commit addresses it.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-613", "sev": "P3", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "uptime-kuma is fixed (catalog a5a729a). The open half is a sweep of all 58 templates for probes that pass on a setup wizard. That needs per-app live inspection, so it is not small. No commit names it.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-615", "sev": "P3", "category": "App updates", "group": "STILL-TRUE-SMALL", "evidence": "Still inert: controller/internal/sync/sync.go:279 clones only 'if _, err := os.Stat(gitDir); os.IsNotExist(err)'. Later cycles run :299 'fetch --depth 1 origin' against the stored remote, and nothing compares cfg.Git.RepoURL with origin (grep 'set-url|remote' in sync.go -> none).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/sync/sync.go", "controller/internal/sync/sync_test.go"], "change": "Before the fetch, run 'git remote set-url origin <buildRepoURL()>' (or read origin and re-clone on mismatch), and log at INFO, masked, when the remote changed. The appended drill-folder residue (stack folders the live catalog lacks) is a separate drill-teardown item and is not part of this fix.", "test": "Test with two local bare repos (git only, no Docker): clone A, switch cfg.Git.RepoURL to B, sync, and assert that the cache HEAD is B's commit. Red-proof against current code.", "minutes": 50}, "not_worth": null, "minutes_spent": 7}
|
||||
{"id": "R-616", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-SMALL", "evidence": "Still true: controller/internal/sync/sync.go:327-331 buildRepoURL injects 'https://%s:%s@' and the clone at :283-288 passes that URL to 'git clone', so it persists as origin. maskRepoURL (:100) masks only log lines. Inert on the fleet while git.token is empty (row).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/sync/sync.go", "controller/internal/sync/sync_test.go"], "change": "Clone and fetch with the credential-free RepoURL, and supply credentials per command via '-c http.extraHeader=Authorization: Basic <b64>' (or GIT_ASKPASS env) only when username+token are set. Pairs naturally with the R-615 set-url fix (set-url to the bare URL).", "test": "Test with a local repo and a token configured: after clone and after a sync, 'git -C cache config remote.origin.url' contains no '@'. Red-proof on current code. The operator token rotation stays a separate operator act.", "minutes": 55}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-622", "sev": "P3", "category": "App updates", "group": "FIXED-BY-LATER-WORK", "evidence": "adventurelog v0.13.0 was diagnosed, fixed and promoted with a two-venue test record in app-catalog-felhom.eu 06ea7da (2026-09-27, 'adventurelog: v0.12.1 -> v0.13.0 with its health and world-data fixes in the same commit (R-655, 09 decision 41)'; bench healthy in 217 s, box 9202 through the guarded Update in 204 s). templates/adventurelog/docker-compose.yml:13 ' image: ghcr.io/seanmorley15/adventurelog-backend:v0.13.0'. The proven step is recorded at templates/adventurelog/.felhom.yml:132 (update_ladder entry from v0.12.1).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-635", "sev": "P3", "category": "Monitoring & notifications", "group": "FIXED-BY-LATER-WORK", "evidence": "The open remainder (app_oom fires once per container run, no escalation) was built in felhom-controller 0054d4b (v0.265.0, 'OOM storm alarm', R-636). controller/internal/notify/notifier.go:746 '\\t\\tn.emit(\"app_oom_storm\", \"error\",' fires once per run when 20 or more kills land in 30 min (:754-764, oomStormKills=20, oomStormWindowMin=30; pinned by TestR636_*). The 79 % headroom and the method lesson are carried by R-462 (per the row).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-645", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Still true for the operator CLI: controller/internal/settings/settings.go:1923-1929 ClearRestoreHold deletes ANY hold reason (including update-failed) with no pin restore, and the flag is in controller/cmd/controller/main.go:94/198. The row lists three candidate shapes with 'none chosen'. Picking one changes operator-path semantics and the capture logic, so it is not a one-hour fix. (The automatic undo is no longer affected, since v0.263.0, per the row.)", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 6}
|
||||
{"id": "R-675", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-SMALL", "evidence": "Unchanged: controller/internal/web/handlers.go:1745 'return head + \"A fájlok a második meghajtó másolatából állíthatók vissza: „Fájlok visszaállítása”.\"'. The branch does not check for a whole copy on the second drive (decision 26, v0.269.0).", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom-controller", "files": ["controller/internal/web/handlers.go", "a handler test in controller/internal/web"], "change": "In missingFileLegsRefusal, when the second drive holds a whole copy of the app (the decision-26 whole-copy check the backup manager already exposes for the restore page), name that whole restore action instead of 'Fájlok visszaállítása'. Keep the existing branch otherwise.", "test": "Three-branch test with a fake backup manager (off-site row / whole copy on drive 2 / tier-2 files only / nothing), asserting each sentence. Red-proof: the whole-copy case fails on current code.", "minutes": 45}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-676", "sev": "P3", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "A watch row, still true: controller/internal/stacks/unhealthy.go:116 skips only 'if st.Deploying || st.Updating || st.HoldReason != \"\" || st.updateHeld {', so a deploy's first start (after Deploying clears) is sampled by decision 28's crash-loop stop. The immich cause is fixed in the catalog (56c4888, 768M, per the row). Covering a slow first start would need a first-start grace decision.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-682", "sev": "P3", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "felhom-controller 7690c27 (v0.296.0): no remove journal in source — grep -rni 'remove.*journal|removeJournal|remove_intent|interrupted remove' controller/*.go returns nothing; git log --grep R-682 empty. Needs a new boot-time journal mechanism.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-683", "sev": "P3", "category": "App updates", "group": "UNCHECKED", "evidence": "Watch item about a power-cut drill outcome; behaviour only a live box shows; git log --grep R-683 empty in controller.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-698", "sev": "P3", "category": "Backup & restore", "group": "NOT-WORTH-IT", "evidence": "Row is an open operator decision among options (a)-(d); no source change implied. No commit references R-698 fix.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "A backup records the image name/digest, not the image; restoring a version the maker deleted from the registry fails at the pull.", "cost": "Options b-d add a new mirror or tens-hundreds of MB per app per copy on every tier, a new part on the recovery path.", "if_never": "A restore of a deleted version fails; the household uses the next copy or a newer version (option a). All 42 ladder digests resolved when measured.", "pick": "close-as-accepted (option a), operator to confirm"}, "minutes_spent": 2}
|
||||
{"id": "R-700", "sev": "P3", "category": "Storage & devices", "group": "FIXED-BY-LATER-WORK", "evidence": "felhom-controller 820e8ef (v0.276.0, R-697/R-700). controller/internal/stacks/migrate.go:789: 'm.logger.Printf(\"[INFO] [stacks] %s: data moved %s -> %s — app.yaml keeps its pin (%d service(s)) and records\"' — persistDriveFlip (migrate.go:763) loads app.yaml and changes only HDD_PATH. Only a live proof on a two-drive box remains (row's own residue).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-704", "sev": "P3", "category": "Apps & catalog", "group": "FIXED-BY-LATER-WORK", "evidence": "felhom-controller 7cba0bf (v0.278.0). controller/internal/api/router.go:918 'func (r *Router) dropLeftoverHold(name, why string) {' calling r.sett.ClearUpdateHold(name); pinned by TestR704_AFreshInstallDropsALeftoverHold. Residue: live proof of the install-time drop only.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-706", "sev": "P3", "category": "Backup & restore", "group": "FIXED-BY-LATER-WORK", "evidence": "felhom-controller 0c702f8 (v0.279.0). controller/internal/api/router.go:946 'if err := r.backupMgr.DeleteOffsiteRestoreCopy(name); err != nil {' inside removeVerificationCopy (R-706); pinned by TestR706_RemovalWithBackupsDeletesTheVerificationCopy. Residue: not seen live.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-717", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "app-catalog templates/opengist/.felhom.yml:42 and templates/wishlist/.felhom.yml:42 carry only signup_block; no after_setup/after_install in either .felhom.yml (grep empty). Fix needs per-app DB writes (opengist sqlite with app stopped) plus live proof.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-723", "sev": "P3", "category": "Monitoring & notifications", "group": "FIXED-BY-LATER-WORK", "evidence": "felhom.eu 80aeac71 (hub v0.126.0). hub/internal/monitor/staleness.go:174 '// R-723 (v0.126.0): a customer's NEW box is not a recovery.'; hub/internal/notify/dispatcher.go:540 '\"suppressed\", \"first hour of a new box (R-723)\", \"operator\"'. Residue: live proof at a real first install.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-724", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Text parts fixed in controller 6be6c53 (v0.283.0). Remaining LAN/gateway read still goes only through the samba container: controller/internal/stacks/guestnet.go:47 'out, err := dockerexec.Command(\"docker\", append([]string{\"exec\", sambaContainer}, args...)...).Output()' — the comment (guestnet.go:18) accepts that reads fail while sharing is off. Needs another read path (design).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-728", "sev": "P3", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "No commit fixes R-728 (git log --grep in felhom.eu only the filing commit 11591f3a). hub/internal/web/configs.go:705 'existing, _ := s.store.GetCustomerConfig(customerID)' is a check-then-act before slow applyOffsite/applyPBSDR; store.go:1592 save is an upsert ('retrieval_password = excluded.retrieval_password'), so two concurrent submits both pass and both mint.", "dup_of": null, "unique_facts": null, "small_fix": {"repo": "felhom.eu (hub)", "files": "hub/internal/web/configs.go (+ a new _test.go)", "change": "Add a per-customerID in-flight guard (sync.Map/mutex set) around the create handler from the duplicate check to the self-bind mint, so a second concurrent submit for the same ID gets the 'already exists' form; optionally disable the submit button on submit in the template.", "test": "Unit test firing two concurrent POSTs for the same customer ID with a blocking applyOffsite seam; assert exactly one 'Customer config created' / one self-bind mint; red-proof by removing the guard.", "minutes": 60}, "not_worth": null, "minutes_spent": 5}
|
||||
{"id": "R-729", "sev": "P3", "category": "Backup & restore", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No route clears the off-site target: controller/internal/web/server.go:744-783 lists /backup/offbox/{config,toggle,enable-all,offer-dismiss,run,reset,status,restore,place,reconstitute,verify-copy/delete,confirm-escrow,inject-password}; /reset (offbox_handlers.go:311) only resets an orphaned repo. New press needs handler + settings clear + template + HU/EN copy + escrow/hub-managed-target interplay — more than an hour.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-733", "sev": "P3", "category": "Process & tooling", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Harness/golden-evidence change plus a decision whether proofs run with swap off; no commit references R-733. Not a source-verifiable single fix.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-738", "sev": "P3", "category": "App updates", "group": "NOT-WORTH-IT", "evidence": "wger part FIXED: app-catalog 7a4ff48 'wger: run the database migrations at start (R-738)'; templates/wger/docker-compose.yml:46 ' - DJANGO_PERFORM_MIGRATIONS=True'. Only the residue stays: the guarded Update's health check reads only the front page.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "The guarded Update's health check sees only an app's front page, so an app that serves its front page while its data is broken passes as done.", "cost": "A per-app data read-back inside the product's update check — a new mechanism across every template (the harness fixture already does this off-box).", "if_never": "Catalog onboarding's fixture read-back keeps catching such breaks before a version reaches the ladder; a box could still report done on a broken update the catalog did not test.", "pick": "close-as-accepted (wger defect fixed; the generic gap is a design note)"}, "minutes_spent": 3}
|
||||
{"id": "R-747", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Lockout shortened: app-catalog a4597cd; templates/mealie/docker-compose.yml:27 ' - SECURITY_USER_LOCKOUT_TIME=1'. Residue still open: hourly lock renewal by a stranger (needs decision 57 option d) and page copy; needs decision + live proof.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-755", "sev": "P3", "category": "Apps & catalog", "group": "DUPLICATE", "evidence": "Still true: templates/wger/docker-compose.yml has no WGER_USE_GUNICORN (grep empty). R-762 (open, read) states 'Owner decides together with R-755 (same server question)' and its fix names 'the gunicorn switch of R-755'.", "dup_of": "R-762", "unique_facts": "wger runs `manage.py runserver` (Django dev server) measured via ps on 9202 2026-10-01; upstream entrypoint.sh:81-87 starts gunicorn only with WGER_USE_GUNICORN=True; the fix needs its own memory watch on bench + box.", "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-756", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKED", "evidence": "Depends on whether 9202's scratch drive is a registered drive — live box state; the row itself says not measured which. Not verifiable from source.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-757", "sev": "P3", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "No commit references R-757. controller/internal/stacks/deploy.go:1339-1349: 'case \"secret\":' ... 'value, err := generateValue(field.Generate)' ... 'appCfg.Env[field.EnvVar] = value' for any missing field of a deployed app, with no exception for fields consumed only by after_install. Fix needs a design (a marker for given-at-install fields or an ask path).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-758", "sev": "P3", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Still true: templates/bookstack/.felhom.yml:18 ' mem_limit: \"512M\"' vs compose limits docker-compose.yml:42 '512M' + :79 '256M'; onboarding/EXISTING-APPS-GAPS.md:26 still lists all 8. No gate (only scripts/onboarding_gaps.py:171 reports it). Not small: raising 8 figures changes the capacity check (decision 22) — which apps fit a box — plus a gate with decoy and a publish.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-762", "sev": "P3", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "templates/wger/docker-compose.yml sets no DJANGO_DEBUG and no static/media server (grep empty); templates/wger/.felhom.yml:18 'lifecycle: hidden' (catalog 55b8c8a). Needs a server design decision (nginx sidecar vs gunicorn+static) and bench+box proof.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-763", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "templates/wger/docker-compose.yml sets neither ALLOW_REGISTRATION nor ALLOW_GUEST_USERS (grep empty); wger hidden (.felhom.yml:18 'lifecycle: hidden'). The env change is tiny but its proof needs a working wger on 9202 (blocked by R-762); best done in the same session as R-762.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-774", "sev": "P3", "category": "Apps & catalog", "group": "STILL-TRUE-NOT-SMALL", "evidence": "templates/karakeep has no Sentry/phone-app sentence (grep -i sentry empty); mail-ON proof needs a hub-enabled live box (demo-hp) — live work.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-775", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Narrowed (Grimmory published behind the family gate). Residue: per-name 15-min lock is hard-coded upstream (no setting) and the reinstall-over-kept-books finding is uninvestigated — needs live investigation.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 2}
|
||||
{"id": "R-776", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Only bookstack has it: templates/bookstack/docker-compose.yml:32 ' - APP_PROXIES=172.16.0.0/12'; grep for TRUSTED_PROXIES/CORE_TRUST_PROXY/IPEXTRACTION/N8N_PROXY_HOPS in kimai, zipline, vikunja, nextcloud, n8n compose returns nothing. Five apps, each needing a live 3.6 re-measure on 9202.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-778", "sev": "P3", "category": "Security & access", "group": "NOT-WORTH-IT", "evidence": "Reasoned-only row; controller client-address code now in controller/internal/web/clientaddr.go (v0.286+). The window exists only during a self-update crash roll-back below 0.286.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "If a box rolls back to a controller older than 0.286, the dashboard's login counter trusts the leftmost forwarded address and can be dodged until the box moves forward.", "cost": "A 0.285.x patch release or a self-update rule refusing to roll back across 0.286 — release work for a rare window.", "if_never": "The window lasts only from a crash roll-back to the next floor delivery; the floor never moves back. A stranger could make more password guesses during that window.", "pick": "close-as-accepted"}, "minutes_spent": 2}
|
||||
{"id": "R-782", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Source agrees: templates/glance/docker-compose.yml:27 seeds glance.yml with no auth: block (grep 'auth' in templates/glance empty); templates/homepage has no HOMEPAGE_ALLOWED_HOSTS (grep empty). Needs live measurement on 9202 and a decision whether a public glance dashboard is intended.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-783", "sev": "P3", "category": "Security & access", "group": "NOT-WORTH-IT", "evidence": "Measured upstream behaviour (better-auth 3 per 10 s keyed on one address); SparkyFitness exposes no env for ipAddressHeaders. No source change possible in the catalog.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "Three wrong SparkyFitness sign-ins by anyone block every visitor's sign-in for about 10 seconds.", "cost": "Needs an upstream setting (better-auth ipAddressHeaders) plus a right-walking reader — not in our control.", "if_never": "A stranger retrying every 10 s can keep the household out of sign-in; one who stops frees it within 10 s. Not forgeable.", "pick": "close-as-accepted (re-open if upstream exposes the setting)"}, "minutes_spent": 2}
|
||||
{"id": "R-785", "sev": "P3", "category": "App updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "templates/sparkyfitness/docker-compose.yml:45 ' image: codewithcj/sparkyfitness_server:v0.17.3' and :89 'codewithcj/sparkyfitness:v0.17.3'. Major-version ladder walk (bench + box), gated on R-784.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-831", "sev": "P3", "category": "Security & access", "group": "NOT-WORTH-IT", "evidence": "Operator decision 73: not rotated by choice. Nothing in source to change.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "The Hetzner storage API token was printed into one session transcript.", "cost": "Three manual console/kubectl steps by the operator (about 10 minutes).", "if_never": "Anyone who obtains that transcript could create, reset or delete Storage Box sub-accounts.", "pick": "keep (rotation is the operator's call; do not close a leaked-secret row silently)"}, "minutes_spent": 1}
|
||||
{"id": "R-836", "sev": "P3", "category": "Box system & updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Live host boot-loader work needing operator-approved reboots and measurement (GRUB env block on ESP, sp5100_tco arming).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-839", "sev": "P3", "category": "Box system & updates", "group": "STILL-TRUE-NOT-SMALL", "evidence": "Gate behaves as described: controller/cmd/controller/main.go:2397 'return false, \"drive \" + hdd + \" is not a live mountpoint\"' with hdd = cfg.Env[\"HDD_PATH\"] (main.go:2379). Which writer put a per-app path in HDD_PATH is undiagnosed — a diagnosis task, not a one-hour fix.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 4}
|
||||
{"id": "R-853", "sev": "P3", "category": "Box system & updates", "group": "NOT-WORTH-IT", "evidence": "Still true: felhom-agent e06ed97, cmd/felhom-agent/main.go:3765 'if !f.at.IsZero() && time.Since(f.at) < 10*time.Minute {' caches a failed guest read; main.go:853 factsReporter needs firstGuest(px). The delay equals the 15-min host report interval, so dropping the failure cache alone does not shorten it; the real fix is a guest-less host facts mode in the wrapper.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "After a boot the box's versions and crash facts reach the hub up to about 15 minutes late.", "cost": "A new guest-less facts mode in the root wrapper plus agent change and release.", "if_never": "Facts arrive one report later after each boot; nothing is lost (the crash guard keeps 7 days).", "pick": "close-as-accepted"}, "minutes_spent": 4}
|
||||
{"id": "R-862", "sev": "P3", "category": "Box system & updates", "group": "UNCHECKED", "evidence": "Waiting on the operator's by-hand bootstrap on Tester 2 through his tunnel; whether done is live-box state. No commit records it (felhom.eu log since 2026-10-04).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-870", "sev": "P3", "category": "Security & access", "group": "NOT-WORTH-IT", "evidence": "Operator ruling 2026-10-05 option B: not rotated now. Nothing in source to change.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": {"what": "Tester 1's two Cloudflare tokens (disposable test customer) were printed into one session transcript.", "cost": "About 15 minutes of Cloudflare dashboard + hub edit by the operator.", "if_never": "Someone with the transcript could change DNS or the tunnel of the disposable enkicsifelhom.hu test zone.", "pick": "keep (operator's call; close when Tester 1 is retired)"}, "minutes_spent": 1}
|
||||
{"id": "R-879", "sev": "P3", "category": "Security & access", "group": "STILL-TRUE-NOT-SMALL", "evidence": "hub/internal/store/store.go:166 'retrieval_password TEXT NOT NULL,' and store.go:1586/1592 write retrieval_password and api_key as given; no seal on them (grep seal near these fields empty). Sealing/hashing three tables with migration is a security change, more than an hour.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 3}
|
||||
{"id": "R-882", "sev": "P3", "category": "Hub & operator", "group": "UNCHECKED", "evidence": "Longhorn instance-manager state on DooPlex (Tier 2, forbidden to touch); live-only, owner operator.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-883", "sev": "P3", "category": "Hub & operator", "group": "UNCHECKED", "evidence": "homelab-manifests repo is not in this workspace (ls /mnt/5_hdd/felhom.eu/git shows only app-catalog-felhom.eu, drills, felhom-agent, felhom-controller, felhom.eu); live DooPlex check (kubectl) is out of scope.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
{"id": "R-886", "sev": "P3", "category": "Monitoring & notifications", "group": "UNCHECKED", "evidence": "DooPlex Alertmanager volume ownership; homelab-manifests not in this workspace and live check not permitted.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "minutes_spent": 1}
|
||||
@@ -0,0 +1,325 @@
|
||||
# Burn-down 2026-10-05 — Part A: every P4 and P3 row checked against live source
|
||||
|
||||
Method: 8 read-only checker agents, oldest id first (P4 then P3), each row against `main` source (no machine reached); every FIXED / DUPLICATE verdict re-checked by the session before closing (a sample of 22 cited lines re-grepped; one (R-274) held back as half-checked; four (R-700, R-704, R-706, R-723) moved to the operator list — their code fix is in, only the live observation the row waits for is missing). Raw per-row JSON: `partA-results.jsonl`.
|
||||
|
||||
Groups: DUPLICATE 2, FIXED-BY-LATER-WORK 29, NOT-WORTH-IT 43, STILL-TRUE-NOT-SMALL 129, STILL-TRUE-SMALL 91, UNCHECKED 23
|
||||
|
||||
| Row | Sev | Group | Evidence (abridged) | Min |
|
||||
|---|---|---|---|---|
|
||||
| R-10 | P4 | STILL-TRUE-SMALL | FIX: After the os.Rename in DumpOne, open filepath.Dir(finalPath) and call a best-effort dir.Sync() (log at DEBUG on error), mirroring atomicPromoteTar in backup.go:948. — felhom-controller@7690c27 controller/internal/appbackup/dbdump.go:364 `if err := tmpFile.Sync(); err != nil {` then :390 `if err := os.Rename(tmpPath, finalPath); err != nil {` with no directory Sync after; the twin at controlle | 4 |
|
||||
| R-25 | P4 | STILL-TRUE-NOT-SMALL | felhom-controller@7690c27 controller/internal/web/storage_handlers.go:153 `uuid := resolveEnrollUUID(ctx, agent, device)` still resolves by device PATH after format, then AssignDisk(uuid) at the next step; FormatResult (controller/internal/agentapi/client.go:384-395) carries DurableID only for the confirmation path, not the new fs UUID. Binding resolve+assign to the format's durable-id needs the a | 6 |
|
||||
| R-76 | P4 | UNCHECKED | Behaviour is FileBrowser-image runtime behaviour (mode/setgid of UI-created folders), only observable on a live box. The image has changed since the finding: controller/internal/infra/infra.go:27 `FileBrowserImage = "gtstef/filebrowser:1.5.6-stable"` (finding was on 1.3.3). The comment at infra.go:207-208 still asserts `umask 002 so folders the customer creates here come out group-writable (2775 w | 5 |
|
||||
| R-89 | P4 | STILL-TRUE-NOT-SMALL | No retention policy object in hub: `grep -rln -i 'retentionpolicy/retention_policy' felhom.eu/hub` returns nothing (felhom.eu@53d8131b). Commercial per-customer policy = money/product decision + new reconciler. | 2 |
|
||||
| R-91 | P4 | UNCHECKED | Whether /srv/pbs-felhom still exists on ep0 is live-only (ep0 is protected; not touched). Source-side: CONTEXT.md:3656 still reads "`/srv/pbs-felhom` is 13 G of dead weight on `/` awaiting R-91's go-ahead". Extra fact found: documentation/runbooks/offsite-endpoint.md:24 still says the datastore `felhom-offsite` is at `/srv/pbs-felhom` and :119 `proxmox-backup-manager datastore create felhom-offsit | 5 |
|
||||
| R-92 | P4 | STILL-TRUE-SMALL | FIX: Add an exact-bytes value to the PBS DR view (e.g. UsedBytesExact rendered as a title= tooltip or a MB-precision string below 10 GB) without changing fmtBytesGB for other callers. — felhom.eu@53d8131b hub/internal/web/pbsdr_box.go:57 and :64 `view.UsedStr = fmtBytesGB(snap.UsedBytes)`; hub/internal/web/offsite_box.go:54 `return fmt.Sprintf("%.1f GB", float64(b)/float64(int64(1)<<30))` — still | 4 |
|
||||
| R-93 | P4 | NOT-WORTH-IT | PICK close-as-accepted (operator word needed: close, or reopen as 'build a drift fixture'); also drop the dead drill-r50 fence in target-selection.md:111 at close: A row about choosing between two fixtures, neither of which exists any more. | 4 |
|
||||
| R-99 | P4 | STILL-TRUE-NOT-SMALL | No phantom-snapshot cleanup in felhom.eu/hub or felhom-agent (grep -i phantom finds only agent runner/test detection code; no removal path). Deletion on a customer datastore is a separate operator ruling per the row — customer data. | 3 |
|
||||
| R-104 | P4 | STILL-TRUE-SMALL | FIX: Add an OffsiteFailLocked class matched by offboxLockRe in ClassifyOffsiteFailure (before transport) and a cause line in OffsiteFailureMessage telling the operator the repository is locked by an interrupted run and how it clears. — felhom-controller@7690c27 controller/internal/backup/offbox.go:193-222 ClassifyOffsiteFailure has cases NoUnits/NoRepo/Transport and `default: return OffsiteFailUnk | 5 |
|
||||
| R-124 | P4 | NOT-WORTH-IT | PICK close-as-accepted: The disaster-recovery recipe writes the PBS root namespace as the word 'root', but PBS itself uses an empty name, so a pasted '--ns root' fails. | 4 |
|
||||
| R-129 | P4 | STILL-TRUE-SMALL | FIX: After one read-only `ssh -o BatchMode=yes demo-hp true` (and reading root's authorized_keys comment to name the key), rewrite nodes.md 'Access' section to the measured truth and drop the R-129 caveat in target-selection.md:111-112 (also update the memory index line). — Docs still say no key: felhom.eu@53d8131b documentation/operations/nodes.md:110 `### Access — there is no baked SSH key` and | 4 |
|
||||
| R-134 | P4 | STILL-TRUE-SMALL | FIX: Extract a pure zoneCandidates(domain) []string that yields the name and every parent down to two labels, and loop resolveZone over it (same order: most specific first). — felhom.eu@53d8131b hub/internal/cloudflare/unblock.go:117 `for _, name := range []string{domain, parentDomain(domain)} {` and :136-141 parentDomain strips exactly one label (`strings.SplitN(domain, ".", 2)`); controller stri | 4 |
|
||||
| R-161 | P4 | NOT-WORTH-IT | PICK close-as-accepted (residual is a deliberate ruling; owner operator): The runtime check that app data lands on a volume is run by hand, not on every push. | 3 |
|
||||
| R-162 | P4 | NOT-WORTH-IT | PICK close-as-accepted: If Docker ever ran on a storage driver where `docker diff` does not work, the persistence gate would refuse to report and blame the prober instead of the driver. | 3 |
|
||||
| R-164 | P4 | STILL-TRUE-NOT-SMALL | Predicate still absent: felhom-controller controller/internal/appbackup/dbdump.go:544 still only WARNs `its accounts table has NO rows`; restore still replays dump + tar (internal/backup/restore_unit.go:114-118 hasReplayableDump). Blocked on a design (live-vs-dump per-table counts). | 3 |
|
||||
| R-169 | P4 | NOT-WORTH-IT | PICK close-as-accepted (row itself says decide only if the window ever costs something): CI only reports after a push lands, because every repo pushes straight to main with no pull request. | 2 |
|
||||
| R-177 | P4 | STILL-TRUE-NOT-SMALL | felhom-controller@7690c27 controller/cmd/controller/main.go:1546 `sched.Daily("fill-watch", "03:30", func(ctx context.Context) error { return fillWatcher.Check() })`; internal/scheduler/scheduler.go:269 has GetJobs but grep finds no RunNow/Trigger method and no run-job route in internal/web. Needs a new operator-gated trigger endpoint (auth surface) — a new mechanism, solve together with R-279. | 4 |
|
||||
| R-184 | P4 | FIXED-BY-LATER-WORK | Fixed by felhom.eu b55fc17d "hub v0.102.0 — refuse to vouch a version that cannot be installed (R-273)" — exactly shape (b), validate at vouch time in the hub. felhom.eu/hub/internal/web/configs.go:1358 `res := s.gitea.PackageDownloadable(ctx, t.pkg, t.version, t.file)` and :1365 `s.logger.Printf("[WARN] artifact vouch REFUSED: %s package %s is NOT downloadable (R-287)", ...)`; tag leg at :1343 Ta | 4 |
|
||||
| R-194 | P4 | NOT-WORTH-IT | PICK close-as-accepted: Proxmox caches permissions, so a removed storage grant can still read as present for seconds to minutes; our self-repair notices only after the cache expires. | 2 |
|
||||
| R-206 | P4 | STILL-TRUE-NOT-SMALL | homelab-manifests@87dfc29 (/home/kisfenyo/git/homelab-manifests): no daemon.json template in homelab-ansible (grep finds only a comment at roles/node_housekeeping/templates/node-housekeeping.sh.j2:17 and homelab-ansible/CLAUDE.md:54). Part (b) was superseded by fc9fbb8 ("correct the expired Docker rationale"): the script now says at :13-20 do NOT add docker calls, the GC policy in daemon.json is t | 4 |
|
||||
| R-207 | P4 | FIXED-BY-LATER-WORK | Fixed by homelab-manifests fc9fbb8 "node_housekeeping: guard DRY_RUN, correct the expired Docker rationale, pin container log rotation". /home/kisfenyo/git/homelab-manifests/homelab-ansible/roles/node_housekeeping/templates/node-housekeeping.sh.j2:137 `if [[ "${DRY_RUN}" == "1" ]]; then` inside write_metrics, :138 logs "file left untouched". | 3 |
|
||||
| R-208 | P4 | STILL-TRUE-SMALL | FIX: Move the ARG VERSION/GIT_COMMIT (controller) and ARG VERSION/BUILD_TIME (hub) declarations down to just above the final `go build` RUN. — felhom-controller@7690c27 controller/Dockerfile:12 `ARG VERSION=dev` and :13 `ARG GIT_COMMIT=unknown` sit above :19 `RUN go mod download // true`; felhom.eu@53d8131b hub/Dockerfile:3 `ARG VERSION=dev`, :4 `ARG BUILD_TIME=unknown` above :9 `RUN go mod downlo | 3 |
|
||||
| R-209a | P4 | UNCHECKED | Live-only: whether DooPlex has rebooted and /var/log/felhom-store-postboot-check.log says PASS. Not read (DooPlex is Tier 2, operator ruled no reboot; this pass touches no machine). No source claim to check. | 2 |
|
||||
| R-210 | P4 | NOT-WORTH-IT | PICK close-as-accepted: 193 old controller/hub images exist only on DooPlex and cannot be re-pulled; the question is whether to delete them. | 2 |
|
||||
| R-213 | P4 | STILL-TRUE-NOT-SMALL | Row is a not-started design (live-vs-backup comparison, then put-back flow), operator-owned; nothing in source to verify against. | 1 |
|
||||
| R-230 | P4 | STILL-TRUE-NOT-SMALL | Owed rulings, not code: (a) bulk-correction ruling on MEMORY.md staleness (MEMORY.md index still carries version literals, e.g. 'ctrl 0.224.0', 'hub 0.109.0'); (c) spec-as-failing-test pilot not started. (b) closed. Operator decision required. | 2 |
|
||||
| R-246 | P4 | STILL-TRUE-NOT-SMALL | felhom.eu@53d8131b hub/internal/store/store.go:3248 `func (s *Store) MarkEscrowStale(hostID string) error {` still has no production caller (grep: only definition + comments at offsite.go:208,216); stale_at still read (store.go:3182 clears it). Ruling owed by operator: evidential setter or retire the column (folds R-248). | 3 |
|
||||
| R-256 | P4 | STILL-TRUE-SMALL | FIX: Rewrite flash.offbox.mgr_unavailable / mgr_unreachable in both languages to say the backup service is not running yet and give a route (try again in a few minutes; if it persists, contact support). — felhom-controller@7690c27 controller/internal/i18n/locales/hu.json:1406 `"flash.offbox.mgr_unavailable": "A mentéskezelő nem elérhető.",` used at controller/internal/web/offbox_handlers.go:54 and | 4 |
|
||||
| R-261 | P4 | STILL-TRUE-SMALL | FIX: Reword the doc comment (selfbind.go:106-110) to say it is a test accessor and name the two tests that pin the auto-mint invariant (selfbind_automint_test.go, customer_delete_test.go) — or, if the operator prefers, add one post-mint production check that logs [WARN] when count != 1. — felhom.eu@53d8131b hub/internal/store/selfbind.go:111 `func (s *Store) CountSelfBindTokens(customerID string) | 3 |
|
||||
| R-263 | P4 | STILL-TRUE-SMALL | FIX: Change the comment to 'the only writer that GRANTS the role' and add a source-scanning test that finds every `.BackupTarget =` assignment in non-test settings code and fails if any other than SetBackupTarget can assign a non-false value. — felhom-controller@7690c27 controller/internal/settings/settings.go:1655 `// from every other. This is the ONLY writer of StoragePath.BackupTarget — registr | 3 |
|
||||
| R-264 | P4 | STILL-TRUE-NOT-SMALL | felhom.eu@53d8131b scripts/wire_contract_gate.py still allowlists the six with _R264: :242 selfupdate_pending, :246 selfupdate_pending_version, :255 restore_tests.mount_parity, :258 restore_tests.mount_inventory, :281 backup.last_db_dump, :282 backup.last_integrity_check. Each reader is a design per the row. | 3 |
|
||||
| R-266 | P4 | STILL-TRUE-NOT-SMALL | felhom-controller@7690c27 controller/internal/report/builder.go:94 `{Mount: "/", Label: "SSD", TotalGB: sysInfo.DiskTotalGB, UsedGB: sysInfo.DiskUsedGB, Percent: sysInfo.DiskPercent},` — no disk_known on the storage entry; hub has no disk_known (grep empty). Two-repo wire change gated by wire_contract_gate.py. | 3 |
|
||||
| R-279 | P4 | STILL-TRUE-NOT-SMALL | No operator/hub path to start an off-site run: grep for offbox run triggers in felhom.eu/hub/internal finds nothing; the only run entry is the customer dashboard handler (felhom-controller controller/internal/web/offbox_handlers.go:270 `if !s.backupMgr.OffboxRunnable() {`). Needs a new operator-authenticated trigger — sibling of R-177, not a duplicate (different job). | 3 |
|
||||
| R-284 | P4 | NOT-WORTH-IT | PICK close-as-accepted (close as not-a-defect): A reported 'almost full' warning on an empty disk; the code shows the warning only below 20% free and hides it by default, so the report was a reading of unrendered HTML. | 5 |
|
||||
| R-285 | P4 | STILL-TRUE-NOT-SMALL | No maintenance/expected-downtime concept in hub: `grep -rln -i 'maintenance/expected_downtime/quiet_until/snooze' felhom.eu/hub/internal` returns nothing. New mechanism (M). | 3 |
|
||||
| R-286 | P4 | STILL-TRUE-SMALL | FIX: Add one paragraph: a positive control must come from a different channel than the measurement (different query path, snapshot, API or clock); give the 2026-08-09 stale-snapshot case as the example. Put it in ONE home (pointer elsewhere). — Lesson (a) not written anywhere: grep -i 'different channel/same channel/independent channel' over documentation/runbooks/workspace-CLAUDE.md, felhom.eu/sk | 6 |
|
||||
| R-287 | P4 | FIXED-BY-LATER-WORK | Deleter established 2026-08-10 (R-267 newest-10 prune, recorded in the row itself); CI fixed by felhom-agent 53d047a "Two guards, one number: bound the published check to the retention it must live with" (R-291). felhom-agent@e06ed97 scripts/check-published-versions.py:101 `RETENTION_FILE = os.path.join(os.path.dirname(os.path.abspath(__file__)), "retention-policy.json")`, :213 `keep = retention_k | 5 |
|
||||
| R-288 | P4 | STILL-TRUE-NOT-SMALL | felhom.eu@53d8131b documentation/architecture/00-capability-map.md is now 210 125 bytes / 30 937 words / 253 lines (`wc`), larger than the 134 642 bytes measured in the row; :38 still reads `*Verified 2026-07-16 against evidence corpus @ felhom.eu tip `4b18cc5``. Restructure is an M doc surgery, operator-owned. | 3 |
|
||||
| R-289 | P4 | FIXED-BY-LATER-WORK | R-182 was closed by felhom.eu ef6ac6fe (2026-08-22, register compression): documentation/backlog/CLOSED-ITEMS.md:474 `/ **R-182** / ... / **CLOSED — SHIPPED** (controller v0.194.0 + hub v0.90.0/.1, 2026-08-03) /`. The residue (digest never seen delivering) was since observed: documentation/audits/DRILL-chaos-night-2026-09-17.md:181 `backup_run_failures` „1 of 12 apps failed to back up in this nigh | 5 |
|
||||
| R-290 | P4 | STILL-TRUE-NOT-SMALL | Gate exists (felhom.eu scripts/check_stands.py) but the map itself still carries the claims: documentation/architecture/00-capability-map.md has 95 'PROVEN-LIVE' occurrences (grep -c); demoting 12 rows or writing walk documents is M and blocked on R-288 per the row. | 3 |
|
||||
| R-291 | P4 | STILL-TRUE-SMALL | FIX: Rewrite the _comment/recorded_by to cite the operator's newest-10 rule (R-267/R-287) instead of 'observed, not a ruling', and drop or correct the non-existent registry-retention.md reader. Keep the min_agent-floor note as the recorded better bound; then close R-291. — felhom-agent/scripts/retention-policy.json still says the 10 is 'NOT a ruling anyone has been able to locate' and recorded_by: | 8 |
|
||||
| R-292 | P4 | STILL-TRUE-SMALL | FIX: Make resolveArtifactSHA return a reason (not-found / unreachable / bad manual sha) and redirect to three distinct flashes (reuse artifact_unverifiable for unreachable, add artifact_version_missing, keep artifact_sha_invalid for a bad typed sha incl. the wrapper sha at :1395). — hub/internal/web/templates/configuration.html:55 still reads 'the Gitea sha lookup failed (version missing / Gitea u | 6 |
|
||||
| R-310 | P4 | STILL-TRUE-SMALL | FIX: Drop the second 'The vouched golden is' sentence when GOLDEN_CHECK_WHY already names it (or drop the version from :3060); add one runbook line: --uninstall needs an interactive terminal; --force does not bypass the typed vmid confirm. — felhom.eu/scripts/felhom-host-install.sh:3060 sets GOLDEN_CHECK_WHY="it is controller $ver, but the vouched golden is $ART_GOLDEN_VER" and :3080-3081 die "... | 6 |
|
||||
| R-315 | P4 | STILL-TRUE-NOT-SMALL | felhom.eu/scripts/wire_contract_gate.py:30-48 still documents the test as a repo-wide literal-tag search ('IT PROVES REACHABILITY OF A NAME'); ROOTS at :88 includes the R-311 escrow/retained root. No receiver-type field-by-field comparison exists. Fix requires resolving receiver mirror types — a new mechanism (M). | 5 |
|
||||
| R-325 | P4 | STILL-TRUE-SMALL | FIX: Import RETRIEVAL_STEMS from ../felhom.eu/scripts/customer_copy_vocab.py (same sibling-path pattern as controller_gates.py:48) and delete the STEMS literal; absent sibling = INCONCLUSIVE exit 2. Follow-up (felhom.eu, separate commit): remove hub_copy_gate.py's drift check, which would then fail to find STEMS. — felhom-controller/controller/scripts/retrieval_promise_gate.py:54 still has its own | 6 |
|
||||
| R-327 | P4 | STILL-TRUE-NOT-SMALL | felhom.eu/documentation/architecture/where-felhom-stands.yaml:126-130 still: id claim.code-naming, title "The same word is used for two different secrets across three surfaces; the email points at a page a rebuilt machine does not show", status: partial. Needs the operator's capability-map ruling first (dataset may not be raised on its own). | 4 |
|
||||
| R-331 | P4 | STILL-TRUE-NOT-SMALL | felhom-controller/controller/internal/agentapi/diskverdict.go:34 uncorrectableFailCount = 64; :33 comment still defers 'growth-rate detection once the box keeps history'. No growth-rate rule found. New mechanism (M). | 4 |
|
||||
| R-336 | P4 | STILL-TRUE-NOT-SMALL | No source change reduces the ep0 poll rate (pvestatd interval is Proxmox-side, not in our repos). Design question (does the hub need a 15-min fill reading) remains; acceptance needs an ep0 access-log measurement. Scaling item, not small. | 3 |
|
||||
| R-337 | P4 | UNCHECKED | Live-only behaviour (WATCHING). From source: felhom-agent/internal/localapi/server.go:518 serves GET /backup/status and :1258 answers from s.pickLatestBackup (the in-memory store), which suggests collection cadence, but the refresh path after an out-of-schedule run was not established within the time box. | 6 |
|
||||
| R-345 | P4 | STILL-TRUE-SMALL | FIX: Delete lines 21-22 (or move them behind an explicitly named opt-in target with a comment). Whether a stale :latest already sits on the registry is a separate live check for a session allowed to query it. — felhom.eu/hub/Makefile:21 'docker tag $(IMAGE):$(VERSION) $(IMAGE):latest' and :22 'docker push $(IMAGE):latest' still present; only commit touching the Makefile is 77b5a4ce (initial). | 3 |
|
||||
| R-346 | P4 | NOT-WORTH-IT | PICK close-as-accepted (audit done, zero instances): A warning that a future reader might anchor an uptime slope on systemd's ActiveEnterTimestamp instead of the process start time. | 4 |
|
||||
| R-348 | P4 | STILL-TRUE-SMALL | FIX: Reword the comment: the backups list IS lost on restart and refills only when a backup runs; the hub's freshness VERDICT is unaffected because it looks back 7 days (hub monitor/deadline.go backupEvidenceLookback, pinned by deadline_anchor_test.go TestCheckBackupDeadlines_RestartBlindWindow_NoEvent). — felhom-agent/internal/backup/store.go:28 still reads '// Backups are unaffected — their fres | 6 |
|
||||
| R-352 | P4 | STILL-TRUE-NOT-SMALL | Placement half is an open operator ruling (SPEC-app-data-placement-2026-08-21.md, 'Viktor rules'); deploy route still has no server-side default: GetDefaultStoragePath has no caller in internal/stacks (grep returns only internal/api/router.go:1137 systemInfo). Point (2) is carried by R-368. | 4 |
|
||||
| R-364 | P4 | STILL-TRUE-SMALL | FIX: A helper that, for a pattern containing a byte >= 0x80, also runs an ASCII anchor (must hit) and a negative control (must miss) and refuses to print a zero unless both behave; documented in the felhom-evidence or ui-hungarian rule as the way to search Hungarian text. — No helper exists: ls felhom.eu/scripts shows nothing grep/accent-related, and no script mentions '0x80' or 'negative control' | 4 |
|
||||
| R-365 | P4 | STILL-TRUE-SMALL | FIX: Set AbandonOverdue when DueAt is past and render a new key ('a törlés esedékes, a következő napi karbantartáskor lefut' / English twin) instead of the future-tense sentence. — felhom-controller/controller/internal/i18n/locales/hu.json:368 'A kérésed szerint a korábbi távoli mentéseidet <strong>{{.AbandonDate}}</strong> napján véglegesen töröljük (még ... nap)'; handlers.go:1164-1167 sets Aban | 5 |
|
||||
| R-367 | P4 | NOT-WORTH-IT | PICK close-as-accepted (or delete it the next time the box is reprovisioned): One old 312 KB paperless database dump on the demo-hp test box sits under the app's old folder name; nothing reads or deletes it. | 4 |
|
||||
| R-368 | P4 | STILL-TRUE-SMALL | FIX: Reword the field comment: the deploy FORM pre-selects this path (templates/deploy.html:614); the deploy API applies no default when HDD_PATH is omitted. (The alternative — a server-side default — is a behaviour change and not small.) — felhom-controller/controller/internal/settings/settings.go:572 'IsDefault bool `json:"is_default,omitempty"` // new apps use this by default' (line moved from | 5 |
|
||||
| R-371 | P4 | NOT-WORTH-IT | PICK close-as-accepted with one line in 07-backup-architecture saying success is silent by design because failure and staleness are alarmed: The weekly off-site backup sends no 'done' event, while the two local tiers do. Failures and an 8-day staleness deadline are already alarmed. | 6 |
|
||||
| R-372 | P4 | NOT-WORTH-IT | PICK close-as-accepted: An optional idea from July: show 'this second-drive copy was never made because its source is missing' separately from 'last copy failed' in the operator screen. | 5 |
|
||||
| R-373 | P4 | FIXED-BY-LATER-WORK | Premise (20G/50G two-volume mismatch, 'nothing sets SysDataGrowGB') was retired by agent v0.120.0 one-data-volume work, commit cd6e267 'v0.120.0 — one data volume (R-165...)'. felhom-agent/internal/reconcile/bringup.go:191 '// SysDataGrowGB is a COMPATIBILITY INPUT since agent v0.120.0 (R-165). There is no longer a second' and :437 'growGB := spec.DataVolGrowGB + spec.SysDataGrowGB'; installer pas | 6 |
|
||||
| R-374 | P4 | NOT-WORTH-IT | PICK close-as-accepted, with one line in the audit saying the three are not recoverable: A July audit says three borderline cases were left unfiled but never named them. | 5 |
|
||||
| R-375 | P4 | UNCHECKED | Requires a read-only check on ep0 (token's datastore audit permission); not verifiable from source and ssh is out of scope for this checker. | 2 |
|
||||
| R-376 | P4 | STILL-TRUE-SMALL | FIX: Carry the same marker legend paragraph into the three documents written after the 2026-08-22 pass; then close the row, since 'mark as sessions touch them' is a standing practice already in the template, not a defect. — Legend present ('not yet classified') in 00..06 and 10 of documentation/architecture/, but MISSING in the newer 08-alarm-ladder.md, 09-update-architecture.md and 11-os-updates. | 5 |
|
||||
| R-377 | P4 | STILL-TRUE-SMALL | FIX: Turn each ruling's opening bold line '**S-NN — TITLE (date ...).**' into a '### S-NN — TITLE (date)' heading, changing no other byte; no compression, no reordering. — felhom.eu/CONTEXT.md:1537 '## Standing rulings' runs to EOF: 189,685 bytes, 0 '###' sub-headings, 153 bullets, 39 distinct S- ids; each ruling starts as a bold paragraph e.g. '**S-39 — "WE DO NOT KNOW" IS NEVER DRAWN AS "FINE".. | 6 |
|
||||
| R-390 | P4 | FIXED-BY-LATER-WORK | Commit 2344589a ('... runbook pveam note'); felhom.eu/documentation/runbooks/RUNBOOK-manual-build.md:154 '2. Run **`pveam update` first** — the `virgin` snapshot's template INDEX is stale too, and a stale index fails as a bogus'. | 3 |
|
||||
| R-391 | P4 | STILL-TRUE-SMALL | FIX: Take the row's second option: state in CLAUDE.md that the catalog REPORT.md carries no observations section by convention (findings go straight to the register), so gate 11 is not needed here. The runner refactor (first option) is the bigger alternative. — app-catalog-felhom.eu/scripts/catalog_gates.py has no SHARED_ / observations entry (grep 'SHARED_/observations' returns nothing); app-cata | 4 |
|
||||
| R-392 | P4 | STILL-TRUE-NOT-SMALL | ls felhom.eu/documentation/architecture shows no agent-tooling/workflow document (00-11 are all product; plus _design-review, _hub-review, _recovery-inventory). Writing a new architecture document is more than an hour and needs the operator's view of the split. | 3 |
|
||||
| R-393 | P4 | NOT-WORTH-IT | PICK close-as-accepted (superseded in practice by unprompted-work.md §2/§4): A proposed skill plus helper script to log every decision an unattended run makes. | 4 |
|
||||
| R-394 | P4 | STILL-TRUE-NOT-SMALL | wc -l felhom.eu/skills/felhom-build-deploy/SKILL.md = 186 (was 179 at filing — grew); scripts/check_skills.py:45 GRANDFATHERED still holds the exemption. Trim needs a session that can verify the build/deploy commands it keeps. | 3 |
|
||||
| R-402 | P4 | STILL-TRUE-NOT-SMALL | No hub Go/template reads last_integrity_ok/_depth (grep in hub/internal returns nothing); still allowlisted at felhom.eu/scripts/wire_contract_gate.py:163 and :220. Needs the operator's decision on what the screen says. | 3 |
|
||||
| R-416 | P4 | STILL-TRUE-SMALL | FIX: Apply the existing RULE 3 duplicate check to CLOSED-ITEMS.md too (suffixed ids like R-88a/R-88b stay distinct), and update the closed_register_gate.py:53 hole list to point at it. — Partly covered: scripts/register_shape_gate.py:120 'RULE 3 — duplicate: {rid} already has a row at line ...' (added in 462ab4a5, R-627) now refuses a duplicate id WITHIN OPEN-ITEMS.md only (REG path at :84 = OPEN- | 7 |
|
||||
| R-418 | P4 | STILL-TRUE-SMALL | FIX: Add the two missing gates to the docstring list and a test that parses the docstring's gate labels and asserts they equal [g[0] for g in GATES]. — It drifted AGAIN: felhom.eu/scripts/repo_gates.py docstring lists 1-14 (+9b) = 15 gates while GATES has 17 — 'script-tests' and 'decoy-coverage' are registered but not listed (python import: len(GATES)=17). No test compares the two. | 5 |
|
||||
| R-420 | P4 | NOT-WORTH-IT | PICK close-as-accepted (add it with the first gate that needs it): The felhom.eu gate runner cannot mark a gate as advisory-only; the controller runner can. | 3 |
|
||||
| R-421 | P4 | STILL-TRUE-NOT-SMALL | Deliberate class row ('stays open as the place the next instance is recorded'); its open instances R-422..R-426 are still open in this batch. Not a fixable item by itself. | 2 |
|
||||
| R-422 | P4 | STILL-TRUE-NOT-SMALL | felhom.eu/scripts/reuse_refs_check.py:41 PATH_RE still ends '\.(?:go/py/html/css/yml/yaml/sh)\b' — no .md. Widening it needs a false-positive walk across all four repos' REUSE.md/CLAUDE.md citations (the row says that pass is the work). | 3 |
|
||||
| R-423 | P4 | STILL-TRUE-SMALL | FIX: Discover website/**/*.html and FAIL on any page not in PAGES (or glob and keep PAGES only as exceptions); flip the decoy test to expect conviction and drop the 'site' EXEMPT entry. — felhom.eu/scripts/site_gates.py:22-27 hardcoded PAGES list; website/ today holds exactly those 9 files, so nothing is missed today, but a new page is not scanned. | 4 |
|
||||
| R-424 | P4 | NOT-WORTH-IT | PICK close-as-accepted (hole stays declared in the gate's docstring): The roadmap gate cannot tell a real defect filed as an 'idea' from a genuine idea. | 3 |
|
||||
| R-425 | P4 | STILL-TRUE-SMALL | FIX: Scan by pattern (templates/backups*.html, *offbox*.go) plus internal/i18n/locales/hu.json, or assert FILES against a discovered set so an unclassified file fails; drop its decoy-coverage exemption. — felhom-controller/controller/scripts/offbox_rename_gate.py:16-20 FILES = backups.html, offbox_handlers.go, offbox.go only; templates/backups_remote.html exists and is not scanned, and customer co | 6 |
|
||||
| R-426 | P4 | STILL-TRUE-NOT-SMALL | felhom.eu/scripts/decoy_coverage_gate.py EXEMPT now has 19 entries (loaded via python): hub-copy is gone, felhom-agent 'release-complete' is new; group (d) gates (hostinstall, wire-contract, due-checks, published, image-resolvable, volume-persistence) all still exempt. | 5 |
|
||||
| R-427 | P4 | FIXED-BY-LATER-WORK | Commit 71b8c8c6 (Backlog triage Part B: '... closed_register_gate RULE 3 refuses a finished row in OPEN-ITEMS'); felhom.eu/scripts/closed_register_gate.py:10 'RULE 3 — (2026-10-03) no row in `OPEN-ITEMS.md` may carry a CLOSED-family word (CLOSED, SHIPPED,'. Of the 12 named rows, R-385/387/341/378/405/88a/88b/123 are now only in CLOSED-ITEMS.md; R-190 and R-352 remain open (partly-closed, as the ro | 5 |
|
||||
| R-437 | P4 | FIXED-BY-LATER-WORK | felhom.eu 71b8c8c6 'Backlog triage Part B: 125 finished rows + 20 id-less rows moved to CLOSED-ITEMS ... closed_register_gate RULE 3 refuses a finished row in OPEN-ITEMS (decoys, seen red) ... register 444 -> 325'. felhom.eu/scripts/closed_register_gate.py:10 'RULE 3 — (2026-10-03) no row in `OPEN-ITEMS.md` may carry a CLOSED-family word'. Gate run today: 'closed-register gate OK — no open work fi | 4 |
|
||||
| R-445 | P4 | NOT-WORTH-IT | PICK close-as-accepted: The hub's per-app memory suggestion can be built from samples of an app that no longer runs anywhere (e.g. a 15-minute test install). | 5 |
|
||||
| R-451 | P4 | STILL-TRUE-NOT-SMALL | felhom-controller/controller/internal/report/types.go ContainerDetailReport still carries only Name/State/CPUPercent/MemoryMB (no image field). Ruled (09 §3 decision 18), build deferred until fleet grows; needs controller payload + hub denormalisation + fleet page across two repos. | 3 |
|
||||
| R-454 | P4 | STILL-TRUE-SMALL | FIX: Add a gofmt -l gate (fails on any listed file, INCONCLUSIVE if gofmt missing) to controller_gates.py, see it red on today's tree, then one gofmt -w formatting commit for the 12 files. — The five files named are now clean (gofmt -l controller/internal/web/ prints nothing), but `gofmt -l controller` in felhom-controller lists 12 OTHER files today: cmd/controller/main.go, internal/agentapi/diskv | 5 |
|
||||
| R-457 | P4 | STILL-TRUE-SMALL | FIX: Read the six candidates; for each date literal that feeds an assertion evaluated against time.Now(), derive it from now (as the R-457 fix did). The faked-future-date CI instrument is a separate, larger idea and should be split out or dropped. — No later commit references R-457 beyond the filing release (felhom-controller 38d28b5 v0.234.0 / 998aa31 REPORT). No faked-clock CI run exists (grep f | 4 |
|
||||
| R-460 | P4 | NOT-WORTH-IT | PICK close-as-accepted: BookStack's uploaded files cannot be checked automatically after an upgrade; only its database can. | 2 |
|
||||
| R-464 | P4 | FIXED-BY-LATER-WORK | Lesson homed and harness uses the correct probe. app-catalog-felhom.eu b7ef0c4 'upgrade-test.py: record the engine's own view of its datadir'; app-catalog-felhom.eu/scripts/upgrade-test.py:278 '"mariadb-upgrade --check-if-upgrade-is-needed --user=root "'. felhom.eu d6837d98 (SPIKE R-459); felhom.eu/documentation/architecture/09-update-architecture.md:1647 '1. **Ask the engine, not the log.** Maria | 4 |
|
||||
| R-488 | P4 | UNCHECKED | The claim is a measured suite runtime (5.5 min); confirming it needs running go test ./internal/backup, which I did not run (read-only; backup tests may reach real docker on DooPlex). No commit after filing (felhom-controller 24d7c54) mentions R-488 or a test-speed change in internal/backup (git log --grep on internal/backup since 2026-09-13: empty), so it is likely still true. | 3 |
|
||||
| R-492 | P4 | STILL-TRUE-SMALL | FIX: Remove Paths.HDDPath, its env binding and each reader's dead global branch, keeping the per-app/discovered fallbacks each reader already uses. — cfg.Paths.HDDPath still defined and read: controller/internal/config/config.go:167 'HDDPath string `yaml:"hdd_path"`', :453 envStr("FELHOM_PATHS_HDD_PATH", &cfg.Paths.HDDPath); readers internal/report/builder.go:69, internal/monitor/healthchec | 4 |
|
||||
| R-494 | P4 | STILL-TRUE-NOT-SMALL | felhom.eu/hub/internal/cloudflare/ holds only unblock.go; no tunnel/DNS creation code in hub (grep cfd_tunnel: none). Building it is a new Cloudflare-API mechanism on the hub; operator ruled it non-blocking. | 3 |
|
||||
| R-501 | P4 | FIXED-BY-LATER-WORK | felhom.eu a4993272 'CLAUDE.md: the CI-check recipe was wrong in two ways, both measured today'. felhom.eu/CLAUDE.md:176 'rows — a run can sit several pages earlier. **Scan every page** and match on `head_sha`; with a'; recipe at CLAUDE.md:166-168 loops every page. | 3 |
|
||||
| R-502 | P4 | STILL-TRUE-SMALL | FIX: Register bootstrap-modes.sh in repo_gates.py behind a docker-available check that reports INCONCLUSIVE (never pass) when docker/the felhom-iso-assistant image is absent, plus a decoy (a broken banner must turn it red). — felhom.eu/scripts/iso/test/bootstrap-modes.sh exists; `grep -rn 'bootstrap-modes/bootstrap_modes' scripts/*.py .gitea/workflows` in felhom.eu returns nothing — no gate or CI | 3 |
|
||||
| R-503 | P4 | NOT-WORTH-IT | PICK close-as-accepted (as DECLINED by ruling; the three measurements stay in the closed row for any future reversal): An idea, offered and not chosen: the installer would pick the disk itself when there is exactly one. | 2 |
|
||||
| R-504 | P4 | UNCHECKED | The claim (iso.felhom.eu/ returns 404) is live-only; I may not curl hosts. felhom.eu/documentation/runbooks/VOLUNTEER-first-hour.md:14 still says '`iso.felhom.eu/` itself still has no index — R-504'. Fix needs a Cloudflare rule the operator owns. Cosmetic; households use felhom.eu/letoltes (website/letoltes.html exists). | 3 |
|
||||
| R-507 | P4 | STILL-TRUE-NOT-SMALL | Needs measuring QEMU input-send-event or a VNC client against a live VM on felhom-pve — a live-machine spike, not a source change. No later commit references R-507. | 2 |
|
||||
| R-525 | P4 | STILL-TRUE-NOT-SMALL | Row itself states it is a new unmeasured mechanism (forwardAuth / Quantum proxy auth) needing a scratch-guest spike. | 1 |
|
||||
| R-526 | P4 | STILL-TRUE-NOT-SMALL | felhom.eu/hub/internal/tenantsync/client.go:110 '// Deprovision DESTROYS the customer's PBS namespace, all its backup groups, and its token — the'; no token-only op exists. Needs a new op on protected ep0 and an operator yes/no. | 3 |
|
||||
| R-527 | P4 | NOT-WORTH-IT | PICK close-as-accepted (with the corrected facts: flag read into LockedFields, enforced only by uncalled UpdateStackConfig): A catalog flag that marks some settings 'locked after install' changes nothing visible: the page makes every setting read-only anyway, and the edit path that would honour the flag is never called. | 8 |
|
||||
| R-532 | P4 | NOT-WORTH-IT | PICK close-as-accepted: Vaultwarden's web page shows a sign-up form even though sign-ups are off; the server then refuses it. | 3 |
|
||||
| R-541 | P4 | STILL-TRUE-NOT-SMALL | Row: needs a new copy/re-key/release mechanism and a design; no later commit references R-541. Far off per re-rank (0.3% full, one pool box). | 2 |
|
||||
| R-544 | P4 | STILL-TRUE-SMALL | FIX: Change the log line to state the effect, e.g. 'host deleted: %s (escrow custody demoted to retained)' when escrow existed, and drop the boolean name from the text. — felhom.eu/hub/internal/web/hosts.go:917 's.logger.Printf("[INFO] host deleted: %s (escrow deleted: %v)", hostID, deleteEscrow)'. | 3 |
|
||||
| R-551 | P4 | NOT-WORTH-IT | PICK close-as-accepted (add 'walk R-546 readiness branches' to the next fresh-install checklist instead): The escrow 'waiting for the agent' screens are proven by tests but never seen on a real box in that state. | 2 |
|
||||
| R-555 | P4 | STILL-TRUE-SMALL | FIX: Strip Go // and /* */ comments and template {{/* */}} before TOKEN_RE in receiver_tokens; add a decoy 'tag named only in a receiver comment must convict'. Newly surfacing tags each become a finding (allowlist with reason or a row). — felhom.eu/scripts/wire_contract_gate.py:451 'def receiver_tokens(repo_root):' tokenises whole files: line 482 'toks.update(TOKEN_RE.findall(fh.read()))' — no com | 3 |
|
||||
| R-564 | P4 | STILL-TRUE-SMALL | FIX: Add split-form Hungarian patterns (állíthatók? vissza, (hoz/szerez/nyit)\w* vissza), register the Hungarian occurrences found (the seven already reviewed in English), and add a planted split-verb decoy. — felhom-controller/controller/scripts/retrieval_promise_gate.py:54 'STEMS = ["visszaállíthat", "visszaszerezhet", "visszahozhat", "visszanyit"]' — joined forms only, no split-verb pattern. | 3 |
|
||||
| R-567 | P4 | STILL-TRUE-SMALL | FIX: Add (eq .Page "storage_init") (eq .Page "storage_attach") to $storageOpen and mark the Meghajtók link active for them. — controller/internal/web/templates/layout.html:83 '{{$storageOpen := or (eq .Page "storage") (eq .Page "storage-network")}}' — storage_init/storage_attach not included; storage_handlers.go:351 'data := s.baseData(tmpl, title)' passes the template name as Page. | 4 |
|
||||
| R-568 | P4 | STILL-TRUE-SMALL | FIX: Sort rows by diskKey(d) (durable id, falling back to name) before returning. — controller/internal/web/disk_health.go:124-152 diskHealthRows appends rows in resp.Disks order; no sort in the file (grep 'sort.' in disk_health.go: none). | 3 |
|
||||
| R-569 | P4 | STILL-TRUE-SMALL | FIX: Add KindErrorf-style sentinels in internal/stacks (protected, not found, not deployed/still running, not orphaned), a statusFor helper per handler family replacing the three Contains blocks. — controller/internal/api/router.go:742 'if strings.Contains(err.Error(), "protected") {', also :745, :1018, :1021, :1024 ('not deployed'/'still running'), :1102, :1105, :1108 ('not orphaned'). | 3 |
|
||||
| R-570 | P4 | STILL-TRUE-NOT-SMALL | Fallback still present: controller/internal/web/handlers.go:1050 'offboxStaleWarningMarker = "nincs mentésre jelölt alkalmazás"'; producer internal/backup/offbox.go:1168. Closing depends on a fleet condition (every box one off-site run on >=0.251.0) — a watch, not a fix. | 3 |
|
||||
| R-571 | P4 | STILL-TRUE-SMALL | FIX: Add a short section to 07 listing the six failure classes, what each means for the customer, and that restic/ssh signatures are external; add an alert-placement paragraph (inline under storage bars vs top banner) to 02. — grep ClassifyOffsiteFailure/PageOnly/Inline in felhom.eu/documentation/architecture/07-backup-architecture.md and 02-controller-module-map.md: no hit. Classifier lives at fe | 3 |
|
||||
| R-574 | P4 | STILL-TRUE-NOT-SMALL | controller/internal/web/handler_debug.go still carries 40 lines with accented Hungarian string literals (grep -cP count); last touched by 0c702f8 v0.279.0, not converted. Labelling ~40 literals page-copy vs payload, adding en/hu keys and parity fixtures exceeds an hour. | 3 |
|
||||
| R-576 | P4 | STILL-TRUE-NOT-SMALL | felhom-controller/controller/scripts/i18n_go_parity.py has no call-site argument-count or '+'-adjacency check (grep verb/argument: only VERB_RE/strip_verbs for text equality, lines 76-78, 196-199). Parsing multi-line Go call arguments reliably from Python with decoys is likely >1 h. | 4 |
|
||||
| R-577 | P4 | STILL-TRUE-NOT-SMALL | Waiting on an operator decision (what the share feature promises a stranger); felhom.eu/documentation/architecture/10-localisation.md table row 'the two guest share pages, the catch-all / a stranger / nobody / **no globe** / — (R-577, the operator's)'. | 2 |
|
||||
| R-579 | P4 | STILL-TRUE-NOT-SMALL | Gate deliberately deferred until R-554 deletes the first-boot wizard; R-554 is still OPEN (OPEN-ITEMS.md:131). Versionless links remain in controller/internal/setup/templates/setup_*.html:8 '<link rel="stylesheet" href="/static/style.css">' (8 files). | 5 |
|
||||
| R-588 | P4 | STILL-TRUE-SMALL | FIX: Name the single home documentation/tests/iso-release-<ver>-<date>/ in the Result-recording section, and add a pointer dir/README for 1.28.0 to its audit record (evidence-backup-promise-2026-09-16/phaseD-iso-gate.txt). Optional existence check left out. — felhom.eu/documentation/runbooks/iso-release-gate.md:319-322 'Result recording' says only 'in the release report' — names no home. documenta | 4 |
|
||||
| R-591 | P4 | STILL-TRUE-SMALL | FIX: Deep-copy Meta.I18n in deepCopyStack (map plus nested values). — The copy is deepCopyStack in controller/internal/stacks/manager.go (row says Copy()); it deep-copies AppConfig (:1137-1148), DeployFields (:1162), OptionalConfig (:1174), Integrations (:1186) and has no I18n line (grep I18n in manager.go: none). Meta.I18n defined at internal/stacks/metadata.go:94. | 4 |
|
||||
| R-594 | P4 | STILL-TRUE-SMALL | FIX: Add ALLOWLIST_EN of (app, path, reason); registered occurrences pass, unregistered convict, and any entry matching nothing is itself a failure. — app-catalog-felhom.eu/scripts/check-copy-i18n.py: grep ALLOWLIST/allowlist — no hit; no way to register a true occurrence. | 3 |
|
||||
| R-599 | P4 | STILL-TRUE-SMALL | FIX: Make both 409 bodies say when the last report arrived and when deletion opens (last report + configured stale threshold), and name the wait in target-selection.md's drill section. — felhom.eu/hub/internal/web/hosts.go:898 'http.Error(w, "Host is ONLINE — deletion is refused (a live agent would receive 401s permanently).", http.StatusConflict)' — no last-report age or opening time. target-sele | 4 |
|
||||
| R-602 | P4 | FIXED-BY-LATER-WORK | felhom.eu e02bc038 'hub v0.119.0 — ... R-596/R-598 closed' added the finding; felhom.eu/documentation/architecture/10-localisation.md:809-812 'the `felhom_lang` cookie and got the **Hungarian** page for `en`. ... The cookie is the right instrument for the anonymous claim page and the **wrong**' and :509 '`langFor`'s order is fixed: `?lang=` → **the household's setting when a session exists**'. Onl | 4 |
|
||||
| R-603 | P4 | STILL-TRUE-SMALL | FIX: Add a test helper that compares against html.EscapeString(want) and use it in the render tests that assert English copy; the bundle gate with a 27-entry allowlist is the larger alternative. — No gate or helper: grep for html.EscapeString(want/'/R-603 in felhom-controller/controller scripts+internal: none. controller/internal/i18n/locales/en.json has 27 lines containing an apostrophe today | 4 |
|
||||
| R-605 | P4 | STILL-TRUE-SMALL | FIX: Give harness-level refusal its own exit code (e.g. 3 = REFUSED) in both scripts and map it to a distinct label in catalog_gates.py. — app-catalog-felhom.eu/scripts/catalog_gates.py:122 'VERDICT = {0: "OK", 1: "FAILED", 2: "INCONCLUSIVE"}' — harness refusal and per-app undetermined both exit 2 and print the same word. | 3 |
|
||||
| R-610 | P4 | NOT-WORTH-IT | PICK close-as-accepted: A power cut landing inside the sub-second `starting` phase has never been measured; all three live cuts landed in `verifying`, which runs the same recovery code. | 5 |
|
||||
| R-617 | P4 | FIXED-BY-LATER-WORK | felhom.eu@462ab4a5 (2026-09-22) documentation/architecture/09-update-architecture.md:1969 "`POST /api/v1/repos/migrate` is the route that works (the project's Gitea tokens carry" - continues at :1970 "`write:repository` but not `write:user`, so `POST /user/repos` answers 403"; recipe at :1979. The one-line note the row asked for exists (in the architecture doc rather than operations/). The optiona | 4 |
|
||||
| R-618 | P4 | STILL-TRUE-NOT-SMALL | Remaining open work = the controller-side idea (let `verifying` accept docker's own `healthy`). grep for docker health status in felhom-controller@7690c27 controller/internal/stacks/update.go returns nothing; no R-618 reference in Go source. It is an undecided design question (operator), not a defect; the three probe fixes (app-catalog@793c4fb) and the gate scripts/check-probe-matches-compose.py a | 4 |
|
||||
| R-619 | P4 | STILL-TRUE-SMALL | FIX: In getDeployFields, copy meta.DeployFields and set Required=true for every field with Type=="password" before writing the response (copy, do not mutate the shared metadata; the web deploy page uses GetDeployFields separately and is untouched). — felhom-controller@7690c27 controller/internal/api/router.go:395 `meta, appCfg, err := r.stackMgr.GetDeployFields(name)` then :402 `"metadata": meta | 6 |
|
||||
| R-621 | P4 | STILL-TRUE-NOT-SMALL | Capture is done: felhom-controller@7690c27 controller/internal/stacks/update.go:1054 `outDir := filepath.Join(dir, "hold-logs", ts)`. The open part (show it on the app page's hold panel / logs fallback) is not built: grep for hold-logs/holdLogs in controller templates and handlers returns only update.go:1050-1082 and undo.go:535,550 (writers). Surfacing needs a page change with HU/EN copy and a de | 4 |
|
||||
| R-624 | P4 | STILL-TRUE-NOT-SMALL | Row's own latest update: remaining class is vaultwarden (closed sign-up by design) and code-server; the open decision is whether the harness may hold an app's admin secret (operator). Not verifiable further from source; needs a decision, not a fix. | 2 |
|
||||
| R-644 | P4 | UNCHECKED | About the live state of gokapi on scratch guest 9202 (crash-loop, deployed:true). Only the box shows it; no ssh allowed. | 1 |
|
||||
| R-652 | P4 | STILL-TRUE-NOT-SMALL | app-catalog@917a779 templates/romm/.felhom.yml:249 still carries `"memory_peak_pct": 80.9` with `"memory_tight": true` and no `memory_basis: anon` (contrast paperless-ngx/.felhom.yml:249 `"memory_basis": "anon"`). Needs a live re-measure of romm plus an undecided cache-thrash rule. | 4 |
|
||||
| R-654 | P4 | NOT-WORTH-IT | PICK close-as-accepted: opengist 1.15 moved its pages under /-/; an old /login bookmark answers 404. The front page redirects correctly. | 3 |
|
||||
| R-687 | P4 | NOT-WORTH-IT | PICK close-as-accepted: Three live proofs a scratch box cannot give (a 3-hour leg, a failing off-site leg, a files_may_change step without a whole copy) plus one log text that names the window's deadline instead of a manually started leg's deadline. | 5 |
|
||||
| R-688 | P4 | NOT-WORTH-IT | PICK close-as-accepted: Deleting a customer does not remove their Cloudflare tunnel and DNS records; the dialog now says so and lists what to remove by hand. | 4 |
|
||||
| R-691 | P4 | STILL-TRUE-NOT-SMALL | Open work = the Tier 2 (second-drive) path of Use/Load is not live-proven; needs a two-drive Tier-0 box (9202 has one drive). Live-only gap, not a source defect; not checkable from source. | 2 |
|
||||
| R-693 | P4 | STILL-TRUE-NOT-SMALL | grep for memory_scales_with_limit / scales_with_limit across app-catalog-felhom.eu, felhom-controller/controller and felhom.eu/scripts returns nothing - no basis that tells growth-to-fill from pressure exists. Needs a design (new harness signal). | 3 |
|
||||
| R-705 | P4 | FIXED-BY-LATER-WORK | The remaining half (manual whole-guest backup) EXISTS and predates the row: felhom-controller bbed5af (v0.47.0, 2026-06-12) 'backups page — whole-guest backup visibility + manual trigger'. Live source @7690c27: controller/internal/web/backup_handlers.go:324 `case r.URL.Path == "/api/guest-backup/trigger" && r.Method == http.MethodPost:`; :340 `if err := s.backupTrigger.TriggerNow(); err != nil {`; | 8 |
|
||||
| R-707 | P4 | STILL-TRUE-NOT-SMALL | Open work = live proof that the gate OPENS for seerr, outline and rallly (needs a media server / e-mail on a test box). Live-only proof gap; the gating itself is in source (catalog 6faf432 per row). | 2 |
|
||||
| R-718 | P4 | STILL-TRUE-SMALL | FIX: Add a key app_info.close_signup_restart ("The app restarts once for this." / HU twin) and render it under the close card when the app has an after_setup env switch (the same SignupNative fact); add the same sentence to the gate-open confirmation where after_setup.env exists. — felhom-controller@7690c27 controller/internal/web/templates/app_info.html:112 `<p>{{T "app_info.close_signup_text"}}< | 6 |
|
||||
| R-719 | P4 | STILL-TRUE-NOT-SMALL | Built: felhom.eu hub/internal/web/selfbind.go:160 "// R-719 (v0.126.0): „Új linket kérek" on an expired or used link." Open part is the operator's review of the changed shape and a live mint+send proof (unit-proven only) - an operator decision, not a CC fix. | 3 |
|
||||
| R-725 | P4 | STILL-TRUE-SMALL | FIX: Reword bind.invalid.body to point at the button below (e.g. 'Ha lejárt, kérj újat lent.' / 'If it has expired, ask for a new one below.'), keeping the operator alternative; optionally drop the ✔ glyph the console font renders as 'V' and update the golden. — felhom.eu hub/internal/i18n/locales/hu.json:79 `"bind.invalid.body": "A hivatkozás 7 napig érvényes. Ha lejárt, kérj újat az ügyfélszolgá | 6 |
|
||||
| R-731 | P4 | STILL-TRUE-NOT-SMALL | The shape-switch control lives only in audit tools: felhom.eu/documentation/audits/catalog-currency-2026-09-30/00-currency.py, 04-analyse.py; no standing currency script in app-catalog-felhom.eu/scripts or felhom.eu/scripts (ls/grep for currency/shape returns only golden_currency_gate.py, which is unrelated). Making it standing means promoting a registry-reading tool with tests - more than an hour | 4 |
|
||||
| R-734 | P4 | STILL-TRUE-NOT-SMALL | grep for '.immich' / hash ignore list in app-catalog-felhom.eu/scripts/*.py and templates/immich/.felhom.yml returns nothing - no exclusion exists. The row says the rule change needs an operator word; calibre-web shows the mark is sometimes right, so the rule needs design. | 3 |
|
||||
| R-739 | P4 | STILL-TRUE-NOT-SMALL | app-catalog@917a779 templates/wanderer/docker-compose.yml:117 `image: getmeili/meilisearch:v1.36.0`; grep MEILI_UPGRADE_DB in the compose returns nothing. Remaining: the template switch, a fixture (PocketBase create refused) and a measured step on the bench - live work. | 3 |
|
||||
| R-759 | P4 | STILL-TRUE-NOT-SMALL | Five checklist rows of wger need live measurement on 9202 (2.5, 3.7, 6.3, 8.2, 9.1); not verifiable from source. | 2 |
|
||||
| R-760 | P4 | STILL-TRUE-SMALL | FIX: Read the vikunja 2.6.0 image config for a HEALTHCHECK/shell; if none and the image has no shell, add a comment saying why there is no compose healthcheck (like adventurelog-frontend's R-655 comment); otherwise add a healthcheck of the family the image supports (REUSE.md §2). No image: line moves, so no catalog_since. — app-catalog@917a779 templates/vikunja/docker-compose.yml: service vikunja | 4 |
|
||||
| R-761 | P4 | STILL-TRUE-SMALL | FIX: Change the comment to name `{slug}-logo.svg` (preferred) and `{slug}-logo.png` (fallback), matching config.go AppLogoURL/AppLogoPNGURL; comment-only. — app-catalog@917a779 templates/paperless-ngx/.felhom.yml:22 `# Logo: {assets.base_url}/assets/{slug}-logo.webp` vs felhom-controller controller/internal/config/config.go:511 `return fmt.Sprintf("/static/assets/%s-logo.svg", slug)` and | 3 |
|
||||
| R-764 | P4 | STILL-TRUE-NOT-SMALL | grep smtp/mail in app-catalog templates/wger/.felhom.yml and docker-compose.yml finds only first_steps text (.felhom.yml:74 'Add meg az email címedet a beállításokban'); no smtp_mapping. A mapping needs a live boot proof with mail off (REUSE.md §2) - more than an hour; wger is hidden. | 3 |
|
||||
| R-766 | P4 | FIXED-BY-LATER-WORK | Hub releases after the assets push (felhom.eu 40f07429, 2026-10-01): hub v0.131.0 (2026-10-04) .. v0.136.0 (d4be9f6f, 2026-10-05). The build copies website assets: felhom.eu scripts/build-hub.sh:98 `cp "${WEBSITE_ASSETS_DIR}"/*-logo.svg "${BUILD_DIR}/assets/" 2>/dev/null // true`, and the hub build workspace /mnt/5_hdd/felhom.eu/build/felhom-hub/workspace/assets/ holds radicale-logo.svg + 3 screen | 8 |
|
||||
| R-768 | P4 | NOT-WORTH-IT | PICK close-as-accepted: Grimoire is not built because upstream rules out public exposure and ships no image for v1.x; the row only watches for that to change. | 1 |
|
||||
| R-769 | P4 | STILL-TRUE-NOT-SMALL | New-app idea waiting on an operator decision (a fork as new upstream, after R-767). No source to check. | 1 |
|
||||
| R-770 | P4 | STILL-TRUE-NOT-SMALL | New-app idea waiting on the operator's go/no-go (CC recommends not building). No source to check. | 1 |
|
||||
| R-771 | P4 | STILL-TRUE-NOT-SMALL | New-app idea waiting on the operator's go/no-go. No source to check. | 1 |
|
||||
| R-779 | P4 | STILL-TRUE-NOT-SMALL | Live proof gap on the real Cloudflare tunnel needing the operator's phone off wifi; not checkable from source. | 1 |
|
||||
| R-781 | P4 | STILL-TRUE-SMALL | FIX: After the clone, delete the real onboarding/<app>.md records (all but _TEMPLATE.md and the exempt wger.md the cases use) from the scratch clone, so genuine cases judge only the records they build; alternative: point the stand-in sibling at the real felhom.eu documentation tree read-only. — app-catalog@917a779 scripts/test_gate_decoys.py:498 `sh(["git", "clone", "-q", "file://" + ROOT, cat], c | 6 |
|
||||
| R-786 | P4 | STILL-TRUE-NOT-SMALL | app-catalog@917a779 onboarding/sparkyfitness.md has 8 '/ open' rows, incl. :18 0.5, :20 0.7, :28 1.6, :29 1.7, :58 5.4, :68 8.3 (plus 0.1 licence = R-784, 2.1 = R-807). Most need live measurement (runtime internet, phone sign-in, second memory watch). | 3 |
|
||||
| R-793 | P4 | NOT-WORTH-IT | PICK close-as-accepted: Four apps ship enterprise/BUSL code that is off as Felhom runs them; the row reminds us never to enable EE features or the -enterprise meilisearch image. | 2 |
|
||||
| R-794 | P4 | STILL-TRUE-NOT-SMALL | app-catalog@917a779 seven redis 7 images: dawarich/docker-compose.yml:164 `image: redis:7.4-alpine`, docmost:86, immich:123, outline:88, nextcloud:103, paperless-ngx:125, romm:136 `image: redis:7-alpine`. Moving each needs a harness-proven ladder step (7 apps). | 3 |
|
||||
| R-796 | P4 | NOT-WORTH-IT | PICK close-as-accepted: MeTube's browser/phone 'send to MeTube' helpers cannot pass the family gate; households paste links in the page. | 2 |
|
||||
| R-797 | P4 | NOT-WORTH-IT | PICK close-as-accepted: CI's single-repo clone cannot check rule 3 of the family-gate gate; it says NOT CHECKED, and the pre-push hook checks it. | 2 |
|
||||
| R-798 | P4 | STILL-TRUE-SMALL | FIX: Remove the dead SWAGGER_ENABLED line (or rename to API_DOCS_ENABLED=false, which v3.5.0 reads and defaults to false). No image: line moves. — app-catalog@917a779 templates/grimmory/docker-compose.yml:29 ` - SWAGGER_ENABLED=false` still present. | 2 |
|
||||
| R-799 | P4 | STILL-TRUE-SMALL | FIX: Add "download_type": "video" to the POST /add body. — app-catalog@917a779 scripts/upgrade_fixtures_box.py:2183 `data=json.dumps({"url": self.URL, "quality": "best", "format": "any", "auto_start": True}), method="POST")` - no download_type. | 2 |
|
||||
| R-804 | P4 | NOT-WORTH-IT | PICK close-as-accepted: plant-it's image repository does not exist; the template is already abandoned and not installable, and no box runs it. | 2 |
|
||||
| R-805 | P4 | STILL-TRUE-NOT-SMALL | app-catalog-felhom.eu scripts/check-volume-persistence.py:342 'if m["class"] == "named-declared" and m.get("files", 0) == 0:' — only named volumes judged empty; binds not. Last change 917a779 (R-788). The rule change itself is small, but it flips Grimmory/komga/paperless-ngx/radarr/sonarr to UNDETERMINED and needs a live re-sweep to regenerate the verdict tables; the row also asks for a decision. | 6 |
|
||||
| R-806 | P4 | STILL-TRUE-SMALL | FIX: Make routed_ports return the traefik loadbalancer.server.scheme (reuse upgrade_boxport LB_SCHEME_RE) and have the GET exercise use that scheme with curl -k for https. The gramps-web :5000 non-answer stays a separate live look (narrow the row to it). — app-catalog-felhom.eu scripts/check-volume-persistence.py:586 'code = _sh(a + [f"http://{ip}:{port}{path}"], timeout=40)' — plain http always; | 6 |
|
||||
| R-807 | P4 | STILL-TRUE-NOT-SMALL | Per-app upload seeds for 13 apps (claper, crafty-controller, dawarich, docmost, gramps-web, immich, outline, sparkyfitness, tandoor, vikunja, wger, wishlist, zipline) + plex/wanderer; each needs a live fixture run. Gate rule at scripts/check-volume-persistence.py:342 still makes empty declared volumes UNDETERMINED (917a779). | 3 |
|
||||
| R-814 | P4 | UNCHECKED | Live Hetzner console state (box 611421 status); not visible in source. Operator action only. | 2 |
|
||||
| R-815 | P4 | UNCHECKED | First GC completion on felhom-offsite is PBS server-side live state; grep of documentation found no GC completion record (DIAG-backup-missed-2026-07-26.md:43 'prune/GC history NOT COLLECTED'). | 3 |
|
||||
| R-816 | P4 | STILL-TRUE-NOT-SMALL | Needs a live exercise of six failure classes on a scratch guest; no source change can close it. | 2 |
|
||||
| R-817 | P4 | STILL-TRUE-SMALL | FIX: Add a dated correction under decision 56: the swap rolls back to the image running when the swap began (controllerswap.go Swap/rollback); the kept previous image is for a hand roll-back. Do not rewrite the ruling itself. — felhom.eu documentation/architecture/09-update-architecture.md:619 '56. **A box keeps the controller image it runs and the one before it** (the self-update's roll-back targ | 8 |
|
||||
| R-818 | P4 | STILL-TRUE-SMALL | FIX: Add a dated correction note under the v0.109.0 entry (and the hub CHANGELOG mentions at :695/:701) naming the real closed rows from CLOSED-ITEMS.md. Also present in felhom-controller/CHANGELOG.md:3107 and :3234 (R-330/R-331 for v0.224.0/v0.225.0) — a second repo; either note it there too or narrow the row. — felhom.eu hub/CHANGELOG.md:759 '## v0.109.0 — the Backup card told every operator tha | 6 |
|
||||
| R-819 | P4 | STILL-TRUE-SMALL | FIX: Let rule 3 accept an id found in CLOSED-ITEMS.md (and check the stand's status agrees), fix the R-273/R-356 dangling ids, register the gate in repo_gates.py with a decoy. — Ran python3 scripts/check_stands.py (read-only): still convicts e.g. 'fail.stolen-machine: register id R-281 is not in OPEN-ITEMS.md', 'fail.customer-self-restore: register id R-356 ...'. grep 'stands' in scripts/repo_gate | 6 |
|
||||
| R-832 | P4 | STILL-TRUE-NOT-SMALL | Deferred roadmap item: a third-location copy of ep0 is money + operator decision (decision 71). Nothing in source to change. | 1 |
|
||||
| R-844 | P4 | STILL-TRUE-NOT-SMALL | Needs a household timeline on the controller, which does not exist (row: 'when the box gets a household timeline'). A new surface, not a fix. | 2 |
|
||||
| R-855 | P4 | STILL-TRUE-SMALL | FIX: Add a small helper (e.g. osSvc.DockerNightsEffective()) mapping negative→0 and 0→2, and print that in the start log. — felhom.eu hub/cmd/hub/main.go:450 'logger.Printf("[INFO] osupdates: the Docker engine set is approved only by the operator, after %d healthy ring-0 night(s)", osSvc.DockerNights)' prints the raw value; internal/osupdates/service.go:169 'negative means none'. | 4 |
|
||||
| R-856 | P4 | STILL-TRUE-NOT-SMALL | Row is marked an operator design question (crash-restart suppression for app mails); not a defect yet. | 1 |
|
||||
| R-857 | P4 | STILL-TRUE-SMALL | FIX: Have newest_baked accept an optional suffix after the date and, for equal versions, prefer the newest bake-log timestamp (or refuse two dirs for one version); add 're-vouch at once after a same-version re-bake' to the runbook. — felhom.eu scripts/golden_currency_gate.py:146 'EVIDENCE_RE = re.compile(r"^golden-(\d+)\.(\d+)\.(\d+)-\d{4}-\d{2}-\d{2}$")' and :237-238 'found.sort() / return found[ | 7 |
|
||||
| R-878 | P4 | STILL-TRUE-NOT-SMALL | Next action is 'measure a large volume first' — a live measurement; the fix direction is a behaviour change to the catch-up. | 2 |
|
||||
| R-881 | P4 | STILL-TRUE-SMALL | FIX: Add an rm -f of /usr/local/sbin/felhom-priv-apply to the uninstall step (tolerate-absent, like felhom-pbs-apply at :1161) and fix the :1679 comment; ships at the next installer tag. — felhom.eu scripts/felhom-host-install.sh: grep 'felhom-priv-apply' → no hit anywhere in the installer (uninstall does not remove it); :1679 '+ guest-hook snippet under /var/lib/vz/snippets/ (agent-installed at r | 4 |
|
||||
| R-884 | P4 | UNCHECKED | Live ArgoCD diff on DooPlex (forbidden to touch here). Related: homelab-manifests mon-system/monitoring.yaml:399-426 is the same prometheus Deployment R-211 concerns. | 2 |
|
||||
| R-885 | P4 | STILL-TRUE-SMALL | FIX: Finish and push the in-progress script_tests_gate.py (walks scripts/ for test_*.py, exit-code verdict, nesting guard) registered in repo_gates.py; coordinate with the session that owns the dirty tree. — On main (b018ca90) scripts/repo_gates.py runs no test_*.py suite. NOTE: the felhom.eu working tree holds UNCOMMITTED work for exactly this row by another session: '?? scripts/script_tests_gate | 6 |
|
||||
| R-30 | P3 | STILL-TRUE-NOT-SMALL | Design change (presence from the Dir-2 long-poll instead of the report clock), size M; no commit with R-30 after aa9c08f0 (filing). | 3 |
|
||||
| R-31 | P3 | STILL-TRUE-NOT-SMALL | felhom.eu hub/internal/web/configs.go:1594 'd, err := s.offsite.ProvisionOffsite(ctx, cfg.CustomerID, in)' still in-request; :1584 detaches from the request context (mid-cancel fixed) but no async/status card. | 5 |
|
||||
| R-35 | P3 | STILL-TRUE-NOT-SMALL | felhom-controller controller/internal/report/config_refresh.go:65 'config-refresh: applied config_version=%d — self-restarting to load it'; sessions are in-memory only: controller/internal/web/auth.go:259-260 's.sessions[token] = &session{'. Hot-apply or persisted sessions is a design change with security weight. | 6 |
|
||||
| R-49 | P3 | STILL-TRUE-NOT-SMALL | app-catalog-felhom.eu templates/immich/.felhom.yml:28-34 backup block has no cache/volume exclusion; row itself says a capture-set exclusion needs its own ruling (data-loss-shaped). | 4 |
|
||||
| R-50b | P3 | FIXED-BY-LATER-WORK | Claim 'fetched via fetch_raw from raw/branch/main — no tag, no pin' no longer true: bee68484 (installer v1.23.0, R-110/R-183) pinned fetch_raw to the vouched agent tag — felhom.eu scripts/felhom-host-install.sh:533 '"$GITEA_BASE/$GITEA_OWNER/$AGENT_REPO/raw/tag/v$ART_AGENT_VER/$path" \' (leg b). Leg (c)-like signed delivery: felhom-agent c9fa2e7 (R-840 config bundle) and configs/test_felhom_config | 7 |
|
||||
| R-78 | P3 | STILL-TRUE-NOT-SMALL | An owed operator decision + spike (local_api authority); not a code defect. | 1 |
|
||||
| R-79 | P3 | STILL-TRUE-NOT-SMALL | felhom-controller controller/internal/monitor/healthcheck.go:100 'fmt.Sprintf("SSD disk usage critical: %.0f%%"', :213 'Protected container not running: %s'; rendered raw at controller/internal/web/alerts.go:241 'Message: issue, // ON THE WIRE ... not ours to translate; slice 3'. Whole-surface, seam needs a spike. | 5 |
|
||||
| R-118 | P3 | STILL-TRUE-SMALL | FIX: Only statfs the mount when s.devicePresent(d.MountPath) is true (else leave capacity zero/unknown); put statfsCapacity behind a seam var for the test. — felhom-agent internal/localapi/disks.go:401 'if total, used, okc := statfsCapacity(d.MountPath); okc {' — no device-presence guard on the union path; devicePresent exists (disks.go:985) and is used just above (:381) for BoundUnderParent. | 6 |
|
||||
| R-121 | P3 | FIXED-BY-LATER-WORK | 3d7a2761 (hub v0.135.0, R-530/R-604 'boxes left behind listed and alarmed'): felhom.eu hub/internal/osupdates/service.go:79 'EventAgentBehind = "agent_behind" // warning, operator' with :180 'AgentBehindAfter: a box runs an agent older than the vouched one this long → an operator alarm' (7 d window, the staleness window the row asked for). | 5 |
|
||||
| R-126 | P3 | STILL-TRUE-SMALL | FIX: Skip IsNetwork() paths in the export-destination list and refuse them in isValidDrivePath for the export POST (keep scanning for import if wanted, via a separate list). — felhom-controller controller/internal/web/handler_export.go:377-386 storageDriveList() appends every s.settings.GetStoragePaths() entry with no IsNetwork() filter; the predicate exists at controller/internal/settings/setting | 6 |
|
||||
| R-127 | P3 | STILL-TRUE-NOT-SMALL | Leg (a) still true: grep 'data_key: true' in app-catalog-felhom.eu templates → only adventurelog, dawarich, homebox, papra, sparkyfitness; n8n N8N_ENCRYPTION_KEY, wanderer POCKETBASE_ENCRYPTION_KEY, calcom CALENDSO_ENCRYPTION_KEY, bookstack APP_KEY unflagged (templates/n8n/.felhom.yml:38 etc.). Leg (b) (regenerated DB password vs restored PGDATA) needs a design choice. Leg (a) alone is a ~45-min c | 6 |
|
||||
| R-130 | P3 | STILL-TRUE-SMALL | FIX: Take the cheap honest branch: rename to RECOMMENDED_MIN_LVM_GIB and reword the warning to 'below the recommended …' (making it refuse would change install behaviour and needs a ruling). — felhom.eu scripts/felhom-host-install.sh:348 'HARD_MIN_LVM_GIB=120 # a useful appliance won't fit below this on local-lvm' and :1760 '... // log_warn "local-lvm free ~${free_gib} GiB < hard min ${HARD_MIN_ | 4 |
|
||||
| R-132 | P3 | UNCHECKED | Whether HUB_PW was rotated is out-of-band operator state; not visible in source. | 1 |
|
||||
| R-136 | P3 | STILL-TRUE-SMALL | FIX: Introduce a const sessionCookieName = "__Host-hub_session" and use it at all five sites (Path=/, Secure, no Domain already hold). Every operator logs in once more; plain-HTTP browser access stops (Basic auth unaffected). — felhom.eu hub/internal/web/server.go:857 'Name: "hub_session",' and readers at server.go:806, :896, :918 and apps.go:351. | 4 |
|
||||
| R-137 | P3 | STILL-TRUE-NOT-SMALL | felhom-controller controller/internal/cloudflare/waf.go:18 'globalRuleDesc = "[felhom-geo] Global"', :21 'appRuleDescPrefix = "[felhom-geo] app:"' — still not namespaced. Two-repo M change. | 3 |
|
||||
| R-138 | P3 | STILL-TRUE-NOT-SMALL | felhom-controller controller/internal/infra/infra.go:157-159 writes CF_DNS_API_TOKEN when d.CFAPIToken != ""; no shared-zone guard in hub (grep shared.zone: none). Needs a policy decision first; no shared zone exists today. | 4 |
|
||||
| R-179 | P3 | STILL-TRUE-SMALL | FIX: In uninstall step 4c, before the umount loop: stop+disable every mnt-felhom\x2ddrives-*.automount/.mount unit, rm their files from /etc/systemd/system, daemon-reload (tolerate-absent; still never umount -l/-f). — felhom.eu scripts/felhom-host-install.sh uninstall section :1121-1157 handles felhom-shared-parent and umounts under /mnt/felhom-drives, but grep 'x2ddrives/automount' → no hit: the | 6 |
|
||||
| R-180 | P3 | STILL-TRUE-SMALL | FIX: In the same pre-flight block, die (byo) / die (appliance) if ARCHIVE_STORAGE is not in PVE_STORAGES, with a message naming --acl-storages. — felhom.eu scripts/felhom-host-install.sh:1800 'if pvesm status --storage "$ARCHIVE_STORAGE" ...' checks existence only; PVE_STORAGES=(local local-lvm felhom-pbs) at :322; no ARCHIVE_STORAGE ∈ PVE_STORAGES assertion. | 5 |
|
||||
| R-190 | P3 | NOT-WORTH-IT | PICK close-as-accepted: A storage permission vanished once on demo-felhom in August and nobody knows why. Since agent 0.124.1 the agent puts it back by itself and mails the operator when it happens. | 4 |
|
||||
| R-200 | P3 | FIXED-BY-LATER-WORK | The remaining half (customer-facing recovery-code form: yell → R form → preview) shipped as the recovery screen: felhom-controller 636c51e 'R-193: the recovery screen — unlocking, and only unlocking (v0.200.0)'; controller/internal/web/templates/recovery.html:82 '<form id="unlock-form" method="POST" action="/recovery/unlock" autocomplete="off">', routed at internal/web/server.go:602. Plumbing half | 8 |
|
||||
| R-211 | P3 | STILL-TRUE-NOT-SMALL | homelab-manifests (/home/kisfenyo/git/homelab-manifests @87dfc29) mon-system/monitoring.yaml:420 'image: prom/prometheus:v3.15.0', :426 '--web.enable-lifecycle'; grep 'reload/checksum/config' → none. The manifest edit is small, but it rolls the production Prometheus on DooPlex (operator territory) and the same Deployment is OutOfSync per R-884 — do the two together. | 6 |
|
||||
| R-231 | P3 | STILL-TRUE-NOT-SMALL | Owner operator; DooPlex /opt/backup/scripts remains host state. Partially touched by cea8502f (scripts/hub-db-backup versioned in felhom.eu, cites R-231) but that covers only the hub-DB push, not /opt/backup/scripts or the same-disk/no-off-site facts. | 4 |
|
||||
| R-235 | P3 | FIXED-BY-LATER-WORK | felhom.eu c033b3b6 'ISO 1.28.0 source: the console stops showing the pairing code once bound (R-535)'. scripts/iso/felhom-bootstrap.sh:538: `print_bound_banner # R-535: replace the pairing code on the console with the truth`. Same defect already CLOSED twice in CLOSED-ITEMS.md as R-535 (line 232) and R-214 (line 200, 'proven on a fresh install'). | 4 |
|
||||
| R-240 | P3 | STILL-TRUE-SMALL | FIX: Replace the producer string at offbox.go:1168 with wording that drops 'Sikeres' but keeps the lowercase marker substring, e.g. 'Ez a futás semmit nem mentett: nincs mentésre jelölt alkalmazás'; update the tests that pin the literal. — felhom-controller/controller/internal/backup/offbox.go:1168: `warns = append(warns, "Sikeres — nincs mentésre jelölt alkalmazás")`; web/handlers.go:1068-1069 re | 8 |
|
||||
| R-242 | P3 | STILL-TRUE-NOT-SMALL | felhom.eu/scripts/golden_currency_gate.py:23-24: 'It does **NOT** check that the golden was **VOUCHED**, because the vouched version lives ONLY in the hub's `hub_settings` table'; :40 'That vouch half is STILL open after 2026-09-13'. Last gate commits 5ef0f52b/ae59c31a did not add a vouch check. | 5 |
|
||||
| R-244 | P3 | STILL-TRUE-NOT-SMALL | felhom.eu/hub: no cascade leg touches app_log_issues — only writers are internal/store/telemetry.go:176-209 (upsert), :537 `DELETE FROM app_log_issues WHERE last_seen < ?`, :556/:575 operator deletes by app/id. cmd/hub/main.go:1004: `if n, err := s.PruneStaleIssues(time.Now().Add(-30 * 24 * time.Hour))` (since a757bee0, hub v0.4.0). | 7 |
|
||||
| R-250 | P3 | STILL-TRUE-NOT-SMALL | felhom.eu/hub/internal/offsite/offsite.go:71-72: `var defaultScanBackoff = []time.Duration{2 * time.Second, 4 * time.Second, 8 * time.Second, 16 * time.Second, 30 * time.Second}`; scanner.go:40 dials plain "tcp" (no A-record preference); no R-250 commit. | 7 |
|
||||
| R-251 | P3 | STILL-TRUE-SMALL | FIX: In offsiteNewestPerTag skip the 'felhom-offbox' marker tag (move the constant into package backup and reuse it from web); consider whether '_shares' should render as a named row or be skipped. — felhom-controller/controller/internal/backup/offbox_inventory.go:102-108: loops `for _, tag := range sn.Tags` and adds every non-empty tag to `newest` — no filter for 'felhom-offbox'. The marker filte | 9 |
|
||||
| R-255 | P3 | STILL-TRUE-NOT-SMALL | felhom-controller: still no page-wide runtime secret-sentinel test — `grep -l sentinel internal/web/*_test.go` hits only edge_safe_status_test.go, i18n_parity_test.go, recovery_test.go; controller/scripts/secret_in_markup_gate.py remains the only all-template net. No commit cites R-255. | 7 |
|
||||
| R-257 | P3 | STILL-TRUE-SMALL | FIX: Rewrite flash.offbox.not_orphaned (hu + en) to say what the customer tried, that it does not apply now, and where to look, without 'offsite'/'elárvult', e.g. 'A távoli mentés rendben van, nincs mit félretenni. Ha gondod van vele, írj nekünk.'; wording sign-off from the operator (owner Viktor). — felhom-controller/controller/internal/i18n/locales/hu.json:1409: `"flash.offbox.not_orphaned": "Az | 5 |
|
||||
| R-262 | P3 | STILL-TRUE-SMALL | FIX: One-repo fix: narrow the comment to say hostBackup is field-for-field and hostRestoreTest is a deliberate SUBSET (lists mount_parity/mount_inventory as not modelled), and add a hub test that names the two agent fields as known-unmodelled so a future addition must edit it. Adding the fields + fixture is a two-repo change (byte-identical golden) and stays a separate choice. — felhom.eu/hub/inte | 7 |
|
||||
| R-269 | P3 | STILL-TRUE-SMALL | FIX: On a map hit, also stat the store and reload when its size differs from loadedSize before answering (one cheap stat per auth), so a rotated-out token is rejected without depending on an unrelated miss. — felhom-agent/internal/localapi/tokenstore.go:173-176: `if vmid, ok := s.byHash[want]; ok { if subtle.ConstantTimeCompare(...) == 1 { return vmid, true } }` — a superseded token's hash is stil | 6 |
|
||||
| R-270 | P3 | STILL-TRUE-NOT-SMALL | felhom-controller/controller/internal/bootstrap/bootstrap.go:255: `if cfg == nil // cfg.LocalAPI.Endpoint != "" {` (fill-only, never refreshes); DetectEndpointDrift (bootstrap.go:369-399) compares only the endpoint. No R-270 commit. | 5 |
|
||||
| R-271 | P3 | STILL-TRUE-NOT-SMALL | felhom-controller/controller/internal/channelhealth/checker.go:152: `if prev != "" && prev != "up" {` (unseeded->up is silent); checker.go:87 alert text still says '(re-bootstrap)'. No R-271 commit. | 4 |
|
||||
| R-274 | P3 | FIXED-BY-LATER-WORK | felhom.eu eb600872 'R-297: installer compares a local golden against the manifest before using it'. scripts/felhom-host-install.sh:3074: `if golden_local_matches_manifest "$GOLDEN_VOLID"; then` (digest vs ART_GOLDEN_SHA, else baked marker vs ART_GOLDEN_VER; otherwise ignores the local golden and fetches, or dies if the operator named it, :3080). Line numbers differ from the triage note (2855-2905) | 5 |
|
||||
| R-275 | P3 | STILL-TRUE-SMALL | FIX: Purge every sibling `"${agent_cfg}".*` (and then rmdir/rm the config dir) instead of `.bak*`; fix the WIPED line wording. — felhom.eu/scripts/felhom-host-install.sh:1059: `for _cfgbak in "${agent_cfg}".bak*; do [[ -e "$_cfgbak" ]] && run rm -f "$_cfgbak"; done` — glob still misses agent.json.campaign8-before, .campaign9-prev, .pre-e-target-move, .pre-prunegate.bak; :814 still claims 'config ( | 6 |
|
||||
| R-276 | P3 | STILL-TRUE-SMALL | FIX: In run_uninstall (full scope): stop+disable wg-quick@wg-felhom and remove /etc/wireguard/wg-felhom.conf via run(); add it to WIPED, and add a KEPT line 'the hub-side WireGuard peer registration — remove it in the operator UI'. — felhom.eu/scripts/felhom-host-install.sh: no reference to wg-felhom/wg-quick anywhere (grep 'wg-quick\/wg-felhom' empty); _uninstall_statement (:807-845) lists neithe | 6 |
|
||||
| R-277 | P3 | STILL-TRUE-SMALL | FIX: For UsageStr use fmtBytesAuto when the usage is below 1 GB (keep GB for the quota and the bar), so a non-empty repo never renders as 0.0 GB. — Part (a) FIXED by felhom.eu f5c9411e 'R-331 (hub half): the Backup card reads `offsite`, not the dead `backup` fields (v0.109.0)' (hub/internal/web/backup_card.go uses fmtBytesAuto :135-142). Part (b) still true: hub/internal/web/offsite_box.go:54 `ret | 8 |
|
||||
| R-282 | P3 | FIXED-BY-LATER-WORK | felhom.eu 4d6ec7c 'hub v0.104.0: ... the hub half of the naming (R-295)' + controller v0.211.0 (R-295 CLOSED, CLOSED-ITEMS.md:207) + R-323 hub v0.105.0. hub/internal/notify/templates.go:204: '// R-295, HUB HALF (2026-08-13). ONE NAME PER SECRET, and it is „Beállító kód".'; hub/internal/claim/engine.go:51 `EmailReenroll EmailKind = "reenroll"` (mail names the setup page a rebuilt box shows). | 5 |
|
||||
| R-283 | P3 | STILL-TRUE-NOT-SMALL | felhom.eu/hub/internal/claim/engine.go:7: '// engine: a rotation bumps the generation (single active code) and NEVER clears claimed_at.'; ReissueForReenroll (engine.go:195-212) rotates and mails but leaves the claim set — the hub still shows the customer as claimed after a guest rebuild. No R-283 commit. | 5 |
|
||||
| R-298 | P3 | UNCHECKED | The template gate is still there: felhom-controller/controller/internal/web/templates/storage.html:364 `if(d.role==='user-data'){` else protected (:368). BUT the agent no longer reclassifies a backup-target drive: felhom-agent/internal/localapi/disks.go:1230-1236 ('WHY THIS IS NOT A ROLE RECLASSIFICATION ... the drive that now holds the whole-guest archives is ALSO the enrolled user-data drive') a | 10 |
|
||||
| R-306 | P3 | STILL-TRUE-SMALL | FIX: Make _state_put (and _state_mark) return 0 when PREFLIGHT_ONLY is true, so a preflight-only run writes nothing; the real run's preflight records ownership again. — felhom.eu/scripts/felhom-host-install.sh:418: `$DRY_RUN && return 0` (only DRY_RUN short-circuits _state_put); :1851/:1854 `_state_put dnsmasq_preexisting yes/no` run unguarded in preflight, while :226 says '--preflight-only: ... n | 6 |
|
||||
| R-314 | P3 | STILL-TRUE-NOT-SMALL | felhom-controller: StopAbandon is called only from cmd/controller/main.go:244 (CLI) — `grep -rn StopAbandon` shows internal/backup/offbox_abandon.go:374 and that CLI call, no web handler. | 4 |
|
||||
| R-317 | P3 | STILL-TRUE-SMALL | FIX: Probe the dnsmasq unit (e.g. /usr/lib/systemd/system/dnsmasq.service or /lib/systemd/system/dnsmasq.service) behind a small stat seam instead of /usr/sbin/dnsmasq. — felhom-agent/internal/lanresolver/lanresolver.go:107: `if _, err := os.Stat("/usr/sbin/dnsmasq"); err != nil { // metadata read, no privilege needed` — still probes the dnsmasq-base file. No R-317 commit. | 4 |
|
||||
| R-330 | P3 | STILL-TRUE-NOT-SMALL | felhom-agent/internal/hub/report.go:406-425 SmartSummary carries reallocated/pending/offline_uncorrectable + NVMe set only — no 187/188/199 fields. Wire change across agent + hub (+ controller), declared M. | 3 |
|
||||
| R-332 | P3 | STILL-TRUE-NOT-SMALL | Closing condition is live-only (a real degrading disk or an injection through agent /disks -> controller -> hub). Last related commits ea16a21b/2fa1efc narrowed the restart half only; no commit records a live Hiba-from-counters verdict. | 3 |
|
||||
| R-333 | P3 | STILL-TRUE-NOT-SMALL | (b) felhom-agent/internal/storage/hostops.go:374: `out, stderr, err := h.runner.Run(ctx, h.bins.Smartctl, "-a", "-j", device)` — no -n standby. (a) still an operator decision (Viktor decides). | 5 |
|
||||
| R-338 | P3 | UNCHECKED | felhom.eu/documentation/operations/nodes.md:86-88 now says demo-hp 'Agent config shape (R-50 island): local_api on 169.254.253.1:8443/vmbr9, guest eth1 169.254.253.2/30', and :84 records demo-hp was reprovisioned (address 192.168.0.87 -> 192.168.0.104, read 2026-09-21). Whether the reprovisioned box is actually on the island is a live-box fact (agent.json, pct config) not provable from source. | 5 |
|
||||
| R-340 | P3 | STILL-TRUE-NOT-SMALL | felhom.eu/scripts/felhom-tenantsync.sh: no health op and no 8007 probe (grep '8007\/health)' empty). Needs an ep0 (protected) script version bump + hub signal (M). | 3 |
|
||||
| R-349 | P3 | STILL-TRUE-NOT-SMALL | felhom-agent/internal/hub/report.go:282-292 reports only WrapperSHA256; no agent binary sha field in the report (grep AgentSHA256 finds only the hub's manifest entry, hub/internal/api/handler.go:2726). R-349 commits 40d857b/910fd911 are the manual correction only. | 4 |
|
||||
| R-350 | P3 | UNCHECKED | Whether the hub operator password was rotated after 2026-08-20 lives only in the hub DB / operator; git log shows no rotation record (only 910fd911 filing it). Not determinable from source. | 3 |
|
||||
| R-362 | P3 | STILL-TRUE-SMALL | FIX: When MkdirAll/write of the restore destination fails with EACCES/ENOENT, check whether the destination's drive is disconnected/decommissioned or no longer a mountpoint, and return a Hungarian error naming the drive instead of the raw permission error. — felhom-controller/controller/internal/backup/offbox_restore.go:380, offbox.go:1929, shares_restore.go:116: `return fmt.Errorf("restore dir: % | 6 |
|
||||
| R-363 | P3 | STILL-TRUE-SMALL | FIX: Replace sched.Daily("fill-watch", "03:30", …) with sched.Every("fill-watch", time.Hour, …) (scheduler.go:104), keep the startup check. — felhom-controller/controller/cmd/controller/main.go:1546: `sched.Daily("fill-watch", "03:30", func(ctx context.Context) error { return fillWatcher.Check() })`. Watcher emits on escalation only with a persisted band (internal/fillwatch/fillwatch.go:133, :231) | 6 |
|
||||
| R-388 | P3 | STILL-TRUE-NOT-SMALL | Product direction, operator's call; recorded as [DESIGN — DIRECTION] in documentation/architecture/08-alarm-ladder.md §8 per the row. Nothing in source to fix. | 2 |
|
||||
| R-401 | P3 | STILL-TRUE-NOT-SMALL | Event-triggered watch row: felhom-controller/controller/internal/backup/offbox_integrity.go:63 `const integrityCheckTimeout = 30 * time.Minute`, :78 `var integritySlowNoticeThreshold = 5 * time.Minute` — unchanged; the trigger (slow WARN on a large store) has not been recorded. | 3 |
|
||||
| R-409 | P3 | STILL-TRUE-NOT-SMALL | felhom-controller/controller/internal/backup/recovery_unit.go:58 `Checksums map[string]string json:"checksums" // sha256 of captured compose/ files`; writers at :147, :152, :202 hash only compose/.felhom.yml/app.yaml — no db_dumps or volume_dumps hash. | 4 |
|
||||
| R-412 | P3 | NOT-WORTH-IT | PICK close-as-accepted: A narrow race: if a recovery unit is destroyed inside an off-site run after its own dump leg, the push ships the just-rebuilt hollow unit; the next run repairs it and the WARN line now says it carried no data. | 4 |
|
||||
| R-433 | P3 | STILL-TRUE-NOT-SMALL | Hetzner answered (ticket per the row): file-level snapshot access is a MAIN-account capability. hub/internal/hetznerapi/hetznerapi.go still has no snapshot read method (only size_snapshots usage, :83-87). The owed proof (read a file from a snapshot with the main account; forced-command append-only key) is a live, credential-bound operator act. | 4 |
|
||||
| R-435 | P3 | STILL-TRUE-NOT-SMALL | felhom.eu/hub/internal/monitor/offsite.go:255: `snapshotDropFraction = 0.5 // more than half the history gone in one step` — no per-tag second signal exists. | 4 |
|
||||
| R-440 | P3 | STILL-TRUE-NOT-SMALL | app-catalog-felhom.eu @917a779: 15 templates still have no `update_ladder:` in .felhom.yml — bentopdf code-server glance gokapi gramps-web homebox homepage jellyfin onlyoffice plant-it plex recipe-importer seerr vaultwarden wanderer (calibre-web got its first step in 53a4a1d). For these, pins float with no recorded digest. | 8 |
|
||||
| R-444 | P3 | STILL-TRUE-NOT-SMALL | No fstrim anywhere in felhom-agent or felhom.eu/scripts (grep 'fstrim' over *.go/*.sh empty). Needs a new periodic host job (agent, privileged, sudoers/bundle) and possibly an operator surface. | 3 |
|
||||
| R-446 | P3 | DUPLICATE | of R-440 — felhom-controller/controller/internal/stacks/updateorder.go:96: `if len(s.CatalogDigests) == 0 // s.CatalogTestedAt.IsZero() { return false }` — blind only for apps with no ladder entry, i.e. the same 15 templates R-440 lists (app-catalog has no update_ladder for them). Both rows close by the same act: each app's first proven ladder step (R-462). | 6 |
|
||||
| R-450 | P3 | FIXED-BY-LATER-WORK | The row's only remainder was 'the other ten PostgreSQL apps need two-venue proof'. R-463 CLOSED 2026-09-30 by felhom.eu 25cb3eb9 ('The last six PostgreSQL apps decided'): 8 of 11 moved by the box's own conversion, 3 (zipline, adventurelog, immich) stay by decision 42; CLOSED-ITEMS.md:223. Source proof: app-catalog templates/docmost/docker-compose.yml:62 'image: postgres:18-alpine' (also rallly:67, | 8 |
|
||||
| R-458 | P3 | NOT-WORTH-IT | PICK close-as-accepted: A frozen (pinned-behind) app can receive a newer health check from .felhom.yml. The only result is a false 'degraded'/dead-app alarm, never data loss, and only for type: api probes with an expect block. | 12 |
|
||||
| R-462 | P3 | STILL-TRUE-NOT-SMALL | Ongoing multi-session work (fixtures and ladders per app). Last progress: catalog e6f3ec2 (2026-09-30), audits/more-night-apps-2026-09-30/. 21 apps still have no ladder (per the row). Not checkable as done from source. | 3 |
|
||||
| R-468 | P3 | STILL-TRUE-NOT-SMALL | A standing pre-customer arrangement, not a defect. The mechanism is live: felhom.eu/scripts/golden_currency_gate.py:137 reads documentation/tests/golden-waiver.yml. The waiver file is currently ABSENT: deleted in felhom.eu 5efe6dae (2026-10-04, 'golden 0.292.0 vouched ... golden waiver deleted'). The row retires only at the first external install, so it stays as a watch. | 4 |
|
||||
| R-469 | P3 | STILL-TRUE-SMALL | FIX: Reword the rule heading at CLAUDE.md:103 and the 'What is NOT lifted' paragraph (:112-114) to say that PostgreSQL crosses a major one app at a time with engine_conversion + both-venue proof (decision 35) and that MySQL stays refused. Then close R-469 citing R-463's closure. Do not delete the gate: it is now the per-app enforcement. — What stood between this row and its close was R-463, now CL | 8 |
|
||||
| R-489 | P3 | FIXED-BY-LATER-WORK | The residual (a unit-restore-recreated volume has no compose label, so the remove answered []) was fixed in felhom-controller 206b035 (v0.268.0, R-658). controller/internal/stacks/delete.go:1089-1090: '// appVolumeSet is every volume the removal accounts for: the ones carrying the project label AND the // ones the app's definition declares that Docker holds by name (R-658, v0.268.0).' delete.go:70 | 5 |
|
||||
| R-498 | P3 | STILL-TRUE-SMALL | FIX: When building the app info view, rewrite each first_steps entry: replace '<word>.DOMAIN' with the stack's real address (installed SUBDOMAIN + customer domain when installed, else the template default subdomain + customer domain). Use the same approach as known_login.go:67. — Still literal: grep -l '\.DOMAIN' over app-catalog templates/*/.felhom.yml -> 58 of 58; e.g. templates/bookstack/.felho | 8 |
|
||||
| R-516 | P3 | STILL-TRUE-NOT-SMALL | By its own text it now waits for a Hungarian walk on a box with a second drive (items 4, 7, 8, 9, 10) and needs a separate row for item 11. That is live-box work, not source. No commit after 2026-09-20 names R-516 as closed. | 3 |
|
||||
| R-521 | P3 | STILL-TRUE-NOT-SMALL | No storage-disconnect suppression of app_start_failed: controller/cmd/controller/main.go:2796 'down := (stacks.IsDownState(st.State) // crashLooping) && !userStopped && !quiesced[st.Name]' (only user-stop/quiesce suppress), and controller/internal/notify/notifier.go:718 emits app_start_failed per newly-down app. It also needs the hub cooldown semantics changed (per-key cooldown outliving storage_r | 6 |
|
||||
| R-522 | P3 | STILL-TRUE-NOT-SMALL | No tunnel-connection signal in the controller: grep for TunnelConnected/cloudflared connection state in controller/internal -> none. The tile comes from container metadata: controller/internal/web/inframeta.go:21-22 '"cloudflared": { DisplayName: "Cloudflare Tunnel"'. Fixing it needs a new state source (cloudflared metrics or the hub-push result) wired into the tile, plus a live internet-cut valid | 5 |
|
||||
| R-531 | P3 | STILL-TRUE-NOT-SMALL | Both measurements are done (audits/evidence-drill-0243-2026-09-16/). What remains is an operator design question about the crash-loop budget (a slow loop every 20 min is never paused). That is an operator decision, not code. | 2 |
|
||||
| R-540 | P3 | STILL-TRUE-NOT-SMALL | Still a single pool box: felhom.eu/hub/cmd/hub/main.go:367 'poolBoxID, _ := strconv.ParseInt(os.Getenv("HETZNER_POOL_BOX_ID"), 10, 64)'. It needs a selection-rule design and eventually a second box (money). No risk today (0.3 % full per the row). | 3 |
|
||||
| R-542 | P3 | UNCHECKED | The controller passes the agent's 'initialize' list through untouched: controller/internal/web/agent_disk_handlers.go:160-162 mergeAttachCandidates only appends to Attach. The agent puts every candidate under initialize: felhom-agent internal/localapi/disks.go:438 'initialize = append(initialize, c) // every unclaimed disk can be initialized'. Whether a REGISTERED in-guest drive still counts as 'u | 8 |
|
||||
| R-545 | P3 | STILL-TRUE-NOT-SMALL | Still no un-configure route: controller/internal/web/server.go:744-783 lists /backup/offbox/{config,toggle,enable-all,offer-dismiss,run,reset,status,restore,place,reconstitute,verify-copy/delete,confirm-escrow,inject-password}; nothing removes a target. The fix is a new destructive customer action (shred data/offbox/) with a hub-escrow refusal check (R-241 rule), Hungarian/English copy and a UI. T | 4 |
|
||||
| R-547 | P3 | STILL-TRUE-SMALL | FIX: Also schedule the fill-watch on an interval (e.g. sched.Every("fill-watch-fast", 15*time.Minute, ...) calling the same fillWatcher.Check) next to the daily run, and log the cadence. Alternative if the operator prefers: state in 08-alarm-ladder.md that a transient full disk is out of scope. — Still daily: controller/cmd/controller/main.go:1546 'sched.Daily("fill-watch", "03:30", func(ctx conte | 8 |
|
||||
| R-548 | P3 | STILL-TRUE-NOT-SMALL | The row's original fix shape SHIPPED: felhom-agent 0722b2c (2026-09-24, R-685) checks before start, internal/backup/runner.go:279 '... so a new one needs about %s (old archives are removed only after a successful backup)'. Shown in the UI by controller 44ae4de (v0.272.0). The 2026-09-30 addendum is still true and is the open part: the local tier is refused for ever (10 refusals on demo-hp) because | 6 |
|
||||
| R-552 | P3 | STILL-TRUE-SMALL | FIX: Add Manager.ClearInterruptedRestore(stack) (delete from opInterrupted + persistRestoreRecordLocked under m.mu). Call it in removeStack next to ClearUpdateHold, and log when it cleared something. — The interrupted-restore notice is cleared only at controller/internal/backup/opstatus.go:61 'delete(m.opInterrupted, stack)' inside BeginRestoreOp. removeStack (controller/internal/api/router.go:955 | 6 |
|
||||
| R-554 | P3 | STILL-TRUE-NOT-SMALL | Still present: controller/internal/setup/ (setup.go, handlers.go, csrf.go, network.go, templates/), and controller/cmd/controller/main.go:328-330 'if setup.NeedsSetup(cfg) { ... runSetupMode(cfg, logger)'. The fix deletes a package and adds a new waiting page (HU+EN copy) with a red-proof. It also needs a check of drill/golden reliance on .needs-setup (controller/internal/web/handler_debug.go refe | 5 |
|
||||
| R-562 | P3 | STILL-TRUE-NOT-SMALL | Needs an operator word on the Hungarian number/date format (the row says so) and a deliberate Hungarian-byte change release with parity re-capture. Not checked further. | 2 |
|
||||
| R-565 | P3 | STILL-TRUE-SMALL | FIX: Add an ASCII Hungarian word regex to TestI18nEnglishPages (seeded from i18n_extract.py ASCII_HU plus the words releases B/C found: mp, db, FIGYELEM, jelenlegi, majd a(z), Konfig, Megtartva, helyi, Befejezve, automatikus, kedd/szerda/szombat, szint), applied after the data mask. — The English page test detects only accented letters: controller/internal/web/i18n_parity_test.go:563 'func huLette | 6 |
|
||||
| R-573 | P3 | FIXED-BY-LATER-WORK | felhom-controller 7c4a33b (v0.258.0, 'the last four Hungarian things an English household met ... R-573 the two channel banners'). controller/internal/web/alerts.go:113 'func (am *AlertManager) SetAgentChannelAlert(down bool, msgKey, msg string) {' and :136 'func (am *AlertManager) SetEndpointDriftAlert(drift bool, msgKey, msg string) {'. Both set MessageKey with msg only as a fail-open fallback. | 4 |
|
||||
| R-575 | P3 | STILL-TRUE-SMALL | FIX: Make memoryVerdict return (refusal error, warningKey string, warningArgs []any) instead of a msgHU string, and render the warning at the deploy answer with the request's language (the Alert/UpdateRefusal pattern). — Still a plain string: controller/internal/stacks/deploy.go:1456 'func (m *Manager) memoryVerdict(newReqMB, newLimitMB, releasedReqMB, releasedLimitMB int) (refusal error, warning | 5 |
|
||||
| R-578 | P3 | STILL-TRUE-NOT-SMALL | The lock-reentrancy guard is still one test in one package: controller/internal/backup/offsite_diag_test.go:190 'TestNoteHelpersAreNotCalledUnderTheSettingsLock'. No other package has one, and there is no gate (grep for R-578 in controller -> none). The fix needs a cross-package AST gate that knows which methods read settings (or a re-entrant read path in Settings). That is a new mechanism, likely | 5 |
|
||||
| R-581 | P3 | STILL-TRUE-SMALL | FIX: In GetCustomers, join on MAX(id) per customer_id instead of MAX(received_at), and add ', id DESC' to the ORDER BY at store.go:1393 and :1446. — Still no tie-break: felhom.eu/hub/internal/store/store.go:1329 'SELECT customer_id, MAX(received_at) as max_time' joined on 'r.received_at = latest.max_time' (:1333). A same-second tie returns BOTH rows (a duplicate customer in GetCustomers), not just | 6 |
|
||||
| R-584 | P3 | NOT-WORTH-IT | PICK close-as-accepted: Helper scripts with the shared DEMO controller password inline were left in a demo guest's /tmp. The cleanup rule exists, but no mechanism enforces it. | 5 |
|
||||
| R-585 | P3 | STILL-TRUE-NOT-SMALL | Still finished Hungarian from callers: controller/internal/notify/notifier.go:408 'n.PushEvent("backup_failed", "error", message, BackupDetails{Error: errMsg})', :487 offbox_enlarge_blocked, :497 db_dump_failed. None of the six types is in convertedProducers (controller/internal/notify/message_customer_test.go:39). The hub still has no customerMessages entry for offbox_enlarge_blocked (felhom.eu/h | 5 |
|
||||
| R-586 | P3 | NOT-WORTH-IT | PICK close-as-accepted: The ISO bootstrap harness runs only at each ISO release gate (G16), not on every push. | 6 |
|
||||
| R-587 | P3 | STILL-TRUE-SMALL | FIX: At the start of build-felhom-iso.sh, refuse (non-zero exit, named reason) when any *.rootpw.txt exists in the output dir. In the SKILL publish block, add a pre-check line that aborts the rclone copy if a *.rootpw.txt is present in the publish source. — The files are gone (ls /mnt/5_hdd/felhom.eu/felhom-iso/out / grep -c rootpw -> 0). The guard is still not built: the publish still relies on t | 6 |
|
||||
| R-593 | P3 | STILL-TRUE-SMALL | FIX: Move 'Az alkalmazás aldomainje' to SUBDOMAIN, give AUTH_SECRET its own description (e.g. 'A munkamenetek aláírásához használt kulcs — ne generáld újra'), add the two English i18n.en descriptions, and re-capture papra's entries in copy_freeze/hu.json with the reason in the commit. — Still wrong: app-catalog-felhom.eu/templates/papra/.felhom.yml:47 ' description: "Az alkalmazás aldomainje"' | 5 |
|
||||
| R-600 | P3 | STILL-TRUE-SMALL | FIX: Call the wgsync reconciler's Trigger() before the COMPLETE log line (nil-safe), and make the line say 'wg peer removal pushed' or 'queued for the next wgsync push' depending on whether a syncer is wired. — Log line unchanged: felhom.eu/hub/internal/web/customer_delete.go:310 's.logger.Printf("[INFO] customer DELETE cascade COMPLETE for %s (journal #%d) — full teardown", customerID, journalID) | 6 |
|
||||
| R-607 | P3 | UNCHECKED | The row asks first for a live reproduction loop (push a tag, sync, read catalog_images on a timer) and has no diagnosis. Why the sync reports 'nincs változás' while the cache moved, and when CatalogImages refreshes, are live-box behaviour I could not settle from source in the time box. No commit after d19f07ea (filing) names R-607 as fixed. | 6 |
|
||||
| R-612 | P3 | STILL-TRUE-NOT-SMALL | The memory half is fixed (catalog a5a729a, 'wishlist 512M (R-612)'). The open half, making a failed first-boot seed visible, needs a new detection mechanism (read the seed's exit/log, or a probe that checks the Role/Group rows). No commit addresses it. | 4 |
|
||||
| R-613 | P3 | STILL-TRUE-NOT-SMALL | uptime-kuma is fixed (catalog a5a729a). The open half is a sweep of all 58 templates for probes that pass on a setup wizard. That needs per-app live inspection, so it is not small. No commit names it. | 3 |
|
||||
| R-615 | P3 | STILL-TRUE-SMALL | FIX: Before the fetch, run 'git remote set-url origin <buildRepoURL()>' (or read origin and re-clone on mismatch), and log at INFO, masked, when the remote changed. The appended drill-folder residue (stack folders the live catalog lacks) is a separate drill-teardown item and is not part of this fix. — Still inert: controller/internal/sync/sync.go:279 clones only 'if _, err := os.Stat(gitDir); os.I | 7 |
|
||||
| R-616 | P3 | STILL-TRUE-SMALL | FIX: Clone and fetch with the credential-free RepoURL, and supply credentials per command via '-c http.extraHeader=Authorization: Basic <b64>' (or GIT_ASKPASS env) only when username+token are set. Pairs naturally with the R-615 set-url fix (set-url to the bare URL). — Still true: controller/internal/sync/sync.go:327-331 buildRepoURL injects 'https://%s:%s@' and the clone at :283-288 passes that U | 5 |
|
||||
| R-622 | P3 | FIXED-BY-LATER-WORK | adventurelog v0.13.0 was diagnosed, fixed and promoted with a two-venue test record in app-catalog-felhom.eu 06ea7da (2026-09-27, 'adventurelog: v0.12.1 -> v0.13.0 with its health and world-data fixes in the same commit (R-655, 09 decision 41)'; bench healthy in 217 s, box 9202 through the guarded Update in 204 s). templates/adventurelog/docker-compose.yml:13 ' image: ghcr.io/seanmorley15/adven | 6 |
|
||||
| R-635 | P3 | FIXED-BY-LATER-WORK | The open remainder (app_oom fires once per container run, no escalation) was built in felhom-controller 0054d4b (v0.265.0, 'OOM storm alarm', R-636). controller/internal/notify/notifier.go:746 '\t\tn.emit("app_oom_storm", "error",' fires once per run when 20 or more kills land in 30 min (:754-764, oomStormKills=20, oomStormWindowMin=30; pinned by TestR636_*). The 79 % headroom and the method lesso | 6 |
|
||||
| R-645 | P3 | STILL-TRUE-NOT-SMALL | Still true for the operator CLI: controller/internal/settings/settings.go:1923-1929 ClearRestoreHold deletes ANY hold reason (including update-failed) with no pin restore, and the flag is in controller/cmd/controller/main.go:94/198. The row lists three candidate shapes with 'none chosen'. Picking one changes operator-path semantics and the capture logic, so it is not a one-hour fix. (The automatic | 6 |
|
||||
| R-675 | P3 | STILL-TRUE-SMALL | FIX: In missingFileLegsRefusal, when the second drive holds a whole copy of the app (the decision-26 whole-copy check the backup manager already exposes for the restore page), name that whole restore action instead of 'Fájlok visszaállítása'. Keep the existing branch otherwise. — Unchanged: controller/internal/web/handlers.go:1745 'return head + "A fájlok a második meghajtó másolatából állíthatók | 5 |
|
||||
| R-676 | P3 | STILL-TRUE-NOT-SMALL | A watch row, still true: controller/internal/stacks/unhealthy.go:116 skips only 'if st.Deploying // st.Updating // st.HoldReason != "" // st.updateHeld {', so a deploy's first start (after Deploying clears) is sampled by decision 28's crash-loop stop. The immich cause is fixed in the catalog (56c4888, 768M, per the row). Covering a slow first start would need a first-start grace decision. | 5 |
|
||||
| R-682 | P3 | STILL-TRUE-NOT-SMALL | felhom-controller 7690c27 (v0.296.0): no remove journal in source — grep -rni 'remove.*journal/removeJournal/remove_intent/interrupted remove' controller/*.go returns nothing; git log --grep R-682 empty. Needs a new boot-time journal mechanism. | 3 |
|
||||
| R-683 | P3 | UNCHECKED | Watch item about a power-cut drill outcome; behaviour only a live box shows; git log --grep R-683 empty in controller. | 1 |
|
||||
| R-698 | P3 | NOT-WORTH-IT | PICK close-as-accepted (option a), operator to confirm: A backup records the image name/digest, not the image; restoring a version the maker deleted from the registry fails at the pull. | 2 |
|
||||
| R-700 | P3 | FIXED-BY-LATER-WORK | felhom-controller 820e8ef (v0.276.0, R-697/R-700). controller/internal/stacks/migrate.go:789: 'm.logger.Printf("[INFO] [stacks] %s: data moved %s -> %s — app.yaml keeps its pin (%d service(s)) and records"' — persistDriveFlip (migrate.go:763) loads app.yaml and changes only HDD_PATH. Only a live proof on a two-drive box remains (row's own residue). | 3 |
|
||||
| R-704 | P3 | FIXED-BY-LATER-WORK | felhom-controller 7cba0bf (v0.278.0). controller/internal/api/router.go:918 'func (r *Router) dropLeftoverHold(name, why string) {' calling r.sett.ClearUpdateHold(name); pinned by TestR704_AFreshInstallDropsALeftoverHold. Residue: live proof of the install-time drop only. | 2 |
|
||||
| R-706 | P3 | FIXED-BY-LATER-WORK | felhom-controller 0c702f8 (v0.279.0). controller/internal/api/router.go:946 'if err := r.backupMgr.DeleteOffsiteRestoreCopy(name); err != nil {' inside removeVerificationCopy (R-706); pinned by TestR706_RemovalWithBackupsDeletesTheVerificationCopy. Residue: not seen live. | 2 |
|
||||
| R-717 | P3 | STILL-TRUE-NOT-SMALL | app-catalog templates/opengist/.felhom.yml:42 and templates/wishlist/.felhom.yml:42 carry only signup_block; no after_setup/after_install in either .felhom.yml (grep empty). Fix needs per-app DB writes (opengist sqlite with app stopped) plus live proof. | 3 |
|
||||
| R-723 | P3 | FIXED-BY-LATER-WORK | felhom.eu 80aeac71 (hub v0.126.0). hub/internal/monitor/staleness.go:174 '// R-723 (v0.126.0): a customer's NEW box is not a recovery.'; hub/internal/notify/dispatcher.go:540 '"suppressed", "first hour of a new box (R-723)", "operator"'. Residue: live proof at a real first install. | 2 |
|
||||
| R-724 | P3 | STILL-TRUE-NOT-SMALL | Text parts fixed in controller 6be6c53 (v0.283.0). Remaining LAN/gateway read still goes only through the samba container: controller/internal/stacks/guestnet.go:47 'out, err := dockerexec.Command("docker", append([]string{"exec", sambaContainer}, args...)...).Output()' — the comment (guestnet.go:18) accepts that reads fail while sharing is off. Needs another read path (design). | 3 |
|
||||
| R-728 | P3 | STILL-TRUE-SMALL | FIX: Add a per-customerID in-flight guard (sync.Map/mutex set) around the create handler from the duplicate check to the self-bind mint, so a second concurrent submit for the same ID gets the 'already exists' form; optionally disable the submit button on submit in the template. — No commit fixes R-728 (git log --grep in felhom.eu only the filing commit 11591f3a). hub/internal/web/configs.go:705 'e | 5 |
|
||||
| R-729 | P3 | STILL-TRUE-NOT-SMALL | No route clears the off-site target: controller/internal/web/server.go:744-783 lists /backup/offbox/{config,toggle,enable-all,offer-dismiss,run,reset,status,restore,place,reconstitute,verify-copy/delete,confirm-escrow,inject-password}; /reset (offbox_handlers.go:311) only resets an orphaned repo. New press needs handler + settings clear + template + HU/EN copy + escrow/hub-managed-target interplay | 4 |
|
||||
| R-733 | P3 | STILL-TRUE-NOT-SMALL | Harness/golden-evidence change plus a decision whether proofs run with swap off; no commit references R-733. Not a source-verifiable single fix. | 1 |
|
||||
| R-738 | P3 | NOT-WORTH-IT | PICK close-as-accepted (wger defect fixed; the generic gap is a design note): The guarded Update's health check sees only an app's front page, so an app that serves its front page while its data is broken passes as done. | 3 |
|
||||
| R-747 | P3 | STILL-TRUE-NOT-SMALL | Lockout shortened: app-catalog a4597cd; templates/mealie/docker-compose.yml:27 ' - SECURITY_USER_LOCKOUT_TIME=1'. Residue still open: hourly lock renewal by a stranger (needs decision 57 option d) and page copy; needs decision + live proof. | 2 |
|
||||
| R-755 | P3 | DUPLICATE | of R-762 — Still true: templates/wger/docker-compose.yml has no WGER_USE_GUNICORN (grep empty). R-762 (open, read) states 'Owner decides together with R-755 (same server question)' and its fix names 'the gunicorn switch of R-755'. | 2 |
|
||||
| R-756 | P3 | UNCHECKED | Depends on whether 9202's scratch drive is a registered drive — live box state; the row itself says not measured which. Not verifiable from source. | 1 |
|
||||
| R-757 | P3 | STILL-TRUE-NOT-SMALL | No commit references R-757. controller/internal/stacks/deploy.go:1339-1349: 'case "secret":' ... 'value, err := generateValue(field.Generate)' ... 'appCfg.Env[field.EnvVar] = value' for any missing field of a deployed app, with no exception for fields consumed only by after_install. Fix needs a design (a marker for given-at-install fields or an ask path). | 3 |
|
||||
| R-758 | P3 | STILL-TRUE-NOT-SMALL | Still true: templates/bookstack/.felhom.yml:18 ' mem_limit: "512M"' vs compose limits docker-compose.yml:42 '512M' + :79 '256M'; onboarding/EXISTING-APPS-GAPS.md:26 still lists all 8. No gate (only scripts/onboarding_gaps.py:171 reports it). Not small: raising 8 figures changes the capacity check (decision 22) — which apps fit a box — plus a gate with decoy and a publish. | 4 |
|
||||
| R-762 | P3 | STILL-TRUE-NOT-SMALL | templates/wger/docker-compose.yml sets no DJANGO_DEBUG and no static/media server (grep empty); templates/wger/.felhom.yml:18 'lifecycle: hidden' (catalog 55b8c8a). Needs a server design decision (nginx sidecar vs gunicorn+static) and bench+box proof. | 2 |
|
||||
| R-763 | P3 | STILL-TRUE-NOT-SMALL | templates/wger/docker-compose.yml sets neither ALLOW_REGISTRATION nor ALLOW_GUEST_USERS (grep empty); wger hidden (.felhom.yml:18 'lifecycle: hidden'). The env change is tiny but its proof needs a working wger on 9202 (blocked by R-762); best done in the same session as R-762. | 2 |
|
||||
| R-774 | P3 | STILL-TRUE-NOT-SMALL | templates/karakeep has no Sentry/phone-app sentence (grep -i sentry empty); mail-ON proof needs a hub-enabled live box (demo-hp) — live work. | 2 |
|
||||
| R-775 | P3 | STILL-TRUE-NOT-SMALL | Narrowed (Grimmory published behind the family gate). Residue: per-name 15-min lock is hard-coded upstream (no setting) and the reinstall-over-kept-books finding is uninvestigated — needs live investigation. | 2 |
|
||||
| R-776 | P3 | STILL-TRUE-NOT-SMALL | Only bookstack has it: templates/bookstack/docker-compose.yml:32 ' - APP_PROXIES=172.16.0.0/12'; grep for TRUSTED_PROXIES/CORE_TRUST_PROXY/IPEXTRACTION/N8N_PROXY_HOPS in kimai, zipline, vikunja, nextcloud, n8n compose returns nothing. Five apps, each needing a live 3.6 re-measure on 9202. | 3 |
|
||||
| R-778 | P3 | NOT-WORTH-IT | PICK close-as-accepted: If a box rolls back to a controller older than 0.286, the dashboard's login counter trusts the leftmost forwarded address and can be dodged until the box moves forward. | 2 |
|
||||
| R-782 | P3 | STILL-TRUE-NOT-SMALL | Source agrees: templates/glance/docker-compose.yml:27 seeds glance.yml with no auth: block (grep 'auth' in templates/glance empty); templates/homepage has no HOMEPAGE_ALLOWED_HOSTS (grep empty). Needs live measurement on 9202 and a decision whether a public glance dashboard is intended. | 3 |
|
||||
| R-783 | P3 | NOT-WORTH-IT | PICK close-as-accepted (re-open if upstream exposes the setting): Three wrong SparkyFitness sign-ins by anyone block every visitor's sign-in for about 10 seconds. | 2 |
|
||||
| R-785 | P3 | STILL-TRUE-NOT-SMALL | templates/sparkyfitness/docker-compose.yml:45 ' image: codewithcj/sparkyfitness_server:v0.17.3' and :89 'codewithcj/sparkyfitness:v0.17.3'. Major-version ladder walk (bench + box), gated on R-784. | 1 |
|
||||
| R-831 | P3 | NOT-WORTH-IT | PICK keep (rotation is the operator's call; do not close a leaked-secret row silently): The Hetzner storage API token was printed into one session transcript. | 1 |
|
||||
| R-836 | P3 | STILL-TRUE-NOT-SMALL | Live host boot-loader work needing operator-approved reboots and measurement (GRUB env block on ESP, sp5100_tco arming). | 1 |
|
||||
| R-839 | P3 | STILL-TRUE-NOT-SMALL | Gate behaves as described: controller/cmd/controller/main.go:2397 'return false, "drive " + hdd + " is not a live mountpoint"' with hdd = cfg.Env["HDD_PATH"] (main.go:2379). Which writer put a per-app path in HDD_PATH is undiagnosed — a diagnosis task, not a one-hour fix. | 4 |
|
||||
| R-853 | P3 | NOT-WORTH-IT | PICK close-as-accepted: After a boot the box's versions and crash facts reach the hub up to about 15 minutes late. | 4 |
|
||||
| R-862 | P3 | UNCHECKED | Waiting on the operator's by-hand bootstrap on Tester 2 through his tunnel; whether done is live-box state. No commit records it (felhom.eu log since 2026-10-04). | 1 |
|
||||
| R-870 | P3 | NOT-WORTH-IT | PICK keep (operator's call; close when Tester 1 is retired): Tester 1's two Cloudflare tokens (disposable test customer) were printed into one session transcript. | 1 |
|
||||
| R-879 | P3 | STILL-TRUE-NOT-SMALL | hub/internal/store/store.go:166 'retrieval_password TEXT NOT NULL,' and store.go:1586/1592 write retrieval_password and api_key as given; no seal on them (grep seal near these fields empty). Sealing/hashing three tables with migration is a security change, more than an hour. | 3 |
|
||||
| R-882 | P3 | UNCHECKED | Longhorn instance-manager state on DooPlex (Tier 2, forbidden to touch); live-only, owner operator. | 1 |
|
||||
| R-883 | P3 | UNCHECKED | homelab-manifests repo is not in this workspace (ls /mnt/5_hdd/felhom.eu/git shows only app-catalog-felhom.eu, drills, felhom-agent, felhom-controller, felhom.eu); live DooPlex check (kubectl) is out of scope. | 1 |
|
||||
| R-886 | P3 | UNCHECKED | DooPlex Alertmanager volume ownership; homelab-manifests not in this workspace and live check not permitted. | 1 |
|
||||
@@ -0,0 +1,6 @@
|
||||
### R262-a: drop skipped from knownUnmodelled
|
||||
--- FAIL: TestR262_RestoreTestFieldsAreAKnownSubset (0.00s)
|
||||
FAIL
|
||||
### R262-b: the hub stops decoding source_tier
|
||||
--- FAIL: TestR262_RestoreTestFieldsAreAKnownSubset (0.00s)
|
||||
FAIL
|
||||
@@ -0,0 +1,9 @@
|
||||
### R263-a: ClearBackupTarget writes true
|
||||
1700: s.StoragePaths[i].BackupTarget = true
|
||||
--- FAIL: TestR263_OnlySetBackupTargetGrantsTheRole (0.24s)
|
||||
r263_backup_target_writers_test.go:73: BackupTarget may be granted outside SetBackupTarget at: [../settings/settings.go:1700:3]
|
||||
FAIL
|
||||
### R263-b: a composite literal BackupTarget: true in another package
|
||||
--- FAIL: TestR263_OnlySetBackupTargetGrantsTheRole (0.28s)
|
||||
r263_backup_target_writers_test.go:73: BackupTarget may be granted outside SetBackupTarget at: [../web/zz_r263_decoy.go:5:50]
|
||||
FAIL
|
||||
@@ -0,0 +1 @@
|
||||
FAIL: site/unlisted-page: decoy PASSED - LIVE HOLE (rc=0)
|
||||
@@ -0,0 +1,16 @@
|
||||
### R885-a: gate ignores a suite's exit code
|
||||
FAIL prints OK, exits 1 want rc=1 got rc=0
|
||||
FAIL failing suite 3 levels deep want rc=1 got rc=0
|
||||
ok empty scripts/ want rc=1 got rc=1
|
||||
ok genuine: two passing suites want rc=0 got rc=0
|
||||
ok nested run steps aside want rc=0 got rc=0
|
||||
script-tests decoys: 3/5
|
||||
rc=1
|
||||
### R885-b: gate passes when it found nothing
|
||||
ok prints OK, exits 1 want rc=1 got rc=1
|
||||
ok failing suite 3 levels deep want rc=1 got rc=1
|
||||
FAIL empty scripts/ want rc=1 got rc=0
|
||||
ok genuine: two passing suites want rc=0 got rc=0
|
||||
ok nested run steps aside want rc=0 got rc=0
|
||||
script-tests decoys: 4/5
|
||||
rc=1
|
||||
@@ -0,0 +1,58 @@
|
||||
===== R-118 RED (guard removed: if s.devicePresent(d.MountPath) -> if true) — 2026-10-05T18:53:22+02:00
|
||||
$ go test ./internal/localapi/ -run TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity -v -count=1
|
||||
=== RUN TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity
|
||||
=== RUN TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/absent
|
||||
disks_device_presence_test.go:212: absent drive advertises total=33697107968 used=4727169024 frac=0.140 — that is the filesystem UNDER the bare mountpoint, not the drive (R-118)
|
||||
=== RUN TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/present
|
||||
--- FAIL: TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity (0.00s)
|
||||
--- FAIL: TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/absent (0.00s)
|
||||
--- PASS: TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/present (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-agent/internal/localapi 0.010s
|
||||
FAIL
|
||||
rc=1
|
||||
===== R-118 GREEN (fix restored)
|
||||
$ go test ./internal/localapi/ -run TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity -v -count=1
|
||||
=== RUN TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity
|
||||
=== RUN TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/absent
|
||||
=== RUN TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/present
|
||||
--- PASS: TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity (0.00s)
|
||||
--- PASS: TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/absent (0.00s)
|
||||
--- PASS: TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/present (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/localapi 0.009s
|
||||
rc=0
|
||||
|
||||
===== R-269 RED (tokenstore.go reverted to HEAD: reload-on-MISS-only Lookup) — 2026-10-05T18:53:55+02:00
|
||||
$ go test ./internal/localapi/ -run TestTokenStore_RotatedOutTokenRejectedFirst -v -count=1
|
||||
=== RUN TestTokenStore_RotatedOutTokenRejectedFirst
|
||||
tokenstore_test.go:247: rotated-out token still authorizes vmid 130 on its first presentation after rotation — Mint's 'any previous token for this guest is revoked' is false across processes (R-269)
|
||||
--- FAIL: TestTokenStore_RotatedOutTokenRejectedFirst (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-agent/internal/localapi 0.008s
|
||||
FAIL
|
||||
rc=1
|
||||
===== R-269 GREEN (fix restored)
|
||||
$ go test ./internal/localapi/ -run TestTokenStore_RotatedOutTokenRejectedFirst -v -count=1
|
||||
=== RUN TestTokenStore_RotatedOutTokenRejectedFirst
|
||||
--- PASS: TestTokenStore_RotatedOutTokenRejectedFirst (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/localapi 0.010s
|
||||
rc=0
|
||||
|
||||
===== R-317 RED (probe reverted to the pre-fix dnsmasq-base binary path usr/sbin/dnsmasq) — 2026-10-05T18:54:37+02:00
|
||||
$ go test ./internal/lanresolver/ -run TestEnsureDnsmasq_BinaryWithoutUnitInstalls -v -count=1
|
||||
=== RUN TestEnsureDnsmasq_BinaryWithoutUnitInstalls
|
||||
ensure_dnsmasq_test.go:80: install was skipped on a dnsmasq-base-only host (binary present, unit absent) — the enable that follows targets a missing unit (R-317). calls: ["/usr/local/sbin/felhom-priv-apply dnsmasq /tmp/felhom-resolver-3076255408.conf felhom-resolver-base.conf" "systemctl enable --now dnsmasq" "systemctl restart dnsmasq"]
|
||||
--- FAIL: TestEnsureDnsmasq_BinaryWithoutUnitInstalls (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-agent/internal/lanresolver 0.006s
|
||||
FAIL
|
||||
rc=1
|
||||
===== R-317 GREEN (fix restored)
|
||||
$ go test ./internal/lanresolver/ -run TestEnsureDnsmasq_BinaryWithoutUnitInstalls -v -count=1
|
||||
=== RUN TestEnsureDnsmasq_BinaryWithoutUnitInstalls
|
||||
--- PASS: TestEnsureDnsmasq_BinaryWithoutUnitInstalls (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/lanresolver 0.004s
|
||||
rc=0
|
||||
@@ -0,0 +1,173 @@
|
||||
=== R-760 red-proof (2026-10-05T18:53:44+02:00) — catalog 29ac711 + working tree
|
||||
--- UNDO: remove the R-760 '# No healthcheck' comment block from templates/vikunja/docker-compose.yml
|
||||
$ python3 scripts/test_healthcheck_explained.py
|
||||
+ [] : service(s) with no compose healthcheck and no '# No healthcheck' comment saying why: ['vikunja/vikunja']
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 2 tests in 0.011s
|
||||
|
||||
FAILED (failures=1)
|
||||
rc=1
|
||||
--- RESTORE
|
||||
$ python3 scripts/test_healthcheck_explained.py
|
||||
Ran 2 tests in 0.011s
|
||||
|
||||
OK
|
||||
rc=0
|
||||
|
||||
=== R-593 red-proof (2026-10-05T18:55:00+02:00) — catalog 29ac711 + working tree
|
||||
--- UNDO: git show HEAD:templates/papra/.felhom.yml > templates/papra/.felhom.yml (the pre-fix papra .felhom.yml)
|
||||
$ python3 scripts/test_deploy_field_descriptions.py
|
||||
- ['papra SUBDOMAIN has no description',
|
||||
- 'papra AUTH_SECRET carries the subdomain sentence',
|
||||
- 'papra SUBDOMAIN has no description'] : ['papra SUBDOMAIN has no description', 'papra AUTH_SECRET carries the subdomain sentence', 'papra SUBDOMAIN has no description']
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 2 tests in 0.025s
|
||||
|
||||
FAILED (failures=1)
|
||||
rc=1
|
||||
--- RESTORE
|
||||
$ python3 scripts/test_deploy_field_descriptions.py
|
||||
Ran 2 tests in 0.025s
|
||||
|
||||
OK
|
||||
rc=0
|
||||
|
||||
=== R-781 red-proof (2026-10-05T18:55:52+02:00) — catalog 29ac711 + working tree
|
||||
--- UNDO: drop the isolate_onboarding_clone(cat) call in onboarding_cases (the pre-fix harness)
|
||||
1
|
||||
$ python3 <runner calling test_gate_decoys.onboarding_cases() only — the whole file also reaches a registry>
|
||||
ok FACT: a new template with NO record rc=1 (expected 1)
|
||||
ok FACT: a record missing id 1.4 rc=1 (expected 1)
|
||||
ok FACT: 1.4 answered only inside an HTML comment rc=1 (expected 1)
|
||||
ok FACT: done with a path that does not exist rc=1 (expected 1)
|
||||
ok FACT: done with an EMPTY directory (the mkdir shape) rc=1 (expected 1)
|
||||
ok FACT: done naming an absent file in the sibling repo rc=1 (expected 1)
|
||||
ok FACT: n/a with an EMPTY reason rc=1 (expected 1)
|
||||
ok FACT: n/a with a two-word reason rc=1 (expected 1)
|
||||
ok FACT: an OPEN row rc=1 (expected 1)
|
||||
ok FACT: opened: backdated before the checklist rc=1 (expected 1)
|
||||
ok FACT: a checklist id the template a new app copies lacks rc=1 (expected 1)
|
||||
ok FACT: an exempt app's record with a done that points nowhere rc=1 (expected 1)
|
||||
FAILS 4
|
||||
FAIL: GENUINE: a complete record (catalog + sibling evidence): rc=1 expected 0; missing ['onboarding gate OK']
|
||||
FAIL: GENUINE: an id added AFTER opened: does not bind: rc=1 expected 0; missing ['onboarding gate OK']
|
||||
FAIL: GENUINE: an exempt app's record may say open: rc=1 expected 0; missing ['exempt app(s) with a record (shape-checked): wger']
|
||||
FAIL: STATED SKIP: sibling repo absent (the CI shape) - printed, not checked: rc=0 expected 0; missing ['felhom.eu/documentation/audits/onb/
|
||||
rc=1
|
||||
--- RESTORE
|
||||
ok GENUINE: a complete record (catalog + sibling evidence) rc=0 (expected 0)
|
||||
ok GENUINE: an id added AFTER opened: does not bind rc=0 (expected 0)
|
||||
ok GENUINE: an exempt app's record may say open rc=0 (expected 0)
|
||||
ok STATED SKIP: sibling repo absent (the CI shape) - printed, not checked rc=0 (expected 0)
|
||||
FAILS 0
|
||||
rc=0
|
||||
|
||||
=== R-806 red-proof (2026-10-05T18:56:53+02:00) — catalog 29ac711 + working tree
|
||||
--- UNDO: exercise_argv back to the pre-fix shape (always http://, no -k) — the old inline argv
|
||||
578: pass # red-proof: no -k
|
||||
581: return a + [f"http://{ip}:{port}{path}"]
|
||||
$ python3 scripts/test_check_volume_persistence.py
|
||||
FAIL: test_https_backend_gets_an_https_url_with_k (__main__.TestRoutedSchemes.test_https_backend_gets_an_https_url_with_k)
|
||||
AssertionError: 'http://10.0.0.5:8443/' != 'https://10.0.0.5:8443/'
|
||||
Ran 54 tests in 0.018s
|
||||
FAILED (failures=1)
|
||||
rc=1
|
||||
--- RESTORE
|
||||
Ran 54 tests in 0.018s
|
||||
|
||||
OK
|
||||
rc=0
|
||||
|
||||
=== R-605 red-proof (2026-10-05T18:58:47+02:00) — catalog 29ac711 + working tree
|
||||
--- UNDO: HARNESS_REFUSED = 2 in both gates (the pre-fix shared code); runner: rc 3 folded back into UNDETERMINED
|
||||
scripts/check-image-resolvable.py:166:HARNESS_REFUSED = 2
|
||||
scripts/check-volume-persistence.py:935:HARNESS_REFUSED = 2
|
||||
124:VERDICT = {0: "OK", 1: "FAILED", 2: "INCONCLUSIVE"}
|
||||
220: refused = [] # red-proof
|
||||
$ python3 scripts/test_check_volume_persistence.py
|
||||
FAIL: test_canary_failure_prints_the_harness_refused_marker (__main__.TestCheckEntryPoint.test_canary_failure_prints_the_harness_refused_marker)
|
||||
AssertionError: 2 != 3
|
||||
Ran 55 tests in 0.017s
|
||||
FAILED (failures=1)
|
||||
rc=1
|
||||
$ python3 scripts/test_check_image_resolvable.py
|
||||
Ran 19 tests in 0.011s
|
||||
OK
|
||||
rc=0
|
||||
$ python3 -m unittest scripts/test_catalog_gates.py SummaryTellsRefusedFromUndecided
|
||||
FAIL: test_a_refused_harness_does_not_read_as_undetermined (test_catalog_gates.SummaryTellsRefusedFromUndecided.test_a_refused_harness_does_not_read_as_undetermined)
|
||||
AssertionError: 'DID-NOT-RUN' not found in '\n==============================================================================\n== summary\n==============================================================================\n image-pins OK (exit 0)\n volume-persistence ERROR (exit 3)\nUNDETERMINED (it ran; some results could not be decided — never a pass): volume-persistence\n'
|
||||
Ran 3 tests in 0.004s
|
||||
FAILED (failures=1)
|
||||
rc=1
|
||||
--- RESTORE
|
||||
$ python3 scripts/test_check_volume_persistence.py
|
||||
Ran 55 tests in 0.018s
|
||||
OK
|
||||
rc=0
|
||||
$ python3 scripts/test_check_image_resolvable.py
|
||||
Ran 19 tests in 0.011s
|
||||
OK
|
||||
rc=0
|
||||
Ran 3 tests in 0.003s
|
||||
OK
|
||||
rc=0
|
||||
|
||||
=== R-605 red-proof, image-resolvable leg re-run (2026-10-05T18:58:59+02:00) — the first run above passed because the test compared against the constant itself (tautology); the tests now pin the literal 3
|
||||
--- UNDO: HARNESS_REFUSED = 2 in check-image-resolvable.py
|
||||
$ python3 scripts/test_check_image_resolvable.py
|
||||
FAIL: test_empty_catalog_is_an_error_not_a_pass (__main__.TestResolvabilityGate.test_empty_catalog_is_an_error_not_a_pass)
|
||||
AssertionError: 2 != 3
|
||||
FAIL: test_untrustworthy_resolver_refuses_to_report (__main__.TestResolvabilityGate.test_untrustworthy_resolver_refuses_to_report)
|
||||
AssertionError: 2 != 3 : a resolver that resolves the canary must abort (3, the harness refused), not pass — and not 2, which reads as 'some pins were throttled' (R-605)
|
||||
Ran 19 tests in 0.013s
|
||||
FAILED (failures=2)
|
||||
rc=1
|
||||
--- RESTORE
|
||||
Ran 19 tests in 0.012s
|
||||
FAILED (failures=2)
|
||||
rc=1
|
||||
|
||||
--- NOTE: the RESTORE above still failed because scripts/__pycache__ held the UNDONE module's bytecode: the undo
|
||||
(3 -> 2) kept the file's size and the restore landed in the same second, so Python's mtime+size cache check
|
||||
accepted the stale .pyc. Restored source confirmed HARNESS_REFUSED = 3; after `touch scripts/check-image-resolvable.py`:
|
||||
$ python3 scripts/test_check_image_resolvable.py
|
||||
Ran 19 tests in 0.011s
|
||||
OK
|
||||
rc=0
|
||||
(lesson for same-size red-proofs: run with PYTHONDONTWRITEBYTECODE=1 / clear __pycache__ between undo and restore)
|
||||
|
||||
=== R-594 red-proof (2026-10-05T19:01:53+02:00) — catalog 29ac711 + working tree (PYTHONDONTWRITEBYTECODE=1)
|
||||
--- UNDO 1: the register is read but never consulted (reg = None) — the pre-fix gate's behaviour
|
||||
1
|
||||
$ python3 scripts/test_gate_decoys.py
|
||||
ok FACT: registered for ANOTHER app - the promise still convicts, the entry is stale rc=1 (expected 1)
|
||||
ok FACT: registered on the path, but the sentence was rewritten (match gone) rc=1 (expected 1)
|
||||
ok FACT: a STALE entry - nothing in the English promises it any more rc=1 (expected 1)
|
||||
ok FACT: a registered promise with a two-word reason rc=1 (expected 1)
|
||||
ok FACT: n/a with a two-word reason rc=1 (expected 1)
|
||||
FAIL: GENUINE: a REGISTERED true retrieval promise passes: rc=1 expected 0; missing ['copy-i18n: OK', '1 registered retrieval promise(s) in ALLOWLIST_EN, 1 used
|
||||
rc=1
|
||||
--- UNDO 2: the STALE check removed (entries never judged live)
|
||||
1
|
||||
$ python3 scripts/test_gate_decoys.py
|
||||
ok GENUINE: a REGISTERED true retrieval promise passes rc=0 (expected 0)
|
||||
ok FACT: a registered promise with a two-word reason rc=1 (expected 1)
|
||||
ok FACT: n/a with a two-word reason rc=1 (expected 1)
|
||||
FAIL: FACT: registered for ANOTHER app - the promise still convicts, the entry is stale: rc=1 expected 1; missing ['STALE entry vaultwarden']
|
||||
FAIL: FACT: registered on the path, but the sentence was rewritten (match gone): rc=1 expected 1; missing ['STALE entry privatebin']
|
||||
FAIL: FACT: a STALE entry - nothing in the English promises it any more: rc=0 expected 1; missing ['STALE entry privatebin']
|
||||
rc=1
|
||||
--- RESTORE
|
||||
$ python3 scripts/test_gate_decoys.py
|
||||
ok GENUINE: a REGISTERED true retrieval promise passes rc=0 (expected 0)
|
||||
ok FACT: registered for ANOTHER app - the promise still convicts, the entry is stale rc=1 (expected 1)
|
||||
ok FACT: registered on the path, but the sentence was rewritten (match gone) rc=1 (expected 1)
|
||||
ok FACT: a STALE entry - nothing in the English promises it any more rc=1 (expected 1)
|
||||
ok FACT: a registered promise with a two-word reason rc=1 (expected 1)
|
||||
ok FACT: n/a with a two-word reason rc=1 (expected 1)
|
||||
catalog gate decoys OK — 137 case(s), every label judged on its fact (R-421)
|
||||
rc=0
|
||||
|
||||
@@ -0,0 +1,505 @@
|
||||
|
||||
==================== R-591 — red-proof ====================
|
||||
Fix undone in: internal/stacks/manager.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestDeepCopyStackI18nIsNotShared -v ./internal/stacks) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestDeepCopyStackI18nIsNotShared
|
||||
r591_copy_i18n_test.go:55: R-591: mutating the copy's data path overlay changed the ORIGINAL — the overlay is shared, not copied
|
||||
r591_copy_i18n_test.go:55: R-591: mutating the copy's initial creds overlay changed the ORIGINAL — the overlay is shared, not copied
|
||||
r591_copy_i18n_test.go:55: R-591: mutating the copy's description overlay changed the ORIGINAL — the overlay is shared, not copied
|
||||
r591_copy_i18n_test.go:55: R-591: mutating the copy's tagline overlay changed the ORIGINAL — the overlay is shared, not copied
|
||||
r591_copy_i18n_test.go:55: R-591: mutating the copy's deploy label overlay changed the ORIGINAL — the overlay is shared, not copied
|
||||
r591_copy_i18n_test.go:55: R-591: mutating the copy's option label overlay changed the ORIGINAL — the overlay is shared, not copied
|
||||
r591_copy_i18n_test.go:55: R-591: mutating the copy's use case overlay changed the ORIGINAL — the overlay is shared, not copied
|
||||
r591_copy_i18n_test.go:55: R-591: mutating the copy's optional group overlay changed the ORIGINAL — the overlay is shared, not copied
|
||||
r591_copy_i18n_test.go:55: R-591: mutating the copy's optional help overlay changed the ORIGINAL — the overlay is shared, not copied
|
||||
r591_copy_i18n_test.go:55: R-591: mutating the copy's integration overlay changed the ORIGINAL — the overlay is shared, not copied
|
||||
r591_copy_i18n_test.go:59: R-591: adding a language to the copy's I18n map added it to the ORIGINAL — the map is shared
|
||||
--- FAIL: TestDeepCopyStackI18nIsNotShared (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestDeepCopyStackI18nIsNotShared -v ./internal/stacks) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestDeepCopyStackI18nIsNotShared
|
||||
--- PASS: TestDeepCopyStackI18nIsNotShared (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
|
||||
|
||||
==================== R-568 — red-proof ====================
|
||||
Fix undone in: internal/web/disk_health.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestDiskHealthRows_OrderIsStable -v ./internal/web) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestDiskHealthRows_OrderIsStable
|
||||
r568_disk_order_test.go:31: R-568: the same two disks render in a different order depending on the agent's order: [nvme0n1 sda] vs [sda nvme0n1]
|
||||
--- FAIL: TestDiskHealthRows_OrderIsStable (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.009s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestDiskHealthRows_OrderIsStable -v ./internal/web) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestDiskHealthRows_OrderIsStable
|
||||
--- PASS: TestDiskHealthRows_OrderIsStable (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.009s
|
||||
|
||||
==================== R-567 — red-proof ====================
|
||||
Fix undone in: internal/web/templates/layout.html (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestStorageWizardPages_OpenTheStorageNavGroup -v ./internal/web) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestStorageWizardPages_OpenTheStorageNavGroup
|
||||
r567_storage_wizard_nav_test.go:21: R-567 storage_init: the storage menu group is not open
|
||||
r567_storage_wizard_nav_test.go:21: R-567 storage_attach: the storage menu group is not open
|
||||
--- FAIL: TestStorageWizardPages_OpenTheStorageNavGroup (0.06s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.072s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestStorageWizardPages_OpenTheStorageNavGroup -v ./internal/web) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestStorageWizardPages_OpenTheStorageNavGroup
|
||||
--- PASS: TestStorageWizardPages_OpenTheStorageNavGroup (0.06s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.069s
|
||||
|
||||
==================== R-363 + R-547 — red-proof ====================
|
||||
Fix undone in: cmd/controller/main.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestFillWatchRunsOnAnInterval -v ./cmd/controller) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestFillWatchRunsOnAnInterval
|
||||
r363_fillwatch_interval_test.go:56: R-363: main.go registers fillWatcher.Check on a periodic (sched.Every, fillWatchInterval) job 0 times, want 1 — without it the fill check is daily only
|
||||
--- FAIL: TestFillWatchRunsOnAnInterval (0.01s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.016s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestFillWatchRunsOnAnInterval -v ./cmd/controller) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestFillWatchRunsOnAnInterval
|
||||
--- PASS: TestFillWatchRunsOnAnInterval (0.01s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.017s
|
||||
|
||||
==================== R-10 — red-proof ====================
|
||||
Fix undone in: internal/appbackup/dbdump.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestDumpOneTo_SyncsTheDumpDirectoryAfterRename -v ./internal/appbackup) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestDumpOneTo_SyncsTheDumpDirectoryAfterRename
|
||||
r10_dump_dirsync_test.go:51: R-10: the dump directory was not fsynced after the rename: synced=[], want [/tmp/TestDumpOneTo_SyncsTheDumpDirectoryAfterRename3850100245/002/unit]
|
||||
--- FAIL: TestDumpOneTo_SyncsTheDumpDirectoryAfterRename (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/appbackup 0.009s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestDumpOneTo_SyncsTheDumpDirectoryAfterRename -v ./internal/appbackup) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestDumpOneTo_SyncsTheDumpDirectoryAfterRename
|
||||
--- PASS: TestDumpOneTo_SyncsTheDumpDirectoryAfterRename (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/appbackup 0.008s
|
||||
|
||||
==================== R-552 (a: the clear itself) — red-proof ====================
|
||||
Fix undone in: internal/backup/restore_record.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR552 -v ./internal/api) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR552_RemoveClearsTheInterruptedRestoreNotice
|
||||
r552_remove_clears_notice_test.go:60: R-552: the removed app's interrupted-restore notice is still listed
|
||||
r552_remove_clears_notice_test.go:67: R-552: the notice comes back after a restart — the clear was not persisted
|
||||
--- FAIL: TestR552_RemoveClearsTheInterruptedRestoreNotice (0.02s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.023s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR552 -v ./internal/api) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR552_RemoveClearsTheInterruptedRestoreNotice
|
||||
--- PASS: TestR552_RemoveClearsTheInterruptedRestoreNotice (0.02s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.024s
|
||||
|
||||
==================== R-552 (b: the wiring in removeStack) — red-proof ====================
|
||||
Fix undone in: internal/api/router.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR552 -v ./internal/api) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR552_RemoveClearsTheInterruptedRestoreNotice
|
||||
r552_remove_clears_notice_test.go:87: R-552: removeStack does not call clearInterruptedRestoreNotice
|
||||
--- FAIL: TestR552_RemoveClearsTheInterruptedRestoreNotice (0.02s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.023s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR552 -v ./internal/api) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR552_RemoveClearsTheInterruptedRestoreNotice
|
||||
--- PASS: TestR552_RemoveClearsTheInterruptedRestoreNotice (0.02s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.027s
|
||||
|
||||
==================== R-251 — red-proof ====================
|
||||
Fix undone in: internal/backup/offbox_inventory.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR251 -v ./internal/backup) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR251_MarkerTagIsNotAnApp
|
||||
r251_marker_tag_test.go:48: R-251: the recovery listing shows [calibre-web felhom-offbox]; want only [calibre-web] — the marker tag is not an app
|
||||
r251_marker_tag_test.go:51: R-251: 2 size calls for one app, want 1
|
||||
r251_marker_tag_test.go:59: R-251: the marker tag is reported as an app with a snapshot time
|
||||
--- FAIL: TestR251_MarkerTagIsNotAnApp (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR251 -v ./internal/backup) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR251_MarkerTagIsNotAnApp
|
||||
--- PASS: TestR251_MarkerTagIsNotAnApp (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s
|
||||
|
||||
==================== R-104 — red-proof ====================
|
||||
Fix undone in: internal/backup/offbox.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR104_SurvivingLockIsNamed -v ./internal/backup) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR104_SurvivingLockIsNamed
|
||||
r104_lock_class_test.go:36: R-104: a lock that survived the self-heal is classed "unknown", want "locked"
|
||||
r104_lock_class_test.go:40: R-104: the Hungarian message does not name the lock: "A távoli mentés ismeretlen okból nem sikerült (1m0s): offbox backup app: exit status 1 (offsite repository is still locked after the self-heal)"
|
||||
r104_lock_class_test.go:44: R-104: the English message does not name the lock: "The remote backup failed for an unknown reason (1m0s): offbox backup app: exit status 1 (offsite repository is still locked after the self-heal)"
|
||||
r104_lock_class_test.go:49: R-104: restic's lock text is classed "unknown", want "locked"
|
||||
--- FAIL: TestR104_SurvivingLockIsNamed (0.01s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.017s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR104_SurvivingLockIsNamed -v ./internal/backup) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR104_SurvivingLockIsNamed
|
||||
--- PASS: TestR104_SurvivingLockIsNamed (0.01s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.017s
|
||||
|
||||
==================== R-619 — red-proof ====================
|
||||
Fix undone in: internal/api/router.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR619 -v ./internal/api) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR619_PasswordFieldIsServedAsRequired
|
||||
r619_password_required_test.go:81: R-619: the password field reaches the wire as required:false, but the deploy refuses without it
|
||||
--- FAIL: TestR619_PasswordFieldIsServedAsRequired (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.009s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR619 -v ./internal/api) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR619_PasswordFieldIsServedAsRequired
|
||||
--- PASS: TestR619_PasswordFieldIsServedAsRequired (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.011s
|
||||
|
||||
==================== R-362 — red-proof ====================
|
||||
Fix undone in: internal/backup/restore_dir_err.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR362 -v ./internal/backup) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR362_DetachedDriveIsNamed
|
||||
r362_restore_drive_gone_test.go:41: R-362: the Hungarian refusal does not name the missing drive: "restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied"
|
||||
r362_restore_drive_gone_test.go:44: R-362: the English refusal does not name the missing drive: "restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied"
|
||||
r362_restore_drive_gone_test.go:49: R-362: a drive the registry marks disconnected is not named: "restore dir: mkdir /mnt/hdd_legacy: permission denied"
|
||||
--- FAIL: TestR362_DetachedDriveIsNamed (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.009s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR362 -v ./internal/backup) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR362_DetachedDriveIsNamed
|
||||
[WARN] [backup] restore dir /mnt/felhom-drives/hdd_1/backups/offsite-restore/app: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied — the drive /mnt/felhom-drives/hdd_1 (Kulso HDD) is not connected; reported as a missing drive (R-362)
|
||||
[WARN] [backup] restore dir /mnt/hdd_legacy/x: mkdir /mnt/hdd_legacy: permission denied — the drive /mnt/hdd_legacy (Regi HDD) is not connected; reported as a missing drive (R-362)
|
||||
--- PASS: TestR362_DetachedDriveIsNamed (0.01s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.017s
|
||||
|
||||
==================== R-675 — red-proof ====================
|
||||
Fix undone in: internal/web/handlers.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR675 -v ./internal/web) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR675_RefusalNamesTheWholeCopy
|
||||
r675_refusal_whole_copy_test.go:33: R-675 whole copy on the second drive: the Hungarian refusal reads "Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a második meghajtó másolatából állíthatók vissza: „Fájlok visszaállítása”."; want it to name "Teljes visszaállítás a másolatból" and not "Fájlok visszaállítása"
|
||||
r675_refusal_whole_copy_test.go:36: R-675 whole copy on the second drive: the English refusal reads "This backup does not hold the files of the app, so we do not restore the database over them — the files stay where they are. The files can be restored from the copy on the second drive: “Restore files”."; want it to name "Full restore from the copy"
|
||||
--- FAIL: TestR675_RefusalNamesTheWholeCopy (0.06s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.073s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR675 -v ./internal/web) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR675_RefusalNamesTheWholeCopy
|
||||
--- PASS: TestR675_RefusalNamesTheWholeCopy (0.07s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.081s
|
||||
|
||||
==================== R-240 — red-proof ====================
|
||||
Fix undone in: internal/backup/offbox.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR240 -v ./internal/backup) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR240_ZeroSelectionRunDoesNotSaySuccess
|
||||
[ERROR] [backup] Database discovery failed: docker ps failed: R-650: refused to run the real "docker ps --format {{.ID}}\t{{.Names}}\t{{.Label \"com.docker.compose.project\"}}\t{{.Image}} --filter status=running" under go test (this host may be production Docker); use a seam or a stub on PATH, or set FELHOM_TEST_REAL_DOCKER=1 deliberately
|
||||
[WARN] [offbox] pre-push dump leg failed (docker ps failed: R-650: refused to run the real "docker ps --format {{.ID}}\t{{.Names}}\t{{.Label \"com.docker.compose.project\"}}\t{{.Image}} --filter status=running" under go test (this host may be production Docker); use a seam or a stub on PATH, or set FELHOM_TEST_REAL_DOCKER=1 deliberately) — continuing with the existing dumps; the snapshot's DB half may be older than its files
|
||||
r240_zero_selection_test.go:30: R-240: the note for a run that saved nothing still calls itself successful: "Sikeres — nincs mentésre jelölt alkalmazás"
|
||||
r240_zero_selection_test.go:33: R-240: the note does not say the run saved nothing: "Sikeres — nincs mentésre jelölt alkalmazás"
|
||||
--- FAIL: TestR240_ZeroSelectionRunDoesNotSaySuccess (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.008s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR240 -v ./internal/backup) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR240_ZeroSelectionRunDoesNotSaySuccess
|
||||
[ERROR] [backup] Database discovery failed: docker ps failed: R-650: refused to run the real "docker ps --format {{.ID}}\t{{.Names}}\t{{.Label \"com.docker.compose.project\"}}\t{{.Image}} --filter status=running" under go test (this host may be production Docker); use a seam or a stub on PATH, or set FELHOM_TEST_REAL_DOCKER=1 deliberately
|
||||
[WARN] [offbox] pre-push dump leg failed (docker ps failed: R-650: refused to run the real "docker ps --format {{.ID}}\t{{.Names}}\t{{.Label \"com.docker.compose.project\"}}\t{{.Image}} --filter status=running" under go test (this host may be production Docker); use a seam or a stub on PATH, or set FELHOM_TEST_REAL_DOCKER=1 deliberately) — continuing with the existing dumps; the snapshot's DB half may be older than its files
|
||||
--- PASS: TestR240_ZeroSelectionRunDoesNotSaySuccess (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.009s
|
||||
|
||||
==================== R-256 (hu copy restored to the old sentence) — red-proof ====================
|
||||
Fix undone in: internal/i18n/locales/hu.json (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR256 -v ./internal/web) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR256_R257_OffboxRefusalsNameARoute
|
||||
r256_r257_offbox_refusals_test.go:37: R-256: the Hungarian refusal names no route or still names the component: "A mentéskezelő nem elérhető."
|
||||
r256_r257_offbox_refusals_test.go:45: R-256: the restore page's twin refusal differs: "A mentések kezelése most nem érhető el. Próbáld újra néhány perc múlva; ha akkor sem megy, keresd a Felhom ügyfélszolgálatát." vs "A mentéskezelő nem elérhető."
|
||||
--- FAIL: TestR256_R257_OffboxRefusalsNameARoute (0.06s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.070s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR256 -v ./internal/web) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR256_R257_OffboxRefusalsNameARoute
|
||||
--- PASS: TestR256_R257_OffboxRefusalsNameARoute (0.06s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.075s
|
||||
|
||||
==================== R-257 (hu copy restored to the old sentence) — red-proof ====================
|
||||
Fix undone in: internal/i18n/locales/hu.json (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR256 -v ./internal/web) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR256_R257_OffboxRefusalsNameARoute
|
||||
r256_r257_offbox_refusals_test.go:67: R-257: the Hungarian refusal still says "offsite": "Az offsite tároló nincs elárvult állapotban."
|
||||
r256_r257_offbox_refusals_test.go:67: R-257: the Hungarian refusal still says "elárvult": "Az offsite tároló nincs elárvult állapotban."
|
||||
r256_r257_offbox_refusals_test.go:71: R-257: the Hungarian refusal does not say why or where to go: "Az offsite tároló nincs elárvult állapotban."
|
||||
--- FAIL: TestR256_R257_OffboxRefusalsNameARoute (0.06s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.071s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR256 -v ./internal/web) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR256_R257_OffboxRefusalsNameARoute
|
||||
--- PASS: TestR256_R257_OffboxRefusalsNameARoute (0.07s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.078s
|
||||
|
||||
==================== R-365 — red-proof ====================
|
||||
Fix undone in: internal/web/handlers.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR365 -v ./internal/web) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR365_OverdueCountdownIsNotFutureTense
|
||||
r365_abandon_overdue_test.go:41: R-365: an overdue countdown does not say the deletion is due since 2026-10-04
|
||||
r365_abandon_overdue_test.go:44: R-365: an overdue countdown still renders a past date in the future tense
|
||||
--- FAIL: TestR365_OverdueCountdownIsNotFutureTense (0.14s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.149s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR365 -v ./internal/web) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR365_OverdueCountdownIsNotFutureTense
|
||||
--- PASS: TestR365_OverdueCountdownIsNotFutureTense (0.15s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.159s
|
||||
|
||||
==================== R-564 — red-proof (gate; Python, so no go test -run) ====================
|
||||
Fix undone: SPLIT_PATTERNS emptied in scripts/retrieval_promise_gate.py; decoy harness run.
|
||||
$ (cd controller && python3 scripts/test_gate_decoys.py | grep split) # FIX UNDONE
|
||||
ok retrieval-promise/split-verb decoy rejected
|
||||
FAIL: retrieval-promise/en-ok: rc=1, expected accept
|
||||
FAIL: retrieval-promise/split-ok: rc=1, expected accept
|
||||
rc=1
|
||||
--- fix restored ---
|
||||
$ (cd controller && python3 scripts/test_gate_decoys.py | grep split) # FIX RESTORED
|
||||
ok retrieval-promise/split-verb decoy rejected
|
||||
ok retrieval-promise/split-ok genuine accepted
|
||||
all 25 controller decoys behaved — labels do not satisfy these gates
|
||||
$ python3 scripts/retrieval_promise_gate.py | head -1
|
||||
retrieval-promise gate OK — 42 surface(s) incl. 1 Go handler file(s), 18 registered claim(s) + 16 English, none unregistered
|
||||
NOTE: the record above is NOT a valid red-proof — with SPLIT_PATTERNS emptied the seven new registrations go stale, so the gate fails for that reason, not because it caught the decoy. The valid red-proof follows.
|
||||
|
||||
==================== R-564 — red-proof (valid) ====================
|
||||
Decoy planted in hu.json: launcher.link_masolasa = 'A régi mentéseid a kóddal bármikor állíthatók vissza.' (a split-verb retrieval PROMISE)
|
||||
$ python3 scripts/<PRE-FIX retrieval_promise_gate.py from HEAD> # FIX UNDONE
|
||||
retrieval-promise gate OK — 42 surface(s) incl. 1 Go handler file(s), 11 registered claim(s) + 16 English, none unregistered
|
||||
rc=0 <- the pre-fix gate PASSES the planted promise (blind)
|
||||
$ python3 scripts/retrieval_promise_gate.py # FIX IN PLACE
|
||||
launcher.html:61 unregistered retrieval claim (állíthatók vissza):
|
||||
RETRIEVAL-PROMISE GATE FAILED: 1 unregistered, 0 stale, across 42 template(s).
|
||||
rc=1 <- convicted
|
||||
--- decoy removed ---
|
||||
$ python3 scripts/retrieval_promise_gate.py
|
||||
retrieval-promise gate OK — 42 surface(s) incl. 1 Go handler file(s), 18 registered claim(s) + 16 English, none unregistered
|
||||
|
||||
==================== R-425 — red-proof (gate) ====================
|
||||
Decoy: a NEW template internal/web/templates/backups_offbox_extra.html containing 'NAS-mentés'.
|
||||
$ python3 scripts/<PRE-FIX offbox_rename_gate.py from HEAD> # FIX UNDONE
|
||||
offbox rename gate OK — Tier-3 is 'Tavoli mentes' everywhere customer-facing
|
||||
rc=0 <- the fixed FILES list never looks at the new file
|
||||
$ python3 scripts/offbox_rename_gate.py # FIX IN PLACE
|
||||
internal/web/templates/backups_offbox_extra.html:1 [NAS-mentés] <p>A NAS-ment\xe9s be\xe1ll\xedt\xe1sa</p>
|
||||
OFFBOX RENAME GATE FAILED: 1 customer-facing NAS-branding string(s) remain
|
||||
rc=1 <- convicted
|
||||
--- decoy removed ---
|
||||
offbox rename gate OK — Tier-3 is 'Tavoli mentes' everywhere customer-facing (25 file(s) + 121 bundle value(s) named by them)
|
||||
|
||||
==================== R-565 (the ' mp' unit the new detector found, put back) — red-proof ====================
|
||||
Fix undone in: internal/web/templates/backups_remote.html (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestI18nEnglishPages$ -v ./internal/web) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestI18nEnglishPages
|
||||
i18n_parity_test.go:709: backups_remote_full: ASCII-only Hungarian word "mp" on the English page (R-565), line 434: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
i18n_parity_test.go:709: backups_remote_incomplete: ASCII-only Hungarian word "mp" on the English page (R-565), line 348: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
i18n_parity_test.go:709: backups_remote_error: ASCII-only Hungarian word "mp" on the English page (R-565), line 333: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
i18n_parity_test.go:709: backups_remote_running: ASCII-only Hungarian word "mp" on the English page (R-565), line 353: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
i18n_parity_test.go:709: backups_remote_pending_agent: ASCII-only Hungarian word "mp" on the English page (R-565), line 339: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
i18n_parity_test.go:709: backups_remote_pending_old: ASCII-only Hungarian word "mp" on the English page (R-565), line 339: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
i18n_parity_test.go:709: backups_remote_stale: ASCII-only Hungarian word "mp" on the English page (R-565), line 337: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
i18n_parity_test.go:709: backups_remote_stale_old: ASCII-only Hungarian word "mp" on the English page (R-565), line 337: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
i18n_parity_test.go:709: backups_remote_escrowed: ASCII-only Hungarian word "mp" on the English page (R-565), line 336: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
i18n_parity_test.go:709: backups_remote_notconf_hub: ASCII-only Hungarian word "mp" on the English page (R-565), line 285: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
i18n_parity_test.go:709: backups_remote_notconf: ASCII-only Hungarian word "mp" on the English page (R-565), line 284: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
i18n_parity_test.go:709: backups_remote_empty: ASCII-only Hungarian word "mp" on the English page (R-565), line 258: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
i18n_parity_test.go:709: backups_remote_offsite_offer_fit: ASCII-only Hungarian word "mp" on the English page (R-565), line 349: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }"
|
||||
--- FAIL: TestI18nEnglishPages (4.23s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 4.243s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestI18nEnglishPages$ -v ./internal/web) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestI18nEnglishPages
|
||||
--- PASS: TestI18nEnglishPages (4.27s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 4.285s
|
||||
|
||||
==================== R-603 (an apostrophe planted in a Go-named English value) — red-proof ====================
|
||||
Fix undone in: internal/i18n/locales/en.json (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR603 -v ./internal/i18n) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR603_GoNamedValuesDoNotHideBehindHTMLEscaping
|
||||
r603_escape_test.go:126: control: a planted apostrophe in err.backup.a_pillanatkep_egy_utvonala_ervenytelen was not caught (got [err.backup.a_pillanatkep_egy_utvonala_ervenytelen note.offsite.fail_locked])
|
||||
--- FAIL: TestR603_GoNamedValuesDoNotHideBehindHTMLEscaping (0.29s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/i18n 0.296s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR603 -v ./internal/i18n) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR603_GoNamedValuesDoNotHideBehindHTMLEscaping
|
||||
--- PASS: TestR603_GoNamedValuesDoNotHideBehindHTMLEscaping (0.34s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/i18n 0.343s
|
||||
|
||||
==================== R-454 — the gate seen RED on the unformatted tree, before the formatting pass ====================
|
||||
$ (cd controller && python3 scripts/gofmt_gate.py)
|
||||
not gofmt-clean: cmd/controller/main.go
|
||||
not gofmt-clean: internal/agentapi/diskverdict.go
|
||||
not gofmt-clean: internal/api/update_reason_test.go
|
||||
not gofmt-clean: internal/appbackup/namespace_root_test.go
|
||||
not gofmt-clean: internal/appbackup/r381_undo_naming_test.go
|
||||
not gofmt-clean: internal/backup/r669_applied_meta_test.go
|
||||
not gofmt-clean: internal/family/family.go
|
||||
not gofmt-clean: internal/infra/infra.go
|
||||
not gofmt-clean: internal/notify/r636_oom_storm_test.go
|
||||
not gofmt-clean: internal/quiesce/tiers_test.go
|
||||
not gofmt-clean: internal/stacks/delete.go
|
||||
not gofmt-clean: internal/stacks/life_records.go
|
||||
GOFMT GATE FAILED: 12 file(s) — run `gofmt -w <file>` (formatting only, no behaviour change)
|
||||
rc=1
|
||||
|
||||
==================== R-208 (controller half: ARGs moved back above go mod download) — red-proof ====================
|
||||
Fix undone in: Dockerfile (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR208 -v ./cmd/controller) # FIX UNDONE
|
||||
rc=0
|
||||
=== RUN TestR208_DockerfileVersionArgsSitBelowModuleDownload
|
||||
--- PASS: TestR208_DockerfileVersionArgsSitBelowModuleDownload (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.009s
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR208 -v ./cmd/controller) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR208_DockerfileVersionArgsSitBelowModuleDownload
|
||||
--- PASS: TestR208_DockerfileVersionArgsSitBelowModuleDownload (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.006s
|
||||
NOTE: NOT CONVICTED above — the test kept the LAST declaration of each ARG, so a duplicate declaration above the download hid behind the one below. Test fixed to keep the FIRST declaration; red-proof re-run below.
|
||||
|
||||
==================== R-208 (re-run after the test fix) — red-proof ====================
|
||||
Fix undone in: Dockerfile (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR208 -v ./cmd/controller) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR208_DockerfileVersionArgsSitBelowModuleDownload
|
||||
r208_dockerfile_args_test.go:50: R-208: ARG VERSION (line 12) is declared above `go mod download` (line 19), so every build with a new value re-downloads the modules
|
||||
r208_dockerfile_args_test.go:50: R-208: ARG GIT_COMMIT (line 13) is declared above `go mod download` (line 19), so every build with a new value re-downloads the modules
|
||||
--- FAIL: TestR208_DockerfileVersionArgsSitBelowModuleDownload (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.009s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR208 -v ./cmd/controller) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR208_DockerfileVersionArgsSitBelowModuleDownload
|
||||
--- PASS: TestR208_DockerfileVersionArgsSitBelowModuleDownload (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.008s
|
||||
|
||||
==================== R-591 follow-up (DataPaths + AfterLoad copies removed) — red-proof ====================
|
||||
Fix undone in: internal/stacks/manager.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestDeepCopyStackMetaSharesNoReference -v ./internal/stacks) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestDeepCopyStackMetaSharesNoReference
|
||||
r591_copy_meta_alias_test.go:21: R-591: deepCopyStack leaves Meta.AfterLoad shared with the original — a write through the copy changes the stack
|
||||
r591_copy_meta_alias_test.go:21: R-591: deepCopyStack leaves Meta.DataPaths shared with the original — a write through the copy changes the stack
|
||||
--- FAIL: TestDeepCopyStackMetaSharesNoReference (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestDeepCopyStackMetaSharesNoReference -v ./internal/stacks) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestDeepCopyStackMetaSharesNoReference
|
||||
--- PASS: TestDeepCopyStackMetaSharesNoReference (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
|
||||
|
||||
==================== R-365 follow-up (layout banner at 0 days; overdue flag removed) — red-proof ====================
|
||||
Fix undone in: internal/web/recovery_handlers.go (the fix text replaced by the pre-fix shape)
|
||||
$ (cd controller && go test -count=1 -run TestR365_Banner -v ./internal/web) # FIX UNDONE
|
||||
rc=1
|
||||
=== RUN TestR365_BannerSaysDueAtZeroDays
|
||||
r365_banner_overdue_test.go:47: R-365: at 0 days the bar does not say the deletion is due:
|
||||
r365_banner_overdue_test.go:50: R-365: at 0 days the bar offers to stop reminders over a due deletion
|
||||
--- FAIL: TestR365_BannerSaysDueAtZeroDays (0.14s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.154s
|
||||
FAIL
|
||||
--- fix restored ---
|
||||
$ (cd controller && go test -count=1 -run TestR365_Banner -v ./internal/web) # FIX RESTORED
|
||||
rc=0
|
||||
=== RUN TestR365_BannerSaysDueAtZeroDays
|
||||
--- PASS: TestR365_BannerSaysDueAtZeroDays (0.12s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.132s
|
||||
|
||||
### CI fix (lead, 2026-10-05 ~18:20Z): the new gofmt gate was INCONCLUSIVE on the CI runner (no Go) — CI run 1371 FAILED
|
||||
Fix: in CI (GITEA_ACTIONS/GITHUB_ACTIONS=true) with no gofmt reachable the gate prints "NOT CHECKED in CI" and exits 0;
|
||||
elsewhere a missing gofmt stays INCONCLUSIVE. Decoys gofmt/ci-without-go and gofmt/dev-without-go (empty PATH).
|
||||
Simulated CI (env -i, empty PATH, GITEA_ACTIONS=true): "gofmt gate NOT CHECKED in CI …" rc=0.
|
||||
Red-proof (CI branch disabled): FAIL: gofmt/ci-without-go: want rc=0 and 'NOT CHECKED in CI', got rc=2.
|
||||
@@ -0,0 +1,4 @@
|
||||
== controller delivery 2026-10-05T18:28:59Z
|
||||
demo-hp gitea.dooplex.hu/admin/felhom-controller:0.297.0 Up 32 seconds (healthy)
|
||||
demo-felhom gitea.dooplex.hu/admin/felhom-controller:0.297.0 Up 35 seconds (healthy)
|
||||
tester-1 (hub customer page, version strings seen): 6 0.297.0 4 0.296.0 4 0.295.0
|
||||
@@ -0,0 +1,11 @@
|
||||
== hub deploy 2026-10-05T18:11:56Z
|
||||
image=gitea.dooplex.hu/admin/felhom-hub:0.137.0
|
||||
sync=Synced health=Healthy op=Succeeded rev=d75ad0fdf3daf3f3b2a1690746d9a6a70ee4104c
|
||||
2026/10/05 20:11:05 [INFO] felhom-hub 0.137.0 starting
|
||||
healthz 200
|
||||
2026/10/05 20:11:06 [INFO] osupdates: the Docker engine set is approved only by the operator, after 2 healthy ring-0 night(s)
|
||||
== System page 2026-10-05T18:12:05Z: root-files / agent cells
|
||||
73: 0 armed unknown 0.142.0 → 0.147.0 (since 2026-10-05)
|
||||
89: 0 armed 0.147.0 0.147.0
|
||||
105: 2 armed 0.147.0 0.147.0
|
||||
121: 2 TRIPPED 2026-10-05T07:57:17Z 0.147.0 0.147.0
|
||||
@@ -0,0 +1,13 @@
|
||||
== agent_update 0.147.0 (sha 642c4d19…) signed with felhom-op-1, ttl 45m, 2026-10-05T17:02:13Z
|
||||
-- demo-hp-bb76ea
|
||||
signed: op=agent_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=b05c992ea1b100af9f53df761c76d440 expires=2026-10-05T17:47:13Z
|
||||
wrote envelope to <scratch>/env-demo-hp-bb76ea-agent_update.json
|
||||
uploaded signed op to the hub jobs queue
|
||||
-- demo-felhom-8363b5
|
||||
signed: op=agent_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=67f8cbeed47a8a107d3809e230837af0 expires=2026-10-05T17:47:13Z
|
||||
wrote envelope to <scratch>/env-demo-felhom-8363b5-agent_update.json
|
||||
uploaded signed op to the hub jobs queue
|
||||
-- tester-1-d70be4
|
||||
signed: op=agent_update host=tester-1-d70be4 guest="" key_id=felhom-op-1 nonce=c0decf168b76b852d71db79378bcb795 expires=2026-10-05T17:47:13Z
|
||||
wrote envelope to <scratch>/env-tester-1-d70be4-agent_update.json
|
||||
uploaded signed op to the hub jobs queue
|
||||
@@ -0,0 +1,14 @@
|
||||
== System page 2026-10-05T17:13:38Z after agent_update: demo-hp, demo-felhom, tester-1 rows read Agent 0.147.0 (root files 0.146.1); Tester-2 0.142.0 → 0.147.0 (offline)
|
||||
== agent_config_update 0.147.0 (bundle sha 326527d0…), 2026-10-05T17:13:38Z
|
||||
-- demo-hp-bb76ea
|
||||
signed: op=agent_config_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=f0d53302c1591959f9577dd2660690d1 expires=2026-10-05T17:58:38Z
|
||||
wrote envelope to <scratch>/env-demo-hp-bb76ea-agent_config_update.json
|
||||
uploaded signed op to the hub jobs queue
|
||||
-- demo-felhom-8363b5
|
||||
signed: op=agent_config_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=c92ab23d76d00821a36ec20018c69b6a expires=2026-10-05T17:58:38Z
|
||||
wrote envelope to <scratch>/env-demo-felhom-8363b5-agent_config_update.json
|
||||
uploaded signed op to the hub jobs queue
|
||||
-- tester-1-d70be4
|
||||
signed: op=agent_config_update host=tester-1-d70be4 guest="" key_id=felhom-op-1 nonce=c51c5f6448ec01d4919408e0590dd1bd expires=2026-10-05T17:58:38Z
|
||||
wrote envelope to <scratch>/env-tester-1-d70be4-agent_config_update.json
|
||||
uploaded signed op to the hub jobs queue
|
||||
@@ -0,0 +1,4 @@
|
||||
== vouch 2026-10-05T17:01:27Z: POST /configuration/artifacts (Basic + X-Felhom-Operator), agent 0.147.0, golden 0.296.0, min_agent 0.131.0
|
||||
HTTP/1.1 303 See Other
|
||||
Location: /configuration?flash=artifacts_set
|
||||
2026/10/05 19:01:56 [INFO] Artifact manifest set: agent=0.147.0 golden=0.296.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d"
|
||||
@@ -0,0 +1,11 @@
|
||||
== vouch 2026-10-05T18:27:46Z: agent 0.147.0, golden 0.297.0, min_agent 0.131.0
|
||||
HTTP/1.1 303 See Other
|
||||
Location: /configuration?flash=artifacts_set
|
||||
== floors 2026-10-05T18:28:16Z: POST /customers/<id>/floor min_controller_version=0.297.0 min_agent=0.131.0
|
||||
demo-hp: Location: /customers/demo-hp?flash=floor_set
|
||||
demo-felhom: Location: /customers/demo-felhom?flash=floor_set
|
||||
tester-1: Location: /customers/tester-1?flash=floor_set
|
||||
2026/10/05 20:28:16 [INFO] Artifact manifest set: agent=0.147.0 golden=0.297.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d"
|
||||
2026/10/05 20:28:16 [INFO] Customer demo-hp controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0")
|
||||
2026/10/05 20:28:17 [INFO] Customer demo-felhom controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0")
|
||||
2026/10/05 20:28:17 [INFO] Customer tester-1 controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0")
|
||||
@@ -0,0 +1,366 @@
|
||||
# felhom.eu burndown2 red-proofs, 2026-10-05 (fix undone -> test fails; fix restored -> test passes)
|
||||
|
||||
### R-277 (offsite row bytes)
|
||||
$ (cd . && go test ./internal/web -run 'TestOffsiteRow_SmallRepoNotZeroGB' -v -count=1)
|
||||
--- fix UNDONE (rc=1):
|
||||
=== RUN TestOffsiteRow_SmallRepoNotZeroGB
|
||||
r277_r92_bytes_test.go:27: 162 KB repo rendered as "0.0 GB", want "162.0 KB"
|
||||
--- FAIL: TestOffsiteRow_SmallRepoNotZeroGB (0.04s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.062s
|
||||
FAIL
|
||||
--- fix RESTORED (rc=0):
|
||||
=== RUN TestOffsiteRow_SmallRepoNotZeroGB
|
||||
--- PASS: TestOffsiteRow_SmallRepoNotZeroGB (0.04s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.064s
|
||||
|
||||
### R-92 (PBS DR exact bytes)
|
||||
$ (cd . && go test ./internal/web -run 'TestPBSDRPanel_ExactBytesShowsSmallDelta' -v -count=1)
|
||||
--- fix UNDONE (rc=1):
|
||||
=== RUN TestPBSDRPanel_ExactBytesShowsSmallDelta
|
||||
r277_r92_bytes_test.go:55: PBS DR panel must show exact bytes:
|
||||
PBS DR datastore</h3>
|
||||
|
||||
<table class="detail-table">
|
||||
<tr><th style="width: 12rem;">Datastore</th><td><code>felhom-offsite</code> (ep0)</td></tr>
|
||||
<tr><th>Capacity</th><td>40.0 GB</td></tr>
|
||||
<tr><th>Used</th><td>8.0 GB · 20% full</td></tr>
|
||||
</table>
|
||||
<div class="bar" style="margin: 0.4rem 0 0.9rem;"><div class="bar-fill bar-ok" style="width: 20%;"></div></div>
|
||||
<table class="detail-table">
|
||||
<tr><th style="width: 12rem;">Polled</th><td>just now</td></tr>
|
||||
</table>
|
||||
|
||||
</section>
|
||||
|
||||
<p class="text-muted" style="margin: 0 0 1rem; font-size: 0.85em;">
|
||||
The endpoint below IS the PBS DR host. Peer allocation and endpoint sync currently use the
|
||||
lowest endpoint id (ep0); per-endpoint allocation is a future work item.
|
||||
</p>
|
||||
|
||||
|
||||
<section class="card" style="margin-bottom: 1.5rem;">
|
||||
<h3>Endpoint</h3>
|
||||
<p class="text-muted">Not configured. Add one below (or via <code>PUT /api/v1/admin/wg/endpoint</code>, runbook: offsite-endpoint.md).</p>
|
||||
</section>
|
||||
|
||||
|
||||
|
||||
<section class="card" style="margin-bottom: 1.5rem;">
|
||||
<h3 id="ep-form-title">Add endpoint</h3>
|
||||
<form method="POST" action="/offsite/endpoints" id="ep-form" onsubmit="return epFormSubmitCheck()"
|
||||
style="display: grid; grid-template-columns: auto 1fr; gap: 0.5rem; align-items: center; max-width: 44em; margin-top: 0.75rem;">
|
||||
<input type="hidden" name="_csrf" value="">
|
||||
<label style="font-size: 0.9em;">Endpoint id</label>
|
||||
<input type="text" name="endpoint_id" id="ep-id" placeholder="ep1" style="padding: 0.3em 0.5em;">
|
||||
<label style="font-size: 0.9em;">DNS name</label>
|
||||
<input type="text" name="dns_name" id="ep-dns" placeholder="ep1.felhom.eu" style="padding: 0.3em 0.5em;">
|
||||
<label style="font-size: 0.9em;">WG port</label>
|
||||
<input type="number" name="wg_port" id="ep-port" min="1" max="65535" placeholder="443" style="padding: 0.3em 0.5em;">
|
||||
--- fix RESTORED (rc=0):
|
||||
=== RUN TestPBSDRPanel_ExactBytesShowsSmallDelta
|
||||
--- PASS: TestPBSDRPanel_ExactBytesShowsSmallDelta (0.07s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.093s
|
||||
|
||||
|
||||
### R-581 (newest report tie-break)
|
||||
$ (cd . && go test ./internal/store -run TestGetCustomers_SameSecondReportsOneRowNewestWins -v -count=1)
|
||||
--- fix UNDONE (rc=1):
|
||||
=== RUN TestGetCustomers_SameSecondReportsOneRowNewestWins
|
||||
r581_newest_report_test.go:32: GetCustomers returned 2 rows for one customer, want exactly 1
|
||||
--- FAIL: TestGetCustomers_SameSecondReportsOneRowNewestWins (0.03s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/store 0.040s
|
||||
FAIL
|
||||
--- fix RESTORED (rc=0):
|
||||
=== RUN TestGetCustomers_SameSecondReportsOneRowNewestWins
|
||||
--- PASS: TestGetCustomers_SameSecondReportsOneRowNewestWins (0.03s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/store 0.040s
|
||||
|
||||
### R-600 (delete cascade WG peer push + honest COMPLETE line)
|
||||
$ (cd . && go test ./internal/web -run 'TestDeleteCascade_TriggersWGPeerPushAndSaysSo' -v -count=1)
|
||||
--- fix UNDONE (rc=1):
|
||||
=== RUN TestDeleteCascade_TriggersWGPeerPushAndSaysSo
|
||||
customer_delete_test.go:642: WG peer-sync triggers = 0, want exactly 1 (the endpoint must drop the peer now, not on the next tick)
|
||||
--- FAIL: TestDeleteCascade_TriggersWGPeerPushAndSaysSo (0.04s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.070s
|
||||
FAIL
|
||||
--- fix RESTORED (rc=0):
|
||||
=== RUN TestDeleteCascade_TriggersWGPeerPushAndSaysSo
|
||||
--- PASS: TestDeleteCascade_TriggersWGPeerPushAndSaysSo (0.05s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.072s
|
||||
|
||||
### R-600 (main wiring of SetWGPeerSync)
|
||||
$ (cd . && go test ./cmd/hub -run TestR600_MainWiresWGPeerSyncIntoWeb -v -count=1)
|
||||
--- fix UNDONE (rc=1):
|
||||
=== RUN TestR600_MainWiresWGPeerSyncIntoWeb
|
||||
r600_wiring_test.go:32: cmd/hub/main.go never calls webServer.SetWGPeerSync(<reconciler>.Trigger) — the delete cascade cannot push the peer removal
|
||||
--- FAIL: TestR600_MainWiresWGPeerSyncIntoWeb (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.024s
|
||||
FAIL
|
||||
--- fix RESTORED (rc=0):
|
||||
=== RUN TestR600_MainWiresWGPeerSyncIntoWeb
|
||||
--- PASS: TestR600_MainWiresWGPeerSyncIntoWeb (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.025s
|
||||
|
||||
### R-599 (ONLINE 409 says when deletion opens; host + cascade)
|
||||
$ (cd . && go test ./internal/web -run TestHostDelete_OnlineRefusalSaysWhenItOpens -v -count=1)
|
||||
--- fix UNDONE (rc=1):
|
||||
=== RUN TestHostDelete_OnlineRefusalSaysWhenItOpens
|
||||
r544_r599_host_delete_test.go:74: host-delete 409 body missing "min ago":
|
||||
Host is ONLINE — deletion is refused (a live agent would receive 401s permanently).
|
||||
r544_r599_host_delete_test.go:74: host-delete 409 body missing "deletion opens at 17:42 UTC":
|
||||
Host is ONLINE — deletion is refused (a live agent would receive 401s permanently).
|
||||
r544_r599_host_delete_test.go:74: host-delete 409 body missing "45m0s":
|
||||
Host is ONLINE — deletion is refused (a live agent would receive 401s permanently).
|
||||
r544_r599_host_delete_test.go:84: cascade ONLINE refusal = 409 "Delete refused: host gone-vm is ONLINE. Decommission the box first — the cascade never deletes a live host.\n", want 409 naming the opening time
|
||||
--- FAIL: TestHostDelete_OnlineRefusalSaysWhenItOpens (0.05s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.066s
|
||||
FAIL
|
||||
--- fix RESTORED (rc=0):
|
||||
=== RUN TestHostDelete_OnlineRefusalSaysWhenItOpens
|
||||
--- PASS: TestHostDelete_OnlineRefusalSaysWhenItOpens (0.04s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.065s
|
||||
|
||||
### R-544 (host-delete log states demotion)
|
||||
$ (cd . && go test ./internal/web -run TestHostDelete_LogSaysEscrowDemotedNotDeleted -v -count=1)
|
||||
--- fix UNDONE (rc=1):
|
||||
=== RUN TestHostDelete_LogSaysEscrowDemotedNotDeleted
|
||||
r544_r599_host_delete_test.go:32: log still says the escrow was deleted:
|
||||
[INFO] host deleted: esc-host (escrow deleted: true)
|
||||
r544_r599_host_delete_test.go:35: log must state the demotion:
|
||||
[INFO] host deleted: esc-host (escrow deleted: true)
|
||||
--- FAIL: TestHostDelete_LogSaysEscrowDemotedNotDeleted (0.04s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.064s
|
||||
FAIL
|
||||
--- fix RESTORED (rc=0):
|
||||
=== RUN TestHostDelete_LogSaysEscrowDemotedNotDeleted
|
||||
--- PASS: TestHostDelete_LogSaysEscrowDemotedNotDeleted (0.04s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.063s
|
||||
|
||||
### R-855 (start log prints effective Docker nights)
|
||||
$ (cd . && go test ./cmd/hub ./internal/osupdates -run 'TestR855_StartLogPrintsEffectiveDockerNights|TestDockerNightsEffective' -v -count=1)
|
||||
--- fix UNDONE (rc=1):
|
||||
=== RUN TestR855_StartLogPrintsEffectiveDockerNights
|
||||
r855_docker_nights_log_test.go:41: the Docker approval-nights start log must print osSvc.DockerNightsEffective(), not the raw field
|
||||
--- FAIL: TestR855_StartLogPrintsEffectiveDockerNights (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.022s
|
||||
=== RUN TestDockerNightsEffective
|
||||
--- PASS: TestDockerNightsEffective (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/osupdates 0.005s
|
||||
FAIL
|
||||
--- fix RESTORED (rc=0):
|
||||
=== RUN TestR855_StartLogPrintsEffectiveDockerNights
|
||||
--- PASS: TestR855_StartLogPrintsEffectiveDockerNights (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.022s
|
||||
=== RUN TestDockerNightsEffective
|
||||
--- PASS: TestDockerNightsEffective (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/osupdates 0.006s
|
||||
|
||||
### R-134 (zone candidates strip progressively; red = the old one-label parentDomain)
|
||||
$ (cd . && go test ./internal/cloudflare -run TestZoneCandidates -v -count=1)
|
||||
--- fix UNDONE (rc=1):
|
||||
=== RUN TestZoneCandidates
|
||||
unblock_test.go:24: zoneCandidates("a.b.felhom.eu") = [a.b.felhom.eu b.felhom.eu], want [a.b.felhom.eu b.felhom.eu felhom.eu]
|
||||
unblock_test.go:24: zoneCandidates("x.y.z.example.co") = [x.y.z.example.co y.z.example.co], want [x.y.z.example.co y.z.example.co z.example.co example.co]
|
||||
--- FAIL: TestZoneCandidates (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/cloudflare 0.006s
|
||||
FAIL
|
||||
--- fix RESTORED (rc=0):
|
||||
=== RUN TestZoneCandidates
|
||||
--- PASS: TestZoneCandidates (0.00s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/cloudflare 0.006s
|
||||
|
||||
|
||||
### R-292 (artifact sha flash names the cause; red = every failure -> artifact_sha_invalid, the old single flash)
|
||||
$ (cd . && go test ./internal/web -run 'TestArtifactSave_FlashNamesTheCause' -v -count=1)
|
||||
--- fix UNDONE (rc=1):
|
||||
=== RUN TestArtifactSave_FlashNamesTheCause
|
||||
=== RUN TestArtifactSave_FlashNamesTheCause/version_not_found
|
||||
r292_artifact_flash_test.go:57: flash = "artifact_sha_invalid", want "artifact_version_missing"
|
||||
=== RUN TestArtifactSave_FlashNamesTheCause/registry_failing
|
||||
r292_artifact_flash_test.go:57: flash = "artifact_sha_invalid", want "artifact_unverifiable"
|
||||
=== RUN TestArtifactSave_FlashNamesTheCause/no_sha_listed
|
||||
r292_artifact_flash_test.go:57: flash = "artifact_sha_invalid", want "artifact_sha_missing"
|
||||
=== RUN TestArtifactSave_FlashNamesTheCause/healthy
|
||||
--- FAIL: TestArtifactSave_FlashNamesTheCause (0.24s)
|
||||
--- FAIL: TestArtifactSave_FlashNamesTheCause/version_not_found (0.06s)
|
||||
--- FAIL: TestArtifactSave_FlashNamesTheCause/registry_failing (0.07s)
|
||||
--- FAIL: TestArtifactSave_FlashNamesTheCause/no_sha_listed (0.07s)
|
||||
--- PASS: TestArtifactSave_FlashNamesTheCause/healthy (0.04s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.264s
|
||||
--- fix RESTORED (rc=0):
|
||||
=== RUN TestArtifactSave_FlashNamesTheCause
|
||||
=== RUN TestArtifactSave_FlashNamesTheCause/version_not_found
|
||||
=== RUN TestArtifactSave_FlashNamesTheCause/registry_failing
|
||||
=== RUN TestArtifactSave_FlashNamesTheCause/no_sha_listed
|
||||
=== RUN TestArtifactSave_FlashNamesTheCause/healthy
|
||||
--- PASS: TestArtifactSave_FlashNamesTheCause (0.17s)
|
||||
--- PASS: TestArtifactSave_FlashNamesTheCause/version_not_found (0.04s)
|
||||
--- PASS: TestArtifactSave_FlashNamesTheCause/registry_failing (0.04s)
|
||||
--- PASS: TestArtifactSave_FlashNamesTheCause/no_sha_listed (0.04s)
|
||||
--- PASS: TestArtifactSave_FlashNamesTheCause/healthy (0.04s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.188s
|
||||
|
||||
### R-725 (expired bind page points at the button)
|
||||
$ (cd . && go test ./internal/web -run TestBindExpiredPage_PointsAtTheButton -v -count=1)
|
||||
--- fix UNDONE (rc=1):
|
||||
=== RUN TestBindExpiredPage_PointsAtTheButton
|
||||
r725_bind_expired_copy_test.go:28: the expired page with a button does not carry the sentence that points at it
|
||||
r725_bind_expired_copy_test.go:31: the expired page with a button still sends the household to support
|
||||
--- FAIL: TestBindExpiredPage_PointsAtTheButton (0.04s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.065s
|
||||
FAIL
|
||||
--- fix RESTORED (rc=0):
|
||||
=== RUN TestBindExpiredPage_PointsAtTheButton
|
||||
--- PASS: TestBindExpiredPage_PointsAtTheButton (0.04s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.057s
|
||||
|
||||
### R-725 (console banner glyph; red = the check mark restored in script + golden)
|
||||
$ (cd . && python3 scripts/iso/test/test_console_glyphs.py)
|
||||
--- fix UNDONE (rc=1):
|
||||
FAIL: console glyphs the Latin-2 font cannot draw (R-725):
|
||||
--- fix RESTORED (rc=0):
|
||||
OK: 58 printf literals + 3 goldens use only console-safe glyphs
|
||||
|
||||
### R-728 (in-flight create guard; red = guard removed, seam kept)
|
||||
$ (cd . && go test ./internal/web -run TestConfigCreate_ConcurrentSubmitCreatesOnce -v -count=1)
|
||||
--- fix UNDONE (rc=1):
|
||||
=== RUN TestConfigCreate_ConcurrentSubmitCreatesOnce
|
||||
r728_create_once_test.go:57: customer created 2 times for one press, want exactly 1:
|
||||
[INFO] Customer config created: tester-2
|
||||
[INFO] self-bind link NOT auto-minted for tester-2 on customer creation: no mailer configured on this hub
|
||||
[INFO] Customer config created: tester-2
|
||||
[INFO] self-bind link NOT auto-minted for tester-2 on customer creation: no mailer configured on this hub
|
||||
--- FAIL: TestConfigCreate_ConcurrentSubmitCreatesOnce (0.04s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.064s
|
||||
FAIL
|
||||
--- fix RESTORED (rc=0):
|
||||
=== RUN TestConfigCreate_ConcurrentSubmitCreatesOnce
|
||||
--- PASS: TestConfigCreate_ConcurrentSubmitCreatesOnce (0.05s)
|
||||
PASS
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.065s
|
||||
|
||||
### R-208 hub half (Dockerfile ARG order; red = ARGs back above go mod download)
|
||||
$ (cd . && python3 scripts/test_dockerfile_arg_order.py)
|
||||
--- fix UNDONE (rc=1):
|
||||
FAIL: hub/Dockerfile (R-208):
|
||||
--- fix RESTORED (rc=0):
|
||||
OK: hub/Dockerfile declares its per-build ARGs below the module download
|
||||
|
||||
|
||||
### R-819 (check_stands accepts CLOSED rows; red = rule 3 reads OPEN-ITEMS.md only)
|
||||
$ (cd . && python3 scripts/check_stands.py)
|
||||
--- fix UNDONE (rc=1):
|
||||
CONVICTED — 34 problem(s):
|
||||
--- fix RESTORED (rc=0):
|
||||
check_stands: OK — every claim cites a source, every citation resolves, and every 'walked' cites a walk.
|
||||
|
||||
### R-819 decoy suite (the genuine closed-row stand must pass; red = OPEN-only rule 3)
|
||||
$ (cd . && python3 scripts/test_gate_decoys.py)
|
||||
--- fix UNDONE (rc=1):
|
||||
ok hub-confirm/subdir decoy rejected
|
||||
ok manifest-bearer/subdir decoy rejected
|
||||
ok observations/R-419 decoy rejected
|
||||
ok observations/genuine-FILED genuine accepted
|
||||
ok observations/genuine-NAF genuine accepted
|
||||
ok reuse-refs/missing-go decoy rejected
|
||||
ok reuse-refs/missing-md (KNOWN HOLE R-422) genuine accepted
|
||||
ok golden-currency empty dir rejected AND named
|
||||
ok closed-register/body-word (BY DESIGN) genuine accepted
|
||||
ok closed-register/verdict-word decoy rejected
|
||||
ok closed-register/unreadable-row decoy rejected
|
||||
ok closed-register/duplicate-closed-id decoy rejected
|
||||
ok closed-register/finished-row-in-open decoy rejected
|
||||
ok closed-register/open-row-closed-word (BY DESIGN) genuine accepted
|
||||
ok one-register/suffix-id-row decoy rejected
|
||||
--- fix RESTORED (rc=0):
|
||||
ok hub-confirm/subdir decoy rejected
|
||||
ok manifest-bearer/subdir decoy rejected
|
||||
ok observations/R-419 decoy rejected
|
||||
ok observations/genuine-FILED genuine accepted
|
||||
ok observations/genuine-NAF genuine accepted
|
||||
ok reuse-refs/missing-go decoy rejected
|
||||
ok reuse-refs/missing-md (KNOWN HOLE R-422) genuine accepted
|
||||
ok golden-currency empty dir rejected AND named
|
||||
ok closed-register/body-word (BY DESIGN) genuine accepted
|
||||
ok closed-register/verdict-word decoy rejected
|
||||
ok closed-register/unreadable-row decoy rejected
|
||||
ok closed-register/duplicate-closed-id decoy rejected
|
||||
ok closed-register/finished-row-in-open decoy rejected
|
||||
ok closed-register/open-row-closed-word (BY DESIGN) genuine accepted
|
||||
ok one-register/suffix-id-row decoy rejected
|
||||
|
||||
### R-857 (golden gate reads the re-bake; red = the old regex + version-only sort)
|
||||
$ (cd . && python3 scripts/test_golden_currency_gate.py)
|
||||
--- fix UNDONE (rc=1):
|
||||
CASE 17 ok (R-857): a later-dated bake of the same version wins
|
||||
FAIL: CASE 16 (R-857): two bakes of one version — the gate did not report the RE-BAKE's sha
|
||||
golden currency gate OK — the newest released controller has a golden (NOTE: this checks the BAKE, not the vouch — see the module docstring)
|
||||
--- fix RESTORED (rc=0):
|
||||
CASE 16 ok (R-857): of two bakes of one version, the re-bake's sha is the one reported
|
||||
CASE 17 ok (R-857): a later-dated bake of the same version wins
|
||||
golden-currency gate self-test OK — a directory name alone cannot satisfy it (R-410), and the waiver is judged on its dates, its row and its direction (2026-09-13)
|
||||
|
||||
|
||||
### R-555 (wire-contract strips comments; red = whole-file tokenising as before)
|
||||
$ (cd . && python3 scripts/test_gate_decoys.py > /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/d.out 2>&1; rc=$?; grep -E 'FAIL: wire|ok wire|behaved|exposed a hole' /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/d.out; exit $rc)
|
||||
--- fix UNDONE (rc=1):
|
||||
ok wire-contract/genuine genuine accepted
|
||||
FAIL: wire-contract/comment: rc=0, want 10 (10 = the comment-only tag convicted, 0 = passed)
|
||||
--- fix RESTORED (rc=0):
|
||||
ok wire-contract/comment decoy rejected
|
||||
ok wire-contract/genuine genuine accepted
|
||||
|
||||
|
||||
### R-364 (hu_grep refuses an untested zero; red = the anchor check disabled; no bytecode cache)
|
||||
$ (cd . && PYTHONDONTWRITEBYTECODE=1 python3 -B scripts/test_hu_grep.py)
|
||||
--- fix UNDONE (rc=1):
|
||||
ok present accented string counted rc=0 1 line(s)
|
||||
ok octal-escaped pattern REFUSED rc=2 REFUSED: the pattern arrived TRANSFORMED ('k\\303\\251rj') — octal esc
|
||||
ok U+FFFD pattern REFUSED rc=2 REFUSED: the pattern arrived TRANSFORMED ('k�rj') — octal escapes or U
|
||||
ok no anchor given REFUSED rc=2 REFUSED: an accented pattern needs --anchor with ASCII text known to b
|
||||
ok NFD file, NFC pattern REFUSED rc=2 REFUSED: 0 for the pattern as typed, but 1 line(s) hold its NFD form —
|
||||
ok tested zero is a zero rc=1 0 line(s) — a TESTED zero: anchor 'Felhom' found on 1 line(s), negativ
|
||||
ok ascii pattern plain zero rc=1 0 line(s)
|
||||
FAIL:
|
||||
blind instrument REFUSED: rc=1 msg="0 line(s) — a TESTED zero: anchor 'NotInTheFile' found on 0 line(s), negative control 0, other normal form 0", want rc=2 containing 'anchor'
|
||||
--- fix RESTORED (rc=0):
|
||||
ok present accented string counted rc=0 1 line(s)
|
||||
ok octal-escaped pattern REFUSED rc=2 REFUSED: the pattern arrived TRANSFORMED ('k\\303\\251rj') — octal esc
|
||||
ok U+FFFD pattern REFUSED rc=2 REFUSED: the pattern arrived TRANSFORMED ('k�rj') — octal escapes or U
|
||||
ok blind instrument REFUSED rc=2 REFUSED: the anchor 'NotInTheFile' was not found either — the instrume
|
||||
ok no anchor given REFUSED rc=2 REFUSED: an accented pattern needs --anchor with ASCII text known to b
|
||||
ok NFD file, NFC pattern REFUSED rc=2 REFUSED: 0 for the pattern as typed, but 1 line(s) hold its NFD form —
|
||||
ok tested zero is a zero rc=1 0 line(s) — a TESTED zero: anchor 'Felhom' found on 1 line(s), negativ
|
||||
ok ascii pattern plain zero rc=1 0 line(s)
|
||||
hu_grep: OK — no untested zero for an accented pattern
|
||||
|
||||
### R-587 (release build refuses a stray *.rootpw.txt; red = the guard body emptied; run under BusyBox PATH too)
|
||||
$ (cd . && PATH=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/bb python3 scripts/iso/test/test_rootpw_guard.py)
|
||||
--- fix UNDONE (rc=1):
|
||||
FAIL (R-587):
|
||||
--- fix RESTORED (rc=0):
|
||||
OK: a release build refuses a *.rootpw.txt in its out dir; a clean dir and a non-release build proceed
|
||||
@@ -0,0 +1,5 @@
|
||||
### R-124 red-proof: PBSRootNamespace back to "root"
|
||||
dr_recipe_test.go:478: root namespace on the wire = "root", want "" (PBS's spelling; no namespace is named "root")
|
||||
--- FAIL: TestR124_RootNamespaceOnTheWireIsPBSSpelling (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-agent/internal/hub 0.008s
|
||||
@@ -0,0 +1,23 @@
|
||||
{"id": "R-76", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Image changed since the 1.3.3 finding: felhom-controller@114ff27 controller/internal/infra/infra.go:27 `FileBrowserImage = \"gtstef/filebrowser:1.5.6-stable\"`. The comment infra.go:207-208 still asserts folders come out '2775 with the parent's setgid' -- the exact claim R-76 measured false on 1.3.3; no test pins it (git log --grep 'R-76|setgid' in controller: only 2026-06 commits). Whether 1.5.6 still drops setgid is runtime behaviour of the image.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- find /mnt/felhom-drives -path '*/userdata/*' -mindepth 3 -maxdepth 5 -type d ! -perm -2000 -printf '%m %u:%g %p\\n'\" | head (any folder a customer made in FileBrowser showing 755 without setgid = still true on 1.5.6)", "minutes_spent": 3}
|
||||
{"id": "R-91", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of deletion found (grep srv/pbs-felhom across felhom.eu). Deleting is on ep0 (protected) and needs an operator word.", "dup_of": null, "unique_facts": "Stale doc to fix in the same commit as the delete: documentation/runbooks/offsite-endpoint.md:24 still says datastore `felhom-offsite` is at `/srv/pbs-felhom`; the real path since 2026-07-27 is `/mnt/pbs-datastore` (RUNBOOK-ep0-datastore-volume-2026-07-27.md:8). CONTEXT.md:3666 is the other line to change.", "small_fix": null, "not_worth": null, "settle_cmd": "ssh root@ep0 'du -sh /srv/pbs-felhom 2>&1; proxmox-backup-manager datastore list; df -h /'", "minutes_spent": 4}
|
||||
{"id": "R-209a", "sev": "P4", "category": "Process & tooling", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Pure live state on DooPlex (whether a reboot has happened and the post-boot check passed). No source claim to test.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "uptime -s; cat /var/log/felhom-store-postboot-check.log; ls -d /var/lib/containerd.pre-move-2026-08-05; df -h /", "minutes_spent": 1}
|
||||
{"id": "R-337", "sev": "P4", "category": "Monitoring & notifications", "group": "NOT-WORTH-IT", "evidence": "The row's first question ('establish the intended refresh path') is answered by source: GET /backup/status reads only the agent's in-memory store (felhom-agent@d833163 internal/localapi/server.go:1258 -> pickLatestBackup :1304-1318), and the ONLY writer is the job goroutine after the whole runner returns: server.go:885 `b, err := tier.Service.BackupWithSnapshotHook(...)` then :901 `s.store.RecordBackup(b)` (grep RecordBackup: no other caller). So it is NOT a collection cadence; the status appears when the agent's own job finishes (WaitTask + archive resolve, internal/backup/runner.go:205-232), and a backup the agent did not run itself never appears. The 4-min demo-hp skew is the gap between PBS writing the manifest and the job returning -- unmeasured.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: measure why demo-hp's job returned ~4 min after the manifest landed; cost: a constructed live repro on demo-hp with task-log timing; if never: during an incident the box's backup status can trail the PBS manifest by minutes after an out-of-schedule run, and a run started outside the agent never shows; pick: close with the source fact above written into the row (status = agent job end, not polling), reopen only if a lag is seen on a scheduled run.", "settle_cmd": null, "minutes_spent": 7}
|
||||
{"id": "R-375", "sev": "P4", "category": "Backup & restore", "group": "NOT-WORTH-IT", "evidence": "The signal (audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:168-171) is `pvesm status` showing felhom-pbs Total/Used/Avail = 0. Nothing in the product consumes those numbers for a PBS target: felhom-agent@d833163 internal/backup/runner.go:265 `if st == nil || st.Type == \"pbs\" || st.Avail <= 0 { return true, \"\" }` (space preflight skips PBS and fails open on 0), and the same report says the hub's gauge reads the real 3.7/97.9 GB.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: confirm on ep0 that the namespace-scoped token lacks Datastore.Audit; cost: a read-only ep0 session; if never: the PVE UI on a box shows 0/0/0 for felhom-pbs, which no Felhom code reads (runner.go:265); pick: close as cosmetic with this pointer.", "settle_cmd": "(if ever wanted) ssh root@ep0 'proxmox-backup-manager acl list' ; on a box: pvesm status --storage felhom-pbs", "minutes_spent": 6}
|
||||
{"id": "R-488", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded `interval := 5 * time.Second` and `time.Sleep(3 * time.Second) // initial settling time`, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go:90, restore_unit.go:471. No commit since 2026-09-13 touching internal/backup mentions R-488/test speed. Runtime (333 s) was NOT re-measured here (read-only).", "dup_of": null, "unique_facts": null, "small_fix": "controller: add Manager fields healthSettle/healthInterval (defaults 3s/5s, set in the constructor) used by waitForHealthy; a test helper (the existing Manager fixture constructor) sets them to 0/10ms. Test: TestWaitForHealthy_DefaultsAreProduction asserts a fresh Manager has 3s/5s (pins production), and measure `go test ./internal/backup` wall time before/after in the commit message (expect the 89 >=1 s tests to drop). Run with the R-650 docker-free seams.", "not_worth": null, "settle_cmd": null, "minutes_spent": 5}
|
||||
{"id": "R-504", "sev": "P4", "category": "Install & onboarding", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Live HTTP behaviour of iso.felhom.eu; curl to hosts is outside this checker. Source side: documentation/runbooks/VOLUNTEER-first-hour.md:14 still says the root has no index (R-504); the download page exists at website/letoltes.html.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "If 404 is confirmed it is cosmetic (row's own re-rank); the operator could close it as accepted rather than add a Cloudflare rule.", "settle_cmd": "curl -sI https://iso.felhom.eu/ | head -1", "minutes_spent": 2}
|
||||
{"id": "R-644", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Live scratch-box state. Source context: app-catalog-felhom.eu templates/gokapi/docker-compose.yml:26-29 seeds config.json with an EMPTY Password only when config.json is absent, then runs `--deployment-password`; a config.json that exists with an empty/plain password (e.g. a restored volume or an interrupted first boot) matches the crash text. Not proven.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- docker ps -a --filter name=gokapi --format '{{.Names}} {{.Status}}'; pct exec 9202 -- grep -s -c '\\\"Password\\\":\\\"\\\"' /opt/docker/stacks/gokapi/config/config.json\"", "minutes_spent": 3}
|
||||
{"id": "R-814", "sev": "P4", "category": "Hub & operator", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Hetzner account state; nothing in source records a deletion.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Hetzner Storage Box API (read-only): GET https://api.hetzner.com/v1/storage_boxes/611421 with the operator's API token (stored out-of-band) -> 404 = deleted, else read .storage_box.status", "minutes_spent": 1}
|
||||
{"id": "R-815", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "PBS server-side state on ep0; no GC completion record in the docs (grep).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh root@ep0 'proxmox-backup-manager garbage-collection status felhom-offsite; proxmox-backup-manager task list --all --limit 20 | grep -i garbage'", "minutes_spent": 2}
|
||||
{"id": "R-884", "sev": "P4", "category": "Monitoring & notifications", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml `prom/prometheus:v3.14.0` -> `v3.15.0` (monitoring.yaml:419), and the `monitoring` Application has no `automated` syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an unsynced Renovate bump, i.e. a full sync = Prometheus 3.14 -> 3.15 upgrade (plus pod restart; R-211: no reloader). Not confirmed live.", "dup_of": null, "unique_facts": "Cause candidate: Renovate 53c6e99 prometheus v3.15.0 merged 2026-10-03, never synced because monitoring has no auto-sync in git.", "small_fix": null, "not_worth": null, "settle_cmd": "sudo kubectl -n mon-system get deploy prometheus -o jsonpath='{.spec.template.spec.containers[0].image}' (v3.14.0 => the diff is the Renovate bump)", "minutes_spent": 5}
|
||||
{"id": "R-132", "sev": "P3", "category": "Security & access", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Whether HUB_PW was rotated is not in source. The hub stores a UI-set password in hub_settings with an updated_at column: felhom.eu hub/internal/store/store.go:2182 key `operator_password_hash`, setSetting :2200-2207 writes `updated_at = datetime('now')`. No commit records a rotation (git log --grep rotate/HUB_PW since 2026-07-31).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "On a copy of the hub DB (hub.db + -wal + -shm): sqlite3 hub.db \"SELECT updated_at FROM hub_settings WHERE key='operator_password_hash'\" -- rotated only if later than 2026-09-18 (the last exposure, R-580 folded here); no row = still the ConfigMap seed", "minutes_spent": 4}
|
||||
{"id": "R-298", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Template gate unchanged: felhom-controller@114ff27 controller/internal/web/templates/storage.html:364 `if(d.role==='user-data'){` else :368 protected, no actions. Dependency R-280 is CLOSED (CLOSED-ITEMS.md:600, v0.211.0). Whether the bug bites depends on the role the agent gives the drive: felhom-agent@d833163 internal/storage/role.go:172-186 -- a dir storage (felhom-backup) is user-data unless its backing device is on the system disk; disks.go:1236-1240 deliberately does not reclassify the backup-target drive. role.go unchanged since 2026-08-09. demo-hp was reprovisioned before 2026-09-21, so the 2026-08-10 topology may no longer hold.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp 'grep -A2 felhom-backup /etc/pve/storage.cfg; lsblk -no PKNAME $(findmnt -no SOURCE /) ; lsblk -no PKNAME $(findmnt -no SOURCE /mnt/nvme-1tb)' (same parent disk => role=system => the page still locks it => R-298 true; different => user-data => not reproducible on this box)", "minutes_spent": 8}
|
||||
{"id": "R-338", "sev": "P3", "category": "Security & access", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"grep -E 'listen_addr|island_' /etc/felhom-agent/agent.json; pct config 9201 | grep ^net; ip -br link show master vmbr9\"", "minutes_spent": 6}
|
||||
{"id": "R-350", "sev": "P3", "category": "Security & access", "group": "DUPLICATE", "evidence": "Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence.", "dup_of": "R-132", "unique_facts": "(1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked.", "small_fix": null, "not_worth": null, "settle_cmd": null, "minutes_spent": 3}
|
||||
{"id": "R-542", "sev": "P3", "category": "Storage & devices", "group": "NOT-WORTH-IT", "evidence": "Still true in source, and by design: felhom-agent@d833163 internal/localapi/disks.go:438 `initialize = append(initialize, c) // every unclaimed disk can be initialized`; a disk mounted under /mnt/felhom-drives counts as UNCLAIMED on purpose (internal/storage/claim.go:88-103, R-220, so drives survive a guest rebuild), and the mkfs wrapper explicitly allows it: configs/felhom-mkfs-guarded.sh:59-60 'Mounts under /mnt/felhom-drives are our own drives (the agent detaches before a re-init) -> allowed'. Controller passes initialize through untouched (agent_disk_handlers.go:158-162). `already_mounted` is only set for controller-contributed stores (agentapi/client.go:447-451, omitempty) -- so 'null' is expected for agent candidates.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: drop felhom-mounted disks from `initialize`; cost: reverses the R-220/re-init design (wrapper comment) and needs a decision on how a registered drive is re-initialized; if never: the raw endpoint lists a registered drive under initialize while the page (customer view) filters it -- only a session reading the raw endpoint is misled; pick: close as by-design, add one comment line at disks.go:438 saying a registered felhom drive is listed here deliberately.", "settle_cmd": null, "minutes_spent": 9}
|
||||
{"id": "R-607", "sev": "P3", "category": "App updates", "group": "STILL-TRUE-SMALL", "evidence": "Diagnosed from source (both questions the row asks). felhom-controller@114ff27 controller/internal/sync/sync.go:236-241 rescans ONLY `if len(newApps) > 0 || len(updated) > 0`, and :257-258 says 'nincs változás' when both are empty. `updated` counts stack-dir copies only (copyTemplates, :447 `updated = append(updated, appName)` after a hash mismatch); for a deployed+pinned app whose catalog moved, renderSource (:457-469 table) copies the STORED definition, so the hash matches and nothing is 'updated' although the git cache moved. CatalogImages is read from the catalog cache only inside ScanStacks (stacks/manager.go:666-672, assigned :690). So (a) the message measures the stack dir, not the catalog; (b) CatalogImages refreshes only on a ScanStacks, which this sync does not trigger. The 29 s nextcloud case (needed several rounds) is not explained by this.", "dup_of": null, "unique_facts": null, "small_fix": "controller sync.go: record the catalog git HEAD before and after gitCloneOrPull; if it moved, call s.rescanFn() even when newApps/updated are empty, and say 'Katalógus frissítve — az alkalmazások nem változtak' (and EN) instead of 'nincs változás'. Test (red first): a Syncer with a fake pull that moves HEAD and a frozen pinned app (renderSource returns the stored definition) asserts rescanFn was called once and the message is not the no-change one; and a no-move pull asserts rescanFn NOT called.", "not_worth": null, "settle_cmd": null, "minutes_spent": 9}
|
||||
{"id": "R-683", "sev": "P3", "category": "App updates", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Evidence in repo: audits/night-2026-09-24/E/round-03-controller.log is the POST-cut log only (first lines 12:05:09Z: update.go:1335 'interrupted in verifying (started 2026-09-24T12:04:21Z)', :589 undo copies '.pre-update-20260924T120426Z'); the pre-cut log that would show a backing-up phase is lost (row says so). No later power-cut-mid-update drill: night-2026-10-04/MORNING-NOTE.md:64 'A1 power cut mid-update (demo-hp): NOT RUN'; DRILL-night-2026-09-25.md:102 was a cut during romm's verifying, not checked for the backup choice.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Not a read-only command: the next power-cut-during-update drill with the controller log saved at arm time, then grep \"phase backing-up\\|Tier-2\" in the saved pre-cut log", "minutes_spent": 7}
|
||||
{"id": "R-756", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 `if !m.DriveLive(hddPath)` -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 `return m.isMountPoint(hddPath)` -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH is compared to a registered storage path (api/router.go:574-575, web/handlers.go:3049, storage_handlers.go:528), i.e. a drive root. The message names `.../scratch_hdd/userdata/calibre-web`, so on 9202 HDD_PATH is a per-app subfolder, which can never be a mount point -> the 409 is certain for that value. Open: whether that HDD_PATH was written by the product (handlers.go:550 prefill from place.Drive) or by the test venue.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- grep -H HDD_PATH /opt/docker/stacks/calibre-web/app.yaml /opt/docker/stacks/grimmory/app.yaml; pct exec 9202 -- findmnt -no TARGET,SOURCE /mnt/felhom-drives/scratch_hdd\" (HDD_PATH = per-app subfolder => product/venue wrote a non-root HDD_PATH; HDD_PATH = drive root and not mounted => scratch drive not registered)", "minutes_spent": 10}
|
||||
{"id": "R-862", "sev": "P3", "category": "Box system & updates", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Waits on the operator's by-hand bootstrap on Tester 2; no commit records it (felhom.eu log since 2026-10-04). runbooks/config-bundle.md:77 'CC sends the bundle by the signed job and reads it back on the System page'; a box behind the vouched bundle for 7 days raises os_config_bundle_behind (:83).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Ask the operator whether the three bootstrap commands ran; read-only proof: the hub's operator view of Tester 2 (OS/config-bundle sha vs the vouched sha), or whether os_config_bundle_behind has fired for it", "minutes_spent": 2}
|
||||
{"id": "R-882", "sev": "P3", "category": "Hub & operator", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Longhorn instance-manager runtime state on DooPlex; nothing in homelab-manifests addresses it (no commit since 87dfc29 names it).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "sudo kubectl -n longhorn-system get pods -l longhorn.io/component=instance-manager -o custom-columns=NAME:.metadata.name,START:.status.startTime ; systemctl show k3s containerd iscsid -p ActiveEnterTimestamp (instance-manager older than a k3s/containerd/iscsid restart => the stale-PID state can recur)", "minutes_spent": 2}
|
||||
{"id": "R-883", "sev": "P3", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "homelab-manifests@87dfc29 still has moving tags (grep image lines without a numeric tag): admin-system/toolbox.yaml:12 nicolaka/netshoot:latest (a bare Pod); calibre-system/cwa.yaml:826 calibre-web-automated:dev; outline-system/outline.yaml:270 minio/minio:latest; tandoor-system/recipe-importer.yaml:26 gitea.dooplex.hu/admin/recipe-importer:latest (pull Always :27); adventurelog-system/adventurelog.yaml:100 and :256 adventurelog-backend/frontend:latest (pull Always :101,:257); jarrs-system/jarr-dev.yaml:311,:345,:630 gitea.dooplex.hu/admin/jarr:latest (pull Always). Zipline fixed (90f60e4, 4c8ec7a). The repo's own rule homelab-manifests/CLAUDE.md:120 'Image tags always pinned'. Helm values files (external-dns, pihole, plex, authentik, cnpg) not checked for tag fields beyond a `tag: latest|dev|empty` grep (no hits).", "dup_of": null, "unique_facts": null, "small_fix": "homelab-manifests only: for each line above, read the running digest/version (`sudo kubectl get pod -n <ns> -o jsonpath='{..imageID}'`), pin that exact tag (or @sha256 for the self-built gitea.dooplex.hu jarr/recipe-importer images, which have no version tags), add the version-checker match-regex annotation per CLAUDE.md:120, and set imagePullPolicy IfNotPresent. Test: `grep -rnE 'image:.*(:latest|:dev)\\s*$' --include=*.yaml .` returns nothing; after ArgoCD sync each pod's imageID equals the pre-change one (no upgrade). Operator-owned DooPlex change: needs the operator's word.", "not_worth": null, "settle_cmd": null, "minutes_spent": 6}
|
||||
{"id": "R-886", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-SMALL", "evidence": "homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root `nobody` image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts silences now survive a restart -- an invariant with no test, which is exactly what this row says broke. Last structural change 58d1cd2 (2026-08-14, 'give alertmanager real storage').", "dup_of": null, "unique_facts": null, "small_fix": "homelab-manifests: add pod `securityContext: {fsGroup: 65534, fsGroupChangePolicy: OnRootMismatch}` (65534 = nobody, the image's user -- confirm with `sudo kubectl -n mon-system exec deploy/alertmanager -- id`) to the alertmanager Deployment. Test (consequence, per CLAUDE.md): create a silence via amtool/API, delete the pod, after it returns the silence is still listed AND the log has no 'Running maintenance failed ... permission denied' within 15 min (positive observable: a 'maintenance done' line).", "not_worth": null, "settle_cmd": null, "minutes_spent": 5}
|
||||
@@ -0,0 +1,9 @@
|
||||
08:27:24
|
||||
('Tester-2-be8404', '0.142.0', '2026-10-04 18:05:48', None)
|
||||
('demo-felhom-8363b5', '0.145.0', '2026-10-05 08:23:56', '0.145.0')
|
||||
('demo-hp-bb76ea', '0.145.0', '2026-10-05 08:27:06', '0.145.0')
|
||||
('tester-1-d70be4', '0.145.0', '2026-10-05 08:27:31', '0.145.0')
|
||||
== demo-hp: agent felhom-agent 0.145.0 | controller 0.295.0 | healthy 20 of 21 | guard 1
|
||||
== felhom-pve: agent felhom-agent 0.145.0 | controller 0.295.0 | healthy 4 of 5 | guard 1
|
||||
Oct 05 10:08:58 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:58.009+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUN
|
||||
9202: gitea.dooplex.hu/admin/felhom-controller:0.295.0 Up 29 minutes (healthy)
|
||||
@@ -0,0 +1,12 @@
|
||||
### 06:58:03 UTC: set 9202's window to 09:05 (Budapest) by POST /backups/window
|
||||
HTTP 303
|
||||
2026/10/05 06:58:05 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 09:05 (next run 2026-10-05 09:05 CEST)
|
||||
2026/10/05 06:58:05 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 10:05 (next run 2026-10-05 10:05 CEST)
|
||||
2026/10/05 06:58:05 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 10:50 (next run 2026-10-05 10:50 CEST)
|
||||
2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job db-dump: rescheduled — recomputing next run
|
||||
2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job tier2-backup: rescheduled — recomputing next run
|
||||
2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job db-dump: next run at 2026-10-05 09:05:00 CEST (waiting 6m54s)
|
||||
2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job tier2-backup: next run at 2026-10-05 10:05:00 CEST (waiting 1h6m54s)
|
||||
2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job offbox-backup: rescheduled — recomputing next run
|
||||
2026/10/05 06:58:05 scheduler.go:67: [DEBUG] [scheduler] daily job offbox-backup: next run at 2026-10-05 10:50:00 CEST (waiting 1h51m54s)
|
||||
2026/10/05 06:58:05 backup_handlers.go:65: [INFO] [web] backup window set to 09:05 (legs 09:05/10:05/10:50)
|
||||
@@ -0,0 +1,9 @@
|
||||
### 07:03:01 pct shutdown 9202
|
||||
status: stopped
|
||||
### 07:08:01 pct start 9202
|
||||
status: running
|
||||
### 07:10:35 controller log after the start
|
||||
2026/10/05 07:08:08 scheduler.go:132: [INFO] [scheduler] Daily job db-dump scheduled for 2026-10-06 09:05 CEST
|
||||
2026/10/05 07:08:08 scheduler.go:132: [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-10-05 10:05 CEST
|
||||
2026/10/05 07:08:08 scheduler.go:132: [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-10-05 10:50 CEST
|
||||
2026/10/05 07:08:08 backup.go:1166: [INFO] [backup] Found 2 DB dump files across drives
|
||||
@@ -0,0 +1 @@
|
||||
('demo-felhom', 'backup_catchup_done', 'info', 'Kimaradt mentés pótolva: a doboz ki volt kapcsolva 10:07-kor, a mentés most elkészült.', 'controller', '2026-10-05 08:25:03')
|
||||
@@ -0,0 +1,12 @@
|
||||
### 07:29:47 9202: controller 0.295.0 by its bootstrap file
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.295.0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.295.0 started 2026-10-05T07:29:50.143207934Z
|
||||
2026/10/05 07:29:50 scheduler.go:149: [INFO] [scheduler] Daily job db-dump scheduled for 2026-10-06 09:05 CEST
|
||||
2026/10/05 07:29:50 nightchain.go:280: [INFO] [catch-up] controller start: no backup leg missed its last scheduled time — nothing to make up
|
||||
{
|
||||
"seeded_at": "2026-10-05T07:29:50.484758959Z",
|
||||
"ended": {},
|
||||
"db_dump_ok": "0001-01-01T00:00:00Z",
|
||||
"last_catch_up": "0001-01-01T00:00:00Z",
|
||||
"banner_dismissed_through": "0001-01-01T00:00:00Z"
|
||||
}
|
||||
@@ -0,0 +1,41 @@
|
||||
### 07:33:00 W = 09:35 Budapest (07:35 UTC). pct shutdown 9202
|
||||
command 'lxc-stop -n 9202 --nokill --timeout 60' failed: exit code 1
|
||||
status: stopped
|
||||
### 07:38:00 pct start 9202
|
||||
status: running
|
||||
### 07:39:06 after the start
|
||||
2026/10/05 07:38:09 scheduler.go:149: [INFO] [scheduler] Daily job db-dump scheduled for 2026-10-06 09:35 CEST
|
||||
2026/10/05 07:38:09 nightchain.go:291: [INFO] [catch-up] controller start: the box missed [db-dump] (last scheduled 2026-10-05 09:35) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night)
|
||||
### 07:53:51 the catch-up
|
||||
2026/10/05 07:43:09 scheduler.go:381: [INFO] [scheduler] Running job: system-health
|
||||
2026/10/05 07:43:09 backup.go:1166: [INFO] [backup] Found 2 DB dump files across drives
|
||||
2026/10/05 07:44:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan
|
||||
2026/10/05 07:46:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan
|
||||
2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: docker-socket-users
|
||||
2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: offsite-credential-retry
|
||||
2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan
|
||||
2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: backup-cache
|
||||
2026/10/05 07:48:09 scheduler.go:381: [INFO] [scheduler] Running job: system-health
|
||||
2026/10/05 07:48:09 backup.go:1166: [INFO] [backup] Found 2 DB dump files across drives
|
||||
2026/10/05 07:50:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan
|
||||
2026/10/05 07:52:09 scheduler.go:381: [INFO] [scheduler] Running job: stack-scan
|
||||
2026/10/05 07:53:09 nightchain.go:332: [INFO] [catch-up] running the missed db-dump leg
|
||||
2026/10/05 07:53:09 scheduler.go:381: [INFO] [scheduler] Running job: offsite-credential-retry
|
||||
2026/10/05 07:53:09 scheduler.go:381: [INFO] [scheduler] Running job: docker-socket-users
|
||||
2026/10/05 07:53:09 scheduler.go:381: [INFO] [scheduler] Running job: backup-cache
|
||||
2026/10/05 07:53:09 scheduler.go:381: [INFO] [scheduler] Running job: system-health
|
||||
2026/10/05 07:53:09 backup.go:1166: [INFO] [backup] Found 2 DB dump files across drives
|
||||
2026/10/05 07:53:09 dbdump.go:410: [INFO] [backup] DB dump: paperless-postgres → paperless-ngx-postgres.sql (412.6 KB, 436ms, 72 tables)
|
||||
2026/10/05 07:53:30 nightchain.go:340: [INFO] [catch-up] done: db-dump in 21s (missed at 2026-10-05 09:35)
|
||||
{
|
||||
"seeded_at": "2026-10-05T07:29:50.484758959Z",
|
||||
"ended": {
|
||||
"db-dump": "2026-10-05T07:53:30.483916689Z"
|
||||
},
|
||||
"db_dump_ok": "2026-10-05T07:53:30.484211296Z",
|
||||
"last_catch_up": "2026-10-05T07:53:30.484384683Z",
|
||||
"banner_dismissed_through": "0001-01-01T00:00:00Z"
|
||||
}
|
||||
2026/10/05 07:38:09 nightchain.go:291: [INFO] [catch-up] controller start: the box missed [db-dump] (last scheduled 2026-10-05 09:35) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night)
|
||||
2026/10/05 07:53:09 nightchain.go:332: [INFO] [catch-up] running the missed db-dump leg
|
||||
2026/10/05 07:53:30 nightchain.go:340: [INFO] [catch-up] done: db-dump in 21s (missed at 2026-10-05 09:35)
|
||||
+23
@@ -0,0 +1,23 @@
|
||||
2026/10/05 07:58:35 main.go:340: [INFO] felhom-controller 0.295.0 starting (customer: demo-hp, domain: enkisfelhom.hu)
|
||||
2026/10/05 07:58:35 scheduler.go:149: [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-10-06 03:30 CEST
|
||||
2026/10/05 07:58:35 nightchain.go:291: [INFO] [catch-up] controller start: the box missed [db-dump tier2 offsite] (last scheduled 2026-10-05 04:15) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night)
|
||||
2026/10/05 08:13:35 nightchain.go:332: [INFO] [catch-up] running the missed db-dump leg
|
||||
2026/10/05 08:13:36 dbdump.go:410: [INFO] [backup] DB dump: paperless-postgres → paperless-ngx-postgres.sql (413.8 KB, 423ms, 72 tables)
|
||||
2026/10/05 08:13:57 nightchain.go:332: [INFO] [catch-up] running the missed tier2 leg
|
||||
2026/10/05 08:13:57 tier2.go:425: [INFO] [backup] Tier 2 copied paperless-ngx → /mnt/sys_drive/felhom-data/backups/secondary/paperless-ngx (82.5 MB, 1 leg(s), 0s) [SSD: state-only]
|
||||
2026/10/05 08:13:57 tier2.go:425: [INFO] [backup] Tier 2 copied privatebin → /mnt/felhom-drives/scratch_hdd/backups/secondary/privatebin (20.3 KB, 0 leg(s), 0s)
|
||||
2026/10/05 08:13:57 tier2.go:476: [INFO] [backup] Tier 2 run complete: 2 app(s) processed (incl. volume-only — F6)
|
||||
2026/10/05 08:13:57 nightchain.go:332: [INFO] [catch-up] running the missed offsite leg
|
||||
2026/10/05 08:13:57 main.go:1385: [INFO] [offbox] no scheduled off-site target on this box — the off-site leg does nothing
|
||||
2026/10/05 08:13:57 nightchain.go:340: [INFO] [catch-up] done: db-dump, offsite, tier2 in 22s (missed at 2026-10-05 04:15)
|
||||
{
|
||||
"seeded_at": "2026-09-30T07:54:19.875183Z",
|
||||
"ended": {
|
||||
"db-dump": "2026-10-05T08:13:57.406392075Z",
|
||||
"offsite": "2026-10-05T08:13:57.69370068Z",
|
||||
"tier2": "2026-10-05T08:13:57.693478189Z"
|
||||
},
|
||||
"db_dump_ok": "2026-10-05T08:13:57.406683255Z",
|
||||
"last_catch_up": "2026-10-05T08:13:57.693858077Z",
|
||||
"banner_dismissed_through": "2026-10-05T00:30:00Z"
|
||||
}
|
||||
@@ -0,0 +1,24 @@
|
||||
### 08:05:01 W = 10:07 Budapest (08:07 UTC). park + stop the controller (apps keep running)
|
||||
felhom-controller
|
||||
4
|
||||
### 08:10:00 start the controller + unpark
|
||||
felhom-controller
|
||||
2026/10/05 08:10:02 [INFO] [scheduler] Daily job db-dump scheduled for 2026-10-06 10:07 CEST
|
||||
### 08:25:04 the catch-up
|
||||
{
|
||||
"seeded_at": "2026-10-05T07:29:41.65020944Z",
|
||||
"ended": {
|
||||
"db-dump": "2026-10-05T08:25:03.369812125Z"
|
||||
},
|
||||
"db_dump_ok": "2026-10-05T08:25:03.370056355Z",
|
||||
"last_catch_up": "2026-10-05T08:25:03.370196251Z",
|
||||
"banner_dismissed_through": "0001-01-01T00:00:00Z"
|
||||
}### 08:25:19 the catch-up lines (this box logs without file:line)
|
||||
2026/10/05 08:10:02 [INFO] [catch-up] controller start: the box missed [db-dump] (last scheduled 2026-10-05 10:07) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night)
|
||||
2026/10/05 08:25:02 [INFO] [catch-up] running the missed db-dump leg
|
||||
2026/10/05 08:25:03 [INFO] [catch-up] done: db-dump in 1s (missed at 2026-10-05 10:07)
|
||||
### 08:25:20 window back to 02:30
|
||||
HTTP 303
|
||||
2026/10/05 08:25:22 [INFO] [web] backup window set to 02:30 (legs 02:30/03:30/04:15)
|
||||
ls: cannot access '/var/lib/felhom-agent/guests/9201/controller-parked': No such file or directory
|
||||
4
|
||||
@@ -0,0 +1,120 @@
|
||||
### RED-PROOF scheduler late-fire guard dropped
|
||||
scheduler_test.go:128: a 6-hour-late fire ran the job (the app-update leg would start at noon)
|
||||
--- FAIL: TestDaily_LateFireIsSkipped (0.00s)
|
||||
--- PASS: TestDaily_LateFireIsSkipped/on_time (0.00s)
|
||||
--- FAIL: TestDaily_LateFireIsSkipped/after_a_suspend (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/scheduler 0.004s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/scheduler 0.456s
|
||||
### RED-PROOF 1 Missed returns nothing (the old behaviour: no catch-up)
|
||||
nightchain_test.go:99: no catch-up was scheduled for a box off across W (missed [])
|
||||
--- FAIL: TestCatchUp_OffAtWThenOn_OneCatchUpAfter15Min (0.00s)
|
||||
nightchain_test.go:135: the catch-up never finished
|
||||
--- FAIL: TestCatchUp_TwoMissedNights_One (3.00s)
|
||||
nightchain_test.go:151: the catch-up never finished
|
||||
--- FAIL: TestCatchUp_PowerCutMidChain_FinishesTheRest (3.00s)
|
||||
nightchain_test.go:178: missed [], want tier2 + offsite of the night of the 4th
|
||||
--- FAIL: TestCatchUp_LegAboutToRun_LeftToNormal (0.00s)
|
||||
nightchain_test.go:189: the catch-up never finished
|
||||
--- FAIL: TestCatchUp_NeverRunsAnythingButBackupLegs (3.00s)
|
||||
nightchain_test.go:206: the catch-up never finished
|
||||
--- FAIL: TestCatchUp_WaitsForAWholeGuestBackup (3.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 12.015s
|
||||
FAIL
|
||||
### RED-PROOF 2 a leg that ran is not checked (Ended ignored)
|
||||
nightchain_test.go:121: a normal night was made up again: [db-dump tier2 offsite]
|
||||
--- FAIL: TestCatchUp_NightRan_DaytimeRestart_None (0.00s)
|
||||
nightchain_test.go:140: after the catch-up nothing is missed any more, got [db-dump tier2 offsite]
|
||||
--- FAIL: TestCatchUp_TwoMissedNights_One (0.00s)
|
||||
nightchain_test.go:153: ran [db-dump tier2 offsite], want only tier2 + offsite
|
||||
--- FAIL: TestCatchUp_PowerCutMidChain_FinishesTheRest (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
FAIL
|
||||
### RED-PROOF 3 the 15-minute delay dropped
|
||||
nightchain_test.go:103: the catch-up must wait 15m0s first, slept []
|
||||
--- FAIL: TestCatchUp_OffAtWThenOn_OneCatchUpAfter15Min (0.00s)
|
||||
nightchain_test.go:208: slept [1m0s 1m0s 1m0s] — want the 15 min delay then 3 one-minute waits for the quiesce
|
||||
--- FAIL: TestCatchUp_WaitsForAWholeGuestBackup (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.008s
|
||||
FAIL
|
||||
### RED-PROOF 4 a pending catch-up does not block a second one
|
||||
nightchain_test.go:133: a second catch-up was scheduled while one was pending: [db-dump tier2 offsite]
|
||||
--- FAIL: TestCatchUp_TwoMissedNights_One (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
FAIL
|
||||
### RED-PROOF 5 the quiesce wait dropped
|
||||
nightchain_test.go:208: slept [15m0s] — want the 15 min delay then 3 one-minute waits for the quiesce
|
||||
--- FAIL: TestCatchUp_WaitsForAWholeGuestBackup (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
FAIL
|
||||
### RED-PROOF 6 the leave-to-normal rule dropped
|
||||
nightchain_test.go:174: the dump due in 20 minutes was put in the catch-up: [db-dump tier2 offsite]
|
||||
--- FAIL: TestCatchUp_LegAboutToRun_LeftToNormal (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
FAIL
|
||||
### RED-PROOF 8 a fresh ledger counts history as missed
|
||||
nightchain_test.go:161: a freshly seeded ledger made up [db-dump tier2 offsite]
|
||||
--- FAIL: TestCatchUp_FreshLedger_None (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.007s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.008s
|
||||
### RED-PROOF 7 the resume watch compares wall with wall (compiles)
|
||||
nightchain_test.go:241: a 7-hour suspend was not seen
|
||||
--- FAIL: TestResumeWatch_SeesASuspend (2.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 2.010s
|
||||
FAIL
|
||||
### RED-PROOF 9 the catch-up runs whatever Legs holds (not only the missed chain legs)
|
||||
nightchain_test.go:153: ran [db-dump tier2 offsite], want only tier2 + offsite
|
||||
--- FAIL: TestCatchUp_PowerCutMidChain_FinishesTheRest (0.00s)
|
||||
nightchain_test.go:191: an app update ran inside a catch-up
|
||||
--- FAIL: TestCatchUp_NeverRunsAnythingButBackupLegs (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.008s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
### RED-PROOF 10 the whole-guest cycle does not wait for a running catch-up
|
||||
r871_catchup_test.go:21: the backup started while a catch-up ran: start=1 stopped=[nextcloud]
|
||||
--- FAIL: TestR871_ScheduledCycleWaitsForCatchUp (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.005s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.005s
|
||||
### RED-PROOF 12 the update leg inside the shared off-site body
|
||||
--- PASS: TestR871_CatchUpWiring (0.01s)
|
||||
r871_catchup_wiring_test.go:44: the update leg is inside the off-site body the catch-up runs
|
||||
--- FAIL: TestR871_CatchUpRunsNoUpdateLeg (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.019s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.019s
|
||||
### RED-PROOF 11 the start trigger not wired (test now keyed on the argument)
|
||||
r871_catchup_wiring_test.go:37: the catch-up's START trigger (Evaluate(ctx, "controller start")) is called 0 times, want 1
|
||||
--- FAIL: TestR871_CatchUpWiring (0.02s)
|
||||
--- PASS: TestR871_CatchUpRunsNoUpdateLeg (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.031s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.026s
|
||||
### RED-PROOF hub: backup_catchup_done not allowed
|
||||
r871_catchup_event_test.go:10: backup_catchup_done is not an allowed event type — the box's catch-up line would be dropped with a 400
|
||||
--- FAIL: TestR871_CatchUpEventIsAllowed (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/api 0.020s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/api 0.019s
|
||||
@@ -0,0 +1,27 @@
|
||||
### 07:54:16 banner on 9202 — the ledger set by hand to a box whose last dump is 3 days old (scratch fixture, said so); window back to 02:30
|
||||
HTTP 303
|
||||
ledger set: last dump 2026-10-02T07:54:19.875183Z
|
||||
2026/10/05 07:54:20 nightchain.go:291: [INFO] [catch-up] controller start: the box missed [db-dump tier2 offsite] (last scheduled 2026-10-05 04:15) — ONE catch-up in 15m0s (backup legs only; app updates wait for a real night)
|
||||
--- the banner as GET /launcher served it (9202, Hungarian household):
|
||||
<div class="alerts-container">
|
||||
<div class="alert-banner alert-banner-warning" id="missed-backup-banner">
|
||||
<span class="alert-icon"><svg class="ico"><use href="#i-triangle-alert"/></svg></span>
|
||||
<span class="alert-message">A legutóbbi mentés 3 nappal ezelőtt készült. Válassz egy olyan időpontot, amikor a doboz általában be van kapcsolva.</span>
|
||||
<span class="alert-actions">
|
||||
<a href="/backups#window_start" class="btn btn-sm btn-primary">Mentési idő módosítása</a>
|
||||
<form method="POST" action="/backups/missed-banner/dismiss" style="display:inline">
|
||||
<input type="hidden" name="_csrf" value="(redacted)"><input type="hidden" name="back" value="/launcher">
|
||||
<button type="submit" class="btn btn-sm btn-outline">Bezárás</button>
|
||||
</form>
|
||||
</span>
|
||||
</div>
|
||||
</div
|
||||
### 07:54:54 POST /backups/missed-banner/dismiss (the Bezárás button)
|
||||
HTTP 302
|
||||
banner on /launcher after closing: 0
|
||||
2026/10/05 07:54:46 missed_backup_banner.go:90: [DEBUG] [web] missed-backup banner: rendered on /launcher (last=2026-10-02T07:54:19Z off_at="" suggest="")
|
||||
2026/10/05 07:54:46 missed_backup_banner.go:90: [DEBUG] [web] missed-backup banner: rendered on /launcher (last=2026-10-02T07:54:19Z off_at="" suggest="")
|
||||
2026/10/05 07:54:57 missed_backup_banner.go:90: [DEBUG] [web] missed-backup banner: rendered on /launcher (last=2026-10-02T07:54:19Z off_at="" suggest="")
|
||||
2026/10/05 07:54:57 missed_backup_banner.go:90: [DEBUG] [web] missed-backup banner: rendered on /launcher (last=2026-10-02T07:54:19Z off_at="" suggest="")
|
||||
2026/10/05 07:54:57 missed_backup_banner.go:100: [INFO] [web] missed-backup banner: closed by the household until the next missed backup time (after 2026-10-05T02:30:00+02:00)
|
||||
"banner_dismissed_through": "2026-10-05T00:30:00Z"
|
||||
@@ -0,0 +1,56 @@
|
||||
### RED-PROOF 1 stale is silent (the old behaviour)
|
||||
banner_test.go:32: not shown although the last backup is 2 days old
|
||||
--- FAIL: TestBanner_ShownWithReasonAndSuggestion (0.00s)
|
||||
banner_test.go:72: the banner did not come back after the next missed backup time
|
||||
--- FAIL: TestBanner_DismissedThenBackAfterANewMiss (0.00s)
|
||||
banner_test.go:92: banner = {Show:false LastBackup:2026-10-02 04:20:00 +0200 CEST DaysAgo:3 OffAt:02:30 Suggest:21:00 MissedAt:2026-10-05 02:30:00 +0200 CEST}, want shown with the off-site copy's
|
||||
--- FAIL: TestBanner_OffsiteTierCounts (0.00s)
|
||||
banner_test.go:107: banner = {Show:false LastBackup:2026-10-03 02:31:00 +0200 CEST DaysAgo:2 OffAt: Suggest: MissedAt:2026-10-05 02:30:00 +0200 CEST}, want shown, no suggestion, no 'off at'
|
||||
--- FAIL: TestBanner_NoPatternNoSuggestion (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.005s
|
||||
FAIL
|
||||
### RED-PROOF 2 dismissal ignored
|
||||
banner_test.go:64: a closed banner came back without a new miss
|
||||
--- FAIL: TestBanner_DismissedThenBackAfterANewMiss (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.005s
|
||||
FAIL
|
||||
### RED-PROOF 4 the seed ignored (banner on upgrade day)
|
||||
banner_test.go:53: shown on the day of the upgrade: {Show:true LastBackup:0001-01-01 00:00:00 +0000 UTC DaysAgo:0 OffAt:02:30 Suggest:21:00 MissedAt:2026-10-05 02:30:00 +0200 CEST}
|
||||
--- FAIL: TestBanner_NotShownBeforeTheRecordKnows (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.005s
|
||||
FAIL
|
||||
### RED-PROOF 5 the off-site tier ignored
|
||||
banner_test.go:92: banner = {Show:false LastBackup:0001-01-01 00:00:00 +0000 UTC DaysAgo:0 OffAt: Suggest: MissedAt:0001-01-01 00:00:00 +0000 UTC}, want shown with the off-site copy's 3 days
|
||||
--- FAIL: TestBanner_OffsiteTierCounts (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.004s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
### RED-PROOF 3 a suggestion from noise (no day count) — test noise now evenings on 2 of 7 days
|
||||
banner_test.go:107: banner = {Show:true LastBackup:2026-10-03 02:31:00 +0200 CEST DaysAgo:2 OffAt: Suggest:21:00 MissedAt:2026-10-05 02:30:00 +0200 CEST}, want shown, no suggestion, no 'off at'
|
||||
--- FAIL: TestBanner_NoPatternNoSuggestion (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.005s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s
|
||||
### RED-PROOF 6 the banner not wired into executeTemplate
|
||||
r871_missed_backup_banner_test.go:37: not shown on /launcher although the last backup is 3 days old
|
||||
--- FAIL: TestR871_BannerShownThenClosedUntilTheNextMiss (0.06s)
|
||||
--- PASS: TestR871_BannerNotShownAfterASuccessfulNight (0.06s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.130s
|
||||
FAIL
|
||||
### RED-PROOF 7 the dismissal not recorded
|
||||
r871_missed_backup_banner_test.go:47: still shown after the household closed it
|
||||
--- FAIL: TestR871_BannerShownThenClosedUntilTheNextMiss (0.06s)
|
||||
--- PASS: TestR871_BannerNotShownAfterASuccessfulNight (0.07s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.145s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.148s
|
||||
@@ -0,0 +1,27 @@
|
||||
### RED-PROOF R-872a: the old skip of a down box
|
||||
r872_down_box_test.go:60: events map[] — a box off at every deadline must raise the missed-backup alarms, down or not
|
||||
--- FAIL: TestR872_DownEveryNightRaisesTheMissedAlarms (0.04s)
|
||||
--- PASS: TestR872_DownWithARecentDumpIsQuiet (0.04s)
|
||||
--- PASS: TestR872_NewDownBoxIsQuiet (0.03s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/monitor 0.118s
|
||||
FAIL
|
||||
### RED-PROOF R-872b: no grace for a new box
|
||||
--- PASS: TestR872_DownEveryNightRaisesTheMissedAlarms (0.04s)
|
||||
--- PASS: TestR872_DownWithARecentDumpIsQuiet (0.04s)
|
||||
r872_down_box_test.go:84: events map[expected_backup_missed:1 expected_dbdump_missed:1] for a box bound a day ago
|
||||
--- FAIL: TestR872_NewDownBoxIsQuiet (0.03s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/monitor 0.113s
|
||||
FAIL
|
||||
### RED-PROOF R-872c: the down dump line shortened to 12 h
|
||||
--- PASS: TestR872_DownEveryNightRaisesTheMissedAlarms (0.04s)
|
||||
r872_down_box_test.go:76: events map[expected_dbdump_missed:1] — a dump 20 h ago and a backup 30 h ago are inside the down box's lines
|
||||
--- FAIL: TestR872_DownWithARecentDumpIsQuiet (0.04s)
|
||||
r872_down_box_test.go:84: events map[expected_backup_missed:1 expected_dbdump_missed:1] for a box bound a day ago
|
||||
--- FAIL: TestR872_NewDownBoxIsQuiet (0.03s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/monitor 0.113s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/monitor 38.508s
|
||||
@@ -0,0 +1,8 @@
|
||||
### RED-PROOF R-873: the weekly household limit dropped
|
||||
r873_liveness_weekly_test.go:40: night 2: household 1 mails (want 0 — told once this week), operator 2 (want every edge)
|
||||
--- FAIL: TestR873_HouseholdHearsAnOutageAtMostWeekly (0.04s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/notify 0.043s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/notify 1.382s
|
||||
@@ -0,0 +1,11 @@
|
||||
Oct 05 09:38:46 demo-felhom felhom-agent[3840353]: time=2026-10-05T09:38:46.189+02:00 level=INFO msg="backup: restore-test scheduler shutting down" reason="context canceled"
|
||||
Oct 05 09:38:46 demo-felhom felhom-agent[3935656]: time=2026-10-05T09:38:46.215+02:00 level=INFO msg="felhom-agent daemon starting" version=0.145.0 host_id=demo-felhom-8363b5 hub_url=https://hub.felhom.eu interval_s=900
|
||||
Oct 05 09:38:46 demo-felhom felhom-agent[3935656]: time=2026-10-05T09:38:46.836+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
|
||||
Oct 05 09:38:48 demo-felhom felhom-agent[3935656]: time=2026-10-05T09:38:48.047+02:00 level=INFO msg="janitor: starting (restore-test scratch retry + stale-lock sweep)" interval=10m0s
|
||||
Oct 05 10:08:46 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:46.837+02:00 level=INFO msg="backup: restore-test first evaluation after start (R-874)" after=30m0s
|
||||
Oct 05 10:08:47 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:47.156+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-backup archive=felhom-backup:backup/vzdump-lxc-9201-2026_10_04-07_49_00.tar.zst landed=2026-10-04T0
|
||||
Oct 05 10:08:47 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:47.948+02:00 level=INFO msg="restore-test: space preflight passed" storage=local-lvm required_bytes=8002555904 avail_bytes=365212749058
|
||||
Oct 05 10:08:47 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:47.972+02:00 level=INFO msg="restore-test: full-fidelity restore params derived from the archive config" scratch=990000 params=4
|
||||
Oct 05 10:08:48 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:08:48.047+02:00 level=INFO msg="janitor: stale-lock sweep deferred — a heavy operation is in flight" busy=restore-test
|
||||
Oct 05 10:09:16 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:09:16.524+02:00 level=INFO msg="restore-test: scratch guest torn down" vmid=990000
|
||||
Oct 05 10:09:16 demo-felhom felhom-agent[3935656]: time=2026-10-05T10:09:16.524+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=felhom-backup:backup/vzdump-lxc-9201-2026_10_04-07_49_00.tar.zst duration_s=29.36536675
|
||||
@@ -0,0 +1,9 @@
|
||||
### RED-PROOF R-874: the first-evaluation timer dropped (back to the bare ticker)
|
||||
r874_first_eval_test.go:38: no evaluation within 3 s of start (FirstEval 30 ms) — a box with short sessions never gets a restore-test
|
||||
--- FAIL: TestR874_FirstEvaluationAfterStart (3.01s)
|
||||
--- PASS: TestR874_CrashLoopNeverEvaluates (0.10s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-agent/internal/backup 3.119s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/backup 0.151s
|
||||
@@ -0,0 +1,8 @@
|
||||
### RED-PROOF R-875: the v0.144.1 text back
|
||||
unsent_test.go:44: health_reason = "sent after the agent stopped mid-pass (R-868)", want the neutral "sent late …"
|
||||
--- FAIL: TestR868_KilledPassIsReportedAtStart (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-agent/internal/osupdate 0.011s
|
||||
FAIL
|
||||
### restored
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/osupdate 1.738s
|
||||
@@ -0,0 +1,12 @@
|
||||
### 07:55:21 before
|
||||
"armed": true,
|
||||
"last_boot_at": "2026-10-05T06:14:33Z",
|
||||
"tripped": false,
|
||||
"unclean_boots_in_window": 0,
|
||||
09:55:22 up 1:40, 1 user, load average: 2.60, 1.97, 1.49
|
||||
20
|
||||
279
|
||||
### 07:55:24 rollback (13 packages, simulated first)
|
||||
0 upgraded, 0 newly installed, 13 downgraded, 0 to remove and 0 not upgraded.
|
||||
0 upgraded, 0 newly installed, 13 downgraded, 0 to remove and 0 not upgraded.
|
||||
audit rc=0
|
||||
@@ -0,0 +1,5 @@
|
||||
07:55:41.640 debug pass started
|
||||
dpkg running at 07:56:02.802: 302936 /usr/bin/dpkg --force-confold --force-confdef --status-fd 24 --no-triggers --unpack --auto-deconfigure /var/cache
|
||||
CRASH at 07:56:03.308
|
||||
Timeout, server 192.168.0.104 not responding.
|
||||
ssh ended 07:56:10
|
||||
@@ -0,0 +1,25 @@
|
||||
ssh back 07:56:56
|
||||
09:56:56 up 0 min, 1 user, load average: 2.97, 0.65, 0.21
|
||||
"armed": true,
|
||||
"last_boot_at": "2026-10-05T07:56:41Z",
|
||||
"last_boot_unclean": true,
|
||||
"tripped": false,
|
||||
"unclean_boots_in_window": 1,
|
||||
kernel.panic = 10
|
||||
--- dpkg at boot
|
||||
audit rc=0
|
||||
1
|
||||
bind9-dnsutils 1:9.20.26-1~deb13u1 ii
|
||||
bind9-host 1:9.20.26-1~deb13u1 ii
|
||||
bind9-libs 1:9.20.26-1~deb13u1 ii
|
||||
libpcre2-8-0 10.46-1~deb13u2 ii
|
||||
libpython3.13-minimal 3.13.5-2+deb13u4 ii
|
||||
libpython3.13-stdlib 3.13.5-2+deb13u4 ii
|
||||
libssh2-1t64 1.11.1-1+deb13u1 ii
|
||||
libssl3t64 3.5.7-1~deb13u2 ii
|
||||
libxml2 2.12.7+dfsg+really2.9.14-2.1+deb13u1 ii
|
||||
openssl 3.5.7-1~deb13u2 ii
|
||||
openssl-provider-legacy 3.5.7-1~deb13u3 ii
|
||||
python3.13 3.13.5-2+deb13u4 ii
|
||||
python3.13-minimal 3.13.5-2+deb13u4 ii
|
||||
08:00:01 apps: 20 of 21 healthy
|
||||
@@ -0,0 +1,25 @@
|
||||
### 08:00:01 the next pass — no person touched dpkg
|
||||
=== felhom-agent 0.145.0 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true block=hub ===
|
||||
"healthy": true,
|
||||
"outcome": "applied",
|
||||
"healthy": true,
|
||||
"outcome": "nothing",
|
||||
"healthy": true,
|
||||
"outcome": "nothing",
|
||||
pass took 1m20.7s
|
||||
Oct 05 10:00:02 demo-hp felhom-os-apply[16364]: os-apply: START release=ring0-20261005T080001Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0
|
||||
Oct 05 10:00:13 demo-hp felhom-os-apply[16840]: os-apply: REPAIR configured=0 journal=1 fixed=0
|
||||
Oct 05 10:00:19 demo-hp felhom-os-apply[17215]: os-apply: PLAN upgrade=12 already=0 not-installed=0 from-snapshot=0
|
||||
Oct 05 10:00:34 demo-hp felhom-os-apply[18302]: os-apply: DONE rc=0 seconds=7.7 upgraded=12 restart-needed=cron,dbus-daemon,sshd,systemd,systemd-journal,systemd-logind,systemd-network reboot-needed=yes
|
||||
Oct 05 10:00:43 demo-hp felhom-os-apply[18833]: os-apply: START release=ring0-20261005T080001Z layer=host lane=fast mode=apply select=pending-fast packages=0
|
||||
Oct 05 10:00:47 demo-hp felhom-os-apply[19271]: os-apply: REPAIR configured=0 journal=0 fixed=0
|
||||
Oct 05 10:00:51 demo-hp felhom-os-apply[19388]: os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0
|
||||
Oct 05 10:00:51 demo-hp felhom-os-apply[19389]: os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)
|
||||
Oct 05 10:01:04 demo-hp felhom-os-apply[20758]: os-apply: START release=ring0-20261005T080001Z layer=docker:9201 lane=slow mode=apply select=pending-docker packages=0 authority=ring0
|
||||
Oct 05 10:01:09 demo-hp felhom-os-apply[20997]: os-apply: REPAIR configured=0 journal=0 fixed=0
|
||||
Oct 05 10:01:14 demo-hp felhom-os-apply[21238]: os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0
|
||||
Oct 05 10:01:14 demo-hp felhom-os-apply[21239]: os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)
|
||||
audit rc=0
|
||||
0
|
||||
0
|
||||
package list == before the rollback (279 lines)
|
||||
@@ -0,0 +1,8 @@
|
||||
--- hub events demo-hp since 07:55 UTC
|
||||
('controller_started', 'info', 'Controller elindult (0.295.0)', 'controller', '2026-10-05 07:58:50')
|
||||
('os_update_applied', 'info', 'System security fixes installed on the box (12 package(s)).', 'hub', '2026-10-05 08:00:42')
|
||||
--- notification_log since 07:55 UTC (all customers)
|
||||
--- os_reports demo-hp newest 3
|
||||
(73, '2026-10-05 08:01:22', 'debug', 'nothing', 1, 'docker')
|
||||
(72, '2026-10-05 08:01:00', 'debug', 'nothing', 1, 'host')
|
||||
(71, '2026-10-05 08:00:42', 'debug', 'applied', 1, 'guest')
|
||||
@@ -0,0 +1,17 @@
|
||||
### RED-PROOF R-876a: the journal does not count (v0.144.1)
|
||||
FAIL: test_the_next_pass_repairs_by_itself_and_finishes (test_felhom_os_apply.CrashLeftTheJournal.test_the_next_pass_repairs_by_itself_and_finishes)
|
||||
AssertionError: 18 not less than 16 : the repair must run before the install
|
||||
Ran 3 tests in 0.032s
|
||||
FAILED (failures=1)
|
||||
### RED-PROOF R-876b: no repair+retry when apt says interrupted
|
||||
FAIL: test_apt_interrupted_is_repaired_and_retried_once (test_felhom_os_apply.CrashLeftTheJournal.test_apt_interrupted_is_repaired_and_retried_once)
|
||||
AssertionError: 3 != 0 : {'failed': {'dpkg_audit': 'clean', 'rc': 100, 'tail': ["E: dpkg was interrupted, you must manually run 'sudo dpkg --configure -a' to correct the problem."]}, 'health_before': {'containers': {'app': {'health': 'healthy', 'id': 'bbb222', 'state': 'running'}, 'felhom-controller': {'health': 'healthy', 'id': 'aaa111', 'state': 'running'}}, 'controller': 'healthy', 'controller_docker_ok': True, 'docker_ok': True, 'network_ok': True}, 'layer': 'guest', 'mode': 'apply', 'pass_seconds': 0.0, 'plan': {'already': 0, 'from_snapshot': 0, 'not_installed': 0, 'upgrade': 2}, 'refused': None, 'release_id': 'os-t1', 'repair': {'clean_after': True, 'fixed': 0, 'half_configured_before': 0, 'journal_before': 0}, 'vmid': 9201}
|
||||
Ran 3 tests in 0.026s
|
||||
FAILED (failures=1)
|
||||
### RED-PROOF R-876c: repair always runs (R-845 speed lost)
|
||||
FAIL: test_a_clean_pass_still_costs_one_state_call (test_felhom_os_apply.CrashLeftTheJournal.test_a_clean_pass_still_costs_one_state_call)
|
||||
AssertionError: 2 != 1 : R-845: a clean pass reads dpkg's state ONCE and runs no repair
|
||||
Ran 3 tests in 0.035s
|
||||
FAILED (failures=1)
|
||||
### restored
|
||||
OK
|
||||
@@ -0,0 +1,3 @@
|
||||
artifacts POST HTTP 303
|
||||
Location: /configuration?flash=artifacts_set
|
||||
2026/10/05 09:48:14 [INFO] Artifact manifest set: agent=0.145.0 golden=0.295.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="78c00adce662d2d966b2ac50ebde46cde1ae225f0107c6a7c02b70ec8ce80c4f"
|
||||
@@ -0,0 +1,3 @@
|
||||
### 07:31:06 found: VM 341 (Tester 1) stopped since the 06:14 UTC demo-hp crash (no onboot)
|
||||
status: running
|
||||
onboot: 1
|
||||
@@ -0,0 +1,38 @@
|
||||
### R1 — SnapshotInto copies hub.db as a file instead of VACUUM INTO
|
||||
=== RUN TestSnapshot_ConsistentWithLiveDBIncludingWAL
|
||||
dbsnap_test.go:104: table hosts: snapshot 0 rows, live 150
|
||||
dbsnap_test.go:104: table hub_settings: snapshot 0 rows, live 2
|
||||
--- FAIL: TestSnapshot_ConsistentWithLiveDBIncludingWAL (0.05s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 0.055s
|
||||
FAIL
|
||||
|
||||
### R2 — no busy check (two runs may overlap)
|
||||
=== RUN TestSnapshot_NeverTwoAtOnce
|
||||
/mnt/5_hdd/felhom.eu/git/felhom.eu/hub/internal/dbsnap/dbsnap_test.go:141 +0x76
|
||||
/mnt/5_hdd/felhom.eu/git/felhom.eu/hub/internal/dbsnap/dbsnap_test.go:153 +0x1a6
|
||||
/mnt/5_hdd/felhom.eu/git/felhom.eu/hub/internal/dbsnap/dbsnap_test.go:141 +0x76
|
||||
/mnt/5_hdd/felhom.eu/git/felhom.eu/hub/internal/dbsnap/dbsnap_test.go:151 +0x34
|
||||
/mnt/5_hdd/felhom.eu/git/felhom.eu/hub/internal/dbsnap/dbsnap_test.go:151 +0x178
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 600.107s
|
||||
FAIL
|
||||
|
||||
### R3 — prune keeps everything
|
||||
=== RUN TestSnapshot_KeepsNewestTwo
|
||||
dbsnap_test.go:130: kept [hub-20261006T000000Z.db hub-20261007T000000Z.db hub-20261008T000000Z.db], want [hub-20261007T000000Z.db hub-20261008T000000Z.db]
|
||||
--- FAIL: TestSnapshot_KeepsNewestTwo (0.07s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 0.075s
|
||||
FAIL
|
||||
|
||||
### R4 — a failed write leaves its .tmp
|
||||
=== RUN TestSnapshot_FailureLeavesNoFile
|
||||
dbsnap_test.go:191: left 1 file(s) behind
|
||||
--- FAIL: TestSnapshot_FailureLeavesNoFile (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 0.007s
|
||||
FAIL
|
||||
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 0.135s
|
||||
|
||||
[exited with code 0]
|
||||
@@ -0,0 +1,26 @@
|
||||
Run 1 (R1–R4) in red-proof-run1.txt. R2 there HUNG (600 s timeout) — the old test blocked forever; it convicted by hanging. The test now fails cleanly; R2 re-run below.
|
||||
|
||||
### R2 (re-run) — no busy check
|
||||
=== RUN TestSnapshot_NeverTwoAtOnce
|
||||
dbsnap_test.go:162: second Make did not return ErrBusy at once: it ran alongside the first
|
||||
--- FAIL: TestSnapshot_NeverTwoAtOnce (2.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 2.007s
|
||||
FAIL
|
||||
|
||||
### R5 — main never schedules the snapshot
|
||||
=== RUN TestR173_MainSchedulesTheDBSnapshot
|
||||
r173_wiring_test.go:42: cmd/hub/main.go never calls scheduleDaily(ctx, "db-snapshot", "02:00", …) — no nightly snapshot
|
||||
--- FAIL: TestR173_MainSchedulesTheDBSnapshot (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.025s
|
||||
FAIL
|
||||
|
||||
### R6 — main has no start-up catch-up
|
||||
=== RUN TestR173_MainSchedulesTheDBSnapshot
|
||||
r173_wiring_test.go:45: cmd/hub/main.go never calls dbsnap.NeedsCatchUp — a pod down at 02:00 leaves a stale snapshot
|
||||
--- FAIL: TestR173_MainSchedulesTheDBSnapshot (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.023s
|
||||
FAIL
|
||||
|
||||
@@ -0,0 +1,10 @@
|
||||
## Longhorn online expansion FAILS — 2026-10-05T12:37:44Z
|
||||
LAST SEEN TYPE REASON OBJECT MESSAGE
|
||||
7m34s Normal ExternalExpanding persistentvolumeclaim/hub-data waiting for an external controller to expand this PVC
|
||||
6m4s Warning VolumeResizeFailed persistentvolumeclaim/hub-data resize volume "pvc-486c9809-4672-4b56-b70e-0bf01d0c3628" by resizer "driver.longhorn.io" failed: rpc error: code = DeadlineExceeded desc = volume pvc-486c9809-4672-4b56-b70e-0bf01d0c3628 expansion from existing capacity 1073741824 to requested capacity 2147483648 failed
|
||||
4s Normal Resizing persistentvolumeclaim/hub-data External resizer is resizing volume pvc-486c9809-4672-4b56-b70e-0bf01d0c3628
|
||||
4s Warning VolumeResizeFailed persistentvolumeclaim/hub-data resize volume "pvc-486c9809-4672-4b56-b70e-0bf01d0c3628" by resizer "driver.longhorn.io" failed: rpc error: code = DeadlineExceeded desc = volume pvc-486c9809-4672-4b56-b70e-0bf01d0c3628 expansion from existing capacity 2147483648 to requested capacity 2147483648 failed
|
||||
engine pvc-486c9809-4672-4b56-b70e-0bf01d0c3628-e-0 spec=2147483648 current=1073741824
|
||||
[pvc-486c9809-4672-4b56-b70e-0bf01d0c3628-e-0] time="2026-10-05T12:37:45.150634697Z" level=error msg="Failed to expand the frontend" func="controller.(*Controller).Expand.func1.1" file="control.go:332" error="device pvc-486c9809-4672-4b56-b70e-0bf01d0c3628: fail to refresh iSCSI initiator: failed to execute: /usr/bin/nsenter [nsenter --mount=/host/proc/196610/ns/mnt --net=/host/proc/196610/ns/net iscsiadm --version], output , stderr nsenter: cannot open /host/proc/196610/ns/mnt: No such file or directory: exit status 1" volume=pvc-486c9809-4672-4b56-b70e-0bf01d0c3628
|
||||
996780 744880 235516 76% /data
|
||||
instance-manager-703c624848c15b89cb295ca89db85615 1/1 Running 0 116d 10.42.0.229 dooplex <none> <none>
|
||||
@@ -0,0 +1,100 @@
|
||||
## BEFORE 2026-10-05T13:10:25Z
|
||||
pvc-0113bbbf-fd5c-441b-9d58-6943ffa8c133 kisfenyo-filebrowser-data kisfenyo-system attached healthy 524288000
|
||||
pvc-0267c9d6-07e4-476c-8866-0731dcd85fff romm-resources arcade-system attached healthy 10737418240
|
||||
pvc-0eae2db7-47f1-477b-b9a9-9968c2d49a6d gokapi-config fileshare-system attached healthy 1073741824
|
||||
pvc-11b257b2-8edc-4bf6-b846-00979feb195f zipline-data zipline-system attached healthy 10737418240
|
||||
pvc-17e3b3cc-9c6a-4d70-85cc-c2ae72e0b4ec plantit-db plantit-system attached healthy 2147483648
|
||||
pvc-1ae51535-0b2b-4975-bd9f-e651cc7a1edd nextcloud-postgresql-data nextcloud-system attached healthy 5368709120
|
||||
pvc-1c47727a-d5ac-4171-ac36-26c2df3d0d63 nextcloud-nextcloud nextcloud-system attached healthy 10737418240
|
||||
pvc-1d4ec717-7ee5-40a0-b093-3a62fd2a81a4 authentik-media auth-system attached healthy 2147483648
|
||||
pvc-202e3bf2-9e61-4c59-ae0c-43fa6a42b403 tailscale-state admin-system attached healthy 1073741824
|
||||
pvc-22343573-266d-492e-bd0a-37296217b523 postgresql-2 database-system attached healthy 53687091200
|
||||
pvc-22c87ddc-afd7-4520-8473-9e4bf2b6425a immich-machine-learning-ssd2 immich-system attached healthy 10737418240
|
||||
pvc-2382e7bb-9514-409b-a2cc-e97c9725512a uptimekuma-data-ssd2 uptimekuma-system attached healthy 5368709120
|
||||
pvc-297c20e6-4ac0-42c4-9c69-24580a7f12ee act-runner-data gitea-system attached healthy 5368709120
|
||||
pvc-2e2e8c4c-4ba3-4dcf-9ed4-78acaf7ef92b sparkyfitness-uploads workout-system attached healthy 5368709120
|
||||
pvc-328d0394-1ca7-47be-8300-2dc8a1740235 paperless-redis paperless-system attached healthy 1073741824
|
||||
pvc-34a1540e-8b89-4c1d-8465-796bdbfb172f pms-config-plex-plex-media-server-0 mediaserver-system attached healthy 37580963840
|
||||
pvc-38e8c624-5717-483b-8da1-613e703e36c8 umami-db-data felhom-system attached healthy 2147483648
|
||||
pvc-3ae955c9-9968-4099-8a7b-af2bd87d9014 immich-valkey-ssd2 immich-system attached healthy 1073741824
|
||||
pvc-3b23c214-ea15-40f8-aac5-61f783bf5beb plantit-uploads plantit-system attached healthy 5368709120
|
||||
pvc-3d3d043a-92e7-4c34-8383-9d08deb96689 crafty-backups crafty-system attached healthy 53687091200
|
||||
pvc-42695a34-d921-4007-b5e5-e40aa728e104 code-server-config code-system attached healthy 2147483648
|
||||
pvc-459e6d2a-92b8-438e-a339-ff7e40089f3d pihole pihole-system attached healthy 5368709120
|
||||
pvc-486c9809-4672-4b56-b70e-0bf01d0c3628 hub-data felhom-system attached healthy 2147483648
|
||||
pvc-49844e6d-dbe7-427a-a021-8a36a5996478 onlyoffice-logs office-system attached healthy 2147483648
|
||||
pvc-4b9796b8-4c76-4c12-b20e-bce7eb720baf outline-redis outline-system attached healthy 1073741824
|
||||
pvc-4c75264d-5ca8-4298-9916-e84fe7130b3b onlyoffice-data office-system attached healthy 5368709120
|
||||
pvc-4e3fd336-9f09-4251-bd90-cdc259ef739e wanderer-db wanderer-system attached healthy 5368709120
|
||||
pvc-4e4aba1e-b414-4016-8895-ecdf0fb04a6b grafana-data mon-system attached healthy 8589934592
|
||||
pvc-4fdd6da2-e795-4fa9-9e36-96f212c7018b audiobookshelf-config audiobookshelf-system attached healthy 5368709120
|
||||
pvc-59667477-c72a-49ea-bf83-b51a672cc727 crafty-import crafty-system attached healthy 5368709120
|
||||
pvc-5acabaa7-6017-4793-b45e-f1f4fd80ab97 code-server-local code-system attached healthy 5368709120
|
||||
pvc-5c436be9-dbec-4b02-ac26-cb27ce341815 onlyoffice-lib office-system attached healthy 5368709120
|
||||
pvc-5cdbcb85-dfb8-480b-9dc3-2e5d03628dd1 dev-jarr-postgres jarrs-system attached healthy 5368709120
|
||||
pvc-5fbd3bbc-fb14-406d-954c-1a232ce313ce gitea-data gitea-system attached healthy 53687091200
|
||||
pvc-63f3b34f-3ac6-41a5-8790-3a27db0b5a2b revfulop-calendar-data orsi-system attached healthy 268435456
|
||||
pvc-69577be6-c9f1-40f7-8edf-89ce7cac7c5b prometheus-data-ssd1 mon-system attached healthy 32212254720
|
||||
pvc-6db10505-5d7d-4997-9f7a-6df13cf847ba glance-helper-data glance-system attached healthy 209715200
|
||||
pvc-6ed74c1a-3f10-4638-9baa-9330cbff980d opengist-data opengist-system attached healthy 5368709120
|
||||
pvc-71cbab32-86f9-40a5-8f3b-220ad626c492 vaultwarden-data vaultwarden-system attached healthy 5368709120
|
||||
pvc-724a5306-f7a4-4b05-ac28-630597b9b8e2 gokapi-data fileshare-system attached healthy 53687091200
|
||||
pvc-72c77aad-fdc4-4551-a79a-7142fae94fa9 serverpackcreator-mongo-data crafty-system attached healthy 42949672960
|
||||
pvc-79478678-07e4-4140-aef1-771ff763ab3f romm-config arcade-system attached healthy 1073741824
|
||||
pvc-7d6dec2c-0cda-43c7-ad1d-44ffebcb8cb6 seerr-config-pvc servarr-system attached healthy 1073741824
|
||||
pvc-7eb1f195-6a35-46c3-bbb1-5bab48209624 prowlarr-config-pvc servarr-system attached healthy 4294967296
|
||||
pvc-7fc61dcd-52f0-45b4-97f1-9c82b28ab714 sparkyfitness-postgres workout-system attached healthy 5368709120
|
||||
pvc-83701410-61f7-4d16-a8f6-158668110685 calcom-redis booking-system attached healthy 1073741824
|
||||
pvc-8401df55-4668-40e3-97d2-d53e2c1090f1 tautulli-config-pvc servarr-system attached healthy 2147483648
|
||||
pvc-85be87dc-3351-4459-9f82-55f8cc598a2f actualbudget-data actualbudget-system attached healthy 5368709120
|
||||
pvc-86a960ec-ac7f-41a1-b595-00deac982ff5 serverpackcreator-data crafty-system attached healthy 42949672960
|
||||
pvc-88b4d7b4-cb5d-48e0-871e-9e150872d0bd calibre-web-automated-config calibre-system attached healthy 10737418240
|
||||
pvc-8931072a-e34e-4333-9ab5-9629b6078505 upsnap-data admin-system attached healthy 1073741824
|
||||
pvc-8af636d2-2d00-46ec-9947-7b3143563c16 dev-jarr-redis jarrs-system attached healthy 1073741824
|
||||
pvc-8ed5eea5-67d4-46fc-9806-6b8fdd61e42e filebrowser-files felhom-system attached healthy 1073741824
|
||||
pvc-8fcb5bc5-865a-4fc3-8bac-b808b1233dbd tandoor-staticfiles tandoor-system attached healthy 1073741824
|
||||
pvc-9a76846d-b1cf-4928-b950-2db5500531e1 wanderer-meilisearch wanderer-system attached healthy 5368709120
|
||||
pvc-9ca6a486-40bb-46ac-ac48-bd49d4399f4b sonarr-config-pvc servarr-system attached healthy 8589934592
|
||||
pvc-a9dd6c7c-19eb-4a0a-935d-d2c30545c378 romm-db arcade-system attached healthy 2147483648
|
||||
pvc-b0e94be5-7cd8-4a4c-9295-901614e9eeb7 crafty-app-config crafty-system attached healthy 2147483648
|
||||
pvc-b10da48e-6fcc-42ef-8414-d5d617028d68 qbittorrent-config-pvc servarr-system attached healthy 1073741824
|
||||
pvc-b359d7de-1c7c-4b6a-9f6d-6776c2f876ef minecraft-tlauncher-data crafty-system attached healthy 21474836480
|
||||
pvc-b8959b5c-6b88-4e65-81f2-0eabbf4a6034 privatebin-data privatebin-system attached healthy 2147483648
|
||||
pvc-bd481dfa-68d6-4f7a-a634-4d3b76dbf57b recipe-importer-data tandoor-system attached healthy 134217728
|
||||
pvc-be9d7813-1881-462e-8951-1bd1315f7dc5 filebrowser-db felhom-system attached healthy 104857600
|
||||
pvc-beb0ad8a-755a-4152-a0af-f7367b32b19c paperless-config paperless-system attached healthy 10737418240
|
||||
pvc-c4dfbc43-2c35-4ef0-9cfe-39b3dc885dae radarrkids-config-pvc servarr-system attached healthy 3221225472
|
||||
pvc-c6b9ab63-24b8-46f2-b52a-06da86393de8 filebrowser-config web-system attached healthy 104857600
|
||||
pvc-c8c0d0f1-c9ea-4ed3-9e22-1e254a026c49 radarr-config-pvc servarr-system attached healthy 8589934592
|
||||
pvc-c911715e-56ee-4d97-836e-89efc448ad9d immich-postgres-ssd2 immich-system attached healthy 10737418240
|
||||
pvc-cbfd55b0-4f44-487d-90fd-9c80b916ced7 crafty-servers crafty-system attached healthy 53687091200
|
||||
pvc-d2ea5b66-d251-4467-9829-14654fab7a2f bookstack-mariadb-ssd2 bookstack-system attached healthy 5368709120
|
||||
pvc-d9f09398-17cf-44b0-b3a2-9475bc773fec adventurelog-postgres adventurelog-system attached healthy 5368709120
|
||||
pvc-e2151b14-2b24-4128-993a-c392c8d9706a code-server-workspace code-system attached healthy 21474836480
|
||||
pvc-e3b26fdc-fd52-4fdf-9fa2-1d1ecfd029e4 termix-data termix-system attached healthy 5368709120
|
||||
pvc-e57c220f-bc92-4815-89c9-0dee35eab3c3 orsi-filebrowser-config orsi-system attached healthy 104857600
|
||||
pvc-e5b8f824-7720-454c-b546-a098608df24a alertmanager-data mon-system attached healthy 1073741824
|
||||
pvc-ea977627-fb0a-4dee-8678-9ceceea448da bookstack-config-ssd2 bookstack-system attached healthy 5368709120
|
||||
pvc-ef87a570-5d8d-45e5-9bbf-c700cbee1508 onlyoffice-db office-system attached healthy 5368709120
|
||||
pvc-fc3d2f7f-6569-473c-acb4-002c613e5f18 filebrowser-data web-system attached healthy 5368709120
|
||||
## settings
|
||||
auto-delete-pod-when-volume-detached-unexpectedly=true
|
||||
auto-salvage=true
|
||||
## pods not Running/Completed
|
||||
## instance managers
|
||||
instance-manager-703c624848c15b89cb295ca89db85615 1/1 Running 0 116d 10.42.0.229 dooplex <none> <none>
|
||||
instance-manager-703c624848c15b89cb295ca89db85615 aio running
|
||||
## 2026-10-05T13:20:10Z DELETE instance-manager pod
|
||||
pod "instance-manager-703c624848c15b89cb295ca89db85615" deleted
|
||||
13:20:21Z instance-manager: 1/1 Running
|
||||
instance-manager-703c624848c15b89cb295ca89db85615 1/1 Running 0 1s 10.42.0.235 dooplex <none> <none>
|
||||
13:22:00Z not attached+healthy: 0
|
||||
## AFTER 2026-10-05T13:35:41Z
|
||||
volumes not attached+healthy: 0
|
||||
pods not Running/Completed:
|
||||
Running but not ready:
|
||||
zipline: was CrashLoop on :latest=4.8.0 (prisma->drizzle refusal); pinned 4.7.0 in homelab-manifests 90f60e4+4c8ec7a; running, 4 migrations applied
|
||||
## hub-data
|
||||
capacity=2Gi conditions=
|
||||
engine current=2147483648
|
||||
2028392 745616 1266392 37% /data
|
||||
@@ -0,0 +1,22 @@
|
||||
## BEFORE sync 2026-10-05T12:29:57Z
|
||||
pvc labels={"app":"hub","recurring-job-group.longhorn.io/default":"disabled"} request=1Gi capacity=1Gi
|
||||
volume labels={"backup-target":"default","longhornvolume":"pvc-486c9809-4672-4b56-b70e-0bf01d0c3628","recurring-job-group.longhorn.io/default":"enabled","setting.longhorn.io/remove-snapshots-during-filesystem-trim":"ignored","setting.longhorn.io/replica-auto-balance":"ignored","setting.longhorn.io/snapshot-data-integrity":"ignored"} size=1073741824 robustness=healthy
|
||||
image=gitea.dooplex.hu/admin/felhom-hub:0.135.0
|
||||
## AFTER sync 2026-10-05T12:30:55Z
|
||||
pvc labels={"app":"hub","recurring-job-group.longhorn.io/default":"enabled"} request=2Gi capacity=1Gi conditions=[{"lastProbeTime":null,"lastTransitionTime":"2026-10-05T12:30:10Z","status":"True","type":"Resizing"}]
|
||||
volume labels={"backup-target":"default","longhornvolume":"pvc-486c9809-4672-4b56-b70e-0bf01d0c3628","recurring-job-group.longhorn.io/default":"enabled","setting.longhorn.io/remove-snapshots-during-filesystem-trim":"ignored","setting.longhorn.io/replica-auto-balance":"ignored","setting.longhorn.io/snapshot-data-integrity":"ignored"} size=2147483648 robustness=healthy
|
||||
image=gitea.dooplex.hu/admin/felhom-hub:0.136.0
|
||||
## hub log (snapshot + version lines)
|
||||
2026/10/05 14:30:14 [INFO] felhom-hub 0.136.0 starting
|
||||
2026/10/05 14:30:14 [INFO] Default controller-version floor: 0.120.0
|
||||
2026/10/05 14:30:15 [INFO] Gitea artifact browser enabled (Day-0 version dropdowns) via http://gitea.gitea-system.svc.cluster.local:3000
|
||||
2026/10/05 14:30:15 [INFO] Registry version checker started (every 6h)
|
||||
2026/10/05 14:30:16 [DEBUG] Registry version check: latest = 0.296.0
|
||||
2026/10/05 14:30:37 [INFO] db-snapshot: next run at 2026-10-06 02:00 CEST (in 11h29m23s)
|
||||
## /data
|
||||
Filesystem 1K-blocks Used Available Use% Mounted on
|
||||
/dev/longhorn/pvc-486c9809-4672-4b56-b70e-0bf01d0c3628
|
||||
996780 507808 472588 52% /data
|
||||
total 124836
|
||||
-rw-r--r-- 1 root root 127840256 Oct 5 14:30 hub-20261005T123037Z.db.tmp
|
||||
-rw-r--r-- 1 root root 1024 Oct 5 14:30 hub-20261005T123037Z.db.tmp-journal
|
||||
@@ -0,0 +1,16 @@
|
||||
## 2026-10-05T12:53:18Z scale hub to 0
|
||||
deployment.apps/hub scaled
|
||||
12:56:30Z volume state=attached
|
||||
13:01:42Z attached expansionRequired=true
|
||||
engine spec=2147483648 current=1073741824
|
||||
30m Warning VolumeResizeFailed persistentvolumeclaim/hub-data resize volume "pvc-486c9809-4672-4b56-b70e-0bf01d0c3628" by resizer "driver.longhorn.io" failed: rpc error: code = DeadlineExceeded desc = volume pvc-486c9809-4672-4b56-b70e-0bf01d0c3628 expansion from existing capacity 1073741824 to requested capacity 2147483648 failed
|
||||
2s Normal Resizing persistentvolumeclaim/hub-data External resizer is resizing volume pvc-486c9809-4672-4b56-b70e-0bf01d0c3628
|
||||
2s Warning VolumeResizeFailed persistentvolumeclaim/hub-data resize volume "pvc-486c9809-4672-4b56-b70e-0bf01d0c3628" by resizer "driver.longhorn.io" failed: rpc error: code = DeadlineExceeded desc = volume pvc-486c9809-4672-4b56-b70e-0bf01d0c3628 expansion from existing capacity 2147483648 to requested capacity 2147483648 failed
|
||||
## 2026-10-05T13:01:50Z volume never detached — attachment tickets:
|
||||
{"volume-expansion-controller-pvc-486c9809-4672-4b56-b70e-0bf01d0c3628":{"generation":0,"id":"volume-expansion-controller-pvc-486c9809-4672-4b56-b70e-0bf01d0c3628","nodeID":"dooplex","parameters":{"disableFrontend":"false"},"type":"volume-expansion-controller"}}
|
||||
## 2026-10-05T13:01:51Z scale hub to 1
|
||||
deployment.apps/hub scaled
|
||||
Waiting for deployment "hub" rollout to finish: 0 of 1 updated replicas are available...
|
||||
deployment "hub" successfully rolled out
|
||||
NAME READY STATUS RESTARTS AGE
|
||||
hub-586df4748f-ppn4x 1/1 Running 0 54s
|
||||
@@ -0,0 +1,44 @@
|
||||
## 2026-10-05T13:37:43Z ACLs
|
||||
## ep0 AFTER 2026-10-05T13:37:43Z
|
||||
## prune jobs
|
||||
[
|
||||
{
|
||||
"comment": "R-82 retention keep-last=2, server-side (box tokens are write-only)",
|
||||
"id": "prune-demo-felhom",
|
||||
"keep-last": 2,
|
||||
"max-depth": 0,
|
||||
"ns": "demo-felhom",
|
||||
"schedule": "03:30",
|
||||
"store": "felhom-offsite"
|
||||
},
|
||||
{
|
||||
"comment": "R-82 retention keep-last=2, server-side (box tokens are write-only)",
|
||||
"id": "prune-demo-hp",
|
||||
"keep-last": 2,
|
||||
"max-depth": 0,
|
||||
"ns": "demo-hp",
|
||||
"schedule": "03:30",
|
||||
"store": "felhom-offsite"
|
||||
},
|
||||
{
|
||||
"comment": "R-173 hub DB copies; operator ns only",
|
||||
"id": "prune-operator-hubdb",
|
||||
"keep-daily": 14,
|
||||
"keep-weekly": 8,
|
||||
"max-depth": 0,
|
||||
"ns": "operator",
|
||||
"schedule": "03:45",
|
||||
"store": "felhom-offsite"
|
||||
}
|
||||
]
|
||||
## gc
|
||||
felhom-offsite sun 04:30
|
||||
## users
|
||||
['felhom@pbs', 'dooplex-hub@pbs', 'root@pam']
|
||||
## tokens (ids only)
|
||||
['dooplex-hub@pbs!restore', 'dooplex-hub@pbs!push']
|
||||
## ACLs naming dooplex-hub
|
||||
/datastore/felhom-offsite/operator dooplex-hub@pbs!restore DatastoreReader True
|
||||
/datastore/felhom-offsite/operator dooplex-hub@pbs DatastoreReader True
|
||||
/datastore/felhom-offsite/operator dooplex-hub@pbs DatastoreBackup True
|
||||
/datastore/felhom-offsite/operator dooplex-hub@pbs!push DatastoreBackup True
|
||||
@@ -0,0 +1,53 @@
|
||||
## ep0 BEFORE 2026-10-05T13:36:16Z host=felhom-hetzner
|
||||
proxmox-backup-server 4.2.8-1 running version: 4.2.5
|
||||
## prune jobs
|
||||
[
|
||||
{
|
||||
"comment": "R-82 retention keep-last=2, server-side (box tokens are write-only)",
|
||||
"id": "prune-demo-hp",
|
||||
"keep-last": 2,
|
||||
"max-depth": 0,
|
||||
"ns": "demo-hp",
|
||||
"schedule": "03:30",
|
||||
"store": "felhom-offsite"
|
||||
},
|
||||
{
|
||||
"comment": "R-82 retention keep-last=2, server-side (box tokens are write-only)",
|
||||
"id": "prune-demo-felhom",
|
||||
"keep-last": 2,
|
||||
"max-depth": 0,
|
||||
"ns": "demo-felhom",
|
||||
"schedule": "03:30",
|
||||
"store": "felhom-offsite"
|
||||
}
|
||||
]
|
||||
## garbage-collection
|
||||
[
|
||||
{
|
||||
"cache-stats": {
|
||||
"hits": 6726,
|
||||
"misses": 8569
|
||||
},
|
||||
"disk-bytes": 11794377956,
|
||||
"disk-chunks": 8570,
|
||||
"duration": 44,
|
||||
"index-data-bytes": 51257596032,
|
||||
"index-file-count": 8,
|
||||
"last-run-endtime": 1791088244,
|
||||
"last-run-state": "OK",
|
||||
"next-run": 1791693000,
|
||||
"pending-bytes": 0,
|
||||
"pending-chunks": 0,
|
||||
"removed-bad": 0,
|
||||
"removed-bytes": 8871579815,
|
||||
"removed-chunks": 7465,
|
||||
"schedule": "sun 04:30",
|
||||
"still-bad": 0,
|
||||
"store": "felhom-offsite",
|
||||
"upid": "UPID:felhom-hetzner:00086AE7:07B20B30:000002A6:6AC1D648:garbage_collection:felhom\\x2doffsite:root@pam:"
|
||||
}
|
||||
]
|
||||
## users (id only)
|
||||
['root@pam', 'felhom@pbs']
|
||||
## operator ns exists?
|
||||
False
|
||||
@@ -0,0 +1,46 @@
|
||||
## 2026-10-05T13:37:04Z changes
|
||||
user created
|
||||
|
||||
thread 'main' (2343121) panicked at /usr/share/cargo/registry/proxmox-router-3.2.8/src/cli/text_table.rs:812:13:
|
||||
not implemented
|
||||
note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace
|
||||
Error: parameter verification failed - 'schedule': unable to parse calendar event at 'daily' - Context("weekday")
|
||||
Usage: proxmox-backup-manager prune-job create <id> --schedule <calendar-event> --store <string> [OPTIONS]
|
||||
|
||||
<id> <string>
|
||||
Job ID.
|
||||
|
||||
--schedule <calendar-event>
|
||||
Run prune job at specified schedule.
|
||||
|
||||
--store <string>
|
||||
Datastore name.
|
||||
|
||||
Optional parameters:
|
||||
|
||||
--comment <string>
|
||||
Comment.
|
||||
--disable <boolean> (default=false)
|
||||
Disable this job.
|
||||
--keep-daily <integer> (1 - N)
|
||||
Number of daily backups to keep.
|
||||
--keep-hourly <integer> (1 - N)
|
||||
Number of hourly backups to keep.
|
||||
--keep-last <integer> (1 - N)
|
||||
Number of backups to keep.
|
||||
--keep-monthly <integer> (1 - N)
|
||||
Number of monthly backups to keep.
|
||||
--keep-weekly <integer> (1 - N)
|
||||
Number of weekly backups to keep.
|
||||
--keep-yearly <integer> (1 - N)
|
||||
Number of yearly backups to keep.
|
||||
--max-depth <integer> (0 - 7)
|
||||
How many levels of namespaces should be operated on (0 == no
|
||||
recursion, empty == automatic full recursion, namespace depths
|
||||
reduce maximum allowed value)
|
||||
--ns <string>
|
||||
Namespace.
|
||||
## 2026-10-05T13:37:12Z check + retry
|
||||
operator ns exists: True
|
||||
Prune job created: prune-operator-hubdb
|
||||
prune job created
|
||||
@@ -0,0 +1,26 @@
|
||||
## promtool, in pod/prometheus-55b675779d-t8c74, 2026-10-05T13:42:17Z
|
||||
### green
|
||||
SUCCESS
|
||||
|
||||
rc=0
|
||||
### red: threshold 26h -> 260h
|
||||
FAILED:
|
||||
alertname: HubDBBackupStale, time: 1d2h40m,
|
||||
exp:[
|
||||
0:
|
||||
Labels:{alertname="HubDBBackupStale", component="backup", instance="dooplex", severity="critical"}
|
||||
Annotations:{description="No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00.", summary="The Felhom hub database has not reached ep0 for 26 h (R-173)"}
|
||||
],
|
||||
got:[]
|
||||
|
||||
|
||||
command terminated with exit code 1
|
||||
rc=0
|
||||
### red: absent() removed
|
||||
SUCCESS
|
||||
|
||||
### red (re-run): absent() removed from HubDBBackupStale — first attempt above did NOT apply (indent mismatch)
|
||||
FAILED:
|
||||
alertname: HubDBBackupStale, time: 40m,
|
||||
Labels:{alertname="HubDBBackupStale", component="backup", severity="critical"}
|
||||
got:[]
|
||||
@@ -0,0 +1,52 @@
|
||||
rule_files: [bf.yml]
|
||||
evaluation_interval: 1m
|
||||
tests:
|
||||
# 1. the push stops: last success at t=0, never again → fires once 26 h + 30 m have passed
|
||||
- interval: 5m
|
||||
input_series:
|
||||
- series: 'felhom_hub_db_backup_last_success_timestamp_seconds{instance="dooplex"}'
|
||||
values: '0+0x400'
|
||||
- series: 'felhom_hub_db_restore_test_last_success_timestamp_seconds{instance="dooplex"}'
|
||||
values: '0+0x400'
|
||||
alert_rule_test:
|
||||
- eval_time: 26h
|
||||
alertname: HubDBBackupStale
|
||||
exp_alerts: []
|
||||
- eval_time: 26h40m
|
||||
alertname: HubDBBackupStale
|
||||
exp_alerts:
|
||||
- exp_labels: {severity: critical, component: backup, instance: dooplex}
|
||||
exp_annotations:
|
||||
summary: "The Felhom hub database has not reached ep0 for 26 h (R-173)"
|
||||
description: "No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00."
|
||||
# 2. healthy: the push succeeds every 24 h → never fires
|
||||
- interval: 1h
|
||||
input_series:
|
||||
- series: 'felhom_hub_db_backup_last_success_timestamp_seconds{instance="dooplex"}'
|
||||
values: '0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 172800 172800 172800'
|
||||
- series: 'felhom_hub_db_restore_test_last_success_timestamp_seconds{instance="dooplex"}'
|
||||
values: '0+0x50'
|
||||
alert_rule_test:
|
||||
- eval_time: 50h
|
||||
alertname: HubDBBackupStale
|
||||
exp_alerts: []
|
||||
# 3. the metric never existed (script never ran) → fires on absent()
|
||||
- interval: 5m
|
||||
input_series:
|
||||
- series: 'up{job="node"}'
|
||||
values: '1+0x30'
|
||||
alert_rule_test:
|
||||
- eval_time: 40m
|
||||
alertname: HubDBBackupStale
|
||||
exp_alerts:
|
||||
- exp_labels: {severity: critical, component: backup}
|
||||
exp_annotations:
|
||||
summary: "The Felhom hub database has not reached ep0 for 26 h (R-173)"
|
||||
description: "No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00."
|
||||
- eval_time: 2h
|
||||
alertname: HubDBRestoreTestStale
|
||||
exp_alerts:
|
||||
- exp_labels: {severity: warning, component: backup}
|
||||
exp_annotations:
|
||||
summary: "The hub database copy on ep0 has not passed a restore test for 8 days (R-173)"
|
||||
description: "The weekly restore test (Sun 04:30) has not succeeded for >8 days. Check `journalctl -u felhom-hub-db-restore-test.service`."
|
||||
@@ -0,0 +1,76 @@
|
||||
### P1 push: integrity_check ignored
|
||||
test_corrupt_copy_is_never_pushed (__main__.Push.test_corrupt_copy_is_never_pushed) ... FAIL
|
||||
FAIL: test_corrupt_copy_is_never_pushed (__main__.Push.test_corrupt_copy_is_never_pushed)
|
||||
AssertionError: 0 == 0
|
||||
Ran 1 test in 0.205s
|
||||
FAILED (failures=1)
|
||||
|
||||
### P2 push: snapshot age ignored
|
||||
test_stale_snapshot_is_not_pushed_again (__main__.Push.test_stale_snapshot_is_not_pushed_again) ... FAIL
|
||||
FAIL: test_stale_snapshot_is_not_pushed_again (__main__.Push.test_stale_snapshot_is_not_pushed_again)
|
||||
AssertionError: 0 == 0
|
||||
Ran 1 test in 0.201s
|
||||
FAILED (failures=1)
|
||||
|
||||
### P3 push: success signal after a failed push
|
||||
test_failed_push_writes_no_signal (__main__.Push.test_failed_push_writes_no_signal) ... FAIL
|
||||
FAIL: test_failed_push_writes_no_signal (__main__.Push.test_failed_push_writes_no_signal)
|
||||
AssertionError: 0 == 0
|
||||
Ran 1 test in 0.206s
|
||||
FAILED (failures=1)
|
||||
|
||||
### P4 push: size check removed
|
||||
test_truncated_copy_is_not_pushed (__main__.Push.test_truncated_copy_is_not_pushed) ... ok
|
||||
Ran 1 test in 0.156s
|
||||
OK
|
||||
|
||||
### P5 push: no encryption flag
|
||||
test_happy_path_pushes_encrypted_to_operator_and_writes_signal (__main__.Push.test_happy_path_pushes_encrypted_to_operator_and_writes_signal) ... ERROR
|
||||
ERROR: test_happy_path_pushes_encrypted_to_operator_and_writes_signal (__main__.Push.test_happy_path_pushes_encrypted_to_operator_and_writes_signal)
|
||||
Ran 1 test in 0.211s
|
||||
FAILED (errors=1)
|
||||
|
||||
### P6 restore: readable console passwords allowed
|
||||
test_readable_console_password_fails (__main__.RestoreTest.test_readable_console_password_fails) ... FAIL
|
||||
FAIL: test_readable_console_password_fails (__main__.RestoreTest.test_readable_console_password_fails)
|
||||
AssertionError: 0 == 0
|
||||
Ran 1 test in 0.357s
|
||||
FAILED (failures=1)
|
||||
|
||||
### P7 restore: uses the push token
|
||||
test_happy_path_writes_signal_with_the_read_only_token (__main__.RestoreTest.test_happy_path_writes_signal_with_the_read_only_token) ... FAIL
|
||||
FAIL: test_happy_path_writes_signal_with_the_read_only_token (__main__.RestoreTest.test_happy_path_writes_signal_with_the_read_only_token)
|
||||
AssertionError: False is not true : {'argv': ['snapshot', 'list', 'host/dooplex-hub', '--ns', 'operator', '--output-format', 'json', '--repository', 'dooplex-hub@pbs!restore@127.0.0.1:18007:felhom-offsite'], 'pw': '/tmp/tmp0xabeyiz/conf/token-push', 'fp': 'aa:bb'}
|
||||
Ran 1 test in 0.371s
|
||||
FAILED (failures=1)
|
||||
|
||||
### P8 restore: copy age ignored
|
||||
test_old_copy_fails (__main__.RestoreTest.test_old_copy_fails) ... FAIL
|
||||
FAIL: test_old_copy_fails (__main__.RestoreTest.test_old_copy_fails)
|
||||
AssertionError: 0 == 0
|
||||
Ran 1 test in 0.352s
|
||||
FAILED (failures=1)
|
||||
|
||||
### P9 push: staged plaintext left behind
|
||||
test_happy_path_pushes_encrypted_to_operator_and_writes_signal (__main__.Push.test_happy_path_pushes_encrypted_to_operator_and_writes_signal) ... FAIL
|
||||
FAIL: test_happy_path_pushes_encrypted_to_operator_and_writes_signal (__main__.Push.test_happy_path_pushes_encrypted_to_operator_and_writes_signal)
|
||||
AssertionError: Lists differ: ['hub.db'] != []
|
||||
Ran 1 test in 0.200s
|
||||
FAILED (failures=1)
|
||||
|
||||
P4 and P5 above: P4 did NOT convict (masked by integrity_check), P5 ERRORED rather than failed. Tests strengthened; re-runs:
|
||||
|
||||
### P4 (re-run) push: size check removed
|
||||
test_truncated_copy_is_not_pushed (__main__.Push.test_truncated_copy_is_not_pushed) ... FAIL
|
||||
FAIL: test_truncated_copy_is_not_pushed (__main__.Push.test_truncated_copy_is_not_pushed)
|
||||
AssertionError: "bytes, the pod's file is" not found in 'felhom-hub-db-backup: FAILED: integrity_check: Error: in prepare, database disk image is malformed (11)\n'
|
||||
Ran 1 test in 0.146s
|
||||
FAILED (failures=1)
|
||||
|
||||
### P5 (re-run) push: no encryption flag
|
||||
test_happy_path_pushes_encrypted_to_operator_and_writes_signal (__main__.Push.test_happy_path_pushes_encrypted_to_operator_and_writes_signal) ... FAIL
|
||||
FAIL: test_happy_path_pushes_encrypted_to_operator_and_writes_signal (__main__.Push.test_happy_path_pushes_encrypted_to_operator_and_writes_signal)
|
||||
AssertionError: '--crypt-mode' not found in ['backup', 'hubdb.pxar:/tmp/tmpq980kr66/state/stage', '--ns', 'operator', '--backup-type', 'host', '--backup-id', 'dooplex-hub', '--keyfile', '/tmp/tmpq980kr66/conf/enc.key', '--repository', 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite'] : push without --crypt-mode
|
||||
Ran 1 test in 0.200s
|
||||
FAILED (failures=1)
|
||||
|
||||
@@ -0,0 +1,16 @@
|
||||
## operator confirmation (inbox screenshot): [FIRING] HubDBBackupStale 16:22 CEST; [RESOLVED] HubDBBackupStale 16:27 CEST
|
||||
## alarm drill start 2026-10-05T13:51:46Z: success file moved aside (absent case)
|
||||
backup_freshness.prom
|
||||
fan_metrics.prom
|
||||
felhom_hub_db_restore.prom
|
||||
node_housekeeping.prom
|
||||
## 2026-10-05T14:23:16Z HubDBBackupStale FIRING (activeAt 13:52:06Z; fired 14:22); Alertmanager: 1 active alert, receiver email-notifications (to: the operator's address); email notifications_total=5 failed_total=0 since the pod started 13:20:48Z
|
||||
## recovery push: unit Result=success
|
||||
felhom-hub-db-backup: snapshot hub-20261005T123037Z.db, 112 min old
|
||||
felhom-hub-db-backup: checked: 369807360 bytes, integrity ok, 4 host(s)
|
||||
felhom-hub-db-backup: pushed hub-20261005T123037Z.db to ep0 (ns operator) in 2 s
|
||||
felhom-hub-db-backup: success signal written
|
||||
felhom_hub_db_backup_last_success_timestamp_seconds 1791210207
|
||||
felhom_hub_db_backup_last_success_bytes 369807360
|
||||
## 2026-10-05T14:23:42Z HubDBBackupStale state=inactive
|
||||
email notifications_total now: 6
|
||||
@@ -0,0 +1,27 @@
|
||||
## manual push via unit, 2026-10-05T13:50:30Z
|
||||
Result=success
|
||||
ExecMainStatus=0
|
||||
2026-10-05T15:50:08+02:00 dooplex systemd[1]: Starting felhom-hub-db-backup.service - Felhom: push the hub DB snapshot to ep0 (R-173)...
|
||||
2026-10-05T15:50:09+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: snapshot hub-20261005T123037Z.db, 79 min old
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: checked: 369807360 bytes, integrity ok, 4 host(s)
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Starting backup: [operator]:host/dooplex-hub/2026-10-05T13:50:18Z
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Client name: dooplex
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Starting backup protocol: Mon Oct 5 15:50:18 2026
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Using encryption key from '/etc/felhom-hub-backup/enc.key'..
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Encryption key fingerprint: b2:19:bf:36:3b:97:3d:6c
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: No previous manifest available.
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Upload directory '/var/lib/felhom-hub-backup/stage' to 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite' as hubdb.pxar.didx
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: hubdb.pxar: had to backup 352.676 MiB of 352.676 MiB (compressed 17.242 MiB) in 6.58 s (average 53.604 MiB/s)
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: Uploaded backup catalog (56 B)
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: Duration: 7.19s
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: End Time: Mon Oct 5 15:50:25 2026
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: pushed hub-20261005T123037Z.db to ep0 (ns operator) in 7 s
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: success signal written
|
||||
2026-10-05T15:50:30+02:00 dooplex systemd[1]: felhom-hub-db-backup.service: Deactivated successfully.
|
||||
2026-10-05T15:50:30+02:00 dooplex systemd[1]: Finished felhom-hub-db-backup.service - Felhom: push the hub DB snapshot to ep0 (R-173).
|
||||
2026-10-05T15:50:30+02:00 dooplex systemd[1]: felhom-hub-db-backup.service: Consumed 9.566s CPU time, 584.8M memory peak.
|
||||
## textfile
|
||||
# HELP felhom_hub_db_backup_last_success_timestamp_seconds Last successful push of the hub DB snapshot to ep0 (R-173).
|
||||
# TYPE felhom_hub_db_backup_last_success_timestamp_seconds gauge
|
||||
felhom_hub_db_backup_last_success_timestamp_seconds 1791208225
|
||||
felhom_hub_db_backup_last_success_bytes 369807360
|
||||
@@ -0,0 +1 @@
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.296.0 Up 6 seconds (healthy)
|
||||
@@ -0,0 +1,3 @@
|
||||
## deploy bookstack 2026-10-05T13:53:26Z (fields: DOMAIN APP_KEY DB_PASSWORD ADMIN_PASSWORD; SUBDOMAIN = catalog default)
|
||||
|
||||
HTTP 202
|
||||
@@ -0,0 +1,36 @@
|
||||
## run 1 (baseline) start 2026-10-05T14:04:36Z
|
||||
|
||||
HTTP 200
|
||||
{"ok":true,"data":{"enabled":true,"running":true}}
|
||||
|
||||
HTTP 200
|
||||
|
||||
## run 1 end 2026-10-05T14:05:17Z: {"ok":true,"data":{"db_dump":{"count":2,"duration":"32.784678377s","last_run":"2026-10-05T14:05:12.109790785Z","success":true},"enabled":true,"running":false}}
|
||||
2026/10/05 14:04:44 data_versions.go:131: [DEBUG] [backup] bookstack: stamped volume-dumps/bookstack_bookstack_config.tar (70144 B) with pins [lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1]
|
||||
2026/10/05 14:04:44 backup.go:848: [DEBUG] [backup] Dumping volume bookstack_bookstack_db_data for bookstack
|
||||
2026/10/05 14:04:45 backup.go:873: [INFO] [backup] Volume dump: bookstack/bookstack_bookstack_db_data → 153.4 MB
|
||||
2026/10/05 14:04:45 data_versions.go:131: [DEBUG] [backup] bookstack: stamped volume-dumps/bookstack_bookstack_db_data.tar (160868864 B) with pins [lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1]
|
||||
2026/10/05 14:04:45 backup.go:990: [INFO] [backup] Restarting bookstack after volume dump
|
||||
2026/10/05 14:04:51 backup.go:974: [INFO] [backup] Stopping paperless-ngx for safe volume dump
|
||||
2026/10/05 14:04:58 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_redis_data for paperless-ngx
|
||||
2026/10/05 14:04:58 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 7.4 MB
|
||||
2026/10/05 14:04:58 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_redis_data.tar (7792128 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine]
|
||||
2026/10/05 14:04:58 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_data for paperless-ngx
|
||||
2026/10/05 14:04:59 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 5.8 MB
|
||||
2026/10/05 14:04:59 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_data.tar (6052352 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine]
|
||||
2026/10/05 14:04:59 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_postgres_data for paperless-ngx
|
||||
2026/10/05 14:04:59 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 68.8 MB
|
||||
2026/10/05 14:04:59 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_postgres_data.tar (72107008 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine]
|
||||
2026/10/05 14:04:59 backup.go:990: [INFO] [backup] Restarting paperless-ngx after volume dump
|
||||
2026/10/05 14:05:11 backup.go:974: [INFO] [backup] Stopping privatebin for safe volume dump
|
||||
2026/10/05 14:05:11 backup.go:848: [DEBUG] [backup] Dumping volume privatebin_privatebin_data for privatebin
|
||||
2026/10/05 14:05:12 backup.go:873: [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 1.5 KB
|
||||
2026/10/05 14:05:12 data_versions.go:131: [DEBUG] [backup] privatebin: stamped volume-dumps/privatebin_privatebin_data.tar (1536 B) with pins [privatebin/pdo:2.0.6@sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5]
|
||||
2026/10/05 14:05:12 backup.go:986: [INFO] [backup] privatebin NOT restarted after the volume dump: the household stopped it meanwhile
|
||||
2026/10/05 14:05:12 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.2 KB total), 3 volume dump(s) (32.785s)
|
||||
2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for bookstack → /mnt/sys_drive/felhom-data/backups/primary/bookstack (images=2, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for paperless-ngx → /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for privatebin → /mnt/sys_drive/felhom-data/backups/primary/privatebin (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0)
|
||||
|
||||
/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/appdata-run.json
|
||||
/var/lib/docker/volumes/felhom-controller-data/_data/data/appdata-run.json
|
||||
@@ -0,0 +1,12 @@
|
||||
## after run 1, 2026-10-05T14:05:34Z
|
||||
### run record
|
||||
{"running":false,"started_at":"0001-01-01T00:00:00Z","interrupted":"0001-01-01T00:00:00Z"}
|
||||
### restore points (bookstack)
|
||||
{"ok":true,"data":[{"time":"2026-10-05T14:04:39Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
|
||||
|
||||
HTTP 200
|
||||
|
||||
### volume dump file times
|
||||
total 157172
|
||||
-rw-r--r-- 1 root root 70144 2026-10-05T14:04:44 bookstack_bookstack_config.tar
|
||||
-rw-r--r-- 1 root root 160868864 2026-10-05T14:04:44 bookstack_bookstack_db_data.tar
|
||||
@@ -0,0 +1,34 @@
|
||||
## run 2 start 2026-10-05T14:05:50Z
|
||||
HTTP 200
|
||||
### watcher
|
||||
match 2026-10-05T14:05:58.856456447Z
|
||||
restarted 2026-10-05T14:05:59.669364637Z rc=0
|
||||
Up 10 seconds (healthy)
|
||||
### record at the moment of the cut
|
||||
{"running":true,"started_at":"2026-10-05T14:05:53.282495446Z","interrupted":"0001-01-01T00:00:00Z"}
|
||||
### record after restart
|
||||
{"running":false,"started_at":"2026-10-05T14:05:53.282495446Z","interrupted":"2026-10-05T14:05:53.282495446Z"}
|
||||
### controller log around the cut (previous + new process)
|
||||
2026/10/05 14:05:53 backup.go:557: [INFO] [backup] Starting database dump run
|
||||
2026/10/05 14:05:53 dbdump.go:208: [INFO] [backup] Discovered 2 databases
|
||||
2026/10/05 14:05:53 backup.go:596: [INFO] [backup] Discovered 2 database(s): paperless-postgres(postgres), bookstack-db(mariadb)
|
||||
2026/10/05 14:05:53 dbdump.go:410: [INFO] [backup] DB dump: paperless-postgres → paperless-ngx-postgres.sql (428.1 KB, 321ms, 72 tables)
|
||||
2026/10/05 14:05:54 dbdump.go:410: [INFO] [backup] DB dump: bookstack-db → bookstack-mariadb.sql (55.9 KB, 310ms, 41 tables)
|
||||
2026/10/05 14:05:54 backup.go:974: [INFO] [backup] Stopping bookstack for safe volume dump
|
||||
2026/10/05 14:05:58 backup.go:873: [INFO] [backup] Volume dump: bookstack/bookstack_bookstack_config → 49.0 KB
|
||||
2026/10/05 14:05:59 main.go:340: [INFO] felhom-controller 0.296.0 starting (customer: demo-hp, domain: enkisfelhom.hu)
|
||||
2026/10/05 14:05:59 appstop_marker.go:278: [WARN] [appstop] crash recovery: an app-data backup (volume dump) (op "volume-dump:bookstack") was interrupted and left 1 app(s) stopped — restarting them: [bookstack]
|
||||
2026/10/05 14:05:59 manager.go:1235: [INFO] [stacks] Starting stack: bookstack
|
||||
2026/10/05 14:06:06 appstop_marker.go:304: [INFO] [appstop] crash recovery: restarted bookstack after the interrupted an app-data backup (volume dump)
|
||||
2026/10/05 14:06:06 sync.go:117: [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s)
|
||||
2026/10/05 14:06:06 sync.go:208: [INFO] [sync] Starting catalog sync
|
||||
2026/10/05 14:06:06 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false)
|
||||
2026/10/05 14:06:06 run_record_wiring.go:27: [WARN] [backup] the app-data backup run started 2026-10-05T14:05:53Z was cut off by the stop — the backup pages say so until the next complete run (R-519)
|
||||
2026/10/05 14:06:06 scheduler.go:223: [INFO] [scheduler] Starting scheduler with 19 jobs
|
||||
2026/10/05 14:06:06 backup.go:1179: [INFO] [backup] Found 3 DB dump files across drives
|
||||
2026/10/05 14:06:06 dbdump.go:208: [INFO] [backup] Discovered 2 databases
|
||||
2026/10/05 14:06:06 [INFO] [backup] Discovered app data: 3 apps
|
||||
2026/10/05 14:06:06 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for bookstack → /mnt/sys_drive/felhom-data/backups/primary/bookstack (images=2, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/10/05 14:06:06 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for paperless-ngx → /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/10/05 14:06:06 backup.go:1250: [INFO] [backup] Backup status cache refreshed
|
||||
2026/10/05 14:06:09 manager.go:1601: [INFO] [stacks] bookstack lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 running Up 3 seconds (health: starting)
|
||||
@@ -0,0 +1,23 @@
|
||||
## after the cut, 2026-10-05T14:06:23Z
|
||||
### GET /backups
|
||||
data-interrupted-run present: 1
|
||||
<div class="alert alert-warning" data-interrupted-run="true">A legutóbbi mentés (2026-10-05 16:05) megszakadt, mert a doboz vagy a vezérlő újraindult. Amit nem fejezett be, annak a korábbi mentése maradt meg — minden visszaállítási pont annyira friss, amennyire a legrégebbi része. A következő teljes mentés után ez az üzenet eltűnik.
|
||||
HTTP 200
|
||||
### GET /backups/apps
|
||||
data-interrupted-run present: 1
|
||||
<div class="alert alert-warning" data-interrupted-run="true">A legutóbbi mentés (2026-10-05 16:05) megszakadt, mert a doboz vagy a vezérlő újraindult. Amit nem fejezett be, annak a korábbi mentése maradt meg — minden visszaállítási pont annyira friss, amennyire a legrégebbi része. A következő teljes mentés után ez az üzenet eltűnik.
|
||||
HTTP 200
|
||||
### negative control: a marker that must NOT be on the page
|
||||
0
|
||||
### restore points (bookstack)
|
||||
{"ok":true,"data":[{"time":"2026-10-05T14:04:44Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
|
||||
### dump file times
|
||||
total 314252
|
||||
-rw-r--r-- 1 root root 50176 2026-10-05T14:05:58 bookstack_bookstack_config.tar
|
||||
-rw-r--r-- 1 root root 160868864 2026-10-05T14:04:44 bookstack_bookstack_db_data.tar
|
||||
-rw-r--r-- 1 root root 160867328 2026-10-05T14:05:58 bookstack_bookstack_db_data.tar.tmp
|
||||
drwxr-xr-x 2 root root 4096 2026-10-05T14:05:54 db-dumps
|
||||
drwxr-xr-x 2 root root 4096 2026-10-05T14:05:58 volume-dumps
|
||||
### bookstack after
|
||||
bookstack Up 27 seconds (healthy)
|
||||
bookstack-db Up 32 seconds (healthy)
|
||||
@@ -0,0 +1,12 @@
|
||||
## run 3 (complete) start 2026-10-05T14:06:44Z
|
||||
HTTP 200
|
||||
## run 3 end 2026-10-05T14:07:28Z: {"ok":true,"data":{"db_dump":{"count":2,"duration":"33.315240638s","last_run":"2026-10-05T14:07:20.615183031Z","success":true},"enabled":true,"running":false}}
|
||||
2026/10/05 14:05:12 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.2 KB total), 3 volume dump(s) (32.785s)
|
||||
2026/10/05 14:07:20 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.9 KB total), 3 volume dump(s) (33.315s)
|
||||
{"running":false,"started_at":"0001-01-01T00:00:00Z","interrupted":"0001-01-01T00:00:00Z"}
|
||||
total 157152
|
||||
-rw-r--r-- 1 root root 50688 2026-10-05T14:06:52 bookstack_bookstack_config.tar
|
||||
-rw-r--r-- 1 root root 160867328 2026-10-05T14:06:53 bookstack_bookstack_db_data.tar
|
||||
/backups data-interrupted-run: 0
|
||||
/backups/apps data-interrupted-run: 0
|
||||
restore points: {"ok":true,"data":[{"time":"2026-10-05T14:06:47Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
|
||||
@@ -0,0 +1,8 @@
|
||||
## teardown 2026-10-05T14:07:48Z
|
||||
stop: HTTP 200
|
||||
remove: HTTP 200
|
||||
containers: 0
|
||||
volumes: 0
|
||||
stackdir: none
|
||||
backups: none
|
||||
/root/.dbody
|
||||
@@ -0,0 +1,9 @@
|
||||
## runbook §3 drill 2026-10-05T14:09:41Z (scratch /var/lib/felhom-hub-backup/sec3.65qk, root 0700)
|
||||
step 1 restore host/dooplex-hub/2026-10-05T13:50:18Z
|
||||
restore complete (352.676 MiB processed in 3.9s, average 90.961 MiB/s)
|
||||
-rw------- 369807360 hub.db
|
||||
step 2 the seal key from Secret/offsite-secret-key → /var/lib/felhom-hub-backup/sec3.65qk/k (0600, not printed)
|
||||
key file bytes: 64
|
||||
step 3 hubdb-check (same key): hosts=4 console_passwords_opened=4 failed=0 absent=0 rc=0
|
||||
control (a random key): hosts=4 console_passwords_opened=0 failed=4 absent=0 hubdb-check: FAILED: not every console password opened with this key rc=0
|
||||
scratch shredded: gone
|
||||
@@ -0,0 +1,7 @@
|
||||
### R7 hubdb-check counts a failed open as opened
|
||||
=== RUN TestCheck_WrongKeyOpensNothing
|
||||
main_test.go:63: got {hosts:2 opened:1 failed:0 absent:1}, want opened=0 failed=1
|
||||
--- FAIL: TestCheck_WrongKeyOpensNothing (0.03s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hubdb-check 0.040s
|
||||
FAIL
|
||||
@@ -0,0 +1,17 @@
|
||||
## snapshot list on ep0 with the READ-ONLY token, 2026-10-05T13:50:40Z
|
||||
+=======================================+=============+=====================================+
|
||||
| snapshot | size | files |
|
||||
+=======================================+=============+=====================================+
|
||||
| host/dooplex-hub/2026-10-05T13:50:18Z | 352.677 MiB | catalog.pcat1 hubdb.pxar index.json |
|
||||
+=======================================+=============+=====================================+
|
||||
unit rc=0
|
||||
Result=success
|
||||
2026-10-05T15:50:41+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: restoring host/dooplex-hub/2026-10-05T13:50:18Z
|
||||
2026-10-05T15:50:46+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: checked: integrity ok, 4 host(s), 4 sealed console password(s), 0 readable
|
||||
2026-10-05T15:50:46+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: success signal written
|
||||
# HELP felhom_hub_db_restore_test_last_success_timestamp_seconds Last successful restore test of the hub DB copy on ep0 (R-173).
|
||||
# TYPE felhom_hub_db_restore_test_last_success_timestamp_seconds gauge
|
||||
felhom_hub_db_restore_test_last_success_timestamp_seconds 1791208246
|
||||
.cache
|
||||
.kube
|
||||
stage
|
||||
@@ -0,0 +1,12 @@
|
||||
## 2026-10-05T13:51:07Z token limits (each line = the tool output)
|
||||
push forget : Error: permission check failed - missing Datastore.Modify|Datastore.Prune on /datastore/felhom-offsite/operator
|
||||
restore forget : Error: permission check failed - missing Datastore.Modify|Datastore.Prune on /datastore/felhom-offsite/operator
|
||||
restore backup : Error: missing permissions 'Datastore.Backup' on '/datastore/felhom-offsite/operator'
|
||||
push list ns root (households): Error: permission check failed - missing Datastore.Audit|Datastore.Backup on /datastore/felhom-offsite
|
||||
restore list ns root (households): Error: permission check failed - missing Datastore.Audit|Datastore.Backup on /datastore/felhom-offsite
|
||||
push restore own copy: Error: missing key - manifest was created with key b2:19:bf:36:3b:97:3d:6c exit-file=absent
|
||||
## paper-key proof: restore with a key file rebuilt from the data field only
|
||||
restore rc=0
|
||||
integrity: ok hosts: 4
|
||||
## and with NO key: Error: missing key - manifest was created with key b2:19:bf:36:3b:97:3d:6c
|
||||
push restore own copy WITH key: restore complete (352.676 MiB processed in 4.6s, average 76.793 MiB/s) file=PRESENT
|
||||
@@ -0,0 +1,8 @@
|
||||
== Part A live, hub 0.135.0, 2026-10-05T09:15:34Z, ClusterIP, Basic auth (password from the credentials file, not printed)
|
||||
POST /configuration/global-floor, Basic, NO header, Origin evil (empty form): 403
|
||||
POST /no-such-route, Basic, NO header: 403
|
||||
POST /no-such-route, Basic + X-Felhom-Operator: cli (passes the gate → router 404): 404
|
||||
POST /no-such-route, header but NO credentials: 401
|
||||
GET /system, Basic, no header: 200
|
||||
2026/10/05 11:15:34 [WARN] CSRF rejected: POST /configuration/global-floor from 10.42.0.1:36344
|
||||
2026/10/05 11:15:34 [WARN] CSRF rejected: POST /no-such-route from 10.42.0.1:41323
|
||||
@@ -0,0 +1,18 @@
|
||||
== RED-PROOF R-135: validateCSRF returns true when there is no cookie (pre-v0.135.0)
|
||||
--- FAIL: TestR135_BasicAuthWithoutHeaderIsRefused (3.41s)
|
||||
r135_csrf_test.go:92: POST /configuration with Basic auth and no X-Felhom-Operator header: 303, want 403
|
||||
r135_csrf_test.go:92: POST /apps/demo/reset-telemetry with Basic auth and no X-Felhom-Operator header: 303, want 403
|
||||
r135_csrf_test.go:92: POST /apps/demo/dismiss-issues with Basic auth and no X-Felhom-Operator header: 400, want 403
|
||||
r135_csrf_test.go:92: POST /offsite/endpoints with Basic auth and no X-Felhom-Operator header: 400, want 403
|
||||
r135_csrf_test.go:92: POST /offsite/endpoints/1/delete with Basic auth and no X-Felhom-Operator header: 404, want 403
|
||||
r135_csrf_test.go:92: POST /appliances/1/bind with Basic auth and no X-Felhom-Operator header: 400, want 403
|
||||
r135_csrf_test.go:92: POST /appliances/1/discard with Basic auth and no X-Felhom-Operator header: 409, want 403
|
||||
r135_csrf_test.go:92: POST /hosts/h1/delete with Basic auth and no X-Felhom-Operator header: 404, want 403
|
||||
r135_csrf_test.go:92: POST /hosts/h1/reveal-recovery-credential with Basic auth and no X-Felhom-Operator header: 404, want 403
|
||||
r135_csrf_test.go:92: POST /hosts/h1/request-logs with Basic auth and no X-Felhom-Operator header: 404, want 403
|
||||
r135_csrf_test.go:92: POST /customers/c1/block with Basic auth and no X-Felhom-Operator header: 404, want 403
|
||||
rc=1
|
||||
|
||||
== restored
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/web (cached)
|
||||
convicted routes: 39
|
||||
@@ -0,0 +1,50 @@
|
||||
# Hub state-changing routes and how each is protected (hub v0.135.0, R-135)
|
||||
|
||||
Every route below is reached through `RequireAuth` → `ServeHTTP`; the CSRF gate is the first thing `ServeHTTP` does for any
|
||||
method other than GET/HEAD/OPTIONS, before the route switch — so the protection is the same for every route, and an
|
||||
unknown path is refused by the gate before it can 404.
|
||||
|
||||
| Route (representative path) | Protected how | Test |
|
||||
|---|---|---|
|
||||
| `POST /configuration` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /apps/demo/reset-telemetry` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /apps/demo/dismiss-issues` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /offsite/endpoints` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /offsite/endpoints/1/delete` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /appliances/1/bind` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /appliances/1/discard` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /hosts/h1/delete` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /hosts/h1/reveal-recovery-credential` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /hosts/h1/request-logs` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /customers/c1/block` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /customers/c1/selfbind-link` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /customers/c1/unblock` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /customers/c1/geo/disable` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /customers/c1/floor` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /customers/c1/create-config` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /customers/c1/request-log-tail` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /configs/new` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /configuration/global-floor` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /configuration/artifacts` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /configuration/password` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /configs/c1/delete` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /configs/c1/edit` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /configs/c1/offsite-reissue` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /configs/c1/claim-resend` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /configs/c1/pbsdr-reissue` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /configs/c1/offsite-freeze` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /configs/c1/regen-password` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /configs/c1/reset` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /offsite/remove-unpinned/c1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /offsite/abandon-cancel/c1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /offsite/window-grant/c1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /offsite/windows-enabled` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /offsite/key-audit` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /os/ring/h1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /os/enabled/h1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /os/approve-now` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /os/approve-docker` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /no-such-route` | the gate, before routing (not a route) | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
|
||||
| `POST /login` | exempt (no session to ride; a wrong password is 401) | — |
|
||||
| `POST /bind/<token>` | exempt (public self-bind; the e-mailed URL token is the capability, rate-limited) | existing `selfbind_test.go` |
|
||||
| `GET` routes | not gated by design (a GET must not change state). **Not audited in this session** for a GET that writes — the route switch sends POST-only actions to handlers that check `MethodPost`, but the GET renderers were not read line by line | `TestR135_GetIsNotGated` |
|
||||
@@ -0,0 +1,6 @@
|
||||
== Part B live, 2026-10-05T09:16:00Z: live hub.db copied to scratch, only prefix + length selected, copy shredded after
|
||||
Tester-2-be8404|enc:v1:|87|2026-10-04 16:07:15
|
||||
demo-felhom-8363b5|enc:v1:|87|2026-07-18 16:30:41
|
||||
demo-hp-bb76ea|enc:v1:|87|2026-07-21 16:24:27
|
||||
tester-1-d70be4|enc:v1:|87|2026-10-04 19:40:27
|
||||
rows NOT sealed: 0
|
||||
@@ -0,0 +1,6 @@
|
||||
== Part B live reveal, 2026-10-05T09:16:15Z: POST /hosts/demo-hp-bb76ea/reveal-recovery-credential (Basic + X-Felhom-Operator), body to a 0600 scratch file, shredded after
|
||||
reveal HTTP 200
|
||||
username root@pam password length 32 set_at 2026-07-21T16:24:27Z
|
||||
the revealed password logs in to demo-hp's Proxmox API (POST /api2/json/access/ticket, root@pam): HTTP 200
|
||||
control, a wrong password: HTTP 401
|
||||
2026/10/05 11:16:15 [INFO] operator revealed break-glass console credential for host demo-hp-bb76ea (user=root@pam, secret 32 chars)
|
||||
@@ -0,0 +1,29 @@
|
||||
== RED-PROOF 1 (R-133): SaveHostRecoveryCredential stores the plaintext (pre-v0.135.0)
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/store [build failed]
|
||||
FAIL
|
||||
|
||||
== RED-PROOF 2 (R-133 wiring): main.go does not call SealLegacyRecoverySecrets
|
||||
=== RUN TestR133_MainSealsLegacyRecoverySecrets
|
||||
r133_wiring_test.go:28: cmd/hub/main.go never calls SealLegacyRecoverySecrets — legacy console passwords stay in plaintext
|
||||
--- FAIL: TestR133_MainSealsLegacyRecoverySecrets (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.028s
|
||||
FAIL
|
||||
|
||||
== restored
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/store 0.139s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.022s
|
||||
|
||||
== RED-PROOF 1 (re-run, compiling): SaveHostRecoveryCredential stores the plaintext (pre-v0.135.0)
|
||||
=== RUN TestR133_RawRowHoldsNoPassword
|
||||
r133_recovery_seal_test.go:29: raw host_recovery.secret is not sealed: "Console-Pw-7741"
|
||||
--- FAIL: TestR133_RawRowHoldsNoPassword (0.04s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/store 0.043s
|
||||
FAIL
|
||||
== restored
|
||||
--- PASS: TestR133_RawRowHoldsNoPassword (0.03s)
|
||||
--- PASS: TestR133_SealLegacyRecoverySecrets (0.03s)
|
||||
--- PASS: TestR133_WrongKeyFailsClosed (0.03s)
|
||||
--- PASS: TestR133_NoKeyRefusesToSave (0.03s)
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/store (cached)
|
||||
@@ -0,0 +1,42 @@
|
||||
== Part C readings on DooPlex, READ ONLY, 2026-10-05T09:17:24Z
|
||||
-- where the hub database lives
|
||||
hub-data pvc-486c9809-4672-4b56-b70e-0bf01d0c3628 1Gi longhorn
|
||||
-rw-r--r-- 1 root root 374534144 Oct 5 11:14 hub.db
|
||||
-rw-r--r-- 1 root root 32768 Oct 5 11:16 hub.db-shm
|
||||
-rw-r--r-- 1 root root 313152 Oct 5 11:16 hub.db-wal
|
||||
973.4M 373.4M 584.1M 39% /data
|
||||
-- the exclusion label: PVC (git, manifests/hub.yaml:47, commit 868e8465 2026-02-16 'updated hub yaml', no reason given) vs the live Longhorn Volume
|
||||
PVC label: disabled
|
||||
Volume label: enabled
|
||||
-- recurring jobs
|
||||
backup-daily backup 0 4 * * * 1 [default]
|
||||
backup-weekly backup 0 5 * * 0 1 [default]
|
||||
-- backups of the hub volume (Longhorn backupstore)
|
||||
2026-10-04T03:05:01Z Completed 708837376
|
||||
2026-10-05T02:06:26Z Completed 713031680
|
||||
-- backup target
|
||||
nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc?nfsOptions=soft,timeo=330,retrans=3 true
|
||||
-- which disk holds the target
|
||||
/dev/sda1
|
||||
/dev/sdb1
|
||||
-- DooPlex's own backup service
|
||||
Mon 2026-10-05 11:30:00 CEST 12min Mon 2026-10-05 11:15:00 CEST 2min 25s ago backup-freshness.timer backup-freshness.service
|
||||
Tue 2026-10-06 03:19:15 CEST 16h Mon 2026-10-05 03:15:21 CEST 8h ago dooplex-backup.timer dooplex-backup.service
|
||||
Result=success
|
||||
ExecMainStatus=0
|
||||
-- what tells anyone when a backup fails
|
||||
80:export NOTIFY_ON_FAILURE="true"
|
||||
81:# export NOTIFY_WEBHOOK_URL="https://your-webhook-url"
|
||||
137: if [ "${NOTIFY_ON_FAILURE}" = "true" ] && [ -n "${NOTIFY_WEBHOOK_URL}" ]; then
|
||||
140: "${NOTIFY_WEBHOOK_URL}" || true
|
||||
prometheus rule backup-freshness-alerts.yml MinecraftBackupStale
|
||||
prometheus rule backup-freshness-alerts.yml BackupFreshnessExporterDead
|
||||
prometheus rule longhorn-alerts.yml LonghornVolumeSpaceCritical
|
||||
prometheus rule longhorn-alerts.yml LonghornVolumeSpaceWarning
|
||||
prometheus rule longhorn-alerts.yml LonghornVolumeDegraded
|
||||
prometheus rule longhorn-alerts.yml LonghornNodeStoragePressure
|
||||
(no rule watches a Longhorn BACKUP's success or age, nor dooplex-backup.service; the only backup-freshness rule is MinecraftBackupStale)
|
||||
-- does anything leave DooPlex for the hub DB? the off-site route that exists today: ep0 PBS reached through felhom-ep0-pbs-tunnel (pull only, ep0 -> DooPlex)
|
||||
active
|
||||
/usr/bin/proxmox-backup-client
|
||||
/usr/bin/sqlite3
|
||||
@@ -0,0 +1,9 @@
|
||||
== Part D live: GET /system on hub 0.135.0 (Basic auth); extracted, no tokens
|
||||
Version floors: Version floors Global controller floor: 0.292.0 · vouched agent: 0.145.0 Customer Own floor Set Global floor moves it? demo-felhom Demo Ügyfél 0.295.0 unknown no — its own floor applies (at or above the global) demo-hp Demo HP 0.295.0 unknown no — its own floor applies (at or above the global) tester-1 Tester 1 0.295.0 unknown no — its own floor applies (at or above the global)
|
||||
Agent cell Tester-2-be8404: [('c-warn', '0.142.0 → 0.145.0 (since 2026-10-05)', '3 minor releases behind — sign an agent_update for this box')]
|
||||
Agent cell demo-felhom-8363b5: []
|
||||
Agent cell demo-hp-bb76ea: []
|
||||
Agent cell tester-1-d70be4: []
|
||||
<td title="current (vouched 0.145.0)">0.145.0</td>
|
||||
<td title="current (vouched 0.145.0)">0.145.0</td>
|
||||
<td title="current (vouched 0.145.0)">0.145.0</td>
|
||||
@@ -0,0 +1,31 @@
|
||||
== RED-PROOF 1 (R-530): alarm block 6 skipped (break out before any host)
|
||||
=== RUN TestAgentAlarm_AfterSevenDaysBehind
|
||||
r530_agent_alarm_test.go:43: the clock must start for the behind box only
|
||||
--- FAIL: TestAgentAlarm_AfterSevenDaysBehind (0.04s)
|
||||
=== RUN TestAgentAlarm_UnknownAndNothingVouchedSayNothing
|
||||
r530_agent_alarm_test.go:76: control: a box on 0.130.0 must start the clock
|
||||
--- FAIL: TestAgentAlarm_UnknownAndNothingVouchedSayNothing (0.04s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/osupdates 0.082s
|
||||
FAIL
|
||||
|
||||
== RED-PROOF 2 (R-604): handleSetGlobalFloor does not call reportFloorHeldBack
|
||||
=== RUN TestR604_GlobalRaiseNamesHeldBackBoxes
|
||||
r604_floor_held_back_test.go:64: no log line for the held-back box:
|
||||
--- FAIL: TestR604_GlobalRaiseNamesHeldBackBoxes (0.05s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.067s
|
||||
FAIL
|
||||
|
||||
== RED-PROOF 3 (R-604 wiring): main.go does not call SetEventEmitter
|
||||
=== RUN TestR604_MainWiresTheWebEventEmitter
|
||||
r133_wiring_test.go:49: cmd/hub/main.go never calls webServer.SetEventEmitter — the R-604 mail is never sent
|
||||
--- FAIL: TestR604_MainWiresTheWebEventEmitter (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.025s
|
||||
FAIL
|
||||
|
||||
== restored
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/osupdates 0.132s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.301s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.029s
|
||||
@@ -0,0 +1,28 @@
|
||||
== R-518 measure, demo-hp guest 9201, controller 0.295.0, 2026-10-05: POST /api/guest-backup/trigger (the button's call) at 09:19:05Z
|
||||
2026/10/05 09:19:07 backup_handlers.go:349: [INFO] [web] manual whole-guest backup triggered (quiesce loop)
|
||||
2026/10/05 09:19:07 quiesce.go:427: [INFO] [quiesce] manual backup requested — quiescing now
|
||||
2026/10/05 09:19:08 quiesce.go:517: [INFO] [quiesce] backup due on 2 tier(s) — quiescing 9 stack(s): [adventurelog bentopdf bookstack calibre-web docmost kimai opengist paperless-ngx privatebin]
|
||||
2026/10/05 09:19:29 quiesce.go:566: [INFO] [quiesce] tier local: backup job backup-9201-1791191969558324187 started — polling
|
||||
2026/10/05 09:23:44 quiesce.go:210: [INFO] [quiesce] a backup cycle is already running — skipping this scheduled check
|
||||
2026/10/05 09:24:09 quiesce.go:643: [INFO] [quiesce] tier local: backup job backup-9201-1791191969558324187 done — next tier may start (app still quiesced)
|
||||
2026/10/05 09:24:09 quiesce.go:307: [INFO] [quiesce] tier felhom-pbs is BUSY — the agent refused the backup because a concurrent heavy operation holds it. This is contention, NOT a failure: the tier stays due and retries in 15m0s (contended for 0s)
|
||||
2026/10/05 09:24:09 quiesce.go:504: [INFO] [quiesce] unquiescing (last tier is busy — deferring to a later cycle): restarting 9 stack(s)
|
||||
-- container StartedAt after the backup (the apps the quiesce stopped):
|
||||
2026-10-05T09:24:10.019684173Z adventurelog-postgres
|
||||
2026-10-05T09:24:10.200085281Z adventurelog-frontend
|
||||
2026-10-05T09:24:15.705236204Z adventurelog
|
||||
2026-10-05T09:24:16.331987016Z bentopdf
|
||||
2026-10-05T09:24:16.974996127Z bookstack-db
|
||||
2026-10-05T09:24:22.669660366Z bookstack
|
||||
2026-10-05T09:24:23.521923837Z calibre-web
|
||||
2026-10-05T09:24:24.437773353Z docmost-postgres
|
||||
2026-10-05T09:24:24.631250114Z docmost-redis
|
||||
2026-10-05T09:24:35.395095325Z docmost
|
||||
2026-10-05T09:24:36.934564476Z kimai-db
|
||||
2026-10-05T09:24:42.753714563Z kimai
|
||||
2026-10-05T09:24:43.466405175Z opengist
|
||||
2026-10-05T09:24:44.478064534Z paperless-redis
|
||||
2026-10-05T09:24:44.688668856Z paperless-postgres
|
||||
2026-10-05T09:24:54.928089575Z paperless-webserver
|
||||
2026-10-05T09:24:55.739635651Z privatebin
|
||||
RESULT: stop began 09:19:08Z (9 stacks quiesced), local tier 09:19:29-09:24:09, PBS tier BUSY (skipped, retried later), last app back 09:24:55Z => longest stop 5 min 47 s, shortest ~5 min 02 s. 21/21 containers running afterwards.
|
||||
@@ -0,0 +1,15 @@
|
||||
09:19:20 running=15 phase=idle
|
||||
09:19:41 running=4 phase=snapshotted
|
||||
09:20:02 running=4 phase=snapshotted
|
||||
09:20:23 running=4 phase=snapshotted
|
||||
09:20:44 running=4 phase=snapshotted
|
||||
09:21:05 running=4 phase=snapshotted
|
||||
09:21:26 running=4 phase=snapshotted
|
||||
09:21:47 running=4 phase=snapshotted
|
||||
09:22:08 running=4 phase=snapshotted
|
||||
09:22:29 running=4 phase=snapshotted
|
||||
09:22:50 running=4 phase=snapshotted
|
||||
09:23:13 running=4 phase=snapshotted
|
||||
09:23:34 running=4 phase=snapshotted
|
||||
09:23:55 running=4 phase=snapshotted
|
||||
09:24:16 running=6 phase=done
|
||||
@@ -0,0 +1,39 @@
|
||||
== RED-PROOF 1 (R-519): runDBDumpsInternal does not call markRunStarted
|
||||
=== RUN TestRunRecord_TheRealRunIsOnRecordWhileItRuns
|
||||
run_record_test.go:78: the run was not on record while it ran — a cut here would go unnoticed
|
||||
--- FAIL: TestRunRecord_TheRealRunIsOnRecordWhileItRuns (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s
|
||||
FAIL
|
||||
|
||||
== RED-PROOF 2 (R-519): the synthesised status says Success: true again
|
||||
=== RUN TestRunRecord_SynthesisedStatusIsNotOKAfterACut
|
||||
run_record_test.go:97: after a cut the synthesised status still reads OK: &{LastRun:2026-10-05 11:34:19.454323277 +0200 CEST m=+0.001336319 Results:[{DB:{ContainerName:adventurelog ContainerID: DBType: DBUser: DBName: StackName:adventurelog} FilePath:adventurelog-postgres.sql Size:0 Duration:0s Error:<nil> Validation:{Valid:false TableCount:0 Error: FileSize:0 ModTime:0001-01-01 00:00:00 +0000 UTC UserTableFound:false UserRows:0 LooksEmpty:false}}] Success:true Duration:0s}
|
||||
--- FAIL: TestRunRecord_SynthesisedStatusIsNotOKAfterACut (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s
|
||||
FAIL
|
||||
|
||||
== RED-PROOF 3 (R-519 wiring): main() does not call loadRunRecordAtStartup
|
||||
=== RUN TestMainWiresRunRecord
|
||||
run_record_wiring_test.go:47: main() never calls loadRunRecordAtStartup — a cut run is never said on the page
|
||||
--- FAIL: TestMainWiresRunRecord (0.01s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.016s
|
||||
FAIL
|
||||
|
||||
== RED-PROOF 4 (R-518): the v0.267.0 copy (12 apps / 8 minutes only)
|
||||
=== RUN TestR518_BackupButtonStatesTheMeasuredDowntime
|
||||
r518_backup_downtime_copy_test.go:33: hu: "kb. 6 perc" appears 0 times, want it on the page AND in the confirm
|
||||
r518_backup_downtime_copy_test.go:33: hu: "nem másodperceket" appears 0 times, want it on the page AND in the confirm
|
||||
r518_backup_downtime_copy_test.go:33: en: "about 6 minutes" appears 0 times, want it on the page AND in the confirm
|
||||
r518_backup_downtime_copy_test.go:33: en: "not seconds" appears 0 times, want it on the page AND in the confirm
|
||||
--- FAIL: TestR518_BackupButtonStatesTheMeasuredDowntime (0.07s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.080s
|
||||
FAIL
|
||||
|
||||
== restored
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.013s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.014s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.080s
|
||||
@@ -0,0 +1,7 @@
|
||||
== 9202 teardown of the throwaway bookstack (installed 09:35:56Z for the R-519 reproduction), 2026-10-05T11:44:31Z
|
||||
stop: HTTP 200
|
||||
remove: HTTP 200
|
||||
containers: 0
|
||||
volumes: 0
|
||||
stackdir: none
|
||||
backups: none
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user