Interim: R-528, R-518, R-892 updated; section counts; STATUS and REPORT (137 -> 138; Parts B and D after 08:30)
gates / gates (push) Successful in 2m49s
gates / gates (push) Successful in 2m49s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,122 +1,81 @@
|
||||
# REPORT — the design build: restore order, out-of-memory alarm, one stop per tier, the family window, wger's files (2026-10-06 evening)
|
||||
# REPORT — night read-back brief, the night-free parts (2026-10-06 evening; Parts B and D follow after 08:30 on 2026-10-07)
|
||||
|
||||
The brief said „start after 08:30 on 2026-10-07". The operator then said „start A, C, E, F now". Parts B (the night
|
||||
read-back) and D (one press, outside 02:00–08:30) wait for tomorrow. **Nothing was delivered to a box tonight**, so the
|
||||
night of 2026-10-06→07 runs controller v0.301.0 and agent v0.149.0, and Part B reads a clean result. The hub was
|
||||
released (it is not a box).
|
||||
|
||||
| Part | Result |
|
||||
|---|---|
|
||||
| **A** — housekeeping | **done** — decisions 151–156 recorded first; R-469 closed by ruling; the agent repo has the shared rule file (identical: `diff` against the controller's copy empty, one md5 across five copies); the Tester 1 walk route is NOT a 30-minute fix → R-892 |
|
||||
| **B** — restore order (R-638) | **done** — measured first on 9202: both main paths SAFE (docmost, romm); two side paths fixed by order, red-proved; exposure 3 not fixable by order → known limit in `07` §6.3 + R-893; R-638 closed |
|
||||
| **C** — out-of-memory alarm (R-528) | **STOPPED by its measurement** — Docker 29.8.2 set `OOMKilled` correctly in all four shapes; nothing built; a decision for you |
|
||||
| **D** — one stop per backup tier (R-518) | **built and delivered, not shown live** — measured first (read-only): the off-site part of a stop is ~2 s; red-proved; the live button test was not possible on 9202 (no agent connection) and the demo boxes take deliveries only → R-518 open for the first night read-back |
|
||||
| **E** — reopen a sign-up for the family window (R-717) | **done for wishlist, proven live; opengist stopped** (no tool in its container) → R-717 narrowed |
|
||||
| **F** — wger's file server (R-762) | **done and proven on both venues; wger stays hidden** — R-762 narrowed to the `runserver` half it carries from R-755 |
|
||||
| **A** — the rulings | **done** — `09` §3 decisions 157 (R-528 option A + the Docker-approval check) and 158 (R-892: VM 341 on the HP box); R-892's "no Proxmox host known" corrected |
|
||||
| **B** — the night read-back (R-518) | **waits for 08:30 2026-10-07** |
|
||||
| **C** — demo-hp's off-site copy | **diagnosed: the premise was wrong — no 10-day gap.** Last copy 2026-10-01 20:15Z (ep0's own listing, verify ok); every box within 7 days. Two real findings: an off-site tier read DUE after an agent restart while its storage was unreachable (R-894, filed), and the alarm's operator mail failed and was never retried (**fixed**, hub v0.140.0) |
|
||||
| **D** — one press, measured | **waits for after 08:30 2026-10-07** |
|
||||
| **E** — the Tester 1 box in the update test (R-892) | **route built and identity matched; the live proof is BLOCKED** — DooPlex's key is not authorized on VM 341, and the permission check refused fetching its vaulted password (the brief's rule: a refusal stops that item) |
|
||||
| **F** — the Docker-approval memory-kill check (R-528) | **built**: hub v0.140.0 LIVE (the approval waits for a passing check); the agent's wrapper merged, unreleased (ships as v0.150.0 after Part B); proven by hand on demo-hp's guest |
|
||||
| small check — wger's 100 % peak | file cache, not a kill (kill counter 0, restarts 0, anon 50.3 %); the memory hint stays |
|
||||
|
||||
| Rows before | Rows after | Opened | Closed |
|
||||
|---|---|---|---|
|
||||
| **137** | **137** | **2** (R-892, R-893) | **2** (R-469, R-638) |
|
||||
| **137** | **138** | **1** (R-894) | **0** (so far) |
|
||||
|
||||
## Baselines and rulings
|
||||
## Part C — what happened on demo-hp, with times
|
||||
|
||||
Verified at the start: felhom.eu `01d5dbd3e0`, controller `31242786af` (v0.300.0), agent `3e8ebeb96c` (v0.149.0), catalog
|
||||
`ecb8552ee9`, register 137. Read: the three designs in `audits/night-burndown-2026-10-05/`, rows R-638, R-528, R-518, R-717,
|
||||
R-762, R-469, `07` §6, `08` (the OOM rung), `09` §3 decisions 47/49. Rulings recorded first as `09` §3 151–156 (ruling 6 is
|
||||
carried as A by the brief; the operator's chat answer was not visible to this session).
|
||||
- **My last report was wrong**: it said „demo-hp had no successful off-site run in 10 days". I read a log view cut by
|
||||
`tail -40` that started on 2026-10-04. The agent's full log and ep0's own listing both show a completed off-site copy
|
||||
on **2026-10-01 20:15Z** (15.2 GB, verify `ok`). With the 7-day cadence the next one is due ~2026-10-08.
|
||||
- ep0, read only (`C/C2-ep0-snapshot-listing.txt`): demo-hp 2026-10-01T20:15Z, demo-felhom 2026-10-06T04:21Z, tester-1
|
||||
2026-10-04T19:57Z, Tester-2 2026-10-04T16:31Z — every box within 7 days.
|
||||
- **2026-10-05 04:25Z (06:25 local), demo-hp:** the off-site storage answered *Can't connect to 10.77.0.1:8007*; the
|
||||
agent had restarted at 02:57Z (04:57 local), 1.5 hours before; its per-tier backup record is in memory only
|
||||
(`internal/backup/store.go`, R-348), so the due-check fell back to an EMPTY record and read the tier DUE although it
|
||||
was not (`internal/localapi/server.go` `newestArchiveOn` → unknown → in-memory). The controller asked; vzdump failed
|
||||
(*could not activate storage 'felhom-pbs'*). By design the controller stops the apps before it asks; whether it did
|
||||
that night is **not known** (the controller's log was lost to a later restart). → **R-894**.
|
||||
- **Who was told:** the hub recorded `whole_guest_backup_failed` (error) at 04:27Z; the household was not mailed (by
|
||||
design: operator only); **the operator's mail FAILED** (Resend: *context deadline exceeded*) and was never retried.
|
||||
It is the only failed operator mail of 692 since 2026-02-16 (`C/C4-hub-failed-mails.txt`).
|
||||
- **Fixed** (hub v0.140.0): a failed operator mail is tried again after 1, 5 and 15 minutes; each try is a row; giving
|
||||
up is an ERROR line. Red-proof `C/red-operator-mail-retry.txt` („the failed mail was never sent again (tries=1)").
|
||||
- Not read: the household's backups page on demo-hp (no session tonight).
|
||||
|
||||
## Part B — R-638
|
||||
## Part E — the Tester 1 box
|
||||
|
||||
**Slice 0, measurement only** (scratch 9202, controller 0.299.0, drill catalog; `audits/design-build-2026-10-06/B/`).
|
||||
Install the OLD definition → seed → the guarded Update migrates (its backing-up phase makes the copy at the old version) →
|
||||
the household's restore of that copy (`POST /backup/restore`, the only copy offered: "helyi").
|
||||
- Identity (`E1-identity-match.txt`): VM 341 `night1004-tester1` on demo-hp, at 192.168.0.154 (found by its MAC);
|
||||
its certificate `CN=felhom.enkicsifelhom.hu`; the agent's own hub report for `tester-1-d70be4` says `host.node=felhom`;
|
||||
its guest at 192.168.0.101 answers `felhom.enkicsifelhom.hu` with the Felhom login (200) and demo-hp's domain 404.
|
||||
- Built (catalog `d63ea35`): `box_walk.py` `TARGETS` (9202, 9201, tester-1 via `-J demo-hp`), `BOX_ADMIN_SEED_GUESTS`
|
||||
adds `("tester-1", "9201")`; tests BoxWalkTargets (red-proved). `operations/nodes.md` has the box.
|
||||
- **Blocked:** `ssh -J demo-hp root@192.168.0.154` → `Permission denied (publickey,password)`. The VM has no guest agent;
|
||||
editing its disk needs a VM stop (a reboot — not allowed). I tried to read the hub's code for revealing the vaulted
|
||||
console password; **the session's permission check refused it**, and I did not try another way. R-892 now asks the
|
||||
operator.
|
||||
|
||||
| app | migration | after restore | `Imported DB dump` | seed | the copy's own dump (control) |
|
||||
|---|---|---|---|---|---|
|
||||
| docmost 0.95.0 → 0.96.0 (pg16) | 42 → 48 tables (+6 `oauth_*`, `public_spaces`, `siem_destinations`) | 42, none of the 6 | yes | read back | 42, equal |
|
||||
| romm 5.0.0 → 5.3.0 (mariadb 11.4) | 27 → 39 (+12) | 27, none of the 12 | yes | read back | 27, equal (views counted) |
|
||||
## Part F — the memory-kill check
|
||||
|
||||
Both went back to the old pinned version and ran. **Said plainly:** romm needed four runs — run 1 refused the install
|
||||
(its data path must be a drive), run 2 hit an empty drill commit, run 3 measured correctly but my control read the wrong
|
||||
file (a `pre-restore-*` dump) and missed three VIEWS; run 4 fixed both. docmost's run 1 installed the live version
|
||||
because the box had not read the drill yet (the script then waited for the old reference).
|
||||
- Hub v0.140.0 (LIVE 17:34Z): `DockerStatus` needs, per ring-0 box, a passing `oom_check` with the set; failed, errored or
|
||||
missing blocks. Red-proof `F/red-hub-docker-approval.txt`; the System page hides the button and says why.
|
||||
- Agent (`acccb66`, unreleased): after a Docker step the wrapper runs a throwaway container from the running
|
||||
controller's image (`--pull never`, `--network none`, 64 MB cap, one 200 MB block); pass = `OOMKilled=true` AND the
|
||||
`oom` event; always removed. 12 red-proofs in `F/`. The helper also found 11 wrapper tests that never ran (a
|
||||
`unittest.main()` mid-file) — moved; all pass.
|
||||
- **By hand on demo-hp's guest** (`F/F1`, `F/F2`): `OOMKilled=true`, exit 137, 21 → 21 containers, none left. The `oom`
|
||||
event was MISSING when the events window ended in the same second as the run, and present (`create attach start
|
||||
oom die`) with the window ending a second later — the wrapper waits 2 s and reads to epoch + 1.
|
||||
- Until agent v0.150.0 reaches the ring-0 boxes, no Docker set can be approved (none is pending).
|
||||
|
||||
**Slices 1 and 2** (controller `9c94568`): `RestoreApp` starts only the database services, replays, then the app; a dump
|
||||
with no identifiable database service is refused before anything is touched; a failed volume step skips the replay
|
||||
(fallback and unit restore). Red: `B/red-slice1-fallback-order.txt` ("the FULL stack was already up when the replay fired
|
||||
(calls before the replay: stop,start)"), `B/red-slice2-no-replay-after-volume-failure.txt`. **Exposure 3** (the off-site
|
||||
rollback loads the newer pre-restore dump over the older volume; the older definition is not put back) cannot be fixed by
|
||||
order — known limit in `07` §6.3, row R-893.
|
||||
## Instruction-file edits
|
||||
|
||||
## Part C — R-528, stopped
|
||||
None tonight.
|
||||
|
||||
`C/C0-oom-signal-measure-9202.txt`: four runs in a throwaway alpine container under a cap. Runs 1–3: the kernel killed the
|
||||
hog AND the main process (the box cannot protect it: `--oom-score-adj -1000` left pid 1 at 0) — exit 137,
|
||||
`OOMKilled=true`. Run 4 (256 MB, `sleep infinity`, one 400 MB block): the container kept running, `oom_kill` 0 → 3,
|
||||
`OOMKilled=true`, three `oom` events. The design's premise — the flag stays false while the counter rises — did not
|
||||
reproduce on Docker 29.8.2 (the false flags of 2026-09-15 were on 29.8.0). By the brief's measure-first rule, nothing was
|
||||
built. Cost read for the record: one exec read of demo-hp's 21 containers 1.6 s wall (`C/C1-…`). `08` records it.
|
||||
## CI
|
||||
|
||||
## Part D — R-518
|
||||
felhom.eu `e0bdd52` → 1448 success (and `eed1dbd` 1446, `01c4a5d` 1447); catalog `d63ea35` → 1445; agent `7e82f32` → 1449;
|
||||
this commit checked after its push.
|
||||
|
||||
**Measured first, read-only:** demo-felhom's night off-site job (2026-10-06): started 06:21:08, `snapshotted` 06:21:10,
|
||||
its one app running at 06:21:18. demo-hp had no successful off-site run in 10 days. A new off-site run from 9202 was not
|
||||
possible (it would write to ep0). **Built** (helper, then corrected by me): one window per tier, resume at that tier's
|
||||
`snapshotted`, the next tier in a later cycle; the button makes the LOCAL copy only. **My correction:** the helper let the
|
||||
button back up the off-site tier when the local storage is absent; the ruling says local only, so it now backs up
|
||||
nothing and stops nothing (red: `D/red-press-no-local-backs-up-nothing.txt`); an agent that flags no tier primary keeps
|
||||
the old first tier (the breaker test's fake agent showed that case). Red: `D/red-first-tier-resume.txt` (restarts
|
||||
sampled at each poll `[0 0 0 0]`), `D/red-manual-local-only.txt`. Page text, hu + en: about 1–1.5 minutes — an estimate
|
||||
from measured parts, not a measured press. **Not shown live** (see the Part table).
|
||||
## What is left for 2026-10-07 after 08:30
|
||||
|
||||
## Part E — R-717
|
||||
|
||||
Controller (`5b86569`): `after_setup` `open_command` / `open_success`; the window marks `opening` on disk first; a failed
|
||||
open closes again and the page says so; the close runs at the window's end, at every start and after a successful
|
||||
update; a failed close retries every 2 min. Six red-proofs in `E/`. Catalog (`a5a5b51`, pushed AFTER the controller
|
||||
reached the three boxes): wishlist closes/opens `system_config.enableSignup` (group `global`) with Node's own sqlite
|
||||
module, `min_controller: "0.301.0"`. **Live on 9202** (the 0.301.0 test image, `E/live.txt`): after setup `false`, a
|
||||
stranger straight at the app 401 "invite only", users 1; window `true`, a family member 200, users 2; window end 13:47:34Z,
|
||||
close 13:47:44Z, `false`, a stranger 401, users 2. **Opengist:** no sqlite tool, no script runtime, and its CLI has only
|
||||
`create-user`, `reset-password`, `toggle-admin` — not doable with this mechanism.
|
||||
|
||||
## Part F — R-762
|
||||
|
||||
Upstream's production compose read first (its `nginx` serves `/static/` and `/media/` from the shared volumes;
|
||||
`prod.env` sets `DJANGO_DEBUG=False`). Here traefik sends only those two paths to a `wger-files` nginx
|
||||
(`nginx:1.30.5-alpine`, 32 MB, stock config, volumes read-only) — every box gate wraps every router, so it is gated
|
||||
like wger. The bench tool could not express a step that adds a service, so I added `--move-to` and
|
||||
`--write-ladder --to-definition` (catalog `c48a0db`, red-proved). **Bench:** proven, wger anon peak 50.3 %, wger-files
|
||||
15.9 %, the page's hashed CSS 200 from wger-files and 404 straight at wger. **9202:** the guarded Update done in 59.4 s,
|
||||
the seed read back, a photo posted before the step 404 → 200 / 178 B / image/png after, CSS 404 → 200, wger-files
|
||||
7.6 MiB. **Said plainly:** my box script crashed after the step when it decoded the PNG as text; the after-checks were
|
||||
re-read binary-safe and the box verdict was written by hand from them (it says so). **Ready to show?** Files and sign-up
|
||||
yes; it still runs Django's `runserver` (upstream's gunicorn runs 3 workers, which do not fit 384 MB) — not ready. Cost:
|
||||
283 MB of collected static files in a named volume, in every wger backup (the checks refuse an unnamed volume).
|
||||
|
||||
## Release and delivery
|
||||
|
||||
Controller **v0.301.0** (`0b1b8d3`, image `sha256:0e80f6af…`), MinAgent 0.131.0. Golden **0.301.0**
|
||||
(`GOLDEN_SHA256=96e94fed…`, Docker pinned to the approved set, registry round trip equal, token grep 0 with a working
|
||||
control). Hub: vouched agent 0.149.0 / golden 0.301.0 / min_agent 0.131.0; floors 0.301.0 for demo-hp, demo-felhom,
|
||||
tester-1 — first saved without a declared MinAgent (allowed: floor = golden), then re-saved with 0.131.0. demo-hp and
|
||||
demo-felhom run 0.301.0 (healthy); tester-1's hub page "Controller elindult (0.301.0)". Global floor and Tester 2 not
|
||||
touched. No agent release.
|
||||
|
||||
## Instruction-file edits (decision 150)
|
||||
|
||||
- `.claude/rules/unprompted-work.md`, all five copies, line 8–9: before „…in `felhom.eu`, `felhom-controller` and
|
||||
`app-catalog-felhom.eu` `.claude/rules/`, and in the workspace root's unversioned `.claude/rules/`; change all four or
|
||||
none." → after: names `felhom-agent` too, „change all five or none". Why: decision 152 added the agent's copy.
|
||||
- `felhom-agent/.claude/rules/unprompted-work.md`: new file (decision 152), identical to the other copies.
|
||||
|
||||
## Fixed without a row
|
||||
|
||||
- `OPEN-ITEMS.md` section headings: every count was stale (e.g. „Install & onboarding — 11 rows" over 5); recomputed.
|
||||
|
||||
## CI, last commit of every repo
|
||||
|
||||
controller `0b1b8d3` → 1439 success; catalog `cf1ed43` → 1438 success (`a5a5b51` checked after its push); agent
|
||||
`de812bc` (checked after its push); felhom.eu `8db4425` → 1440 (checked), and this commit (checked after the push).
|
||||
|
||||
## Teardown
|
||||
|
||||
Machine: 9202 back on the live catalog (`controller.yaml` byte-identical), on the released 0.301.0 (was 0.299.0), every
|
||||
test app removed through the product (no volume left), the test images removed by name, the same six containers as at
|
||||
the start. Bench 9401 stopped. Drill VM: build guest destroyed, files shredded, off, `virgin`. Drill catalog = live.
|
||||
Host: nothing else. Hub: the vouch and three floors; two read-only DB copies, deleted.
|
||||
Part B (read both demo boxes' night under v0.301.0), Part D (one press on a demo box, measured; the page text), the
|
||||
agent v0.150.0 release + bundle and its delivery, the controller release for Part D's text, and Part E's live proof if
|
||||
the operator authorizes the key. Teardown tonight: none needed (no box provisioned; the hub DB copies deleted; the
|
||||
by-hand check containers removed — 0 left).
|
||||
|
||||
@@ -2,8 +2,24 @@
|
||||
|
||||
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.**
|
||||
|
||||
**Updated 2026-10-06 16:30: demo-hp, demo-felhom and Tester 1 run controller 0.301.0 (hub 0.139.0, agent 0.149.0).
|
||||
The open-items list is at 137. Report: `REPORT.md`.**
|
||||
**Updated 2026-10-06 19:45: hub 0.140.0; demo-hp, demo-felhom and Tester 1 run controller 0.301.0, agent 0.149.0.
|
||||
The open-items list is at 138. Report: `REPORT.md` (interim — the night read-back follows after 08:30).**
|
||||
|
||||
## Night (2026-10-06): the night-free parts done; the rest after 08:30
|
||||
|
||||
- **demo-hp's off-site copy is NOT 10 days old** — my last report was wrong (I read a cut log). Its last off-site copy
|
||||
is from 2026-10-01; every box is within its 7 days.
|
||||
- **One alarm was lost:** a failed off-site try on 2026-10-05 reached the hub, but its mail to you failed and was never
|
||||
sent again. **Fixed:** the hub now retries a failed mail. The try itself should not have happened (a design question,
|
||||
on the list).
|
||||
- **The Docker update now needs a memory-kill proof** before you can approve a new set (your answer). The hub part is
|
||||
live; the box part ships tomorrow after the night read-back.
|
||||
- **The Tester 1 box:** found and matched (the VM on the HP box), and the test tool knows the route. It cannot log in
|
||||
yet: my key is not on that box. See below.
|
||||
|
||||
**Needs you (none urgent):**
|
||||
1. **R-892** — put DooPlex's SSH key on the Tester 1 box (one line, from its console), or allow me to read its
|
||||
vaulted password. Then the update test runs there. If nothing: it stays on the scratch box only.
|
||||
|
||||
## Evening (2026-10-06): the designs built; the list at 137
|
||||
|
||||
|
||||
@@ -146,7 +146,7 @@ stopping line that lies.
|
||||
| **R-683** | App updates | P3 | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | — | — | CC |
|
||||
| **R-785** | App updates | P3 | **[P3-LOW] SparkyFitness is pinned 11 releases and a major behind upstream (v0.17.3; upstream v1.7.3, v1.6.0 dated 2026-07-24).** READ 2026-10-01 (`audits/visitors-2026-10-01/C/bench/C1-previous-tag.txt`). **Needs:** an update walk 0.17 → 1.x through the ladder (bench + box), after R-784 is decided. | **OPEN — rank P3-LOW; owner: CC (after R-784)** | — | — | CC |
|
||||
|
||||
## Backup & restore — 32 rows (P2 7, P3 11, P4 14)
|
||||
## Backup & restore — 33 rows (P2 7, P3 12, P4 14)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
@@ -155,7 +155,7 @@ stopping line that lies.
|
||||
| **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-05 — owner Viktor.** (b) partly: the hub database now leaves DooPlex nightly, encrypted, to ep0 (R-173); everything else in DooPlex's backup still stays on the box. (a) partly: the hub copy alarms through Prometheus (`HubDBBackupStale`); `notify_failure` is still a no-op for the rest. (c)–(h) unchanged. **READY** for the rest | — | — | operator |
|
||||
| **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY (L) — NEW 2026-08-12, RANK 1** | R-198, R-199, R-224, R-241 | Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. **Until one of those, the honest position is that retention is an operator-only capability.** At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked | operator + CC |
|
||||
| **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC |
|
||||
| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. | — | Read back the first night runs on demo-hp and demo-felhom under v0.301.0 (the agent's `snapshotted` line, the apps' StartedAt, the controller's window lines); close with that evidence. If nothing: the row stays open | CC |
|
||||
| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. **2026-10-06 (night), from Part C:** demo-hp's off-site tier was NOT overdue — its last copy is 2026-10-01 20:15Z (ep0's listing, verify ok), so with the 7-day cadence it is due ~2026-10-08; the night of 2026-10-06→07 is most likely local-only on both demo boxes (demo-felhom's off-site landed 2026-10-06 04:21Z). The two-tier night under the new rule is then ~2026-10-08 on demo-hp. | — | Read back the 2026-10-07 night (local tier) and the ~2026-10-08 night (both tiers on demo-hp); measure one press (Part D) | CC |
|
||||
| **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. | **OPEN — filed 2026-10-06** | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC |
|
||||
| **R-894** | Backup & restore | P3 | **After an agent restart, an UNREADABLE off-site storage makes the off-site tier look DUE, so the box asks for a copy that cannot be made.** MEASURED 2026-10-05 on demo-hp, read 2026-10-06 (`audits/readback-2026-10-07/C/`): the last off-site copy was 2026-10-01 20:15Z (ep0's own listing, verify `ok`), so the 7-day tier was NOT due; the agent had restarted at 04:57 local; at 06:25 `GET …/storage/felhom-pbs/content` answered 500 *Can't connect to 10.77.0.1:8007*; `newestArchiveOn` returned `unknown` and fell back to the in-memory record (`internal/localapi/server.go`), which a restart empties (`internal/backup/store.go` — memory only, R-348); so the tier read DUE, the controller requested it, and vzdump failed (*could not activate storage 'felhom-pbs'*). By design the controller stops the apps before it asks (`07` §6.4); whether it did that night is NOT KNOWN (the controller's log was lost to a later restart). The hub got `whole_guest_backup_failed` (error); its operator mail then failed (fixed in the hub this session: a failed operator mail is retried). The code's rule is deliberate: an unreadable storage must not suppress a backup. | **OPEN — filed 2026-10-06** | a design: keep the newest success per tier on disk (as `RestoreTestState` does) so the fallback is the last known copy, not "never"; or report an unreachable storage so no app is stopped for it | Decide the fallback; build it in the agent with a test that restarts the agent and then cannot read the storage; measure once | CC |
|
||||
| **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by `MaxRemove` and the hub's count check, not prevented.** | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC |
|
||||
@@ -233,7 +233,7 @@ stopping line that lies.
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **R-243** | Monitoring & notifications | P2 | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. **⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces.** 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted `offsite_delivery_stuck` (**warning**) and wrote an **operator-channel** `notification_log` row recording `offsite_credential_restaged` / status **REFUSED** with an accurate reason — *"the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193"*. So on the **regressed-apply** shape the operator IS told, promptly and correctly, and this row's *"skips the applied shape"* does not apply. The gap stands for a box that reaches the held state **without** a prior working tier in its report history. **Recorded so the row is not read wider than it measures.** | **READY** — owner Viktor | — | — | operator |
|
||||
| **R-528** | Monitoring & notifications | P2 | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** **2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely.** On `tester-1-022354` (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed `app_oom` (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was `write CONNECTION_CLOSED immich-postgres:5432` and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: `audits/evidence-chaos-night-2026-09-17/round-2.txt`. **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-528.md`** — for the operator. **2026-10-06: STOPPED BEFORE ANY CODE — the measurement contradicted the design** (decision 155 said A then C). On scratch 9202, Docker 29.8.2, the `OOMKilled` flag was TRUE in all four shapes tried: a child process killed while the container kept running (`oom_kill` 0 → 3, flag true) and three main-process kills (exit 137, flag true); an `oom` event each time. The false flags of 2026-09-15 were on Docker 29.8.0. Cost read for the record: one exec read of 21 containers on demo-hp = 1.6 s. `audits/design-build-2026-10-06/`C/. **2026-10-06 18:24: operator ruling, option A (decision 157):** nothing is built; the row stays open as a WATCH for a box whose Docker reports a false flag. **Added:** a Docker engine set may not be approved until it reports a memory kill correctly on the boxes that ran it (built in this row's next session). | — | Watch; build the Docker-approval memory-kill check (decision 157) | CC |
|
||||
| **R-528** | Monitoring & notifications | P2 | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** **2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely.** On `tester-1-022354` (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed `app_oom` (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was `write CONNECTION_CLOSED immich-postgres:5432` and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: `audits/evidence-chaos-night-2026-09-17/round-2.txt`. **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-528.md`** — for the operator. **2026-10-06: STOPPED BEFORE ANY CODE — the measurement contradicted the design** (decision 155 said A then C). On scratch 9202, Docker 29.8.2, the `OOMKilled` flag was TRUE in all four shapes tried: a child process killed while the container kept running (`oom_kill` 0 → 3, flag true) and three main-process kills (exit 137, flag true); an `oom` event each time. The false flags of 2026-09-15 were on Docker 29.8.0. Cost read for the record: one exec read of 21 containers on demo-hp = 1.6 s. `audits/design-build-2026-10-06/`C/. **2026-10-06 18:24: operator ruling, option A (decision 157):** nothing is built; the row stays open as a WATCH for a box whose Docker reports a false flag. **Added:** a Docker engine set may not be approved until it reports a memory kill correctly on the boxes that ran it (built in this row's next session). **2026-10-06 (night): the Docker-approval memory-kill check is BUILT** — hub v0.140.0 (LIVE: the approval waits for a passing `oom_check` on every ring-0 box; a failed, errored or missing one blocks) and the agent's wrapper (`felhom-agent` `acccb66`, unreleased — ships as v0.150.0 after the 2026-10-07 read-back). Proven by hand on demo-hp's guest: `OOMKilled=true`, exit 137, the `oom` event — seen only with an events window ending after the run (the wrapper waits 2 s). `audits/readback-2026-10-07/`F/. | — | Release agent v0.150.0 + bundle and deliver it; then the row is a WATCH only | CC |
|
||||
| **R-79** | Monitoring & notifications | P3 | **`report.Issues` / `report.Warnings` are English on customer-facing surfaces** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN** — the issues and warnings travel to the hub as sentences; changing them is the two-repo spike the row itself names. | — | **Whole-surface, not a one-off** (DIAG §6): every producer is English — `"SSD/HDD disk usage critical"`, `"Docker: %v"`, `"Protected container not running: %s"`, and all six `Warnings` strings. They render on the customer's Hungarian dashboard, and the `health_critical` path has reached the **customer** email channel three times historically. Deliberately NOT bundled into R-77: a copy sweep across every producer would have buried two safety fixes in string churn, and the seam is not obvious — translate at the producer, or at the render/notification boundary where operator-English and customer-Hungarian already diverge? Pick the seam in a spike; the strings are mechanical after. | CC |
|
||||
| **R-211** | Monitoring & notifications | P3 | **Prometheus has no config-reloader — a rules change reaches the pod and is never read** | **READY (S) — NEW 2026-08-05** | — | Found while verifying R-205 rather than by looking for it. The `mon-system/prometheus` Deployment runs **one** container (`prom/prometheus:v3.12.0`) with **no `configmap-reload`/`prometheus-config-reloader` sidecar**. After the ArgoCD sync the updated `node-housekeeping-alerts.yml` was present **inside the pod** (`grep -c "and on(instance)"` → 3 on the mounted symlink) while the Prometheus **rules API still served the old expression** — for **4+ minutes**, with no error anywhere. It only took effect after an explicit `POST /-/reload`. **The consequence is general, not specific to R-205: every rule edit in this repo since the stack was built has silently not applied until something happened to restart the pod** — so "committed and synced" has never meant "in force", and ArgoCD reporting `Synced/Healthy` is true and beside the point. `--web.enable-lifecycle` IS already set, so the fix is small: add a reloader sidecar watching the ConfigMap, or a `checksum/config` pod annotation so a rules change rolls the pod. **Same class as the four *built-but-never-wired* seams** — the control exists, nothing walks it | CC |
|
||||
| **R-333** | Monitoring & notifications | P3 | **Two disk-health questions the deploy raised and did NOT act on.** **(a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe.** They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but **measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba**, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as **Hiba** — the single worst outcome this feature can produce. **(b) The agent runs bare `smartctl -a -j` with no `-n standby`** (`felhom-agent/internal/storage/hostops.go:368`), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only **3375 load cycles in 60505 hours** (~one per 18h), i.e. that duty cycle barely spins down at all | **READY (S each) — NEW 2026-08-14** | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on `critical_warning`; (b) add `-n standby` to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) |
|
||||
@@ -283,7 +283,7 @@ stopping line that lies.
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** | — | — | CC |
|
||||
| **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** **RE-DIAGNOSED 2026-10-05 (round 2), from the logs that survive:** (1) **„lost in a runner restart" does NOT fit** — the runner pod last restarted 13:24:42Z (`restartCount 5`, all around the 13:20Z Longhorn restart), the lost attempts started 1.5–2.5 h later. (2) **FOUR attempts were lost, not two:** controller job 1357 (start 15:05:41Z → failed 15:18:38Z), felhom.eu job 1359 (15:15:37 → 15:28:38 — the previous session blamed that one on the BusyBox fault; the runner never ran it), job 1361 (15:33:21 → 15:43:38) and its API re-run (15:46:48 → 15:58:38). None has a `task` line in the runner log; every one was failed at a :38-second mark on a 5-minute step, 10–13 min after it was handed out — **the shape of Gitea's periodic „zombie task" stop** (a task assigned to a runner that never reports is failed after ~10 min; no log exists because none was written). (3) The runner's task ids are NOT job id + 1 (the controller re-run was task 1363). (4) **Gitea's own log for the window is gone** — the pod log starts 16:16:32Z (rotated), so the assignment side cannot be read. **Likely mechanism, NOT proven:** the runner's fetch-task request timed out on its side after Gitea had already assigned the task, so the task was orphaned. In that same hour this session polled Gitea's jobs API hard (15 pages every 15 s per wait loop) and Gitea logged „slow" requests — a plausible load cause, and the session's own. Mitigation taken: the session's CI waiter now polls once a minute. **Nothing changed on DooPlex.** **MECHANISM SEEN 2026-10-05 17:15–17:28Z, with Gitea's own log (round 2):** catalog run 1368 (`4828dc7`) — 17:15:14 the job is marked started; 17:15:16 `router: slow POST /api/actions/runner.v1.RunnerService/FetchTask for 10.42.0.42 (the runner), elapsed 3192ms`, then `UpdateRepoRunsNumbers … context canceled` and `GetActionWorkflow: EOF` — **the runner abandoned its fetch after Gitea had assigned the task**; the runner log has no line for task 1371; 17:28:39 `actions/clear_tasks.go:174 stopTasks() [W] Cannot transfer logs of task 1371` — Gitea's zombie-task stop. **The load at that minute:** an outside crawler (216.73.216.78) walking commit pages and `archive/*.tar.gz`, and THIS session's CI waiter, whose 15-page job listings took 13–31 s each. An API re-run passed in 7 s. **Done in-session:** the waiter now asks `GET …/actions/runs?head_sha=<sha>` once a minute (1 s). **Not done (DooPlex, the operator's):** the runner's fetch timeout and Gitea's exposure to the crawler. | **OPEN** **DATED CHECK 2026-10-12 (DUE-CHECKS):** if no job was lost since 2026-10-05 16:00Z (no completed job whose runner log has no `task` line / whose log API answers `file does not exist`), close. **NIGHT WATCH 2026-10-05/06 (burn-down night): 2 jobs lost of ~30 runs** — felhom.eu run 1384 (`4aa4d837`, 21:23→21:33Z, no log) and felhom-controller run 1401 (`c67b26be`, 00:45→00:58Z, no log); each re-run once through the API and each passed (2 m 05 s, 57 s). The night's waiter made one filtered call a minute. So the 2026-10-12 close condition („no job lost since 2026-10-05 16:00Z") is already NOT met. `audits/night-burndown-2026-10-05/r887-lost-jobs.txt`. | — | Operator: decide whether to raise the act-runner fetch timeout and/or rate-limit the public Gitea pages the crawler walks; meanwhile re-run a lost job via `POST /repos/admin/<repo>/actions/runs/<id>/rerun`. Keep the 2026-10-12 check | operator |
|
||||
| **R-892** | Process & tooling | P4 | **The update test's box walk cannot reach the Tester 1 box, so decision 149's admin seed there cannot be used.** `app-catalog-felhom.eu/scripts/box_walk.py` drives guests only on demo-hp (its `HP`, `ssh` + `pct exec`) and reaches the app by the guest's LAN address. Read 2026-10-06 (evening): the Tester 1 box (hub host `tester-1-d70be4`) has no SSH alias in DooPlex's `~/.ssh/config` and no entry in `operations/nodes.md`. **Corrected 2026-10-06 18:24 (`09` §3 decision 158):** its Proxmox host IS known — it runs as **VM 341 on the HP box** (`ssh hp`), recorded in `audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt`; the session that filed this did not find that file. Not a 30-minute fix: it needs the box's location and an operator-approved route first. | **OPEN — filed 2026-10-06** **2026-10-06 18:24: operator ruling — yes, CC may reach it by SSH (decision 158).** | — | Build the route (a jump through `hp`, the target in `box_walk.py`'s own table), prove one step there, add the box to `operations/nodes.md` | CC |
|
||||
| **R-892** | Process & tooling | P4 | **The update test's box walk cannot reach the Tester 1 box, so decision 149's admin seed there cannot be used.** `app-catalog-felhom.eu/scripts/box_walk.py` drives guests only on demo-hp (its `HP`, `ssh` + `pct exec`) and reaches the app by the guest's LAN address. Read 2026-10-06 (evening): the Tester 1 box (hub host `tester-1-d70be4`) has no SSH alias in DooPlex's `~/.ssh/config` and no entry in `operations/nodes.md`. **Corrected 2026-10-06 18:24 (`09` §3 decision 158):** its Proxmox host IS known — it runs as **VM 341 on the HP box** (`ssh hp`), recorded in `audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt`; the session that filed this did not find that file. Not a 30-minute fix: it needs the box's location and an operator-approved route first. | **OPEN — filed 2026-10-06** **2026-10-06 18:24: operator ruling — yes, CC may reach it by SSH (decision 158).** **2026-10-06 (night): the route is BUILT in the box walk** (catalog `d63ea35`: `TARGET=tester-1`, `ssh -J demo-hp root@192.168.0.154`, guest 9201, `felhom.enkicsifelhom.hu`; `BOX_ADMIN_SEED_GUESTS` has it) and the identity is matched (the agent's report `host.node=felhom` = VM 341's certificate; the guest answers its domain 200, demo-hp's 404). **Blocked:** DooPlex's key is not authorized on VM 341 (`Permission denied (publickey,password)`); the VM has no guest agent and a disk edit needs a VM stop (a reboot, not allowed); fetching its vaulted password from the hub was refused by the session's permission check. `audits/readback-2026-10-07/` | DooPlex's key on VM 341 | Operator: authorize DooPlex's public key on VM 341 (one line in `/root/.ssh/authorized_keys`, from its console), or allow CC to read the vaulted password; then prove one step there | operator |
|
||||
| **R-206** | Process & tooling | P4 | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC |
|
||||
| **R-209a** | Process & tooling | P4 | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator + CC |
|
||||
| **R-230** | Process & tooling | P4 | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | — | — | operator |
|
||||
|
||||
Reference in New Issue
Block a user