second night: scratch guest 9202 built (R-481 CLOSED, persists); controller v0.242.0 delivered (R-487 R-491 R-490 R-476 R-456 CLOSED, R-489 re-scoped); R-492 filed; rotation restarted from bentopdf; morning note
gates / gates (push) Successful in 19s

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-13 23:06:21 +02:00
parent 72ee053a9e
commit 41590f8ee6
26 changed files with 586 additions and 46 deletions
+11
View File
@@ -40,6 +40,17 @@ on demo-hp — `documentation/audits/slice4-2026-09-13/`, design `09-update-arch
- **One host, one customer** (schema: `hosts.host_id` PK, one `customer_id`; one agent per box). A
scratch guest "enrolled as its own customer" on a box that already belongs to a customer is not
something the product can do (R-481, 2026-09-13); a second guest of the same customer is.
- **The scratch guest is LXC 9202 on demo-hp and it persists** (operator ruling option 1, built
2026-09-13 night, R-481 CLOSED). A second guest of the demo-hp customer, restored from the vouched
golden onto the NVMe `dir` storage at `/mnt/hdd_1`, hub OFF, tunnel OFF, agent OFF, off-site OFF,
self-update OFF — so the floor does not reach it and its controller image is set by hand in
`/etc/felhom-controller-image` (the one place hand-setting is allowed). Never bind the real data
drive into it; never start cloudflared there (a second connector would serve the public domain).
Two choices were CC's and are reversible — decided by CC unattended, operator may reverse: no
data-drive bind, no tunnel. Reach, shape and rebuild recipe: `operations/nodes.md`.
- **The local backup lists are keyed on the DRIVES, not on what is deployed** (controller v0.242.0,
R-487 — the same rule R-237 set for the off-site list). A removed app whose unit was kept is listed
with its restore; a unit on a data drive is opened where it sits.
- **A release reaches the fleet by floor between golden bakes when its MinAgent is declared with the
floor** (ruling 2026-09-13, hub v0.112.0, R-472 CLOSED). An undeclared floor above the golden is still
held, and the forms refuse it. v0.239.0 arrived on both demo boxes this way in ~15 s. Never hand-deploy
+33 -33
View File
@@ -1,45 +1,45 @@
# REPORT — the four evening items before the second night (2026-09-13)
# REPORT — the second night: scratch guest built, controller v0.242.0, rotation restarted (2026-09-13/14)
*Overwritten each session. The night's own report (adventurelog rotation, v0.240.0) is preserved as
`audits/nightly-2026-09-13-adventurelog/` and the register; the controller report is
`felhom-controller/REPORT.md`, the catalog's `app-catalog-felhom.eu/REPORT.md`.*
*Overwritten each session. Evidence: `documentation/audits/nightly-2026-09-13b-bentopdf/` (the guest
build and the bentopdf walk) and `documentation/audits/v0242-2026-09-14/` (red-proofs, floor, three
live passes). Controller detail: `felhom-controller/REPORT.md`.*
## 1. The rules file
**Rules:** `.claude/rules/unprompted-work.md`. **Architecture named:** `01-topology-and-trust.md` §2
(the scratch guest is a second guest of the same customer), `07-backup-architecture.md` §6.3 and the
R-237 store-keyed list rule, `09-update-architecture.md` (holds). **Baselines:** felhom.eu `72ee053`,
controller `3e81330` (v0.241.0), catalog `6d6eec3`.
`.claude/rules/unprompted-work.md` copied byte-identical into `felhom-controller`, `felhom.eu` and
`app-catalog-felhom.eu` (documents-only pushes). `instructions_gate.py` refused the first push: a rule
file needs `paths:` or `unconditional: true`; all four copies (root included) now start with the
three-line `unconditional: true` frontmatter — sha256 `d1aa3af83ae3b220…` everywhere.
## Step 0 — R-481, operator ruling option 1 → LXC 9202 on demo-hp, persists
## 2. R-481 — the scratch guest (ruled; NOT built, and why)
Restored from the vouched golden 0.236.0 onto a `dir` storage re-added at `/mnt/hdd_1`
(`nvme-scratch`), sized like 9201, unprivileged; the demo-hp customer seeded with hub OFF, tunnel OFF,
agent OFF, off-site OFF, self-update OFF; image set by hand (the one place allowed). Disposition in
the guest, on the host and on the hub side (`operations/nodes.md`, the register). Two CC decisions,
tagged in `CONTEXT.md` and `01-topology-and-trust.md`: no real-data-drive bind, no cloudflared.
The ruling ("a second enrolled LXC … enrolled as a scratch customer") collides with the model: the hub
keys a host to one customer (`hosts.host_id` PK) and the box runs one agent with one host identity;
re-running the installer with another customer-id would rewrite demo-hp's enrolment. Options and a
recommendation are in the row and STATUS; 9201's shape and the disk placement (`pct move-volume`
onto a `dir` storage at `/mnt/hdd_1`) are recorded for whichever is chosen.
## Step 1–3 — controller v0.242.0 (`d698ce3`, docs `406755f`), one release, floor-delivered
## 3. R-483 — photos (re-ranked P2, catalog)
R-487 (removed app listed with its restore; the lists are keyed on the drives), R-491 (removal
clears the update hold), R-490 (`/api/system/info` reachable + fallback), R-489 (`volumes_removed`
difference — **half: a restore-recreated volume has no compose label and is missed; row kept open,
measured**), R-476 (Tier-2 copy dated by its data), R-456 (rule pinned). Red-proofs ×11, full gate
green, floor 0.242.0 with MinAgent 0.129.0: demo-hp +16 s, demo-felhom +17 s. Live on 9202 with an
opengist throwaway through the endpoints the UI invokes (no browser on DooPlex); the first pass was
refused 409 (a running app must be stopped before removal) and repeated.
Diffed against `homelab-manifests` `adventurelog-system/adventurelog.yaml`: the k3s ingress sends
`/media`, `/static`, `/admin`, `/accounts` to the backend **service port 80**; the catalog sent
everything to the frontend. Two catalog cuts (`3172258`, `ed62cfd`): a backend traefik router for the
four prefixes, then port 80 instead of 8000 — the backend image runs nginx in front of gunicorn and
Django serves a photo by `X-Accel-Redirect`, so port 8000 returned 200 with an empty body. Applied to
the operator's own instance through the guarded Update; proven without a browser
(`audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt`): `GET /media/images/<id>.webp` with a
session → 200 `image/webp`, RIFF/WEBP magic; anonymous 403 (the app's rule). **Confirmed by the operator in a browser at 21:49 — closed.** The frontend's 500 on scripted multipart uploads is upstream and
recorded in the row; browser uploads work.
## Step 2 — rotation reset to bentopdf on 9202: clean
## 4. R-479 — tier order for bind-data apps (ruled; controller v0.241.0)
Front door, use (browser-side tool, no data route — recorded, not faked), backup + second copy,
remove-with-data keep-backups, Tier-2 unit restore, guarded update (same pin; backed up first because
the restore rewrote `deployed_at`, R-478), remove-everything. No new rows from the walk.
Ruling 9 in `09-update-architecture.md` §3. Built, tested, red-proofed; delivered by the floor and
live-checked on a nextcloud throwaway — see `felhom-controller/REPORT.md` and
`audits/v0241-2026-09-13/`.
## Register
## Observations
210 open / 194 closed → **205 / 200**. Closed: R-481, R-487, R-491, R-490, R-476, R-456. Re-scoped:
R-489. Opened: R-492 (delete `cfg.Paths.HDDPath`). Security note: demo-hp's retrieval passphrase
appeared once in a tool output on DooPlex during the guest build (redaction list missed the key).
0. **A removed app keeps its update hold; a reinstall would start held.** FILED: R-491
## Teardown, three layers
1. **The controller's evening release rides the same session as the night's — two releases in one calendar day, one per session.** NOT-A-FINDING: the rule is per session; both are recorded with their MinAgent lines.
2. **`homelab-manifests` is not in the workspace; its manifests were read from Gitea's API.** NOT-A-FINDING: the workspace CLAUDE.md says non-felhom repos are unrelated; a read-only fetch was enough.
Machine: throwaways (opengist, bentopdf) removed with data and backups; 9202 persists on purpose.
Host: the `nvme-scratch` storage and 9202 stay (disposition recorded). Hub: floor raised; nothing else.
+42
View File
@@ -1,5 +1,47 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-14 (morning note, the second night) — the scratch guest is built, one controller release, the rotation restarted.**
**Decisions I took.** (1) The scratch guest on the HP was built as you ruled: a second guest under
the HP's own customer, on the fast internal disk, sized like the main one, kept on purpose and written
up in all three places. Two properties of it were my call and you may reverse them: it never sees your
real data drive, and it never starts the public tunnel — a second connector would serve the public
domain from a throwaway. (2) A finding I had already listed as fixed turned out half-fixed when
measured live, and the rule is one release per night, so I kept the line open with the exact measurement
instead of shipping a second release. (3) The removed-app listing was a medium-priority line, but it
needed no ruling, touched no customer data and was the rotation's own finding, so I took it into
tonight's release.
**What I exercised.** The rotation restarted from its first standing app on the scratch guest: front
door, use, backup, second copy, remove-with-data, full restore from the second copy, the guarded
update, remove-everything — all clean. Then a throwaway app for the release proof.
**What broke, and whether I fixed it.** Six lines fixed and shipped in one controller release,
delivered by the floor in 16 and 17 seconds, proven on the scratch guest: an app you removed while
keeping its backup now shows on the backup pages with a button that reinstalls it (before, the way
back existed only as a hidden endpoint, and a backup kept on a data drive could not be found at all);
a removal now clears the "held after a failed update" mark; the memory card on the monitoring page can
render; the second copy is dated by its data rather than by a file that only moves when the app's
definition changes; and a boot rule is now pinned by a test. Half-fixed: the removal's list of
deleted volumes is right for a freshly installed app and empty for one that came back from a restore,
because the restore recreates the volume without the label the list looks for. Measured, kept open.
One security slip to know about: while building the scratch guest, the HP's retrieval passphrase was
printed once into a tool output here on DooPlex. Nothing left the machine.
**Rows.** Opened 1 (delete the empty drive-path setting nothing reads any more). Closed 6 (the
scratch guest, the removed-app listing, the hold left behind, the monitoring card, the second-copy
date, the boot rule). Re-scoped 1 (deleted-volume list). Register: 210 open / 194 closed before,
205 open / 200 closed after.
**Needs you.** (1) Open the backup page once in a browser after removing a throwaway app with its
backup kept, and press the new button — strict screen coverage is yours. **If you do nothing:** the
feature stays proven at the endpoint level only. (2) The passphrase slip: re-issue the HP's retrieval
passphrase from the hub when convenient. **If you do nothing:** the old one stays valid; the exposure
is one line in this session's local record. (3) The scratch guest stays up and idle. **If you do
nothing:** it costs the HP about one and a half gigabytes of memory and nothing else.
---
**Updated 2026-09-13 (evening, before the night) — your four items.**
**Decisions I took.** (1) The rules file now sits in the workspace root and all three repos,
@@ -64,6 +64,14 @@ customer box.
one host (a company environment) is **not precluded** — the agent manages a *set* of
guests. The only multi-tenant-specific work deferred to "if it becomes real" is resource
fairness (per-guest disk/RAM/CPU quotas).
- **A scratch guest is a second guest of the SAME customer, unenrolled** (operator ruling
2026-09-13, R-481; built as LXC 9202 on demo-hp, persists). The hub ties one host to one
customer, so a second enrolled customer on a box is not a thing the product does. The scratch
guest runs with the hub, the tunnel, the agent link, off-site and self-update all off, and its
controller image is set by hand — the one place that is allowed. Two of its properties were
*decided by CC unattended — operator may reverse*: **it never binds the customer's real data
drive** (a throwaway must not be able to reach real data) and **it never starts cloudflared**
(a second connector would serve the public domain from a scratch box).
---
@@ -277,7 +277,7 @@ what the unit holds** — for an app whose data is a bind mount that is the defi
(R-479). **Removal and the tiers (controller v0.240.0):** removing an app with its backups KEPT keeps
the unit, the Tier-2 mirror AND the Tier-2 record, so „Teljes visszaállítás" still works afterwards
(R-486); „Mentési adatok törlése" deletes the unit, every mirror and the app's backup preferences, and
never touches off-site snapshots (R-474). A removed app is listed on neither backup page (R-487, open). **For a bind-data app the update's tier order is second drive → off-site → own unit (R-479, v0.241.0), because the unit does not hold the files.**
never touches off-site snapshots (R-474). A removed app whose unit was kept is listed on both local backup pages with its restore since controller v0.242.0 (R-487): **the local lists are keyed on the drives, not on what is deployed** — the rule R-237 set for the off-site list — and the restore opens the unit where it sits, a data drive included. **For a bind-data app the update's tier order is second drive → off-site → own unit (R-479, v0.241.0), because the unit does not hold the files.**
### 6.1 The four tiers, as configured on the live fleet
@@ -0,0 +1,17 @@
=== scratch guest 9202 — first contact 2026-09-13T20:24:41Z ===
HTTP 200
{"ok":true,"data":[{"name":"actualbudget","meta":{"display_name":"ActualBudget","description":"Személyes pénzügyek és költségvetés kezelése","category":"finance","subdomain":"budget",
--- sync the catalog ---
HTTP 200
{"ok":true,"data":{"ok":true,"message":"Sablonok naprakészek — nincs változás"},"message":"Sablonok naprakészek — nincs változás"}
stacks known: 55 | deployed: []
--- guest facts ---
VMID Status Lock Name
9201 running demo-hp
9202 running demo-hp-scratch
felhom-controller filebrowser traefik
0
SCRATCH GUEST — R-481, operator ruling 2026-09-13 (option 1)
The nightly rotation throwaway host under the demo-hp customer. Hub OFF, tunnel OFF, agent OFF, off-site OFF.
@@ -0,0 +1,60 @@
=== NIGHT 2 — bentopdf as a customer on the SCRATCH guest 9202 — 2026-09-13T20:25:37Z ===
controller: gitea.dooplex.hu/admin/felhom-controller:0.241.0 | catalog pin: image: ghcr.io/alam00000/bentopdf:v2.8.6
--- install ---
HTTP 202 {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
after: state=running updating=False phase=None err='' hold='' | ghcr.io/alam00000/bentopdf:v2.8.6 Up 14 seconds (healthy)
--- front door: GET / and the app's assets through traefik ---
GET / -> 200 text/html 72677
GET /index.html -> 301 text/html 169
USE: BentoPDF is a browser-side toolbox — every PDF is processed in the browser and never sent to the server. There is no data route and no API; nothing to put in and nothing to read back. Recorded, not faked.
--- backup through the backups page: backup now, then the 2nd copy ---
HTTP 200 {"ok":true,"message":"Mentés elindítva"}
idle after 5 s
HTTP 200 {"ok":true,"message":"2. mentés elindítva"}
/mnt/sys_drive/felhom-data/backups/primary/bentopdf:
2026-09-13T20:26 1864 compose/.felhom.yml
2026-09-13T20:26 291 compose/app.yaml
2026-09-13T20:26 1109 compose/docker-compose.yml
2026-09-13T20:26 1009 manifest.json
/mnt/felhom-drives/scratch_hdd/backups/secondary/bentopdf:
2026-09-13T20:26 1864 recovery-unit/compose/.felhom.yml
2026-09-13T20:26 1109 recovery-unit/compose/docker-compose.yml
2026-09-13T20:26 291 recovery-unit/compose/app.yaml
2026-09-13T20:26 1009 recovery-unit/manifest.json
2026-09-13T20:26 1 .felhom-tier2-layout
card: {"ok":true,"data":{"stack":"bentopdf","backup_paths":[{"path":"/mnt/sys_drive/felhom-data/backups/primary/bentopdf","size_bytes":12465,"size_human":"24K","exists":true},{"path":"/mnt/felhom-drives/scratch_hdd/backups/secondary/bentopdf","size_bytes":16562,"size_human":"32K","exists":true}],"has_back
--- disaster: remove with data, keep backups; the way back: Teljes visszaállítás (Tier 2) ---
HTTP 200 {"ok":true,"data":{"removed":"bentopdf","volumes_removed":null,"hdd_paths_removed":[],"hdd_paths_preserved":[],"hdd_note":"Az alkalmazás nem tárolt saját adatot külső meghajtón, így ott nem volt mit törölni."},"message":"Stack bentopdf removed"}
removed: state=not_deployed updating=False phase=None err='' hold='' | front door now: 404 text/plain; charset=utf-8 19
HTTP/2 302
location: /backups/apps?flash=Teljes+vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl.
restore-status: {"running": false, "op": "tier2-unit-restore", "stack": "bentopdf", "started_at": "2026-09-13T20:26:23.141536724Z", "last": {"op": "tier2-unit-restore", "stack": "bentopdf", "ok": true, "message": "A(z) bentopdf: a beállítások visszaálltak — az alkalmazás újraindult. FIGYELEM: ez a mentés csak a beállításokat tartalmazta, adatot nem. Az alkalmazás adatai NEM álltak vissza ebből a mentésből. A viss
after restore: state=running updating=False phase=None err='' hold='' | front door: 200 text/html 72677
--- guarded Update (same pin: no edge in the catalog) ---
--- POST /api/stacks/bentopdf/update at 2026-09-13T20:26:34Z ---
HTTP 202
{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"}
+ 0s phase backing-up
+ 3s phase done
end (2026-09-13T20:26:38Z, +3s): state=running updating=False phase=done err='' hold=''
2026/09/13 20:26:35 update.go:391: [INFO] [stacks] update bentopdf: accepted — guarded update started
2026/09/13 20:26:35 update.go:820: [INFO] [stacks] update bentopdf: phase checking
2026/09/13 20:26:35 update.go:508: [INFO] [stacks] update bentopdf: no usable copy on any tier — younger than 24h0m0s and not older than this install's deploy (2026-09-13T20:26:23Z) (found: Tier 2 (second drive) at 2026-09-13T20:26:13Z (0s old); Tier 1 (own recovery unit) at 2026-09-13T20:26:07Z (
2026/09/13 20:26:35 update.go:820: [INFO] [stacks] update bentopdf: phase backing-up
2026/09/13 20:26:35 update.go:523: [INFO] [stacks] update bentopdf: precondition met after the backup — Tier 2 (second drive) copy from 2026-09-13T20:26:35Z
2026/09/13 20:26:35 update.go:820: [INFO] [stacks] update bentopdf: phase safety-dump
2026/09/13 20:26:35 update.go:538: [INFO] [stacks] update bentopdf: safety dump done (0 file(s)) []
2026/09/13 20:26:35 update.go:820: [INFO] [stacks] update bentopdf: phase pinning
2026/09/13 20:26:35 pin.go:362: [INFO] [stacks] update bentopdf: pin advanced to the catalog's current definition (bentopdf=ghcr.io/alam00000/bentopdf:v2.8.6)
2026/09/13 20:26:35 update.go:820: [INFO] [stacks] update bentopdf: phase pulling
2026/09/13 20:26:35 update.go:820: [INFO] [stacks] update bentopdf: phase starting
2026/09/13 20:26:36 update.go:820: [INFO] [stacks] update bentopdf: phase verifying
2026/09/13 20:26:36 update.go:614: [INFO] [stacks] update bentopdf: healthy after 0s (the app's health check passed)
2026/09/13 20:26:36 update.go:620: [INFO] [stacks] update bentopdf: DONE in 1s
--- remove with data AND backups; drive clean? ---
HTTP 200 {"ok":true,"data":{"removed":"bentopdf","volumes_removed":null,"hdd_paths_removed":[],"hdd_paths_preserved":[],"hdd_note":"Az alkalmazás nem tárolt saját adatot külső meghajtón, így ott nem volt mit törölni.","backup_paths_removed":["/mnt/sys_drive/felhom-data/backups/primary/bentopdf (24K)","/mnt/felhom-drives/scratch_hdd/backups/secondary/bentopdf (16.2 KB)"]},"message":"Stack bentopdf removed"}
/mnt/sys_drive/felhom-data/backups/primary/bentopdf: ABSENT
/mnt/felhom-drives/scratch_hdd/backups/secondary/bentopdf: ABSENT
containers: 0 | prefs: None None | card: state=not_deployed updating=False phase=done err='' hold=''
=== done 2026-09-13T20:26:46Z ===
@@ -0,0 +1,54 @@
=== FLOOR RAISE: 0.242.0 with min_agent 0.129.0 (R-487 R-491 R-490 R-489 R-476) ===
T0 POST at 2026-09-13T20:51:13Z
HTTP/1.1 303 See Other
Location: /configuration?flash=floor_set
DB floor + declared: id="global-floor-input" name="min_controller_version" value="0.242.0" | name="min_agent" value="0.129.0" placeholder="MinAgent from |
demo-hp on 0.242.0 at 2026-09-13T20:51:30Z (+16s): gitea.dooplex.hu/admin/felhom-controller:0.242.0 Up 7 seconds (healthy)
demo-felhom on 0.242.0 at 2026-09-13T20:51:31Z (+17s): gitea.dooplex.hu/admin/felhom-controller:0.242.0 Up 10 seconds (healthy)
--- hub log since T0 (managed floor) ---
2026/09/13 22:51:14 [INFO] Global controller-version floor set to "0.242.0" (declared MinAgent "0.129.0")
2026/09/13 22:51:16 [INFO] managed floor SERVED for demo-felhom: floor 0.242.0, agent requirement "0.129.0" from declared (golden 0.236.0)
2026/09/13 22:51:17 [INFO] managed floor SERVED for demo-hp: floor 0.242.0, agent requirement "0.129.0" from declared (golden 0.236.0)
--- demo-hp: controller log (SetFloor / self-update / version) ---
2026/09/13 20:51:23 updater.go:96: [DEBUG] [selfupdate] VerifyStartup: checking update state in /opt/docker/felhom-controller/data
2026/09/13 20:51:23 updater.go:96: [DEBUG] [selfupdate] VerifyStartup: pending update found — target=0.242.0 previous=0.241.0
2026/09/13 20:51:23 updater.go:809: [INFO] [selfupdate] Post-update startup: update successful (0.241.0 → 0.242.0)
2026/09/13 20:51:23 main.go:655: [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30)
2026/09/13 20:51:23 scheduler.go:102: [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s)
2026/09/13 20:51:23 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="selfupdate-check" interval=6h0m0s totalJobs=18
2026/09/13 20:51:23 offsiteapply.go:151: [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply
2026/09/13 20:51:24 client.go:67: [DEBUG] [agentapi] agent version seen: 0.130.0 (was "")
2026/09/13 20:51:28 builder.go:40: [DEBUG] [report] BuildReport: starting — version=0.242.0, storagePaths=1
2026/09/13 20:51:29 updater.go:96: [DEBUG] [selfupdate] SetFloor: floor "" → "0.242.0"
2026/09/13 20:51:29 updater.go:96: [DEBUG] [selfupdate] maybeAutoUpdate: current 0.242.0 >= floor 0.242.0 — no action
2026/09/13 20:51:33 offsiteapply.go:151: [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.242.0 (we are 0.242.0), no managed update running
--- demo-hp: bootstrap service journal ---
Sep 13 20:51:22 demo-hp systemd[1]: felhom-controller-bootstrap.service: Deactivated successfully.
Sep 13 20:51:22 demo-hp systemd[1]: Stopped felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
Sep 13 20:51:22 demo-hp systemd[1]: Stopping felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
Sep 13 20:51:22 demo-hp systemd[1]: Starting felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
Sep 13 20:51:22 demo-hp felhom-controller-bootstrap.sh[2438508]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.242.0 from /etc/felhom-bootstrap/bootstrap.json (hostname=demo-hp)
Sep 13 20:51:22 demo-hp felhom-controller-bootstrap.sh[2438582]: 9e9e3a9987db6cab7c7d8d413a64ce73695fe5b9f090f554022ab02520ebe0b7
Sep 13 20:51:23 demo-hp felhom-controller-bootstrap.sh[2438508]: [ctrl-bootstrap] controller started
Sep 13 20:51:23 demo-hp systemd[1]: Finished felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
--- demo-felhom: controller log (SetFloor / self-update / version) ---
2026/09/13 20:51:22 [INFO] [selfupdate] Post-update startup: update successful (0.241.0 → 0.242.0)
2026/09/13 20:51:22 [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30)
2026/09/13 20:51:22 [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s)
2026/09/13 20:51:22 [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply
2026/09/13 20:51:32 [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.242.0 (we are 0.242.0), no managed update running
--- demo-felhom: bootstrap service journal ---
Sep 13 20:51:21 demo-felhom systemd[1]: felhom-controller-bootstrap.service: Deactivated successfully.
Sep 13 20:51:21 demo-felhom systemd[1]: Stopped felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
Sep 13 20:51:21 demo-felhom systemd[1]: Stopping felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
Sep 13 20:51:21 demo-felhom systemd[1]: Starting felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)...
Sep 13 20:51:21 demo-felhom felhom-controller-bootstrap.sh[3573553]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.242.0 from /etc/felhom-bootstrap/bootstrap.json (hostname=demo-felhom)
Sep 13 20:51:21 demo-felhom felhom-controller-bootstrap.sh[3573602]: eaa532876b7d3df1f2ab46965d08245f2b3eaff97f8ba9b4f6d93e799876a38f
Sep 13 20:51:21 demo-felhom felhom-controller-bootstrap.sh[3573553]: [ctrl-bootstrap] controller started
Sep 13 20:51:21 demo-felhom systemd[1]: Finished felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount).
SUMMARY T0=2026-09-13T20:51:13Z {'demo-hp': ('2026-09-13T20:51:30Z', 16), 'demo-felhom': ('2026-09-13T20:51:31Z', 17)} not done: []
@@ -0,0 +1,90 @@
=== v0.242.0 LIVE on the scratch guest 9202 — 2026-09-13T20:52:22Z ===
controller: gitea.dooplex.hu/admin/felhom-controller:0.242.0 Up 35 seconds (healthy)
--- R-490: GET /api/system/info (the monitoring card's fetch) ---
HTTP HTTP 200 | hdd_configured=True hdd_total_gb=937.8 used_mem_mb=0
--- install opengist (throwaway) ---
HTTP 202 {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
after: state=running updating=False phase=None err='' hold='' | front door: 302 0
--- backup now → the primary unit ---
HTTP 200 {"ok":true,"message":"Mentés elindítva"}
idle after 5 s
/mnt/sys_drive/felhom-data/backups/primary/opengist:
2026-09-13T20:52 1941 compose/.felhom.yml
2026-09-13T20:52 292 compose/app.yaml
2026-09-13T20:52 1260 compose/docker-compose.yml
2026-09-13T20:52 181248 volume-dumps/opengist_opengist_data.tar
2026-09-13T20:52 1041 manifest.json
/mnt/felhom-drives/scratch_hdd/backups/secondary/opengist: ABSENT
--- R-491: seed an UPDATE hold (settings file edit on the scratch guest + controller restart; the hold's origin is not what is under test) ---
positive control — held: state=running updating=False phase=None err='' hold='A(z) opengist frissítése 2026-09-13 22:52-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 22:52 — ez a másolat a'
settings has hold: True
--- R-489 + R-491: remove with data, KEEP backups ---
HTTP HTTP 409
volumes_removed=None backup_paths_removed=None hdd_paths_removed=None
log:
after removal: state=running updating=False phase=None err='' hold='A(z) opengist frissítése 2026-09-13 22:52-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 22:52 — ez a másolat a' | settings still has hold: True
docker volume ls: opengist_opengist_data
--- R-487: the removed app on the two pages and the picker API ---
/backups/apps has the row: False | Megtartva badge: False | restore form: False
row says last backup: None
GET /api/backup/snapshots → HTTP 200 {"ok":true,"data":[{"time":"2026-09-13T20:52:42Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
/backups/restore picker lists it: False
--- R-487: Visszaállítás a mentésből (POST /backup/restore, the row's button) ---
HTTP/2 302
location: /backups/restore?flash=Vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl.
restore-status: {"running": false, "op": "restore", "stack": "opengist", "started_at": "2026-09-13T20:53:18.897214019Z", "last": {"op": "restore", "stack": "opengist", "ok": true, "message": "A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.", "finished_at": "2026-09-13T20:53:28.201829063Z"}, "last_recent": true}
after restore: state=running updating=False phase=None err='' hold='' | front door: 302 0
log: 2026/09/13 20:53:28 handlers.go:1542: [INFO] [web] Restore completed (async): stack=opengist in 9.304495819s (volumes 1/1, dbs 0/0)
--- R-476: two captures, one definition → the Tier-2 date must be the DATA's ---
HTTP 200 {"ok":true,"message":"Mentés elindítva"}
capture A manifest
HTTP 200 {"ok":true,"message":"Mentés elindítva"}
capture B manifest (same = definition unchanged)
newest dump: -rw-r--r-- 1 root root 181248 2026-09-13T20:54Z opengist_opengist_data.tar
HTTP 200 {"ok":true,"message":"2. mentés elindítva"}
/mnt/sys_drive/felhom-data/backups/primary/opengist:
2026-09-13T20:52 1941 compose/.felhom.yml
2026-09-13T20:52 292 compose/app.yaml
2026-09-13T20:52 1260 compose/docker-compose.yml
2026-09-13T20:54 181248 volume-dumps/opengist_opengist_data.tar
2026-09-13T20:52 1041 manifest.json
/mnt/felhom-drives/scratch_hdd/backups/secondary/opengist:
2026-09-13T20:52 1941 recovery-unit/compose/.felhom.yml
2026-09-13T20:52 1260 recovery-unit/compose/docker-compose.yml
2026-09-13T20:52 292 recovery-unit/compose/app.yaml
2026-09-13T20:54 181248 recovery-unit/volume-dumps/opengist_opengist_data.tar
2026-09-13T20:52 1041 recovery-unit/manifest.json
2026-09-13T20:54 1 .felhom-tier2-layout
Tier-2 unit confirm sentences on the page: []
--- teardown: remove with data AND backups ---
HTTP HTTP 409
volumes_removed=None backup_paths_removed=None
/mnt/sys_drive/felhom-data/backups/primary/opengist:
2026-09-13T20:52 1941 compose/.felhom.yml
2026-09-13T20:52 292 compose/app.yaml
2026-09-13T20:52 1260 compose/docker-compose.yml
2026-09-13T20:54 181248 volume-dumps/opengist_opengist_data.tar
2026-09-13T20:52 1041 manifest.json
/mnt/felhom-drives/scratch_hdd/backups/secondary/opengist:
2026-09-13T20:52 1941 recovery-unit/compose/.felhom.yml
2026-09-13T20:52 1260 recovery-unit/compose/docker-compose.yml
2026-09-13T20:52 292 recovery-unit/compose/app.yaml
2026-09-13T20:54 181248 recovery-unit/volume-dumps/opengist_opengist_data.tar
2026-09-13T20:52 1041 recovery-unit/manifest.json
2026-09-13T20:54 1 .felhom-tier2-layout
volumes left: opengist_opengist_data
/backups/apps still lists it: True | Megtartva anywhere: False
settings restore_holds: none
=== done 2026-09-13T20:55:01Z ===
live-rc=0
@@ -0,0 +1,75 @@
=== v0.242.0 LIVE on 9202, second pass (the first was refused 409: a running app must be stopped before removal) — 2026-09-13T20:56:25Z ===
controller: gitea.dooplex.hu/admin/felhom-controller:0.242.0 Up 3 minutes (healthy) | state=running updating=False phase=None err='' hold=''
--- R-491: seed an UPDATE hold (settings file edit + controller restart; the hold's ORIGIN is not what is under test) ---
positive control — held: A(z) opengist frissítése 2026-09-13 22:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazá | settings has hold: True
--- R-489 + R-491: stop, then remove with data, KEEP backups ---
stop: HTTP 200 {"ok":true,"message":"Stack opengist stop completed"}
HTTP 200
volumes_removed=[] backup_paths_removed=None hdd_paths_removed=[]
log: 2026/09/13 20:56:57 router.go:916: [INFO] [api] remove opengist: its update hold is cleared with it (R-491)
after removal: state=not_deployed updating=False phase=None err='' hold='' | settings still has hold: False
docker volume ls: none
/mnt/sys_drive/felhom-data/backups/primary/opengist:
2026-09-13T20:52 1941 compose/.felhom.yml
2026-09-13T20:52 292 compose/app.yaml
2026-09-13T20:52 1260 compose/docker-compose.yml
2026-09-13T20:54 181248 volume-dumps/opengist_opengist_data.tar
2026-09-13T20:52 1041 manifest.json
/mnt/felhom-drives/scratch_hdd/backups/secondary/opengist:
2026-09-13T20:52 1941 recovery-unit/compose/.felhom.yml
2026-09-13T20:52 1260 recovery-unit/compose/docker-compose.yml
2026-09-13T20:52 292 recovery-unit/compose/app.yaml
2026-09-13T20:54 181248 recovery-unit/volume-dumps/opengist_opengist_data.tar
2026-09-13T20:52 1041 recovery-unit/manifest.json
2026-09-13T20:54 1 .felhom-tier2-layout
--- R-487: the removed app on the two pages and the picker API ---
/backups/apps row: True | Megtartva badge: True | restore form: True
row says last backup: 2026-09-13 22:54
GET /api/backup/snapshots → HTTP 200 {"ok":true,"data":[{"time":"2026-09-13T20:54:42Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
/backups/restore picker lists it: True
--- R-487: Visszaállítás a mentésből (the row's button = POST /backup/restore) ---
HTTP/2 302
location: /backups/restore?flash=Vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl.
restore-status: {"op": "restore", "stack": "opengist", "ok": true, "message": "A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.", "finished_at": "2026-09-13T20:57:13.803999331Z"}
after restore: state=running updating=False phase=None err='' hold='' | front door: 302 0
log: 2026/09/13 20:57:13 handlers.go:1542: [INFO] [web] Restore completed (async): stack=opengist in 9.118788015s (volumes 1/1, dbs 0/0)
row is a normal deployed row again (no Megtartva): True
--- R-476: two captures under one definition → the Tier-2 date must be the DATA's ---
HTTP 200 {"ok":true,"message":"Mentés elindítva"}
capture A manifest "created_at": "2026-09-13T20:52:42Z" at 2026-09-13T20:57:22Z
HTTP 200 {"ok":true,"message":"Mentés elindítva"}
capture B manifest "created_at": "2026-09-13T20:52:42Z" at 2026-09-13T20:58:34Z (same = definition unchanged)
newest dump: -rw-r--r-- 1 root root 181248 2026-09-13T20:58:28Z opengist_opengist_data.tar
HTTP 200 {"ok":true,"message":"2. mentés elindítva"}
/mnt/sys_drive/felhom-data/backups/primary/opengist:
2026-09-13T20:52 1941 compose/.felhom.yml
2026-09-13T20:52 292 compose/app.yaml
2026-09-13T20:52 1260 compose/docker-compose.yml
2026-09-13T20:58 181248 volume-dumps/opengist_opengist_data.tar
2026-09-13T20:52 1041 manifest.json
/mnt/felhom-drives/scratch_hdd/backups/secondary/opengist:
2026-09-13T20:52 1941 recovery-unit/compose/.felhom.yml
2026-09-13T20:52 1260 recovery-unit/compose/docker-compose.yml
2026-09-13T20:52 292 recovery-unit/compose/app.yaml
2026-09-13T20:58 181248 recovery-unit/volume-dumps/opengist_opengist_data.tar
2026-09-13T20:52 1041 recovery-unit/manifest.json
2026-09-13T20:58 1 .felhom-tier2-layout
Tier-2 unit confirm for opengist: második meghajtón lévő másolattal. Ami a másolat óta keletkezett, elveszik. A másolat kelte: 2026-09-13 22:58. A mellette lévő „Fájlok visszaállítása” ezzel szemben csak a hiányzó fájlokat pótolja, és semmit nem ír felül. Az alkalmazás a művelet idejére leáll.
--- teardown: stop, remove with data AND backups ---
stop: HTTP 200 {"ok":true,"message":"Stack opengist stop completed"}
HTTP 200
volumes_removed=[] backup_paths_removed=['/mnt/sys_drive/felhom-data/backups/primary/opengist (208K)', '/mnt/felhom-drives/scratch_hdd/backups/secondary/opengist (197.4 KB)']
/mnt/sys_drive/felhom-data/backups/primary/opengist: ABSENT
/mnt/felhom-drives/scratch_hdd/backups/secondary/opengist: ABSENT
volumes left: none
/backups/apps still lists it: False | Megtartva anywhere: False
settings restore_holds: none
docker ps: felhom-controller filebrowser traefik
=== done 2026-09-13T20:58:52Z ===
@@ -0,0 +1,19 @@
=== R-489 cause on 9202 — 2026-09-13T21:01:40Z ===
A) fresh compose-created volume
HTTP 202 {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
labels: {"com.docker.compose.config-hash":"5f89c3baba6dff6769fa699a804f2c56a643af979b55241e2af324dfde6be51c","com.docker.compose.project":"opengist","com.docker.compose.version":"5.5.0","com.docker.compose.volume":"opengist_data"}
remove-all → HTTP 200 ["opengist_opengist_data"]
left: none
B) volume recreated by a unit restore
HTTP 202 {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
HTTP 200 {"ok":true,"message":"Mentés elindítva"}
HTTP/2 302
location: /backups/restore?flash=Vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl.
restore: A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.
labels: null
remove-all → HTTP 200 []
left: none
units: /mnt/sys_drive/felhom-data/backups/primary/opengist: ABSENT
/mnt/felhom-drives/scratch_hdd/backups/secondary/opengist: ABSENT
=== done 2026-09-13T21:02:40Z ===
@@ -0,0 +1,11 @@
MUTATION in internal/backup/tier2_restore.go:
- if !c.UnitLegPreserved && c.UnitDataDate != "" && c.UnitDataDate > c.UnitPackageDate {
+ if false && !c.UnitLegPreserved && c.UnitDataDate != "" && c.UnitDataDate > c.UnitPackageDate {
=== RUN TestR476_UnitRestoreDateNamesTheDataTimeWhenTheLegWasRefreshed
r476_unit_data_date_test.go:13: refreshed leg: got "2026-09-12T02:15:29Z" preserved=false, want the dump's time
--- FAIL: TestR476_UnitRestoreDateNamesTheDataTimeWhenTheLegWasRefreshed (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.005s
FAIL
exit=1
@@ -0,0 +1,20 @@
MUTATION in internal/backup/removed_units.go:
- if m.stackProvider == nil {
return nil
}
deployed
+ if m.stackProvider == nil || true {
return nil
}
deployed
=== RUN TestR487_ListRemovedAppUnits_ListsKeptUnitsOfUndeployedApps
[INFO] [settings] No settings.json found, using defaults
[INFO] [settings] Added storage path: /tmp/TestR487_ListRemovedAppUnits_ListsKeptUnitsOfUndeployedApps3160582406/004
[INFO] [settings] Settings saved
r487_removed_units_test.go:88: want [drive-app gone-app], got []
--- FAIL: TestR487_ListRemovedAppUnits_ListsKeptUnitsOfUndeployedApps (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s
FAIL
exit=1
@@ -0,0 +1,12 @@
MUTATION in internal/backup/restore_points.go:
- if u, found := m.RemovedAppUnitFor(stackName); found {
+ if u, found := m.RemovedAppUnitFor(stackName); found && false {
=== RUN TestR487_ListRestorePointsFindsARemovedAppsUnit
[INFO] [settings] No settings.json found, using defaults
r487_removed_units_test.go:112: want found with one point, got found=false pts=[]
--- FAIL: TestR487_ListRestorePointsFindsARemovedAppsUnit (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.006s
FAIL
exit=1
@@ -0,0 +1,11 @@
MUTATION in internal/web/templates/backups_restore.html:
- <option value="{{.StackName}}" data-removed="true"
+ <option value="{{.StackName}}" data-gone="true"
=== RUN TestR487_RestorePickerListsRemovedApps
r487_removed_row_test.go:46: removed app must be an option in the picker
--- FAIL: TestR487_RestorePickerListsRemovedApps (0.03s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.038s
FAIL
exit=1
@@ -0,0 +1,16 @@
MUTATION in internal/backup/restore_unit.go:
- if u, found := m.RemovedAppUnitFor(stackName); found {
return u.UnitDir
+ if u, found := m.RemovedAppUnitFor(stackName); found && false {
return u.UnitDir
=== RUN TestR487_PrimaryUnitDirForNamesTheRemovedUnitWhereItSits
[INFO] [settings] No settings.json found, using defaults
[INFO] [settings] Added storage path: /tmp/TestR487_PrimaryUnitDirForNamesTheRemovedUnitWhereItSits2078383903/004
[INFO] [settings] Settings saved
r487_removed_units_test.go:133: unit dir: got /tmp/TestR487_PrimaryUnitDirForNamesTheRemovedUnitWhereItSits2078383903/001/felhom-data/backups/primary/drive-app want /tmp/TestR487_PrimaryUnitDirForNamesTheRemovedUnitWhereItSits2078383903/004/backups/primary/drive-app
--- FAIL: TestR487_PrimaryUnitDirForNamesTheRemovedUnitWhereItSits (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.006s
FAIL
exit=1
@@ -0,0 +1,15 @@
MUTATION in internal/web/handlers.go:
- for _, u := range s.backupMgr.ListRemovedAppUnits() {
rows = append(rows, AppBackupRow{
+ for _, u := range s.backupMgr.ListRemovedAppUnits() {
_ = u
continue
rows = append(rows, AppBackupRow{
=== RUN TestR487_BuildAppBackupRowsAppendsRemovedApps
r487_removed_row_test.go:109: want one removed row for gone-app, got []
--- FAIL: TestR487_BuildAppBackupRowsAppendsRemovedApps (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.007s
FAIL
exit=1
@@ -0,0 +1,11 @@
MUTATION in internal/web/templates/backups_apps.html:
- <input type="hidden" name="snapshot_id" value="helyi">
+
=== RUN TestR487_BackupsPageListsARemovedAppWithItsRestore
r487_removed_row_test.go:28: removed row must render "name=\"snapshot_id\" value=\"helyi\""
--- FAIL: TestR487_BackupsPageListsARemovedAppWithItsRestore (0.02s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.032s
FAIL
exit=1
@@ -0,0 +1,15 @@
MUTATION in /mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/stacks/delete.go:
- removed := []string{}
for _, v := range before {
+ var removed []string
for _, v := range before {
=== RUN TestR489_RemovedVolumesIsADifferenceAndNeverNull
r489_volumes_removed_test.go:20: no volumes must read as [] — got {"removed":"","volumes_removed":null,"hdd_paths_removed":null,"hdd_paths_preserved":null}
--- FAIL: TestR489_RemovedVolumesIsADifferenceAndNeverNull (0.00s)
=== RUN TestR489_ProjectVolumesReadsTheDockerListing
--- PASS: TestR489_ProjectVolumesReadsTheDockerListing (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.008s
FAIL
exit=1
@@ -0,0 +1,11 @@
MUTATION in /mnt/5_hdd/felhom.eu/git/felhom-controller/controller/cmd/controller/main.go:
- mux.Handle("/api/system/info", webServer.RequireAuth(webServer.CsrfProtect(http.HandlerFunc(apiRouter.ServeHTTP))))
+ _ = 0 // RED-PROOF: exact mount removed
=== RUN TestR490_SystemInfoIsMountedAheadOfTheWebPrefix
r475_wiring_test.go:111: /api/system/info must be mounted on the API router ahead of the web layer's /api/system/ prefix (R-490)
--- FAIL: TestR490_SystemInfoIsMountedAheadOfTheWebPrefix (0.01s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.016s
FAIL
exit=1
@@ -0,0 +1,11 @@
MUTATION in internal/api/router.go:
- if p := r.sett.GetDefaultStoragePath(); p != "" {
+ if p := r.sett.GetDefaultStoragePath(); p != "" && false {
=== RUN TestR490_SystemInfoFallsBackToTheDefaultStoragePath
r474_remove_wiring_test.go:119: hdd_configured must be true with a default storage path registered; body {"ok":true,"data":{"sync_status":null,"system":{"total_mem_mb":64272,"used_mem_mb":0,"avail_mem_mb":64272,"mem_percent":0,"disk_total_gb":444.60704803466797,"disk_used_gb":215.30562591552734,"disk_avail_gb":206.64654159545898,"disk_percent":48.42604876986543,"disk_known":true,"hdd_configured":false,"hdd_known":false,"cpu_percent":0,"load_avg_1":4.9,"load_avg_5":4.46,"load_avg_15":4.54,"temperature_celsius":48,"temperature_source":"x86_pkg_temp"}}}
--- FAIL: TestR490_SystemInfoFallsBackToTheDefaultStoragePath (0.15s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.158s
FAIL
exit=1
@@ -0,0 +1,11 @@
MUTATION in /mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/api/router.go:
- if cleared, err := r.sett.ClearUpdateHold(name); err != nil {
+ if cleared, err := false, error(nil); err != nil {
=== RUN TestR491_RemovalClearsTheUpdateHold
r474_remove_wiring_test.go:87: removeStack must clear the app's update hold (R-491), or a reinstall starts held
--- FAIL: TestR491_RemovalClearsTheUpdateHold (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.011s
FAIL
exit=1
+6
View File
@@ -281,3 +281,9 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-465** | **`cfg.Paths.HDDPath` — empty on every box — still had six readers; were any inert?** AUDITED 2026-09-13 on demo-hp (registered drive `/mnt/felhom-drives/hdd_1`, `hdd_path` absent, no `FELHOM_PATHS_HDD_PATH`). Five of six fall back before the value matters: `report/builder.go:69` and `monitor/healthcheck.go:35` take `storagePaths[0]`; `web/server.go:740` (`primaryHDDPath`) takes the default storage path; `main.go:511` (metrics) takes the default storage path; `main.go:347` passes it only as the auto-discovery FALLBACK, and discovery seeds from the apps. **One is inert AND unreachable:** `api/router.go:981` (`systemInfo`, `GET /api/system/info`) reads the empty value with no fallback (`hdd_configured:false` forever) — and the endpoint itself is shadowed: the web layer's `ServeSystemAPI` claims `/api/system/*` and answers **404 „ismeretlen végpont"** for everything but the two memory routes (measured live). Its only consumer is the monitoring page's memory-distribution card, which therefore never renders — split out as **R-490**. Conclusion: the global can be deleted once R-490 is fixed; no report field, health check or metric depends on it. | **CLOSED 2026-09-13 — AUDITED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
| **R-479** | **For a bind-data app the Tier-1 route back restored settings only, and the hold did not say so.** Operator ruling 2026-09-13; closed in controller **v0.241.0** (`3e81330`): an app with classified binds walks second drive → off-site → own unit (`UpdateTierOrderFor`), and the hold sentence ends with what the chosen copy holds (`RestoreHold.CopyHolds`). Delivered by the floor (16 s / 18 s). Proven live on demo-hp with a nextcloud throwaway on the registered drive, Tier 2 off: the failed update held it naming „saját meghajtó, … — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem." Red-proof: a layout-blind order fails the bind case. The 2→3→1 preference itself is unit-proven (a live off-site run touches the standing apps' leg and was not done). `audits/v0241-2026-09-13/` | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` |
| **R-483** | **adventurelog photos uploaded but rendered as broken „Uploaded content" on demo-hp (P2, catalog).** Closed 2026-09-13: the k3s ingress routes `/media`, `/static`, `/admin`, `/accounts` to the backend service on port **80** — the nginx inside the backend image that serves Django's `X-Accel-Redirect` media; the catalog routed everything to the frontend. Two catalog cuts (`3172258` router, `ed62cfd` port 80 — the first cut hit gunicorn and returned empty 200s). Applied to the operator's instance through the guarded Update; proven headless (`GET /media/…webp` with a session → 200 `image/webp`, RIFF/WEBP) and **confirmed by the operator in a browser at 21:49** (two photos render). Scripted multipart uploads through the frontend's `/api` proxy still 500 (upstream `RequestContentLengthMismatchError`); browser uploads work — not a template matter. `audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt` | **CLOSED 2026-09-13 — PROVEN-LIVE, operator-confirmed** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` |
| **R-456** | **A partly-dead stack is not a boot orphan, written down nowhere (P3).** Closed in controller **v0.242.0** (`d698ce3`) by pinning the rule: an absent member does not make a stack degraded, a present-but-dead member does (`internal/bootrecon/r456_partly_dead_test.go`). Design unchanged. | **CLOSED 2026-09-14 — PINNED** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
| **R-476** | **The Mentések page dated a Tier-2 copy from the unit manifest, which moves only with the definition (P3).** Closed in controller **v0.242.0** (`d698ce3`): `Tier2Coverage.UnitDataDate` (newest dump in the mirrored unit) is what a refreshed leg names; a PRESERVED package keeps the manifest date (R-403). Unit-proven with the measured shape (manifest 09-12, dump 09-13); live on 9202 two captures under one definition dated the copy by the second capture's dump. Red-proof: the manifest-only date fails. `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
| **R-481** | **No scratch guest on demo-hp for the nightly rotation (P2).** Operator ruling 2026-09-13, option 1. **BUILT the same night:** LXC **9202** `demo-hp-scratch` on demo-hp, restored from the vouched golden 0.236.0 onto a `dir` storage re-added at `/mnt/hdd_1` (`nvme-scratch`), sized like 9201 (7 cores / 25 898 MB / 32 G + 70 G, unprivileged), the demo-hp customer seeded with hub OFF, tunnel OFF, agent OFF, off-site OFF, self-update OFF; image set by hand (allowed only there); a claimed `settings.json` with the demo password and a scratch second drive. **It persists on purpose.** Disposition in all three layers: the guest (`/etc/felhom-scratch-disposition`), the host (`pct` description, tags `scratch,r481`, the bootstrap file) and the hub side (`operations/nodes.md`, this row — the hub has no customer notes field and the guest is outside the felhom pool). Two traps written into nodes.md: `pct restore` wants the golden as a `backup` volume; directories made on the raw volume from the host must be chowned to 100000. The rotation restarted from bentopdf on it. `audits/nightly-2026-09-13b-bentopdf/14-scratch-guest-built.txt` | **CLOSED 2026-09-13 — BUILT, persists** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
| **R-487** | **A removed app whose backups were kept was listed on neither backup page (P2).** Closed in controller **v0.242.0** (`d698ce3`): the local lists are keyed on the DRIVES the way R-237 keyed the off-site list on the store — `ListRemovedAppUnits` walks `backups/primary/` on the system path and every connected registered drive; the Mentések page lists the unit after the deployed rows („Eltávolítva — visszaállítható", one action), the Visszaállítás picker lists it in its own group, `GET /api/backup/snapshots` answers for it, and the restore opens the unit where it sits (`primaryUnitDirFor` — a unit kept on a data drive was unreachable before, the fallback named the system path). Proven live on the scratch guest 9202 with an opengist throwaway: removed with data, backups kept → row + picker + API answered; „Visszaállítás a mentésből" reinstalled it running. Red-proofs: lister inert, picker 404, wrong unit dir, row not built, row not rendered, picker not rendered — all fail. `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN-LIVE** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
| **R-490** | **The monitoring page's memory-distribution card never rendered — `/api/system/info` was 404 (P3).** Closed in controller **v0.242.0** (`d698ce3`): an exact-path mount ahead of the web layer's `/api/system/` prefix, and `systemInfo` reads the default storage path like every other reader of the empty global. Live on 9202: 200 with the drive figures. Red-proofs: mount removed, fallback removed — both fail. **The global's deletion stays deferred → R-492.** `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN-LIVE** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
| **R-491** | **Removing an app left its update hold in the store, so a reinstall started held (P2).** Closed in controller **v0.242.0** (`d698ce3`): `removeStack` clears an UPDATE hold (`Settings.ClearUpdateHold`, never an R-379 restore hold), logged. Proven live on 9202: a held opengist removed → the store no longer carries the hold, the app redeployed without refusal. Red-proof: the removal without the clear fails the wiring test. `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN-LIVE** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
+2 -7
View File
@@ -682,7 +682,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
| **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 | **READY — rank P3-LOW; owner: CC** |
| **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** |
| **R-456** | **[P3-LOW] A partly-dead stack is not a boot orphan, and that is written down nowhere.** MEASURED 2026-09-02 on demo-hp while validating v0.233.0: `docker rm -f bookstack` (leaving `bookstack-db` running) then a controller restart produced `Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]` — **bookstack was NOT selected**, although the app container was gone and `desired_state: running` was recorded. Removing `bookstack-db` as well made the whole stack orphaned and the very next pass repaired it in 6.3 s. **So `bootrecon.isBootOrphan` requires the stack as a WHOLE to be down; one live member is enough to make it invisible to the reconciler.** **NOT called a defect, and the reason is part of the row:** `StateDegraded` IS in `IsDownState`, and the crash-loop/dead-app alarm path (`classifyRunStates`) does count a degraded stack as down — so the customer IS told; it is the automatic REPAIR that does not fire, and there may be a good reason (repairing half a stack while its DB is live is not obviously safe). **What is certain is that nobody has written the rule down**, so the next session re-derives it the same way this one did — by watching a reconciliation not happen, which is an absent observable and the weakest possible evidence. Either state the rule in `02-controller-module-map.md` with a test pinning it, or change it. Owner: **CC.** `tests/VALIDATION-update-slice12-2026-09-02.md` §2.2 **HALF DONE 2026-09-13:** the rule is written in `02-controller-module-map.md` ("Boot recovery reads desired"); the pinning test is owed to the next controller release (tonight's was spent on v0.240.0). | **READY — rank P3-LOW; owner: CC** |
| **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** |
| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 | **READY — rank P3-LOW; owner: CC** |
| **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 | **READY — rank P3-LOW; owner: CC** |
@@ -694,13 +693,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). | **READY — unblocked by R-448; rank P3-LOW; owner: CC (removal is a deliberate act)** |
| **R-476** | **[P3-LOW] The Mentések page names a Tier-2 copy's date from the unit MANIFEST, which moves only when the app's DEFINITION changes — so it can undersell a fresh copy by a day or more.** MEASURED on demo-hp 2026-09-13: bookstack's mirror held `bookstack-mariadb.sql` written 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z; the Tier-2 run succeeded at 01:30Z. A capture rewrites the manifest only when checksums, dump NAMES or controller version change, and nightly dumps keep their names. R-403 chose the package date so a PRESERVED package is never shown as fresh — correct — but for a normal run it is the older, flattering-in-reverse date. The update (slice 4) deliberately ages the copy by the last successful copy instead (`Tier2RestorePoint.ProvenCopyTime`), so the two can name different dates. Fix shape: record the unit's DATA time (newest dump mtime) in the manifest, or name `LastSuccess` when the leg was not preserved. | **READY — rank P3-LOW; owner: CC** |
| **R-481** | **[P2-MEDIUM] There is no scratch guest on demo-hp, so the nightly rotation cannot restore a throwaway "into a scratch guest", and the nine standing apps cannot be tonight's throwaway at all.** MEASURED 2026-09-13 (`pct list` / `qm list` on demo-hp: only 9201). The rotation brief needs a second controller guest for two steps — the cross-guest restore, and a throwaway deploy of an app that is already standing on 9201 (the stack name collides, and the standing apps may not be touched). Every earlier cross-guest walk built a whole appliance from the published ISO (VM 323/325, hours each) and enrolled it as a new customer; a `pct clone` of 9201 would carry demo-hp's identity, tunnel and hub enrolment. **Decision-shaped:** which route makes the scratch guest — a persistent second LXC on demo-hp born from the golden template and enrolled as its own customer (`nightly-scratch`), or an ISO-built appliance per night. Until then the rotation restores in place (remove → restore from the unit on the same guest) and skips the standing nine (`runbooks/nightly-rotation.md`). **RULED 2026-09-13:** a second enrolled LXC on demo-hp, disk on the NVMe path (`/mnt/hdd_1`), sized like 9201, enrolled as a scratch CUSTOMER with its disposition recorded. **MEASURED BEFORE BUILDING — the ruling collides with the product's own model:** the hub's `hosts` table keys a host to exactly ONE customer (`host_id` PK, `customer_id NOT NULL`) and the box runs ONE agent with ONE `host_id`; a second customer on the same Proxmox host would need a second host identity, and the installer (`felhom-host-install.sh --customer-id … --vmid …`) run with another customer-id on demo-hp would rewrite the existing agent's identity — i.e. break the standing demo-hp enrolment. So "enrolled as a scratch customer" is not something the product can do on a box that already belongs to a customer. **What the product CAN do today, two options:** (a) **a second guest of the demo-hp customer** (the `guests` table is per host+vmid; the agent's provision writes a per-vmid bootstrap): supported by the data model, but both controllers report as customer demo-hp and the customer page shows one controller — the standing box's monitoring flips between the two unless the scratch guest's controller is told no hub (unenrolled scratch, disposition recorded on demo-hp's customer page); (b) **a separate scratch HOST** — a nested Proxmox VM on demo-hp (the ISO appliance route the earlier walks used, VM 323/325) enrolled as its own customer: fully enrolled, fully isolated, but an appliance to build (hours) and keep. Disk placement is the same under both: the installer has no rootfs-storage flag (only `--rootfs-grow`, `--archive-storage`), so the guest lands on `local-lvm` and is moved with `pct move-volume` to a `dir` storage created at `/mnt/hdd_1` — a post-provision step, reversible. 9201's shape for "sized like 9201": 7 cores, 25 898 MB, rootfs 32 G + mp0 70 G on local-lvm, unprivileged. **Recommendation: (a) with the hub left out** — it is the reversible one and it gives the rotation its restore target tomorrow; (b) if "enrolled" is the point. Not built tonight: §1 of the rules — a decision that changes what the product promises about hosts is not CC's. | **WAITING-ON-OPERATOR — the ruling cannot be executed as stated; two options below; rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
| **R-487** | **[P2-MEDIUM] A removed app whose backups were kept is listed on NEITHER backup page, so the restore that brings it back has no button — the customer's remove-by-mistake route exists only as an endpoint.** MEASURED 2026-09-13 on demo-hp (nightly rotation, `adventurelog` removed with backups kept, unit + mirror on disk): `GET /backups/apps` and `GET /backups/restore` contain the string `adventurelog` zero times; `POST /backup/restore stack_name=adventurelog snapshot_id=helyi` then restored it in 22 s with the data byte-identical. Cause: `buildAppBackupRows` walks `status.AppDataInfo` = `DiscoverAppData` over DEPLOYED stacks only. The off-site list had exactly this defect and was fixed by keying it on the store (R-237, v0.204.0); the local and Tier-2 lists were not. **Fix shape:** list every app with a recovery unit on a registered drive (`ListRestorePoints` over the primary dirs), marking removed ones „eltávolítva — visszaállítható"; the unit restore already reinstalls (R-253). Not a design reversal — the same rule R-237 set. Evidence: `audits/nightly-2026-09-13-adventurelog/05b-restore-tier1.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
| **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** |
| **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). | **READY — rank P3-LOW; owner: CC** |
| **R-490** | **[P3-LOW] The monitoring page's „Memória-eloszlás" card has never rendered: its `fetch('/api/system/info')` is answered 404 by the web layer's `ServeSystemAPI`, which claims all of `/api/system/*` and knows only the two memory routes.** MEASURED 2026-09-13 on demo-hp: `GET /api/system/info` → `{"error":"ismeretlen végpont","ok":false}`; `monitoring.html` shows the card (`display:none` by default) only when that fetch returns `used_mem_mb`, so it stays hidden on every box. The API router's `systemInfo` handler (`internal/api/router.go:981`) is unreachable, and it is also the one reader of the always-empty `cfg.Paths.HDDPath` with no fallback (R-465). **Fix shape:** let `ServeSystemAPI` fall through to the API router for unknown `/api/system/*` paths (or route `/api/system/info` explicitly), give `systemInfo` the same default-storage-path fallback the other readers have, and pin the card with a render test; then delete the global (R-465's deferred deletion). Next controller release. Evidence: `audits/nightly-2026-09-13-adventurelog/` (audit notes in the R-465 closure). | **READY — rank P3-LOW; owner: CC** |
| **R-491** | **[P2-MEDIUM] Removing an app leaves its update hold in the store, so a reinstall under the same name starts HELD.** MEASURED 2026-09-13 on demo-hp (v0.241.0, R-479 live check): nextcloud was held after a failed update, then removed with data and backups; `settings.json` still carried `restore_holds.nextcloud` (`reason: update_failed`, `copy_tier: 1`). `fillHoldReason` hides the sentence for a not-deployed app (R-480), but every start gate reads the store, so the next deploy of `nextcloud` would be refused as held with a sentence about a backup that no longer exists. Cleared by hand with `-clear-restore-hold`. **Fix shape:** `removeStack` clears an UPDATE hold (never an R-379 restore hold, which stays operator-cleared) — same place the prefs are forgotten; a test that deploys after a held removal. Next controller release. | **READY — rank P2-MEDIUM; owner: CC** |
| **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). **PARTLY SHIPPED in v0.242.0 (`d698ce3`), measured live on 9202 the same night:** the difference is computed and a fresh compose-created volume IS reported (`["opengist_opengist_data"]`), but the listing filters on the compose project LABEL and a volume recreated by a unit restore (`docker volume create <name>`, `restore.go:154`) carries no labels — compose still removes it and the response says `[]` (`audits/v0242-2026-09-14/19-R489-cause.txt`). **Remaining fix:** list by the `<project>_` name prefix as well (union), or label the recreated volume as compose would. | **READY — rank P3-LOW; owner: CC (residual)** |
| **R-492** | **[P3-LOW] `cfg.Paths.HDDPath` is empty on every box and still has readers; delete it.** R-465 audited its six readers and found every one falling back; R-490 (v0.242.0) gave the last one, `systemInfo`, the same fallback. The global now carries no information on any box and its deletion was deferred twice. **Fix shape:** remove the field, its env binding and the readers' fallback branches; a build proves nothing reads it. Next controller release. | **READY — rank P3-LOW; owner: CC** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
+19
View File
@@ -185,6 +185,25 @@ to demo-hp's below (swap the host_id). So a lost N100 key is recoverable the sam
Details, and the two hard rules that apply to any host node, in `operations/tailscale.md`.
### demo-hp scratch guest — LXC 9202 `demo-hp-scratch` (R-481, built 2026-09-13)
- **Disposition:** the nightly rotation's throwaway host. Under the **demo-hp customer** (same domain,
same dashboard password), **hub reporting OFF, Cloudflare tunnel OFF, agent local API OFF, off-site
OFF, self-update OFF** (the image is set by hand in `/etc/felhom-controller-image`). **Persists across
nights on purpose.** Apps on it are throwaways; nothing on it is a customer promise. Not in the felhom
pool, so the hub never sees it. Tags `scratch,r481`; the same text sits in the guest at
`/etc/felhom-scratch-disposition`, in the host's `pct` description and in its bootstrap file.
- **Reach:** `https://192.168.0.114` with the demo-hp Host names (`felhom.enkisfelhom.hu`,
`<sub>.enkisfelhom.hu`) — LAN only. Helper: the scratch twin of `ctl.sh`.
- **Shape:** 7 cores, 25 898 MB cap, rootfs 32 G + data 70 G, both on the `nvme-scratch` dir storage
at `/mnt/hdd_1` (the runbook's NVMe path; a `dir` storage was re-added there). Unprivileged.
"Second drive" for Tier 2: host `/mnt/hdd_1/scratch-drives/scratch_hdd` mounted at
`/mnt/felhom-drives/scratch_hdd` (mp8), registered as the default storage path.
- **Rebuild from nothing:** restore the golden (`local:backup/felhom-golden-<ver>.tar.zst`) with
`pct restore … --storage nvme-scratch`, add mp0/mp8/mp9 as above, fix ownership from the host with
the unprivileged mapping (100000), seed `controller.yaml` (hub/tunnel/agent/off-site off) and a
claimed `settings.json`, then start `felhom-controller-bootstrap.service`.
## What is NOT enrolled here (deliberately)
- **No PBS datastore, no offsite target** on demo-hp; the DR tier is the N100's.
+5 -5
View File
@@ -7,13 +7,13 @@ guarded Update if the catalog offers one, removal with "delete my data", and a c
surprise is a register row before the next step. An app with no non-browser front door is recorded,
not faked. Tick with the date; the next night takes the first unticked app.
**The nine standing apps of demo-hp come first, and they cannot be tonight's throwaway on guest 9201:**
a throwaway deploy shares the stack name with the standing one, and the standing apps may not be
touched. They are ticked only when a scratch guest exists to host the throwaway (R-481).
**The nine standing apps of demo-hp come first.** They cannot be a throwaway on guest 9201 (same stack
name), so they are walked on the **scratch guest 9202** (R-481, built 2026-09-13: under the demo-hp
customer, hub off, tunnel off, `https://192.168.0.114` with the same Host names and password).
## Standing on demo-hp (blocked on a scratch guest — R-481)
## Standing on demo-hp — walked on the scratch guest 9202
- [ ] bentopdf
- [x] bentopdf — 2026-09-13 (night 2, scratch guest 9202): install, front door 200, no data route (browser-side toolbox — recorded, not faked), backup now + Tier 2, remove-keep → Tier-2 restore back, same-version Update (backed up first: the restore's deploy time is newer than the copies, R-478 by design), remove-all clean. No new rows. `audits/nightly-2026-09-13b-bentopdf/`.
- [ ] bookstack
- [ ] calibre-web
- [ ] docmost