Files
felhom.eu/REPORT-release-2026-10-10.md
T
admin e0be6e7cd6
gates / gates (push) Successful in 6m13s
release 2026-10-10: CI runner restart (operator yes) evidence; report
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-10 11:11:31 +02:00

64 lines
4.9 KiB
Markdown

# REPORT — 2026-10-10: the waiting fixes released (a security hole closed), and the kernel approval rule
| Part | What | Result |
|---|---|---|
| A1 | Hub 0.145.0 (operator present) | Deployed 07:51Z, `Deployment/hub` only; `/healthz` 200, `/system` 200; **live check of the hole: PASS** |
| A2 | Controller 0.305.0; agent | Controller on demo-hp, demo-felhom, Tester 1 (floors 07:58:50Z, all three healthy by 07:59:08Z). Agent: no change since 0.154.0 — no release |
| A3 | MAIL-HOLD after the deploy | Marker absent, no banner, 0 MAIL-HOLD log lines |
| B | Kernel approval rule (`09` §3 decision 195) | Built, red-proved, deployed; the System page offers **7.0.14-22** — not clicked |
**Register: before 137 · after 137 · opened 0 · closed 0.** R-921 and R-922 marked released; R-925 got two measured
consequences.
## A1 — hub 0.145.0
- Contents: the SECURITY fix (`/preferences`, `/notify` refuse another household's key), R-922 (`email_cleared`),
MAIL-HOLD, two log lines without the address, the kernel rule. `go test ./...` rc 0; CI 1624 (code), 1625 (manifest).
- **Image push:** DooPlex's saved registry login is stale since the R-925 rotation (token endpoint 401 for it, 200 for
the current one) — the push used a one-off `DOCKER_CONFIG` login in the scratchpad, password file→stdin, shredded after.
- **Deploy:** the `felhom` app also showed 3 Secrets + `Deployment/umami` OutOfSync (R-925's de-gitting); a whole-app sync
would have pushed git's view over the rotated Secrets, so only `Deployment/hub` was synced. Pod image 0.145.0, log
`felhom-hub 0.145.0 starting`.
- **Live check (two channels):** Tester 1's key (read file→file from its controller.yaml, 64 chars, never printed):
own household, own settings → **200**; demo-felhom's household, demo-felhom's own settings → **403 „customer_id does
not match the key"**. Second channel: both households' stored rows hashed from a hub DB copy (with -wal/-shm) before
the deploy and after.
- **My mistake, corrected:** the control request stored Tester 1's event list as `["null"]` — my baseline script read
the stored JSON `null` as a list holding the word null. No real event was enabled by it. Restored with the same
own-household request carrying `enabled_events: null`; both rows then **equal the pre-deploy baseline** (hashes).
## A2 — controller 0.305.0
R-921 pre-check + R-922 `email_cleared`. CI 1627; image in the registry (anonymous 200). Floors 0.305.0 with MinAgent
0.131.0 for demo-hp, demo-felhom, tester-1; global floor and Tester 2 untouched. Read-back per box: sudo
`felhom-priv-apply controller-image 9201` → `WROTE … 0.305.0` → agent „new controller healthy" (host journal) and
`docker ps` 0.305.0 healthy (guest).
## B — the kernel rule
`KernelStatus` now offers the newest kernel in every ring-0 box's set of kernels booted healthily after a night stage;
per box and kernel only the newest ended step counts. 6 tests (`kernel_approval_test.go`); against the old function two
fail (today's case: „booted different kernels" where -22 was due; and the fell-back case). Behaviour change to note: a
box whose NEWEST step fell back on -23 no longer blocks the button — -22 (healthy earlier) is offered instead.
Live: the System page shows „Approve kernel set" (2 packages, first seen 07:52Z); the hub DB's candidate is
`proxmox-kernel-7.0` + `proxmox-kernel-7.0.14-22-pve-signed`. `11` §5.11 and `09` §3 updated; poster facts: no fact
changed (it says only „the operator approves on the System page").
## Instruction-file edit (rule 5)
`skills/felhom-build-deploy/SKILL.md`, hub step 4: added — read what is OutOfSync before syncing; if more than
`Deployment/hub`, sync only it (command given); a 401 push from DooPlex is the stale login (R-925). Why: both measured today.
## CI runner restarted (operator yes in chat)
The runner took no job after 07:55:50Z; job 1628 (this session's docs push `a406efc7`) was dropped twice the R-887 way
(every step `failure`, no runner task). With the operator's yes: `rollout restart deploy/act-runner` 08:50:59Z; the new
runner declared at 08:51:02Z, took task 1642 at 08:51:29Z, and job 1628 ended **success**
(`audits/release-2026-10-10/ci-runner-restart.txt`).
## Not done
R-922 live (a real household clear) and R-921 live (a two-tier night) — not exercised today. No release of the agent.
## Decisions for the operator
1. **Refresh DooPlex's registry login** (`docker login gitea.dooplex.hu` with the new admin password, once, at your
keyboard; or a scoped push token). Pick: do it. **If you do nothing:** every build session must use a one-off login,
and a session that does not know this stops at the push.
2. **„Approve kernel set" for 7.0.14-22** — yours to click. Pick: click it (both demo boxes booted -22 healthily; ring 1
still takes it only by a signed job). **If you do nothing:** ring 1 gets no kernel; tonight demo-hp moves to -23, and
once it boots -23 healthily the button will offer -23 instead.