Files
felhom.eu/REPORT-rulings-2026-10-01.md

7.3 KiB
Raw Permalink Blame History

REPORT — the operator's three rulings of 2026-10-01 built (day session)

Evidence: documentation/audits/rulings-2026-10-01/ (A–D, T, tools) · documentation/audits/retest-2026-10/ (the monthly run) · golden: documentation/tests/golden-0.285.0-2026-10-01/. Architecture read: 09 §3 decisions 30, 45, 52, 53 and §6.5; 03-host-agent.md (the controller swap) with agent internal/localapi/controllerswap.go; 07 (R-698); runbooks/monthly-floating-retest.md; audits/night-rulings-2026-09-30/. Baselines (live Gitea ~07:10 CEST): controller a70c398 (0.284.2), agent d766666 (0.138.0), felhom.eu 3159892, catalog efd492d. Register 383 rows by register_shape_gate's method; highest id R-748; last decision 53.

The Part table

Part done / not done / changed why
Rulings 54–56 done — recorded first in 09 §3 and CONTEXT (2076bf9) before any work
A — the full re-test done — nextcloud 3b59dfb and sonarr 1a37032 re-tested on both venues and written, pushed through the pre-push gates first start refused in one minute: the script checked the bench for its own files before copying them (R-749, fixed 9e53205)
A1 runbook done — every app by default, --engines-only the switch, the standing brief, cost —
A4 linuxserver cost done — see below —
A5 STATUS standing line done — "last run 2026-10-01, next due ~2026-11-01" —
B — two controller versions done — controller v0.285.0, floor 0.285.0 (MinAgent 0.131.0 declared) measured first: the agent rolls back to the RUNNING image
B3 live done on 9202 (version-order fallback) and both demo boxes (the swap record) 9202's self-update is off (no hub), so the record path was shown on the demo boxes
— R-751 added — a nil-stack panic in the update clean-up, found by the full suite, fixed in the same release a panic in a goroutine ends the controller
C — mealie done — decided by CC unattended (decision 57): SECURITY_USER_LOCKOUT_TIME=1, catalog a4597cd; 9202 proof 120 min no fix stops an hourly renewal without changing the login name (option d, left open)
C other apps done (read, not measured) — four more lockable: R-752 source reading by a sub-agent at each pinned tag
D1 — R-746 done 804884a, red-proofed (unit + live registry) —
D2 — R-744 done 9fc7052, proven on 9202 at 1.10.1 outline has no newer release, so no step
E — release, floor, golden done — demo boxes on 0.285.0 within ~6 s; golden 0.285.0 baked, round-trip identical, vouched; the gate prints OK —

Claims in the brief that turned out wrong (or right)

  1. "previousImage … handed to the agent's SwapController" — wrong. SwapController(ctx, target) carries the target only; previousImage goes into update-state.json. The agent reads /etc/felhom-controller-image when the swap begins and rolls back to THAT — the running image. On all three boxes it equals <image>:<current version>, so the practical outcome matches; the mechanism does not.
  2. "The sweep skips every controller image today" — right, twice over: the sweep's repositories never include the controller's, and it skips any repo containing felhom-controller.
  3. "nextcloud and sonarr still differ" — right, and only those two (bookstack, radarr, code-server did not differ today).
  4. "mealie's lock is per account" — right (login_attemps / locked_at on the user). Also: the lock outlives its hours until mealie's HOURLY job resets the counter, so 1 means 1–2 h (measured 120 min).
  5. "The golden's image is not needed after first boot" — right after the first swap; until then it IS the running image (kept as such). A whole-guest restore brings its own Docker store (mp0 backup=1).
  6. "Is an old controller tag still in the registry?" — no, below 0.213.0 (2026-08-12): 0.201.0 answers 404. Nobody recorded what removed them (R-750). This session deleted no registry tag.
  7. "Live on 9202 … after the release" for the record path — 9202 has self-update off; the record path ran on the demo boxes.

Part A — the run

app result bench box
nextcloud 34.0.4-apache DONE — written 3b59dfb 05:19–05:32 UTC (proven, peak 23.7 %) ~5 min: old digest installed, seeded, the guarded Update ran the re-test step, new digest running, read back, badge "Naprakész"
sonarr 4.0.20 (linuxserver) DONE — written 1a37032 05:38–05:50 (proven, peak 13.9 %) ~3 min, same steps

Monthly cost. ~17 min per app (bench ~13 incl. the 10-minute watch, box ~4) plus ~15 min to set up and tear down. The four linuxserver apps with ladders (bookstack, radarr, sonarr, code-server) are rebuilt weekly, but the monthly run tests only the day's digest — at most 4 re-tests a month from them (~70 min of bench, ~16 min of 9202), not 16. Acceptable.

Part B — images before → after

box controller images Docker images (system df) /var/lib/docker used
9202 5 → 2 3.17 → 3.10 GB (layers are shared) 3.5 → 3.4 GB
demo-hp 9201 84 → 2 (82 deleted) 15.2 → 11.65 GB 18 → 15 GB
N100 9201 76 → 2 (74 deleted) 6.07 → 1.11 GB 6.0 → 1.2 GB

Each kept 0.285.0 (running) and 0.284.2 (previous). Red-proofs (B/B1-red-proofs.txt): previous dropped → 0.284.2 deleted; record ignored → 0.283.1 deleted; no swap check → deleted while swapping; no in-use check → a used image deleted; no success check → a failed swap's image named previous; no nil-stack check → panic. The first in-use red attempt removed a line and did not build; redone with the condition disabled (recorded).

Part C — mealie

Settings (v3.28.0 source): SECURITY_MAX_LOGIN_ATTEMPTS 5, SECURITY_USER_LOCKOUT_TIME 24 (hours), per account, lifted by an hourly job; POST /api/admin/users/unlock exists but the only admin is the locked account. Fix and 9202 proof: see decision 57 (C/C1-mealie-lockout-1h.txt: right password 200 → 5 wrong 401 → 423 → right password 423 for 120 min → 200; a wrong one 401 after). Other apps (R-752): calibre-web-automated (per username, 3/min and 40/day), wger (per IP = traefik's, 30 min, everyone), Grafana (per account, 5 min), BookStack (e-mail|IP, 60 s); gokapi and claper cannot.

Rows

383 → 387. Opened R-749 (re-test start), R-750 (old registry tags gone), R-751 (update clean-up crash), R-752 (four lockable apps). Closed R-743, R-744, R-745, R-746, R-749, R-751. Narrowed R-747.

Teardown

  • Machine: 9202 back on the live catalog (repo_url read back), the same six containers as at the start, controller 0.285.0 (the release); the apps this session installed (nextcloud, sonarr, outline, mealie) removed through the product. Bench 9401 destroyed with its template. Drill VM: CT 9100 destroyed, token/runner/log shredded, qemu exited, reverted to virgin. Drill catalog reset to live (a4597cd), image lines identical.
  • Host: demo-hp pct list = 9201, 9202 (as at the start); the N100 untouched except the floor's controller update.
  • Hub: two form saves — the floor (0.285.0, MinAgent 0.131.0) and the vouch (golden 0.285.0); one floor attempt without the credential answered 302 /login and stored nothing.