Both rulings built: same-tag re-tests (decision 52), image retention + the install hold (controller 0.284.2); golden 0.284.2
gates / gates (push) Successful in 27s

- Decision 52: catalog 6a3ead9 (re-test entries, gates, decoys, the monthly command); proven end to end on 9202
  through the leg; runbook monthly-floating-retest.md; nothing to re-test on the engine lines today.
- Decision 53 + R-741: controller v0.284.2 (0.284.0/0.284.1 never floored — two wiring faults found live on 9202);
  floor 0.284.2; one-time sweep 9202 26.6 -> 5.7 GB, demo-hp 24.3 -> 13.5 GB; the install hold proven as a stranger.
- Golden 0.284.2 baked, round-trip identical, vouched (agent 0.138.0, min_agent 0.131.0); the gate prints OK.
- Rows 377 -> 383: opened R-743..R-748, closed R-736, R-737, R-740, R-741, R-748; narrowed R-739, R-698, R-446.
- register_shape_gate: a lettered id (R-88a) is a row too (R-748), with a decoy seen red.

Evidence: documentation/audits/night-rulings-2026-09-30/, documentation/tests/golden-0.284.2-2026-09-30/.
Report: REPORT-night-rulings-2026-09-30.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-30 23:26:21 +02:00
parent 1de0f7c294
commit 3159892fab
69 changed files with 4642 additions and 58 deletions
+39 -50
View File
@@ -2,65 +2,54 @@
**Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).**
**Updated 2026-09-30 (evening). Both demo boxes run controller 0.283.1 and host agent 0.138.0. Hub 0.126.0. New installs get golden 0.283.1 with agent 0.138.0. No controller, agent or hub release today.**
**Updated 2026-09-30 (late evening). Both demo boxes run controller 0.284.2 and host agent 0.138.0. Hub 0.126.0. New installs get golden 0.284.2 with agent 0.138.0.**
**Tester-2 — read only, from the hub.** The customer record exists. Tester-2's box has not registered yet, so there are no apps to read.
**Tester-2 — read only, from the hub.** The customer record exists. Tester-2's box has not registered yet.
## One decision for you
## Your two decisions of tonight — built
**Question: should a box also take, at night, a security fix that the app's maker ships under the same name?**
Database images like `postgres:18-alpine` or `redis:7-alpine` get fixes without a new name. Seven of the eight such names
we use on Docker Hub got a new push in the last 30 days.
- **Security fixes under the same name (your A).** The catalog can now re-test a database image's new fix and record
it as a tested step. A box then takes it at night by itself. Tested from start to end on the scratch box: the night
run took the fix, the app kept its data, and the page went back to "up to date". **Today there was nothing to re-test**
for the database and redis images.
- **Old program files on the disk (your A).** A box now keeps each app's current program and the one before it, and
deletes the older ones. It never deletes one that another app still uses. On the first start it cleaned up once:
**the HP demo box freed 11 GB**, the scratch box 21 GB. All apps kept running.
Today **no box gets these fixes at all** — not at night, and not by the Update button either. The reason: the catalog
only records a test when the name changes, so the box never learns that a newer, tested image exists. The page says
"up to date".
## What else I did, and it worked
- **Option A — the night takes a tested fix.** The catalog re-tests the same name at the new image (bench and scratch
box, as today), and records it. The box then takes it at night by itself, with its backup and undo. Measured: the box
already does this with no change; only the catalog side is new work (about one evening to build). Cost after that:
about 20 app re-tests a month, machine time only, run by day.
- **Option B — keep it manual, say so honestly.** No new work. The fixes arrive only when we move an app to a new name.
Database engines rarely change name, so their fixes may wait for months. I would correct the text of the earlier
decision, which says the night would take them.
**If nothing is decided:** nothing breaks. Database images stay at the build of the day we tested them.
**Who is blocked:** nobody today. **My recommendation: A**, database and redis names first — that is where security
fixes land, and the box side already works.
## What I did, and it worked
- **More apps update themselves at night: 35 now (plus 3 when a full copy exists), up from 30 (+2) at the start of the
evening.** New: gitea, wger, crafty-controller, uptime-kuma, zipline. calibre-web counts as "with a full copy", because its
update rewrites the book library's own database file.
- **Twelve updates tested and published**, each on the test bench and on the scratch box: calibre-web, gitea, wger,
crafty-controller, uptime-kuma, zipline (their first tested updates; zipline in two steps), and emby, ghost, home-assistant, outline, rallly.
Each of these apps is now on its maker's newest version in its line.
- **calibre-web (books) and gitea now have tests.** calibre-web: a book goes in through its Upload button and is read
back. gitea: its first-run form is filled in the way a household does it.
- **immich's older in-between update is fixed.** It still gave the database 512 MB and it does reload the big list of
places. Tested again with 768 MB and no swap: 0 kills. No box runs immich today.
- **The short login risk is closed.** After an install, an app with a known default password is now closed to
strangers until the box has changed that password. Tested on two apps: no stranger got in.
- **wger's phone app can log in** now (its missing key is made at the first start).
- **New controller 0.284.2** is on both demo boxes, and **a new golden 0.284.2** is baked and vouched. The golden check
is green (not just waived).
## What broke, and what I did
- **wger: an update would have broken it.** After the update the app looked fine, but nobody could log in. Its database
was never upgraded, because the catalog forgot one setting. **Fixed** in the catalog, and tested again: it works. No
box runs wger.
- **zipline: its newest version cannot be reached in one jump.** The scratch box saw it fail and **put the old version
back by itself in 20 seconds** — the undo worked on a real failure. Going through the version in between works; that
is now the path.
- **The scratch box's disk filled with old app images.** Removing an app or updating it never deletes the old image.
On a household box this would slowly fill the disk until updates are refused. Written down; not fixed today.
- **For a few seconds after a fresh install, calibre-web's well-known default login works.** The box changes it right
after, but the app is already reachable. Written down; not fixed today.
- **wanderer cannot run on the test bench**, so it was not tested. Written down.
- **The clean-up first missed most images.** Two mistakes, both found on the scratch box before any demo box got the
release: it did not run on the Remove button, and it could not see images stored without a name. Both fixed and
tested. That is why the version is 0.284.2 (0.284.0 and 0.284.1 never went out).
- **wanderer** can now run on the test bench, but its update is not tested yet: I could not create test data in it.
Also found: its search engine needs one extra setting before any update, or it refuses to start.
- **Two exact versions were rebuilt by their makers under the same name** (nextcloud and sonarr). The monthly re-test
finds them. Tonight I only did the database images, as you ruled.
- **mealie locks the account after five wrong passwords** — also when a stranger types them. Written down.
- **Old controller versions** still pile up on each box (about 50). Your rule covered apps only. Written down.
**Rows.** Tonight: 3 closed (all three found and fixed tonight), 6 narrowed, 8 opened. The list went from 369 to 377 rows. *(The brief said 366; the
afternoon session had already added three.)*
**Rows.** Tonight: 5 closed, 3 narrowed, 6 opened. The list went from 377 to 383 rows (counted with every row, including three whose number has a letter — the old check skipped those).
## What needs you
1. **The decision above** (same-name security fixes). If you do nothing, nothing changes.
2. **The image clean-up on a box** will need your word on what the box keeps (the running image and the one before it,
for the undo). If you do nothing, disks fill slowly; nothing breaks tonight.
3. **No golden bake is due.** The weekly bake stays around 7 October.
1. **Who presses the monthly re-test?** It is one command, but it needs the test bench and the scratch box set up
first, so it cannot run by itself (the steps are written down). Options: CC does it in a monthly session, or you
start it. If nobody presses it, same-name security fixes wait until the next time we move that app.
2. **Should the monthly re-test cover every app, or only databases?** Exact versions get rebuilt too (nextcloud,
sonarr). If you do nothing, only databases are re-tested.
3. **Old controller versions:** keep only the current and the previous one? If you do nothing, about 150 MB per release
keeps piling up.
4. **No golden bake is due** — tonight's bake carries the newest controller.
## Standing steps
- **Monthly:** the same-name security re-test (runbook `monthly-floating-retest.md`).
- **Weekly:** the golden bake (around 7 October).