Files
felhom.eu/documentation/runbooks/monthly-floating-retest.md
T
admin 3159892fab
gates / gates (push) Successful in 27s
Both rulings built: same-tag re-tests (decision 52), image retention + the install hold (controller 0.284.2); golden 0.284.2
- Decision 52: catalog 6a3ead9 (re-test entries, gates, decoys, the monthly command); proven end to end on 9202
  through the leg; runbook monthly-floating-retest.md; nothing to re-test on the engine lines today.
- Decision 53 + R-741: controller v0.284.2 (0.284.0/0.284.1 never floored — two wiring faults found live on 9202);
  floor 0.284.2; one-time sweep 9202 26.6 -> 5.7 GB, demo-hp 24.3 -> 13.5 GB; the install hold proven as a stranger.
- Golden 0.284.2 baked, round-trip identical, vouched (agent 0.138.0, min_agent 0.131.0); the gate prints OK.
- Rows 377 -> 383: opened R-743..R-748, closed R-736, R-737, R-740, R-741, R-748; narrowed R-739, R-698, R-446.
- register_shape_gate: a lettered id (R-88a) is a row too (R-748), with a decoy seen red.

Evidence: documentation/audits/night-rulings-2026-09-30/, documentation/tests/golden-0.284.2-2026-09-30/.
Report: REPORT-night-rulings-2026-09-30.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 23:26:21 +02:00

3.3 KiB

RUNBOOK — the monthly re-test of same-name security fixes (09 §3 decision 52, R-740)

What it does. An image such as postgres:18-alpine or redis:7-alpine gets security fixes under the SAME name. A box takes such a fix at night only when the catalog has re-tested the tag at the new digest on both venues and written it as a ladder step. This runbook is that re-test, once a month. One command does the work; the steps around it set up the two venues and tear them down.

Who runs it: a person or a CC session, from DooPlex, once a month (the operator decides who presses — STATUS). Why not a cron job (measured 2026-09-30): it needs a fresh bench LXC on demo-hp, scratch guest 9202 pointed at the drill catalog, a drill reset (a force-push — the permission check refused it once and the operator allowed it), and pushes to the LIVE catalog. None of that should happen with nobody watching.

1. Look first (read-only, 1 minute)

cd /mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu && git pull -q
python3 scripts/retest-floating.py --dry-run --engines-only   # the ruled start: database and redis lines
python3 scripts/retest-floating.py --dry-run                  # everything, for the record

nothing to re-test today ends the month. Otherwise go on.

2. The bench (LXC 9401 on demo-hp)

The recipe in audits/more-night-apps-2026-09-30/bench/B1-bench-create.txt (60 GB disk — 40 GB filled up on 2026-09-30), swap 0 (the stricter venue, R-733). retest-floating.py syncs the catalog's scripts and templates to it itself.

3. The box (scratch guest 9202)

  1. Reset the drill to the live catalog: git -C /mnt/5_hdd/felhom.eu/drill/app-catalog-drill fetch live && git reset --hard live/main && git push -f origin main (force-push: ask if the permission check refuses).
  2. Point 9202 at the drill (09 §6.5): the repoint.py drill of the latest audit's tools/ (saves controller.yaml, sets the drill URL + credentials, removes the catalog cache, restarts). Quote repo_url read back.
  3. export SC=<a 0600 scratch dir> holding .ctlpw (9202's dashboard password — never committed).

4. The run

python3 scripts/retest-floating.py --engines-only --push \
    --evidence $SC/evidence --evidence-rel felhom.eu/documentation/audits/retest-<YYYY-MM>

Per app: bench (the full method, 10-minute memory watch), box (fresh install at the OLD tested digest, seed, the re-test entry in the drill, the guarded Update, read-back, the running digest must be the NEW one), then the writer, the catalog gates and one commit (pushed with --push; the pre-push gates run). A failure stops that app and never the list; the summary names each app DONE or STOPPED with its reason. Copy $SC/evidence to felhom.eu/documentation/audits/retest-<YYYY-MM>/.

5. Teardown (three layers, stated)

Machine: 9202 back on the live catalog (repoint.py restore), apps the run installed are removed by it. Host: pct destroy 9401 --purge. Hub: nothing touched. Reset the drill again (§3.1).

What proves it works (2026-09-30)

audits/night-rulings-2026-09-30/A/e2e/: docmost at an OLDER redis:7-alpine digest on 9202, the re-test on the bench, "run tonight's chain now" → the leg's docmost: step pressed … step ended done after 95.0 s, the new digest running, the data read back, the badge back to "Naprakész".