Commit Graph

879 Commits

Author SHA1 Message Date
admin 593312daf4 night 2026-09-25: Part C nights 5-6, Part D floor + arrival, Part E bench/box proofs (n8n, mealie), teardown of 9202 and bench
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 23:50:35 +02:00
admin a7971c4822 night 2026-09-25 (in progress): part 7 shipped as controller v0.271.0 — decisions 31-33, evidence A/B/C/F, register (R-685..R-687; closed R-672 R-673 R-684 R-680 R-678)
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 23:04:43 +02:00
admin 75ff26408a tools: liveR669.py (the R-669 live proof)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 17:06:23 +02:00
admin d53bd08442 R-672 session: audit README (not done first, wrong claims), STATUS, capability map, register 344->341, topic REPORT
gates / gates (push) Successful in 28s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 17:06:09 +02:00
admin 54bff69f88 R-672/R-673: 03 + 08 + register + CONTEXT (restore test off, 9201 repaired, agent v0.133.0, hub v0.124.0); R-684
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 16:29:54 +02:00
admin f5e11277f6 r672-2026-09-24: evidence (restore test off, 9201 repair, red-proofs, live cases)
gates / gates (push) Successful in 22s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 16:26:28 +02:00
admin 2d24931597 hub v0.124.0: a thin pool is critical at 90% (data or metadata), one alarm per pool per 6 h (R-672)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 16:26:17 +02:00
admin 1a72fef87f night 2026-09-24: CI by head_sha, final progress line
gates / gates (push) Successful in 26s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 14:56:48 +02:00
admin e924a17632 night 2026-09-24: record the test-password slip; runner state out of the evidence tree
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 14:55:35 +02:00
admin f2234c0fdd night 2026-09-24: drop E/chaos-state.json (drill test-account passwords of removed apps); ignore it
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 14:55:19 +02:00
admin 83e436f609 night 2026-09-24: findings doc, decisions 29-30, STATUS/CONTEXT, register 335->344 (7 closed), teardown evidence, floor 0.269.1
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 14:55:06 +02:00
admin 4502af6bb1 night 2026-09-24: chaos rounds 1-8 evidence; 07/08/09/capability map for decisions 26-28 + digests; R-681
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 14:21:14 +02:00
admin 3208cb2083 night 2026-09-24: Part E schedule + pairing committed before round 1
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 13:34:56 +02:00
admin 6ebe95c13e night 2026-09-24: Part C spike (09 §6.4.2 build brief), demo-hp pool incident evidence, rows R-668..R-680
gates / gates (push) Successful in 41s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 13:30:30 +02:00
admin 1cedf94d6b night 2026-09-24: A1/A2/A3/B live evidence; chaos schedule drawn (seed 20260924)
gates / gates (push) Successful in 29s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 13:05:29 +02:00
admin f853f43b67 night 2026-09-24: evidence so far (phase 0, A1 spike, A3 measurements, red-proofs)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 11:38:16 +02:00
admin 99c0709cbe hub v0.123.0: app_stopped_unhealthy (decision 28); 09 decisions 26-28 recorded
gates / gates (push) Successful in 31s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 11:38:13 +02:00
admin 86c4b9a0d8 2026-09-24: 09 decisions 24-25, part 5 shipped; the whole-copy truth table; fourth suppression; rows; STATUS; evidence
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 08:52:42 +02:00
admin 68bb1ea394 evidence: ladder-2026-09-24 — floor 0.267.0 delivered, first red-proofs
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 07:33:15 +02:00
admin 3e58c184f6 night shift 2026-09-23: the record, the register, the morning note
gates / gates (push) Successful in 27s
DRILL-night-2026-09-23.md: Parts A-E. 09 §3 decisions 21 (operator word),
22 and 23 (CC unattended, operator may reverse); §6.4 parts 4 and 6
(catalog half) shipped; §6.1a residuals R-658/R-659. Register 330 -> 336:
R-651..R-660 opened (R-658 and R-659 P1), R-650/R-640/R-499/R-626 closed.
Capability map, nightly rotation (opengist), STATUS (one question: the
floor), CONTEXT, REPORT. The floor stays 0.266.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 00:24:04 +02:00
admin 77335625fe night 2026-09-23: round 11's second update is romm's engine step (changed before round 1)
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 22:05:05 +02:00
admin cf8dec8480 night 2026-09-23: the chaos schedule (seed 20260923) and its app table, written before round 1
gates / gates (push) Successful in 26s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 22:04:37 +02:00
admin 11e37ee807 night 2026-09-23: Part C evidence (11 moves across 10 apps so far)
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 22:02:17 +02:00
admin 3898580fa8 night 2026-09-23: evidence for the wishlist/opengist/ghost/romm moves
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 21:42:24 +02:00
admin 0030cbd1de night 2026-09-23: evidence so far (Parts A, B, C in progress)
gates / gates (push) Successful in 26s
The records the catalog's ladder entries cite (apps/<app>/bench and
apps/<app>/verdict.json), the drill tools, the spike, R-626's run,
R-612/R-613 red-proofs, the chaos schedule (seed 20260923) drawn
BEFORE round 1. Docs and register follow at the end of the night.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 21:33:54 +02:00
admin 899ae18b95 R-649 closed: a failed install removes what it started (operator ruling); floor 0.266.0
gates / gates (push) Successful in 29s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 18:14:38 +02:00
admin 210394ec6e Clean-up evening: R-634 cause fixed, held badge, OOM storm; floor 0.265.0
gates / gates (push) Successful in 26s
R-634, R-625, R-636, R-647, R-648 closed; R-649 (operator question) and
R-650 opened. Open rows 334 -> 331. 08 §6.2 storm rung; 09 §6.4 parts
8-9 SHIPPED. Evidence: audits/cleanup-2026-09-23/.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 17:40:01 +02:00
admin 2644f9c5e7 R-634: the mechanism, diagnosed before any code (whole-box backup stops and restarts a DEPLOYING app)
gates / gates (push) Successful in 26s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 16:41:57 +02:00
admin 21f17ed32b The household is told: controller v0.264.0 + hub v0.120.0 proven live, floor 0.264.0
gates / gates (push) Successful in 28s
09 §6.4 parts 2-3 SHIPPED. R-606, R-620, R-646 closed; R-647 (three
leftovers) and R-648 (whole-box backup press in the harness) opened.
Open rows 335 -> 334. Evidence: audits/undo-fleet-2026-09-23/.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 14:39:31 +02:00
admin 05ea21e918 The undo, built and proven live: controller v0.263.2 (09 decision 15)
gates / gates (push) Successful in 25s
- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two
  live-only defects, §6.4 part 1 SHIPPED.
- Capability map: a failed update is undone by the box - PROVEN-LIVE.
- Live evidence on 9202: three apps undone by the product with seeds before
  the backup, after it and seconds before the press read back; cut-off copy
  held honestly; power cut during the undo resumed; manual press after undo.
- Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643
  ruled; R-646 opened. STATUS asks the floor question.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 12:25:13 +02:00
admin 5a349d9884 Undo bake-off: copy the folder wins (09 §3 decisions 19-20, §6.1a)
gates / gates (push) Successful in 26s
- 09 §3: decision 19 (the copy method is chosen by a bake-off) and 20 (the
  full-system backup waits for the update leg, inside its window; built later).
- Bake-off on 9202, docmost / romm / vikunja: both methods pass every case;
  the folder copy wins because an app with no database server gets no dump,
  so dump-and-load would need the folder copy anyway. 1-5 s extra downtime,
  ~420 MB/s, disk = the volumes.
- R-645 filed: lifting an update hold by hand lets the recovery unit be
  re-captured with the failed definition within seconds.

Documents and evidence only; product code follows in the controller.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 10:28:56 +02:00
admin 4c92beab8f Update arc: the undo and ladder spiked, the build plan for the 2026-09-23 rulings
gates / gates (push) Successful in 26s
- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost
  (PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came
  back with data written before AND after the backup. The product's loader
  cannot do it: over a migrated PG database it fails on the new tables'
  foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads
  rc 0 into an empty database. No-DB apps have no last-second copy.
- Part 2: one press jumps A -> C; the box's catalog clone is depth 1.
  Ladder format recommended: update_ladder in .felhom.yml, not git history.
- Part 3: memory watch red-proof results (harness change in the catalog repo).
- Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643).
- Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT.

No product code. Live catalog untouched; 9202 back on it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 08:54:41 +02:00
admin 805ad1e962 docs(update arc): record the operator's rulings of 2026-09-23 (09 §3 decisions 11-18)
gates / gates (push) Successful in 29s
- 09 §3: decisions 11 (one window = a leg of the backup chain), 12 (automatic,
  per-box switch on by default), 13 (the test decides, not the tag - replaces
  decision 3's "never across a major"), 14 (the ladder), 15 (the box undoes a
  failed update - replaces §6.1's no-auto-undo), 16 (Postgres majors converted
  by the box), 17 (digests), 18 (fleet view, later).
- 09 §3b marked ANSWERED with a pointer per question; kept as the reasoning.
- 09 §6.2 rewritten to the ruled shape; §6.1 abort paragraph and §4 point at 15.
- Register: R-450, R-451, R-446, R-463 cite the decisions.
- STATUS: the seven questions no longer wait on the operator.

Documents only. No product code.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 07:55:04 +02:00
admin 267dcad01b floor raised 0.261.0 -> 0.262.1, MinAgent 0.131.0 declared
gates / gates (push) Successful in 26s
Both demo boxes took it in under 20 seconds and read healthy. drill-r50 is DOWN and correctly HELD:
its agent is 0.129.0, below the declared MinAgent, so the hub refuses to hand it a controller it
cannot run (R-472's guard, working).

Boxes below the floor: 5 -> 3. The three that remain are drill-r50 and the two guests the hub does
not hear from, including scratch 9202 which runs hub.enabled: false.

Read back from the hub rather than from the POST's own answer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 22:10:38 +02:00
admin 22439b0e43 v0.262.0 + v0.262.1 live on 9202: four scenarios proven, R-630/R-633/R-621/R-614 closed
gates / gates (push) Successful in 30s
A (R-630, the P1): paperless-ngx, same app same button - failed at +313.0s with the app stopped
under v0.261.0, done at +53.4s now, with no "no probe container" warning because the explicit
healthcheck.container resolved the target.

B (R-633): the busy guard fires - RemoveStack REFUSED (busy): a backup or restore is running - and
the live proof caught it answering HTTP 500, because router.go maps remove errors by grepping the
error TEXT. v0.262.1 makes it a typed error and a 409; re-proven live.

C (R-634 half): a half-state with app.yaml on disk and deployed=false answered 200, leftovers NONE.
Under v0.261.0 the same call said "not deployed".

F (R-614): phase done before the remove, no phase at all after redeploying the same name.

Also: the first B run proved NOTHING and nearly went down as a pass - the refusal came from the
pre-existing "still running" check, not the new guard. Recorded.

09 6.1 and 8.8, the capability map, and STATUS updated. R-625 and R-634's mechanism are named as
owed, not half-done. Register 325.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 21:53:53 +02:00
admin 1de6aaf904 corrections: calibre-web restore was REFUSED, calcom was INCONCLUSIVE
gates / gates (push) Successful in 25s
Both from their own records, not re-run.

calibre-web: the flash_error refusal is in its log; that walk predated the refusal-capture code.
The report's own prose already said so while the table disagreed.

calcom: the restore was accepted and the app read running; the classifier then saw "starting", a
settling state it did not list beside running/unhealthy, and fell through to failed. The brief
supposed a read-back artefact - that is wrong, and restore_state_seen says so.

Restores correctly refused: 2 -> 3.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 20:18:37 +02:00
admin 737694c603 CORRECTION: the app_oom alarm DID fire - R-635 was wrong, R-636 opened
gates / gates (push) Successful in 26s
The operator produced the mails. The controller HAS an OOM detector (main.go:821), it emits
app_oom (notifier.go:726), the hub allow-lists it (dispatcher.go:636) and delivered it to the
OPERATOR channel - two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard.
CUSTOMER skipped, correctly.

I asserted an absence without opening the hub's Events or Notifications tab, reasoning instead from
a memory note that said the signal was UNPROVEN - not that it was missing. That is R-628's shape
again, from the same hand, four days later.

R-636: the real defect is the signal's SHAPE. notifier.go:715-724 keys on container|startedAt and
emits once per container lifetime, so 4530 worker kills over six hours produced exactly one
warning-level mail - indistinguishable from one transient kill. The magnitude was already collected
(App Telemetry: RomM 5023 errors, 632 warnings) but nothing turns it into a louder event.

Memory note lxc-docker-oom-signals-unreliable corrected with the positive reading.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 20:13:41 +02:00
admin 8efd2d00df romm OOM storm on demo-hp: fixed, measured, closed (R-635)
gates / gates (push) Successful in 26s
Found because the operator heard the fans. romm 5.3.0 was promoted that morning; the update read
`done` and the app ran clean for two hours, then OOM-crash-looped for six - 4530 worker SIGKILLs,
~500% CPU, host load 5.2 while otherwise idle, and nothing alarmed.

Raising the limit to 768M was still a guess and fixed nothing (memory.peak hit exactly 768 MiB).
Measured instead: ~216 MiB per warm uvicorn worker, so the image's default of 4 workers needs
~882 MiB. /init reads WEB_SERVER_CONCURRENCY; set to 2.

Proven under load, not just at idle: 26,645 requests over 300 s, memory 416-614 MiB against 768,
trending down, zero SIGKILLs, OOMKilled false. Idle CPU 500% -> 1.64%.

The first soak measured nothing - it was pointed at the scratch-guest subdomain, every request
404'd at traefik in 9 ms, and the counter reported 14,026 successes. Positive and negative controls
are now asserted before any load is driven.

Carried into R-462: `proven` has meant "the update applied and the data survived", not "the new
version runs".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 18:04:49 +02:00
admin 186546d562 THE TWENTY-EIGHT: every app no drill had touched, walked in one night
gates / gates (push) Successful in 27s
All 28 walked on scratch guest 9202 against the private drill catalog. 26 deployed, 6 proven,
5 inconclusive, 14 with no upstream edge, 1 failed honestly (outline 1.9.1->1.10.1, HELD with the
right sentence), 2 undeployable - one (plant-it) by design, refused by the lifecycle gate, proven
live for the first time. Each app also got the half the update night skipped: a restore from its own
copy with the seed read back again - 21 restored, 2 correctly REFUSED per 07 6.2.

R-630 RAISED TO P1 by measurement: a stack with NO probe container does not skip verifying - it
waits out the full health timeout and HOLDS, stopping an app whose three containers read healthy.
The controller's own words: "not healthy within 5m0s (last: no probe container)".

R-633 opened: a remove sent during a restore reports success and leaves a container restarting with
a live public route. The product already refuses that clash for update and for restore, naming the
blocker; remove has no such guard.

R-634 opened: an app can be running, healthy and serving while recorded as deployed=false, and is
then unremovable. Reproducible alone on sparkyfitness; concurrency-linked on two others.

R-631 and R-632 CLOSED. Register 321 -> 323. Seven interventions, six of them my own harness -
named, with what each cost. No product code. The live catalog's image: lines are byte-identical to
the start of the night.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 16:00:56 +02:00
admin a975cfde5b probe fix, the gate, and the promotion train (R-618 closed, R-630..632 opened)
gates / gates (push) Successful in 28s
Part 1: tandoor/zipline/wger probes corrected in the catalog and red-proofed live on 9202 in both
directions - "Nem egeszseges" with the front door serving 200, then "Fut" after the real sync with
no redeploy. tandoor's failed edge re-walked: done at +41.1s where it was failed at +361.9s.

Part 2: fifteen proven versions on the live catalog, one commit per app; the guarded Update pressed
on four apps on demo-hp, all four done.

Opened: R-630 (paperless-ngx's probe has never run on any box - a silent absence, worse than the
wrong probe that was found in one night), R-631 (five templates no static rule can judge),
R-632 (28 of 53 templates never deployed by any drill). Closed: R-618.

Register 318 -> 321. No product code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 11:10:58 +02:00
admin 462ab4a5ff Part 0: repair the register, gate its shape, and record Hetzner's answers (R-627, R-628, R-629)
gates / gates (push) Successful in 29s
THE REPAIR. Last night's append regex ate the state cells of R-446 and R-458, left them as a stray
fourth cell on duplicated copies of R-626 and R-625, and split the table with blank lines. The
register read 317 rows for 315 findings. Both cells restored from the cells that carried them, the
two duplicates deleted, and 15 blank lines that split the register into 12 separate markdown tables
removed. Every row's text is byte-identical afterwards, proven by diff; no row added or removed.

THE RED-PROOF FOUND AN OLDER INSTANCE: R-254 lost its state cell on 2026-08-08 (59527d0) and had
rendered without a State column for 45 days. Its own verdict sentence was the cell; it has it back.

THE GATE. scripts/register_shape_gate.py, gate 14 in repo_gates.py and reached by the pre-push
hook: a row that does not end with `|` (an eaten state cell), a duplicated id, or a blank line
splitting the table. Four decoys in test_gate_decoys.py — three convicting on the exact damage
shapes, one asserting a healthy register still passes. The red-proof corrected the gate twice: a
first draft counted CELLS and convicted 125 innocent rows (register cells carry literal `|` in
prose and shell snippets, so a row cannot be split on `|`), and it skipped malformed rows before
counting ids, reporting 5 duplicates where there were 2. It then immediately caught a blank line
left by this session's own next insert.

R-618's rank now reads P1-HIGH in both its title and its state cell.

HETZNER ANSWERED, AND I FIRST SAID THEY HAD NOT. The operator supplied ticket #2026090103040671.
Q1: with the MAIN account, files and directories can be downloaded from a snapshot (Storage Box
docs govern, not the Storage Share FAQ); a restore reverts the whole box. Q2: --append-only is
enforced by pinning `command="rclone serve restic --stdio --append-only path/to/repo"` to the key
in authorized_keys — so append-only IS expressible on a Storage Box, which is what R-95 was blocked
on. Neither is measured; R-436's due check now asks for the measurement, not the question.

My error is filed as R-628 and is precise: the query was fine — a control returns 201 threads and
reaches SENT and TRASH — and the thread genuinely is not in this mailbox. What was invented was the
step from "absent here" to "Hetzner has not answered". An absent record in one place cannot answer
what someone else did. The rule is now in the Gmail-access memory.

R-629: the drill repo inherited has_actions from the migrate call and mailed the operator 47 CI
failures overnight, into the mailbox that was carrying real off-site alarms. Actions disabled and
verified; 09 §6.5 now makes it a step of creating a drill repo. The 47 mails are the operator's to
clear: subject:"gates FAILED in admin/app-catalog-drill".

Gates: repo_gates.py --fast — all 16 OK, including the new one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 10:22:37 +02:00
admin d27663dd91 Update night: record the one later catalog commit, and re-prove no image line moved
gates / gates (push) Successful in 26s
The teardown was taken before the vikunja verb correction (test code only), so the live catalog's
main now reads d4392e2a10f4 rather than 4463243f2e09. The check that matters was re-run after it:
git diff f5f6a152b513 origin/main -- templates/ is EMPTY. Not one image: line moved on the live
catalog at any point in the night.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 23:14:52 +02:00
admin 1ee14ce166 Update night: the final PROGRESS lines — teardown, documents, CI green by id
gates / gates (push) Successful in 24s
Both CI runs confirmed by job id: app-catalog-felhom.eu job 830 and felhom.eu job 870, each
'completed' / 'success', matched on head_sha. unproven.py --summary did not move and that is
stated rather than left for the reader to notice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:35:42 +02:00
admin 8d786f7940 Update night 2026-09-21: the full record, twelve rows, and the answers to five of the seven questions
gates / gates (push) Successful in 28s
The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:`
line is proven identical to before.

WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own
guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at
any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN
front door, with a negative control on every readback. Ten of the fourteen printed a verbatim
migration line. Up from the three apps this project had ever measured.

THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not
answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by
STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples
across five minutes, docker's own healthcheck green, and was then stopped and the household sent
to a restore they did not need. zipline and wger are the same defect, both confirmed live. The
gate that catches all three is static and cheap: both health checks already sit in the same file.

WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) —
which needed a purpose-built image store, because the rule that makes automatic updates safe is
the same rule that refuses the obvious way to break one. MariaDB across a major through the real
button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted,
with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at
once and all ended honest. And the two EARLY power-cut phases nobody had cut in.

TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was
measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English
household in English and they do not, and R-446/R-458 are both narrower than their rows state.

Two instrument fixes were needed before anything could be trusted: the unattended caller turned
every success into a timeout (R-623), and one of my own reproductions was wrong and is kept
labelled with what it actually measured.

Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond
the floor the operator asked for.

Gates: repo_gates.py --fast, all 15 OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:34:24 +02:00
admin 9c69b3ff07 Update night: Phases 2-4 evidence — both engines, the unattended HOLD, and five new findings
gates / gates (push) Successful in 27s
Evidence off the machine at the end of the phases that produced it (R-320). Teardown follows.

PHASE 2 — the two database engines, through the REAL Update button:
- MariaDB 11.6 -> 12.3 on nextcloud: PROVEN, and pressed through the button for the first time.
  All four SPIKE-r459 observables: the datadir's own record moved 11.6.2 -> 12.3.3; the engine
  itself says "already upgraded ... no need to run mariadb-upgrade again"; the entrypoint says
  "Major version upgrade detected ... Check required!" and then STARTED and FINISHED it (not the
  `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); and the engine took its own
  pre-upgrade backup, 631 905 B. The seeded Nextcloud account read back.
- PostgreSQL 16 -> 17 on docmost: FAILED exactly as R-463 predicted and nobody had measured.
  5.1 s to held; the pin named 17 while nothing ran; the restore brought it back in 29.1 s.
  The engine's REFUSAL LINE was destroyed by failAndHold before any probe could read it, so it
  was REPRODUCED INDEPENDENTLY with a control on every step (R-320).

PHASE 3 — the bad days. B1 produced THE UNATTENDED HOLD, which this project has never had: the
caller pressed once with nobody watching, the app held after 312.9 s, and passes 2 and 3 pressed
nothing. B2 put the pin back on a pull failure in 1.0 s. B3 refused `busy` six times. B4 showed
there is NO single-flight — 5 of 5 updates ran at once and all ended honest. B5 cut the power in
`backing-up` and the box recovered itself and said so. B7 refused under the 2 GB floor. B9 found
R-458's risk narrower than the row states.

PHASE 4 — every badge on the box is TRUE, and the held app answers all four of Q4's questions.

FINDINGS, five new and three corrections to existing rows. The one that matters: R-618 is P1 —
two templates name a health probe the app does not answer, and because the guarded update waits
on that same probe, a SUCCESSFUL update ends by STOPPING a working app. Measured: tandoor served
HTTP 200 on the new version at four samples across five minutes and was then stopped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:13:57 +02:00
admin da20722e76 Update night 2026-09-21: Phase 0 and Phase 1 evidence, the drill method, and two instrument fixes
gates / gates (push) Successful in 27s
INTERIM CHECKPOINT — evidence off the machine at the end of the phase that produced it (R-320),
not at the end of the session. Phases 2-5 follow in a later commit.

Phase 0, all three mechanisms proven with their controls:
- the fleet floor to 0.261.0 with its declared MinAgent — both demo boxes in 13 s, the hub
  logging `managed floor SERVED ... from declared (golden 0.258.0)`.
- a PRIVATE DRILL CATALOG (admin/app-catalog-drill), so that broken, dummy, cross-repo and
  engine-major edges can be measured without the live catalog ever carrying one. Positive
  control quoted, and two negative controls: the live catalog's main and both real boxes'
  caches unchanged.
- a throwaway image store on the scratch guest, which is what makes an UNATTENDED HOLD
  measurable at all: an edge that PASSES the within-a-major test and still fails.
  CompareImageRefs was proven to order host:port/ references by RUNNING it (4 positive cases
  + 1 negative control), not by reading it.

Phase 1: real within-a-major upstream edges walked on guest 9202 through the product's own
guarded Update, each app seeded and read back through its OWN front door (R-156), with a
per-edge verdict record in 09's shape. `inconclusive` is never collapsed into `failed`.

TWO INSTRUMENT FIXES, both in this repo's own evidence code:
- 00-api-recipe.md said the app page is /app/<n>; it is /apps/<n>, and every call it described
  404s. Corrected, with the session-expiry note that cost the same time.
- unattended-caller.py's follow() read update_phase/updating off the API ENVELOPE, so both were
  always None and EVERY followed update ran to its 900 s timeout and was then recorded
  `timeout` and never-press-again. Fixed before B1 relied on it. R-623.

No controller, agent or hub code was written. The live catalog carries no broken reference.

Gates: repo_gates.py --fast — all 15 OK, exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 21:17:46 +02:00
admin c85262111c The update arc's two missing measurements, the lock, and the floor to 0.260.0
gates / gates (push) Successful in 23s
Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all
down or blocked; both demo boxes SERVED.

Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and
two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live
compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s
after the cut decision and the seeded data read back intact — so the branch that is
one step from old-binary-on-migrated-database is now evidence, not argument.
Instrument limit stated: `starting` lasts under a second; all three landed in
`verifying`, which RecoverUpdates handles in the same branch.

Part 3 (R-611) — the night the previous session skipped without saying so. An app
updated with nobody pressing anything; a terminally-refused app was pressed exactly
once and never again over three passes. The unattended HOLD was NOT produced: the
within-a-major rule correctly refused the broken edge before it was attempted, so
Q4 still rests on the attended hold from slice 4. Said plainly rather than implied.

Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh
install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard),
R-614 (stale update phase survives a redeploy). R-520's pointer corrected.

Catalog: two drill pairs, both reverted; every image line byte-identical to
ff9717d3. The alpine:3.20 negative control a security review flagged is cleared.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 15:00:09 +02:00
admin d19f07ea04 File R-607 (a sync that says 'no change' while the cache moves) and sharpen the audit
gates / gates (push) Successful in 26s
Found by the live run, not by reading: POST /api/sync answered 'nincs valtozas'
while the box's catalog cache HAD moved, and catalog_images stayed stale until a
separate rescan. Since CatalogImages is the one input CatalogOrder compares
against, the badge answers from a stale catalog for that window — and the session
nearly recorded a stale tag-ok badge as proof of the R-524 ahead arm.

Neither half is isolated, so the row records the observation, not a diagnosis.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 13:14:10 +02:00
admin 0c263c77f2 Update arc resumed: the state measured, R-524/R-520/R-589/R-469 closed, seven questions put to the operator
gates / gates (push) Successful in 24s
Phase 0 — measured, never estimated:
- both demo boxes: 10 apps, 0 behind, 0 unknown
- 46 of 58 exact catalog pins are behind upstream; 39 within a major, 7 across
- 6 of 7 measurable floating pins have been repushed since the catalog set them
  (R-446 is no longer theoretical)
- the "23 of 66 floating pins" figure repeated in four places was STALE; recounted
  to 10, with the definition written down beside it

Three claims in the brief corrected, named first:
- R-589 was NOT open — it shipped in v0.258.0; only the row was stale
- the chaos-night canary is NOT a defect — both gates refused to certify by design
- the hub half of the report confirmed, with the nuance that the raw payload is
  stored whole, so Slice 7 is cheaper than the row implies

Closed: R-524 (controller v0.260.0, proven live in both languages), R-520 (power cut
during a REAL version change — the pin goes back, the app runs, the page says so),
R-589, R-469 (MariaDB half). Filed: R-605, R-606. R-462's stale scope corrected.

09 gains §3 decision 10 (decided by CC unattended — operator may reverse), §3b with
the seven questions in the decision shape, §6.2/6.3 the two open slices, and §6.4 an
update night costed from R-462's real numbers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 13:13:23 +02:00
admin bcdd5b2058 floor 0.259.0 raised; R-601 withdrawn as FALSE; R-604 filed
gates / gates (push) Successful in 28s
R-601 said demo-hp was unreachable. The operator looked at the hub and said it
was online. It was, and had been up four and a half weeks, reporting every few
minutes. Both of my SSH routes pointed at stale addresses: `demo-hp` at a
tailnet peer for a box that has no tailscale installed at all, and `demo-hp-lan`
at 192.168.0.87 when the box is statically on .104 since a reprovision. The hub
had carried the right address in every host report, and `ip neigh` on felhom-pve
had .104 four lines above the .87 I quoted — I searched that output for the
address I expected instead of reading it for the address that was there.

Both ssh entries repointed and verified; nodes.md corrected, including that the
tailnet route for this box does not exist.

The hunt then found R-604, which is the real defect: demo-hp carried a
per-customer floor override of 0.243.0 left over from the 2026-09-16 drill, so
it had silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0.
`managed floor SERVED` fires once per change by design, so a box behind a static
override is silent for ever and its silence is indistinguishable from a box that
already logged. Cleared; demo-hp self-updated to 0.259.0 in under four minutes
and its claim page now answers "Wrong or expired code" in English.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 09:13:57 +02:00