a975cfde5b33de033e47db1bedbeeff9faadf84c
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1ee14ce166 |
Update night: the final PROGRESS lines — teardown, documents, CI green by id
gates / gates (push) Successful in 24s
Both CI runs confirmed by job id: app-catalog-felhom.eu job 830 and felhom.eu job 870, each 'completed' / 'success', matched on head_sha. unproven.py --summary did not move and that is stated rather than left for the reader to notice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
8d786f7940 |
Update night 2026-09-21: the full record, twelve rows, and the answers to five of the seven questions
gates / gates (push) Successful in 28s
The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:` line is proven identical to before. WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN front door, with a negative control on every readback. Ten of the fourteen printed a verbatim migration line. Up from the three apps this project had ever measured. THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples across five minutes, docker's own healthcheck green, and was then stopped and the household sent to a restore they did not need. zipline and wger are the same defect, both confirmed live. The gate that catches all three is static and cheap: both health checks already sit in the same file. WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) — which needed a purpose-built image store, because the rule that makes automatic updates safe is the same rule that refuses the obvious way to break one. MariaDB across a major through the real button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted, with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at once and all ended honest. And the two EARLY power-cut phases nobody had cut in. TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English household in English and they do not, and R-446/R-458 are both narrower than their rows state. Two instrument fixes were needed before anything could be trusted: the unattended caller turned every success into a timeout (R-623), and one of my own reproductions was wrong and is kept labelled with what it actually measured. Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond the floor the operator asked for. Gates: repo_gates.py --fast, all 15 OK. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
9c69b3ff07 |
Update night: Phases 2-4 evidence — both engines, the unattended HOLD, and five new findings
gates / gates (push) Successful in 27s
Evidence off the machine at the end of the phases that produced it (R-320). Teardown follows. PHASE 2 — the two database engines, through the REAL Update button: - MariaDB 11.6 -> 12.3 on nextcloud: PROVEN, and pressed through the button for the first time. All four SPIKE-r459 observables: the datadir's own record moved 11.6.2 -> 12.3.3; the engine itself says "already upgraded ... no need to run mariadb-upgrade again"; the entrypoint says "Major version upgrade detected ... Check required!" and then STARTED and FINISHED it (not the `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); and the engine took its own pre-upgrade backup, 631 905 B. The seeded Nextcloud account read back. - PostgreSQL 16 -> 17 on docmost: FAILED exactly as R-463 predicted and nobody had measured. 5.1 s to held; the pin named 17 while nothing ran; the restore brought it back in 29.1 s. The engine's REFUSAL LINE was destroyed by failAndHold before any probe could read it, so it was REPRODUCED INDEPENDENTLY with a control on every step (R-320). PHASE 3 — the bad days. B1 produced THE UNATTENDED HOLD, which this project has never had: the caller pressed once with nobody watching, the app held after 312.9 s, and passes 2 and 3 pressed nothing. B2 put the pin back on a pull failure in 1.0 s. B3 refused `busy` six times. B4 showed there is NO single-flight — 5 of 5 updates ran at once and all ended honest. B5 cut the power in `backing-up` and the box recovered itself and said so. B7 refused under the 2 GB floor. B9 found R-458's risk narrower than the row states. PHASE 4 — every badge on the box is TRUE, and the held app answers all four of Q4's questions. FINDINGS, five new and three corrections to existing rows. The one that matters: R-618 is P1 — two templates name a health probe the app does not answer, and because the guarded update waits on that same probe, a SUCCESSFUL update ends by STOPPING a working app. Measured: tandoor served HTTP 200 on the new version at four samples across five minutes and was then stopped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
da20722e76 |
Update night 2026-09-21: Phase 0 and Phase 1 evidence, the drill method, and two instrument fixes
gates / gates (push) Successful in 27s
INTERIM CHECKPOINT — evidence off the machine at the end of the phase that produced it (R-320), not at the end of the session. Phases 2-5 follow in a later commit. Phase 0, all three mechanisms proven with their controls: - the fleet floor to 0.261.0 with its declared MinAgent — both demo boxes in 13 s, the hub logging `managed floor SERVED ... from declared (golden 0.258.0)`. - a PRIVATE DRILL CATALOG (admin/app-catalog-drill), so that broken, dummy, cross-repo and engine-major edges can be measured without the live catalog ever carrying one. Positive control quoted, and two negative controls: the live catalog's main and both real boxes' caches unchanged. - a throwaway image store on the scratch guest, which is what makes an UNATTENDED HOLD measurable at all: an edge that PASSES the within-a-major test and still fails. CompareImageRefs was proven to order host:port/ references by RUNNING it (4 positive cases + 1 negative control), not by reading it. Phase 1: real within-a-major upstream edges walked on guest 9202 through the product's own guarded Update, each app seeded and read back through its OWN front door (R-156), with a per-edge verdict record in 09's shape. `inconclusive` is never collapsed into `failed`. TWO INSTRUMENT FIXES, both in this repo's own evidence code: - 00-api-recipe.md said the app page is /app/<n>; it is /apps/<n>, and every call it described 404s. Corrected, with the session-expiry note that cost the same time. - unattended-caller.py's follow() read update_phase/updating off the API ENVELOPE, so both were always None and EVERY followed update ran to its 900 s timeout and was then recorded `timeout` and never-press-again. Fixed before B1 relied on it. R-623. No controller, agent or hub code was written. The live catalog carries no broken reference. Gates: repo_gates.py --fast — all 15 OK, exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |