22439b0e43e04ea60e676befc9ceef4eb23095dc
18 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
22439b0e43 |
v0.262.0 + v0.262.1 live on 9202: four scenarios proven, R-630/R-633/R-621/R-614 closed
gates / gates (push) Successful in 30s
A (R-630, the P1): paperless-ngx, same app same button - failed at +313.0s with the app stopped under v0.261.0, done at +53.4s now, with no "no probe container" warning because the explicit healthcheck.container resolved the target. B (R-633): the busy guard fires - RemoveStack REFUSED (busy): a backup or restore is running - and the live proof caught it answering HTTP 500, because router.go maps remove errors by grepping the error TEXT. v0.262.1 makes it a typed error and a 409; re-proven live. C (R-634 half): a half-state with app.yaml on disk and deployed=false answered 200, leftovers NONE. Under v0.261.0 the same call said "not deployed". F (R-614): phase done before the remove, no phase at all after redeploying the same name. Also: the first B run proved NOTHING and nearly went down as a pass - the refusal came from the pre-existing "still running" check, not the new guard. Recorded. 09 6.1 and 8.8, the capability map, and STATUS updated. R-625 and R-634's mechanism are named as owed, not half-done. Register 325. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
186546d562 |
THE TWENTY-EIGHT: every app no drill had touched, walked in one night
gates / gates (push) Successful in 27s
All 28 walked on scratch guest 9202 against the private drill catalog. 26 deployed, 6 proven, 5 inconclusive, 14 with no upstream edge, 1 failed honestly (outline 1.9.1->1.10.1, HELD with the right sentence), 2 undeployable - one (plant-it) by design, refused by the lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a restore from its own copy with the seed read back again - 21 restored, 2 correctly REFUSED per 07 6.2. R-630 RAISED TO P1 by measurement: a stack with NO probe container does not skip verifying - it waits out the full health timeout and HOLDS, stopping an app whose three containers read healthy. The controller's own words: "not healthy within 5m0s (last: no probe container)". R-633 opened: a remove sent during a restore reports success and leaves a container restarting with a live public route. The product already refuses that clash for update and for restore, naming the blocker; remove has no such guard. R-634 opened: an app can be running, healthy and serving while recorded as deployed=false, and is then unremovable. Reproducible alone on sparkyfitness; concurrency-linked on two others. R-631 and R-632 CLOSED. Register 321 -> 323. Seven interventions, six of them my own harness - named, with what each cost. No product code. The live catalog's image: lines are byte-identical to the start of the night. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
a975cfde5b |
probe fix, the gate, and the promotion train (R-618 closed, R-630..632 opened)
gates / gates (push) Successful in 28s
Part 1: tandoor/zipline/wger probes corrected in the catalog and red-proofed live on 9202 in both directions - "Nem egeszseges" with the front door serving 200, then "Fut" after the real sync with no redeploy. tandoor's failed edge re-walked: done at +41.1s where it was failed at +361.9s. Part 2: fifteen proven versions on the live catalog, one commit per app; the guarded Update pressed on four apps on demo-hp, all four done. Opened: R-630 (paperless-ngx's probe has never run on any box - a silent absence, worse than the wrong probe that was found in one night), R-631 (five templates no static rule can judge), R-632 (28 of 53 templates never deployed by any drill). Closed: R-618. Register 318 -> 321. No product code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
462ab4a5ff |
Part 0: repair the register, gate its shape, and record Hetzner's answers (R-627, R-628, R-629)
gates / gates (push) Successful in 29s
THE REPAIR. Last night's append regex ate the state cells of R-446 and R-458, left them as a stray
fourth cell on duplicated copies of R-626 and R-625, and split the table with blank lines. The
register read 317 rows for 315 findings. Both cells restored from the cells that carried them, the
two duplicates deleted, and 15 blank lines that split the register into 12 separate markdown tables
removed. Every row's text is byte-identical afterwards, proven by diff; no row added or removed.
THE RED-PROOF FOUND AN OLDER INSTANCE: R-254 lost its state cell on 2026-08-08 (
|
||
|
|
8d786f7940 |
Update night 2026-09-21: the full record, twelve rows, and the answers to five of the seven questions
gates / gates (push) Successful in 28s
The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:` line is proven identical to before. WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN front door, with a negative control on every readback. Ten of the fourteen printed a verbatim migration line. Up from the three apps this project had ever measured. THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples across five minutes, docker's own healthcheck green, and was then stopped and the household sent to a restore they did not need. zipline and wger are the same defect, both confirmed live. The gate that catches all three is static and cheap: both health checks already sit in the same file. WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) — which needed a purpose-built image store, because the rule that makes automatic updates safe is the same rule that refuses the obvious way to break one. MariaDB across a major through the real button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted, with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at once and all ended honest. And the two EARLY power-cut phases nobody had cut in. TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English household in English and they do not, and R-446/R-458 are both narrower than their rows state. Two instrument fixes were needed before anything could be trusted: the unattended caller turned every success into a timeout (R-623), and one of my own reproductions was wrong and is kept labelled with what it actually measured. Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond the floor the operator asked for. Gates: repo_gates.py --fast, all 15 OK. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
c85262111c |
The update arc's two missing measurements, the lock, and the floor to 0.260.0
gates / gates (push) Successful in 23s
Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all down or blocked; both demo boxes SERVED. Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s after the cut decision and the seeded data read back intact — so the branch that is one step from old-binary-on-migrated-database is now evidence, not argument. Instrument limit stated: `starting` lasts under a second; all three landed in `verifying`, which RecoverUpdates handles in the same branch. Part 3 (R-611) — the night the previous session skipped without saying so. An app updated with nobody pressing anything; a terminally-refused app was pressed exactly once and never again over three passes. The unattended HOLD was NOT produced: the within-a-major rule correctly refused the broken edge before it was attempted, so Q4 still rests on the attended hold from slice 4. Said plainly rather than implied. Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard), R-614 (stale update phase survives a redeploy). R-520's pointer corrected. Catalog: two drill pairs, both reverted; every image line byte-identical to ff9717d3. The alpine:3.20 negative control a security review flagged is cleared. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
0c263c77f2 |
Update arc resumed: the state measured, R-524/R-520/R-589/R-469 closed, seven questions put to the operator
gates / gates (push) Successful in 24s
Phase 0 — measured, never estimated: - both demo boxes: 10 apps, 0 behind, 0 unknown - 46 of 58 exact catalog pins are behind upstream; 39 within a major, 7 across - 6 of 7 measurable floating pins have been repushed since the catalog set them (R-446 is no longer theoretical) - the "23 of 66 floating pins" figure repeated in four places was STALE; recounted to 10, with the definition written down beside it Three claims in the brief corrected, named first: - R-589 was NOT open — it shipped in v0.258.0; only the row was stale - the chaos-night canary is NOT a defect — both gates refused to certify by design - the hub half of the report confirmed, with the nuance that the raw payload is stored whole, so Slice 7 is cheaper than the row implies Closed: R-524 (controller v0.260.0, proven live in both languages), R-520 (power cut during a REAL version change — the pin goes back, the app runs, the page says so), R-589, R-469 (MariaDB half). Filed: R-605, R-606. R-462's stale scope corrected. 09 gains §3 decision 10 (decided by CC unattended — operator may reverse), §3b with the seven questions in the decision shape, §6.2/6.3 the two open slices, and §6.4 an update night costed from R-462's real numbers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
72ee053a9e |
evening: rules file in place; R-483 CLOSED (catalog, operator-confirmed); R-479 CLOSED (controller v0.241.0, proven live); R-481 blocked on the one-host-one-customer model; R-491 filed
gates / gates (push) Successful in 18s
|
||
|
|
321770d9d6 |
R-452 CLOSED: the catalog-since gate (hook-enforced); 09 §8.2 limitation lifted; STATUS note updated
gates / gates (push) Successful in 18s
|
||
|
|
5e8a82c3c4 |
night 2026-09-13/14: first "be a customer" rotation (adventurelog) — 7 defects found, 13 rows closed
gates / gates (push) Successful in 18s
New runbooks/nightly-rotation.md; observations_gate.py reads every section (R-471); target-selection.md names real paths (R-461); R-93 carries the fact that drill-r50 is gone. Register: R-473/R-474/R-466/R-471/R-453/R-461 and v0.240.0's R-477/R-478/R-480/R-482/R-484/R-485/R-486 closed; R-481, R-483, R-487, R-488, R-489 opened. 09 §6.1, 07 §6, CONTEXT, STATUS note. Evidence: audits/nightly-2026-09-13-adventurelog/, audits/v0240-2026-09-13/. |
||
|
|
681c3d6a6d |
docs: rulings 7 and 8 shipped and proven live (R-470/R-472/R-475 CLOSED); R-477..R-480 opened
gates / gates (push) Successful in 21s
Hub v0.112.0 serves a floor above the golden with a declared MinAgent;
controller v0.239.0 reached both demo boxes by that floor in 14 s and 15 s
and updates on any backup tier. 09 §3 decisions 7 and 8, §6/§6.1; 07 §6
line; capability map row; STATUS items 15/16 done and the cadence line
corrected; CONTEXT; register: R-470/R-472/R-475 compressed to CLOSED-ITEMS
(full text at
|
||
|
|
5ef0f52bcd |
Slice 4 shipped (R-448/R-443/R-439 CLOSED, proven live); R-472..R-476; the floor-between-bakes claim corrected
gates / gates (push) Successful in 19s
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure. Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/). Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the runbook, STATUS, CONTEXT, R-468 and the gate docstring. Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
ae59c31a84 |
R-459 CLOSED (MariaDB converts itself, proven by harness + live), golden 0.236.0 (R-467), the golden waiver (R-468)
Operator rulings 2026-09-13, both shipped the same day: - MariaDB finishes its own conversion (catalog eec1228/bd32830/3525e35). Harness E3/E3b `proven` with engine_state_after "already upgraded to 12.3.3-MariaDB [exit=1]", the skip line gone, C3 still `failed`; landed on demo-hp through the real 15-min cycle, nothing recreated, one deliberate restart logged "MariaDB upgrade not required" with the app serving. Evidence: documentation/audits/r459-close-2026-09-13/. The engine-major rule + gate keep every engine inside its major until Slice 4 (R-448) — removal tracked as R-469. - Goldens on a cadence, not per release. golden_currency_gate.py reads a dated waiver (documentation/tests/golden-waiver.yml, <= 14 days, row-bound): valid + BEHIND -> loud advisory, exit 0; expired -> red again naming the date; UNRECORDED (R-385) never covered; malformed -> 2, never 0. Tests cases 5-15 incl. the R-421 decoy; red-proof old-vs-new on the real behind tree. R-242's vouch half stays open. Cadence in RUNBOOK-manual-build.md §4.2 + the checklist. - Golden 0.236.0 baked, round-tripped, vouched, floor raised 0.232.0 -> 0.236.0 (documentation/tests/golden-0.236.0-2026-09-13/) — the last per-release bake; the waiver was issued AFTER it landed. No --no-verify anywhere in this session. Rows: R-459 CLOSED, R-467 CLOSED, R-242 narrowed; R-468/R-469/R-470/R-471 opened. 09 §3 gains decisions 5 and 6; STATUS items 11 and 12 closed; CONTEXT records the cadence ruling. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
d6837d98ee |
SPIKE R-459: the skipped MariaDB conversion is stable, and the trade it implied does not exist
gates / gates (push) Successful in 20s
Outcome A, qualified. Not B and not C. It does not degrade: 5 of 5 restarts of 12.3 on an 11.6 datadir, readback passed every time, mariadb_upgrade_info unchanged, the entrypoint line never escalated past [Note]. It also never heals - the engine answers 'Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!' on every start and will forever. The trade R-459 was expected to produce is not real. Converting properly SUCCEEDS across the multi-major jump, takes 7 seconds, backs up the system database unasked - and putting 11.6 back afterwards STILL starts and serves the data. So the operator is being handed a cheap correction, not a choice between a correct engine and a reversible one. The exit-code polarity was measured rather than read: 0 means the upgrade IS needed, 1 means it is not. Assuming either the flag name or the polarity would have inverted the headline. And run without credentials the same command returns a confident-looking FATAL ERROR that is an auth failure. R-464: after converting and going back, the entrypoint prints 'MariaDB upgrade not required' on a state the same engine calls an unsupported downgrade. The obvious cheap instrument for R-459 would have been to grep for that line, and it would have reported fine for the broken case. R-463: the PostgreSQL analogue, deliberately NOT measured here. 11 templates, 8 on postgres:16-alpine, register grep for pg_upgrade returns zero. The two engines fail in OPPOSITE directions - MariaDB skips quietly, Postgres refuses to start - so that one cannot hide; it presents as eight apps down at once. No template changed. Teardown all three layers, hub checked rather than asserted, local-lvm 30.53 percent before and after. |
||
|
|
a1a6c73fe1 |
SPIKE: an upgrade test that runs again — and a real defect in our own bookstack template
gates / gates (push) Successful in 19s
R-449. Until today one app upgrade out of 53 had ever been measured, by hand, and the whole update arc was designed against that single data point. C3 first: the negative control, whose TO image exits immediately, came back failed. That is what makes the greens mean anything, and it cost 556s because a negative is only honest if it waits out the full settle window. Seven edges, three apps. All five real catalog upgrades kept the customer's data. The finding that changes an assumption the arc was carrying: whether an upgrade can be UNDONE is a property of the individual APP, not of upgrades. Docmost refuses - 'corrupted migrations: previously executed migration 20260213T085259-notifications is missing' - and privatebin does not. That reproduces the Nextcloud result on a second app by a DIFFERENT mechanism, so the struck word 'rollback' now rests on two measurements instead of one. The finding nobody was looking for, R-459: our own bookstack template moves MariaDB across a major and sets no MARIADB_* env at all, so the engine logs that the datadir upgrade it requires is being skipped, and serves anyway. The cause is assigned rather than guessed - the app half alone produces no upgrade line, both edges that move the engine produce it - which is exactly what decomposing E3 into E3a and E3b was for. It also explains why E3's abort looked like it worked: the datadir was never converted. Whether that ever breaks is NOT established, and the row says so. Also opened: R-460 (bookstack's file half cannot be seeded headlessly), R-461 (target-selection.md names a venue that does not exist and fences a VM that is gone), R-462 (the widening, costed with this run's real numbers - and the cost is dominated by fixtures, which do not amortise). Teardown all three layers, hub checked rather than asserted. local-lvm read 30.50 percent before and after. The capability map was deliberately NOT edited: this measured apps, not the product. |
||
|
|
417df06f35 |
slice 3 docs: the ruling, the shipped mechanism, and four rows closed
gates / gates (push) Successful in 17s
09-update-architecture.md gains the fourth dated operator ruling (2026-09-06, Option 1) and its section 5 is rewritten from a proposed shape into the shipped one: the pin, the stored definition, the render table, the four writers, the startup ordering, and the trap this slice set for slice 2 - the live compose file is now the frozen one, so a badge comparing against it would answer Naprakesz on exactly the apps that are behind. 02-controller-module-map.md said 'copy compose + .felhom.yml'. That stopped being true today, so it is corrected, and the two sections describing the old seam now carry a banner saying they describe v0.234.0 and below - kept because every box under v0.235.0 still behaves that way and because they are the measured account of why it changed. R-447, R-441, R-438 and R-455 closed and compressed into CLOSED-ITEMS; R-458 opened for the .felhom.yml asymmetry, with what would settle it by measurement. Live evidence: two real catalog pushes travelling the real 15-minute cycle, both reverted, the tree byte-identical afterwards. The restart that used to take 18.3 seconds and pull a new image now takes 0.1 seconds and pulls nothing. |
||
|
|
bc47dd4ef9 |
v0.234.0: a known limitation written on 2026-09-02 was a defect by the next morning
gates / gates (push) Successful in 18s
The operator looked at demo-felhom and found OpenGist - up 15 hours, running exactly the catalog pin, showing no badge at all. 09-update-architecture.md had recorded that as an accepted limitation the day before: 'the fleet view fills in gradually'. On a quiet box gradually means never, and a feature that fills itself in on an event nobody triggers is, on the quiet installations, not shipped. That limitation row is now struck with the reason kept. The living document gains slice 1b, the two admission rules of the backfill (it never overwrites, and it refuses to seed a partial observation because the badge reads a service-count mismatch as BEHIND), and the note that the same field having two writers with two different admission rules is deliberate. Live evidence added: all nine apps already had records by the time 0.234.0 was ready, so the natural fleet state could no longer exercise the new code - said plainly rather than papered over. The pre-0.233.0 shape was recreated on demo-hp by stripping two records; the backfill re-seeded exactly those two with digests matching independently-read ground truth and left the other seven alone. The refusal half was deliberately NOT staged live: it needs a degraded app, and manufacturing one risks the false-customer-email class that already cost 61 mails (R-330). Unit-tested with a red-proof, and recorded as unproven-live. R-457: a test that hardcodes a date and asserts an age derived from it is green only on the day it is written. Mine was, and it went red overnight. Six other files carry both a date literal and time.Now() - named as candidates, not accused. |
||
|
|
6035dfcc3a |
09-update-architecture.md: the update path finally has a document, and it is a living one
gates / gates (push) Successful in 17s
R-438's document half. It records how an update works AS MEASURED, quotes the RestartStack comment that proves the restart half was CHOSEN (a design decision is not a defect), carries the three operator rulings of 2026-09-02, strikes the word 'rollback' (once a migration has run the old image will not start), states the target shape, and lists the seven slices with a status each. R-438 and R-440 amended and BOTH STAY OPEN: the mechanism is documented, not changed. Nothing closed, so CLOSED-ITEMS.md is untouched. Eight new register rows, 194 -> 202: R-446 (Naprakesz can be false for the 23 floating pins), R-447..R-451 (one per remaining slice, with a rank and an owner), R-452 (no gate enforces catalog_since - the runner fetches at --depth 1), and R-453 (the vaulted dashboard password is stale on BOTH demo boxes, which is what stopped the badge render from being validated live). Live evidence for slices 1 and 2 in documentation/tests/. The record is PROVEN LIVE through the boot reconciler on demo-hp - one entry per compose service, digests matching ground truth read independently. The badge RENDER is not, and the five attempts are listed rather than summarised. |