version-travel: session record, evidence B/C/D/R, 09 decision 42, 07 §6.5 decision, register 339 -> 336, STATUS (D4)
gates / gates (push) Successful in 26s
gates / gates (push) Successful in 26s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -672,7 +672,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 **— UPDATE NIGHT 2026-09-21:** **MEASURED 2026-09-21 (update night), leg B9, and the row's risk is NARROWER than it states.** A `.felhom.yml`-only change (a health check for a path only a newer version would serve) was pushed to a FROZEN `bentopdf` — installed `v2.8.6`, catalog ahead. §5.4's asymmetry is confirmed live: the new `.felhom.yml` reached the box while the compose `image:` line stayed `v2.8.6`. **But no false alarm was produced**: ten samples over two minutes all read `state=running` with the front door at `200`. The reason is the probe's own semantics, not luck — `healthprobe.go:258-261` treats **any response** as healthy for `type: http`, and the bogus path answers 404, which is a response. **So this row's false-alarm risk exists only for `type: api` probes carrying an `expect` block**, where the status is compared; for every `type: http` template and every `type: api` without `expect`, a newer version's path is invisible to the probe. The row's actual claim — the failure direction is a false alarm, never data loss — stands and is now measured. Evidence: `audits/update-night-2026-09-21/21-B9-frozen-app-newer-felhomyml.md`. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 **-- UPDATE NIGHT 2026-09-21:** bookstack's edge was walked again on 2026-09-21 and is again **half-proven**: the database half read back through `php artisan` with its own negative control, the file half untouched. The limitation is unchanged and is now measured on the box as well as on the harness. Two more apps joined the same class tonight for a different reason (R-624). | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 **ROW CORRECTED 2026-09-21: the scope is NOT open and this row said it was.** It read *"VIKTOR rules on scope"*; the operator ruled on 2026-09-13 (`09` §3 decision 6) that the upgrade test goes to **ALL** apps through the nightly rotation, explicitly *not* "database apps first". What is open is the WORK, not the scope. An ORDER inside that ruling — the 15 database services first, because that is where a wrong answer costs data rather than uptime — is costed as a drill brief in `09` §6.4: legs A–E ≈ **21–34 CC-hours** plus ~25–30 GB on a scratch host, with legs C (a PostgreSQL `pg_upgrade` rehearsal) and E (one automatic night on a throwaway) the two that unblock a decision. **-- UPDATE NIGHT 2026-09-21:** **The count moved from 3 apps to 21 EDGES ACROSS 19 APPS.** The update night walked real within-a-major upstream edges on scratch guest 9202 through the product's own guarded Update, each app seeded and read back through its OWN front door with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive**. Proven: actualbudget, audiobookshelf, bookstack, docmost, grafana, home-assistant, mealie, n8n, navidrome, papra, privatebin, romm, vikunja, and nextcloud's MariaDB engine major. **Ten of the fourteen printed a verbatim migration line**, so the database really was rewritten and the data still read back. Box-side fixtures for 20 apps now exist at `audits/update-night-2026-09-21/fixtures.py`, and four (actualbudget, navidrome, audiobookshelf, vikunja) are ported into `app-catalog-felhom.eu/scripts/upgrade_fixtures.py` with seven new EDGES (U1-U7) so the same edges can be run on the harness venue **with their ABORT step**, which the box deliberately does not offer. **OWED, stated so it is not mistaken for done:** the U1-U7 harness RUNS (the code is in, the runs are not), and fixtures for the four inconclusive apps, of which two (vaultwarden, zipline) cannot be seeded at all while the catalog rightly closes their sign-up (see R-624). **-- 2026-09-23, the RomM lesson is IN THE HARNESS:** `upgrade-test.py` v2 runs every edge that read back under light load for `--soak` seconds (default 600) and reads the kernel's `oom_kill` counter host-side; a kill or restart turns `proven` into `failed`, a peak over 80 % of the limit adds `memory_tight`. **Red-proof:** romm 5.0.0 → 5.3.0 on the template AS PROMOTED (512M, four workers) — seeded, migrated, read back, and then **OOM-killed at +76 s** under four light callers, verdict `failed` (Docker's own OOMKilled read true here; restarts stayed 0, which is why the walk never saw it). Ten minutes is ample for this failure; demo-hp's first kill came at two hours only because nothing was loading it. Evidence `audits/update-rulings-2026-09-23/harness/`. **— NIGHT 2026-09-23:** the bench and the box now share ONE fixture set — the box walk's fixtures run on the bench through `upgrade_boxport.py` (ported verbatim into the catalog), plus a new wishlist fixture and fixes for opengist 1.15, komga (`/api/v2/users/me`) and nextcloud (wait for `occ status`). 14 apps / 15 edges tried; 12 proven on both venues and published with their test records (`audits/DRILL-night-2026-09-23.md` Part C). | **READY — rank P2-MEDIUM; owner: CC (scope already ruled, `09` §3 decision 6)** |
|
||||
| **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 **-- UPDATE NIGHT 2026-09-21:** **Measured 2026-09-21, both halves.** (a) What a household sees TODAY: the guarded Update of `postgres:16-alpine` to `17-alpine` on docmost ended **`failed` in 5.1 s**, the app stopped and held, **the pin naming 17 while `installed_images` still said 16 and nothing ran**, the data intact, and the restore the hold sentence names back in **29.1 s**. The engine's refusal line had to be REPRODUCED independently because `failAndHold` destroyed it (R-621): *FATAL: database files are incompatible with server / DETAIL: The data directory was initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir was still `16` afterwards, and the same copy started under 16 holding 48 tables as the positive control. (b) The conversion **COSTED** on a real seeded 49 MB / 48-table datadir: `pg_dumpall` **2.6 s / 132 201 B**, fresh 17 plus replay **6.5 s / 48 tables restored**, the app up on 17 saying *Database connection successful*, **the seeded account read back**, total **155.9 s of which ~9 s is engine work**. `pg_upgrade` was NOT run: it needs both majors' binaries in one image and no such image exists here. Full paragraph: `audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`. **-- RULED 2026-09-23 (`09` §3 decision 16):** PostgreSQL majors are converted BY THE BOX as a guarded-update step — save everything from the old engine, start the new one empty, load it back, check. Each of the eleven apps is proven on the test bench before the catalog may move it; the engine-major gate stays until then. **-- 2026-09-25 (evening): docmost CONVERTED by the box** (16 → 18, controller v0.273.0; catalog `afd3a60` moved it with the conversion mark; proven on the bench and on 9202 incl. three undo cases). Measured: `postgres:18` refuses even an EMPTY volume at `/var/lib/postgresql/data` — an 18 step moves the mount. Open: the other ten, one proof each (calcom, adventurelog/postgis, claper, rallly, immich's own image, zipline, outline, paperless-ngx, tandoor, sparkyfitness on 15). | **NARROWED — ten apps left; owner: CC; the gate stays per app** |
|
||||
| **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 **-- UPDATE NIGHT 2026-09-21:** **Measured 2026-09-21, both halves.** (a) What a household sees TODAY: the guarded Update of `postgres:16-alpine` to `17-alpine` on docmost ended **`failed` in 5.1 s**, the app stopped and held, **the pin naming 17 while `installed_images` still said 16 and nothing ran**, the data intact, and the restore the hold sentence names back in **29.1 s**. The engine's refusal line had to be REPRODUCED independently because `failAndHold` destroyed it (R-621): *FATAL: database files are incompatible with server / DETAIL: The data directory was initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir was still `16` afterwards, and the same copy started under 16 holding 48 tables as the positive control. (b) The conversion **COSTED** on a real seeded 49 MB / 48-table datadir: `pg_dumpall` **2.6 s / 132 201 B**, fresh 17 plus replay **6.5 s / 48 tables restored**, the app up on 17 saying *Database connection successful*, **the seeded account read back**, total **155.9 s of which ~9 s is engine work**. `pg_upgrade` was NOT run: it needs both majors' binaries in one image and no such image exists here. Full paragraph: `audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`. **-- RULED 2026-09-23 (`09` §3 decision 16):** PostgreSQL majors are converted BY THE BOX as a guarded-update step — save everything from the old engine, start the new one empty, load it back, check. Each of the eleven apps is proven on the test bench before the catalog may move it; the engine-major gate stays until then. **-- 2026-09-25 (evening): docmost CONVERTED by the box** (16 → 18, controller v0.273.0; catalog `afd3a60` moved it with the conversion mark; proven on the bench and on 9202 incl. three undo cases). Measured: `postgres:18` refuses even an EMPTY volume at `/var/lib/postgresql/data` — an 18 step moves the mount. Open: the other ten, one proof each (calcom, adventurelog/postgis, claper, rallly, immich's own image, zipline, outline, paperless-ngx, tandoor, sparkyfitness on 15). **-- 2026-09-27: paperless-ngx (to 18) and tandoor (to 17) moved, each proven on the bench and on 9202 through the product's conversion (`09` §3 decision 42, `audits/version-travel-2026-09-26/B/`); with docmost that is 3 of the 11. Not moved: calcom, claper, outline, rallly, zipline (no seed route checked, or none: R-624), immich and the PostGIS family (out of scope).** | **NARROWED — ten apps left; owner: CC; the gate stays per app** |
|
||||
| **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** |
|
||||
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). **HALF-LIFTED 2026-09-21, catalog `5ff36d098cbc`.** Slice 4 shipped 2026-09-13, so the rule's own expiry condition is met — **for MariaDB**: the four `mariadb:` sidecars have both halves they need, a verified backup in front of the Update (any tier since v0.239.0) and `MARIADB_AUTO_UPGRADE=1` whose conversion the harness WATCHED run on E3/E3b with the seeded data read back after. **PostgreSQL and MySQL stay refused** — postgres performs no `pg_upgrade` and REFUSES to start on an older major's datadir across eleven templates (R-463); a backup is a route BACK, not a conversion. The refusal text now cites R-463 instead of the shipped R-448. **R-450's second half is enforced in its place:** a MariaDB major must be the ONLY image move in its template in that commit (the bookstack `0b73e5e` shape — two migrations behind one edge). The gate now PRINTS what it allowed, by name — a lifted rule that goes quiet is a lifted rule nobody can audit. Two new decoy cases; two red-proofs, each seen to fail; 40 cases green. **What remains of this row:** the PostgreSQL half, which is R-463's to clear — see `09` §3b **Q5**. **-- NARROWED 2026-09-25 (evening), catalog `6a4a5f0`, `09` §3 decision 35:** the PostgreSQL half now passes ONE app at a time — only a template whose ladder entry for the step is proven on BOTH venues and carries `engine_conversion` (the box converts it, controller v0.273.0), as the only image move in its commit. Every other PostgreSQL app stays refused; the postgis family is judged now (it was not). CLAUDE.md rule text updated the same commit. Decoys + red-proof: `audits/night-2026-09-26/C/`. | **PARTLY CLOSED 2026-09-21 — MariaDB lifted; the PostgreSQL half stands until R-463** |
|
||||
@@ -802,19 +802,14 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** |
|
||||
| **R-652** | **[P3-LOW] The memory watch counted the kernel's file cache as the app's memory.** MEASURED 2026-09-23 night on the bench: nextcloud 34.0.4 read **100 %** of its 1 GiB and immich's PostgreSQL **100 %**, each with **0** kernel `oom_kill`s — `memory.peak` includes page cache, which the kernel drops before it kills anything. Under the watch as built (R-635 follow-up) both would be marked `memory_tight`, and the gate would demand a raised `mem_limit` — a customer-box capacity figure — for cache. **Done the same night (09 §3 decision 22, CC — operator may reverse):** the watch samples the app's own memory (`anon` of `memory.stat`) every 15 s; the mark and the ladder's `memory_peak_pct` read it where measured; the cgroup peak stays beside it (`memory_cgroup_peak_pct`). **Open:** romm's backfilled entry carries M1's 80.9 % cgroup peak (measured before the anon sample existed) — re-measure it on its next step; and decide whether an app whose anon is low but whose cgroup stays pinned at its limit (cache thrash) should be marked at all. Evidence: `audits/night-2026-09-23/apps/nextcloud/bench-1024M/`, `apps/immich/bench-noanon/`. | **READY — P3; owner: CC (catalog harness)** |
|
||||
| **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** MEASURED 2026-09-23 night: 1.13 serves `/login`, `/register`, `/all`; 1.15.2 answers **404** on all three and serves `/-/login`, `/-/register`, `/-/all`; `/` redirects to `/-/all`. The app, its data and its probe (`/healthcheck`) are fine, and a household arriving at the root lands correctly — only a deep link breaks. 1.15 also marks its session cookie `Secure`. **Needs:** a line in opengist's `app_info` if the operator wants households told; nothing in the product. Evidence: `apps/opengist-oldfixture/`, `apps/opengist/`. | **READY — P3; owner: operator (copy decision) / CC (writes it)** |
|
||||
| **R-655** | **[P2-MEDIUM] adventurelog v0.13.0 cannot become healthy in the catalog's template — its update is undone on every box.** MEASURED 2026-09-23 night, both venues. v0.13.0's FRONTEND image adds its own `HEALTHCHECK` (`node -e fetch('http://127.0.0.1:3000/health')`), and `/health` answers **503 `{"ok":false,"backend":"unreachable"}`** unless the backend's `/health/` answers OK to the frontend's own request (read from the image's `django-proxy` chunk: `fetch(${getServerEndpoint()}/health/)`). On the bench the backend served `/api/` 200, the seed read back, the migration ran (6 lines) — and the frontend stayed `unhealthy` for 420 s; on 9202 the guarded Update went `verifying` → **`undoing` → `undone`** (under the drill's 90 s timeout). **MEASURED LATER THE SAME NIGHT — two causes, and the first hypothesis was wrong.** (1) The backend's `/health/` answers `200 {"ok": true}` to the frontend (read on a fresh v0.13.0 install); the unhealthy frontend is **our template's own healthcheck override**, `["CMD", "/nodejs/bin/node", …]` — v0.12.1's distroless image keeps node there, v0.13.0 moved it to `/usr/bin/node`, and Docker's health log reads `exec: "/nodejs/bin/node": no such file or directory` 19 times in a row. (2) With the override removed, a second bench run hit a harder wall: v0.13.0's backend runs `download-countries` at EVERY start, fetching world data from the internet, and a cut-short download (`ijson.common.IncompleteJSONError: Incomplete JSON content`) crash-loops the entrypoint — the backend never became healthy in 15 min. The first bench run's download had succeeded. **So v0.13.0's boot depends on an outside download, and whether an update succeeds depends on it too.** **The catalog did NOT move adventurelog.** **Needs:** the override dropped (or pointed at `/usr/bin/node`) in the SAME commit as the image move; and a measured answer on whether `download-countries` can be skipped or pre-seeded (an env switch, or the data in the volume) before the edge is re-proven on both venues. Evidence: `audits/night-2026-09-23/apps/adventurelog/`. **-- MEASURED 2026-09-25 (read from both images on the bench, `audits/night-2026-09-26/F/F3-01-entrypoint.txt`):** v0.13.0's entrypoint has an env switch, `SKIP_WORLD_DATA=1`, that skips `download-countries` (v0.12.1 has none). Both versions use the SAME dataset file, `countries+regions+states-v3.1.json`, kept in the `adventurelog_media` volume, and download only when it is absent — so an UPDATE re-uses the file v0.12.1 left. **The likely crash chain:** v0.12.1 exits only on 137 and ignores any other download failure, so a cut-off file it saved stays; v0.13.0 parses it with `ijson` and fails at every start. **Options for the move (not taken tonight):** (a) `SKIP_WORLD_DATA=1` in the template — no internet dependence at update, but a FRESH install gets no world data; (b) keep the download and have the harness prove the update with the file present (the pre-seed is the volume itself) plus a check that a truncated file is replaced (`--force` is the command's own repair). The healthcheck override fix (`/usr/bin/node`) still rides the image move. | **READY — P2; owner: CC (catalog)** |
|
||||
| **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** `missingFileLegsRefusal` predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. | **READY — P3; owner: CC (controller)** |
|
||||
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). | **OPEN — P3, gaps (1)–(4) only; owner: CC** |
|
||||
| **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` **-- HALF DONE 2026-09-25 (hub v0.125.0):** the dialog no longer promises a Cloudflare removal; the preview lists what the operator removes by hand, by domain (the tunnel, the DNS records), never the token — proven live on the hub (`audits/night-2026-09-26/F/`). The Cloudflare leg itself is NOT built. | **NARROWED — the Cloudflare leg only; owner: operator (decide if it is wanted) / CC (build)** |
|
||||
| **R-689** | **[P3-LOW] demo-hp's scheduled restore test picks the golden template in `local:backup/` as a backup and fails every 6 h.** Read 2026-09-25 (agent 0.134.0): `restore-test tier is DUE … archive=local:backup/felhom-golden-0.236.0.tar.zst … newest settled archive (landed 2026-09-13T20:17:01Z)` then `scheduled restore-test FAILED … extractconfig … HTTP 403: permission denied at /vms/ (missing privilege VM.Backup)` — 2026-09-24 10:36, 2026-09-25 04:57 and 10:57. The golden is not a backup of any guest; the real archive (`vzdump-lxc-9201-2026_09_24-21_59_25`) was younger than the 24 h settle. So the restore test proves nothing on demo-hp and logs an ERROR each cycle. **Fix direction:** the restore test selects only `vzdump-*` archives of a known guest (or the golden moves out of `local:backup/`). Agent-side; not built (the agent is untouched this session). `audits/night-2026-09-26/part0/P0-demo-hp-restoretest-golden.txt` | **READY — P3; owner: CC (agent)** |
|
||||
| **R-691** | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). | **OPEN — P3; owner: CC (needs a small decision for (1))** |
|
||||
| **R-691** | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). **-- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay `:ro`, nothing on disk changes (CC-unattended decision, `07` §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, `audits/version-travel-2026-09-26/D3/`. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy.** | **OPEN — P3; owner: CC — NARROWED to (2)** |
|
||||
| **R-693** | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** |
|
||||
| **R-694** | **[P3-LOW] Loading kept data (and every unit restore) regenerates a withheld login secret — does the household's shown password still work?** Seen 2026-09-25 on 9202 (E5, nextcloud): `generated replacement for [NEXTCLOUD_ADMIN_PASSWORD] — the credential was reset (old value unrecoverable)`. The unit deliberately carries no internet-reachable admin login (D5). For nextcloud the real admin password lives in the loaded DATABASE, so the new env value is likely inert — but if the app page shows the regenerated value as "your password", it is a false one. **Not measured:** what the page shows after a load, per app. `audits/night-2026-09-26/E/E5-5-use.txt` | **OPEN — P3; owner: CC (measure first)** |
|
||||
| **R-695** | **[P3-LOW] Two kept-data Deletes in the same second can leave a self-perpetuating empty kept folder.** Seen 2026-09-25 on 9202 (v0.274.0, teardown): two Deletes at 13:35:16 each started `SyncFileBrowserMounts` in a goroutine; one restarted the file browser with the OLD bind list, Docker recreated the deleted folder `kept/nextcloud/2026-09-25_141014` EMPTY (root, 13:35:17), and the next sync LISTED that empty folder as a kept item and kept binding it — so the bind recreated the folder and the folder kept the bind. Broken by removing the folder and one controller restart. Harm: an empty 0-byte row and a stale empty source in the view; no data. **Fix direction:** single-flight the file-browser sync (the second waits and re-reads), and never list or bind an EMPTY dated kept folder. `audits/night-2026-09-26/G/T1-teardown-9202.txt` | **OPEN — P3; owner: CC** |
|
||||
| **R-696** | **[P2-MEDIUM] Tier 1's "proven at" is the unit MANIFEST's refresh time, not its data's — so the kept pre-conversion copy was released on a backup taken BEFORE the conversion.** Seen 2026-09-26 on demo-hp (v0.274.0, the automatic leg): docmost converted 16 → 18 at 02:18:52Z; at 02:20:16Z the recovery unit was re-captured (its pins changed) — `manifest.json` `created_at` 02:20:16Z — while its `db-dumps/docmost-postgres.sql` is from **02:15:01Z, on PostgreSQL 16**. `UpdateRestorePoints` reads Tier 1's time from `ListRestorePoints` (the unit's time), so at 02:30:16Z `ReleaseConversionCopies` logged "a backup proven on 18: Tier 1 at 02:20:16Z" and removed the 16 datadir copy. **Harm tonight: low** — the 02:15 logical dump holds the rows the conversion's check proved equal, and the next night's dump is on 18. **The class is wider:** the update's own precondition reads the same time, and a unit is re-captured on every controller release or pin change (`CaptureRecoveryUnit`), so a stale dump can read as minutes old (9202 2026-09-25 11:06: "Tier 1 copy 2m0s old" from the 11:04 restart, dumps older). On 9202 the release was right by luck (dump 11:57 > conversion 11:13). **Fix direction:** Tier 1's time = the newest DATA in the unit (its db/volume dumps' own times), never the manifest's; the release additionally requires a dump of the converted service written after the conversion. Presence-is-not-success, `CLAUDE.md`. `audits/night-2026-09-26/G/G4-demo-hp-docmost-after.txt` **-- A1 SPIKE 2026-09-26 (9202, controller 0.274.0, `audits/version-travel-2026-09-26/A1/README.md`) — the class is WIDER than the time: the refresh re-captures the unit's DEFINITION too, so the unit pairs new pins with old data.** Measured both shapes: an app step (docmost 0.95.0 → 0.96.0) — the refresh 2 min after the update wrote 0.96.0 into `compose/` over 0.95.0's dump; a restore in that window came back whole only because docmost migrated the old data forward at its first start (four migrations, no guard, no undo). An engine step (PostgreSQL 16 → 18) — the refresh wrote the 18 definition over the 16 dump and 16 datadir tar; a restore in that window poured the 16 datadir back, `postgres:18` REFUSED it, the replay timed out and **the app was left DOWN**. The second-drive mirror, written before the update, brought it back whole at 16. From source: the OFF-SITE restore never writes the definition at all (it uses the live compose), so after any update it mixes versions until the next night's off-site run. Being fixed as Part A of the version-travel brief (controller v0.275.0). | **OPEN — P2; owner: CC** |
|
||||
| **R-697** | **[P3-LOW] A restore drops the `conversion_copy` record but not the kept pre-conversion volume — the copy is then never released.** Seen 2026-09-26 on 9202 (A1, controller 0.274.0): after docmost's 16 → 18 conversion `app.yaml` held `conversion_copy` (`docmost_docmost_postgres_data.pre-update-20260926T073003Z`); a restore from the own unit, then one from the second drive, both left `conversion_copy=None` while the volume stayed. `PersistUnitRedeployConfig` builds a fresh `AppConfig` (Deployed, DeployedAt, Env, LockedFields), so every record kept beside the env is dropped by a restore — the hourly release reads the record, so the volume is orphaned. **Harm: disk only** (the size of the old datadir), never data. **Fix direction:** carry `conversion_copy` across the restore's app.yaml rewrite (the release then removes it when a backup on the new major is proven), and say in the log when a restore supersedes one. `audits/version-travel-2026-09-26/A1/S4-restore-window-step2.txt`, `S5-restore-tier2-window.txt` | **OPEN — P3; owner: CC** |
|
||||
| **R-698** | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. | **OPEN — P3; owner: operator (a decision), CC measures** |
|
||||
|
||||
|
||||
Reference in New Issue
Block a user