09 decisions 40-41 (operator rulings 2026-09-26); A1 spike: the unit's definition and data drift apart (R-696 widened), R-697 filed
gates / gates (push) Successful in 25s
gates / gates (push) Successful in 25s
This commit is contained in:
@@ -814,7 +814,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-693** | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** |
|
||||
| **R-694** | **[P3-LOW] Loading kept data (and every unit restore) regenerates a withheld login secret — does the household's shown password still work?** Seen 2026-09-25 on 9202 (E5, nextcloud): `generated replacement for [NEXTCLOUD_ADMIN_PASSWORD] — the credential was reset (old value unrecoverable)`. The unit deliberately carries no internet-reachable admin login (D5). For nextcloud the real admin password lives in the loaded DATABASE, so the new env value is likely inert — but if the app page shows the regenerated value as "your password", it is a false one. **Not measured:** what the page shows after a load, per app. `audits/night-2026-09-26/E/E5-5-use.txt` | **OPEN — P3; owner: CC (measure first)** |
|
||||
| **R-695** | **[P3-LOW] Two kept-data Deletes in the same second can leave a self-perpetuating empty kept folder.** Seen 2026-09-25 on 9202 (v0.274.0, teardown): two Deletes at 13:35:16 each started `SyncFileBrowserMounts` in a goroutine; one restarted the file browser with the OLD bind list, Docker recreated the deleted folder `kept/nextcloud/2026-09-25_141014` EMPTY (root, 13:35:17), and the next sync LISTED that empty folder as a kept item and kept binding it — so the bind recreated the folder and the folder kept the bind. Broken by removing the folder and one controller restart. Harm: an empty 0-byte row and a stale empty source in the view; no data. **Fix direction:** single-flight the file-browser sync (the second waits and re-reads), and never list or bind an EMPTY dated kept folder. `audits/night-2026-09-26/G/T1-teardown-9202.txt` | **OPEN — P3; owner: CC** |
|
||||
| **R-696** | **[P2-MEDIUM] Tier 1's "proven at" is the unit MANIFEST's refresh time, not its data's — so the kept pre-conversion copy was released on a backup taken BEFORE the conversion.** Seen 2026-09-26 on demo-hp (v0.274.0, the automatic leg): docmost converted 16 → 18 at 02:18:52Z; at 02:20:16Z the recovery unit was re-captured (its pins changed) — `manifest.json` `created_at` 02:20:16Z — while its `db-dumps/docmost-postgres.sql` is from **02:15:01Z, on PostgreSQL 16**. `UpdateRestorePoints` reads Tier 1's time from `ListRestorePoints` (the unit's time), so at 02:30:16Z `ReleaseConversionCopies` logged "a backup proven on 18: Tier 1 at 02:20:16Z" and removed the 16 datadir copy. **Harm tonight: low** — the 02:15 logical dump holds the rows the conversion's check proved equal, and the next night's dump is on 18. **The class is wider:** the update's own precondition reads the same time, and a unit is re-captured on every controller release or pin change (`CaptureRecoveryUnit`), so a stale dump can read as minutes old (9202 2026-09-25 11:06: "Tier 1 copy 2m0s old" from the 11:04 restart, dumps older). On 9202 the release was right by luck (dump 11:57 > conversion 11:13). **Fix direction:** Tier 1's time = the newest DATA in the unit (its db/volume dumps' own times), never the manifest's; the release additionally requires a dump of the converted service written after the conversion. Presence-is-not-success, `CLAUDE.md`. `audits/night-2026-09-26/G/G4-demo-hp-docmost-after.txt` | **OPEN — P2; owner: CC** |
|
||||
| **R-696** | **[P2-MEDIUM] Tier 1's "proven at" is the unit MANIFEST's refresh time, not its data's — so the kept pre-conversion copy was released on a backup taken BEFORE the conversion.** Seen 2026-09-26 on demo-hp (v0.274.0, the automatic leg): docmost converted 16 → 18 at 02:18:52Z; at 02:20:16Z the recovery unit was re-captured (its pins changed) — `manifest.json` `created_at` 02:20:16Z — while its `db-dumps/docmost-postgres.sql` is from **02:15:01Z, on PostgreSQL 16**. `UpdateRestorePoints` reads Tier 1's time from `ListRestorePoints` (the unit's time), so at 02:30:16Z `ReleaseConversionCopies` logged "a backup proven on 18: Tier 1 at 02:20:16Z" and removed the 16 datadir copy. **Harm tonight: low** — the 02:15 logical dump holds the rows the conversion's check proved equal, and the next night's dump is on 18. **The class is wider:** the update's own precondition reads the same time, and a unit is re-captured on every controller release or pin change (`CaptureRecoveryUnit`), so a stale dump can read as minutes old (9202 2026-09-25 11:06: "Tier 1 copy 2m0s old" from the 11:04 restart, dumps older). On 9202 the release was right by luck (dump 11:57 > conversion 11:13). **Fix direction:** Tier 1's time = the newest DATA in the unit (its db/volume dumps' own times), never the manifest's; the release additionally requires a dump of the converted service written after the conversion. Presence-is-not-success, `CLAUDE.md`. `audits/night-2026-09-26/G/G4-demo-hp-docmost-after.txt` **-- A1 SPIKE 2026-09-26 (9202, controller 0.274.0, `audits/version-travel-2026-09-26/A1/README.md`) — the class is WIDER than the time: the refresh re-captures the unit's DEFINITION too, so the unit pairs new pins with old data.** Measured both shapes: an app step (docmost 0.95.0 → 0.96.0) — the refresh 2 min after the update wrote 0.96.0 into `compose/` over 0.95.0's dump; a restore in that window came back whole only because docmost migrated the old data forward at its first start (four migrations, no guard, no undo). An engine step (PostgreSQL 16 → 18) — the refresh wrote the 18 definition over the 16 dump and 16 datadir tar; a restore in that window poured the 16 datadir back, `postgres:18` REFUSED it, the replay timed out and **the app was left DOWN**. The second-drive mirror, written before the update, brought it back whole at 16. From source: the OFF-SITE restore never writes the definition at all (it uses the live compose), so after any update it mixes versions until the next night's off-site run. Being fixed as Part A of the version-travel brief (controller v0.275.0). | **OPEN — P2; owner: CC** |
|
||||
| **R-697** | **[P3-LOW] A restore drops the `conversion_copy` record but not the kept pre-conversion volume — the copy is then never released.** Seen 2026-09-26 on 9202 (A1, controller 0.274.0): after docmost's 16 → 18 conversion `app.yaml` held `conversion_copy` (`docmost_docmost_postgres_data.pre-update-20260926T073003Z`); a restore from the own unit, then one from the second drive, both left `conversion_copy=None` while the volume stayed. `PersistUnitRedeployConfig` builds a fresh `AppConfig` (Deployed, DeployedAt, Env, LockedFields), so every record kept beside the env is dropped by a restore — the hourly release reads the record, so the volume is orphaned. **Harm: disk only** (the size of the old datadir), never data. **Fix direction:** carry `conversion_copy` across the restore's app.yaml rewrite (the release then removes it when a backup on the new major is proven), and say in the log when a restore supersedes one. `audits/version-travel-2026-09-26/A1/S4-restore-window-step2.txt`, `S5-restore-tier2-window.txt` | **OPEN — P3; owner: CC** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
Clearing a row means the check was DONE and its result recorded in that R-row —
|
||||
|
||||
Reference in New Issue
Block a user