night watch: demo-hp converted docmost by itself; R-696 filed; teardown recorded
gates / gates (push) Successful in 26s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-26 04:42:47 +02:00
parent 82eef7b7cf
commit b280bb5d59
11 changed files with 1292 additions and 10 deletions
+1
View File
@@ -814,6 +814,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-693** | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** |
| **R-694** | **[P3-LOW] Loading kept data (and every unit restore) regenerates a withheld login secret — does the household's shown password still work?** Seen 2026-09-25 on 9202 (E5, nextcloud): `generated replacement for [NEXTCLOUD_ADMIN_PASSWORD] — the credential was reset (old value unrecoverable)`. The unit deliberately carries no internet-reachable admin login (D5). For nextcloud the real admin password lives in the loaded DATABASE, so the new env value is likely inert — but if the app page shows the regenerated value as "your password", it is a false one. **Not measured:** what the page shows after a load, per app. `audits/night-2026-09-26/E/E5-5-use.txt` | **OPEN — P3; owner: CC (measure first)** |
| **R-695** | **[P3-LOW] Two kept-data Deletes in the same second can leave a self-perpetuating empty kept folder.** Seen 2026-09-25 on 9202 (v0.274.0, teardown): two Deletes at 13:35:16 each started `SyncFileBrowserMounts` in a goroutine; one restarted the file browser with the OLD bind list, Docker recreated the deleted folder `kept/nextcloud/2026-09-25_141014` EMPTY (root, 13:35:17), and the next sync LISTED that empty folder as a kept item and kept binding it — so the bind recreated the folder and the folder kept the bind. Broken by removing the folder and one controller restart. Harm: an empty 0-byte row and a stale empty source in the view; no data. **Fix direction:** single-flight the file-browser sync (the second waits and re-reads), and never list or bind an EMPTY dated kept folder. `audits/night-2026-09-26/G/T1-teardown-9202.txt` | **OPEN — P3; owner: CC** |
| **R-696** | **[P2-MEDIUM] Tier 1's "proven at" is the unit MANIFEST's refresh time, not its data's — so the kept pre-conversion copy was released on a backup taken BEFORE the conversion.** Seen 2026-09-26 on demo-hp (v0.274.0, the automatic leg): docmost converted 16 → 18 at 02:18:52Z; at 02:20:16Z the recovery unit was re-captured (its pins changed) — `manifest.json` `created_at` 02:20:16Z — while its `db-dumps/docmost-postgres.sql` is from **02:15:01Z, on PostgreSQL 16**. `UpdateRestorePoints` reads Tier 1's time from `ListRestorePoints` (the unit's time), so at 02:30:16Z `ReleaseConversionCopies` logged "a backup proven on 18: Tier 1 at 02:20:16Z" and removed the 16 datadir copy. **Harm tonight: low** — the 02:15 logical dump holds the rows the conversion's check proved equal, and the next night's dump is on 18. **The class is wider:** the update's own precondition reads the same time, and a unit is re-captured on every controller release or pin change (`CaptureRecoveryUnit`), so a stale dump can read as minutes old (9202 2026-09-25 11:06: "Tier 1 copy 2m0s old" from the 11:04 restart, dumps older). On 9202 the release was right by luck (dump 11:57 > conversion 11:13). **Fix direction:** Tier 1's time = the newest DATA in the unit (its db/volume dumps' own times), never the manifest's; the release additionally requires a dump of the converted service written after the conversion. Presence-is-not-success, `CLAUDE.md`. `audits/night-2026-09-26/G/G4-demo-hp-docmost-after.txt` | **OPEN — P2; owner: CC** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
Clearing a row means the check was DONE and its result recorded in that R-row —