diff --git a/CONTEXT.md b/CONTEXT.md index 9f2ea46e..49237fec 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,29 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## 2026-09-23 — the operator rules on automatic updates (`09` §3 decisions 11–18); two spikes say what the build needs + +**Operator rulings, not CC decisions.** One window = a leg of the backup chain (11); automatic, per-box +switch on by default (12); **the test decides, not the tag** (13, replaces decision 3's "never across +a major"); the ladder, one tested step at a time (14); **the box undoes a failed update itself** (15, +replaces §6.1's no-auto-undo); PostgreSQL majors converted by the box (16); digests recorded by the +catalog and pulled exactly (17); fleet view later (18). §3b is kept, marked ANSWERED. + +**The two mechanisms were spiked the same day, by hand, on 9202** (`audits/update-rulings-2026-09-23/`): +- **The undo works** — docmost and romm, whose OLD versions refuse migrated data, came back with data + written before AND after the backup, ≈16 s and ≈38 s. **But not with the loader the product has:** + `ImportDump` over a migrated PostgreSQL database FAILS on the new tables' foreign keys, and over + MariaDB leaves the new tables behind (R-638). A truncated PostgreSQL copy loads rc 0 into an EMPTY + database and `ValidateDump` would accept it (R-640). No-DB apps have no last-second copy (R-641). + Build list: R-637. +- **The ladder is absent**: one press jumped 2.3.0 → 2.5.0; the box's catalog clone is `--depth 1` + and cannot see steps. Recommended format: `update_ladder:` in `.felhom.yml`, each intermediate step + carrying its own definition — not the git history. +- **The harness watches memory** (`upgrade-test.py` v2): the RomM template as promoted fails at +76 s. + +**Build order and costs:** `09` §6.4 (≈22 evenings). **One open point for the operator:** the chain as +ruled leaves the update leg ≤15 min a night (R-643). + ## 2026-09-20 — localisation slice 5: the app catalog's copy model is MEASURED, not proposed (R-560) `architecture/10-localisation.md` §7 was a proposal; it is now a fact, with the numbers it was diff --git a/REPORT.md b/REPORT.md index 79b228d7..78f96570 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,161 +1,119 @@ -# REPORT — the update arc's two missing measurements, one lock, and the floor to 0.260.0 +# REPORT — update arc: the operator's rulings recorded, the undo and ladder spiked, the memory watch, the build plan -2026-09-21 (evening). Repos touched: **felhom.eu** (floor, docs, register, evidence), -**felhom-controller v0.261.0**, **app-catalog-felhom.eu** (two drill pairs, both reverted). -Architecture read first and named: `documentation/architecture/09-update-architecture.md` §3, §3b, -§4, §6.1, §6.2, §6.4, §8. +2026-09-23. Repos touched: **felhom.eu** (docs, register, STATUS, CONTEXT, evidence), +**app-catalog-felhom.eu** (`scripts/` only — the memory watch), **admin/app-catalog-drill** (drill +commits, reset to live `main` at the end). **felhom-controller, felhom-agent, hub: read only.** +Architecture read first and named: `documentation/architecture/09-update-architecture.md` (all of it), +`07-backup-architecture.md` §6. +Baselines verified live before starting: controller `b9deec19077b`, agent `d9864a94bf62`, felhom.eu +`267dcad01bcf`, catalog `02844ae0a579` — all equal to the brief. --- -## 1. NOT DONE / CHANGED FROM THE BRIEF — first, because that is the point of this session +## 1. Not done, or changed from the brief — first | item | state | |---|---| -| **Part 0** floor to 0.260.0 | **done** | -| **Part 1** the three cuts | **done, but NOT as specified.** All three landed in `verifying`, never in `starting` — see §2. The brief allowed this explicitly and asked that it be said. | -| **Part 2** the lock + `reason` on the wire | **done**, five red-proofs, proven live | -| **Part 3** the unattended night | **done for the success night and the no-retry proof. The unattended HOLD was NOT produced** — see §4. | -| **Part 4** docs and rows | **done** | -| the caller script's ≤150-line budget | **167 lines.** Over by 17, not trimmed: the excess is the within-a-major rule and its comment, the one part of that file that must be readable. | -| Scenario C run by the measuring agent | **run by the coordinator instead.** The agent was stood down mid-session after two long intervals with no evidence written; the coordinator ran C and captured A and B independently. Stated because it changes who measured what. | -| a second drill bump/revert pair | **used.** The brief permits it "if a fifth move is truly needed" and asks that it be named. It was: DRILL 2 (`ae08a037fd68`) added one real edge and one deliberately failing edge for Part 3. | +| Part 0 — rulings into `09` | **done**, commit `805ad1e` (documents only) | +| Part 1 — the undo, four cases + the wrong case | **done, with three changes.** (a) **The fourth case, "files on disk", was measured on vikunja's attachment (a file in a volume), not on a bind-mounted drive folder:** romm's two drive folders stayed EMPTY throughout — they measured nothing and are reported as unmeasured. (b) **Starting the app by hand needed its decrypted secrets; the session's safety guard refused that, and it was not worked around.** The product's own Start was used instead, after lifting the hold with the operator CLI + a controller restart. (c) The wrong case was run on BOTH engines, and on PostgreSQL with both the product's loader and the fixed one — the fixed one produced the session's most important finding (a truncated copy loads rc 0). | +| Part 2 — the ladder | **done** | +| Part 3 — the memory watch + red-proof | **done**; harness v2. `C3`, the harness's standing negative control, was **NOT run**: its template's `container_name: privatebin` collides with the privatebin the controller runs on 9202. The memory watch has its own pair instead — M1old must fail, M1 must pass. | +| Part 4 — the build plan | **done**, `09` §6.4 — with **one open point for the operator** the brief did not expect (§6) | -**Two instrumentation failures of my own, recorded because they cost evidence:** -1. The first unattended run's stdout was piped through `tail`, which buffers, and the run was later - killed — **the caller's own log for the Scenario F press was lost.** The outcome survived on the - box; the log did not. The second run wrote straight to a file. -2. A background security review flagged the deliberately-broken `vikunja → alpine:3.20` catalog edge - as a supply-chain change. **It was right to.** Accepted deliberately — no customer or demo box runs - vikunja, a deployed app is frozen at its own pin since v0.235.0, and it is the documented C3-class - control — but the window is now closed by the revert, and it is named here rather than left in a - tool notification. +**Claims in the brief that turned out wrong:** +1. *"A safety dump exists for every app class"* — **false.** An app with no database server gets none + (`update safety dump for vikunja: the app has no database — nothing to copy (no-op)`), measured. +2. *"The box keeps a git clone of the catalog"* with history — **false.** Depth 1 on both demo + guests, `rev-list --count HEAD` = 1; `sync.go:283`/`:300` clone and fetch `--depth 1`. +3. *"`stacks.update_window` is unread"* — **true, and the grep was widened** from `config.go` + + `setup/handlers.go` to the whole controller repo: the only other hits are + `configs/controller.yaml.example` and the i18n base file. No Go code reads it. +4. The register held **329** row lines by `grep -c '^| \*\*R-'`, not 326; the highest id was R-636 as stated. ---- +## 2. Part 0 — the rulings -## 2. Claims in the brief that turned out wrong +`09` §3 gains decisions **11–18** in the existing shape; decision 3's second half and §6.1's abort +paragraph are marked REPLACED with pointers; §4 says why the undo is not a rollback; §3b is kept, +headed ANSWERED, each question pointing at its decision; §6.2 rewritten to the ruled shape; the slices +table updated. Register: R-450, R-451, R-446, R-463 cite the decisions. -**§2.4's open question — does `backupMgr.IsRunning()` cover the update's `backing-up` phase?** -**YES.** `RunAppBackupNow` calls `acquireRunning` (`internal/backup/update_guard.go:333`), so that one -phase was already protected. The gap was `checking`, `safety-dump`, `pinning`, `pulling`, `starting` -and `verifying`. **The live lock probe landed in `safety-dump`**, i.e. squarely in the previously -unprotected window rather than in the one that was already covered. Everything else §2.4 asserted -held at source. +## 3. Part 1 — the undo, by hand -**§2.3 — the resume path was READ, not measured. It is now measured**, three times, and it does what -it said. +Full evidence and tables: `documentation/audits/update-rulings-2026-09-23/README.md`. -**The stopped-guest byte path in my own brief to the measuring agent was wrong.** -`/var/lib/lxc/9202/rootfs/var/lib/felhom/...` is an empty mountpoint while 9202 is stopped, because -`mp0` is a separate raw volume; `pct mount 9202` does attach it. I had corrected one trap and -introduced a second. The positive control caught it. - -**"Four qualifying apps exist" — held.** vikunja, uptime-kuma, wishlist, glance: single-container, no -database sidecar, none on either demo box, all four target tags verified to exist upstream first. - ---- - -## 3. Part 0 — the floor - -Raised to **0.260.0** with MinAgent **0.131.0** declared (above the vouched golden 0.258.0, so the -declaration carries it — §3 decision 7). Hub log: - -``` -[INFO] Global controller-version floor set to "0.260.0" (declared MinAgent "0.131.0") -[INFO] managed floor SERVED for demo-felhom: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0) -[INFO] managed floor SERVED for demo-hp: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0) -``` - -Blast radius, read from the hub before saving: **3 boxes below** — `drill-r50` (0.213.0, BLOCKED), -`peti-felhom` (0.115.0, DOWN), `tester-1` (0.245.0, DOWN). None was reachable, so none moved; they -take it when they return. Both demo boxes were already on 0.260.0 by hand and now hold it by floor. - ---- - -## 4. Part 1 — the three cuts (R-610, CLOSED) - -Guest 9202, controller v0.260.0, four throwaway apps seeded through their own front doors first. - -| | A — vikunja | B — uptime-kuma | C — wishlist | +| | docmost (PostgreSQL) | romm (MariaDB) | vikunja (volume, no DB server) | |---|---|---|---| -| edge | 2.3.0 → 2.6.0 | 2.4.0 → 2.5.0 | v0.66.0 → v0.67.0 | -| cut | `pct stop` | `pct stop` | **controller container only** | -| phase at decision | `starting` | `verifying` | `starting` | -| cut latency | 3 759 ms | 3 016 ms | **1 675 ms** | -| phase it died in | `verifying` | `verifying` | `verifying` | -| recovery | resumed, healthy 0 s | resumed, healthy 5 s | resumed, healthy 10 s | -| total | DONE 1 m 26 s | DONE 1 m 0 s | DONE 51 s | -| four observables | **agree** | **agree** | **agree** | -| seeded data | **read back intact** | **read back intact** | not re-read (gap) | +| held after | 95.3 s | 102.6 s | 93.3 s | +| old version on migrated data | **refuses** | **refuses** | starts | +| safety dump | 135 816 B, DB only | 62 943 B, DB only | **none** | +| product loader (`ImportDump` semantics) | **FAILS**, rc 3, foreign keys | rc 0, **12 tables left behind** | — | +| fixed load (empty schema + copy, one transaction) | rc 0, 1.38 s | (product loader sufficed) | — | +| undo → healthy on the OLD probe | ≈ 16 s | ≈ 38 s | ≈ 1 s | +| data before / after the backup | yes / **yes** | yes / **yes** | yes / **yes** (attachment too) | -**The dangerous case was genuinely exercised, and a log line proves it rather than an assumption.** -vikunja's own log: `Ran all migrations successfully` / `Vikunja version v2.6.0` at **12:28:26.881 -UTC — 0.64 s after the cut decision and ~0.4 s before the guest stopped answering.** The 2.6.0 schema -migration had already been applied to the customer's database when the power went. Recovery resumed -**forward**, so old-binary-on-migrated-database never happened — **but this branch is one step from -it**, and that is now evidence for §4's "no automatic rollback" ruling rather than argument for it. +**Does a product path load a safety dump back?** Yes — `rollbackSafetyDump`, but only the off-site +restore calls it; the update never reads its own dump, and `failAndHold` deletes the pre-update +definition copies. -**Instrument limit:** `starting` lasts well under a second on this box. Three attempts, two cut -mechanisms, all landed in `verifying`. No phase was faked. `RecoverUpdates` handles `starting` and -`verifying` in **one branch**, so all three exercise the arm under test. A cut inside `starting` -itself needs an in-process fault injector. +**The wrong case:** PostgreSQL + product loader → rc 3, nothing changed (honest hold). **PostgreSQL + +the fixed atomic loader + a half-length copy → rc 0 and an EMPTY database (0 users, 0 constraints, +0 indexes) that still shows 42 tables and 48 ledger rows** — no hold, a dishonest success. MariaDB + +half copy → rc 1, half the tables already replaced (not atomic). The whole copies carry an end marker +the truncated ones lack; `ValidateDump` does not check it. -**Scenario C also showed the other apps were undisturbed** by the controller restart — uptime-kuma, -vikunja, glance and filebrowser all kept their uptime. Only the app the update was itself recreating -restarted. +## 4. Part 2 — the ladder ---- +One press on vikunja two steps behind: **2.3.0 → 2.5.0 in 9.5 s; 2.4.0 never ran.** The box cannot see +2.4.0 (depth-1 clone). Recommended format: `update_ladder:` in `.felhom.yml`, intermediate steps with +their own definition — **not** git history, because romm's image-moving commit is the definition that +OOM-looped on demo-hp. Full comparison in the audit. -## 5. Part 2 — v0.261.0, proven live with a control +## 5. Part 3 — the memory watch -See `felhom-controller/REPORT.md` for the code. The live proof is the part worth repeating: +`upgrade-test.py` v2: after a successful readback, `--soak` seconds (default 600) of light load +(4 callers), sampling every 15 s the kernel's `oom_kill` counter read host-side from the container's +cgroup, the peak, the limit, the restarts and Docker's OOMKilled flag. Kill or restart → `failed`; +peak > 80 % → mark `memory_tight`. New `Romm` fixture; edges `M1` (current template) and `M1old` +(the template as promoted, `15f9ebf`). -| probe | result | +- **M1old (red-proof): `failed` — first OOM kill at +76 s**, peak 512 MiB = 100 % of the limit, + restarts 0 (the container kept running — the shape that hid it on demo-hp), abort `refuses`. + **Ten minutes is ample for this failure under load.** +- **M1 (positive control):** **proven** + mark `memory_tight` — 608.5 s under 11 429 requests (5 712 × 200, 5 717 × 401): **0 kernel OOM kills, 0 restarts**, peak 621 MiB = **81 %** of 768 MiB — the watch passes the fix and still flags the thin headroom R-635 left open. + +## 6. Part 4 — the build plan + +`09` §6.4: ten parts, **≈ 22 evenings**, recommended order undo → sentences in the household's language ++ notifier honesty → test record + gate + memory check → ladder → digests → the automatic leg → R-636 → +R-625 → PostgreSQL conversion; fleet view deferred by ruling. **One open point needs the operator +(R-643):** as ruled, the update leg sits between the off-site leg (W+105m) and the full-system gate +(W+2h) — at most 15 minutes a night. Recommendation: the full-system backup waits for the leg inside +its own four-hour window. + +## 7. Rows + +**Opened (8):** R-637 (build the undo), R-638 (the loader cannot replay over a newer schema — and the +restore the hold names is UNMEASURED after a real schema migration), R-639 (pre-update copies deleted +on hold), R-640 (a truncated PostgreSQL copy loads rc 0 into an empty database), R-641 (no-DB apps have +no last-second copy), R-642 (Start returns 200 over a crash loop), R-643 (the ≤15-minute leg, +operator), R-644 (gokapi crash-looping on 9202 at session start, not caused here). +**Updated:** R-446, R-450, R-451, R-462, R-463. **Closed:** none. **329 → 337.** + +## 8. Teardown — three layers + +| layer | state | |---|---| -| manual self-update, **no** app update running | „A frissítés nem érhető el (nincs gazda-ügynök)" — the **agent** refusal | -| manual self-update, app update **in flight** (`safety-dump`) | „Egy alkalmazás frissítése éppen folyamatban van…" — **our** refusal | +| **machine — guest 9202** | `controller.yaml` restored from the saved copy and read back identical (live catalog, no `update:` block); catalog cache re-cloned from the live repo (`02844ae`); docmost, romm, vikunja removed through the product (no containers, no volumes); romm's drive folder (kept by the product, R-442) removed by name; `/opt/upg` and every temp file in `/root` removed; the eight test images removed **by name** — `vikunja:2.4.0` was **absent**, i.e. never pulled, which is the jump seen from a second side; **no `prune`**. Containers afterwards: the same three apps as at the start (gokapi still crash-looping — R-644, pre-existing). `82-teardown-guest.txt` | +| **host — demo-hp** | nothing provisioned; only transient `/tmp` files, removed | +| **hub** | nothing touched — no hub call was made | +| **drill repo** | reset to live `main` `02844ae0a579`; `has_actions: false`; image lines identical to live. The local clone's push URL to the LIVE catalog was disabled at the start. `81-teardown-drill-repo.txt` | -**The sentence changed.** Guest 9202 has no host agent, so `TriggerUpdate` refuses either way — which -makes it the perfect negative control, because the new check sits *before* the agent check. Two -sentences, one probe, and no swap could reach a machine. Self-update was enabled for the probe and -**restored to `false`** from a copy taken first; verified by re-reading the file. +**One instrumentation slip, recorded:** a background watcher and the first teardown both wrote the same +temporary script file on demo-hp at the same moment, so the first teardown never ran (its output file +held the watcher's lines). Caught by reading the file, re-run after the watcher ended; the second run +is the one recorded. -**The reverse direction — an app update refused while the controller swaps — is NOT staged live.** It -is covered by a red-proofed consequence test. Staging it would need a real swap and a host agent this -guest does not have. Stated as a gap. +**Fences:** DooPlex, Peti's box, ep0, the demo guests' apps, `drill-r50`, `tester-1` and the hub were +not touched. The live catalog's `main` was `02844ae0a579` before and after. ---- - -## 6. Part 3 — the unattended night (R-611, CLOSED) - -**Success:** `uptime-kuma` 2.5.0 → 2.5.1 applied with nobody pressing anything. - -**No-retry:** after the revert left all four apps *ahead* of the catalog, the caller pressed each -**exactly once**, was refused `downgrade` (terminal), and pressed nothing across two further passes. -`never_again=['glance','uptime-kuma','vikunja','wishlist']`, `outcomes={}`. That is R-524 and R-609 -working together, unattended. - -**The unattended HOLD was never produced, and the reason matters:** the only failing edge available -(`vikunja → alpine:3.20`) was **correctly refused by the within-a-major rule before it was ever -attempted**. The rule that makes automatic updates safe is the same rule that refuses the obvious way -to break one. Measuring it needs an image that passes the version test and still fails health. -**§3b Q4 therefore still rests on the ATTENDED hold from slice 4.** - ---- - -## 7. Rows, and the catalog - -**Closed:** R-608, R-609, R-610, R-611. **Opened:** R-612 (P1 — wishlist unusable on a fresh install -and the error is a lie), R-613 (P2 — uptime-kuma reports healthy on its setup wizard), R-614 (P3 — -stale update phase survives a redeploy). **Corrected:** R-520's closing pointer now names R-610. -Register 302 → **307**. - -**Catalog:** DRILL `573e41f5`, DRILL 2 `ae08a037`, **REVERT `f5f6a152`**. Every `image:` line in -`templates/` is byte-identical to the pre-drill `ff9717d3` — `git diff` over those paths is **0 -lines**. `catalog_since` reads 2026-09-21 on the four rather than the older dates, because the gate -requires an image move to carry the day's date in either direction and a revert is a move. - -**Teardown, three layers.** *Machine:* guest 9202 left running on v0.261.0 with the four throwaway -apps still deployed and healthy (glance, uptime-kuma, vikunja, wishlist) — they are the fixture for -the remaining Q4 work and removing them would cost the next session the seeding. *Host:* demo-hp -untouched apart from 9202; **guest 9201 never touched.** *Hub:* nothing provisioned, nothing -enrolled; 9202 reports to no hub by design. `/tmp/.ctlpw` shredded. +**`unproven.py --summary`:** walked 20 / partial 17 / built 14 / missing 4 — **not walked 35 of 55, unchanged.** diff --git a/STATUS.md b/STATUS.md index 8733d85a..97da5e2c 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,32 +1,21 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-23 (morning) — your seven answers on automatic updates are written into the design notes as rulings.** No later session will ask them again. The two new pieces your answers need — the automatic undo and the step-by-step climb — are being tested by hand on the scratch machine today, before anyone builds them. The build plan comes back to you part by part. - ---- - -**Updated 2026-09-22 (late) — I fixed the six faults the two drill nights found in the update, delete and hold machinery, and shipped the six app versions you approved. One thing needs your word: whether the fleet moves to the new controller.** +**Updated 2026-09-23 — your seven answers on automatic updates are now written rules. I tested the two new pieces by hand on the scratch machine. The build plan is ready for you, part by part.** **Decisions I took on my own: none.** -**The one that mattered most is fixed and proven.** An app with no health check used to be **shut down by a successful update** — the machine waited five minutes for a check that could never arrive, then stopped a working app. Paperless-ngx, same app, same button: **before, it failed after 5 minutes and the app went dark. Now it finishes in 53 seconds and keeps running.** +**The automatic undo works — but not with the tool the machine has today.** I broke three real updates on purpose. For two of the apps, the old version refused to start on the data the new version had changed. After I loaded the copy taken seconds before the update, all three apps came back with all their data, including what was written after the nightly backup. It took 16 and 38 seconds. **But the machine's current way of loading a copy fails on one app and leaves junk behind on another.** A fix is known and tested: empty the database first, then load, as one step. -**Five more, all proven on the test machine.** -- **Deleting an app while it is being backed up or restored is now refused**, with a plain sentence telling you to wait — instead of quietly tearing it down and leaving a ghost behind. -- **A delete now checks its own work.** The machine watches for 25 seconds afterwards and removes anything that comes back, and says whether it verified. -- **An app the machine has lost track of can now be deleted.** Before, if its record went wrong, no button worked and only a command line could clear it. -- **A failed update now keeps the app's own log** before shutting it down. Twice we lost the only evidence of why. -- **Deleting an app clears its old update status**, so a fresh install of the same app no longer shows a stale "Updated". +**One dangerous finding.** A cut-off database copy loads as a "success" — into an empty database. The machine's copy checker would accept it. The fix is simple: check that the copy has its end marker before loading. This also protects today's restores. -**The six versions you approved are live on the catalogue** — Emby, Ghost, Immich, Radarr, Sonarr, Termix. **None of them is installed on either demo machine**, so nothing updated; they simply show as available. +**The step-by-step climb does not exist yet.** Today a machine two versions behind jumps straight to the newest one; the middle version never runs. The machine also cannot see old versions: it keeps only the newest copy of the catalogue. I propose a list of tested steps in each app's catalogue file. -**What I did not do, and it is on purpose.** Two items from the plan are untouched and named rather than half-finished: finding out *why* an app's record goes wrong in the first place (I fixed the consequence, not the cause), and making a held app stop offering an Update button it will refuse. +**The memory lesson from RomM is now in the test bench.** After an update, the bench runs the app for 10 minutes and watches memory. RomM's old setup failed in 76 seconds, so the bench would have caught it. -**What went wrong on my side.** I lost **44 minutes** to my own progress-watchers: they waited for a build that had already succeeded, because each was watching for a name its own command contained. The same bug cost me a pile of stuck watchers earlier in the day. It is now written down as a rule so it does not happen a third time. I also nearly recorded one test as passing when it had proved nothing — the refusal I saw came from an older rule, not the new one. I caught it and re-ran it properly. +**What needs you — two things.** +1. **The build plan: about 22 evenings in 10 parts.** I recommend starting with the undo (4 evenings). It also makes the manual Update button safer on its own. If you do nothing, nothing is built and updates stay manual. +2. **One question about the night schedule.** As ruled, updates run after the off-site copy and before the full-system backup. That gap is at most 15 minutes a night. Option A (my pick): the full-system backup waits for updates, still inside its own 4-hour window. Option B: keep the 15 minutes; a machine far behind takes weeks to catch up. If you do nothing, part 7 of the plan waits. -**Rows opened and closed.** Four closed, one narrowed to what is still unknown. The list stands at 325. +**Rows opened and closed.** Eight opened, none closed. The list went from 329 to 337. -**The fleet is on the new controller — you said yes and it is done.** I raised the floor to **0.262.1** with the required agent version declared alongside it. **Both demo machines picked it up in under twenty seconds** and are running healthy. The drill machine is switched off and is **held back on purpose**: its helper software is older than the new controller needs, so the machine refuses to give it a version it cannot run. That is the guard working, not a failure. Peti's machine is parked and would take it only if it ever comes back online. - -**What needs you: nothing.** - -**Nothing on your own machine or the off-site box was touched. The demo machines were not touched — they only see the six new version badges.** +**Nothing on your own machine, Peti's machine or the off-site box was touched. The demo machines were not touched. Only the scratch machine was used, and it is back on the real catalogue.** diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 2ab57153..fb7e4995 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -777,6 +777,27 @@ the harness has proven it, is slice 6's. **Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6. +### 6.1a The undo (decision 15) — SPIKED BY HAND 2026-09-23, not built + +Evidence: `audits/update-rulings-2026-09-23/README.md`. Three real migrating edges on 9202, each made +to fail a deliberately wrong probe, each held by today's product, each then undone by hand. + +| | docmost (PostgreSQL) | romm (MariaDB) | vikunja (SQLite in a volume) | +|---|---|---|---| +| old version on the migrated data, nothing loaded | **refuses** (migration ledger) | **refuses** (alembic revision) | starts and serves | +| the safety dump | DB only, holds the post-backup write | DB only, holds it | **none — no-op** | +| undo, load + start → healthy | **≈ 16 s** | **≈ 38 s** | ≈ 1 s | +| data written before AND after the backup read back | yes / yes | yes / yes | yes / yes | + +**The undo works — and not with the loader the product has.** `ImportDump` over a migrated +PostgreSQL database FAILS (the new version's foreign keys block the dump's own drops); over MariaDB it +succeeds and leaves the new version's tables behind. The load that worked empties the schema and loads +the copy in one transaction. **And a truncated PostgreSQL copy loads with exit 0 into an EMPTY +database** — the copy's completion marker must be checked first, which `ValidateDump` does not do. A +product path that loads a safety dump back exists (`rollbackSafetyDump`), but only the off-site +restore calls it. The eight things the build must add are listed in the audit; §6.4 part 1 prices +them. + **The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0 were hand-deployed to the demo guests. **RESOLVED by §3 decision 7 (hub v0.112.0):** v0.239.0 reached @@ -816,10 +837,19 @@ evidence; Slice 6 puts the same shape in the catalog. "seed_read_before": true, "seed_read_after": true, "healthy_after": true, "migration_observed": "verbatim log line, or null", "abort": "starts-and-serves | refuses | starts-data-gone | not-attempted", + "memory": {"soak_s": 600, "containers": {"": {"limit": 0, "peak": 0, "peak_pct": 0.0, + "oom_kills": 0, "restarts": 0}}, "first_kill": null}, + "marks": ["memory_tight"], "abort_detail": "the refusal quoted verbatim, or null", "duration_s": 0, "measured_at": "RFC3339", "evidence": "relative path"} ``` +**Harness version 2 (2026-09-23, R-635) adds `memory` and `marks`.** After a successful readback the +harness runs the new version for `--soak` seconds (default 600) under light load and reads the +kernel's own `oom_kill` counter host-side. A kill or a restart turns `proven` into `failed`; a peak +above 80 % of the compose limit adds `memory_tight`. The two fields are the test record's memory half +(decision 13). + **`inconclusive` is a first-class verdict and must never be collapsed into `failed`.** "We could not measure it" and "it does not work" are different facts, and only one of them is about the app. **`migration_observed` is a quoted line, never an inference from timing** — the value of both the @@ -900,7 +930,16 @@ spend) or the leg gets almost no time. That is a build choice inside decision 11 **How far — one step at a time (decision 14).** A box two steps behind applies step A→B, then B→C, each the full guarded update, each with its own health check and undo. A failed step stops the -ladder for that app. The ladder's format is §6.4 / `audits/update-rulings-2026-09-23/`. +ladder for that app. **Measured 2026-09-23:** today one press jumps A → C and B never runs, and the +box cannot see B at all — its catalog clone is `--depth 1` (`sync.go:283`, `:300`; one commit +visible on both demo guests). + +**The ladder's format — recommended, not ruled** (`audits/update-rulings-2026-09-23/README.md` Part 2): +an `update_ladder:` list in `.felhom.yml`, one entry per step — `from`/`to` refs per service, the +digest per ref, the test record, the marks — and, for every step but the last, the step's OWN +complete definition in `templates//steps/.yml`. **Not the git history:** romm's image-moving +commit `15f9ebf` is the definition that OOM-looped on demo-hp; the step that works is its images with +the later `f4eb94f` template, and no commit holds that pair. **When it fails — undo, then hold only if the undo fails (decision 15).** The household is told on the app page and by **one** mail, in the box's language (R-606 is a precondition — an automatic @@ -949,7 +988,47 @@ Rank stays P3-LOW at two enrolled boxes. It rises with the fleet, and §2 of the that looks like today: the only way to answer *"is the fleet current?"* was to read both boxes' files by hand. -### 6.4 The update night — a drill brief outline, costed from R-462's real numbers +### 6.4 The build order for the 2026-09-23 rulings (PLAN — each part returns to the operator for go/no-go) + +Costed in **CC-evenings** (one evening ≈ one unattended session: build, red-proofs, live proof on 9202, +release). Written from the two spikes and the memory watch of 2026-09-23 +(`audits/update-rulings-2026-09-23/`), not from source reading alone. **Risk to customer data** is +what the part can do to a household's data if it is wrong, not how likely that is. + +| # | part | rulings / rows | cost | depends on | risk to customer data | +|---|---|---|---|---|---| +| **1** | **The undo.** Keep the pre-update copies (compose, applied, pin, **old `.felhom.yml`**) until the undo is over; in `failAndHold`: pin back → DB up alone → **validate the copy's completion marker** → **empty-then-load in one transaction** (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → **health with the OLD probe** → `undone`, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. | 15; the audit's 8-point list | **4** | — | **HIGH by nature** — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. **It also makes the manual button safer on its own**, which is why it goes first. | +| **2** | **The update sentences in the household's language** (R-606) and a mail when an automatic update is undone or held. | R-606, 15 | **1** | — | none | +| **3** | **A disabled notifier says so** (R-620), so the mail of part 2 can be measured on a scratch box at all. | R-620 | **0.5** | — | none | +| **4** | **The test record + the catalog gate + the memory check.** The harness writes the ladder entry (below) from its verdict record, including the memory watch's peak and marks; the gate refuses an image move with no entry, an entry with a `failed` verdict, or one with no memory watch; `CompareImageRefs`' rule moves here as the push-time safety net. **Backfill:** one entry per current pin — the 21 proven moves from their records, every other pin `needs_person: "never tested"`, which is honest and keeps them manual. A version move re-checks `mem_limit` against the watch's peak (the RomM follow-up: gate, not checklist, because the watch now produces the number). | 13, R-635 follow-up | **2.5** | the memory watch (shipped 2026-09-23) | none on a box — catalog-side only | +| **5** | **The ladder on the box.** Read `update_ladder:` from the clone, find the installed step, apply ONE step with its OWN definition (`steps/.yml`, the last step the current template), repeat next night; a failed step stops the ladder for that app. `CatalogOrder` compares refs with the digest stripped (see part 7). | 14 | **2.5** | 1, 4 | medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape | +| **6** | **Digests.** The catalog records `sha256` per pin at push time (`check-image-resolvable.py` already resolves it); the box compares it for the badge and renders `name:tag@sha256:…` when present. **Measured 2026-09-23 on 9202:** Docker and Compose both pull and run `redis:7-alpine@sha256:858f…`, and refuse a digest that does not exist (`audits/update-rulings-2026-09-23/70-…`). **Build trap, read from source:** `splitImageRef` returns "unorderable" for ANY ref containing `@` (`updateorder.go:134`), so the digest must be split off before ordering or every digest-pinned app reads Unknown. A digest gone upstream fails the PULL — Scenario E, pin back, nothing ran. | 17, R-446 | **2** | 4 (the entry carries the digest) | low | +| **7** | **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has FINISHED (legs are clock-scheduled today, not chained — a completion signal is new), one app at a time (there is no single-flight, §3b Q4), `app_update.unattended` default ON, `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it. **See the one open point below.** | 11, 12 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe | +| **8** | **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none | +| **9** | **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none | +| **10** | **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside | +| **11** | **Fleet view** — per compose service: installed ref, catalog ref, badge state in the report; the hub lists boxes behind. | 18, R-451 | 2 | — | none — **deferred by the ruling** until the fleet grows | + +**Recommended order: 1 → 2 + 3 → 4 → 5 → 6 → 7 → 8 → 9 → 10**, part 11 when the fleet grows. **Total +for 1–10: ≈ 22 evenings.** The automatic caller (7) is deliberately late: it is the only part that +acts with nobody watching, and every part before it is what makes that safe. Parts 8 and 9 are small +and independent and can fill any short evening. + +**The one open point the build cannot settle alone — part 7, inside decision 11.** The ruled chain is +*off-site copy → updates → full-system backup*. Today the off-site leg starts at **W+105m** and the +full-system backup's gate opens at **W+2h** (`quiesce.go` `gateOpenOffsetMin = 120`, span to W+6h). So +the update leg has **at most 15 minutes**, and none on a night the off-site copy runs long — while one +step takes ~1 min when it works and ~2–6 min when it fails and is undone. + +| option | cost | +|---|---| +| **the full-system backup waits for the update leg, inside its own window; the leg stops starting new steps at W+5h** | the full-system backup starts later on update nights, still inside its four-hour window, with an hour kept; one more interlock between two nightly jobs | +| the leg stops at W+2h as the chain stands | ≤ 15 min a night — about ten steps on a good night, none on a slow one; a box far behind takes weeks to climb | + +**Recommendation: the first.** It keeps the ruling's order and its promise that the full-system backup +is never skipped for an update; only the start time inside its existing window moves. + +### 6.4.1 (record) The update night — the drill brief that preceded the rulings, costed and re-costed **The ruling is decision 6: all 53 apps, through the nightly rotation.** This is an ORDER inside that ruling, not a scope change. The database apps go first because they are the ones where a wrong answer diff --git a/documentation/audits/update-rulings-2026-09-23/01-repoint-to-drill.txt b/documentation/audits/update-rulings-2026-09-23/01-repoint-to-drill.txt new file mode 100644 index 00000000..f2c89786 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/01-repoint-to-drill.txt @@ -0,0 +1,10 @@ +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git + sync_interval: 15m + token: + username: "admin" +hub: +update: + health_timeout: 90s + diff --git a/documentation/audits/update-rulings-2026-09-23/02-three-controls.txt b/documentation/audits/update-rulings-2026-09-23/02-three-controls.txt new file mode 100644 index 00000000..31979947 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/02-three-controls.txt @@ -0,0 +1,26 @@ +sync: 429 +rescan: 200 +--- control 1: 9202 follows the drill catalog +cache=/var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache +02844ae decoys: eight cases for the probe target rule (R-630) +https://@gitea.dooplex.hu/admin/app-catalog-drill.git + image: docmost/docmost:0.96.0 + +docmost catalog_images= None +vikunja catalog_images= None +romm catalog_images= None +--- control 2: demo-hp guest 9201 still follows the live catalog +02844ae decoys: eight cases for the probe target rule (R-630) +https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + +--- control 3: live catalog main unchanged +02844ae0a579e8bf66882d01598c2212264a1b78 refs/heads/main + +=== retry after the 429 (a sync was already in flight at startup) +sync: 200 {'ok': True, 'data': {'ok': True, 'updated': ['docmost', 'romm', 'vikunja'], 'message': 'Sablonok frissítve — frissítve: +rescan: 200 +193001a DRILL: FROM states for the undo spike (docmost 0.95.0, vikunja 2.3.0, romm 5.0.0) + image: docmost/docmost:0.95.0 + image: vikunja/vikunja:2.3.0 + image: rommapp/romm:5.0.0 + diff --git a/documentation/audits/update-rulings-2026-09-23/50-part2-clone-depth.txt b/documentation/audits/update-rulings-2026-09-23/50-part2-clone-depth.txt new file mode 100644 index 00000000..41acc454 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/50-part2-clone-depth.txt @@ -0,0 +1,15 @@ +### guest 9202 (currently on the DRILL catalog) +shallow: true +/var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache/.git/shallow +has .git/shallow +commits reachable from HEAD: 1 +commits touching templates/docmost/docker-compose.yml: +419f243 DRILL docmost: docmost/docmost:0.95.0 -> docmost/docmost:0.96.0 + ... total 1 +template B readable at an older commit? + +### guest 9201 demo-hp (the LIVE catalog, never repointed) +shallow: true +commits: 1 +romm compose history: 1 +02844ae 2026-09-22 decoys: eight cases for the probe target rule (R-630) diff --git a/documentation/audits/update-rulings-2026-09-23/51-part2-ladder-today.txt b/documentation/audits/update-rulings-2026-09-23/51-part2-ladder-today.txt new file mode 100644 index 00000000..f348e676 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/51-part2-ladder-today.txt @@ -0,0 +1,24 @@ +08:25:27 === Part 2.1 — today's behaviour with a box TWO catalog steps behind +08:25:30 installed now: {'vikunja': 'vikunja/vikunja:2.3.0'} +08:25:30 [5] drill commit 610ed1a66dee: vikunja vikunja/vikunja:2.6.0 -> vikunja/vikunja:2.4.0 (push rc=0) +08:25:31 [5] drill commit c71807f5d47e: vikunja vikunja/vikunja:2.4.0 -> vikunja/vikunja:2.5.0 (push rc=0) +08:25:31 drill commits: step B 610ed1a66dee (2.4.0), step C c71807f5d47e (2.5.0) +08:25:36 badge caught up after 4.5 s; catalog_images = {'vikunja': 'vikunja/vikunja:2.5.0'} +08:25:36 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:25:36 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:25:37 + 0.6s phase=pulling label=Új verzió letöltése… err=None hold=None +08:25:40 + 3.7s phase=starting label=Indítás az új verzióval… err=None hold=None +08:25:41 + 4.7s phase=verifying label=Működés ellenőrzése… err=None hold=None +08:25:45 + 9.5s phase=done label=Frissítve err=None hold=None +08:25:45 {"final_phase": "done", "update_error": null, "hold_reason": null, "state": "running", "duration_s": 9.5} +08:25:48 AFTER ONE PRESS: {"pinned_images": {"vikunja": "vikunja/vikunja:2.5.0"}, "installed_images": {"vikunja": "vikunja/vikunja:2.5.0"}, "catalog_images": {"vikunja": "vikunja/vikunja:2.5.0"}, "live_compose_image_lines": ["image: vikunja/vikunja:2.5.0"], "docker_inspect": ["vikunja vikunja/vikunja:2.5.0 running=true restarts=0"]} +08:25:50 time=2026-09-23T08:25:40.726+02:00 level=INFO msg="Running migrations…" +time=2026-09-23T08:25:40.747+02:00 level=INFO msg="Ran all migrations successfully." +time=2026-09-23T08:25:40.757+02:00 level=INFO msg="Vikunja version v2.5.0" +2026/09/23 06:22:30 pin.go:362: [INFO] [stacks] update vikunja: pin advanced to the catalog's current definition (vikunja=vikunja/vikunja:2.6.0) +2026/09/23 06:25:36 update.go:623: [INFO] [stacks] update vikunja: precondition met — Tier 1 (own recovery unit) copy from 2026-09-23T06:24:24Z (1m0s old, limit 24h0m0s) +2026/09/23 06:25:36 pin.go:362: [INFO] [stacks] update vikunja: pin advanced to the catalog's current definition (vikunja=vikunja/vikunja:2.5.0) +2026/09/23 06:25:45 update.go:737: [INFO] [stacks] update vikunja: DONE in 9s + +08:25:51 vikunja: readback of the seeded project http=200 ok=True +08:25:51 seed A after the jump: True diff --git a/documentation/audits/update-rulings-2026-09-23/60-part1-teardown.txt b/documentation/audits/update-rulings-2026-09-23/60-part1-teardown.txt new file mode 100644 index 00000000..1e70ac9b --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/60-part1-teardown.txt @@ -0,0 +1,18 @@ +08:26:52 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'} +08:27:24 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_ +08:27:32 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost' +08:27:34 [X] stop -> 200 {'ok': True, 'message': 'Stack romm stop completed'} +08:27:39 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/romm tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss +08:27:39 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead +08:28:06 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'romm', 'volumes_removed': ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'], 'hdd_paths_removed': [], 'hdd_pat +08:28:13 [X] after remove: deployed=False leftovers='/opt/docker/stacks/romm' +08:28:14 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'} +08:28:45 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [ +08:28:53 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja' +08:29:15 docmost: containers=0 volumes=0 stackdir-appyaml=0 +romm: containers=0 volumes=0 stackdir-appyaml=0 +vikunja: containers=0 volumes=0 stackdir-appyaml=0 + +08:29:15 docmost deployed= False hold= None +08:29:15 romm deployed= False hold= None +08:29:15 vikunja deployed= False hold= None diff --git a/documentation/audits/update-rulings-2026-09-23/70-digest-pull-by-tag-and-digest.txt b/documentation/audits/update-rulings-2026-09-23/70-digest-pull-by-tag-and-digest.txt new file mode 100644 index 00000000..fb4f375b --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/70-digest-pull-by-tag-and-digest.txt @@ -0,0 +1,10 @@ +installed: redis@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 +ref with tag AND digest: redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 +docker.io/library/redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 +pull rc=0 + image: redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 +Redis server v=7.4.11 sha=00000000:0 malloc=jemalloc-5.3.0 bits=64 build=40ff01a501d8e4b6 +negative control — a digest that does not exist: +Error response from daemon: manifest for redis@sha256:0000000000000000000000000000000000000000000000000000000000000000 not found: manifest unknown: manifest unknown +rc=0 +NOTE (added after the run): the 'rc=0' on the last line is the exit code of the 'tail' in that pipe, NOT of docker pull — the refusal is the daemon's own sentence above it. diff --git a/documentation/audits/update-rulings-2026-09-23/80-teardown-repoint-live.txt b/documentation/audits/update-rulings-2026-09-23/80-teardown-repoint-live.txt new file mode 100644 index 00000000..1112dab6 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/80-teardown-repoint-live.txt @@ -0,0 +1,17 @@ +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + sync_interval: 15m + token: + username: "" +hub: +0 + +sync: 200 {'ok': True, 'data': {'ok': True, 'message': 'Sablonok naprakészek — nincs változás'}, 'message': 'S +02844ae decoys: eight cases for the probe target rule (R-630) +https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git +0 +controller.yaml == the saved pre-rulings copy + +02844ae0a579e8bf66882d01598c2212264a1b78 refs/heads/main + diff --git a/documentation/audits/update-rulings-2026-09-23/81-teardown-drill-repo.txt b/documentation/audits/update-rulings-2026-09-23/81-teardown-drill-repo.txt new file mode 100644 index 00000000..57a188a9 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/81-teardown-drill-repo.txt @@ -0,0 +1,7 @@ +02844ae decoys: eight cases for the probe target rule (R-630) +405a127 probe target: explicit container for paperless-ngx and immich; ambiguity refused (R-630) +dcb6b1d CHANGELOG: the six proven moves (R-462) +8898b1d termix: 2.5.0 -> 2.8.0 (proven on 9202, R-462) +0b283d2 sonarr: 4.0.19 -> 4.0.20 (proven on 9202, R-462) +b7b0479 radarr: 6.3.0 -> 6.4.4 (proven on 9202, R-462) +(drill commits made this session, now reset away: 193001a 419f243 57a4be3 88613e2 610ed1a c71807f; drill HEAD == origin == live == 02844ae0a579; image lines identical) diff --git a/documentation/audits/update-rulings-2026-09-23/82-teardown-guest.txt b/documentation/audits/update-rulings-2026-09-23/82-teardown-guest.txt new file mode 100644 index 00000000..269069dd --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/82-teardown-guest.txt @@ -0,0 +1,22 @@ +containers now: +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.262.1 Up 13 minutes (healthy) +filebrowser gtstef/filebrowser:1.3.3-stable Up 36 minutes (healthy) +gokapi f0rc3/gokapi:v1.9.6 Restarting (1) 33 seconds ago +paperless-postgres postgres:16-alpine Up 32 minutes (healthy) +paperless-redis redis:7-alpine Up 32 minutes (healthy) +paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 Up 32 minutes (healthy) +privatebin privatebin/pdo:2.0.6 Up 32 minutes (healthy) +traefik traefik:v3.6.7 Up 35 hours +/opt/upg removed +removed docmost/docmost:0.95.0 +removed docmost/docmost:0.96.0 +removed rommapp/romm:5.0.0 +removed rommapp/romm:5.3.0 +removed vikunja/vikunja:2.3.0 +absent vikunja/vikunja:2.4.0 +removed vikunja/vikunja:2.5.0 +removed vikunja/vikunja:2.6.0 +romm drive folder removed +calibre-web documents downloads emby immich jellyfin komga media nextcloud paperless-ngx plex radarr roms sonarr +no leftover volumes +/dev/loop1 69G 43G 23G 67% /var/lib/felhom diff --git a/documentation/audits/update-rulings-2026-09-23/README.md b/documentation/audits/update-rulings-2026-09-23/README.md new file mode 100644 index 00000000..485b65fe --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/README.md @@ -0,0 +1,178 @@ +# Update rulings 2026-09-23 — the undo spike, the ladder spike, the memory watch + +Venue: scratch guest **9202** on demo-hp, controller v0.262.1, pointed at the **drill catalog** +(`admin/app-catalog-drill`, reset to live `02844ae0a579` first) with `update.health_timeout: 90s`. +The live catalog carried no test reference at any point (control 3 in `02-three-controls.txt`). +Architecture document for the area: `architecture/09-update-architecture.md` (§3 decisions 11–18, +§6.1a, §6.4). + +**Method honesty.** The product has no undo path, so the undo was performed BY HAND, in the order +the product would take it, each step timed. Where a product path exists it was used: the Update, the +Start, the Remove, the backup. Lifting the hold used the operator CLI `--clear-restore-hold` plus a +controller restart (the only exit that exists today). **An attempt to decrypt the app's secrets so +compose could be run by hand was refused by the session's safety guard; that route was dropped, not +worked around** — the product's own Start supplies the secrets. + +--- + +## Part 1 — the automatic undo, by hand + +### Does a product path already load a safety dump back? + +**Yes, one — and the update never calls it.** `backup.Manager.rollbackSafetyDump` +(`internal/backup/offbox_reconstitute.go:413`) re-applies a `pre-restore-*` undo copy through +`ImportDump` — but only inside `ReconstituteFromOffsite`. `runGuardedUpdate` → `failAndHold` +(`internal/stacks/update.go:766`) writes the safety dump in phase 3 and never reads it again. +**And `failAndHold` deletes the journal's pre-update definition copies** (`removePreUpdateCopies`), +so after a hold the old definition survives only in the recovery unit's `compose/` directory. It was +there in all three cases because a backup preceded each update; that is not guaranteed. + +### The four cases — each a REAL migration that then failed a deliberately wrong probe + +| case | edge (drill) | what migrated | the safety dump | old version on the migrated data, nothing loaded | the load | undo → healthy | seed A (before backup) | seed B (after backup) | +|---|---|---|---|---|---|---|---|---| +| **PostgreSQL** docmost | 0.95.0 → 0.96.0, held 95.3 s | 4 migrations, 42 → 48 tables (`docmost-31`) | 135 816 B, DB only, **holds B** (tier copy does not) | **REFUSES** — *corrupted migrations: previously executed migration 20260824T211732-page-title-trgm-index is missing* | product semantics **FAILS** (below); fixed load **1.38 s** | **14.8 s** | yes | **yes** | +| **MariaDB** romm | 5.0.0 → 5.3.0, held 102.6 s | alembic 0095 → 0128, 27 → 39 tables | 62 943 B, DB only, **holds B** | **REFUSES** — *Can't locate revision identified by '0128_hltb_main_story_column'* | product semantics **1.25 s**, rc 0 — **12 new tables left behind** | **36.5 s** | yes | **yes** | +| **Volume data, no DB server** vikunja | 2.3.0 → 2.6.0, held 93.3 s | *Ran all migrations successfully* (SQLite in a volume) | **none** — *the app has no database — nothing to copy (no-op)* | **STARTS AND SERVES** in 0.7 s, A, B and B's attachment read back | — | 0.7 s | yes | **yes** | +| vikunja, **if it had refused** | the only other copy: the tier unit's volume tar | — | — | — | volume put back from the tier copy **0.86 s** | — | yes | **NO — lost** | +| **Files on disk** | vikunja's attachment (a file in the files volume, written after the backup); romm's two drive folders | none touched the files | the undo copy never holds files | the attachment read back after the update AND after the undo | — | — | — | attachment yes | + +Three different apps, three mechanisms of refusal now measured (Nextcloud's version check §4, docmost's +migration ledger, RomM's alembic revision) — and in each refusing case **the undo made the old version +start**, because it met the data it knew. + +**romm's drive folders were empty before, after the update and after the undo** (tree hash +`e3b0c442…` throughout) — so they measured nothing, and are reported as unmeasured, not as "untouched". + +### The finding that shapes decision 15's build: the product's loader cannot undo a migration + +`ImportDump` replays a `pg_dump --clean --if-exists` file over the live database. **Over a database +the new version migrated, that FAILS**: the new version created six tables (`oauth_*`, +`public_spaces`, `siem_destinations`) whose foreign keys point at old tables, and the dump's own +`DROP … workspaces_pkey` is refused — *cannot drop constraint workspaces_pkey on table +public.workspaces because other objects depend on it* — rc 3 in 0.40 s, database unchanged +(`docmost-45`). The same loader on MariaDB "succeeds" (`FOREIGN_KEY_CHECKS=0`) and leaves the new +version's **12 tables** behind; RomM 5.0.0 happens to ignore them. + +**What worked:** empty the schema and load the copy **in ONE transaction** — +`DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump, `psql --single-transaction +ON_ERROR_STOP=1`: rc 0 in 1.38 s, 42 tables, 48 ledger rows, `pg_trgm` and `unaccent` back. + +### The wrong case — a truncated undo copy + +| engine / loader | exit | what it left | is the outcome honest? | +|---|---|---|---| +| PostgreSQL, product semantics | rc 3 | unchanged (it failed on the foreign key before reaching the cut) | yes — nothing moved; the hold's named copy is intact | +| **PostgreSQL, the fixed atomic load** | **rc 0 (!)** | **42 tables and 48 ledger rows — and 0 users, 0 spaces, 0 constraints, 0 indexes** (`docmost-47`, measured in a scratch database beside the real one) | **NO.** psql treats end-of-file inside a `COPY` as end of data and commits. The undo would report success, the old version would start on an EMPTY database, and a health check would pass on it. No hold. | +| MariaDB, product semantics | rc 1 in 0.94 s | **half-replaced**: alembic back to 0095, later tables still the new version's — neither state | partly — the load is not transactional; the hold names the tier copy, which can still bring the app back | + +**The check that separates them already exists in the files:** the whole PostgreSQL copy ends with +`-- PostgreSQL database dump complete` (and a `\unrestrict` line), the MariaDB copy with +`-- Dump completed`; the truncated copies have neither. **`ValidateDump` does not look for them** — it +checks the header and one `CREATE TABLE` (`internal/appbackup/dbdump.go:415`, read from source), so it +would accept both truncated copies. + +### Seconds + +| | find the copy | old definition + pin back | load | start → healthy (old probe) | undo total | +|---|---|---|---|---|---| +| docmost | 0.04 s | 0.03 s | 1.38 s | 14.8 s | **≈ 16 s** | +| romm | 0.04 s | 0.04 s | 1.25 s | 36.5 s | **≈ 38 s** | +| vikunja | — | 0.02 s | none | 0.7 s | **≈ 1 s** | + +Plus the failing health wait that precedes any undo (`update.health_timeout`, 90 s here, 5 min by +default). Lifting the hold by hand cost 15.8 s (CLI + restart) and is NOT part of a product undo, +which would never hold in the first place. + +### What the product must add — the list decision 15's build starts from + +1. **Keep the pre-update copies until the undo is over** — compose, applied definition, pin **and the + old `.felhom.yml`**. `failAndHold` deletes the first three today; the fourth was never kept. +2. **The undo step inside `failAndHold`:** `pinBack` (exists) → DB service up alone → validated load → + full start → **health check with the OLD `.felhom.yml` probe** (the new one may name a port the old + version does not answer — it did in this spike by construction) → `undone`, or HOLD. +3. **The loader must empty the database first, atomically.** PostgreSQL: drop and recreate exactly + the schemas the dump creates, in the same transaction as the load. MariaDB: drop every table first + (`FOREIGN_KEY_CHECKS=0`); its DDL is not transactional, so a failed load is a HOLD with a sentence + saying the database is in neither state. +4. **Validate the copy's completion marker before loading**, both engines. A truncated PostgreSQL copy + is otherwise a silent, successful load of an empty database. +5. **Apps with no database server need their own last-second copy**: a tar of the data volumes at + safety-dump time (vikunja: 2.2 MB, put back in 0.86 s). Without it, an app whose old version + refuses loses everything written since the last backup. Vikunja's old version happens to start — + **a per-app fact the test record should carry**, not a rule. +6. **Files on disk:** nothing measured was touched by a migration, and the undo does nothing to files. + A step that does rewrite files must say so — decision 13's *files may change* mark is that place. +7. **The Start button says `start completed` (200) while the app crash-loops** (docmost and romm, both + negative controls). The undo's success must be the probe, never the start's return. +8. **No retry:** after an undo the badge reads „Frissítés elérhető" again, because the catalog is still + ahead. The caller must remember the failed step (decision 15), or it presses it every night. + +--- + +## Part 2 — the ladder + +### Today's behaviour, measured (`51-part2-ladder-today.txt`) + +vikunja installed at **2.3.0**; the drill catalog then got step **B = 2.4.0** (`610ed1a66dee`) and +step **C = 2.5.0** (`c71807f5d47e`); one sync, one press. **Result: `done` in 9.5 s, pin, installed +record, live compose and container all `vikunja/vikunja:2.5.0`; 2.4.0 never ran.** The box jumps. + +### Can the box see the steps? No (`50-part2-clone-depth.txt`) + +Both 9202 (drill) and 9201 (live) hold a **shallow clone of depth 1**: `is-shallow-repository true`, +`rev-list --count HEAD` = **1**, one commit visible per template. Source agrees: +`sync.go:283` clones `--depth 1`, `sync.go:300` fetches `--depth 1`. **The brief's premise that "the +box keeps a git clone" with history is wrong** — it keeps a clone of the newest commit only. + +### The format — two options, one recommended + +**Size is not the cost.** The whole catalog history is 529 KiB packed, 286 commits. + +**The cost is that history does not contain the tested steps.** romm's compose has **16 commits, 3 +of which move an image**; the step that works today is *15f9ebf's images with f4eb94f's template* +(two workers, 768M) — **the commit that moved the image is the one that OOM-looped on demo-hp**. A box +walking history would apply the definition that is known to be broken. + +| option | what it is | cost | +|---|---|---| +| **A — `update_ladder:` in `.felhom.yml` (recommended)** | per app, one entry per step: `from` and `to` refs per service, the digest per ref (decision 17), the test record (verdict, date, harness version, memory peak), the marks (`files_may_change`, `needs_person: ""`), and — for every step that is NOT the last — the step's own complete definition in `templates//steps/.yml`. The last step is the current `docker-compose.yml`. | every image move adds an entry (the catalog gate enforces it: no entry, no move — decision 13); an intermediate definition that needs a fix is fixed in its step file too, or the step is retired; entries older than the support window are pruned. `.felhom.yml` already flows to frozen apps (§5.4), so the box sees the ladder while frozen; step files are read from the clone. | +| B — the box walks the catalog's git history | deepen the clone, treat each image-moving commit as a step | every box carries full history; a history rewrite breaks every box; a commit is not a tested step and cannot be made one retroactively; **and it applies the broken intermediate definition measured above** | + +The test record and the marks of decision 13 live in the same entry, so one gate reads one place. + +--- + +## Part 3 — the memory watch in the harness (`app-catalog-felhom.eu/scripts/upgrade-test.py` v2) + +After an edge reads back, the harness runs the new version for `--soak` seconds (default 600) under +four light callers and samples every 15 s, per container: memory against the compose limit, the +**kernel's own `oom_kill` counter read host-side** (`/sys/fs/cgroup/system.slice/docker-.scope/ +memory.events` — works on images with no shell, and does not depend on Docker's OOMKilled flag, which +has read false for real kills on this kind of guest), the peak, and restarts. A kill or a restart → +`failed`; a peak above 80 % → mark `memory_tight`. + +Run on 9202 in `/opt/upg` with raw compose (no controller in the path), after Part 1's apps were +removed so the fixed `container_name`s could not collide. + +| edge | template | verdict | memory | +|---|---|---|---| +| **M1old** — the red-proof | as promoted, catalog `15f9ebf`: 512M, four workers | **failed** | seeded, migrated (2 lines), read back — then **first kernel OOM kill at +76 s**, peak 512 MiB = **100 %**, restarts **0** (the container kept running, which is how it hid on demo-hp), Docker's OOMKilled read true here; abort `refuses` | +| **M1** — the positive control | current: 768M, two workers | **proven** + mark `memory_tight` | 608.5 s under 11 429 requests (5 712 × 200, 5 717 × 401): **0 kernel OOM kills, 0 restarts**, peak 621 MiB = **81 %** of 768 MiB — the watch passes the fix and still flags the thin headroom R-635 left open | + +**Is ten minutes long enough?** For this failure, under load, yes by a wide margin: +76 s. demo-hp's +first kill came two hours in only because nothing was loading the app. What ten minutes cannot see is +a leak that grows over days; that is a monitoring question (R-636), not a test-bench one. + +Evidence: `harness/M1old.log`, `harness/evidence/M1old/` (verdict, memory samples, logs), and the same +for M1. + +--- + +## Controls and teardown + +- `02-three-controls.txt` — 9202 followed the drill catalog; 9201 stayed on the live one; the live + catalog's `main` was `02844ae0a579` before and after. +- `60-part1-teardown.txt` — docmost, romm, vikunja removed through the product; romm's drive data was + kept by the product (R-442's fail-closed refusal on 9202, as in every earlier drill) and removed by + name at the end. diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-10-deploy-seed.txt b/documentation/audits/update-rulings-2026-09-23/docmost-10-deploy-seed.txt new file mode 100644 index 00000000..24665666 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-10-deploy-seed.txt @@ -0,0 +1,18 @@ +07:58:58 deploy docmost at the drill FROM pin +07:58:58 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +08:00:39 [1] deployed, controller state=unhealthy, pinned={'docmost': 'docmost/docmost:0.95.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} +08:00:39 [1] NOTE: the controller's own state is 'unhealthy', not 'running' — recorded, not treated as a failure; the fixture's front-door wait is the real gate +08:00:39 deployed: True +08:00:41 {"pinned_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "installed_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "catalog_images": null, "live_compose_image_lines": ["image: docmost/docmost:0.95.0", "image: postgres:16-alpine", "image: redis:7-alpine"], "docker_inspect": ["docmost docmost/docmost:0.95.0 running=true restarts=0", "docmost-postgres postgres:16-alpine running=true restarts=0", "docmost-redis redis:7-alpine running=true restarts=0"]} +08:00:42 docmost: /api/auth/setup http=200 rc=0 +08:00:42 seed A: True +08:00:42 docmost: login as the seeded user http=200 ok=True +08:00:42 C1 A reads back: True +08:00:48 login 200 {"success":true,"status":200} +nR5cCI6IkpXVCJ9.eyJzdWIiOiIwMWEwY2NkYS0xZWViLTc4NTItOGRjNi1lMjNmZmExMzE0ODEiLCJlbWFpbCI6ImRyaWxsLTQ2NjFhMWVhQGdhdGUuaW52YWxpZCIsIndvcmtzcGFjZUlkIjoiMDFhMGNjZGEtMWVmMi03OGU1LTkwZjUtMDhlYTEyNDI5Mzk1IiwidHlwZSI6ImFjY2VzcyIsInNlc3Npb25JZCI6IjAxYTBjY2RhLTM5N2MtNzQ1ZS1iZmI2LWYyZWZmZDg0YWQ0MiIsImlhdCI6MTc5MDE0MzI0OCwiZXhwIjoxNzk3OTE5MjQ4LCJpc3MiOiJEb2Ntb3N0In0.uvcxRF50aQGZA_DniKLWO9BhFV8rEI1AYsOmNVxlUNU + +08:00:59 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +08:01:39 [4] backup idle; last=None +08:01:39 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0ccdb-009f-7975-a992-d0486e2b67e4","name":"drillB3c2580","description":"","slug":"drillb3c2580","logo":null,"visibility":"private","defaultRol +08:01:40 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False) +08:01:40 B reads back: True diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-20-break.txt b/documentation/audits/update-rulings-2026-09-23/docmost-20-break.txt new file mode 100644 index 00000000..951a18e4 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-20-break.txt @@ -0,0 +1,16 @@ +08:01:53 pg before update (tables | last 3 migrations | count): +42 +20260620T010047-personal-spaces +20260529T125146-bases +20260509T121236-labels +48 + +08:01:54 [5] drill commit 419f24385ddf: docmost docmost/docmost:0.95.0 -> docmost/docmost:0.96.0 (push rc=0) + + checks: + - type: http + port: 3999 + +# --- English copy (localisation slice 5, R-560) ------------------------- +08:01:58 badge caught up after 4.5 s +08:01:58 {"hu": [{"title": "Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.", "text": "Frissítés elérhető — ma"}], "en": [{"title": "A newer version of this app is available. Select the Update button to start it.", "text": "Update available — today"}]} diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-30-update-holds.txt b/documentation/audits/update-rulings-2026-09-23/docmost-30-update-holds.txt new file mode 100644 index 00000000..df329d4f --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-30-update-holds.txt @@ -0,0 +1,24 @@ +08:02:05 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:02:05 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:02:06 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None +08:02:07 + 2.1s phase=starting label=Indítás az új verzióval… err=None hold=None +08:02:08 + 3.1s phase=verifying label=Működés ellenőrzése… err=None hold=None +08:03:40 + 95.3s phase=failed label=A frissítés nem sikerült err=A(z) docmost frissítése 2026-09-23 08:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 08:01 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) docmost frissítése 2026-09-23 08:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 08:01 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. +08:03:40 {"final_phase": "failed", "update_error": "A(z) docmost frissítése 2026-09-23 08:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 08:01 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "hold_reason": "A(z) docmost frissítése 2026-09-23 08:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 08:01 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 95.4} +08:03:43 {"pinned_images": {"docmost": "docmost/docmost:0.96.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "installed_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "catalog_images": {"docmost": "docmost/docmost:0.96.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: docmost/docmost:0.96.0", "image: postgres:16-alpine", "image: redis:7-alpine"], "docker_inspect": []} +08:03:45 total 28 +drwxr-xr-x 3 root root 4096 Sep 23 06:03 . +drwxr-xr-x 57 root root 4096 Sep 13 20:22 .. +-rw-r--r-- 1 root root 3998 Sep 23 06:01 .felhom.yml +-rw------- 1 root root 1184 Sep 23 06:02 app.yaml +-rw-r--r-- 1 root root 3105 Sep 23 06:02 applied-compose.yml +-rw-r--r-- 1 root root 3105 Sep 23 06:02 docker-compose.yml +drwxr-xr-x 3 root root 4096 Sep 23 06:03 hold-logs + +hold-logs: +20260923T060338Z + +hold-logs/20260923T060338Z: +compose-logs.txt + + diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-31-after-hold-files.txt b/documentation/audits/update-rulings-2026-09-23/docmost-31-after-hold-files.txt new file mode 100644 index 00000000..75f90943 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-31-after-hold-files.txt @@ -0,0 +1,8 @@ +=== migration lines in the kept hold log +docmost | {"level":"info","time":"2026-09-23T06:02:19.400Z","pid":45,"hostname":"8ddf85c23281","context":"DatabaseMigrationService","msg":"Migration \"20260824T211732-page-title-trgm-index\" executed successfully"} +docmost | {"level":"info","time":"2026-09-23T06:02:19.400Z","pid":45,"hostname":"8ddf85c23281","context":"DatabaseMigrationService","msg":"Migration \"20260825T022612-oauth\" executed successfully"} +docmost | {"level":"info","time":"2026-09-23T06:02:19.400Z","pid":45,"hostname":"8ddf85c23281","context":"DatabaseMigrationService","msg":"Migration \"20260902T121326-siem-destinations\" executed successfully"} +docmost | {"level":"info","time":"2026-09-23T06:02:19.400Z","pid":45,"hostname":"8ddf85c23281","context":"DatabaseMigrationService","msg":"Migration \"20260904T171920-public-spaces\" executed successfully"} +=== where is the unit +/mnt/sys_drive/felhom-data/backups/primary/docmost +/var/lib/felhom/sys_drive/felhom-data/backups/primary/docmost diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-40-undo-step1-find-dump.txt b/documentation/audits/update-rulings-2026-09-23/docmost-40-undo-step1-find-dump.txt new file mode 100644 index 00000000..730bc5af --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-40-undo-step1-find-dump.txt @@ -0,0 +1,33 @@ +=== step 1: the safety dump the update wrote (unit: /mnt/sys_drive/felhom-data/backups/primary/docmost) +/mnt/sys_drive/felhom-data/backups/primary/docmost: +total 24 +drwxr-xr-x 5 root root 4096 2026-09-23 06:01:37.035847153 +0000 . +drwxr-xr-x 7 root root 4096 2026-09-23 06:00:59.201405588 +0000 .. +drwxr-xr-x 2 root root 4096 2026-09-23 06:01:37.035847153 +0000 compose +drwxr-xr-x 2 root root 4096 2026-09-23 06:02:05.524179635 +0000 db-dumps +-rw-r--r-- 1 root root 1298 2026-09-23 06:01:37.035847153 +0000 manifest.json +drwxr-xr-x 2 root root 4096 2026-09-23 06:01:02.101439434 +0000 volume-dumps + +/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps: +total 276 +drwxr-xr-x 2 root root 4096 2026-09-23 06:02:05.524179635 +0000 . +drwxr-xr-x 5 root root 4096 2026-09-23 06:01:37.035847153 +0000 .. +-rw-r--r-- 1 root root 134941 2026-09-23 06:00:59.482408868 +0000 docmost-postgres.sql +-rw-r--r-- 1 root root 135816 2026-09-23 06:02:05.511179483 +0000 pre-restore-20260923T060205Z-docmost-postgres.sql + +/mnt/sys_drive/felhom-data/backups/primary/docmost/volume-dumps: +total 50608 +drwxr-xr-x 2 root root 4096 2026-09-23 06:01:02.101439434 +0000 . +drwxr-xr-x 5 root root 4096 2026-09-23 06:01:37.035847153 +0000 .. +-rw-r--r-- 1 root root 51780096 2026-09-23 06:01:01.084427565 +0000 docmost_docmost_postgres_data.tar +-rw-r--r-- 1 root root 25088 2026-09-23 06:01:01.592631070 +0000 docmost_docmost_redis_data.tar +-rw-r--r-- 1 root root 1536 2026-09-23 06:01:01.967437870 +0000 docmost_docmost_storage.tar +newest undo copy: /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260923T060205Z-docmost-postgres.sql +size: 135816 bytes +holds: 42 CREATE TABLE, 42 COPY blocks, drop-first: 42 DROP TABLE IF EXISTS +migration ledger rows in the dump: +48 +seed B (space created after the backup) in the dump: 1 +seed B in the unit's canonical dump (the tier copy): 0 +unit compose image: +step 1 took 0.04s diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-41-undo-step2-definition-back.txt b/documentation/audits/update-rulings-2026-09-23/docmost-41-undo-step2-definition-back.txt new file mode 100644 index 00000000..4f8aca05 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-41-undo-step2-definition-back.txt @@ -0,0 +1,19 @@ +total 20 +drwxr-xr-x 2 root root 4096 Sep 23 06:01 . +drwxr-xr-x 5 root root 4096 Sep 23 06:01 .. +-rw-r--r-- 1 root root 3998 Sep 23 06:01 .felhom.yml +-rw------- 1 root root 445 Sep 23 06:01 app.yaml +-rw-r--r-- 1 root root 3105 Sep 23 06:01 docker-compose.yml +16: image: docmost/docmost:0.95.0 +56: image: postgres:16-alpine +80: image: redis:7-alpine +=== the pre-update copies the journal kept: 0 (failAndHold removes them) +=== pin now: +pinned_images: + docmost: docmost/docmost:0.95.0 + docmost-postgres: postgres:16-alpine + docmost-redis: redis:7-alpine +16: image: docmost/docmost:0.95.0 +56: image: postgres:16-alpine +80: image: redis:7-alpine +step 2 took 0.03s diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-42-undo-step2b-lift-hold.txt b/documentation/audits/update-rulings-2026-09-23/docmost-42-undo-step2b-lift-hold.txt new file mode 100644 index 00000000..cd7eaa7e --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-42-undo-step2b-lift-hold.txt @@ -0,0 +1,14 @@ +81: port: 3000 +/usr/local/bin/felhom-controller --config /opt/docker/felhom-controller/controller.yaml +flag provided but not defined: -list-restore-holds +Usage of /usr/local/bin/felhom-controller: + -abandon-extend int + R-241 (operator): extend a running abandonment countdown by N days from now, then exit. Refuses when no countdown is running. + -abandon-status + NOW RESTART THE CONTROLLER, or it will keep refusing to start the app: + systemctl restart felhom-controller-bootstrap.service + Check the app's data first — the undo copies are in its unit's db-dumps dir. +flag provided but not defined: -list-restore-holds +Usage of /usr/local/bin/felhom-controller: + -abandon-extend int +step 2b (hold lifted + controller restart) took 15.80s diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-43-negative-control-old-on-migrated.txt b/documentation/audits/update-rulings-2026-09-23/docmost-43-negative-control-old-on-migrated.txt new file mode 100644 index 00000000..7eed92f5 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-43-negative-control-old-on-migrated.txt @@ -0,0 +1,16 @@ +08:08:07 hold_reason after lift: None state: degraded pinned: {'docmost': 'docmost/docmost:0.95.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} +08:08:09 product start -> 200 {'ok': True, 'message': 'Stack docmost start completed'} +08:08:57 after 45 s: docmost docmost/docmost:0.95.0 Restarting (1) Less than a second ago +docmost-postgres postgres:16-alpine Up 58 seconds (healthy) +docmost-redis redis:7-alpine Up 58 seconds (healthy) + +08:08:59 {"level":"error","time":"2026-09-23T06:08:47.595Z","pid":45,"hostname":"c89a9567fa85","context":"DatabaseMigrationService","msg":"Failed to run database migration. Exiting program."} +{"level":"error","time":"2026-09-23T06:08:47.595Z","pid":45,"hostname":"c89a9567fa85","context":"DatabaseMigrationService","err":{"type":"Error","message":"corrupted migrations: previously executed migration 20260824T211732-page-title-trgm-index is missing","stack":"Error: corrupted migrations: previously executed migration 20260824T211732-page-title-trgm-index is missing\n at #ensureNoMissingMigrations (/app/node_modules/.pnpm/kysely@0.28.17/node_modules/kysely/dist/cjs/migration/migrator.js:495:23)\n at #getState (/app/node_modules/.pnpm/kysely@0.28.17/node_modules/kysely/dist/cjs/migration/migrator.js:447:40)\n at process.processTicksAndRejections (node:internal/process/task_queues:103:5)\n at async run (/app/node_modules/.pnpm/kysely@0.28.17/node_modules/kysely/dist/cjs/migration/migrator.js:417:31)\n at async /app/node_modules/.pnpm/kysely@0.28.17/node_modules/kysely/dist/cjs/kysely.js:578:32\n at async DefaultConnectionProvider.provideConnection (/app/node_modules/.pnpm/kysely@0.28.17/node_modules/kysely/dist/cjs/driver/default-connection-provider.js:12:20)\n at async #migrate (/app/node_modules/.pnpm/kysely@0.28.17/node_modules/kysely/dist/cjs/migration/migrator.js:273:20)\n at async MigrationService.migrateToLatest (/app/apps/server/dist/database/services/migration.service.js:36:36)\n at async DatabaseModule.onApplicationBootstrap (/app/apps/server/dist/database/database.module.js:59:13)\n at async callModuleBootstrapHook (/app/node_modules/.pnpm/@nestjs+core@11.1.27_@nestjs+common@11.1.27_class-transformer@0.5.1_class-validator@0.1_d57b2ebf04f40b3fdc1575d79e91a4e1/node_modules/@nestjs/core/hooks/on-app-bootstrap.hook.js:51:9)\n at async NestApplication.callBootstrapHook (/app/node_modules/.pnpm/@nestjs+core@11.1.27_@nestjs+common@11.1.27_class-transformer@0.5.1_class-validator@0.1_d57b2ebf04f40b3fdc1575d79e91a4e1/node_modules/@nestjs/core/nest-application-context.js:274:13)\n at async NestApplication.init (/app/node_modules/.pnpm/@nestjs+core@11.1.27_@nestjs+common@11.1.27_class-transformer@0.5.1_class-validator@0.1_d57b2ebf04f40b3fdc1575d79e91a4e1/node_modules/@nestjs/core/nest-application.js:107:9)\n at async NestApplication.listen (/app/node_modules/.pnpm/@nestjs+core@11.1.27_@nestjs+common@11.1.27_class-transformer@0.5.1_class-validator@0.1_d57b2ebf04f40b3fdc1575d79e91a4e1/node_modules/@nestjs/core/nest-application.js:177:13)\n at async bootstrap (/app/apps/server/dist/main.js:122:5)"},"msg":"corrupted migrations: previously executed migration 20260824T211732-page-title-trgm-index is missing"} +Exit status 1 + ELIFECYCLE  Command failed with exit code 1. +{"level":"error","time":"2026-09-23T06:08:56.732Z","pid":45,"hostname":"c89a9567fa85","context":"DatabaseMigrationService","msg":"Failed to run database migration. Exiting program."} +{"level":"error","time":"2026-09-23T06:08:56.732Z","pid":45,"hostname":"c89a9567fa85","context":"DatabaseMigrationService","err":{"type":"Error","message":"corrupted migrations: previously executed migration 20260824T211732-page-title-trgm-index is missing","stack":"Error: corrupted migrations: previously executed migration 20260824T211732-page-title-trgm-index is missing\n at #ensureNoMissingMigrations (/app/node_modules/.pnpm/kysely@0.28.17/node_modules/kysely/dist/cjs/migration/migrator.js:495:23)\n at #getState (/app/node_modules/.pnpm/kysely@0.28.17/node_modules/kysely/dist/cjs/migration/migrator.js:447:40)\n at process.processTicksAndRejections (node:internal/process/task_queues:103:5)\n at async run (/app/node_modules/.pnpm/kysely@0.28.17/node_modules/kysely/dist/cjs/migration/migrator.js:417:31)\n at async /app/node_modules/.pnpm/kysely@0.28.17/node_modules/kysely/dist/cjs/kysely.js:578:32\n at async DefaultConnectionProvider.provideConnection (/app/node_modules/.pnpm/kysely@0.28.17/node_modules/kysely/dist/cjs/driver/default-connection-provider.js:12:20)\n at async #migrate (/app/node_modules/.pnpm/kysely@0.28.17/node_modules/kysely/dist/cjs/migration/migrator.js:273:20)\n at async MigrationService.migrateToLatest (/app/apps/server/dist/database/services/migration.service.js:36:36)\n at async DatabaseModule.onApplicationBootstrap (/app/apps/server/dist/database/database.module.js:59:13)\n at async callModuleBootstrapHook (/app/node_modules/.pnpm/@nestjs+core@11.1.27_@nestjs+common@11.1.27_class-transformer@0.5.1_class-validator@0.1_d57b2ebf04f40b3fdc1575d79e91a4e1/node_modules/@nestjs/core/hooks/on-app-bootstrap.hook.js:51:9)\n at async NestApplication.callBootstrapHook (/app/node_modules/.pnpm/@nestjs+core@11.1.27_@nestjs+common@11.1.27_class-transformer@0.5.1_class-validator@0.1_d57b2ebf04f40b3fdc1575d79e91a4e1/node_modules/@nestjs/core/nest-application-context.js:274:13)\n at async NestApplication.init (/app/node_modules/.pnpm/@nestjs+core@11.1.27_@nestjs+common@11.1.27_class-transformer@0.5.1_class-validator@0.1_d57b2ebf04f40b3fdc1575d79e91a4e1/node_modules/@nestjs/core/nest-application.js:107:9)\n at async NestApplication.listen (/app/node_modules/.pnpm/@nestjs+core@11.1.27_@nestjs+common@11.1.27_class-transformer@0.5.1_class-validator@0.1_d57b2ebf04f40b3fdc1575d79e91a4e1/node_modules/@nestjs/core/nest-application.js:177:13)\n at async bootstrap (/app/apps/server/dist/main.js:122:5)"},"msg":"corrupted migrations: previously executed migration 20260824T211732-page-title-trgm-index is missing"} +Exit status 1 + ELIFECYCLE  Command failed with exit code 1. + +08:08:59 front door http: 404 diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-44-wrong-case-truncated-dump.txt b/documentation/audits/update-rulings-2026-09-23/docmost-44-wrong-case-truncated-dump.txt new file mode 100644 index 00000000..9847db19 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-44-wrong-case-truncated-dump.txt @@ -0,0 +1,11 @@ +app container stopped (DB kept up): exited +DB state BEFORE any load (the migrated state): 48 tables | ledger rows: 52 | newest: 20260904T171920-public-spaces +truncated copy: 67908 of 135816 bytes +load of the TRUNCATED copy: rc=3 in 0.34s +constraint siem_destinations_workspace_id_fkey on table public.siem_destinations depends on index public.workspaces_pkey +constraint public_spaces_workspace_id_fkey on table public.public_spaces depends on index public.workspaces_pkey +HINT: Use DROP ... CASCADE to drop the dependent objects too. +DB state AFTER the failed load: 48 tables | ledger rows: 52 | newest: 20260904T171920-public-spaces +old app on it: exited restarts=0 exit=1 +"msg":"corrupted migrations: previously executed migration 20260824T211732-page-title-trgm-index is missing" +wrong case total 31.37s diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-45-full-dump-product-semantics.txt b/documentation/audits/update-rulings-2026-09-23/docmost-45-full-dump-product-semantics.txt new file mode 100644 index 00000000..ce47d9a9 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-45-full-dump-product-semantics.txt @@ -0,0 +1,13 @@ +=== the WHOLE undo copy, loaded exactly as the product's ImportDump loads it (psql -v ON_ERROR_STOP=1 --single-transaction) +rc=3 in 0.40s +ERROR: cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it +DETAIL: constraint oauth_clients_workspace_id_fkey on table public.oauth_clients depends on index public.workspaces_pkey +HINT: Use DROP ... CASCADE to drop the dependent objects too. +DB state after: 48 tables | ledger rows: 52 | newest: 20260904T171920-public-spaces +=== the tables the NEW version created that the old dump does not know +oauth_authorization_codes +oauth_clients +oauth_grants +oauth_tokens +public_spaces +siem_destinations diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-46-atomic-truncated-then-full.txt b/documentation/audits/update-rulings-2026-09-23/docmost-46-atomic-truncated-then-full.txt new file mode 100644 index 00000000..d2f2e654 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-46-atomic-truncated-then-full.txt @@ -0,0 +1,10 @@ +=== WRONG CASE, atomic form: empty the schema + the TRUNCATED copy, ONE transaction +rc=0 in 0.81s +DETAIL: drop cascades to table kysely_migration +DB state after the failed atomic load (must be the MIGRATED state, untouched): 42 tables | ledger rows: 48 | newest: 20260620T010047-personal-spaces + +=== RIGHT CASE, atomic form: empty the schema + the WHOLE copy, ONE transaction +rc=0 in 1.38s +DETAIL: drop cascades to extension pg_trgm +DB state after (expect 42 tables, 48 ledger rows, newest 20260620T010047-personal-spaces): 42 tables | ledger rows: 48 | newest: 20260620T010047-personal-spaces +extensions: plpgsql,pg_trgm,unaccent diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-47-what-a-truncated-success-leaves.txt b/documentation/audits/update-rulings-2026-09-23/docmost-47-what-a-truncated-success-leaves.txt new file mode 100644 index 00000000..9686d392 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-47-what-a-truncated-success-leaves.txt @@ -0,0 +1,9 @@ +full rc=0 +trunc rc=0 +== scratch_full: tables=42 constraints(pk/fk/unique)=189 indexes=147 users=1 spaces=2 ledger=48 +== scratch_trunc: tables=42 constraints(pk/fk/unique)=0 indexes=0 users=0 spaces=0 ledger=48 +== where the cut fell: rifiers; Type: TABLE DATA; Schema: public; Owner: - -- COPY public.page_verifiers (id, page_verification_id, user_id, is_primary, added_by_id, created_at) FROM +== the completion marker pg_dump writes at the end: + full copy : 0 + truncated : 0 +scratch databases dropped diff --git a/documentation/audits/update-rulings-2026-09-23/docmost-48-undo-step4-start-and-readback.txt b/documentation/audits/update-rulings-2026-09-23/docmost-48-undo-step4-start-and-readback.txt new file mode 100644 index 00000000..8ac32510 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/docmost-48-undo-step4-start-and-readback.txt @@ -0,0 +1,10 @@ +08:11:30 started + +08:11:42 health (the OLD version's own probe: http :3000 answered through the front door) = True after 14.8s +08:11:44 docmost/docmost:0.95.0 running restarts=0 +"msg":"Nest application successfully started" + +08:11:45 docmost: login as the seeded user http=200 ok=True +08:11:45 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False) +08:11:45 READBACK after the undo: seed A (before the backup)=True seed B (after the backup, only in the safety dump)=True +08:11:47 {"pinned_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "installed_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "catalog_images": {"docmost": "docmost/docmost:0.96.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: docmost/docmost:0.95.0", "image: postgres:16-alpine", "image: redis:7-alpine"], "docker_inspect": ["docmost docmost/docmost:0.95.0 running=true restarts=0", "docmost-postgres postgres:16-alpine running=true restarts=0", "docmost-redis redis:7-alpine running=true restarts=0"]} diff --git a/documentation/audits/update-rulings-2026-09-23/evidence-docmost/hold-log-compose-logs.txt b/documentation/audits/update-rulings-2026-09-23/evidence-docmost/hold-log-compose-logs.txt new file mode 100644 index 00000000..2e117723 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/evidence-docmost/hold-log-compose-logs.txt @@ -0,0 +1,41 @@ +docmost | $ pnpm --filter ./apps/server run start:prod +docmost | $ cross-env NODE_ENV=production node dist/main +docmost-postgres | +docmost-redis | 1:C 23 Sep 2026 08:01:02.584 # WARNING Memory overcommit must be enabled! Without it, a background save or replication may fail under low memory condition. Being disabled, it can also cause failures without low memory condition, see https://github.com/jemalloc/jemalloc/issues/1328. To fix this issue add 'vm.overcommit_memory = 1' to /etc/sysctl.conf and then reboot or run the command 'sysctl vm.overcommit_memory=1' for this to take effect. +docmost-redis | 1:C 23 Sep 2026 08:01:02.584 * oO0OoO0OoO0Oo Redis is starting oO0OoO0OoO0Oo +docmost-postgres | PostgreSQL Database directory appears to contain a database; Skipping initialization +docmost-redis | 1:C 23 Sep 2026 08:01:02.584 * Redis version=7.4.11, bits=64, commit=00000000, modified=0, pid=1, just started +docmost-postgres | +docmost-redis | 1:C 23 Sep 2026 08:01:02.584 * Configuration loaded +docmost-postgres | 2026-09-23 08:01:02.662 CEST [1] LOG: starting PostgreSQL 16.15 on x86_64-pc-linux-musl, compiled by gcc (Alpine 15.2.0) 15.2.0, 64-bit +docmost | (node:45) ExperimentalWarning: localStorage is not available because --localstorage-file was not provided. +docmost-postgres | 2026-09-23 08:01:02.663 CEST [1] LOG: listening on IPv4 address "0.0.0.0", port 5432 +docmost | (Use `node --trace-warnings ...` to show where the warning was created) +docmost-postgres | 2026-09-23 08:01:02.663 CEST [1] LOG: listening on IPv6 address "::", port 5432 +docmost | {"level":"info","time":"2026-09-23T06:02:18.725Z","pid":45,"hostname":"8ddf85c23281","context":"RedisModule","msg":"default: the connection was successfully established"} +docmost-redis | 1:M 23 Sep 2026 08:01:02.584 * Increased maximum number of open files to 10032 (it was originally set to 1024). +docmost | {"level":"info","time":"2026-09-23T06:02:18.968Z","pid":45,"hostname":"8ddf85c23281","context":"DatabaseModule","msg":"Establishing database connection"} +docmost | {"level":"info","time":"2026-09-23T06:02:18.996Z","pid":45,"hostname":"8ddf85c23281","context":"DatabaseModule","msg":"Database connection successful"} +docmost-redis | 1:M 23 Sep 2026 08:01:02.584 * monotonic clock: POSIX clock_gettime +docmost | {"level":"info","time":"2026-09-23T06:02:19.400Z","pid":45,"hostname":"8ddf85c23281","context":"DatabaseMigrationService","msg":"Migration \"20260824T211732-page-title-trgm-index\" executed successfully"} +docmost-postgres | 2026-09-23 08:01:02.670 CEST [1] LOG: listening on Unix socket "/var/run/postgresql/.s.PGSQL.5432" +docmost-postgres | 2026-09-23 08:01:02.679 CEST [29] LOG: database system was shut down at 2026-09-23 08:01:00 CEST +docmost-postgres | 2026-09-23 08:01:02.687 CEST [1] LOG: database system is ready to accept connections +docmost-redis | 1:M 23 Sep 2026 08:01:02.586 * Running mode=standalone, port=6379. +docmost-redis | 1:M 23 Sep 2026 08:01:02.586 * Server initialized +docmost-redis | 1:M 23 Sep 2026 08:01:02.586 * Reading RDB base file on AOF loading... +docmost-redis | 1:M 23 Sep 2026 08:01:02.586 * Loading RDB produced by version 7.4.11 +docmost-redis | 1:M 23 Sep 2026 08:01:02.586 * RDB age 52 seconds +docmost-redis | 1:M 23 Sep 2026 08:01:02.586 * RDB memory usage when created 0.90 Mb +docmost-redis | 1:M 23 Sep 2026 08:01:02.586 * RDB is base AOF +docmost-redis | 1:M 23 Sep 2026 08:01:02.586 * Done loading RDB, keys loaded: 0, keys expired: 0. +docmost | {"level":"info","time":"2026-09-23T06:02:19.400Z","pid":45,"hostname":"8ddf85c23281","context":"DatabaseMigrationService","msg":"Migration \"20260825T022612-oauth\" executed successfully"} +docmost | {"level":"info","time":"2026-09-23T06:02:19.400Z","pid":45,"hostname":"8ddf85c23281","context":"DatabaseMigrationService","msg":"Migration \"20260902T121326-siem-destinations\" executed successfully"} +docmost | {"level":"info","time":"2026-09-23T06:02:19.400Z","pid":45,"hostname":"8ddf85c23281","context":"DatabaseMigrationService","msg":"Migration \"20260904T171920-public-spaces\" executed successfully"} +docmost-redis | 1:M 23 Sep 2026 08:01:02.586 * DB loaded from base file appendonly.aof.1.base.rdb: 0.000 seconds +docmost | {"level":"info","time":"2026-09-23T06:02:19.465Z","pid":45,"hostname":"8ddf85c23281","context":"NestApplication","msg":"Nest application successfully started"} +docmost-redis | 1:M 23 Sep 2026 08:01:02.587 * DB loaded from incr file appendonly.aof.1.incr.aof: 0.001 seconds +docmost | {"level":"info","time":"2026-09-23T06:02:19.478Z","pid":45,"hostname":"8ddf85c23281","context":"NestApplication","msg":"Listening on http://127.0.0.1:3000 / https://docs.enkisfelhom.hu"} +docmost-redis | 1:M 23 Sep 2026 08:01:02.587 * DB loaded from append only file: 0.001 seconds +docmost-redis | 1:M 23 Sep 2026 08:01:02.587 * Opening AOF incr file appendonly.aof.1.incr.aof on server start +docmost-redis | 1:M 23 Sep 2026 08:01:02.587 * Ready to accept connections tcp diff --git a/documentation/audits/update-rulings-2026-09-23/evidence-romm-hold-log.txt b/documentation/audits/update-rulings-2026-09-23/evidence-romm-hold-log.txt new file mode 100644 index 00000000..7b8cb6dd --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/evidence-romm-hold-log.txt @@ -0,0 +1,116 @@ +romm | INFO: [RomM][init][2026-09-23 08:15:25] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:15:25] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:15:25] | |__) |___ _ __ ___ | \ / | +romm-redis | 1:C 23 Sep 2026 08:13:58.183 # WARNING Memory overcommit must be enabled! Without it, a background save or replication may fail under low memory condition. Being disabled, it can also cause failures without low memory condition, see https://github.com/jemalloc/jemalloc/issues/1328. To fix this issue add 'vm.overcommit_memory = 1' to /etc/sysctl.conf and then reboot or run the command 'sysctl vm.overcommit_memory=1' for this to take effect. +romm-db | 2026-09-23 08:13:58+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:11.4.13+maria~ubu2404 started. +romm | INFO: [RomM][init][2026-09-23 08:15:25] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:15:25] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:15:25] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:15:25] +romm | INFO: [RomM][init][2026-09-23 08:15:25] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:15:25] +romm-db | 2026-09-23 08:13:58+02:00 [Warn] [Entrypoint]: /sys/fs/cgroup///memory.pressure not writable, functionality unavailable to MariaDB +romm-redis | 1:C 23 Sep 2026 08:13:58.183 * oO0OoO0OoO0Oo Redis is starting oO0OoO0OoO0Oo +romm-db | 2026-09-23 08:13:58+02:00 [Note] [Entrypoint]: Switching to dedicated user 'mysql' +romm-db | 2026-09-23 08:13:58+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:11.4.13+maria~ubu2404 started. +romm-db | 2026-09-23 08:13:58+02:00 [Note] [Entrypoint]: MariaDB upgrade not required +romm-redis | 1:C 23 Sep 2026 08:13:58.183 * Redis version=7.4.11, bits=64, commit=00000000, modified=0, pid=1, just started +romm-db | 2026-09-23 8:13:58 0 [Note] Starting MariaDB 11.4.13-MariaDB-ubu2404 source revision 170b1d70737be6f134448f51713cdc1ae215b420 server_uid w6iGLIckToxP3WCDRLUzm1Y/EIs= as process 1 +romm-redis | 1:C 23 Sep 2026 08:13:58.183 * Configuration loaded +romm-redis | 1:M 23 Sep 2026 08:13:58.184 * Increased maximum number of open files to 10032 (it was originally set to 1024). +romm-redis | 1:M 23 Sep 2026 08:13:58.184 * monotonic clock: POSIX clock_gettime +romm-redis | 1:M 23 Sep 2026 08:13:58.185 * Running mode=standalone, port=6379. +romm-redis | 1:M 23 Sep 2026 08:13:58.186 * Server initialized +romm-redis | 1:M 23 Sep 2026 08:13:58.186 * Reading RDB base file on AOF loading... +romm-redis | 1:M 23 Sep 2026 08:13:58.186 * Loading RDB produced by version 7.4.11 +romm-redis | 1:M 23 Sep 2026 08:13:58.186 * RDB age 103 seconds +romm-redis | 1:M 23 Sep 2026 08:13:58.186 * RDB memory usage when created 0.90 Mb +romm-redis | 1:M 23 Sep 2026 08:13:58.186 * RDB is base AOF +romm-redis | 1:M 23 Sep 2026 08:13:58.186 * Done loading RDB, keys loaded: 0, keys expired: 0. +romm-redis | 1:M 23 Sep 2026 08:13:58.186 * DB loaded from base file appendonly.aof.1.base.rdb: 0.000 seconds +romm-redis | 1:M 23 Sep 2026 08:13:58.256 * DB loaded from incr file appendonly.aof.1.incr.aof: 0.070 seconds +romm | INFO: [RomM][init][2026-09-23 08:15:25] Version: 5.3.0 +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: Compressed tables use zlib 1.3 +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: Number of transaction pools: 1 +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: Using crc32 + pclmulqdq instructions +romm-db | 2026-09-23 8:13:58 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: Using Linux native AIO +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: innodb_buffer_pool_size_max=8388608m, innodb_buffer_pool_size=128m +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: Completed initialization of buffer pool +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: File system buffers for log disabled (block size=512 bytes) +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: End of log at LSN=3042523 +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: Opened 3 undo tablespaces +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: 128 rollback segments in 3 undo tablespaces are active. +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: Setting file './ibtmp1' size to 12.000MiB. Physically writing the file full; Please wait ... +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: File './ibtmp1' size is now 12.000MiB. +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: log sequence number 3042523; transaction id 2444 +romm | INFO: [RomM][init][2026-09-23 08:15:25] +romm-redis | 1:M 23 Sep 2026 08:13:58.256 * DB loaded from append only file: 0.070 seconds +romm-redis | 1:M 23 Sep 2026 08:13:58.256 * Opening AOF incr file appendonly.aof.1.incr.aof on server start +romm | INFO: [RomM][init][2026-09-23 08:15:25] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm-redis | 1:M 23 Sep 2026 08:13:58.256 * Ready to accept connections tcp +romm-redis | 1:M 23 Sep 2026 08:14:59.034 * 10000 changes in 60 seconds. Saving... +romm-redis | 1:M 23 Sep 2026 08:14:59.036 * Background saving started by pid 51 +romm-redis | 51:C 23 Sep 2026 08:14:59.155 * DB saved on disk +romm-redis | 51:C 23 Sep 2026 08:14:59.156 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 0 MB +romm-redis | 1:M 23 Sep 2026 08:14:59.238 * Background saving terminated with success +romm-db | 2026-09-23 8:13:58 0 [Note] Plugin 'FEEDBACK' is disabled. +romm | INFO: [RomM][init][2026-09-23 08:15:25] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:15:25] Running database migrations +romm | INFO: [RomM][init][2026-09-23 08:15:41] Database migrations succeeded +romm | INFO: [RomM][startup][2026-09-23 08:15:53] Running startup tasks +romm | INFO: [RomM][startup][2026-09-23 08:15:53] Cleared 2 job(s) left behind by the old scheduler +romm | INFO: [RomM][startup][2026-09-23 08:15:53] Initializing cache with fixtures data +romm | INFO: [RomM][startup][2026-09-23 08:15:53] Startup tasks completed +romm | INFO: [RomM][init][2026-09-23 08:15:54] Starting backend +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: Loading buffer pool(s) from /var/lib/mysql/ib_buffer_pool +romm | INFO: [RomM][init][2026-09-23 08:15:54] Starting RQ cron scheduler +romm-db | 2026-09-23 8:13:58 0 [Note] Plugin 'wsrep-provider' is disabled. +romm-db | 2026-09-23 8:13:58 0 [Note] InnoDB: Buffer pool(s) load completed at 260923 8:13:58 +romm-db | 2026-09-23 8:14:01 0 [Note] Server socket created on IP: '0.0.0.0', port: '3306'. +romm-db | 2026-09-23 8:14:01 0 [Note] Server socket created on IP: '::', port: '3306'. +romm-db | 2026-09-23 8:14:01 0 [Note] mariadbd: Event Scheduler: Loaded 0 events +romm-db | 2026-09-23 8:14:01 0 [Note] mariadbd: ready for connections. +romm-db | Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution +romm-db | 2026-09-23 8:15:24 16 [Warning] Aborted connection 16 to db: 'romm' user: 'romm' host: '172.21.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:15:24 15 [Warning] Aborted connection 15 to db: 'romm' user: 'romm' host: '172.21.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:15:41 31 [Warning] Aborted connection 31 to db: 'romm' user: 'romm' host: '172.21.0.4' (Got an error reading communication packets) +romm | INFO: [RomM][init][2026-09-23 08:15:54] Starting RQ worker +romm | INFO: [RomM][init][2026-09-23 08:15:54] Starting RQ scan worker +romm | INFO: [RomM][init][2026-09-23 08:15:56] Starting nginx +romm | 2026/09/23 08:15:56 [notice] 128#128: js vm init njs: 00007A0B0A878B00 +romm | INFO: [RomM][init][2026-09-23 08:15:56] 🚀 RomM is now available at http://0.0.0.0:8080 +romm | 08:15:56 Worker bfc3044001f04323ab0fcd969f27f134: started with PID 104, version 2.12.0 +romm | 08:15:56 Worker bfc3044001f04323ab0fcd969f27f134: subscribing to channel rq:pubsub:bfc3044001f04323ab0fcd969f27f134 +romm | 08:15:56 *** Listening on high, default, low... +romm | 08:15:56 Acquired scheduler lock for low +romm | 08:15:56 Acquired scheduler lock for high +romm | 08:15:56 Acquired scheduler lock for default +romm | 08:15:56 Loading cron configuration from tasks.cron_config +romm | 08:15:56 Scheduler for low, high, default started with PID 143 +romm | 08:15:57 Worker 1c07db4039a94c7db8fac0f5a05b08ea: started with PID 107, version 2.12.0 +romm | 08:15:57 Worker 1c07db4039a94c7db8fac0f5a05b08ea: subscribing to channel rq:pubsub:1c07db4039a94c7db8fac0f5a05b08ea +romm | 08:15:57 *** Listening on scans... +romm | 08:15:57 Acquired scheduler lock for scans +romm | 08:15:57 Scheduler for scans started with PID 146 +romm | INFO: [RomM][cron_config][2026-09-23 08:16:05] Scheduled 'build_recommendations' at '30 5 * * *' +romm | INFO: [RomM][cron_config][2026-09-23 08:16:05] Scheduled 'cleanup_zip_cache' at '0 4 * * *' +romm | INFO: [RomM][cron_config][2026-09-23 08:16:05] Scheduled 'cleanup_netplay' at '*/30 * * * *' +romm | INFO: [RomM][cron_config][2026-09-23 08:16:05] Scheduled 'cleanup_upload_tmp' at '0 * * * *' +romm | 08:16:05 Registered 'tasks.tasks.run_task_by_name' to run on low with cron schedule '30 5 * * *' +romm | 08:16:05 Registered 'tasks.tasks.run_task_by_name' to run on low with cron schedule '0 4 * * *' +romm | 08:16:05 Registered 'tasks.tasks.run_task_by_name' to run on low with cron schedule '*/30 * * * *' +romm | 08:16:05 Registered 'tasks.tasks.run_task_by_name' to run on low with cron schedule '0 * * * *' +romm | 08:16:05 Successfully registered 4 cron jobs from 'tasks.cron_config' +romm | 08:16:05 CronScheduler 02b0b707a3fd:101:7e5404: starting... +romm | 08:16:05 CronScheduler 02b0b707a3fd:101:7e5404: registering birth... +romm | INFO: [RomM][nginx][2026-09-23 08:16:00] 127.0.0.1 | - | GET / 200 | 5878 | Unknown Unknown | 0.000 +romm | INFO: [RomM][server][2026-09-23 08:16:12] Started server process [124] +romm | INFO: [RomM][on][2026-09-23 08:16:12] Waiting for application startup. +romm | INFO: [RomM][on][2026-09-23 08:16:12] Application startup complete. +romm | INFO: [RomM][logs][2026-09-23 08:16:12] Log stream forwarder started +romm | INFO: [RomM][server][2026-09-23 08:16:12] Started server process [125] +romm | INFO: [RomM][on][2026-09-23 08:16:12] Waiting for application startup. +romm | INFO: [RomM][on][2026-09-23 08:16:12] Application startup complete. +romm | INFO: [RomM][nginx][2026-09-23 08:16:30] 127.0.0.1 | - | GET / 200 | 5878 | Unknown Unknown | 0.000 diff --git a/documentation/audits/update-rulings-2026-09-23/fixtures.py b/documentation/audits/update-rulings-2026-09-23/fixtures.py new file mode 100644 index 00000000..3a681e96 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/fixtures.py @@ -0,0 +1,1008 @@ +#!/usr/bin/env python3 +"""Box-side seed/verify fixtures for walk.py, guest 9202. + +THE ONE RULE (R-156), carried verbatim from `app-catalog-felhom.eu/scripts/upgrade_fixtures.py`: +*nothing is ever seeded into a volume by hand.* Every seed here goes in through the app's OWN +interface — its HTTP API through the household's real front door (traefik, `Host: .`), +or its own CLI running inside its own container. A raw SQL INSERT or a planted file is never used. + +If an app has no non-browser route, its fixture returns None and the edge is recorded +`inconclusive — no non-browser seed route`, WITH WHAT WAS TRIED. That is a result, not a gap. + +Each fixture: + seed(w, sub, say) -> an opaque token, or None + verify(w, sub, tok, say) -> True / False +verify() must ask the APP, never the filesystem: a migration is supposed to rewrite files. +Where a fixture can prove itself (a negative control that must read as absent) it does so on EVERY +call, so a readback that has broken into always saying "found" fails instead of passing everything. +""" +import base64, json, re, secrets, time + + +def _gx(w, container, *cmd, timeout=240): + """Run a command inside the app's OWN container on 9202 (its own CLI, not our SQL).""" + import shlex + line = " ".join(shlex.quote(c) for c in cmd) + return w.guest(f"docker exec {container} {line} 2>&1", timeout=timeout) + + +# ============================================================================================= +class PrivateBin: + """PrivateBin's own JSON API. A paste is a POST and reading it back is a GET — an + application-level round trip. File-backed, no database: this single seed IS the file half.""" + sub = "paste" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200",)): + return None + marker = "upg-" + secrets.token_hex(8) + ct = base64.b64encode(marker.encode()).decode() + body = json.dumps({ + "v": 2, + "adata": [[base64.b64encode(secrets.token_bytes(16)).decode(), + base64.b64encode(secrets.token_bytes(8)).decode(), + 100000, 256, 128, "aes", "gcm", "none"], "plaintext", 0, 0], + "ct": ct, "meta": {"expire": "never"}}) + rc, code, out = w.app_curl(sub, "/", "-H", "X-Requested-With: JSONHttpRequest", + "-H", "Content-Type: application/json", + data=body, method="POST") + try: + j = json.loads(out) + except Exception: + say(f" privatebin: POST returned non-JSON (http {code}): {out[:200]}") + return None + if j.get("status") != 0 or not j.get("id"): + say(f" privatebin: POST refused: {out[:250]}") + return None + say(f" privatebin: seeded paste id={j['id']}") + return {"id": j["id"], "marker": ct} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200",), tries=36): + return False + # negative control, every call: a paste id that cannot exist must NOT read back + rc, code, out = w.app_curl(sub, "/?pasteid=" + secrets.token_hex(8), + "-H", "X-Requested-With: JSONHttpRequest") + if t["marker"] in out: + say(" privatebin: READBACK UNUSABLE — a paste id that cannot exist returned the marker") + return False + rc, code, out = w.app_curl(sub, "/?pasteid=" + t["id"], + "-H", "X-Requested-With: JSONHttpRequest") + got = code == "200" and t["marker"] in out + say(f" privatebin: readback http={code} marker_present={got}") + return got + + +# ============================================================================================= +class Docmost: + """Docmost's own REST API: create the first workspace+user, then prove the account survives by + asking the app to AUTHENTICATE it. Login is version-stable across the API churn.""" + sub = "docs" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302", "404")): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + body = json.dumps({"workspaceName": "drill", "name": "drill", "email": email, "password": pw}) + rc, code, out = w.app_curl(sub, "/api/auth/setup", "-H", "Content-Type: application/json", + data=body, method="POST") + say(f" docmost: /api/auth/setup http={code} rc={rc}") + if code not in ("200", "201"): + say(f" docmost: setup refused: {out[:250]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302", "404"), tries=36): + return False + # negative control: a password that was never set must NOT authenticate + bad = json.dumps({"email": t["email"], "password": "definitely-" + secrets.token_hex(8)}) + rc, code, _ = w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code in ("200", "201"): + say(" docmost: READBACK UNUSABLE — a wrong password authenticated") + return False + body = json.dumps({"email": t["email"], "password": t["pw"]}) + rc, code, out = w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=body, method="POST") + ok = code in ("200", "201") + say(f" docmost: login as the seeded user http={code} ok={ok}") + if not ok: + say(f" docmost: login body {out[:200]}") + return ok + + +# ============================================================================================= +class BookStack: + """BookStack mints no API token without a browser, so BOTH halves go through `php artisan` — + BookStack's OWN CLI, inside its own container, against its own User model. + + The exit code carries no information here (`bookstack:reset-mfa` exits 1 for a user it FOUND + and for one it did not), so the discriminator is the OUTPUT: the positive sentence required and + the not-found sentence required absent. The negative control runs on every verify. + + LIMITATION (R-460): this seeds the DATABASE half only. The FILE half needs the API token the + app cannot mint headlessly — so a bookstack edge is at best HALF-proven here. + """ + sub = "wiki" + + def _artisan(self, w, *args): + for path in ("/app/www/artisan", "/var/www/html/artisan"): + out = _gx(w, "bookstack", "php", path, *args) + if "Could not open input file" not in out: + return " ".join(out.split()) + return " ".join(out.split()) + + def _lookup(self, w, email): + out = self._artisan(w, "bookstack:reset-mfa", f"--email={email}") + found = f"Email: {email}" in out + missing = "could not be found" in out + if found == missing: + return None, out + return found, out + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + out = self._artisan(w, "bookstack:create-admin", f"--email={email}", + f"--name=drill-{secrets.token_hex(3)}", f"--password={pw}") + say(f" bookstack: artisan create-admin :: {out[:140]}") + if "successfully created" not in out: + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + say(" bookstack: the app never served /login") + return False + absent, _ = self._lookup(w, f"nobody-{secrets.token_hex(6)}@gate.invalid") + if absent is not False: + say(f" bookstack: READBACK UNUSABLE — an email that cannot exist did not read absent ({absent})") + return False + found, out = self._lookup(w, t["email"]) + say(f" bookstack: readback of the seeded account found={found} :: {out[:140]}") + return found is True + + +# ============================================================================================= +class Gitea: + """Gitea's own admin CLI creates the first user; its own REST API (basic auth) then creates a + repository and reads it back. Both are the app's own interfaces.""" + sub = "git" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302")): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = _gx(w, "gitea", "su", "git", "-c", + f"gitea admin user create --username {user} --password {pw} " + f"--email {user}@gate.invalid --admin --must-change-password=false") + say(f" gitea: admin user create :: {' '.join(out.split())[:140]}") + if "has been successfully created" not in out and "successfully created" not in out: + return None + repo = "drillrepo" + secrets.token_hex(3) + rc, code, body = w.app_curl(sub, "/api/v1/user/repos", "-u", f"{user}:{pw}", + "-H", "Content-Type: application/json", + data=json.dumps({"name": repo, "private": True}), method="POST") + say(f" gitea: create repo http={code}") + if code not in ("201", "200"): + say(f" gitea: repo refused {body[:200]}") + return None + return {"user": user, "pw": pw, "repo": repo} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=36): + return False + rc, code, _ = w.app_curl(sub, f"/api/v1/repos/{t['user']}/nope{secrets.token_hex(4)}", + "-u", f"{t['user']}:{t['pw']}") + if code == "200": + say(" gitea: READBACK UNUSABLE — a repo that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/v1/repos/{t['user']}/{t['repo']}", + "-u", f"{t['user']}:{t['pw']}") + ok = code == "200" and t["repo"] in body + say(f" gitea: readback of the seeded repo http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Navidrome: + """Navidrome's own REST API: create the first admin through /auth/createAdmin, then prove the + account survives by logging in through the same door.""" + sub = "music" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302")): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, out = w.app_curl(sub, "/auth/createAdmin", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw}), method="POST") + say(f" navidrome: createAdmin http={code}") + if code not in ("200", "201"): + say(f" navidrome: refused {out[:200]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=36): + return False + bad = json.dumps({"username": t["user"], "password": "wrong-" + secrets.token_hex(6)}) + rc, code, _ = w.app_curl(sub, "/auth/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code in ("200", "201"): + say(" navidrome: READBACK UNUSABLE — a wrong password authenticated") + return False + body = json.dumps({"username": t["user"], "password": t["pw"]}) + rc, code, out = w.app_curl(sub, "/auth/login", "-H", "Content-Type: application/json", + data=body, method="POST") + ok = code in ("200", "201") + say(f" navidrome: login as the seeded user http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Vaultwarden: + """Vaultwarden's own account API: register an account, then prove it survives by asking the app + to issue a token for it (its own login endpoint, the household's own route).""" + sub = "vault" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/alive", want=("200",)): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + # Vaultwarden stores an already-hashed master key; the value is opaque to the server. + key = base64.b64encode(secrets.token_bytes(32)).decode() + body = json.dumps({"email": email, "name": "drill", "masterPasswordHash": key, + "key": "0." + base64.b64encode(secrets.token_bytes(48)).decode(), + "kdf": 0, "kdfIterations": 600000}) + rc, code, out = w.app_curl(sub, "/api/accounts/register", + "-H", "Content-Type: application/json", + data=body, method="POST") + say(f" vaultwarden: register http={code}") + if code not in ("200", "204"): + say(f" vaultwarden: refused {out[:250]}") + return None + return {"email": email, "key": key} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/alive", want=("200",), tries=36): + return False + def login(pwhash): + return w.app_curl(sub, "/identity/connect/token", + "-H", "Content-Type: application/x-www-form-urlencoded", + data=("grant_type=password&scope=api%20offline_access" + f"&client_id=web&deviceType=9&deviceIdentifier=drill" + f"&deviceName=drill&username={t['email']}&password={pwhash}"), + method="POST") + rc, code, _ = login(base64.b64encode(secrets.token_bytes(32)).decode()) + if code == "200": + say(" vaultwarden: READBACK UNUSABLE — a wrong master key authenticated") + return False + rc, code, out = login(t["key"].replace("+", "%2B").replace("=", "%3D").replace("/", "%2F")) + ok = code == "200" and "access_token" in out + say(f" vaultwarden: token for the seeded account http={code} ok={ok}") + if not ok: + say(f" vaultwarden: body {out[:200]}") + return ok + + +# ============================================================================================= +class Django: + """A Django app's OWN management CLI, inside its own container, against its own User model. + + Same category as BookStack's `php artisan`: the app's own code and its own ORM, never a raw SQL + INSERT and never a planted file (R-156). `createsuperuser --noinput` is Django's own documented + non-interactive route, and the readback asks the SAME ORM whether the account exists. + + THE FIXTURE PROVES ITSELF ON EVERY CALL: each verify() also asks for a username that cannot + exist and requires the answer False. A readback that has broken into always saying True + therefore fails instead of passing everything. + + LIMITATION, recorded rather than papered over: this seeds the DATABASE half only. An app whose + data is also FILES (adventurelog's images) has a file half this fixture does not touch. + """ + + def __init__(self, container, sub, ready_path="/", ready=("200", "302", "301", "404"), + python="python", workdir=None): + # `python` and `workdir` are per-app because the image decides them: adventurelog's + # interpreter is on PATH, tandoor ships a VENV and the bare `python` cannot import Django + # at all ("Couldn't import Django. Are you sure it's installed…"). Measured, not guessed. + self.container = container + self.sub = sub + self.ready_path = ready_path + self.ready = ready + self.python = python + self.workdir = workdir + + def _wd(self): + return f"-w {self.workdir} " if self.workdir else "" + + def _manage(self, w, code): + # -c is passed to `manage.py shell`; the app's own shell, its own ORM. + return w.guest( + f"docker exec {self._wd()}{self.container} {self.python} manage.py shell " + f"-c {json.dumps(code)} 2>&1", timeout=300) + + def _exists(self, w, username): + # ONE LINE, semicolon-separated. A `\n` inside a double-quoted shell argument reaches + # python as a literal backslash-n and is a SyntaxError — which is exactly how the first + # adventurelog run read as `inconclusive`. The fixture refused to guess, which is right, + # but the instrument was the thing that was broken. + out = self._manage(w, ( + "from django.contrib.auth import get_user_model; " + f"print('DRILL_ANSWER=' + str(get_user_model().objects.filter(username={username!r}).exists()))" + )) + m = re.search(r"DRILL_ANSWER=(True|False)", out) + return (m.group(1) == "True") if m else None, " ".join(out.split())[-300:] + + def seed(self, w, sub, say): + if not w.wait_app(sub, self.ready_path, want=self.ready, tries=90): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = w.guest( + f"docker exec -e DJANGO_SUPERUSER_PASSWORD={pw} {self._wd()}{self.container} " + f"{self.python} manage.py createsuperuser --noinput " + f"--username {user} --email {user}@gate.invalid 2>&1", timeout=300) + say(f" {self.container}: createsuperuser :: {' '.join(out.split())[:160]}") + got, detail = self._exists(w, user) + if got is not True: + say(f" {self.container}: the account did not appear in the app's own ORM :: {detail[:200]}") + return None + say(f" {self.container}: seeded superuser {user}") + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, self.ready_path, want=self.ready, tries=90): + say(f" {self.container}: the app never served {self.ready_path}") + return False + absent, detail = self._exists(w, "nobody" + secrets.token_hex(6)) + if absent is not False: + say(f" {self.container}: READBACK UNUSABLE — a username that cannot exist did not " + f"read as absent ({absent}) :: {detail[:200]}") + return False + found, detail = self._exists(w, t["user"]) + say(f" {self.container}: readback of the seeded account found={found}") + if found is not True: + say(f" {self.container}: :: {detail[:250]}") + return found is True + + +# ============================================================================================= +class Nextcloud: + """Nextcloud's OWN admin CLI, `occ`, inside its own container: its own code, its own user + backend. Not a SQL INSERT and not a planted file (R-156). + + `occ user:info` is the readback, and it PROVES ITSELF on every call: a uid that cannot exist + must answer "user not found". A readback that has broken into always succeeding therefore + fails instead of passing everything. + + This is the app chosen for the MariaDB engine-major edge (`09` §3 decision 5, R-469 lifted): + the app image does NOT move, only the `mariadb:` sidecar, so the edge carries exactly one + migration and a failure is readable. + """ + sub = "cloud" + + def _occ(self, w, *args, timeout=420): + import shlex + line = " ".join(shlex.quote(a) for a in args) + return w.guest(f"docker exec -u www-data nextcloud php occ {line} 2>&1", timeout=timeout) + + def _info(self, w, uid): + out = self._occ(w, "user:info", uid) + flat = " ".join(out.split()) + if "user not found" in flat.lower() or "could not be found" in flat.lower(): + return False, flat + if f"user_id: {uid}" in flat or f"- user_id: {uid}" in flat or f"user_id: {uid}" in out: + return True, flat + return None, flat + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/status.php", want=("200",), tries=120): + return None + uid = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = w.guest( + f"docker exec -u www-data -e OC_PASS={pw} nextcloud php occ user:add " + f"--password-from-env --display-name={uid} {uid} 2>&1", timeout=420) + say(f" nextcloud: occ user:add :: {' '.join(out.split())[:160]}") + got, flat = self._info(w, uid) + if got is not True: + say(f" nextcloud: the account did not appear via occ user:info :: {flat[:220]}") + return None + say(f" nextcloud: seeded user {uid}") + return {"uid": uid, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/status.php", want=("200",), tries=120): + say(" nextcloud: the app never served /status.php") + return False + absent, flat = self._info(w, "nobody" + secrets.token_hex(6)) + if absent is not False: + say(f" nextcloud: READBACK UNUSABLE — a uid that cannot exist did not read absent " + f"({absent}) :: {flat[:200]}") + return False + found, flat = self._info(w, t["uid"]) + say(f" nextcloud: readback of the seeded user found={found}") + if found is not True: + say(f" nextcloud: :: {flat[:250]}") + return found is True + + +# ============================================================================================= +class Grafana: + """Grafana's own HTTP API as the admin the DEPLOY created. The password is the one the + controller showed the household — read from the app's own `app.yaml`, not invented — and the + data (a folder) goes in and comes back through the app's own REST API.""" + sub = "grafana" + + def _auth(self, w, name="grafana"): + # app.yaml stores this ENCRYPTED (`ENC:…`), so it cannot be read back off the box — which + # is correct, and is why the harness uses the value IT generated for the deploy. + pw = (w.GENERATED.get(name) or {}).get("GF_SECURITY_ADMIN_PASSWORD") or "admin" + return f"admin:{pw}" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + return None + au = self._auth(w) + title = "drill-" + secrets.token_hex(5) + rc, code, body = w.app_curl(sub, "/api/folders", "-u", au, + "-H", "Content-Type: application/json", + data=json.dumps({"title": title}), method="POST") + say(f" grafana: create folder http={code}") + if code not in ("200", "201"): + say(f" grafana: refused {body[:220]}") + return None + try: + uid = json.loads(body)["uid"] + except Exception: + say(f" grafana: no uid in {body[:200]}") + return None + return {"uid": uid, "title": title} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + return False + au = self._auth(w) + rc, code, _ = w.app_curl(sub, "/api/folders/nope" + secrets.token_hex(5), "-u", au) + if code == "200": + say(" grafana: READBACK UNUSABLE — a folder uid that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/folders/{t['uid']}", "-u", au) + ok = code == "200" and t["title"] in body + say(f" grafana: readback of the seeded folder http={code} ok={ok}") + return ok + + +# ============================================================================================= +class AudiobookShelf: + """audiobookshelf's own /init endpoint creates the first root account; its own /login proves + the account survived. Both are the app's own API.""" + sub = "audiobooks" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/status", want=("200",), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/init", "-H", "Content-Type: application/json", + data=json.dumps({"newRoot": {"username": user, "password": pw}}), + method="POST") + say(f" audiobookshelf: /init http={code}") + if code not in ("200", "204"): + say(f" audiobookshelf: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/status", want=("200",), tries=72): + return False + bad = json.dumps({"username": t["user"], "password": "wrong-" + secrets.token_hex(6)}) + rc, code, _ = w.app_curl(sub, "/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code == "200": + say(" audiobookshelf: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = w.app_curl(sub, "/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], "password": t["pw"]}), + method="POST") + ok = code == "200" and t["user"] in body + say(f" audiobookshelf: login as the seeded root http={code} ok={ok}") + return ok + + +# ============================================================================================= +class ActualBudget: + """Actual's own bootstrap API sets the server password; its own login proves it survived.""" + sub = "budget" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return None + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/account/bootstrap", + "-H", "Content-Type: application/json", + data=json.dumps({"password": pw}), method="POST") + say(f" actualbudget: /account/bootstrap http={code} :: {body[:140]}") + if code not in ("200", "201") or '"status":"ok"' not in body: + return None + return {"pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return False + def login(p): + return w.app_curl(sub, "/account/login", "-H", "Content-Type: application/json", + data=json.dumps({"loginMethod": "password", "password": p}), + method="POST") + rc, code, body = login("wrong-" + secrets.token_hex(6)) + if '"status":"ok"' in body: + say(" actualbudget: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = '"status":"ok"' in body + say(f" actualbudget: login with the seeded password http={code} ok={ok}") + if not ok: + say(f" actualbudget: body {body[:200]}") + return ok + + +# ============================================================================================= +class Mealie: + """Mealie ships a documented first-run admin. We log in as it through the app's own OAuth-style + token endpoint, create a recipe through the app's own API, and read the recipe back.""" + sub = "recipes" + + def _token(self, w, sub, pw="MyPassword"): + rc, code, body = w.app_curl( + sub, "/api/auth/token", "-H", "Content-Type: application/x-www-form-urlencoded", + data=f"username=changeme%40example.com&password={pw}", method="POST") + if code != "200": + return None, f"http={code} {body[:200]}" + try: + return json.loads(body)["access_token"], "" + except Exception: + return None, body[:200] + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/app/about", want=("200",), tries=90): + return None + tok, why = self._token(w, sub) + if not tok: + say(f" mealie: could not authenticate as the first-run admin :: {why}") + return None + name = "drill-" + secrets.token_hex(5) + rc, code, body = w.app_curl(sub, "/api/recipes", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"name": name}), method="POST") + say(f" mealie: create recipe http={code}") + if code not in ("200", "201"): + say(f" mealie: refused {body[:220]}") + return None + slug = body.strip().strip('"') + return {"slug": slug, "name": name} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/app/about", want=("200",), tries=90): + return False + tok, why = self._token(w, sub) + if not tok: + say(f" mealie: could not authenticate after the update :: {why}") + return False + rc, code, _ = w.app_curl(sub, "/api/recipes/nope" + secrets.token_hex(5), + "-H", f"Authorization: Bearer {tok}") + if code == "200": + say(" mealie: READBACK UNUSABLE — a slug that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/recipes/{t['slug']}", + "-H", f"Authorization: Bearer {tok}") + ok = code == "200" and t["name"] in body + say(f" mealie: readback of the seeded recipe http={code} ok={ok}") + return ok + + +# ============================================================================================= +class N8n: + """n8n's own owner-setup API creates the first account; its own login proves it survived.""" + sub = "auto" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/healthz", want=("200",), tries=90): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill" + secrets.token_hex(8) + "1" + rc, code, body = w.app_curl(sub, "/rest/owner/setup", "-H", "Content-Type: application/json", + data=json.dumps({"email": email, "firstName": "drill", + "lastName": "drill", "password": pw}), + method="POST") + say(f" n8n: /rest/owner/setup http={code}") + if code not in ("200", "201"): + say(f" n8n: refused {body[:220]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/healthz", want=("200",), tries=90): + return False + def login(p): + return w.app_curl(sub, "/rest/login", "-H", "Content-Type: application/json", + data=json.dumps({"emailOrLdapLoginId": t["email"], "password": p}), + method="POST") + rc, code, _ = login("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" n8n: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = code == "200" and t["email"] in body + say(f" n8n: login as the seeded owner http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Zipline: + """Zipline's own setup/login API. Zipline 4 creates the first user through its own endpoint.""" + sub = "img" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/healthcheck", want=("200",), tries=90): + if not w.wait_app(sub, "/", want=("200", "302", "307"), tries=30): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + for path in ("/api/auth/register", "/api/auth/setup"): + rc, code, body = w.app_curl(sub, path, "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw}), + method="POST") + say(f" zipline: {path} http={code} :: {body[:160]}") + if code in ("200", "201"): + return {"user": user, "pw": pw} + say(" zipline: neither register nor setup accepted a first user") + return None + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302", "307"), tries=60): + return False + def login(p): + return w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], "password": p}), + method="POST") + rc, code, _ = login("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" zipline: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = code == "200" + say(f" zipline: login as the seeded user http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Vikunja: + """Vikunja's own REST API: register a user, log in, create a project, read the project back. + Four calls, all the app's own front door.""" + sub = "tasks" + + def _token(self, w, sub, t, pw=None): + rc, code, body = w.app_curl(sub, "/api/v1/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], + "password": pw or t["pw"]}), method="POST") + if code != "200": + return None, f"http={code} {body[:160]}" + try: + return json.loads(body)["token"], "" + except Exception: + return None, body[:160] + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/v1/info", want=("200",), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/v1/register", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw, + "email": f"{user}@gate.invalid"}), + method="POST") + say(f" vikunja: register http={code}") + if code not in ("200", "201"): + say(f" vikunja: refused {body[:220]}") + return None + t = {"user": user, "pw": pw} + tok, why = self._token(w, sub, t) + if not tok: + say(f" vikunja: could not log in after registering :: {why}") + return None + title = "drill-" + secrets.token_hex(5) + # Vikunja CREATES with PUT, not POST — a POST answers `405 Method Not Allowed`, which + # reads like a broken fixture and is really the wrong verb. Measured 2026-09-21. + rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"title": title}), method="PUT") + say(f" vikunja: create project http={code}") + if code not in ("200", "201"): + say(f" vikunja: project refused {body[:220]}") + return None + t["title"] = title + t["pid"] = json.loads(body).get("id") + return t + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/v1/info", want=("200",), tries=72): + return False + bad, why = self._token(w, sub, t, pw="wrong-" + secrets.token_hex(6)) + if bad: + say(" vikunja: READBACK UNUSABLE — a wrong password authenticated") + return False + tok, why = self._token(w, sub, t) + if not tok: + say(f" vikunja: the seeded account no longer authenticates :: {why}") + return False + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{t['pid']}", + "-H", f"Authorization: Bearer {tok}") + ok = code == "200" and t["title"] in body + say(f" vikunja: readback of the seeded project http={code} ok={ok}") + return ok + + +# ============================================================================================= +class OpenGist: + """Opengist's own sign-up and sign-in FORMS. + + Two things had to be measured. Its sign-up is CSRF-protected: a bare POST answers 500 with an + HTML page, which reads like a broken app and is really a missing token — fetch the form, keep + its cookie, send its `_csrf` back. And its REST API refuses the account's own password + (`401 {"message":"Bad crendentials"}`) because it wants a token the app will not mint without a + browser. So the SEEDED DATA is the account itself and the READBACK is a real sign-in, which is + the same shape the docmost and navidrome fixtures use. + + LIMITATION, recorded rather than papered over: this is the DATABASE half. A gist's CONTENT is + not seeded, because that needs the API token above. + """ + sub = "gist" + + def _form(self, w, sub, path, jar, fields): + rc, code, html = w.app_curl(sub, path, "-b", jar, "-c", jar) + m = re.search(r'name="_csrf"[^>]*value="([^"]+)"', html or "") + if not m: + return None, f"no _csrf on {path} (http={code})" + body = "&".join([f"_csrf={m.group(1)}"] + [f"{k}={v}" for k, v in fields.items()]) + rc, code, out = w.app_curl(sub, path, "-b", jar, "-c", jar, + "-H", "Content-Type: application/x-www-form-urlencoded", + data=body, method="POST") + return code, out + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + jar = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, out = self._form(w, sub, "/register", jar, {"username": user, "password": pw}) + say(f" opengist: /register (with its own _csrf) http={code}") + if code not in ("200", "302", "303"): + say(f" opengist: refused {str(out)[:200]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + # Wait for the LOGIN FORM, not for the root page. Measured 2026-09-21: immediately after a + # successful update the root answers while /login does not yet carry its `_csrf`, so the + # sign-in silently fails and the app looks like it lost the account. It had not. + if not w.wait_app(sub, "/login", want=("200",), tries=72): + say(" opengist: /login never came back after the update") + return False + for _ in range(24): + rc, code, html = w.app_curl(sub, "/login") + if code == "200" and '_csrf' in (html or ""): + break + time.sleep(5) + jar = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, _ = self._form(w, sub, "/login", jar, + {"username": t["user"], "password": "wrong-" + secrets.token_hex(5)}) + rc, c2, home = w.app_curl(sub, "/", "-b", jar) + if t["user"] in (home or ""): + say(" opengist: READBACK UNUSABLE — a wrong password signed in") + return False + jar2 = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, _ = self._form(w, sub, "/login", jar2, {"username": t["user"], "password": t["pw"]}) + rc, c2, home = w.app_curl(sub, "/", "-b", jar2) + ok = t["user"] in (home or "") + say(f" opengist: sign-in as the seeded account http={code} name_on_page={ok}") + return ok + + +# ============================================================================================= +class Papra: + """Papra's own e-mail sign-up and sign-in endpoints.""" + sub = "papra" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + if not w.wait_app(sub, "/", want=("200", "302"), tries=30): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/auth/sign-up/email", + "-H", "Content-Type: application/json", + data=json.dumps({"email": email, "password": pw, + "name": "drill"}), method="POST") + say(f" papra: sign-up http={code}") + if code not in ("200", "201"): + say(f" papra: refused {body[:220]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return False + def signin(p): + return w.app_curl(sub, "/api/auth/sign-in/email", + "-H", "Content-Type: application/json", + data=json.dumps({"email": t["email"], "password": p}), method="POST") + rc, code, _ = signin("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" papra: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = signin(t["pw"]) + ok = code == "200" + say(f" papra: sign-in as the seeded account http={code} ok={ok}") + return ok + + +# ============================================================================================= +class HomeAssistant: + """Home Assistant's own onboarding API creates the owner account and hands back a code the + same API exchanges for a token. Both are the app's own documented non-browser route.""" + sub = "ha" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=120): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/onboarding/users", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "name": "drill", "username": user, + "password": pw, "language": "en"}), + method="POST") + say(f" home-assistant: /api/onboarding/users http={code}") + if code not in ("200", "201"): + say(f" home-assistant: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def _login(self, w, sub, user, pw): + """The app's own login flow: start it, then answer it. A 200 with a step_id of + `mfa`/`init` means the credentials were REFUSED; only `create_entry` is a pass.""" + rc, code, body = w.app_curl(sub, "/auth/login_flow", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "handler": ["homeassistant", None], + "redirect_uri": f"https://{sub}.felhom.invalid/", + "type": "authorize"}), method="POST") + if code not in ("200", "201"): + return None, f"flow start http={code} {body[:160]}" + try: + fid = json.loads(body)["flow_id"] + except Exception: + return None, body[:160] + rc, code, body = w.app_curl(sub, f"/auth/login_flow/{fid}", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "username": user, "password": pw}), + method="POST") + try: + j = json.loads(body) + except Exception: + return None, body[:160] + return (j.get("result") if j.get("type") == "create_entry" else None), body[:200] + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=120): + return False + bad, why = self._login(w, sub, t["user"], "wrong-" + secrets.token_hex(6)) + if bad: + say(" home-assistant: READBACK UNUSABLE — a wrong password authenticated") + return False + good, why = self._login(w, sub, t["user"], t["pw"]) + ok = bool(good) + say(f" home-assistant: login as the seeded owner ok={ok}") + if not ok: + say(f" home-assistant: {why}") + return ok + + +# ============================================================================================= +class Romm: + """RomM's own user API, driven the way RomM's own front end drives it. + + Three things had to be measured rather than guessed, and each one answered a 403 or a 422 that + looked like a different fault: RomM sets a **`romm_csrftoken` cookie** on any GET and requires + it back in an **`x-csrftoken` header** (a bare POST is `403 CSRF token verification failed`, + which reads like an auth problem); the fields go in the **JSON body**, not the query string (a + query-string POST is `422 Field required` for every field it was just given); and `email` is + required alongside username, password and role. + + On a fresh install with no admin the first `POST /api/users` is accepted unauthenticated; + afterwards it is not — which is what makes the readback (`POST /api/login` as that user) a real + authentication rather than a repeat of the seed. + + LIMITATION: this is the DATABASE half. RomM's other half is the ROM library on the drive, which + this does not populate. + """ + sub = "arcade" + + def _csrf(self, w, sub): + jar = f"/tmp/romm-{secrets.token_hex(4)}.jar" + w.app_curl(sub, "/api/heartbeat", "-c", jar) + out = w.sh(["bash", "-lc", f"grep -i csrf {jar} | awk '{{print $7}}'"]).stdout or "" + return jar, out.strip() + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/heartbeat", want=("200",), tries=120): + if not w.wait_app(sub, "/", want=("200", "302"), tries=30): + return None + jar, tok = self._csrf(w, sub) + if not tok: + say(" romm: no romm_csrftoken cookie was set on /api/heartbeat") + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl( + sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "email": f"{user}@gate.invalid", + "password": pw, "role": "admin"}), method="POST") + say(f" romm: POST /api/users http={code}") + if code not in ("200", "201"): + say(f" romm: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/heartbeat", want=("200",), tries=120): + return False + jar, tok = self._csrf(w, sub) + rc, code, _ = w.app_curl(sub, "/api/login", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{t['user']}:wrong-{secrets.token_hex(5)}", method="POST") + if code == "200": + say(" romm: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = w.app_curl(sub, "/api/login", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{t['user']}:{t['pw']}", method="POST") + ok = code == "200" + say(f" romm: login as the seeded user http={code} ok={ok}") + if not ok: + say(f" romm: body {body[:200]}") + return ok + + +FIXTURES = { + "home-assistant": HomeAssistant(), + "romm": Romm(), + "vikunja": Vikunja(), + "opengist": OpenGist(), + "papra": Papra(), + "mealie": Mealie(), + "n8n": N8n(), + "zipline": Zipline(), + "grafana": Grafana(), + "audiobookshelf": AudiobookShelf(), + "actualbudget": ActualBudget(), + "nextcloud": Nextcloud(), + "adventurelog": Django("adventurelog", "travel", "/admin/login/"), + "tandoor": Django("tandoor", "recipes", "/accounts/login/", + python="/opt/recipes/venv/bin/python", workdir="/opt/recipes"), + "privatebin": PrivateBin(), + "docmost": Docmost(), + "bookstack": BookStack(), + "gitea": Gitea(), + "navidrome": Navidrome(), + "vaultwarden": Vaultwarden(), +} diff --git a/documentation/audits/update-rulings-2026-09-23/harness/M1.log b/documentation/audits/update-rulings-2026-09-23/harness/M1.log new file mode 100644 index 00000000..9d23d938 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/M1.log @@ -0,0 +1,123 @@ +[06:37:57] M1: deploying romm at FROM {'romm': 'rommapp/romm:5.0.0'} +[06:38:45] FROM settled=True in 36.7s :: {"romm": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-db": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-redis": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}} +[06:38:53] romm: POST /api/users http=201 +[06:38:55] romm: login as the seeded user http=200 ok=True +[06:38:55] C1 (seed reads back BEFORE): True +[06:38:55] M1: swapping to TO {'romm': 'rommapp/romm:5.3.0'} +[06:38:58] TO up -d rc=0 +[06:39:29] TO settled=True in 31.6s :: {"romm": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-db": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-redis": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}} +[06:39:29] migration lines observed: 2 +[06:39:46] romm: login as the seeded user http=200 ok=True +[06:39:46] RESULT (seed reads back AFTER): True +[06:39:46] memory watch: 600s, 4 callers on 8 path(s) at 172.18.0.6:8080 +[06:40:01] + 15s romm=615M/768M peak=620M kills=0 rs=0 romm-db=140M/384M peak=146M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=291 +[06:40:16] + 30s romm=615M/768M peak=620M kills=0 rs=0 romm-db=141M/384M peak=146M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=584 +[06:40:31] + 46s romm=615M/768M peak=620M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=877 +[06:40:47] + 61s romm=616M/768M peak=620M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=1169 +[06:41:02] + 76s romm=616M/768M peak=620M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=1459 +[06:41:17] + 91s romm=615M/768M peak=620M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=1751 +[06:41:32] + 106s romm=616M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=2042 +[06:41:47] + 122s romm=449M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=2334 +[06:42:03] + 137s romm=536M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=14M/128M peak=37M kills=0 rs=0 reqs=2431 +[06:42:18] + 152s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=2719 +[06:42:33] + 167s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=14M/128M peak=37M kills=0 rs=0 reqs=3012 +[06:42:48] + 182s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=3302 +[06:43:03] + 198s romm=614M/768M peak=621M kills=0 rs=0 romm-db=142M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=3594 +[06:43:19] + 213s romm=615M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=3885 +[06:43:34] + 228s romm=615M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=14M/128M peak=37M kills=0 rs=0 reqs=4175 +[06:43:49] + 243s romm=498M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=4466 +[06:44:04] + 258s romm=615M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=4753 +[06:44:19] + 274s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=5046 +[06:44:35] + 289s romm=478M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=5337 +[06:44:50] + 304s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=5626 +[06:45:05] + 319s romm=615M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=5918 +[06:45:20] + 334s romm=614M/768M peak=621M kills=0 rs=0 romm-db=142M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=6210 +[06:45:35] + 350s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=6500 +[06:45:51] + 365s romm=536M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=6790 +[06:46:06] + 380s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=16M/128M peak=37M kills=0 rs=0 reqs=7079 +[06:46:21] + 395s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=7369 +[06:46:36] + 410s romm=555M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=7653 +[06:46:51] + 426s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=7942 +[06:47:07] + 441s romm=615M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=8235 +[06:47:22] + 456s romm=614M/768M peak=621M kills=0 rs=0 romm-db=142M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=8527 +[06:47:37] + 471s romm=614M/768M peak=621M kills=0 rs=0 romm-db=142M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=8819 +[06:47:52] + 486s romm=615M/768M peak=621M kills=0 rs=0 romm-db=142M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=9111 +[06:48:07] + 502s romm=526M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=9401 +[06:48:23] + 517s romm=613M/768M peak=621M kills=0 rs=0 romm-db=142M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=9690 +[06:48:38] + 532s romm=456M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=9983 +[06:48:53] + 547s romm=608M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=14M/128M peak=37M kills=0 rs=0 reqs=10273 +[06:49:08] + 563s romm=608M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=17M/128M peak=37M kills=0 rs=0 reqs=10563 +[06:49:24] + 578s romm=608M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=17M/128M peak=37M kills=0 rs=0 reqs=10855 +[06:49:39] + 593s romm=490M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=17M/128M peak=37M kills=0 rs=0 reqs=11139 +[06:49:54] + 608s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=16M/128M peak=37M kills=0 rs=0 reqs=11429 +[06:49:54] memory watch: killed=False tight=['romm'] requests=11429 codes={'200': 5712, '401': 5717} +[06:49:54] M1: ABORT — putting the FROM images back +[06:53:05] ABORT: the app did NOT come back (rc=0, 182.4s) +{ + "harness_version": 2, + "edge": "M1", + "app": "romm", + "note": "catalog move 15f9ebf 5.0.0 -> 5.3.0 on the CURRENT template (768M, 2 workers)", + "from": { + "romm": "rommapp/romm:5.0.0" + }, + "to": { + "romm": "rommapp/romm:5.3.0" + }, + "verdict": "proven", + "seed_read_before": true, + "seed_read_after": true, + "healthy_after": true, + "migration_observed": "romm | \u001b[0;32mINFO: \u001b[0;34m[RomM]\u001b[0;95m[init]\u001b[0;36m[2026-09-23 08:38:58]\u001b[0;00m Running database migrations", + "abort": "refuses", + "abort_detail": "to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:49:34 148 [Warning] Aborted connection 148 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:50:01 171 [Warning] Aborted connection 171 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:50:01 165 [Warning] Aborted connection 165 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:50:01 164 [Warning] Aborted connection 164 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:50:01 168 [Warning] Aborted connection 168 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)", + "engine_state_after": null, + "memory": { + "soak_s": 608.5, + "requested_s": 600, + "requests": 11429, + "codes": { + "200": 5712, + "401": 5717 + }, + "first_kill": null, + "containers": { + "romm": { + "limit": 805306368, + "peak": 651239424, + "peak_pct": 0.809, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + }, + "romm-db": { + "limit": 402653184, + "peak": 155947008, + "peak_pct": 0.387, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + }, + "romm-redis": { + "limit": 134217728, + "peak": 39243776, + "peak_pct": 0.292, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + } + }, + "unmeasured": [] + }, + "marks": [ + "memory_tight" + ], + "duration_s": 31.6, + "measured_at": "2026-09-23T06:53:05Z", + "evidence": "evidence/M1", + "total_s": 908.0 +} +rc=0 diff --git a/documentation/audits/update-rulings-2026-09-23/harness/M1old.log b/documentation/audits/update-rulings-2026-09-23/harness/M1old.log new file mode 100644 index 00000000..4eea8826 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/M1old.log @@ -0,0 +1,93 @@ +[06:30:39] M1old: deploying romm at FROM {'romm': 'rommapp/romm:5.0.0'} +[06:31:27] FROM settled=True in 36.7s :: {"romm": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-db": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-redis": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}} +[06:31:34] romm: POST /api/users http=201 +[06:31:36] romm: login as the seeded user http=200 ok=True +[06:31:36] C1 (seed reads back BEFORE): True +[06:31:36] M1old: swapping to TO {'romm': 'rommapp/romm:5.3.0'} +[06:31:39] TO up -d rc=0 +[06:32:11] TO settled=True in 31.5s :: {"romm": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-db": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-redis": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}} +[06:32:11] migration lines observed: 2 +[06:32:44] romm: login as the seeded user http=200 ok=True +[06:32:44] RESULT (seed reads back AFTER): True +[06:32:44] memory watch: 600s, 4 callers on 8 path(s) at 172.18.0.6:8080 +[06:33:00] + 15s romm=511M/512M peak=512M kills=0 rs=0 romm-db=140M/384M peak=145M kills=0 rs=0 romm-redis=15M/128M peak=30M kills=0 rs=0 reqs=286 +[06:33:15] + 30s romm=509M/512M peak=512M kills=0 rs=0 romm-db=140M/384M peak=146M kills=0 rs=0 romm-redis=15M/128M peak=30M kills=0 rs=0 reqs=575 +[06:33:30] + 46s romm=510M/512M peak=512M kills=0 rs=0 romm-db=140M/384M peak=146M kills=0 rs=0 romm-redis=15M/128M peak=30M kills=0 rs=0 reqs=864 +[06:33:45] + 61s romm=508M/512M peak=512M kills=0 rs=0 romm-db=140M/384M peak=146M kills=0 rs=0 romm-redis=15M/128M peak=30M kills=0 rs=0 reqs=1157 +[06:34:00] + 76s romm=509M/512M peak=512M kills=1 rs=0 romm-db=140M/384M peak=146M kills=0 rs=0 romm-redis=15M/128M peak=30M kills=0 rs=0 reqs=1404 +[06:34:00] memory watch: STOPPING EARLY — {'t': 76.0, 'container': 'romm', 'oom_kills': 1, 'restarts': 0} +[06:34:01] memory watch: killed=True tight=['romm'] requests=1405 codes={'401': 704, '200': 701} +[06:34:01] VERDICT -> failed: the new version was OOM-killed or restarted under light load +[06:34:01] M1old: ABORT — putting the FROM images back +[06:37:15] ABORT: the app did NOT come back (rc=0, 182.4s) +{ + "harness_version": 2, + "edge": "M1old", + "app": "romm", + "note": "the same move on the template AS PROMOTED (512M, 4 workers) \u2014 must FAIL the memory watch", + "from": { + "romm": "rommapp/romm:5.0.0" + }, + "to": { + "romm": "rommapp/romm:5.3.0" + }, + "verdict": "failed", + "seed_read_before": true, + "seed_read_after": true, + "healthy_after": true, + "migration_observed": "romm | \u001b[0;32mINFO: \u001b[0;34m[RomM]\u001b[0;95m[init]\u001b[0;36m[2026-09-23 08:31:39]\u001b[0;00m Running database migrations", + "abort": "refuses", + "abort_detail": "ection 33 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:33:51 41 [Warning] Aborted connection 41 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:33:51 32 [Warning] Aborted connection 32 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:34:11 36 [Warning] Aborted connection 36 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:34:11 42 [Warning] Aborted connection 42 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:34:11 49 [Warning] Aborted connection 49 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)", + "engine_state_after": null, + "memory": { + "soak_s": 76.5, + "requested_s": 600, + "requests": 1405, + "codes": { + "401": 704, + "200": 701 + }, + "first_kill": { + "t": 76.0, + "container": "romm", + "oom_kills": 1, + "restarts": 0 + }, + "containers": { + "romm": { + "limit": 536870912, + "peak": 536875008, + "peak_pct": 1.0, + "oom_kills": 1, + "restarts": 0, + "oomkilled_flag": true, + "measured": true + }, + "romm-db": { + "limit": 402653184, + "peak": 153600000, + "peak_pct": 0.381, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + }, + "romm-redis": { + "limit": 134217728, + "peak": 32137216, + "peak_pct": 0.239, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + } + }, + "unmeasured": [] + }, + "marks": [], + "duration_s": 31.5, + "measured_at": "2026-09-23T06:37:15Z", + "evidence": "evidence/M1old", + "total_s": 395.7 +} +rc=0 diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/abort-refusal.txt b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/abort-refusal.txt new file mode 100644 index 00000000..2eab94ac --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/abort-refusal.txt @@ -0,0 +1,106 @@ +romm-redis | 1:C 23 Sep 2026 08:37:58.125 # WARNING Memory overcommit must be enabled! Without it, a background save or replication may fail under low memory condition. Being disabled, it can also cause failures without low memory condition, see https://github.com/jemalloc/jemalloc/issues/1328. To fix this issue add 'vm.overcommit_memory = 1' to /etc/sysctl.conf and then reboot or run the command 'sysctl vm.overcommit_memory=1' for this to take effect. +romm-redis | 1:C 23 Sep 2026 08:37:58.125 * oO0OoO0OoO0Oo Redis is starting oO0OoO0OoO0Oo +romm-redis | 1:C 23 Sep 2026 08:37:58.125 * Redis version=7.4.11, bits=64, commit=00000000, modified=0, pid=1, just started +romm-redis | 1:C 23 Sep 2026 08:37:58.125 * Configuration loaded +romm-redis | 1:M 23 Sep 2026 08:37:58.125 * Increased maximum number of open files to 10032 (it was originally set to 1024). +romm-redis | 1:M 23 Sep 2026 08:37:58.125 * monotonic clock: POSIX clock_gettime +romm-redis | 1:M 23 Sep 2026 08:37:58.128 * Running mode=standalone, port=6379. +romm-redis | 1:M 23 Sep 2026 08:37:58.128 * Server initialized +romm-redis | 1:M 23 Sep 2026 08:37:58.136 * Creating AOF base file appendonly.aof.1.base.rdb on server start +romm-redis | 1:M 23 Sep 2026 08:37:58.145 * Creating AOF incr file appendonly.aof.1.incr.aof on server start +romm-redis | 1:M 23 Sep 2026 08:37:58.145 * Ready to accept connections tcp +romm-redis | 1:M 23 Sep 2026 08:38:59.091 * 10000 changes in 60 seconds. Saving... +romm-redis | 1:M 23 Sep 2026 08:38:59.093 * Background saving started by pid 52 +romm-redis | 52:C 23 Sep 2026 08:38:59.225 * DB saved on disk +romm | INFO: [RomM][init][2026-09-23 08:51:18] +romm-redis | 52:C 23 Sep 2026 08:38:59.226 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 1 MB +romm-redis | 1:M 23 Sep 2026 08:38:59.295 * Background saving terminated with success +romm-redis | 1:M 23 Sep 2026 08:44:00.006 * 100 changes in 300 seconds. Saving... +romm-redis | 1:M 23 Sep 2026 08:44:00.007 * Background saving started by pid 233 +romm-redis | 233:C 23 Sep 2026 08:44:00.119 * DB saved on disk +romm-redis | 233:C 23 Sep 2026 08:44:00.120 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 1 MB +romm-redis | 1:M 23 Sep 2026 08:44:00.210 * Background saving terminated with success +romm-redis | 1:M 23 Sep 2026 08:49:01.065 * 100 changes in 300 seconds. Saving... +romm-redis | 1:M 23 Sep 2026 08:49:01.066 * Background saving started by pid 408 +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: 128 rollback segments in 3 undo tablespaces are active. +romm | INFO: [RomM][init][2026-09-23 08:51:18] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:51:18] +romm | INFO: [RomM][init][2026-09-23 08:51:18] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:51:18] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:51:18] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:51:27] Failed to run database migrations +romm-redis | 408:C 23 Sep 2026 08:49:01.191 * DB saved on disk +romm-redis | 408:C 23 Sep 2026 08:49:01.192 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 1 MB +romm | INFO: [RomM][init][2026-09-23 08:51:40] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:51:40] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:51:40] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:51:40] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:51:40] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:51:40] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:51:40] +romm-redis | 1:M 23 Sep 2026 08:49:01.268 * Background saving terminated with success +romm | INFO: [RomM][init][2026-09-23 08:51:40] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:51:40] +romm | INFO: [RomM][init][2026-09-23 08:51:40] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:51:40] +romm | INFO: [RomM][init][2026-09-23 08:51:40] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:51:40] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:51:40] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:51:49] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:52:15] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:52:15] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:52:15] | |__) |___ _ __ ___ | \ / | +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Setting file './ibtmp1' size to 12.000MiB. Physically writing the file full; Please wait ... +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: File './ibtmp1' size is now 12.000MiB. +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: log sequence number 45568; transaction id 14 +romm-db | 2026-09-23 8:38:04 0 [Note] Plugin 'FEEDBACK' is disabled. +romm-db | 2026-09-23 8:38:04 0 [Note] Plugin 'wsrep-provider' is disabled. +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Loading buffer pool(s) from /var/lib/mysql/ib_buffer_pool +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Buffer pool(s) load completed at 260923 8:38:04 +romm-db | 2026-09-23 8:38:05 0 [Note] Server socket created on IP: '0.0.0.0', port: '3306'. +romm-db | 2026-09-23 8:38:05 0 [Note] Server socket created on IP: '::', port: '3306'. +romm-db | 2026-09-23 8:38:05 0 [Note] mariadbd: Event Scheduler: Loaded 0 events +romm-db | 2026-09-23 8:38:05 0 [Note] mariadbd: ready for connections. +romm-db | Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution +romm-db | 2026-09-23 8:38:27 5 [Warning] Aborted connection 5 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:38:56 16 [Warning] Aborted connection 16 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm | INFO: [RomM][init][2026-09-23 08:52:15] | _ // _ \| '_ ` _ \| |\/| | +romm-db | 2026-09-23 8:38:56 15 [Warning] Aborted connection 15 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm | INFO: [RomM][init][2026-09-23 08:52:15] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:52:15] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:52:15] +romm | INFO: [RomM][init][2026-09-23 08:52:15] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:52:15] +romm | INFO: [RomM][init][2026-09-23 08:52:15] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:52:15] +romm | INFO: [RomM][init][2026-09-23 08:52:15] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:52:15] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:52:15] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:52:23] Failed to run database migrations +romm-db | 2026-09-23 8:39:12 19 [Warning] Aborted connection 19 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:41:46 35 [Warning] Aborted connection 35 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:41:46 30 [Warning] Aborted connection 30 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:41:54 44 [Warning] Aborted connection 44 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:41:54 29 [Warning] Aborted connection 29 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:43:44 62 [Warning] Aborted connection 62 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:43:44 61 [Warning] Aborted connection 61 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:43:44 63 [Warning] Aborted connection 63 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:44:31 75 [Warning] Aborted connection 75 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:44:31 66 [Warning] Aborted connection 66 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:45:42 91 [Warning] Aborted connection 91 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:45:42 90 [Warning] Aborted connection 90 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:46:26 115 [Warning] Aborted connection 115 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:46:26 100 [Warning] Aborted connection 100 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:48:00 118 [Warning] Aborted connection 118 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:48:00 123 [Warning] Aborted connection 123 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:48:36 137 [Warning] Aborted connection 137 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:48:36 128 [Warning] Aborted connection 128 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:49:34 153 [Warning] Aborted connection 153 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:49:34 148 [Warning] Aborted connection 148 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:50:01 171 [Warning] Aborted connection 171 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:50:01 165 [Warning] Aborted connection 165 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:50:01 164 [Warning] Aborted connection 164 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:50:01 168 [Warning] Aborted connection 168 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/abort-states.json b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/abort-states.json new file mode 100644 index 00000000..498e0098 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/abort-states.json @@ -0,0 +1,20 @@ +{ + "romm": { + "status": "restarting", + "health": "unhealthy", + "restarts": 0, + "exit": 1 + }, + "romm-db": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + }, + "romm-redis": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + } +} \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/compose-final.log b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/compose-final.log new file mode 100644 index 00000000..8bf306c8 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/compose-final.log @@ -0,0 +1,288 @@ +romm-db | 2026-09-23 08:37:58+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:11.4.13+maria~ubu2404 started. +romm-db | 2026-09-23 08:37:58+02:00 [Warn] [Entrypoint]: /sys/fs/cgroup///memory.pressure not writable, functionality unavailable to MariaDB +romm-db | 2026-09-23 08:37:58+02:00 [Note] [Entrypoint]: Switching to dedicated user 'mysql' +romm-db | 2026-09-23 08:37:58+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:11.4.13+maria~ubu2404 started. +romm-db | 2026-09-23 08:37:58+02:00 [Note] [Entrypoint]: Initializing database files +romm-db | 2026-09-23 8:37:58 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 08:38:01+02:00 [Note] [Entrypoint]: Database files initialized +romm-db | 2026-09-23 08:38:01+02:00 [Note] [Entrypoint]: Starting temporary server +romm-db | 2026-09-23 08:38:01+02:00 [Note] [Entrypoint]: Waiting for server startup +romm-db | 2026-09-23 8:38:01 0 [Note] Starting MariaDB 11.4.13-MariaDB-ubu2404 source revision 170b1d70737be6f134448f51713cdc1ae215b420 server_uid CoZjczEYgGSu3N2tEniVTdlmtSg= as process 90 +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Compressed tables use zlib 1.3 +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Number of transaction pools: 1 +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Using crc32 + pclmulqdq instructions +romm-db | 2026-09-23 8:38:01 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Using Linux native AIO +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: innodb_buffer_pool_size_max=8388608m, innodb_buffer_pool_size=128m +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Completed initialization of buffer pool +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: File system buffers for log disabled (block size=512 bytes) +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: End of log at LSN=45568 +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Opened 3 undo tablespaces +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: 128 rollback segments in 3 undo tablespaces are active. +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Setting file './ibtmp1' size to 12.000MiB. Physically writing the file full; Please wait ... +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: File './ibtmp1' size is now 12.000MiB. +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: log sequence number 45568; transaction id 14 +romm-db | 2026-09-23 8:38:01 0 [Note] Plugin 'FEEDBACK' is disabled. +romm-db | 2026-09-23 8:38:01 0 [Note] Plugin 'wsrep-provider' is disabled. +romm-db | 2026-09-23 8:38:02 0 [Note] mariadbd: Event Scheduler: Loaded 0 events +romm-db | 2026-09-23 8:38:02 0 [Note] mariadbd: ready for connections. +romm-db | Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 0 mariadb.org binary distribution +romm-db | 2026-09-23 08:38:03+02:00 [Note] [Entrypoint]: Temporary server started. +romm-db | 2026-09-23 08:38:03+02:00 [Note] [Entrypoint]: Creating database romm +romm-db | 2026-09-23 08:38:03+02:00 [Note] [Entrypoint]: Creating user romm +romm-db | 2026-09-23 08:38:03+02:00 [Note] [Entrypoint]: Giving user romm access to schema romm +romm-db | 2026-09-23 08:38:03+02:00 [Note] [Entrypoint]: Securing system users (equivalent to running mysql_secure_installation) +romm-db | +romm-db | 2026-09-23 08:38:03+02:00 [Note] [Entrypoint]: Stopping temporary server +romm-db | 2026-09-23 8:38:03 0 [Note] mariadbd (initiated by: unknown): Normal shutdown +romm-db | 2026-09-23 8:38:03 0 [Note] InnoDB: FTS optimize thread exiting. +romm-db | 2026-09-23 8:38:03 0 [Note] InnoDB: Starting shutdown... +romm-db | 2026-09-23 8:38:03 0 [Note] InnoDB: Dumping buffer pool(s) to /var/lib/mysql/ib_buffer_pool +romm-db | 2026-09-23 8:38:03 0 [Note] InnoDB: Buffer pool(s) dump completed at 260923 8:38:03 +romm-db | 2026-09-23 8:38:03 0 [Note] InnoDB: Removed temporary tablespace data file: "./ibtmp1" +romm-db | 2026-09-23 8:38:03 0 [Note] Shutdown completed; log sequence number 45568; transaction id 15 +romm-db | 2026-09-23 8:38:04 0 [Note] mariadbd: Shutdown complete +romm-db | 2026-09-23 08:38:04+02:00 [Note] [Entrypoint]: Temporary server stopped +romm-db | +romm-db | 2026-09-23 08:38:04+02:00 [Note] [Entrypoint]: MariaDB init process done. Ready for start up. +romm-db | +romm-db | 2026-09-23 8:38:04 0 [Note] Starting MariaDB 11.4.13-MariaDB-ubu2404 source revision 170b1d70737be6f134448f51713cdc1ae215b420 server_uid CoZjczEYgGSu3N2tEniVTdlmtSg= as process 1 +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Compressed tables use zlib 1.3 +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Number of transaction pools: 1 +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Using crc32 + pclmulqdq instructions +romm-db | 2026-09-23 8:38:04 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Using Linux native AIO +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: innodb_buffer_pool_size_max=8388608m, innodb_buffer_pool_size=128m +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Completed initialization of buffer pool +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: File system buffers for log disabled (block size=512 bytes) +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: End of log at LSN=45568 +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Opened 3 undo tablespaces +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: 128 rollback segments in 3 undo tablespaces are active. +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Setting file './ibtmp1' size to 12.000MiB. Physically writing the file full; Please wait ... +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: File './ibtmp1' size is now 12.000MiB. +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: log sequence number 45568; transaction id 14 +romm-db | 2026-09-23 8:38:04 0 [Note] Plugin 'FEEDBACK' is disabled. +romm-db | 2026-09-23 8:38:04 0 [Note] Plugin 'wsrep-provider' is disabled. +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Loading buffer pool(s) from /var/lib/mysql/ib_buffer_pool +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Buffer pool(s) load completed at 260923 8:38:04 +romm-db | 2026-09-23 8:38:05 0 [Note] Server socket created on IP: '0.0.0.0', port: '3306'. +romm-db | 2026-09-23 8:38:05 0 [Note] Server socket created on IP: '::', port: '3306'. +romm-db | 2026-09-23 8:38:05 0 [Note] mariadbd: Event Scheduler: Loaded 0 events +romm-db | 2026-09-23 8:38:05 0 [Note] mariadbd: ready for connections. +romm-db | Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution +romm-db | 2026-09-23 8:38:27 5 [Warning] Aborted connection 5 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:38:56 16 [Warning] Aborted connection 16 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:38:56 15 [Warning] Aborted connection 15 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:39:12 19 [Warning] Aborted connection 19 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:41:46 35 [Warning] Aborted connection 35 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:41:46 30 [Warning] Aborted connection 30 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:41:54 44 [Warning] Aborted connection 44 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:41:54 29 [Warning] Aborted connection 29 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:43:44 62 [Warning] Aborted connection 62 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:43:44 61 [Warning] Aborted connection 61 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:43:44 63 [Warning] Aborted connection 63 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:44:31 75 [Warning] Aborted connection 75 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:44:31 66 [Warning] Aborted connection 66 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:45:42 91 [Warning] Aborted connection 91 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:45:42 90 [Warning] Aborted connection 90 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:46:26 115 [Warning] Aborted connection 115 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:46:26 100 [Warning] Aborted connection 100 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:48:00 118 [Warning] Aborted connection 118 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm | INFO: [RomM][init][2026-09-23 08:50:03] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:50:03] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:50:03] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:50:03] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:50:03] | | \ \ (_) | | | | | | | | | +romm-redis | 1:C 23 Sep 2026 08:37:58.125 # WARNING Memory overcommit must be enabled! Without it, a background save or replication may fail under low memory condition. Being disabled, it can also cause failures without low memory condition, see https://github.com/jemalloc/jemalloc/issues/1328. To fix this issue add 'vm.overcommit_memory = 1' to /etc/sysctl.conf and then reboot or run the command 'sysctl vm.overcommit_memory=1' for this to take effect. +romm | INFO: [RomM][init][2026-09-23 08:50:03] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:50:03] +romm-redis | 1:C 23 Sep 2026 08:37:58.125 * oO0OoO0OoO0Oo Redis is starting oO0OoO0OoO0Oo +romm | INFO: [RomM][init][2026-09-23 08:50:03] The beautiful, powerful, self-hosted Rom manager and player +romm-db | 2026-09-23 8:48:00 123 [Warning] Aborted connection 123 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm | INFO: [RomM][init][2026-09-23 08:50:03] +romm-db | 2026-09-23 8:48:36 137 [Warning] Aborted connection 137 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-redis | 1:C 23 Sep 2026 08:37:58.125 * Redis version=7.4.11, bits=64, commit=00000000, modified=0, pid=1, just started +romm-db | 2026-09-23 8:48:36 128 [Warning] Aborted connection 128 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:49:34 153 [Warning] Aborted connection 153 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-redis | 1:C 23 Sep 2026 08:37:58.125 * Configuration loaded +romm-db | 2026-09-23 8:49:34 148 [Warning] Aborted connection 148 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-redis | 1:M 23 Sep 2026 08:37:58.125 * Increased maximum number of open files to 10032 (it was originally set to 1024). +romm-db | 2026-09-23 8:50:01 171 [Warning] Aborted connection 171 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-redis | 1:M 23 Sep 2026 08:37:58.125 * monotonic clock: POSIX clock_gettime +romm-db | 2026-09-23 8:50:01 165 [Warning] Aborted connection 165 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm | INFO: [RomM][init][2026-09-23 08:50:03] Version: 5.0.0 +romm-db | 2026-09-23 8:50:01 164 [Warning] Aborted connection 164 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:50:01 168 [Warning] Aborted connection 168 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm | INFO: [RomM][init][2026-09-23 08:50:03] +romm | INFO: [RomM][init][2026-09-23 08:50:03] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:50:03] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:50:03] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:50:11] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:50:12] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:50:12] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:50:12] | |__) |___ _ __ ___ | \ / | +romm-redis | 1:M 23 Sep 2026 08:37:58.128 * Running mode=standalone, port=6379. +romm-redis | 1:M 23 Sep 2026 08:37:58.128 * Server initialized +romm-redis | 1:M 23 Sep 2026 08:37:58.136 * Creating AOF base file appendonly.aof.1.base.rdb on server start +romm-redis | 1:M 23 Sep 2026 08:37:58.145 * Creating AOF incr file appendonly.aof.1.incr.aof on server start +romm-redis | 1:M 23 Sep 2026 08:37:58.145 * Ready to accept connections tcp +romm-redis | 1:M 23 Sep 2026 08:38:59.091 * 10000 changes in 60 seconds. Saving... +romm-redis | 1:M 23 Sep 2026 08:38:59.093 * Background saving started by pid 52 +romm-redis | 52:C 23 Sep 2026 08:38:59.225 * DB saved on disk +romm-redis | 52:C 23 Sep 2026 08:38:59.226 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 1 MB +romm-redis | 1:M 23 Sep 2026 08:38:59.295 * Background saving terminated with success +romm-redis | 1:M 23 Sep 2026 08:44:00.006 * 100 changes in 300 seconds. Saving... +romm-redis | 1:M 23 Sep 2026 08:44:00.007 * Background saving started by pid 233 +romm | INFO: [RomM][init][2026-09-23 08:50:12] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:50:12] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:50:12] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:50:12] +romm | INFO: [RomM][init][2026-09-23 08:50:12] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:50:12] +romm | INFO: [RomM][init][2026-09-23 08:50:12] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:50:12] +romm | INFO: [RomM][init][2026-09-23 08:50:12] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:50:12] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:50:12] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:50:21] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:50:21] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:50:21] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:50:21] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:50:21] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:50:21] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:50:21] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:50:21] +romm | INFO: [RomM][init][2026-09-23 08:50:21] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:50:21] +romm | INFO: [RomM][init][2026-09-23 08:50:21] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:50:21] +romm | INFO: [RomM][init][2026-09-23 08:50:21] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:50:21] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:50:21] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:50:30] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:50:31] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:50:31] | __ \ | \/ | +romm-redis | 233:C 23 Sep 2026 08:44:00.119 * DB saved on disk +romm | INFO: [RomM][init][2026-09-23 08:50:31] | |__) |___ _ __ ___ | \ / | +romm-redis | 233:C 23 Sep 2026 08:44:00.120 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 1 MB +romm | INFO: [RomM][init][2026-09-23 08:50:31] | _ // _ \| '_ ` _ \| |\/| | +romm-redis | 1:M 23 Sep 2026 08:44:00.210 * Background saving terminated with success +romm | INFO: [RomM][init][2026-09-23 08:50:31] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:50:31] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:50:31] +romm | INFO: [RomM][init][2026-09-23 08:50:31] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:50:31] +romm | INFO: [RomM][init][2026-09-23 08:50:31] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:50:31] +romm | INFO: [RomM][init][2026-09-23 08:50:31] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm-redis | 1:M 23 Sep 2026 08:49:01.065 * 100 changes in 300 seconds. Saving... +romm | INFO: [RomM][init][2026-09-23 08:50:31] REDIS_HOST is set, not starting internal valkey-server +romm-redis | 1:M 23 Sep 2026 08:49:01.066 * Background saving started by pid 408 +romm | INFO: [RomM][init][2026-09-23 08:50:31] Running database migrations +romm-redis | 408:C 23 Sep 2026 08:49:01.191 * DB saved on disk +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm-redis | 408:C 23 Sep 2026 08:49:01.192 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 1 MB +romm | ERROR: [RomM][init][2026-09-23 08:50:39] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:50:40] _____ __ __ +romm-redis | 1:M 23 Sep 2026 08:49:01.268 * Background saving terminated with success +romm | INFO: [RomM][init][2026-09-23 08:50:40] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:50:40] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:50:40] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:50:40] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:50:40] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:50:40] +romm | INFO: [RomM][init][2026-09-23 08:50:40] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:50:40] +romm | INFO: [RomM][init][2026-09-23 08:50:40] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:50:40] +romm | INFO: [RomM][init][2026-09-23 08:50:40] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:50:40] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:50:40] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:50:49] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:50:51] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:50:51] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:50:51] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:50:51] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:50:51] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:50:51] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:50:51] +romm | INFO: [RomM][init][2026-09-23 08:50:51] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:50:51] +romm | INFO: [RomM][init][2026-09-23 08:50:51] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:50:51] +romm | INFO: [RomM][init][2026-09-23 08:50:51] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:50:51] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:50:51] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:51:00] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:51:03] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:51:03] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:51:03] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:51:03] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:51:03] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:51:03] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:51:03] +romm | INFO: [RomM][init][2026-09-23 08:51:03] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:51:03] +romm | INFO: [RomM][init][2026-09-23 08:51:03] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:51:03] +romm | INFO: [RomM][init][2026-09-23 08:51:03] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:51:03] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:51:03] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:51:12] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:51:18] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:51:18] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:51:18] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:51:18] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:51:18] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:51:18] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:51:18] +romm | INFO: [RomM][init][2026-09-23 08:51:18] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:51:18] +romm | INFO: [RomM][init][2026-09-23 08:51:18] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:51:18] +romm | INFO: [RomM][init][2026-09-23 08:51:18] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:51:18] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:51:18] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:51:27] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:51:40] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:51:40] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:51:40] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:51:40] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:51:40] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:51:40] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:51:40] +romm | INFO: [RomM][init][2026-09-23 08:51:40] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:51:40] +romm | INFO: [RomM][init][2026-09-23 08:51:40] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:51:40] +romm | INFO: [RomM][init][2026-09-23 08:51:40] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:51:40] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:51:40] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:51:49] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:52:15] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:52:15] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:52:15] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:52:15] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:52:15] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:52:15] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:52:15] +romm | INFO: [RomM][init][2026-09-23 08:52:15] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:52:15] +romm | INFO: [RomM][init][2026-09-23 08:52:15] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:52:15] +romm | INFO: [RomM][init][2026-09-23 08:52:15] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:52:15] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:52:15] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:52:23] Failed to run database migrations diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/engine-state.json b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/engine-state.json new file mode 100644 index 00000000..ec747fa4 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/engine-state.json @@ -0,0 +1 @@ +null \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/memory-samples.json b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/memory-samples.json new file mode 100644 index 00000000..2ee6ffa3 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/memory-samples.json @@ -0,0 +1,1442 @@ +[ + { + "t": 15.2, + "containers": { + "romm": { + "limit": 805306368, + "current": 644911104, + "peak": 650174464, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 147722240, + "peak": 153554944, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16240640, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 291 + }, + { + "t": 30.4, + "containers": { + "romm": { + "limit": 805306368, + "current": 645398528, + "peak": 650174464, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148389888, + "peak": 153554944, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15970304, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 584 + }, + { + "t": 45.6, + "containers": { + "romm": { + "limit": 805306368, + "current": 645423104, + "peak": 650731520, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 147873792, + "peak": 155402240, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15966208, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 877 + }, + { + "t": 60.8, + "containers": { + "romm": { + "limit": 805306368, + "current": 646082560, + "peak": 650731520, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148119552, + "peak": 155402240, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16453632, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 1169 + }, + { + "t": 76.0, + "containers": { + "romm": { + "limit": 805306368, + "current": 646246400, + "peak": 650809344, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148426752, + "peak": 155402240, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16089088, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 1459 + }, + { + "t": 91.2, + "containers": { + "romm": { + "limit": 805306368, + "current": 645873664, + "peak": 650809344, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 147931136, + "peak": 155402240, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15867904, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 1751 + }, + { + "t": 106.4, + "containers": { + "romm": { + "limit": 805306368, + "current": 646631424, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 147947520, + "peak": 155758592, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15872000, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 2042 + }, + { + "t": 121.7, + "containers": { + "romm": { + "limit": 805306368, + "current": 471539712, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148180992, + "peak": 155758592, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15839232, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 2334 + }, + { + "t": 136.9, + "containers": { + "romm": { + "limit": 805306368, + "current": 562036736, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 147943424, + "peak": 155758592, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15568896, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 2431 + }, + { + "t": 152.1, + "containers": { + "romm": { + "limit": 805306368, + "current": 644702208, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148242432, + "peak": 155758592, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15769600, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 2719 + }, + { + "t": 167.3, + "containers": { + "romm": { + "limit": 805306368, + "current": 643891200, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 147968000, + "peak": 155758592, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15491072, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 3012 + }, + { + "t": 182.5, + "containers": { + "romm": { + "limit": 805306368, + "current": 644644864, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148213760, + "peak": 155758592, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16015360, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 3302 + }, + { + "t": 197.7, + "containers": { + "romm": { + "limit": 805306368, + "current": 644153344, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 149135360, + "peak": 155758592, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15998976, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 3594 + }, + { + "t": 212.9, + "containers": { + "romm": { + "limit": 805306368, + "current": 645292032, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148365312, + "peak": 155758592, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16257024, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 3885 + }, + { + "t": 228.1, + "containers": { + "romm": { + "limit": 805306368, + "current": 644964352, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148373504, + "peak": 155758592, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15486976, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 4175 + }, + { + "t": 243.3, + "containers": { + "romm": { + "limit": 805306368, + "current": 522743808, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148602880, + "peak": 155758592, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15740928, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 4466 + }, + { + "t": 258.5, + "containers": { + "romm": { + "limit": 805306368, + "current": 645410816, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148410368, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16048128, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 4753 + }, + { + "t": 273.7, + "containers": { + "romm": { + "limit": 805306368, + "current": 644169728, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148455424, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16060416, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 5046 + }, + { + "t": 288.9, + "containers": { + "romm": { + "limit": 805306368, + "current": 502222848, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148459520, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16568320, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 5337 + }, + { + "t": 304.1, + "containers": { + "romm": { + "limit": 805306368, + "current": 644100096, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148205568, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15810560, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 5626 + }, + { + "t": 319.3, + "containers": { + "romm": { + "limit": 805306368, + "current": 644923392, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148480000, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16056320, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 5918 + }, + { + "t": 334.5, + "containers": { + "romm": { + "limit": 805306368, + "current": 644542464, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 149020672, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15822848, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 6210 + }, + { + "t": 349.7, + "containers": { + "romm": { + "limit": 805306368, + "current": 644681728, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148238336, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15794176, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 6500 + }, + { + "t": 364.9, + "containers": { + "romm": { + "limit": 805306368, + "current": 562343936, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148451328, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16056320, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 6790 + }, + { + "t": 380.1, + "containers": { + "romm": { + "limit": 805306368, + "current": 644440064, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148656128, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16834560, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 7079 + }, + { + "t": 395.3, + "containers": { + "romm": { + "limit": 805306368, + "current": 644263936, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148471808, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15839232, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 7369 + }, + { + "t": 410.5, + "containers": { + "romm": { + "limit": 805306368, + "current": 582758400, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148258816, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15818752, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 7653 + }, + { + "t": 425.7, + "containers": { + "romm": { + "limit": 805306368, + "current": 644472832, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148733952, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16105472, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 7942 + }, + { + "t": 440.9, + "containers": { + "romm": { + "limit": 805306368, + "current": 645042176, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148516864, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16330752, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 8235 + }, + { + "t": 456.1, + "containers": { + "romm": { + "limit": 805306368, + "current": 644759552, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 149245952, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15740928, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 8527 + }, + { + "t": 471.3, + "containers": { + "romm": { + "limit": 805306368, + "current": 644665344, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148963328, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16326656, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 8819 + }, + { + "t": 486.5, + "containers": { + "romm": { + "limit": 805306368, + "current": 644902912, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 149028864, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16171008, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 9111 + }, + { + "t": 501.7, + "containers": { + "romm": { + "limit": 805306368, + "current": 552001536, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148541440, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15814656, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 9401 + }, + { + "t": 516.9, + "containers": { + "romm": { + "limit": 805306368, + "current": 643719168, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 149135360, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16080896, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 9690 + }, + { + "t": 532.1, + "containers": { + "romm": { + "limit": 805306368, + "current": 479137792, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148455424, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16330752, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 9983 + }, + { + "t": 547.3, + "containers": { + "romm": { + "limit": 805306368, + "current": 637800448, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148414464, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 15560704, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 10273 + }, + { + "t": 562.6, + "containers": { + "romm": { + "limit": 805306368, + "current": 637820928, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148226048, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 18468864, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 10563 + }, + { + "t": 577.8, + "containers": { + "romm": { + "limit": 805306368, + "current": 638476288, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148525056, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 18460672, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 10855 + }, + { + "t": 593.0, + "containers": { + "romm": { + "limit": 805306368, + "current": 514371584, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148238336, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 17948672, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 11139 + }, + { + "t": 608.2, + "containers": { + "romm": { + "limit": 805306368, + "current": 644366336, + "peak": 651239424, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 148221952, + "peak": 155947008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 17686528, + "peak": 39243776, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 11429 + } +] \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/migration-lines.txt b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/migration-lines.txt new file mode 100644 index 00000000..d96fc1bd --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/migration-lines.txt @@ -0,0 +1,2 @@ +romm | INFO: [RomM][init][2026-09-23 08:38:58] Running database migrations +romm | INFO: [RomM][init][2026-09-23 08:39:12] Database migrations succeeded \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/run.log b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/run.log new file mode 100644 index 00000000..1ed57226 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/run.log @@ -0,0 +1,55 @@ +[06:37:57] M1: deploying romm at FROM {'romm': 'rommapp/romm:5.0.0'} +[06:38:45] FROM settled=True in 36.7s :: {"romm": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-db": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-redis": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}} +[06:38:53] romm: POST /api/users http=201 +[06:38:55] romm: login as the seeded user http=200 ok=True +[06:38:55] C1 (seed reads back BEFORE): True +[06:38:55] M1: swapping to TO {'romm': 'rommapp/romm:5.3.0'} +[06:38:58] TO up -d rc=0 +[06:39:29] TO settled=True in 31.6s :: {"romm": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-db": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-redis": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}} +[06:39:29] migration lines observed: 2 +[06:39:46] romm: login as the seeded user http=200 ok=True +[06:39:46] RESULT (seed reads back AFTER): True +[06:39:46] memory watch: 600s, 4 callers on 8 path(s) at 172.18.0.6:8080 +[06:40:01] + 15s romm=615M/768M peak=620M kills=0 rs=0 romm-db=140M/384M peak=146M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=291 +[06:40:16] + 30s romm=615M/768M peak=620M kills=0 rs=0 romm-db=141M/384M peak=146M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=584 +[06:40:31] + 46s romm=615M/768M peak=620M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=877 +[06:40:47] + 61s romm=616M/768M peak=620M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=1169 +[06:41:02] + 76s romm=616M/768M peak=620M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=1459 +[06:41:17] + 91s romm=615M/768M peak=620M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=1751 +[06:41:32] + 106s romm=616M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=2042 +[06:41:47] + 122s romm=449M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=2334 +[06:42:03] + 137s romm=536M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=14M/128M peak=37M kills=0 rs=0 reqs=2431 +[06:42:18] + 152s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=2719 +[06:42:33] + 167s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=14M/128M peak=37M kills=0 rs=0 reqs=3012 +[06:42:48] + 182s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=3302 +[06:43:03] + 198s romm=614M/768M peak=621M kills=0 rs=0 romm-db=142M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=3594 +[06:43:19] + 213s romm=615M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=3885 +[06:43:34] + 228s romm=615M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=14M/128M peak=37M kills=0 rs=0 reqs=4175 +[06:43:49] + 243s romm=498M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=4466 +[06:44:04] + 258s romm=615M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=4753 +[06:44:19] + 274s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=5046 +[06:44:35] + 289s romm=478M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=5337 +[06:44:50] + 304s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=5626 +[06:45:05] + 319s romm=615M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=5918 +[06:45:20] + 334s romm=614M/768M peak=621M kills=0 rs=0 romm-db=142M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=6210 +[06:45:35] + 350s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=6500 +[06:45:51] + 365s romm=536M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=6790 +[06:46:06] + 380s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=16M/128M peak=37M kills=0 rs=0 reqs=7079 +[06:46:21] + 395s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=7369 +[06:46:36] + 410s romm=555M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=7653 +[06:46:51] + 426s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=7942 +[06:47:07] + 441s romm=615M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=8235 +[06:47:22] + 456s romm=614M/768M peak=621M kills=0 rs=0 romm-db=142M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=8527 +[06:47:37] + 471s romm=614M/768M peak=621M kills=0 rs=0 romm-db=142M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=8819 +[06:47:52] + 486s romm=615M/768M peak=621M kills=0 rs=0 romm-db=142M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=9111 +[06:48:07] + 502s romm=526M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=9401 +[06:48:23] + 517s romm=613M/768M peak=621M kills=0 rs=0 romm-db=142M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=9690 +[06:48:38] + 532s romm=456M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=15M/128M peak=37M kills=0 rs=0 reqs=9983 +[06:48:53] + 547s romm=608M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=14M/128M peak=37M kills=0 rs=0 reqs=10273 +[06:49:08] + 563s romm=608M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=17M/128M peak=37M kills=0 rs=0 reqs=10563 +[06:49:24] + 578s romm=608M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=17M/128M peak=37M kills=0 rs=0 reqs=10855 +[06:49:39] + 593s romm=490M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=17M/128M peak=37M kills=0 rs=0 reqs=11139 +[06:49:54] + 608s romm=614M/768M peak=621M kills=0 rs=0 romm-db=141M/384M peak=148M kills=0 rs=0 romm-redis=16M/128M peak=37M kills=0 rs=0 reqs=11429 +[06:49:54] memory watch: killed=False tight=['romm'] requests=11429 codes={'200': 5712, '401': 5717} +[06:49:54] M1: ABORT — putting the FROM images back +[06:53:05] ABORT: the app did NOT come back (rc=0, 182.4s) \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/to-full.log b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/to-full.log new file mode 100644 index 00000000..8c83aec4 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/to-full.log @@ -0,0 +1,134 @@ +romm-db | 2026-09-23 08:37:58+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:11.4.13+maria~ubu2404 started. +romm-db | 2026-09-23 08:37:58+02:00 [Warn] [Entrypoint]: /sys/fs/cgroup///memory.pressure not writable, functionality unavailable to MariaDB +romm-db | 2026-09-23 08:37:58+02:00 [Note] [Entrypoint]: Switching to dedicated user 'mysql' +romm-db | 2026-09-23 08:37:58+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:11.4.13+maria~ubu2404 started. +romm-db | 2026-09-23 08:37:58+02:00 [Note] [Entrypoint]: Initializing database files +romm-db | 2026-09-23 8:37:58 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 08:38:01+02:00 [Note] [Entrypoint]: Database files initialized +romm-db | 2026-09-23 08:38:01+02:00 [Note] [Entrypoint]: Starting temporary server +romm-db | 2026-09-23 08:38:01+02:00 [Note] [Entrypoint]: Waiting for server startup +romm-db | 2026-09-23 8:38:01 0 [Note] Starting MariaDB 11.4.13-MariaDB-ubu2404 source revision 170b1d70737be6f134448f51713cdc1ae215b420 server_uid CoZjczEYgGSu3N2tEniVTdlmtSg= as process 90 +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Compressed tables use zlib 1.3 +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Number of transaction pools: 1 +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Using crc32 + pclmulqdq instructions +romm-db | 2026-09-23 8:38:01 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Using Linux native AIO +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: innodb_buffer_pool_size_max=8388608m, innodb_buffer_pool_size=128m +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Completed initialization of buffer pool +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: File system buffers for log disabled (block size=512 bytes) +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: End of log at LSN=45568 +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Opened 3 undo tablespaces +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: 128 rollback segments in 3 undo tablespaces are active. +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: Setting file './ibtmp1' size to 12.000MiB. Physically writing the file full; Please wait ... +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: File './ibtmp1' size is now 12.000MiB. +romm-db | 2026-09-23 8:38:01 0 [Note] InnoDB: log sequence number 45568; transaction id 14 +romm-db | 2026-09-23 8:38:01 0 [Note] Plugin 'FEEDBACK' is disabled. +romm-db | 2026-09-23 8:38:01 0 [Note] Plugin 'wsrep-provider' is disabled. +romm-db | 2026-09-23 8:38:02 0 [Note] mariadbd: Event Scheduler: Loaded 0 events +romm-db | 2026-09-23 8:38:02 0 [Note] mariadbd: ready for connections. +romm-db | Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 0 mariadb.org binary distribution +romm-db | 2026-09-23 08:38:03+02:00 [Note] [Entrypoint]: Temporary server started. +romm-db | 2026-09-23 08:38:03+02:00 [Note] [Entrypoint]: Creating database romm +romm-db | 2026-09-23 08:38:03+02:00 [Note] [Entrypoint]: Creating user romm +romm-db | 2026-09-23 08:38:03+02:00 [Note] [Entrypoint]: Giving user romm access to schema romm +romm-db | 2026-09-23 08:38:03+02:00 [Note] [Entrypoint]: Securing system users (equivalent to running mysql_secure_installation) +romm-db | +romm-db | 2026-09-23 08:38:03+02:00 [Note] [Entrypoint]: Stopping temporary server +romm-db | 2026-09-23 8:38:03 0 [Note] mariadbd (initiated by: unknown): Normal shutdown +romm-db | 2026-09-23 8:38:03 0 [Note] InnoDB: FTS optimize thread exiting. +romm-db | 2026-09-23 8:38:03 0 [Note] InnoDB: Starting shutdown... +romm-db | 2026-09-23 8:38:03 0 [Note] InnoDB: Dumping buffer pool(s) to /var/lib/mysql/ib_buffer_pool +romm-db | 2026-09-23 8:38:03 0 [Note] InnoDB: Buffer pool(s) dump completed at 260923 8:38:03 +romm-db | 2026-09-23 8:38:03 0 [Note] InnoDB: Removed temporary tablespace data file: "./ibtmp1" +romm-db | 2026-09-23 8:38:03 0 [Note] Shutdown completed; log sequence number 45568; transaction id 15 +romm-db | 2026-09-23 8:38:04 0 [Note] mariadbd: Shutdown complete +romm-db | 2026-09-23 08:38:04+02:00 [Note] [Entrypoint]: Temporary server stopped +romm-db | +romm-db | 2026-09-23 08:38:04+02:00 [Note] [Entrypoint]: MariaDB init process done. Ready for start up. +romm-db | +romm-db | 2026-09-23 8:38:04 0 [Note] Starting MariaDB 11.4.13-MariaDB-ubu2404 source revision 170b1d70737be6f134448f51713cdc1ae215b420 server_uid CoZjczEYgGSu3N2tEniVTdlmtSg= as process 1 +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Compressed tables use zlib 1.3 +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Number of transaction pools: 1 +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Using crc32 + pclmulqdq instructions +romm-db | 2026-09-23 8:38:04 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Using Linux native AIO +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: innodb_buffer_pool_size_max=8388608m, innodb_buffer_pool_size=128m +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Completed initialization of buffer pool +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: File system buffers for log disabled (block size=512 bytes) +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: End of log at LSN=45568 +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Opened 3 undo tablespaces +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: 128 rollback segments in 3 undo tablespaces are active. +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Setting file './ibtmp1' size to 12.000MiB. Physically writing the file full; Please wait ... +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: File './ibtmp1' size is now 12.000MiB. +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: log sequence number 45568; transaction id 14 +romm-db | 2026-09-23 8:38:04 0 [Note] Plugin 'FEEDBACK' is disabled. +romm-db | 2026-09-23 8:38:04 0 [Note] Plugin 'wsrep-provider' is disabled. +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Loading buffer pool(s) from /var/lib/mysql/ib_buffer_pool +romm-db | 2026-09-23 8:38:04 0 [Note] InnoDB: Buffer pool(s) load completed at 260923 8:38:04 +romm-db | 2026-09-23 8:38:05 0 [Note] Server socket created on IP: '0.0.0.0', port: '3306'. +romm-db | 2026-09-23 8:38:05 0 [Note] Server socket created on IP: '::', port: '3306'. +romm-db | 2026-09-23 8:38:05 0 [Note] mariadbd: Event Scheduler: Loaded 0 events +romm-db | 2026-09-23 8:38:05 0 [Note] mariadbd: ready for connections. +romm-db | Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution +romm-db | 2026-09-23 8:38:27 5 [Warning] Aborted connection 5 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:38:56 16 [Warning] Aborted connection 16 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:38:56 15 [Warning] Aborted connection 15 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:39:12 19 [Warning] Aborted connection 19 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm | INFO: [RomM][init][2026-09-23 08:38:58] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:38:58] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:38:58] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:38:58] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:38:58] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:38:58] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:38:58] +romm | INFO: [RomM][init][2026-09-23 08:38:58] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:38:58] +romm | INFO: [RomM][init][2026-09-23 08:38:58] Version: 5.3.0 +romm | INFO: [RomM][init][2026-09-23 08:38:58] +romm | INFO: [RomM][init][2026-09-23 08:38:58] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:38:58] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:38:58] Running database migrations +romm | INFO: [RomM][init][2026-09-23 08:39:12] Database migrations succeeded +romm | INFO: [RomM][startup][2026-09-23 08:39:24] Running startup tasks +romm | INFO: [RomM][startup][2026-09-23 08:39:24] Cleared 2 job(s) left behind by the old scheduler +romm | INFO: [RomM][startup][2026-09-23 08:39:24] Initializing cache with fixtures data +romm | INFO: [RomM][startup][2026-09-23 08:39:24] Startup tasks completed +romm | INFO: [RomM][init][2026-09-23 08:39:25] Starting backend +romm | INFO: [RomM][init][2026-09-23 08:39:25] Starting RQ cron scheduler +romm | INFO: [RomM][init][2026-09-23 08:39:25] Starting RQ worker +romm | INFO: [RomM][init][2026-09-23 08:39:25] Starting RQ scan worker +romm | INFO: [RomM][init][2026-09-23 08:39:28] Starting nginx +romm | 2026/09/23 08:39:28 [notice] 123#123: js vm init njs: 0000759641C76B00 +romm | INFO: [RomM][init][2026-09-23 08:39:28] 🚀 RomM is now available at http://0.0.0.0:8080 +romm | 08:39:28 Loading cron configuration from tasks.cron_config +romm | 08:39:28 Worker 8819d4ef33764a81a0fb0ff40cdca8d3: started with PID 104, version 2.12.0 +romm | 08:39:28 Worker 8819d4ef33764a81a0fb0ff40cdca8d3: subscribing to channel rq:pubsub:8819d4ef33764a81a0fb0ff40cdca8d3 +romm | 08:39:28 *** Listening on high, default, low... +romm | 08:39:28 Acquired scheduler lock for high +romm | 08:39:28 Acquired scheduler lock for default +romm | 08:39:28 Acquired scheduler lock for low +romm | 08:39:28 Scheduler for high, default, low started with PID 142 +romm-redis | 1:C 23 Sep 2026 08:37:58.125 # WARNING Memory overcommit must be enabled! Without it, a background save or replication may fail under low memory condition. Being disabled, it can also cause failures without low memory condition, see https://github.com/jemalloc/jemalloc/issues/1328. To fix this issue add 'vm.overcommit_memory = 1' to /etc/sysctl.conf and then reboot or run the command 'sysctl vm.overcommit_memory=1' for this to take effect. +romm | 08:39:28 Worker e80205cfb5aa4cd2b78fd01218b778cf: started with PID 107, version 2.12.0 +romm-redis | 1:C 23 Sep 2026 08:37:58.125 * oO0OoO0OoO0Oo Redis is starting oO0OoO0OoO0Oo +romm-redis | 1:C 23 Sep 2026 08:37:58.125 * Redis version=7.4.11, bits=64, commit=00000000, modified=0, pid=1, just started +romm-redis | 1:C 23 Sep 2026 08:37:58.125 * Configuration loaded +romm-redis | 1:M 23 Sep 2026 08:37:58.125 * Increased maximum number of open files to 10032 (it was originally set to 1024). +romm-redis | 1:M 23 Sep 2026 08:37:58.125 * monotonic clock: POSIX clock_gettime +romm-redis | 1:M 23 Sep 2026 08:37:58.128 * Running mode=standalone, port=6379. +romm-redis | 1:M 23 Sep 2026 08:37:58.128 * Server initialized +romm-redis | 1:M 23 Sep 2026 08:37:58.136 * Creating AOF base file appendonly.aof.1.base.rdb on server start +romm-redis | 1:M 23 Sep 2026 08:37:58.145 * Creating AOF incr file appendonly.aof.1.incr.aof on server start +romm-redis | 1:M 23 Sep 2026 08:37:58.145 * Ready to accept connections tcp +romm-redis | 1:M 23 Sep 2026 08:38:59.091 * 10000 changes in 60 seconds. Saving... +romm-redis | 1:M 23 Sep 2026 08:38:59.093 * Background saving started by pid 52 +romm-redis | 52:C 23 Sep 2026 08:38:59.225 * DB saved on disk +romm-redis | 52:C 23 Sep 2026 08:38:59.226 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 1 MB +romm-redis | 1:M 23 Sep 2026 08:38:59.295 * Background saving terminated with success +romm | 08:39:28 Worker e80205cfb5aa4cd2b78fd01218b778cf: subscribing to channel rq:pubsub:e80205cfb5aa4cd2b78fd01218b778cf +romm | 08:39:28 *** Listening on scans... +romm | 08:39:28 Acquired scheduler lock for scans +romm | 08:39:28 Scheduler for scans started with PID 145 diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/to-states.json b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/to-states.json new file mode 100644 index 00000000..1c7ad7de --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/to-states.json @@ -0,0 +1,20 @@ +{ + "romm": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + }, + "romm-db": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + }, + "romm-redis": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + } +} \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/verdict.json b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/verdict.json new file mode 100644 index 00000000..c92e02ee --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1/verdict.json @@ -0,0 +1,67 @@ +{ + "harness_version": 2, + "edge": "M1", + "app": "romm", + "note": "catalog move 15f9ebf 5.0.0 -> 5.3.0 on the CURRENT template (768M, 2 workers)", + "from": { + "romm": "rommapp/romm:5.0.0" + }, + "to": { + "romm": "rommapp/romm:5.3.0" + }, + "verdict": "proven", + "seed_read_before": true, + "seed_read_after": true, + "healthy_after": true, + "migration_observed": "romm | \u001b[0;32mINFO: \u001b[0;34m[RomM]\u001b[0;95m[init]\u001b[0;36m[2026-09-23 08:38:58]\u001b[0;00m Running database migrations", + "abort": "refuses", + "abort_detail": "to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:49:34 148 [Warning] Aborted connection 148 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:50:01 171 [Warning] Aborted connection 171 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:50:01 165 [Warning] Aborted connection 165 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:50:01 164 [Warning] Aborted connection 164 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:50:01 168 [Warning] Aborted connection 168 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)", + "engine_state_after": null, + "memory": { + "soak_s": 608.5, + "requested_s": 600, + "requests": 11429, + "codes": { + "200": 5712, + "401": 5717 + }, + "first_kill": null, + "containers": { + "romm": { + "limit": 805306368, + "peak": 651239424, + "peak_pct": 0.809, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + }, + "romm-db": { + "limit": 402653184, + "peak": 155947008, + "peak_pct": 0.387, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + }, + "romm-redis": { + "limit": 134217728, + "peak": 39243776, + "peak_pct": 0.292, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + } + }, + "unmeasured": [] + }, + "marks": [ + "memory_tight" + ], + "duration_s": 31.6, + "measured_at": "2026-09-23T06:53:05Z", + "evidence": "evidence/M1", + "total_s": 908.0 +} \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/abort-refusal.txt b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/abort-refusal.txt new file mode 100644 index 00000000..8a64b020 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/abort-refusal.txt @@ -0,0 +1,101 @@ +romm-redis | 1:C 23 Sep 2026 08:30:39.879 # WARNING Memory overcommit must be enabled! Without it, a background save or replication may fail under low memory condition. Being disabled, it can also cause failures without low memory condition, see https://github.com/jemalloc/jemalloc/issues/1328. To fix this issue add 'vm.overcommit_memory = 1' to /etc/sysctl.conf and then reboot or run the command 'sysctl vm.overcommit_memory=1' for this to take effect. +romm-db | 2026-09-23 8:30:44 0 [Note] Shutdown completed; log sequence number 45568; transaction id 15 +romm | INFO: [RomM][init][2026-09-23 08:35:28] +romm-redis | 1:C 23 Sep 2026 08:30:39.880 * oO0OoO0OoO0Oo Redis is starting oO0OoO0OoO0Oo +romm-redis | 1:C 23 Sep 2026 08:30:39.880 * Redis version=7.4.11, bits=64, commit=00000000, modified=0, pid=1, just started +romm-redis | 1:C 23 Sep 2026 08:30:39.880 * Configuration loaded +romm-redis | 1:M 23 Sep 2026 08:30:39.880 * Increased maximum number of open files to 10032 (it was originally set to 1024). +romm-redis | 1:M 23 Sep 2026 08:30:39.880 * monotonic clock: POSIX clock_gettime +romm-redis | 1:M 23 Sep 2026 08:30:39.881 * Running mode=standalone, port=6379. +romm-redis | 1:M 23 Sep 2026 08:30:39.882 * Server initialized +romm | INFO: [RomM][init][2026-09-23 08:35:28] Version: 5.0.0 +romm-redis | 1:M 23 Sep 2026 08:30:39.886 * Creating AOF base file appendonly.aof.1.base.rdb on server start +romm | INFO: [RomM][init][2026-09-23 08:35:28] +romm-redis | 1:M 23 Sep 2026 08:30:39.899 * Creating AOF incr file appendonly.aof.1.incr.aof on server start +romm | INFO: [RomM][init][2026-09-23 08:35:28] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm-redis | 1:M 23 Sep 2026 08:30:39.899 * Ready to accept connections tcp +romm | INFO: [RomM][init][2026-09-23 08:35:28] REDIS_HOST is set, not starting internal valkey-server +romm-db | 2026-09-23 8:30:44 0 [Note] mariadbd: Shutdown complete +romm | INFO: [RomM][init][2026-09-23 08:35:28] Running database migrations +romm-redis | 1:M 23 Sep 2026 08:31:40.061 * 10000 changes in 60 seconds. Saving... +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:35:37] Failed to run database migrations +romm-redis | 1:M 23 Sep 2026 08:31:40.062 * Background saving started by pid 46 +romm | INFO: [RomM][init][2026-09-23 08:35:50] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:35:50] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:35:50] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:35:50] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:35:50] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:35:50] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:35:50] +romm | INFO: [RomM][init][2026-09-23 08:35:50] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:35:50] +romm | INFO: [RomM][init][2026-09-23 08:35:50] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:35:50] +romm | INFO: [RomM][init][2026-09-23 08:35:50] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Temporary server stopped +romm | INFO: [RomM][init][2026-09-23 08:35:50] REDIS_HOST is set, not starting internal valkey-server +romm-db | +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: MariaDB init process done. Ready for start up. +romm-db | +romm-db | 2026-09-23 8:30:44 0 [Note] Starting MariaDB 11.4.13-MariaDB-ubu2404 source revision 170b1d70737be6f134448f51713cdc1ae215b420 server_uid Psh/IuH2qRBz9PeEZH//qqIAdts= as process 1 +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Compressed tables use zlib 1.3 +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Number of transaction pools: 1 +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Using crc32 + pclmulqdq instructions +romm-db | 2026-09-23 8:30:44 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-redis | 46:C 23 Sep 2026 08:31:40.211 * DB saved on disk +romm-db | create_uring failed: falling back to libaio +romm-redis | 46:C 23 Sep 2026 08:31:40.212 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 1 MB +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Using Linux native AIO +romm-redis | 1:M 23 Sep 2026 08:31:40.264 * Background saving terminated with success +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: innodb_buffer_pool_size_max=8388608m, innodb_buffer_pool_size=128m +romm | INFO: [RomM][init][2026-09-23 08:35:50] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:35:59] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:36:25] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:36:25] | __ \ | \/ | +romm-redis | 1:M 23 Sep 2026 08:36:41.033 * 100 changes in 300 seconds. Saving... +romm-redis | 1:M 23 Sep 2026 08:36:41.034 * Background saving started by pid 226 +romm-redis | 226:C 23 Sep 2026 08:36:41.158 * DB saved on disk +romm-redis | 226:C 23 Sep 2026 08:36:41.159 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 0 MB +romm-redis | 1:M 23 Sep 2026 08:36:41.236 * Background saving terminated with success +romm | INFO: [RomM][init][2026-09-23 08:36:25] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:36:25] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:36:25] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:36:25] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:36:25] +romm | INFO: [RomM][init][2026-09-23 08:36:25] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:36:25] +romm | INFO: [RomM][init][2026-09-23 08:36:25] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:36:25] +romm | INFO: [RomM][init][2026-09-23 08:36:25] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:36:25] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:36:25] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:36:34] Failed to run database migrations +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Completed initialization of buffer pool +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: File system buffers for log disabled (block size=512 bytes) +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: End of log at LSN=45568 +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Opened 3 undo tablespaces +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: 128 rollback segments in 3 undo tablespaces are active. +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Setting file './ibtmp1' size to 12.000MiB. Physically writing the file full; Please wait ... +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: File './ibtmp1' size is now 12.000MiB. +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: log sequence number 45568; transaction id 14 +romm-db | 2026-09-23 8:30:44 0 [Note] Plugin 'FEEDBACK' is disabled. +romm-db | 2026-09-23 8:30:44 0 [Note] Plugin 'wsrep-provider' is disabled. +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Loading buffer pool(s) from /var/lib/mysql/ib_buffer_pool +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Buffer pool(s) load completed at 260923 8:30:44 +romm-db | 2026-09-23 8:30:48 0 [Note] Server socket created on IP: '0.0.0.0', port: '3306'. +romm-db | 2026-09-23 8:30:48 0 [Note] Server socket created on IP: '::', port: '3306'. +romm-db | 2026-09-23 8:30:48 0 [Note] mariadbd: Event Scheduler: Loaded 0 events +romm-db | 2026-09-23 8:30:48 0 [Note] mariadbd: ready for connections. +romm-db | Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution +romm-db | 2026-09-23 8:31:08 5 [Warning] Aborted connection 5 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:31:38 15 [Warning] Aborted connection 15 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:31:54 18 [Warning] Aborted connection 18 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:33:51 33 [Warning] Aborted connection 33 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:33:51 41 [Warning] Aborted connection 41 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:33:51 32 [Warning] Aborted connection 32 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:34:11 36 [Warning] Aborted connection 36 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:34:11 42 [Warning] Aborted connection 42 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:34:11 49 [Warning] Aborted connection 49 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/abort-states.json b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/abort-states.json new file mode 100644 index 00000000..498e0098 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/abort-states.json @@ -0,0 +1,20 @@ +{ + "romm": { + "status": "restarting", + "health": "unhealthy", + "restarts": 0, + "exit": 1 + }, + "romm-db": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + }, + "romm-redis": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + } +} \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/compose-final.log b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/compose-final.log new file mode 100644 index 00000000..6a4f8697 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/compose-final.log @@ -0,0 +1,265 @@ +romm-redis | 1:C 23 Sep 2026 08:30:39.879 # WARNING Memory overcommit must be enabled! Without it, a background save or replication may fail under low memory condition. Being disabled, it can also cause failures without low memory condition, see https://github.com/jemalloc/jemalloc/issues/1328. To fix this issue add 'vm.overcommit_memory = 1' to /etc/sysctl.conf and then reboot or run the command 'sysctl vm.overcommit_memory=1' for this to take effect. +romm-redis | 1:C 23 Sep 2026 08:30:39.880 * oO0OoO0OoO0Oo Redis is starting oO0OoO0OoO0Oo +romm-redis | 1:C 23 Sep 2026 08:30:39.880 * Redis version=7.4.11, bits=64, commit=00000000, modified=0, pid=1, just started +romm | INFO: [RomM][init][2026-09-23 08:34:12] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:34:12] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:34:12] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:34:12] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:34:12] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:34:12] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:34:12] +romm | INFO: [RomM][init][2026-09-23 08:34:12] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:34:12] +romm | INFO: [RomM][init][2026-09-23 08:34:12] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:34:12] +romm | INFO: [RomM][init][2026-09-23 08:34:12] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:34:12] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:34:12] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:34:21] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:34:21] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:34:21] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:34:21] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:34:21] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:34:21] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:34:21] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:34:21] +romm | INFO: [RomM][init][2026-09-23 08:34:21] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:34:21] +romm | INFO: [RomM][init][2026-09-23 08:34:21] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:34:21] +romm | INFO: [RomM][init][2026-09-23 08:34:21] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:34:21] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:34:21] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:34:30] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:34:31] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:34:31] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:34:31] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:34:31] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:34:31] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:34:31] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:34:31] +romm | INFO: [RomM][init][2026-09-23 08:34:31] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:34:31] +romm | INFO: [RomM][init][2026-09-23 08:34:31] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:34:31] +romm | INFO: [RomM][init][2026-09-23 08:34:31] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:34:31] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:34:31] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:34:40] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:34:40] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:34:40] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:34:40] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:34:40] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:34:40] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:34:40] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:34:40] +romm | INFO: [RomM][init][2026-09-23 08:34:40] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:34:40] +romm | INFO: [RomM][init][2026-09-23 08:34:40] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:34:40] +romm | INFO: [RomM][init][2026-09-23 08:34:40] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:34:40] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:34:40] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:34:49] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:34:50] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:34:50] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:34:50] | |__) |___ _ __ ___ | \ / | +romm-db | 2026-09-23 08:30:40+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:11.4.13+maria~ubu2404 started. +romm | INFO: [RomM][init][2026-09-23 08:34:50] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:34:50] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:34:50] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:34:50] +romm-db | 2026-09-23 08:30:40+02:00 [Warn] [Entrypoint]: /sys/fs/cgroup///memory.pressure not writable, functionality unavailable to MariaDB +romm | INFO: [RomM][init][2026-09-23 08:34:50] The beautiful, powerful, self-hosted Rom manager and player +romm-db | 2026-09-23 08:30:40+02:00 [Note] [Entrypoint]: Switching to dedicated user 'mysql' +romm-db | 2026-09-23 08:30:40+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:11.4.13+maria~ubu2404 started. +romm-db | 2026-09-23 08:30:40+02:00 [Note] [Entrypoint]: Initializing database files +romm-db | 2026-09-23 8:30:40 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 08:30:42+02:00 [Note] [Entrypoint]: Database files initialized +romm-db | 2026-09-23 08:30:42+02:00 [Note] [Entrypoint]: Starting temporary server +romm-db | 2026-09-23 08:30:42+02:00 [Note] [Entrypoint]: Waiting for server startup +romm-db | 2026-09-23 8:30:42 0 [Note] Starting MariaDB 11.4.13-MariaDB-ubu2404 source revision 170b1d70737be6f134448f51713cdc1ae215b420 server_uid Psh/IuH2qRBz9PeEZH//qqIAdts= as process 92 +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Compressed tables use zlib 1.3 +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Number of transaction pools: 1 +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Using crc32 + pclmulqdq instructions +romm-db | 2026-09-23 8:30:42 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Using Linux native AIO +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: innodb_buffer_pool_size_max=8388608m, innodb_buffer_pool_size=128m +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Completed initialization of buffer pool +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: File system buffers for log disabled (block size=512 bytes) +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: End of log at LSN=45568 +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Opened 3 undo tablespaces +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: 128 rollback segments in 3 undo tablespaces are active. +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Setting file './ibtmp1' size to 12.000MiB. Physically writing the file full; Please wait ... +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: File './ibtmp1' size is now 12.000MiB. +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: log sequence number 45568; transaction id 14 +romm-db | 2026-09-23 8:30:42 0 [Note] Plugin 'FEEDBACK' is disabled. +romm-db | 2026-09-23 8:30:42 0 [Note] Plugin 'wsrep-provider' is disabled. +romm-db | 2026-09-23 8:30:43 0 [Note] mariadbd: Event Scheduler: Loaded 0 events +romm-db | 2026-09-23 8:30:43 0 [Note] mariadbd: ready for connections. +romm-db | Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 0 mariadb.org binary distribution +romm-db | 2026-09-23 08:30:43+02:00 [Note] [Entrypoint]: Temporary server started. +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Creating database romm +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Creating user romm +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Giving user romm access to schema romm +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Securing system users (equivalent to running mysql_secure_installation) +romm-db | +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Stopping temporary server +romm-db | 2026-09-23 8:30:44 0 [Note] mariadbd (initiated by: unknown): Normal shutdown +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: FTS optimize thread exiting. +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Starting shutdown... +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Dumping buffer pool(s) to /var/lib/mysql/ib_buffer_pool +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Buffer pool(s) dump completed at 260923 8:30:44 +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Removed temporary tablespace data file: "./ibtmp1" +romm-db | 2026-09-23 8:30:44 0 [Note] Shutdown completed; log sequence number 45568; transaction id 15 +romm-redis | 1:C 23 Sep 2026 08:30:39.880 * Configuration loaded +romm | INFO: [RomM][init][2026-09-23 08:34:50] +romm-redis | 1:M 23 Sep 2026 08:30:39.880 * Increased maximum number of open files to 10032 (it was originally set to 1024). +romm | INFO: [RomM][init][2026-09-23 08:34:50] Version: 5.0.0 +romm-redis | 1:M 23 Sep 2026 08:30:39.880 * monotonic clock: POSIX clock_gettime +romm | INFO: [RomM][init][2026-09-23 08:34:50] +romm-redis | 1:M 23 Sep 2026 08:30:39.881 * Running mode=standalone, port=6379. +romm-redis | 1:M 23 Sep 2026 08:30:39.882 * Server initialized +romm-redis | 1:M 23 Sep 2026 08:30:39.886 * Creating AOF base file appendonly.aof.1.base.rdb on server start +romm-redis | 1:M 23 Sep 2026 08:30:39.899 * Creating AOF incr file appendonly.aof.1.incr.aof on server start +romm-db | 2026-09-23 8:30:44 0 [Note] mariadbd: Shutdown complete +romm-redis | 1:M 23 Sep 2026 08:30:39.899 * Ready to accept connections tcp +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Temporary server stopped +romm-redis | 1:M 23 Sep 2026 08:31:40.061 * 10000 changes in 60 seconds. Saving... +romm-db | +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: MariaDB init process done. Ready for start up. +romm-db | +romm | INFO: [RomM][init][2026-09-23 08:34:50] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:34:50] REDIS_HOST is set, not starting internal valkey-server +romm-db | 2026-09-23 8:30:44 0 [Note] Starting MariaDB 11.4.13-MariaDB-ubu2404 source revision 170b1d70737be6f134448f51713cdc1ae215b420 server_uid Psh/IuH2qRBz9PeEZH//qqIAdts= as process 1 +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Compressed tables use zlib 1.3 +romm-redis | 1:M 23 Sep 2026 08:31:40.062 * Background saving started by pid 46 +romm-redis | 46:C 23 Sep 2026 08:31:40.211 * DB saved on disk +romm-redis | 46:C 23 Sep 2026 08:31:40.212 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 1 MB +romm-redis | 1:M 23 Sep 2026 08:31:40.264 * Background saving terminated with success +romm-redis | 1:M 23 Sep 2026 08:36:41.033 * 100 changes in 300 seconds. Saving... +romm | INFO: [RomM][init][2026-09-23 08:34:50] Running database migrations +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Number of transaction pools: 1 +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Using crc32 + pclmulqdq instructions +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm-db | 2026-09-23 8:30:44 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-redis | 1:M 23 Sep 2026 08:36:41.034 * Background saving started by pid 226 +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Using Linux native AIO +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: innodb_buffer_pool_size_max=8388608m, innodb_buffer_pool_size=128m +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Completed initialization of buffer pool +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: File system buffers for log disabled (block size=512 bytes) +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: End of log at LSN=45568 +romm | ERROR: [RomM][init][2026-09-23 08:34:59] Failed to run database migrations +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Opened 3 undo tablespaces +romm-redis | 226:C 23 Sep 2026 08:36:41.158 * DB saved on disk +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: 128 rollback segments in 3 undo tablespaces are active. +romm-redis | 226:C 23 Sep 2026 08:36:41.159 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 0 MB +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Setting file './ibtmp1' size to 12.000MiB. Physically writing the file full; Please wait ... +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: File './ibtmp1' size is now 12.000MiB. +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: log sequence number 45568; transaction id 14 +romm-db | 2026-09-23 8:30:44 0 [Note] Plugin 'FEEDBACK' is disabled. +romm-db | 2026-09-23 8:30:44 0 [Note] Plugin 'wsrep-provider' is disabled. +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Loading buffer pool(s) from /var/lib/mysql/ib_buffer_pool +romm-redis | 1:M 23 Sep 2026 08:36:41.236 * Background saving terminated with success +romm | INFO: [RomM][init][2026-09-23 08:35:01] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:35:01] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:35:01] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:35:01] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:35:01] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:35:01] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:35:01] +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Buffer pool(s) load completed at 260923 8:30:44 +romm-db | 2026-09-23 8:30:48 0 [Note] Server socket created on IP: '0.0.0.0', port: '3306'. +romm-db | 2026-09-23 8:30:48 0 [Note] Server socket created on IP: '::', port: '3306'. +romm-db | 2026-09-23 8:30:48 0 [Note] mariadbd: Event Scheduler: Loaded 0 events +romm-db | 2026-09-23 8:30:48 0 [Note] mariadbd: ready for connections. +romm-db | Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution +romm-db | 2026-09-23 8:31:08 5 [Warning] Aborted connection 5 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:31:38 15 [Warning] Aborted connection 15 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm | INFO: [RomM][init][2026-09-23 08:35:01] The beautiful, powerful, self-hosted Rom manager and player +romm-db | 2026-09-23 8:31:54 18 [Warning] Aborted connection 18 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:33:51 33 [Warning] Aborted connection 33 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:33:51 41 [Warning] Aborted connection 41 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:33:51 32 [Warning] Aborted connection 32 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:34:11 36 [Warning] Aborted connection 36 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:34:11 42 [Warning] Aborted connection 42 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:34:11 49 [Warning] Aborted connection 49 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm | INFO: [RomM][init][2026-09-23 08:35:01] +romm | INFO: [RomM][init][2026-09-23 08:35:01] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:35:01] +romm | INFO: [RomM][init][2026-09-23 08:35:01] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:35:01] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:35:01] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:35:09] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:35:13] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:35:13] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:35:13] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:35:13] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:35:13] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:35:13] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:35:13] +romm | INFO: [RomM][init][2026-09-23 08:35:13] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:35:13] +romm | INFO: [RomM][init][2026-09-23 08:35:13] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:35:13] +romm | INFO: [RomM][init][2026-09-23 08:35:13] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:35:13] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:35:13] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:35:22] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:35:28] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:35:28] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:35:28] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:35:28] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:35:28] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:35:28] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:35:28] +romm | INFO: [RomM][init][2026-09-23 08:35:28] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:35:28] +romm | INFO: [RomM][init][2026-09-23 08:35:28] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:35:28] +romm | INFO: [RomM][init][2026-09-23 08:35:28] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:35:28] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:35:28] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:35:37] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:35:50] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:35:50] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:35:50] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:35:50] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:35:50] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:35:50] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:35:50] +romm | INFO: [RomM][init][2026-09-23 08:35:50] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:35:50] +romm | INFO: [RomM][init][2026-09-23 08:35:50] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:35:50] +romm | INFO: [RomM][init][2026-09-23 08:35:50] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:35:50] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:35:50] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:35:59] Failed to run database migrations +romm | INFO: [RomM][init][2026-09-23 08:36:25] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:36:25] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:36:25] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:36:25] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:36:25] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:36:25] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:36:25] +romm | INFO: [RomM][init][2026-09-23 08:36:25] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:36:25] +romm | INFO: [RomM][init][2026-09-23 08:36:25] Version: 5.0.0 +romm | INFO: [RomM][init][2026-09-23 08:36:25] +romm | INFO: [RomM][init][2026-09-23 08:36:25] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:36:25] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:36:25] Running database migrations +romm | FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +romm | ERROR: [RomM][init][2026-09-23 08:36:34] Failed to run database migrations diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/engine-state.json b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/engine-state.json new file mode 100644 index 00000000..ec747fa4 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/engine-state.json @@ -0,0 +1 @@ +null \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/memory-samples.json b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/memory-samples.json new file mode 100644 index 00000000..a90dae00 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/memory-samples.json @@ -0,0 +1,182 @@ +[ + { + "t": 15.2, + "containers": { + "romm": { + "limit": 536870912, + "current": 536096768, + "peak": 536875008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 146841600, + "peak": 152616960, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16109568, + "peak": 32137216, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 286 + }, + { + "t": 30.4, + "containers": { + "romm": { + "limit": 536870912, + "current": 534298624, + "peak": 536875008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 147464192, + "peak": 153600000, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16093184, + "peak": 32137216, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 575 + }, + { + "t": 45.6, + "containers": { + "romm": { + "limit": 536870912, + "current": 535674880, + "peak": 536875008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 147525632, + "peak": 153600000, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16617472, + "peak": 32137216, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 864 + }, + { + "t": 60.8, + "containers": { + "romm": { + "limit": 536870912, + "current": 533225472, + "peak": 536875008, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 147542016, + "peak": 153600000, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16547840, + "peak": 32137216, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 1157 + }, + { + "t": 76.0, + "containers": { + "romm": { + "limit": 536870912, + "current": 534589440, + "peak": 536875008, + "oom_kill": 1, + "restarts": 0, + "oomkilled_flag": true, + "status": "running", + "cgroup": true + }, + "romm-db": { + "limit": 402653184, + "current": 147644416, + "peak": 153600000, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "romm-redis": { + "limit": 134217728, + "current": 16531456, + "peak": 32137216, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 1404 + } +] \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/migration-lines.txt b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/migration-lines.txt new file mode 100644 index 00000000..739ca809 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/migration-lines.txt @@ -0,0 +1,2 @@ +romm | INFO: [RomM][init][2026-09-23 08:31:39] Running database migrations +romm | INFO: [RomM][init][2026-09-23 08:31:54] Database migrations succeeded \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/run.log b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/run.log new file mode 100644 index 00000000..c4624124 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/run.log @@ -0,0 +1,22 @@ +[06:30:39] M1old: deploying romm at FROM {'romm': 'rommapp/romm:5.0.0'} +[06:31:27] FROM settled=True in 36.7s :: {"romm": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-db": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-redis": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}} +[06:31:34] romm: POST /api/users http=201 +[06:31:36] romm: login as the seeded user http=200 ok=True +[06:31:36] C1 (seed reads back BEFORE): True +[06:31:36] M1old: swapping to TO {'romm': 'rommapp/romm:5.3.0'} +[06:31:39] TO up -d rc=0 +[06:32:11] TO settled=True in 31.5s :: {"romm": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-db": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "romm-redis": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}} +[06:32:11] migration lines observed: 2 +[06:32:44] romm: login as the seeded user http=200 ok=True +[06:32:44] RESULT (seed reads back AFTER): True +[06:32:44] memory watch: 600s, 4 callers on 8 path(s) at 172.18.0.6:8080 +[06:33:00] + 15s romm=511M/512M peak=512M kills=0 rs=0 romm-db=140M/384M peak=145M kills=0 rs=0 romm-redis=15M/128M peak=30M kills=0 rs=0 reqs=286 +[06:33:15] + 30s romm=509M/512M peak=512M kills=0 rs=0 romm-db=140M/384M peak=146M kills=0 rs=0 romm-redis=15M/128M peak=30M kills=0 rs=0 reqs=575 +[06:33:30] + 46s romm=510M/512M peak=512M kills=0 rs=0 romm-db=140M/384M peak=146M kills=0 rs=0 romm-redis=15M/128M peak=30M kills=0 rs=0 reqs=864 +[06:33:45] + 61s romm=508M/512M peak=512M kills=0 rs=0 romm-db=140M/384M peak=146M kills=0 rs=0 romm-redis=15M/128M peak=30M kills=0 rs=0 reqs=1157 +[06:34:00] + 76s romm=509M/512M peak=512M kills=1 rs=0 romm-db=140M/384M peak=146M kills=0 rs=0 romm-redis=15M/128M peak=30M kills=0 rs=0 reqs=1404 +[06:34:00] memory watch: STOPPING EARLY — {'t': 76.0, 'container': 'romm', 'oom_kills': 1, 'restarts': 0} +[06:34:01] memory watch: killed=True tight=['romm'] requests=1405 codes={'401': 704, '200': 701} +[06:34:01] VERDICT -> failed: the new version was OOM-killed or restarted under light load +[06:34:01] M1old: ABORT — putting the FROM images back +[06:37:15] ABORT: the app did NOT come back (rc=0, 182.4s) \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/to-full.log b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/to-full.log new file mode 100644 index 00000000..5e954c36 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/to-full.log @@ -0,0 +1,133 @@ +romm-db | 2026-09-23 08:30:40+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:11.4.13+maria~ubu2404 started. +romm-db | 2026-09-23 08:30:40+02:00 [Warn] [Entrypoint]: /sys/fs/cgroup///memory.pressure not writable, functionality unavailable to MariaDB +romm-db | 2026-09-23 08:30:40+02:00 [Note] [Entrypoint]: Switching to dedicated user 'mysql' +romm | INFO: [RomM][init][2026-09-23 08:31:39] _____ __ __ +romm | INFO: [RomM][init][2026-09-23 08:31:39] | __ \ | \/ | +romm | INFO: [RomM][init][2026-09-23 08:31:39] | |__) |___ _ __ ___ | \ / | +romm | INFO: [RomM][init][2026-09-23 08:31:39] | _ // _ \| '_ ` _ \| |\/| | +romm | INFO: [RomM][init][2026-09-23 08:31:39] | | \ \ (_) | | | | | | | | | +romm | INFO: [RomM][init][2026-09-23 08:31:39] |_| \_\___/|_| |_| |_|_| |_| +romm | INFO: [RomM][init][2026-09-23 08:31:39] +romm | INFO: [RomM][init][2026-09-23 08:31:39] The beautiful, powerful, self-hosted Rom manager and player +romm | INFO: [RomM][init][2026-09-23 08:31:39] +romm | INFO: [RomM][init][2026-09-23 08:31:39] Version: 5.3.0 +romm | INFO: [RomM][init][2026-09-23 08:31:39] +romm | INFO: [RomM][init][2026-09-23 08:31:39] No OpenTelemetry environment variables found, disabling OpenTelemetry SDK +romm | INFO: [RomM][init][2026-09-23 08:31:39] REDIS_HOST is set, not starting internal valkey-server +romm | INFO: [RomM][init][2026-09-23 08:31:39] Running database migrations +romm | INFO: [RomM][init][2026-09-23 08:31:54] Database migrations succeeded +romm | INFO: [RomM][startup][2026-09-23 08:32:06] Running startup tasks +romm | INFO: [RomM][startup][2026-09-23 08:32:06] Cleared 2 job(s) left behind by the old scheduler +romm | INFO: [RomM][startup][2026-09-23 08:32:06] Initializing cache with fixtures data +romm | INFO: [RomM][startup][2026-09-23 08:32:06] Startup tasks completed +romm | INFO: [RomM][init][2026-09-23 08:32:07] Starting backend +romm | INFO: [RomM][init][2026-09-23 08:32:07] Starting RQ cron scheduler +romm | INFO: [RomM][init][2026-09-23 08:32:07] Starting RQ worker +romm | INFO: [RomM][init][2026-09-23 08:32:07] Starting RQ scan worker +romm | INFO: [RomM][init][2026-09-23 08:32:09] Starting nginx +romm | 2026/09/23 08:32:09 [notice] 119#119: js vm init njs: 00007E35328F6B00 +romm | INFO: [RomM][init][2026-09-23 08:32:09] 🚀 RomM is now available at http://0.0.0.0:8080 +romm | 08:32:09 Worker 4e9ae2d061064c5d9b285b68af387d4f: started with PID 107, version 2.12.0 +romm | 08:32:09 Worker 4e9ae2d061064c5d9b285b68af387d4f: subscribing to channel rq:pubsub:4e9ae2d061064c5d9b285b68af387d4f +romm | 08:32:09 Loading cron configuration from tasks.cron_config +romm | 08:32:09 *** Listening on scans... +romm | 08:32:09 Worker 55f7c62f88dd43c29e0873eb10c480ad: started with PID 104, version 2.12.0 +romm | 08:32:09 Acquired scheduler lock for scans +romm | 08:32:09 Worker 55f7c62f88dd43c29e0873eb10c480ad: subscribing to channel rq:pubsub:55f7c62f88dd43c29e0873eb10c480ad +romm | 08:32:09 *** Listening on high, default, low... +romm | 08:32:09 Acquired scheduler lock for high +romm | 08:32:09 Acquired scheduler lock for default +romm | 08:32:09 Acquired scheduler lock for low +romm | 08:32:09 Scheduler for scans started with PID 145 +romm | 08:32:09 Scheduler for high, default, low started with PID 144 +romm-db | 2026-09-23 08:30:40+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:11.4.13+maria~ubu2404 started. +romm-db | 2026-09-23 08:30:40+02:00 [Note] [Entrypoint]: Initializing database files +romm-db | 2026-09-23 8:30:40 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 08:30:42+02:00 [Note] [Entrypoint]: Database files initialized +romm-db | 2026-09-23 08:30:42+02:00 [Note] [Entrypoint]: Starting temporary server +romm-db | 2026-09-23 08:30:42+02:00 [Note] [Entrypoint]: Waiting for server startup +romm-db | 2026-09-23 8:30:42 0 [Note] Starting MariaDB 11.4.13-MariaDB-ubu2404 source revision 170b1d70737be6f134448f51713cdc1ae215b420 server_uid Psh/IuH2qRBz9PeEZH//qqIAdts= as process 92 +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Compressed tables use zlib 1.3 +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Number of transaction pools: 1 +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Using crc32 + pclmulqdq instructions +romm-db | 2026-09-23 8:30:42 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Using Linux native AIO +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: innodb_buffer_pool_size_max=8388608m, innodb_buffer_pool_size=128m +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Completed initialization of buffer pool +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: File system buffers for log disabled (block size=512 bytes) +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: End of log at LSN=45568 +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Opened 3 undo tablespaces +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: 128 rollback segments in 3 undo tablespaces are active. +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: Setting file './ibtmp1' size to 12.000MiB. Physically writing the file full; Please wait ... +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: File './ibtmp1' size is now 12.000MiB. +romm-db | 2026-09-23 8:30:42 0 [Note] InnoDB: log sequence number 45568; transaction id 14 +romm-db | 2026-09-23 8:30:42 0 [Note] Plugin 'FEEDBACK' is disabled. +romm-db | 2026-09-23 8:30:42 0 [Note] Plugin 'wsrep-provider' is disabled. +romm-db | 2026-09-23 8:30:43 0 [Note] mariadbd: Event Scheduler: Loaded 0 events +romm-db | 2026-09-23 8:30:43 0 [Note] mariadbd: ready for connections. +romm-db | Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 0 mariadb.org binary distribution +romm-db | 2026-09-23 08:30:43+02:00 [Note] [Entrypoint]: Temporary server started. +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Creating database romm +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Creating user romm +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Giving user romm access to schema romm +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Securing system users (equivalent to running mysql_secure_installation) +romm-db | +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Stopping temporary server +romm-db | 2026-09-23 8:30:44 0 [Note] mariadbd (initiated by: unknown): Normal shutdown +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: FTS optimize thread exiting. +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Starting shutdown... +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Dumping buffer pool(s) to /var/lib/mysql/ib_buffer_pool +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Buffer pool(s) dump completed at 260923 8:30:44 +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Removed temporary tablespace data file: "./ibtmp1" +romm-db | 2026-09-23 8:30:44 0 [Note] Shutdown completed; log sequence number 45568; transaction id 15 +romm-db | 2026-09-23 8:30:44 0 [Note] mariadbd: Shutdown complete +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: Temporary server stopped +romm-db | +romm-db | 2026-09-23 08:30:44+02:00 [Note] [Entrypoint]: MariaDB init process done. Ready for start up. +romm-db | +romm-db | 2026-09-23 8:30:44 0 [Note] Starting MariaDB 11.4.13-MariaDB-ubu2404 source revision 170b1d70737be6f134448f51713cdc1ae215b420 server_uid Psh/IuH2qRBz9PeEZH//qqIAdts= as process 1 +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Compressed tables use zlib 1.3 +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Number of transaction pools: 1 +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Using crc32 + pclmulqdq instructions +romm-db | 2026-09-23 8:30:44 0 [Warning] mariadbd: io_uring_queue_init() failed with EPERM: sysctl kernel.io_uring_disabled has the value 2, or 1 and the user of the process is not a member of sysctl kernel.io_uring_group. (see man 2 io_uring_setup). +romm-db | create_uring failed: falling back to libaio +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Using Linux native AIO +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: innodb_buffer_pool_size_max=8388608m, innodb_buffer_pool_size=128m +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Completed initialization of buffer pool +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: File system buffers for log disabled (block size=512 bytes) +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: End of log at LSN=45568 +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Opened 3 undo tablespaces +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: 128 rollback segments in 3 undo tablespaces are active. +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Setting file './ibtmp1' size to 12.000MiB. Physically writing the file full; Please wait ... +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: File './ibtmp1' size is now 12.000MiB. +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: log sequence number 45568; transaction id 14 +romm-db | 2026-09-23 8:30:44 0 [Note] Plugin 'FEEDBACK' is disabled. +romm-db | 2026-09-23 8:30:44 0 [Note] Plugin 'wsrep-provider' is disabled. +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Loading buffer pool(s) from /var/lib/mysql/ib_buffer_pool +romm-db | 2026-09-23 8:30:44 0 [Note] InnoDB: Buffer pool(s) load completed at 260923 8:30:44 +romm-db | 2026-09-23 8:30:48 0 [Note] Server socket created on IP: '0.0.0.0', port: '3306'. +romm-db | 2026-09-23 8:30:48 0 [Note] Server socket created on IP: '::', port: '3306'. +romm-db | 2026-09-23 8:30:48 0 [Note] mariadbd: Event Scheduler: Loaded 0 events +romm-db | 2026-09-23 8:30:48 0 [Note] mariadbd: ready for connections. +romm-db | Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution +romm-db | 2026-09-23 8:31:08 5 [Warning] Aborted connection 5 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:31:38 15 [Warning] Aborted connection 15 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-db | 2026-09-23 8:31:54 18 [Warning] Aborted connection 18 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets) +romm-redis | 1:C 23 Sep 2026 08:30:39.879 # WARNING Memory overcommit must be enabled! Without it, a background save or replication may fail under low memory condition. Being disabled, it can also cause failures without low memory condition, see https://github.com/jemalloc/jemalloc/issues/1328. To fix this issue add 'vm.overcommit_memory = 1' to /etc/sysctl.conf and then reboot or run the command 'sysctl vm.overcommit_memory=1' for this to take effect. +romm-redis | 1:C 23 Sep 2026 08:30:39.880 * oO0OoO0OoO0Oo Redis is starting oO0OoO0OoO0Oo +romm-redis | 1:C 23 Sep 2026 08:30:39.880 * Redis version=7.4.11, bits=64, commit=00000000, modified=0, pid=1, just started +romm-redis | 1:C 23 Sep 2026 08:30:39.880 * Configuration loaded +romm-redis | 1:M 23 Sep 2026 08:30:39.880 * Increased maximum number of open files to 10032 (it was originally set to 1024). +romm-redis | 1:M 23 Sep 2026 08:30:39.880 * monotonic clock: POSIX clock_gettime +romm-redis | 1:M 23 Sep 2026 08:30:39.881 * Running mode=standalone, port=6379. +romm-redis | 1:M 23 Sep 2026 08:30:39.882 * Server initialized +romm-redis | 1:M 23 Sep 2026 08:30:39.886 * Creating AOF base file appendonly.aof.1.base.rdb on server start +romm-redis | 1:M 23 Sep 2026 08:30:39.899 * Creating AOF incr file appendonly.aof.1.incr.aof on server start +romm-redis | 1:M 23 Sep 2026 08:30:39.899 * Ready to accept connections tcp +romm-redis | 1:M 23 Sep 2026 08:31:40.061 * 10000 changes in 60 seconds. Saving... +romm-redis | 1:M 23 Sep 2026 08:31:40.062 * Background saving started by pid 46 +romm-redis | 46:C 23 Sep 2026 08:31:40.211 * DB saved on disk +romm-redis | 46:C 23 Sep 2026 08:31:40.212 * Fork CoW for RDB: current 1 MB, peak 1 MB, average 1 MB +romm-redis | 1:M 23 Sep 2026 08:31:40.264 * Background saving terminated with success diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/to-states.json b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/to-states.json new file mode 100644 index 00000000..1c7ad7de --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/to-states.json @@ -0,0 +1,20 @@ +{ + "romm": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + }, + "romm-db": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + }, + "romm-redis": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + } +} \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/verdict.json b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/verdict.json new file mode 100644 index 00000000..5b163a5c --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/harness/evidence/M1old/verdict.json @@ -0,0 +1,70 @@ +{ + "harness_version": 2, + "edge": "M1old", + "app": "romm", + "note": "the same move on the template AS PROMOTED (512M, 4 workers) \u2014 must FAIL the memory watch", + "from": { + "romm": "rommapp/romm:5.0.0" + }, + "to": { + "romm": "rommapp/romm:5.3.0" + }, + "verdict": "failed", + "seed_read_before": true, + "seed_read_after": true, + "healthy_after": true, + "migration_observed": "romm | \u001b[0;32mINFO: \u001b[0;34m[RomM]\u001b[0;95m[init]\u001b[0;36m[2026-09-23 08:31:39]\u001b[0;00m Running database migrations", + "abort": "refuses", + "abort_detail": "ection 33 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:33:51 41 [Warning] Aborted connection 41 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:33:51 32 [Warning] Aborted connection 32 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:34:11 36 [Warning] Aborted connection 36 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:34:11 42 [Warning] Aborted connection 42 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)\nromm-db | 2026-09-23 8:34:11 49 [Warning] Aborted connection 49 to db: 'romm' user: 'romm' host: '172.20.0.4' (Got an error reading communication packets)", + "engine_state_after": null, + "memory": { + "soak_s": 76.5, + "requested_s": 600, + "requests": 1405, + "codes": { + "401": 704, + "200": 701 + }, + "first_kill": { + "t": 76.0, + "container": "romm", + "oom_kills": 1, + "restarts": 0 + }, + "containers": { + "romm": { + "limit": 536870912, + "peak": 536875008, + "peak_pct": 1.0, + "oom_kills": 1, + "restarts": 0, + "oomkilled_flag": true, + "measured": true + }, + "romm-db": { + "limit": 402653184, + "peak": 153600000, + "peak_pct": 0.381, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + }, + "romm-redis": { + "limit": 134217728, + "peak": 32137216, + "peak_pct": 0.239, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + } + }, + "unmeasured": [] + }, + "marks": [], + "duration_s": 31.5, + "measured_at": "2026-09-23T06:37:15Z", + "evidence": "evidence/M1old", + "total_s": 395.7 +} \ No newline at end of file diff --git a/documentation/audits/update-rulings-2026-09-23/repoint.py b/documentation/audits/update-rulings-2026-09-23/repoint.py new file mode 100644 index 00000000..110d1f2b --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/repoint.py @@ -0,0 +1,60 @@ +#!/usr/bin/env python3 +"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config. + +`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is +`controller.yaml.pre-rulings` (NOT the older `.pre-28`, which a restore must never pick up). +""" +import re, sys, io +sys.path.insert(0, '.') +import walk as w + +VOL = "/var/lib/docker/volumes/felhom-controller-data/_data" +DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git" + + +def creds(): + for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"): + m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l) + if m: + return m.group(1), m.group(2) + raise SystemExit("no admin credential") + + +def to_drill(): + u, t = creds() + print(w.guest(f""" +set -e +test -f {VOL}/controller.yaml.pre-rulings || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-rulings +python3 - <<'PY' +import re +p = "{VOL}/controller.yaml" +s = open(p).read() +s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M) +if not re.search(r'^update:', s, re.M): + s += "update:\\n health_timeout: 90s\\n" +open(p, "w").write(s) +PY +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +sleep 15 +grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: /' +grep -A2 '^update:' {VOL}/controller.yaml +""")) + + +def restore(): + print(w.guest(f""" +set -e +cp -p {VOL}/controller.yaml.pre-rulings {VOL}/controller.yaml +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +sleep 15 +grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: /' +grep -c '^update:' {VOL}/controller.yaml || true +""")) + + +if __name__ == "__main__": + to_drill() if sys.argv[1] == "drill" else restore() diff --git a/documentation/audits/update-rulings-2026-09-23/romm-10-deploy-seed.txt b/documentation/audits/update-rulings-2026-09-23/romm-10-deploy-seed.txt new file mode 100644 index 00000000..40a7ee9d --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/romm-10-deploy-seed.txt @@ -0,0 +1,22 @@ +08:12:14 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/romm'] +08:12:14 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['HDD_PATH'] +08:12:14 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +08:13:05 [1] deployed, controller state=running, pinned={'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'} +08:13:05 deployed: True +08:13:07 {"pinned_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "installed_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "catalog_images": null, "live_compose_image_lines": ["image: rommapp/romm:5.0.0", "image: mariadb:11.4", "image: redis:7-alpine"], "docker_inspect": ["romm rommapp/romm:5.0.0 running=true restarts=0", "romm-db mariadb:11.4 running=true restarts=0", "romm-redis redis:7-alpine running=true restarts=0"]} +08:13:09 bind sources: ['/mnt/felhom-drives/scratch_hdd/userdata/romm/userdata/roms', '/mnt/felhom-drives/scratch_hdd/userdata/romm/appdata/romm/resources'] +08:13:11 romm: POST /api/users http=201 +08:13:13 romm: login as the seeded user http=200 ok=True +08:13:13 C1 A: True +08:13:13 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +08:14:13 [4] backup idle; last=None +08:14:13 romm seed B: POST /api/users (as the admin) http=404 404 page not found + +08:14:13 B reads back: None +08:14:16 /mnt/felhom-drives/scratch_hdd/userdata/romm/userdata/roms: files=0 sum=e3b0c44298fc1c14 +/mnt/felhom-drives/scratch_hdd/userdata/romm/appdata/romm/resources: files=0 sum=e3b0c44298fc1c14 + +08:14:35 seed B retry — the first try met traefik's 404 while the backup's volume dump had the app stopped +08:14:47 romm seed B: POST /api/users (as the admin) http=201 {"id":2,"username":"drillbecef29","email":"drillbecef29@gate.invalid","enabled":true,"role":"user","permission_group_id" +08:14:47 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False) +08:14:47 B reads back: True diff --git a/documentation/audits/update-rulings-2026-09-23/romm-20-break-and-update.txt b/documentation/audits/update-rulings-2026-09-23/romm-20-break-and-update.txt new file mode 100644 index 00000000..9ec926a3 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/romm-20-break-and-update.txt @@ -0,0 +1,19 @@ +08:15:12 mariadb before update: tables=27 alembic=0095_virtual_collections_source users=2 +08:15:13 [5] drill commit 57a4be3d2832: romm rommapp/romm:5.0.0 -> rommapp/romm:5.3.0 (push rc=0) +08:15:18 badge caught up after 4.5 s +08:15:18 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:15:18 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:15:19 + 1.0s phase=pulling label=Új verzió letöltése… err=None hold=None +08:15:20 + 2.1s phase=starting label=Indítás az új verzióval… err=None hold=None +08:15:25 + 7.2s phase=verifying label=Működés ellenőrzése… err=None hold=None +08:17:00 + 102.6s phase=failed label=A frissítés nem sikerült err=A(z) romm frissítése 2026-09-23 08:17-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 08:14 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem. hold=A(z) romm frissítése 2026-09-23 08:17-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 08:14 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem. +08:17:01 {"final_phase": "failed", "hold_reason": "A(z) romm frissítése 2026-09-23 08:17-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 08:14 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.", "state": "stopped", "duration_s": 102.6} +08:17:03 {"pinned_images": {"romm": "rommapp/romm:5.3.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "installed_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "catalog_images": {"romm": "rommapp/romm:5.3.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: rommapp/romm:5.3.0", "image: mariadb:11.4", "image: redis:7-alpine"], "docker_inspect": []} +08:17:05 /mnt/felhom-drives/scratch_hdd/userdata/romm/userdata/roms: files=0 sum=e3b0c44298fc1c14 +/mnt/felhom-drives/scratch_hdd/userdata/romm/appdata/romm/resources: files=0 sum=e3b0c44298fc1c14 + +08:17:07 20260923T061656Z +3 +ls: cannot access '/mnt/sys_drive/felhom-data/backups/primary/romm/db-dumps/': No such file or directory +/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm/db-dumps + diff --git a/documentation/audits/update-rulings-2026-09-23/romm-40-undo-steps-1-2-2b.txt b/documentation/audits/update-rulings-2026-09-23/romm-40-undo-steps-1-2-2b.txt new file mode 100644 index 00000000..21f037e2 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/romm-40-undo-steps-1-2-2b.txt @@ -0,0 +1,35 @@ +=== what the new version printed about its schema (kept hold log) +romm-db | 2026-09-23 08:13:58+02:00 [Note] [Entrypoint]: MariaDB upgrade not required +romm | INFO: [RomM][init][2026-09-23 08:15:25] Running database migrations +romm | INFO: [RomM][init][2026-09-23 08:15:41] Database migrations succeeded +=== step 1: the safety dump +/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm/db-dumps: +total 136 +drwxr-sr-x 2 root 1000 4096 06:15:18 . +drwxr-sr-x 5 root 1000 4096 06:14:09 .. +-rw-r--r-- 1 root 1000 62943 06:15:18 pre-restore-20260923T061518Z-romm-mariadb.sql +-rw-r--r-- 1 root 1000 62748 06:13:13 romm-mariadb.sql + +/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm/volume-dumps: +total 170856 +drwxr-sr-x 2 root 1000 4096 06:13:57 . +drwxr-sr-x 5 root 1000 4096 06:14:09 .. +-rw-r--r-- 1 root 1000 2560 06:13:56 romm_romm_config.tar +-rw-r--r-- 1 root 1000 160428544 06:13:56 romm_romm_db_data.tar +-rw-r--r-- 1 root 1000 14506496 06:13:57 romm_romm_redis_data.tar +undo copy: /mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm/db-dumps/pre-restore-20260923T061518Z-romm-mariadb.sql size=62943 +holds: CREATE TABLE=24 DROP TABLE IF EXISTS=27 FOREIGN_KEY_CHECKS=0 set: 1 completion marker: 1 +seed B in the undo copy: 1; in the tier copy: 0 +step 1 0.04s +=== step 2: old definition + pin back (from the unit; the journal's pre-update copies: 0) + image: rommapp/romm:5.0.0 + port: 8080 +pinned_images: + romm: rommapp/romm:5.0.0 + romm-db: mariadb:11.4 + romm-redis: redis:7-alpine +step 2 0.04s +=== step 2b: lift the hold (operator CLI) + controller restart +[INFO] [settings] Loaded settings from /opt/docker/felhom-controller/data/settings.json +[INFO] [settings] restore hold CLEARED for romm +restarted diff --git a/documentation/audits/update-rulings-2026-09-23/romm-43-negative-control.txt b/documentation/audits/update-rulings-2026-09-23/romm-43-negative-control.txt new file mode 100644 index 00000000..180b3f23 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/romm-43-negative-control.txt @@ -0,0 +1,16 @@ +08:17:47 hold after lift: None +08:17:58 product start -> 200 {'ok': True, 'message': 'Stack romm start completed'} +08:19:01 romm rommapp/romm:5.0.0 Up 2 seconds (health: starting) +romm-db mariadb:11.4 Up About a minute (healthy) +romm-redis redis:7-alpine Up About a minute (healthy) +INFO: [RomM][init][2026-09-23 08:17:58] Running database migrations +FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +ERROR: [RomM][init][2026-09-23 08:18:07] Failed to run database migrations +INFO: [RomM][init][2026-09-23 08:18:07] Running database migrations +FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +ERROR: [RomM][init][2026-09-23 08:18:16] Failed to run database migrations +INFO: [RomM][init][2026-09-23 08:18:17] Running database migrations +FAILED: Can't locate revision identified by '0128_hltb_main_story_column' + +08:19:01 front door /api/heartbeat: 404 +08:19:03 mariadb now (migrated): tables=39 alembic=0128_hltb_main_story_column users=2 diff --git a/documentation/audits/update-rulings-2026-09-23/romm-44-loads-truncated-then-full.txt b/documentation/audits/update-rulings-2026-09-23/romm-44-loads-truncated-then-full.txt new file mode 100644 index 00000000..abc17533 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/romm-44-loads-truncated-then-full.txt @@ -0,0 +1,15 @@ +app container stopped, DB up +before any load (migrated): tables=39 alembic=0128_hltb_main_story_column users=2 +=== WRONG CASE: truncated copy (31471 of 62943 bytes), loaded as the product's ImportDump does (mariadb < file) +rc=1 in 0.94s + +ERROR 1064 (42000) at line 754: You have an error in your SQL syntax; check the manual that corresponds to your MariaDB server version for the right syntax to use near '`sa' at line 1 +after the truncated load: tables=39 alembic=0095_virtual_collections_source users=2 +=== RIGHT CASE: the whole copy, same loader +rc=0 in 1.25s +after the whole load: tables=39 alembic=0095_virtual_collections_source users=2 +tables the NEW version created that remain (not in the copy): +memory_card_versions memory_cards music_favorite_tracks music_playlist_tracks music_playlists rom_file_doc_meta rom_file_user rom_identity_keys rom_similarity roms_facets roms_metadata sibling_roms streaming_container_adoptions virtual_collection_roms virtual_collections +views now: roms_metadata virtual_collections sibling_roms +views in the copy: 12 lines mention VIEW; CREATE ... VIEW statements: 3 +base tables now: 36 diff --git a/documentation/audits/update-rulings-2026-09-23/romm-48-start-and-readback.txt b/documentation/audits/update-rulings-2026-09-23/romm-48-start-and-readback.txt new file mode 100644 index 00000000..c58d57af --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/romm-48-start-and-readback.txt @@ -0,0 +1,16 @@ +08:20:00 CORRECTION to the listing above: roms_metadata, virtual_collections, sibling_roms are the copy's own VIEWS (not CREATE TABLE lines) — the real leftovers are 12 base tables (36 now vs 24 in the copy) +08:20:03 started + +08:20:37 health (the OLD probe's port 8080 via the front door /api/heartbeat) = True after 36.5s +08:20:39 rommapp/romm:5.0.0 running restarts=0 +INFO: [RomM][init][2026-09-23 08:19:14] Running database migrations +FAILED: Can't locate revision identified by '0128_hltb_main_story_column' +ERROR: [RomM][init][2026-09-23 08:19:22] Failed to run database migrations +INFO: [RomM][init][2026-09-23 08:20:03] Running database migrations + +08:20:41 romm: login as the seeded user http=200 ok=True +08:20:41 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False) +08:20:41 READBACK after the undo: seed A=True seed B (only in the safety dump)=True +08:20:43 /mnt/felhom-drives/scratch_hdd/userdata/romm/userdata/roms: files=0 sum=e3b0c44298fc1c14 +/mnt/felhom-drives/scratch_hdd/userdata/romm/appdata/romm/resources: files=0 sum=e3b0c44298fc1c14 + diff --git a/documentation/audits/update-rulings-2026-09-23/spike.py b/documentation/audits/update-rulings-2026-09-23/spike.py new file mode 100644 index 00000000..f0de8d13 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/spike.py @@ -0,0 +1,160 @@ +#!/usr/bin/env python3 +"""spike.py — Part 1 of the 2026-09-23 brief: the AUTOMATIC UNDO, performed BY HAND on guest 9202. + +EVIDENCE, NOT PRODUCT. Every product act goes through the endpoints the UI invokes (deploy, backup, +sync, rescan, update, remove). The UNDO itself has no product path yet (that is what is being +spiked), so it is performed by hand with plain docker/compose inside the guest, using exactly the +steps the product would take, each one timed. + +State between stages lives in state-.json so each stage can be run, read, and only then +followed by the next (the hand undo needs a person looking at what the previous step left). +""" +import json, os, sys, time, re +sys.path.insert(0, ".") +import walk as w +import fixtures as fx + +HERE = os.path.dirname(os.path.abspath(__file__)) +FX = {"docmost": fx.Docmost(), "vikunja": fx.Vikunja(), "romm": fx.Romm()} +SUB = {"docmost": "docs", "vikunja": "tasks", "romm": "arcade"} + + +def st_path(app): + return os.path.join(HERE, f"state-{app}.json") + + +def load(app): + return json.load(open(st_path(app))) if os.path.exists(st_path(app)) else {} + + +def save(app, s): + json.dump(s, open(st_path(app), "w"), indent=2, ensure_ascii=False) + + +def ts(): + return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) + + +# ---- a SECOND seed, written AFTER the backup and BEFORE the update. It is the discriminator: only +# the pre-pin safety dump can hold it — the backup tier copy was taken before it existed. So if it +# reads back after the undo, the undo used the safety dump; if only A reads back, it used the tier. +def docmost_seed_b(sub, A): + jar = "/tmp/dm.jar" + w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json", + data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST") + name = "drillB" + os.urandom(3).hex() + rc, code, out = w.app_curl(sub, "/api/spaces/create", "-b", jar, "-H", "Content-Type: application/json", + data=json.dumps({"name": name, "slug": name.lower()}), method="POST") + w.say(f" docmost seed B: /api/spaces/create http={code} {out[:160]}") + return {"space": name} if code in ("200", "201") else None + + +def docmost_verify_b(sub, A, B): + jar = "/tmp/dm.jar" + rc, code, out = w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json", + data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST") + if code not in ("200", "201"): + w.say(f" docmost B: cannot log in (http={code})"); return False + rc, code, out = w.app_curl(sub, "/api/spaces", "-b", jar, "-H", "Content-Type: application/json", + data="{}", method="POST") + ok = code in ("200", "201") and B["space"] in out + neg = "drillBnever" in out + w.say(f" docmost B: /api/spaces http={code} seeded-space-listed={ok} (negative control listed={neg})") + return ok and not neg + + +def break_edge(app, frm, to, port_from, port_to): + """The failing edge: a REAL migrating image move, plus — in the DRILL template only — the + health probe pointed at a port the app does not answer. Both in one drill commit.""" + fy = f"{w.DRILL}/templates/{app}/.felhom.yml" + f = open(fy).read() + m = re.search(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )" + str(port_from) + r"\b", f) + assert m, "probe port not found" + f = f[:m.end() - len(str(port_from))] + str(port_to) + f[m.end():] + open(fy, "w").write(f) + h = w.drill_bump(app, frm, to) + return h + + +def pg_state(container, db, user): + return w.guest(f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1; " + f"docker exec {container} psql -U {user} -d {db} -Atc \"select name from kysely_migration order by name desc limit 3\" 2>&1; " + f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from kysely_migration\" 2>&1") + + +def romm_seed_b(sub, A): + """A SECOND RomM user, created by the first (admin) one — written after the backup.""" + jar, tok = FX["romm"]._csrf(w, sub) + user = "drillb" + os.urandom(3).hex() + rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{A['user']}:{A['pw']}", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "email": f"{user}@gate.invalid", + "password": "Drill-" + os.urandom(8).hex(), "role": "viewer"}), + method="POST") + w.say(f" romm seed B: POST /api/users (as the admin) http={code} {body[:120]}") + return {"user": user} if code in ("200", "201") else None + + +def romm_verify_b(sub, A, B): + jar, tok = FX["romm"]._csrf(w, sub) + rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{A['user']}:{A['pw']}") + ok = code == "200" and B["user"] in body + neg = "drillbnever" in body + w.say(f" romm B: GET /api/users http={code} seeded-user-listed={ok} (negative control listed={neg})") + return ok and not neg + + +def tree(paths): + cmd = "; ".join(f"echo \"{p}: files=$(find {p} -type f 2>/dev/null | wc -l) sum=$(find {p} -type f -exec sha256sum {{}} + 2>/dev/null | sort | sha256sum | cut -c1-16)\"" for p in paths) + return w.guest(cmd) + + +def my_state(container="romm-db", db="romm"): + # The root password is used INSIDE the container from its own env — it never leaves it. + q = lambda sql: f"docker exec {container} sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" {db}' 2>&1 | tr '\\n' ' '" + return w.guest(f"""echo -n "tables=$({q("select count(*) from information_schema.tables where table_schema=database()")}) " +echo -n "alembic=$({q("select version_num from alembic_version")}) " +echo -n "users=$({q("select count(*) from users")})" +""") + + +def vik_seed_b(sub, A): + """A second project, created AFTER the backup — plus a task with a real ATTACHMENT (a file the + app writes into its files volume), uploaded through the app's own attachment API.""" + tok, why = FX["vikunja"]._token(w, sub, A) + title = "drillB-" + os.urandom(4).hex() + rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", data=json.dumps({"title": title}), method="PUT") + if code not in ("200", "201"): + w.say(f" vikunja B: project refused {code} {body[:120]}"); return None + pid = json.loads(body)["id"] + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{pid}/tasks", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", data=json.dumps({"title": "task-" + title}), method="PUT") + tid = json.loads(body)["id"] if code in ("200", "201") else None + content = "drill attachment " + os.urandom(8).hex() + fn = "/tmp/vik-att.txt"; open(fn, "w").write(content) + rc, code2, body2 = w.app_curl(sub, f"/api/v1/tasks/{tid}/attachments", "-H", f"Authorization: Bearer {tok}", + "-F", f"files=@{fn}", method="PUT") + w.say(f" vikunja seed B: project http=200 task={tid} attachment upload http={code2} {body2[:120]}") + return {"title": title, "pid": pid, "tid": tid, "att": content} + + +def vik_verify_b(sub, A, B): + tok, why = FX["vikunja"]._token(w, sub, A) + if not tok: + w.say(f" vikunja B: cannot log in {why}"); return {"project": False, "attachment": False} + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{B['pid']}", "-H", f"Authorization: Bearer {tok}") + proj = code == "200" and B["title"] in body + rc, code, body = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments", "-H", f"Authorization: Bearer {tok}") + att_ok = False + try: + atts = json.loads(body) + if atts: + aid = atts[0]["id"] + rc, c3, b3 = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments/{aid}", "-H", f"Authorization: Bearer {tok}") + att_ok = c3 == "200" and B["att"] in b3 + except Exception as e: + w.say(f" vikunja B: attachments list unreadable http={code} {body[:120]}") + w.say(f" vikunja B: project readback={proj} attachment content readback={att_ok}") + return {"project": proj, "attachment": att_ok} diff --git a/documentation/audits/update-rulings-2026-09-23/vikunja-10-deploy-seed.txt b/documentation/audits/update-rulings-2026-09-23/vikunja-10-deploy-seed.txt new file mode 100644 index 00000000..4c12a6b9 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/vikunja-10-deploy-seed.txt @@ -0,0 +1,12 @@ +08:21:03 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +08:21:08 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'} +08:21:08 deployed: True +08:21:11 vikunja: register http=200 +08:21:11 vikunja: create project http=201 +08:21:11 vikunja: readback of the seeded project http=200 ok=True +08:21:11 C1 A: True +08:21:12 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +08:22:12 [4] backup idle; last=None +08:22:17 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drill44f4ec","created":"2026-09 +08:22:17 vikunja B: project readback=True attachment content readback=True +08:22:17 B reads back: {'project': True, 'attachment': True} diff --git a/documentation/audits/update-rulings-2026-09-23/vikunja-20-break-and-update.txt b/documentation/audits/update-rulings-2026-09-23/vikunja-20-break-and-update.txt new file mode 100644 index 00000000..110b3b17 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/vikunja-20-break-and-update.txt @@ -0,0 +1,21 @@ +08:22:26 [5] drill commit 88613e2affbe: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0) +08:22:30 badge caught up after 4.5 s +08:22:30 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:22:30 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:22:31 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None +08:22:32 + 2.1s phase=verifying label=Működés ellenőrzése… err=None hold=None +08:24:04 + 93.3s phase=failed label=A frissítés nem sikerült err=A(z) vikunja frissítése 2026-09-23 08:24-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 08:22 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) vikunja frissítése 2026-09-23 08:24-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 08:22 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. +08:24:04 {"final_phase": "failed", "hold_reason": "A(z) vikunja frissítése 2026-09-23 08:24-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 08:22 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 93.3} +08:24:08 vikunja | time=2026-09-23T08:22:32.956+02:00 level=INFO msg="Running migrations…" +vikunja | time=2026-09-23T08:22:33.018+02:00 level=INFO msg="Ran all migrations successfully." +2026/09/23 06:22:30 update_guard.go:452: [INFO] [backup] update safety dump for vikunja: the app has no database — nothing to copy (no-op) +2026/09/23 06:22:30 update.go:655: [INFO] [stacks] update vikunja: safety dump done (0 file(s)) [] +unit=/mnt/sys_drive/felhom-data/backups/primary/vikunja +ls: cannot access '/mnt/sys_drive/felhom-data/backups/primary/vikunja/db-dumps': No such file or directory +/mnt/sys_drive/felhom-data/backups/primary/vikunja/volume-dumps: +total 2244 +drwxr-xr-x 2 root root 4096 06:22:11 . +drwxr-xr-x 4 root root 4096 06:22:12 .. +-rw-r--r-- 1 root root 1536 06:22:10 vikunja_vikunja_data.tar +-rw-r--r-- 1 root root 2285568 06:22:11 vikunja_vikunja_db.tar + diff --git a/documentation/audits/update-rulings-2026-09-23/vikunja-40-undo-steps-2-2b.txt b/documentation/audits/update-rulings-2026-09-23/vikunja-40-undo-steps-2-2b.txt new file mode 100644 index 00000000..731792c4 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/vikunja-40-undo-steps-2-2b.txt @@ -0,0 +1,6 @@ + image: vikunja/vikunja:2.3.0 +pinned_images: + vikunja: vikunja/vikunja:2.3.0 +step 2 0.02s +[INFO] [settings] restore hold CLEARED for vikunja +restarted diff --git a/documentation/audits/update-rulings-2026-09-23/vikunja-43-old-on-migrated-sqlite.txt b/documentation/audits/update-rulings-2026-09-23/vikunja-43-old-on-migrated-sqlite.txt new file mode 100644 index 00000000..4534bb4c --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/vikunja-43-old-on-migrated-sqlite.txt @@ -0,0 +1,10 @@ +08:24:39 product start (old 2.3.0 on the 2.6.0-migrated SQLite, NOTHING loaded) -> 200 +08:24:39 answers on its own port: True after 0.7s +08:24:41 vikunja/vikunja:2.3.0 running restarts=0 +time=2026-09-23T08:24:39.591+02:00 level=INFO msg="Running migrations…" +time=2026-09-23T08:24:39.602+02:00 level=INFO msg="Ran all migrations successfully." +time=2026-09-23T08:24:39.608+02:00 level=INFO msg="Vikunja version v2.3.0" + +08:24:42 vikunja: readback of the seeded project http=200 ok=True +08:24:42 vikunja B: project readback=True attachment content readback=True +08:24:42 READBACK: A=True B={'project': True, 'attachment': True} diff --git a/documentation/audits/update-rulings-2026-09-23/vikunja-44-fallback-volume-copy.txt b/documentation/audits/update-rulings-2026-09-23/vikunja-44-fallback-volume-copy.txt new file mode 100644 index 00000000..44c1b764 --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/vikunja-44-fallback-volume-copy.txt @@ -0,0 +1,13 @@ +drwxr-xr-x root/root 0 2026-09-23 06:21 ./ +-rw-r--r-- root/root 2245432 2026-09-23 06:21 ./vikunja.db-wal +-rw-r--r-- root/root 32768 2026-09-23 06:21 ./vikunja.db-shm +-rw-r--r-- root/root 4096 2026-09-23 06:21 ./vikunja.db +-rw-r--r-- 1 root root 4096 Sep 23 06:21 vikunja.db +-rw-r--r-- 1 root root 32768 Sep 23 06:21 vikunja.db-shm +-rw-r--r-- 1 root root 2245432 Sep 23 06:21 vikunja.db-wal +volume put back from the tier copy + start: 0.86s +08:25:01 vikunja: readback of the seeded project http=200 ok=True +08:25:01 vikunja B: attachments list unreadable http=404 {"code":4002,"message":"This task does not exist"} + +08:25:01 vikunja B: project readback=False attachment content readback=False +08:25:01 READBACK from the TIER copy (the only copy this app has): A=True B={'project': False, 'attachment': False} <- B was written after the backup diff --git a/documentation/audits/update-rulings-2026-09-23/walk.py b/documentation/audits/update-rulings-2026-09-23/walk.py new file mode 100644 index 00000000..9b088d0b --- /dev/null +++ b/documentation/audits/update-rulings-2026-09-23/walk.py @@ -0,0 +1,493 @@ +#!/usr/bin/env python3 +"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints. + +EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses: + POST /api/stacks//deploy · POST /api/backup/run · POST /api/sync · POST /api/stacks/rescan + POST /api/stacks//update · POST /api/stacks//remove +and reads GET /api/stacks/. No controller code exists for it. + +The walk, per `09` §6.4 and the update-night brief §4: + 1 deploy from the DRILL catalog at the LIVE pin + 2 seed through the app's OWN front door (R-156: never a volume, never SQL) + 3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing + 4 „Mentés most" + 5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages + 6 press the guarded Update, record every phase with timestamps + 7 read the seed back through the front door + 8 the four version observables side by side + 9 write the verdict record in `09`'s JSON shape + +`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`. +""" +import argparse, json, os, re, subprocess, sys, time +from datetime import datetime, timezone + +SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad" +EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-rulings-2026-09-23" +DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill" +BASE = "https://192.168.0.114" +HOSTHDR = "Host: felhom.enkisfelhom.hu" +DOMAIN = "enkisfelhom.hu" +HP = "demo-hp" + +LOG = [] + + +def say(*a): + line = " ".join(str(x) for x in a) + ts = datetime.now().strftime("%H:%M:%S") + print(f"{ts} {line}", flush=True) + LOG.append(f"{ts} {line}") + + +def sh(args, timeout=300, inp=None): + try: + return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp) + except (subprocess.TimeoutExpired, OSError) as e: + return subprocess.CompletedProcess(args, 124, "", f"{e}") + + +def guest(script, timeout=600): + """Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting).""" + r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP, + "cat > /tmp/w.sh; pct push 9202 /tmp/w.sh /tmp/w.sh >/dev/null 2>&1; " + "pct exec 9202 -- bash /tmp/w.sh; rm -f /tmp/w.sh"], + timeout=timeout, inp=script) + return r.stdout or "" + + +def login(): + pw = open(f"{SC}/.ctlpw").read().strip() + sh(["curl", "-sk", "-D", f"{SC}/hdr.txt", "-o", "/dev/null", "-H", HOSTHDR, + "-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"]) + h = open(f"{SC}/hdr.txt").read() + m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I) + if not m: + sys.exit("login failed: no session cookie") + open(f"{SC}/sess.txt", "w").write(m.group(0)) + r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"]) + c = re.search(r'/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN. + + Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because + a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password + (grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only + reason this was safe to discover by running it (live-probes rule). + + A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on + the scratch drive first — the same act the drive browser performs for a household. + """ + code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields") + fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or [] + values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub} + made = [] + for f in fields: + ev, ty = f.get("env_var"), f.get("type") + if ev in values: + continue + # `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses + # when the caller sends none, deliberately ("the user needs to know their password"), + # while `.felhom.yml` declares `required: false` and the API serves that verbatim. A + # caller that trusts the contract gets a 400. Measured tonight on grafana; filed. + if not f.get("required") and ty != "password": + continue # the controller generates the optional secrets itself + if ty == "path": + p = f"{DRIVE}/{name}" + values[ev] = p + made.append(p) + elif ty in ("secret", "password"): + import secrets as _s + values[ev] = "Drill-" + _s.token_hex(12) + GENERATED.setdefault(name, {})[ev] = values[ev] + elif f.get("default"): + values[ev] = f["default"] + else: + values[ev] = f"drill-{name}" + if made: + guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made)) + say(f" [1] made the drive paths this app requires: {made}") + extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")] + if extra: + say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}") + return values + + +def deploy(name, sub, extra_values=None): + st = stack(name) + if st.get("deployed"): + say(f" [1] {name} already deployed — reusing") + return True + values = deploy_values(name, sub) + if extra_values: + values.update(extra_values) + code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values}) + say(f" [1] deploy -> {code} {str(d)[:120]}") + if code != "202": + return False + # WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the + # container `healthy` while the controller's own state read `unhealthy` — a gate on `running` + # alone therefore times out on an app that is up. The state is RECORDED rather than required; + # the real gate is the fixture's own `wait_app`, which asks whether the APP answers. + seen = None + for _ in range(90): + time.sleep(5) + st = stack(name) + seen = st.get("state") + # `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21: + # tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm + # read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the + # deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark + # (`runComposeDeploy` writes it), so that is what to wait for. + pins = (st.get("app_config") or {}).get("pinned_images") + if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"): + say(f" [1] deployed, controller state={seen}, " + f"pinned={(st.get('app_config') or {}).get('pinned_images')}") + if seen != "running": + say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, " + f"not treated as a failure; the fixture's front-door wait is the real gate") + return True + say(f" [1] never became deployed (last controller state={seen!r})") + return False + + +def backup_now(name): + code, d = ctl("POST", "/api/backup/run") + say(f" [4] „Mentés most\" -> {code} {str(d)[:160]}") + for _ in range(90): + time.sleep(5) + c2, s = ctl("GET", "/api/backup/status") + dd = s.get("data") or {} + if not dd.get("running", False): + say(f" [4] backup idle; last={dd.get('last_run') or dd.get('last_db_dump')}") + return True + say(" [4] backup still running after 7.5 min — carrying on") + return False + + +def drill_bump(app, frm, to, service_hint=None): + """Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates). + + `frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in + two images (adventurelog's backend and frontend) moves both in one edge, while its engine + sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move + every service that carries the app's own version and no others. + """ + comp = f"{DRILL}/templates/{app}/docker-compose.yml" + fy = f"{DRILL}/templates/{app}/.felhom.yml" + s = open(comp).read() + froms = [x.strip() for x in frm.split(",") if x.strip()] + tos = [x.strip() for x in to.split(",") if x.strip()] + if len(froms) != len(tos): + say(f" [5] from/to lists differ in length: {froms} vs {tos}") + return None + for f1, t1 in zip(froms, tos): + if f"image: {f1}" not in s: + say(f" [5] FROM ref not found in compose: {f1}") + return None + s = s.replace(f"image: {f1}", f"image: {t1}") + open(comp, "w").write(s) + f = open(fy).read() + today = datetime.now().strftime("%Y-%m-%d") + f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M) + open(fy, "w").write(f) + sh(["git", "-C", DRILL, "add", "-A"]) + sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"]) + r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120) + h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip() + say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})") + return h + + +def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5): + """Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP. + + R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and + `catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not + enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the + Update that followed moved nothing and still reported "Frissitve". So when the caller knows + which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the + NUMBER R-607 asks for and has never had. + """ + t0 = time.time() + ctl("POST", "/api/sync") + time.sleep(2) + ctl("POST", "/api/stacks/rescan") + time.sleep(2) + if not expect_app or not expect_ref: + return None + for i in range(tries): + cat = stack(expect_app).get("catalog_images") or {} + if expect_ref in cat.values(): + waited = round(time.time() - t0, 1) + if i: + say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up " + f"to {expect_ref} — R-607's window, measured") + return waited + time.sleep(delay) + ctl("POST", "/api/sync") + time.sleep(1) + ctl("POST", "/api/stacks/rescan") + say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — " + f"catalog_images = {stack(expect_app).get('catalog_images')}") + return None + + +def badges(name): + out = {} + for lang, suffix in (("hu", ""), ("en", "?lang=en")): + h = page(f"/apps/{name}{suffix}") + m = re.findall(r']*title="([^"]*)"[^>]*>([^<]*)<', h) + out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3] + return out + + +def press_update(name, poll=1.0, cap_s=1800): + code, d = ctl("POST", f"/api/stacks/{name}/update") + say(f" [6] Update -> {code} {str(d)[:220]}") + if code not in ("202", "200"): + return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0} + phases, seen, t0 = [], None, time.time() + while time.time() - t0 < cap_s: + st = stack(name) + ph = st.get("update_phase") + if ph != seen: + seen = ph + rec = {"t": round(time.time() - t0, 1), "phase": ph, + "label": st.get("update_phase_label"), "updating": st.get("updating"), + "error": st.get("update_error"), "hold": st.get("hold_reason")} + phases.append(rec) + say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} " + f"err={rec['error']} hold={rec['hold']}") + if not st.get("updating") and ph in ("done", "failed", None) and time.time() - t0 > 3: + break + time.sleep(poll) + st = stack(name) + return {"accepted": True, "http": code, "phases": phases, + "duration_s": round(time.time() - t0, 1), + "final_phase": st.get("update_phase"), "update_error": st.get("update_error"), + "hold_reason": st.get("hold_reason"), "state": st.get("state")} + + +def observables(name): + st = stack(name) + ac = st.get("app_config") or {} + live = guest(f""" +grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//' +echo '---inspect---' +for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do + echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}' +done +""") + a, _, b = live.partition("---inspect---") + return { + "pinned_images": ac.get("pinned_images"), + "installed_images": {k: (v.get("ref") if isinstance(v, dict) else v) + for k, v in (ac.get("installed_images") or {}).items()}, + "catalog_images": st.get("catalog_images"), + "live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()], + "docker_inspect": [x for x in b.strip().splitlines() if x.strip()], + } + + +def app_logs(name, lines=400): + """The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is + one string with escaped newlines — a scan over the envelope sees a single enormous line and + finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96 + rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines.""" + code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}") + if isinstance(d, dict): + data = d.get("data") + if isinstance(data, dict) and isinstance(data.get("logs"), str): + return data["logs"] + if isinstance(d.get("_raw"), str): + return d["_raw"] + return str(d) + + +def write_verdict(rec, appdir): + os.makedirs(appdir, exist_ok=True) + p = os.path.join(appdir, "verdict.json") + json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False) + say(f" [9] verdict {rec['verdict']} -> {p}") + + +def remove(name): + """Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint + refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up.""" + c1, d1 = ctl("POST", f"/api/stacks/{name}/stop") + say(f" [X] stop -> {c1} {str(d1)[:100]}") + for _ in range(24): + time.sleep(5) + if stack(name).get("state") != "running": + break + code, d = ctl("POST", f"/api/stacks/{name}/remove", + {"remove_hdd_data": True, "remove_backups": True}) + say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}") + if code == "409": + # R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive + # path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202 + # `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits + # this. The household's other choice — remove the app, KEEP the data — is accepted, and the + # harness takes it, then tidies its own directory by name at teardown. + say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)" + " — removing the app and KEEPING the drive data instead") + code, d = ctl("POST", f"/api/stacks/{name}/remove", + {"remove_hdd_data": False, "remove_backups": True}) + say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}") + time.sleep(5) + st = stack(name) + left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; " + f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'") + say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}") + return code + + +def app_env(name, key): + """Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the + app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller + shows them the same value. Data still goes in through the app's own front door.""" + out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1") + if ":" in out: + return out.split(":", 1)[1].strip().strip('"').strip("'") + return "" + + +def snapshots(name): + """The restorable copies the backups page offers for this app.""" + code, d = ctl("GET", f"/api/backup/snapshots?stack={name}") + data = d.get("data") if isinstance(d, dict) else None + if isinstance(data, dict): + for k in ("snapshots", "items", "restore_points"): + if isinstance(data.get(k), list): + return data[k] + return data if isinstance(data, list) else [] + + +def restore(name, snapshot_id=None, wait_s=1200): + """The household's own way out: the „Visszaállítás a mentésből" button on the backups page. + + A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`, + `snapshot_id` — because that is the button the sentence tells them to press. + """ + snaps = snapshots(name) + if snapshot_id is None: + if not snaps: + say(f" [R] no restorable copy offered for {name}") + return {"ok": False, "why": "no snapshot offered", "snapshots": snaps} + first = snaps[0] + snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id") + say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)") + sess = open(f"{SC}/sess.txt").read().strip() + csrf = open(f"{SC}/csrf.txt").read().strip() + r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}", + "-X", "POST", + "--data-urlencode", f"_csrf={csrf}", + "--data-urlencode", f"stack_name={name}", + "--data-urlencode", f"snapshot_id={snapshot_id}", + f"{BASE}/backup/restore"], timeout=180) + head = (r.stdout or "").split("\n")[0].strip() + loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")] + say(f" [R] POST /backup/restore -> {head} {loc[:1]}") + t0 = time.time() + last = None + while time.time() - t0 < wait_s: + code, d = ctl("GET", "/api/backup/restore-status") + dd = d.get("data") or {} + cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message")) + if cur != last: + say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}") + last = cur + if not dd.get("running", False) and time.time() - t0 > 5: + break + time.sleep(2) + st = stack(name) + say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} " + f"phase={st.get('update_phase')}") + return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps, + "http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1), + "state_after": st.get("state"), "hold_after": st.get("hold_reason"), + "observables_after": observables(name)} diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index c5b1388d..beb68f6a 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -666,14 +666,14 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** | | **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | | **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** | -| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` **MEASURED 2026-09-21, and the blind spot is not one or two pins.** `audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3: the catalog carries **10 floating pins of 66** (recounted — the old "23" was stale), and **6 of the 7 measurable engine pins have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`; only `mariadb:11.6` has not. The 8th (immich's own ghcr build) is UNMEASURED — ghcr exposes no anonymous last-modified timestamp. **So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.** The fix does NOT need the box to query a registry: the catalog can record each pin's digest at push time (`check-image-resolvable.py` already resolves it) and the box compares digests. Put to the operator as `09` §3b **Q6**, recommended YES — the cheapest real improvement on the arc's list. **— UPDATE NIGHT 2026-09-21:** **MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it REFINES the row in two ways rather than merely confirming it.** §8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins were read as `installed_images` records them and compared against the upstream digests measured the same night: `postgres:16-alpine` → **`sha256:721873c34ceb9…` on the box and `sha256:721873c34ceb9…` upstream**, and `redis:7-alpine` → **`sha256:858f009f9709c…` both sides**. **Identical. So the badge „Naprakész" is TRUE for this box**, and the app reads correctly. **(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for a box that pulled BEFORE the tag moved; a box deployed after the repush holds the current image and its badge is right. R-446's "six repushed pins" measured the tag against the date the CATALOG set it, which is the right measure for *the catalog* and not for *a box*. **(2) The producer Q6 needs ALREADY EXISTS on the box.** `installed_images` records a real `digest` per service (`installed.go` §7.1) — the box knows exactly what it is running. What it cannot do is COMPARE, because the catalog carries no digest to compare against. That is Q6's proposal, and this is a concrete confirmation that only the catalog half is missing. Evidence: `audits/update-night-2026-09-21/23-B8-floating-pin.txt`. **-- RULED 2026-09-23 (`09` §3 decision 17):** YES — the catalog records the image digest of every pin at push time; the box compares against it and, where the catalog carries one, pulls **that exact image**, which makes a floating tag reproducible, not only the badge honest. *Pull-by-digest while the definition names a tag is a claim to verify in the build, not a ruling on mechanism.* | **READY TO BUILD — owner: CC; `09` §6.4** | -| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 **HALF SHIPPED 2026-09-21 (catalog `5ff36d098cbc`): the second half — an engine change gets its OWN edge — is now ENFORCED** by `check-engine-major.py`, which refuses a commit moving a MariaDB major together with any other image move in that template, naming what it was bundled with. The FIRST half (automatic within a major) is Slice 6 and needs four operator answers — `09` §3b **Q1–Q4**, with the shape it would take in `09` §6.2. **The urgency is now measured:** 46 of the catalog's 58 exact pins are behind upstream and **39 of those are within a major** — the population the 2026-09-02 ruling already says may move without a human. **-- RULED 2026-09-23 (operator, `09` §3 decisions 11–15):** Q1–Q4 answered. The update is a leg of the backup chain after off-site and before the full-system backup (11); automatic with a per-box switch ON by default (12); **the TEST decides, not the tag** — the box applies every step the catalog holds because the catalog holds only tested steps, and `CompareImageRefs` moves to the catalog gate (13, REPLACES "never across a major"); a box behind climbs **one tested step at a time** (14); **the box UNDOES a failed update itself** — old definition + the pre-pin safety dump + health check again, HOLD only if the undo fails (15, REPLACES §6.1's no-auto-undo). The undo and the ladder were SPIKED the same day before any build (`audits/update-rulings-2026-09-23/`); build order and costs in `09` §6.4. | **READY TO BUILD — owner: CC; `09` §6.4 part by part, each part returns to the operator for go/no-go** | +| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` **MEASURED 2026-09-21, and the blind spot is not one or two pins.** `audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3: the catalog carries **10 floating pins of 66** (recounted — the old "23" was stale), and **6 of the 7 measurable engine pins have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`; only `mariadb:11.6` has not. The 8th (immich's own ghcr build) is UNMEASURED — ghcr exposes no anonymous last-modified timestamp. **So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.** The fix does NOT need the box to query a registry: the catalog can record each pin's digest at push time (`check-image-resolvable.py` already resolves it) and the box compares digests. Put to the operator as `09` §3b **Q6**, recommended YES — the cheapest real improvement on the arc's list. **— UPDATE NIGHT 2026-09-21:** **MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it REFINES the row in two ways rather than merely confirming it.** §8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins were read as `installed_images` records them and compared against the upstream digests measured the same night: `postgres:16-alpine` → **`sha256:721873c34ceb9…` on the box and `sha256:721873c34ceb9…` upstream**, and `redis:7-alpine` → **`sha256:858f009f9709c…` both sides**. **Identical. So the badge „Naprakész" is TRUE for this box**, and the app reads correctly. **(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for a box that pulled BEFORE the tag moved; a box deployed after the repush holds the current image and its badge is right. R-446's "six repushed pins" measured the tag against the date the CATALOG set it, which is the right measure for *the catalog* and not for *a box*. **(2) The producer Q6 needs ALREADY EXISTS on the box.** `installed_images` records a real `digest` per service (`installed.go` §7.1) — the box knows exactly what it is running. What it cannot do is COMPARE, because the catalog carries no digest to compare against. That is Q6's proposal, and this is a concrete confirmation that only the catalog half is missing. Evidence: `audits/update-night-2026-09-21/23-B8-floating-pin.txt`. **-- RULED 2026-09-23 (`09` §3 decision 17):** YES — the catalog records the image digest of every pin at push time; the box compares against it and, where the catalog carries one, pulls **that exact image**, which makes a floating tag reproducible, not only the badge honest. *Pull-by-digest while the definition names a tag is a claim to verify in the build, not a ruling on mechanism.* **-- 2026-09-23:** pull-by-digest MEASURED on 9202 — Docker and Compose both pull and run `redis:7-alpine@sha256:858f…` and refuse a digest that does not exist. **Build trap, read from source:** `splitImageRef` returns "unorderable" for any ref containing `@` (`stacks/updateorder.go:134`), so a digest-carrying pin must have its digest split off before ordering or every such app reads Unknown. `09` §6.4 part 6. | **READY TO BUILD — owner: CC; `09` §6.4** | +| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 **HALF SHIPPED 2026-09-21 (catalog `5ff36d098cbc`): the second half — an engine change gets its OWN edge — is now ENFORCED** by `check-engine-major.py`, which refuses a commit moving a MariaDB major together with any other image move in that template, naming what it was bundled with. The FIRST half (automatic within a major) is Slice 6 and needs four operator answers — `09` §3b **Q1–Q4**, with the shape it would take in `09` §6.2. **The urgency is now measured:** 46 of the catalog's 58 exact pins are behind upstream and **39 of those are within a major** — the population the 2026-09-02 ruling already says may move without a human. **-- RULED 2026-09-23 (operator, `09` §3 decisions 11–15):** Q1–Q4 answered. The update is a leg of the backup chain after off-site and before the full-system backup (11); automatic with a per-box switch ON by default (12); **the TEST decides, not the tag** — the box applies every step the catalog holds because the catalog holds only tested steps, and `CompareImageRefs` moves to the catalog gate (13, REPLACES "never across a major"); a box behind climbs **one tested step at a time** (14); **the box UNDOES a failed update itself** — old definition + the pre-pin safety dump + health check again, HOLD only if the undo fails (15, REPLACES §6.1's no-auto-undo). The undo and the ladder were SPIKED the same day before any build (`audits/update-rulings-2026-09-23/`); build order and costs in `09` §6.4. **-- SPIKED 2026-09-23 (`audits/update-rulings-2026-09-23/`):** the undo works by hand on three real migrating edges and needs eight product additions (R-637..R-642); **the ladder is measured absent** — one press on a box two steps behind jumped vikunja 2.3.0 → 2.5.0 in 9.5 s and 2.4.0 never ran, and the box cannot see intermediate steps at all because its catalog clone is `--depth 1` (`sync.go:283`/`:300`, one commit visible on both demo guests). The ladder's recommended format is an `update_ladder:` list in `.felhom.yml` with each intermediate step's own definition, NOT the git history (romm's image-moving commit is the definition that OOM-looped). The chain's update leg has ≤15 min as ruled (R-643). Build order `09` §6.4. | **READY TO BUILD — owner: CC; `09` §6.4 part by part, each part returns to the operator for go/no-go** | | **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 **BOTH SIDES VERIFIED 2026-09-21, and it is cheaper than this row implies.** The controller's payload carries no image (`internal/report/types.go` L98–103) and the hub's `Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts** — but **the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent. What is missing is the denormalisation and the page, not the transport. Shape in `09` §6.2–6.3; the payload question is `09` §3b **Q7**. **-- RULED 2026-09-23 (`09` §3 decision 18):** the report carries, per compose service (database included), the installed reference, the catalog reference and the badge state. **Built later, when the fleet grows** — Q7's recommendation, confirmed. | **RULED — build deferred until the fleet grows; owner: CC** | | **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** | | **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** | | **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 **— UPDATE NIGHT 2026-09-21:** **MEASURED 2026-09-21 (update night), leg B9, and the row's risk is NARROWER than it states.** A `.felhom.yml`-only change (a health check for a path only a newer version would serve) was pushed to a FROZEN `bentopdf` — installed `v2.8.6`, catalog ahead. §5.4's asymmetry is confirmed live: the new `.felhom.yml` reached the box while the compose `image:` line stayed `v2.8.6`. **But no false alarm was produced**: ten samples over two minutes all read `state=running` with the front door at `200`. The reason is the probe's own semantics, not luck — `healthprobe.go:258-261` treats **any response** as healthy for `type: http`, and the bogus path answers 404, which is a response. **So this row's false-alarm risk exists only for `type: api` probes carrying an `expect` block**, where the status is compared; for every `type: http` template and every `type: api` without `expect`, a newer version's path is invisible to the probe. The row's actual claim — the failure direction is a false alarm, never data loss — stands and is now measured. Evidence: `audits/update-night-2026-09-21/21-B9-frozen-app-newer-felhomyml.md`. | **READY — rank P3-LOW; owner: CC** | | **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 **-- UPDATE NIGHT 2026-09-21:** bookstack's edge was walked again on 2026-09-21 and is again **half-proven**: the database half read back through `php artisan` with its own negative control, the file half untouched. The limitation is unchanged and is now measured on the box as well as on the harness. Two more apps joined the same class tonight for a different reason (R-624). | **READY — rank P3-LOW; owner: CC** | -| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 **ROW CORRECTED 2026-09-21: the scope is NOT open and this row said it was.** It read *"VIKTOR rules on scope"*; the operator ruled on 2026-09-13 (`09` §3 decision 6) that the upgrade test goes to **ALL** apps through the nightly rotation, explicitly *not* "database apps first". What is open is the WORK, not the scope. An ORDER inside that ruling — the 15 database services first, because that is where a wrong answer costs data rather than uptime — is costed as a drill brief in `09` §6.4: legs A–E ≈ **21–34 CC-hours** plus ~25–30 GB on a scratch host, with legs C (a PostgreSQL `pg_upgrade` rehearsal) and E (one automatic night on a throwaway) the two that unblock a decision. **-- UPDATE NIGHT 2026-09-21:** **The count moved from 3 apps to 21 EDGES ACROSS 19 APPS.** The update night walked real within-a-major upstream edges on scratch guest 9202 through the product's own guarded Update, each app seeded and read back through its OWN front door with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive**. Proven: actualbudget, audiobookshelf, bookstack, docmost, grafana, home-assistant, mealie, n8n, navidrome, papra, privatebin, romm, vikunja, and nextcloud's MariaDB engine major. **Ten of the fourteen printed a verbatim migration line**, so the database really was rewritten and the data still read back. Box-side fixtures for 20 apps now exist at `audits/update-night-2026-09-21/fixtures.py`, and four (actualbudget, navidrome, audiobookshelf, vikunja) are ported into `app-catalog-felhom.eu/scripts/upgrade_fixtures.py` with seven new EDGES (U1-U7) so the same edges can be run on the harness venue **with their ABORT step**, which the box deliberately does not offer. **OWED, stated so it is not mistaken for done:** the U1-U7 harness RUNS (the code is in, the runs are not), and fixtures for the four inconclusive apps, of which two (vaultwarden, zipline) cannot be seeded at all while the catalog rightly closes their sign-up (see R-624). | **READY — rank P2-MEDIUM; owner: CC (scope already ruled, `09` §3 decision 6)** | +| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 **ROW CORRECTED 2026-09-21: the scope is NOT open and this row said it was.** It read *"VIKTOR rules on scope"*; the operator ruled on 2026-09-13 (`09` §3 decision 6) that the upgrade test goes to **ALL** apps through the nightly rotation, explicitly *not* "database apps first". What is open is the WORK, not the scope. An ORDER inside that ruling — the 15 database services first, because that is where a wrong answer costs data rather than uptime — is costed as a drill brief in `09` §6.4: legs A–E ≈ **21–34 CC-hours** plus ~25–30 GB on a scratch host, with legs C (a PostgreSQL `pg_upgrade` rehearsal) and E (one automatic night on a throwaway) the two that unblock a decision. **-- UPDATE NIGHT 2026-09-21:** **The count moved from 3 apps to 21 EDGES ACROSS 19 APPS.** The update night walked real within-a-major upstream edges on scratch guest 9202 through the product's own guarded Update, each app seeded and read back through its OWN front door with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive**. Proven: actualbudget, audiobookshelf, bookstack, docmost, grafana, home-assistant, mealie, n8n, navidrome, papra, privatebin, romm, vikunja, and nextcloud's MariaDB engine major. **Ten of the fourteen printed a verbatim migration line**, so the database really was rewritten and the data still read back. Box-side fixtures for 20 apps now exist at `audits/update-night-2026-09-21/fixtures.py`, and four (actualbudget, navidrome, audiobookshelf, vikunja) are ported into `app-catalog-felhom.eu/scripts/upgrade_fixtures.py` with seven new EDGES (U1-U7) so the same edges can be run on the harness venue **with their ABORT step**, which the box deliberately does not offer. **OWED, stated so it is not mistaken for done:** the U1-U7 harness RUNS (the code is in, the runs are not), and fixtures for the four inconclusive apps, of which two (vaultwarden, zipline) cannot be seeded at all while the catalog rightly closes their sign-up (see R-624). **-- 2026-09-23, the RomM lesson is IN THE HARNESS:** `upgrade-test.py` v2 runs every edge that read back under light load for `--soak` seconds (default 600) and reads the kernel's `oom_kill` counter host-side; a kill or restart turns `proven` into `failed`, a peak over 80 % of the limit adds `memory_tight`. **Red-proof:** romm 5.0.0 → 5.3.0 on the template AS PROMOTED (512M, four workers) — seeded, migrated, read back, and then **OOM-killed at +76 s** under four light callers, verdict `failed` (Docker's own OOMKilled read true here; restarts stayed 0, which is why the walk never saw it). Ten minutes is ample for this failure; demo-hp's first kill came at two hours only because nothing was loading it. Evidence `audits/update-rulings-2026-09-23/harness/`. | **READY — rank P2-MEDIUM; owner: CC (scope already ruled, `09` §3 decision 6)** | | **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 **-- UPDATE NIGHT 2026-09-21:** **Measured 2026-09-21, both halves.** (a) What a household sees TODAY: the guarded Update of `postgres:16-alpine` to `17-alpine` on docmost ended **`failed` in 5.1 s**, the app stopped and held, **the pin naming 17 while `installed_images` still said 16 and nothing ran**, the data intact, and the restore the hold sentence names back in **29.1 s**. The engine's refusal line had to be REPRODUCED independently because `failAndHold` destroyed it (R-621): *FATAL: database files are incompatible with server / DETAIL: The data directory was initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir was still `16` afterwards, and the same copy started under 16 holding 48 tables as the positive control. (b) The conversion **COSTED** on a real seeded 49 MB / 48-table datadir: `pg_dumpall` **2.6 s / 132 201 B**, fresh 17 plus replay **6.5 s / 48 tables restored**, the app up on 17 saying *Database connection successful*, **the seeded account read back**, total **155.9 s of which ~9 s is engine work**. `pg_upgrade` was NOT run: it needs both majors' binaries in one image and no such image exists here. Full paragraph: `audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`. **-- RULED 2026-09-23 (`09` §3 decision 16):** PostgreSQL majors are converted BY THE BOX as a guarded-update step — save everything from the old engine, start the new one empty, load it back, check. Each of the eleven apps is proven on the test bench before the catalog may move it; the engine-major gate stays until then. | **READY TO BUILD — owner: CC; `09` §6.4; the gate stays until all eleven are proven** | | **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** | | **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** | @@ -806,6 +806,14 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-634** | **[P1-HIGH] An app can be RUNNING, HEALTHY and serving while the controller records it as not deployed — and in that state the household cannot remove it through the product at all.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, on **two independent apps in one night**: `outline` and `sparkyfitness`. **The controller's own words, in order.** `outline` deploy accepted 11:54:47; **`11:55:51 StopStack outline: current state=deploying deployed=true containers=0`** — a stop while the stack is still deploying; `11:56:09 SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, **0 encrypted**, 3 sensitive fields` (the two saves before it both read `3 encrypted`); then **`11:57:16` and `12:02:16 Health probe outline: API GET :3000/_health -> 200`** — the app is up and answering its own health endpoint; and **`12:02:19 StopStack outline: current state=running deployed=false containers=3`**. Three containers, state `running`, health 200, **`deployed=false`**. The remove then answers **`RemoveStack outline: state=not_deployed, deployed=false, orphaned=false, deploying=false` -> `[ERROR] Remove failed for outline: stack "outline" is not deployed`** for BOTH the remove-with-data and the remove-keeping-data call, while `ScanStacks` goes on finding the stack every ten seconds. `app.yaml` survives with `desired_state: stopped` and an **empty `installed_images`**. `sparkyfitness` produced the identical shape 13 minutes earlier. **WHAT THIS COSTS A HOUSEHOLD:** an app that works is invisible to the product as an installation — no badge, no update, no backup selection, and **no way to delete it**; the only exit is a shell. It is the mirror of R-633 (there the record is gone and the container remains; here the container is fine and the record is gone) and it is the **worse** of the two, because the app is serving customer traffic the whole time. **THE MECHANISM IS NOT DIAGNOSED, and this row says so rather than guessing.** What was tried: the controller's full container log for both apps (the sequence above), `app.yaml` on disk, `docker ps -a`, and `GET /api/stacks/`. What was NOT done: reading `runComposeDeploy`'s pin-write path — this was an unattended run and the brief forbade product code. **The one discriminator worth running first:** both apps were walked while two other walks ran concurrently, and `POST /api/backup/run` is box-wide, so a backup or restore for a NEIGHBOURING app was in flight. A serial re-walk is queued tonight; if it reproduces alone, concurrency is not the cause. **A THIRD THING THE SAME LOG SHOWS, recorded here because it is one line away:** at `11:55:57` the health probe dialled **`http://outline-postgres:3000/_health`** — during startup, with the exactly-named container not yet running, `findProbeContainer`'s PREFIX fallback latched onto the POSTGRES sidecar and probed port 3000 on it. Transient, and it resolved once `outline` came up, but it is the same function R-630 is about. Evidence: `audits/the-28-2026-09-22/apps/half-state-outline-sparkyfitness.txt`. **THE SERIAL RE-WALK WAS RUN THE SAME NIGHT AND IT SPLITS THIS ROW IN TWO — recorded here rather than left as the first reading.** Walked again one at a time, with no other walk running: **`outline` deployed normally and removed clean**, and **`crafty-controller` deployed normally, updated `4.10.7 → 4.11.0` to `done`, restored and removed clean.** So for those two the half-state did NOT reproduce alone, and concurrency — a box-wide `POST /api/backup/run` or a restore in flight for a NEIGHBOURING app — is implicated rather than the deploy path itself. **`sparkyfitness` reproduced EXACTLY, alone, in 534 s**: deploy accepted, never reached `deployed` with a pin, `app.yaml` left with `desired_state: stopped` and an empty `installed_images`, and **both remove calls refused with `stack "sparkyfitness" is not deployed`.** **So the row stands, at one reproducible app instead of three, and the honest split is:** (a) `sparkyfitness` has a deploy that does not finish and leaves a record the product cannot clear — reproducible, P1; (b) under concurrent work the same unremovable half-state can be reached by apps that are otherwise fine, which is the more alarming half because those apps were **running, healthy and serving** while recorded as not deployed. **Neither half is diagnosed** — `runComposeDeploy`'s pin write was not read, because the brief forbade product code. **THE BOUNDED HALF IS FIXED in v0.262.0; the MECHANISM IS STILL NOT DIAGNOSED, and this row stays open for it.** `RemoveStack` refused on `!stack.Deployed` — a FLAG — while the machine plainly had containers, a compose file and an `app.yaml`. It now asks whether anything EXISTS (`halfStateEvidence`): containers, a compose file, or an `app.yaml` are each enough, and it removes what exists with every other fence unchanged (R-442 stays fail-closed). **The household must always be able to remove what the box shows them** — that is true whatever the cause of the bad record. Red-proof seen failing. **What is NOT done:** why `deployed` stays false while containers run. `runComposeDeploy`'s pin-write path was not read against a fresh reproduction in this session. **THE UNREMOVABLE HALF IS PROVEN LIVE ON 9202:** `sparkyfitness` rebuilt in its exact measured shape — `app.yaml` and a compose file on disk, no containers, `deployed: false`, `state: not_deployed` — and the remove answered **200** with `leftovers: NONE`. Under v0.261.0 the identical call answered `stack "sparkyfitness" is not deployed`. | **OPEN — P1 for the MECHANISM only; the unremovable half is CLOSED in v0.262.0 and proven live** | | **R-635** | **[P1-HIGH] `romm 5.3.0` does not fit the memory the template gives it, and the guarded Update called that a success — the app has been OOM-crash-looping on demo-hp for six hours at ~500% CPU.** FOUND 2026-09-22 17:37 **because the operator heard the fans**, which is the only reason it was found at all. `romm` was promoted `5.0.0 -> 5.3.0` on the live catalog that morning (`audits/PROBE-FIX-2026-09-22.md`) after the edge was PROVEN on scratch guest 9202, and the guarded Update was then pressed on demo-hp guest 9201 at **09:08:22Z**, reaching **`done` in 74.8 s** with the app `running`. **It ran clean for two hours.** The first worker kill is at **11:09:20Z**; by 15:38Z there had been **4,530** of them — `Worker (pid:…) was sent SIGKILL! Perhaps out of memory?` — with `docker inspect` reading **`OOMKilled: true`**, the container pinned at **457 MiB of its 512 MiB limit**, and `docker stats` showing **499.51% CPU**. The host's load average sat at **5.2 while otherwise idle**. `rq_cron` is killed and restarted every few seconds in a permanent storm. **THREE THINGS THIS ESTABLISHES, and the third is the one that changes how promotions are judged.** (1) The template's own comment says *`RAM: ~300MB (mem_limit: 1024M total — romm 512M + mariadb 384M + redis 128M)`*; **5.3.0 needs more than 512M and the template was not re-sized when the version moved.** A version move is not only an `image:` line. (2) **The update reported `done` and the app reads `running`**, because nginx answers `GET /` with 200 while the gunicorn workers behind it are being killed — a THIRD variant of the R-618/R-630 theme: the probe is right, the port is right, and the answer is still a false green. (3) **THE ALARM DID FIRE, AND MY FIRST WRITE-UP OF THIS ROW SAID IT DID NOT — CORRECTED 2026-09-22 BY THE OPERATOR, WHO PRODUCED THE MAILS.** The controller HAS an OOM detector (`main.go:821`, `[deadapp] romm: container romm was OOM-killed`, evaluated every 30 s), it emits **`app_oom`** (`notifier.go:726`), the hub allow-lists it (`dispatcher.go:636`) and dispatched it to the OPERATOR channel as **SENT** — `admin@felhom.eu` received **`[Felhom] ⚠️ demo-hp: app_oom`** at **11:09 CEST** and again at **17:48 CEST**, each carrying the app name, the Hungarian sentence and the dashboard link. The CUSTOMER channel is **SKIPPED**, which is correct. **I asserted an absence without looking at the instrument** — the hub's own Events and Notifications tabs show all of it — and I did it by reasoning from a memory note (`lxc-docker-oom-signals-unreliable`, R-528) instead of reading the hub. **That is R-628's shape exactly, four days on, and from the same hand.** **WHAT IS ACTUALLY WRONG, and it is narrower and real:** `notifier.go:715-726` keys the alarm on `container|startedAt` in an `oomSeen` map and emits **once per container lifetime**. So **six hours of continuous thrashing — 4,530 worker kills — produced exactly ONE mail**, severity `warning`, never escalating, while the app went on reading `running`. **A six-hour storm is indistinguishable from a single transient kill.** The hub's App Telemetry panel did carry the magnitude (RomM: **5,023 errors, 632 warnings**, peak 855 MB) but nothing turns that into a second, louder signal. So the fix worth having is not a detector — there is one — but an ESCALATION: a warning that repeats for hours should stop looking like a warning that happened once. **AND THE LESSON FOR R-462's METHOD, which is the real cost:** every `proven` verdict in the update night and in the twenty-eight measures the app for the **minutes of the walk**, not for a day of running. `romm` passed its walk, was seeded, read back and restored — and broke two hours later. **`proven` currently means "the update applied and the data survived", NOT "the new version runs".** Needs: decide between rolling the catalog back to 5.0.0 and raising romm's `mem_limit` (measured, not guessed); and a soak longer than a walk before any future promotion. **FIXED AND MEASURED 2026-09-22, in two steps, and the FIRST step was still a guess.** *Step 1 (operator's choice):* the limit was raised 512M → **768M** (`app-catalog-felhom.eu@886956d`). It slowed the kills from ~12/min to ~7/min and **stopped nothing** — 37 SIGKILLs in five minutes, `OOMKilled` still true, and the cgroup's own `memory.peak` read **exactly 768 MiB**: it hit the new ceiling and died there. *Step 2, from a MEASUREMENT instead:* the per-process RSS inside the container reads **~216 MiB per warm uvicorn worker**, so the image's default of four workers plus the master needs **~882 MiB** before nginx and the job runner — more than any sensible limit for this box. **The lever was in the image all along:** `/init:143` runs `--workers "${WEB_SERVER_CONCURRENCY:-4}"`. **Four workers is a SERVER default on an appliance serving one household.** Setting **`WEB_SERVER_CONCURRENCY=2`** (`app-catalog-felhom.eu@f4eb94f`, limit left at 768M) and applying it through the product's own Update button fixed it. **PROVEN UNDER LOAD, not just at idle:** 6 concurrent callers driven at romm through the household's own route for 300 s — **26,645 requests** (9,687 × 200, 16,958 × 401 on the auth-gated endpoints), CPU a steady **~200%** (exactly two workers saturated, by design), memory oscillating **416–614 MiB against the 768 MiB limit and trending DOWN**, and **zero** SIGKILLs, `OOMKilled: false`, `RestartCount: 0` throughout. At idle afterwards: **1.64% CPU**, 610 MiB, host load falling from 5.2 to 2.4. **AND THE FIRST SOAK MEASURED NOTHING, which is worth more than the second one:** it was pointed at `arcade.enkisfelhom.hu` — the scratch-guest fixture's default subdomain — while this box deployed romm at `jatek`. Every request 404'd at traefik in 9 ms, romm sat idle at 0.64% CPU, and the counter cheerfully reported **14,026 successful requests**. It was caught only because 0.64% CPU under load is not believable. **A positive AND a negative control are now asserted before any load is driven** (`jatek` must not 404; a nonsense host must). R-96 rule 3, in a new surface. **WHAT STAYS OPEN, and it is the part that outlives romm:** 610 MiB of 768 MiB is **79%** — it works with ~158 MiB of headroom and the soak never exceeded 614 MiB, but it is not generous, and nothing watches it. **And the method lesson for R-462:** a version move is not only an `image:` line — the new version's SHAPE (worker counts, per-worker footprint) has to be measured too, and a walk lasting minutes cannot see a ceiling reached in two hours. Every `proven` verdict in the update night and in the twenty-eight means *"the update applied and the data survived"*, **not** *"the new version runs"*. Evidence: `audits/probe-fix-2026-09-22/romm-soak.json`, `romm-soak.out`. | **CLOSED 2026-09-22 — two workers, 768M, proven under 26,645 requests; the 79% headroom and the `proven`-means-minutes lesson are carried into R-462** | | **R-636** | **[P2-MEDIUM] An app that has been out of memory for six hours sends the same single warning an app that hiccuped once sends.** FOUND 2026-09-22 when the operator produced the alarm mails I had wrongly written off as absent (R-635). **The detector is correct and works.** `main.go:821` re-checks every 30 s and logs `[deadapp] romm: container romm was OOM-killed`; `notifier.go:726` emits **`app_oom`** at severity `warning`; the hub allow-lists it (`dispatcher.go:636`) and delivered it to the OPERATOR channel — two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard. CUSTOMER is SKIPPED, correctly. **The defect is the SHAPE of the signal, not its absence.** `notifier.go:715-724` keys on `container|startedAt` in an `oomSeen` map and returns early on a repeat, so the alarm fires **once per container lifetime**. romm's 09:08 container produced **4,530 worker kills over six hours and exactly one mail**; the 15:47 container produced one more. Severity never escalates, the app goes on reading `running`, and `08` §4 rightly does not put it in `IsDownState`. **So a six-hour storm that pins five cores is indistinguishable, in the operator's inbox, from one transient kill at 3 a.m.** — and the operator, who had two correct mails, still found the fault by hearing the fans. **The once-per-lifetime rule is right in itself** (it is what stops a crash loop from mailing 4,530 times, which is R-629's lesson) — what is missing is the second, louder signal when the same key keeps re-firing. **The magnitude IS already collected:** the hub's App Telemetry panel read **RomM: 5,023 errors, 632 warnings, peak 855 MB against a 1280M limit** while the same panel showed every other app at 0. Nothing turns that into an event. **Candidate shapes, none chosen here:** escalate `app_oom` to `error` when the same `container|startedAt` key re-fires past a threshold; or a periodic digest for a key still firing after N minutes; or let App Telemetry raise its own event when an app's error count crosses a bound. All are hub/controller code. Evidence: the operator's screenshots of the Events, Notifications and App Telemetry tabs plus the two mails; `felhom-controller/controller/internal/notify/notifier.go:715-726`. | **OPEN — P2; owner: CC; product code, so not fixed unattended** | +| **R-637** | **[P2-MEDIUM] BUILD THE UNDO (`09` §3 decision 15) — the eight things the 2026-09-23 spike says the product must add before a failed update can put itself back.** SPIKED BY HAND 2026-09-23 on 9202, three real migrating edges each made to fail a deliberately wrong probe: docmost (PostgreSQL) and romm (MariaDB) — whose OLD versions REFUSE the migrated data (*corrupted migrations …*, *Can't locate revision …*) — and vikunja (SQLite in a volume), whose old version starts. **The undo worked on all three**: data written before AND after the backup read back through each app's own front door, ≈16 s / ≈38 s / ≈1 s after the failing health wait. **The list:** (1) keep the pre-update copies — compose, applied, pin AND the old `.felhom.yml` — until the undo is over (`failAndHold` deletes the first three, the fourth was never kept, R-639); (2) the undo inside `failAndHold`: pin back → DB up alone → validated load → start → health with the OLD probe → `undone`, else HOLD; (3) empty-then-load atomically, because the loader the product has cannot undo a migration (R-638); (4) check the copy's completion marker first (R-640); (5) a volume copy at safety-dump time for apps with no database server (R-641); (6) files: nothing measured touched them — a step that does says so through decision 13's *files may change* mark; (7) the undo's success is the probe, never the Start's return (R-642); (8) remember the failed step so the caller never re-presses it. `09` §6.1a, §6.4 part 1 (4 evenings). Evidence: `audits/update-rulings-2026-09-23/README.md`. | **READY — owner: CC; `09` §6.4 part 1, returns to the operator for go/no-go** | +| **R-638** | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. | **OPEN — P2; owner: CC; measure the named restore after a real schema migration first** | +| **R-639** | **[P3-LOW] After a failed update is held, the previous definition survives only in the recovery unit — `failAndHold` deletes the journal's copies, and the old `.felhom.yml` was never kept.** READ FROM SOURCE and CONFIRMED 2026-09-23: `failAndHold` (`stacks/update.go:766`) ends with `clearJournal` + `removePreUpdateCopies`; in all three spike cases `ls /*pre-update*` counted 0 after the hold. The old definition was taken from the unit's `compose/` directory each time — present only because a backup preceded the update. The old `.felhom.yml` matters as much: its probe is the one the OLD version answers, and the catalog's copy flows to the app regardless (§5.4). Needed by R-637. Evidence: `audits/update-rulings-2026-09-23/README.md`, `docmost-41`, `romm-40`. | **OPEN — P3; owner: CC; lands with R-637** | +| **R-640** | **[P2-MEDIUM] A TRUNCATED PostgreSQL copy loads with exit 0 into an EMPTY database — and `ValidateDump` would accept it.** MEASURED 2026-09-23 on 9202 in a scratch database beside docmost's: the first half of the 135 816-byte undo copy, loaded with `psql -v ON_ERROR_STOP=1 --single-transaction`, returned **rc 0** and left **42 tables and 48 migration rows — and 0 users, 0 spaces, 0 constraints, 0 indexes**; psql treats end-of-file inside a `COPY` as end of data and commits. The whole copy ends with `-- PostgreSQL database dump complete` (plus a `\unrestrict` line); the truncated one does not. `ValidateDump` (`appbackup/dbdump.go:415`, read from source) checks the header and a `CREATE TABLE` only. **An undo or a restore fed a truncated copy would report success, start the app on an empty database, and pass a health check.** MariaDB's truncated load failed (rc 1) but NON-atomically — half the tables already replaced; its copy carries `-- Dump completed`. Fix: check the engine's completion marker before any load. Evidence: `audits/update-rulings-2026-09-23/README.md`, `docmost-46`, `docmost-47`, `romm-44`. | **OPEN — P2; owner: CC; lands with R-637, also guards the restore paths** | +| **R-641** | **[P2-MEDIUM] An app with no database server has NO last-second copy — the update's safety dump is a no-op — so if its old version refuses the migrated data, everything written since the last backup is lost by any route back.** MEASURED 2026-09-23 on 9202 with vikunja (SQLite in a volume): the controller logged *update safety dump for vikunja: the app has no database — nothing to copy (no-op)*. Vikunja 2.3.0 happens to start on 2.6.0's migrated SQLite, so nothing was lost; the fallback was measured anyway — the tier unit's volume tar put back in 0.86 s brought back the data written before the backup and **NOT the project and attachment written after it**. The fix is a tar of the data volumes at safety-dump time (2.2 MB here). Needed by R-637. Evidence: `audits/update-rulings-2026-09-23/README.md`, `vikunja-40`..`-44`. | **OPEN — P2; owner: CC; lands with R-637** | +| **R-642** | **[P3-LOW] `POST /api/stacks/{name}/start` answers 200 *Stack … start completed* while the app is crash-looping.** MEASURED 2026-09-23 on 9202 twice: docmost 0.95.0 and romm 5.0.0 started on data their newer versions had migrated — both refused and restarted in a loop (`Restarting (1)`, front door 404) behind a 200. The same false-green class as R-443 (closed for the Update) and R-635, on the Start. It matters now because an undo (R-637) must never read the start's return as success. Evidence: `audits/update-rulings-2026-09-23/README.md`, `docmost-43`, `romm-43`. | **OPEN — P3; owner: CC** | +| **R-643** | **[P2-MEDIUM] The ruled chain leaves the automatic update leg AT MOST 15 MINUTES a night.** FOUND 2026-09-23 while writing the build plan for `09` §3 decision 11 (*updates after the off-site copy, before the full-system backup*). The off-site leg starts at W+105m (`cmd/controller/main.go:961`) and the full-system backup's gate opens at W+2h (`quiesce/quiesce.go:656`, span to W+6h); the legs are clock-scheduled, not chained. One step takes ~1 min when it works and ~2–6 min when it fails and is undone. Options and the recommendation (the full-system backup waits for the leg inside its own window; the leg stops starting steps at W+5h) are in `09` §6.4. | **WAITING-ON-OPERATOR — `09` §6.4's one open point; owner: CC once answered** | +| **R-644** | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** |