diff --git a/CONTEXT.md b/CONTEXT.md index 3eae55a3..50121495 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -20,9 +20,10 @@ on demo-hp — `documentation/audits/slice4-2026-09-13/`, design `09-update-architecture.md` §6.1. **Decisions a later session must not re-derive:** -- **The precondition is the existing verified backup** (ruling 2026-09-02): `backup.Tier2UnitRestorePoint`, - the predicate the destructive Tier-2 unit restore already uses — extracted, not copied. Tier-2-only by - specification; that leaves apps with no Tier-2 copy un-updatable (R-475, operator decision). +- **The precondition is the existing verified backup** (ruling 2026-09-02) **on ANY tier** (ruling + 2026-09-13, controller v0.239.0, R-475 CLOSED): `backup.Manager.UpdateRestorePoints` walks Tier 2 → + Tier 1 → Tier 3 and the update leans on the first copy younger than `backup_max_age`; nothing + anywhere → back up first. `Tier2UnitRestorePoint` stays the Tier-2 half and the page's predicate. - **Age a copy by its last SUCCESSFUL Tier-2 copy, never by the unit manifest's `created_at`** — measured: the manifest moves only when the definition changes (R-476 is the page's side of this). - **No automatic rollback** — measured per-app in `SPIKE-upgrade-test-2026-09-06`. Health failure → stop + @@ -30,8 +31,10 @@ on demo-hp — `documentation/audits/slice4-2026-09-13/`, design `09-update-arch successful unit restore lifts it. - **Anything that writes a restore point skips an app that is held OR updating** (capture, Tier 2, volume dump) — the updating half was found live (v0.238.1). -- **A release does not reach the fleet by floor between golden bakes** — the hub holds a floor above the - vouched golden (R-472, operator decision). Slice 4's three releases were hand-deployed. +- **A release reaches the fleet by floor between golden bakes when its MinAgent is declared with the + floor** (ruling 2026-09-13, hub v0.112.0, R-472 CLOSED). An undeclared floor above the golden is still + held, and the forms refuse it. v0.239.0 arrived on both demo boxes this way in ~15 s. Never hand-deploy + a release to route around this — the floor IS the delivery. ## Goldens move to a CADENCE, and the gate learns to read a dated waiver (2026-09-13, R-468 / R-242 / R-467) @@ -41,7 +44,7 @@ design, and in August that produced **25 goldens in 26 days** plus thirteen decl bypasses (R-404/R-417). A guard bypassed that often teaches everyone to bypass it. Every release still raises the FLOOR, so the demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. -**⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.** +**⚠ CORRECTED THE SAME DAY (R-472): the hub held any floor above the vouched golden, so releases between bakes reached the demo boxes only by hand-deploy. RESOLVED the same afternoon (hub v0.112.0): a floor above the golden is served when saved with the release's declared MinAgent.** **The mechanism, because a rule without one is a wish (R-242 recurred the day after it was written):** `documentation/tests/golden-waiver.yml` — `issued`, `expires` (≤ 14 days, enforced by the gate's diff --git a/REPORT.md b/REPORT.md index 7c17eff8..7fe984b0 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,35 +1,40 @@ -# REPORT — update arc slice 4 documents, and a correction to this morning's golden-cadence page (2026-09-13) +# REPORT — two rulings: the floor carries a release without a golden, and every backup counts (2026-09-13) *Overwritten each session. Nothing durable lives only here — every finding below has a register row. -The implementation report is `felhom-controller/REPORT.md`.* +The controller report is `felhom-controller/REPORT.md`.* -> **The Update button is a guarded job, shipped as controller v0.237.0–v0.238.1 and proven live on -> demo-hp.** This repo carries its documents: `09-update-architecture.md` §6.1, the register, the -> capability map, ROADMAP, CONTEXT and STATUS. **It also carries a correction to what this repo said -> this morning:** under the weekly golden cadence a release does NOT reach the fleet by floor. +> **Both rulings are shipped and proven live.** Hub v0.112.0 serves a floor above the vouched golden +> when the release's MinAgent is declared with it. Controller v0.239.0 reached both demo boxes by that +> floor alone, in 14 s and 15 s, and lets an app update on any backup tier. ## What changed here -- `documentation/architecture/09-update-architecture.md` — slice 4 SHIPPED; §6.1 describes the sequence - as shipped, the two knobs, the abort decision (not built, by measurement), the v0.238.1 finding and - the live proof; §8.6 closed. -- `documentation/backlog/OPEN-ITEMS.md` — **R-448, R-443, R-439 CLOSED** and compressed to - `CLOSED-ITEMS.md`; **R-472..R-476 opened**; R-469 unblocked, not lifted. 214 open / 174 closed; - 431 689 bytes, from 439 828. -- `documentation/architecture/00-capability-map.md` — "Update is GUARDED" row, PROVEN-LIVE, citing - `audits/slice4-2026-09-13/`. -- `documentation/backlog/ROADMAP.md` — slice 4 collapsed. -- `documentation/audits/slice4-2026-09-13/` — live evidence, red-proof outputs, gate results. -- `STATUS.md` — item 14 (plain words), items 15 and 16 (two operator decisions). -- **The correction.** `RUNBOOK-manual-build.md` §4.2, `STATUS.md`, `CONTEXT.md`, R-468's row and the - `golden_currency_gate.py` docstring said a release still reaches the demo boxes by floor in about - 20 seconds between bakes. **The hub holds a floor above the vouched golden**, so that is false. Each - now says so and points at R-472. `scripts/CHANGELOG.md` records it. +- **Hub v0.112.0** (`f181efd`, deployed by `2f5d3af` + ArgoCD sync; Synced/Healthy, image 0.112.0). + Declared MinAgent stored beside both floors; `ResolveManagedFloor` uses it above the golden; both + forms refuse a floor above the golden without it (`floor_needs_min_agent`); Hosts badge and a + `managed floor SERVED … from ` log. Vouch path and R-120 gate untouched. +- **Runbooks:** `publish-train-rules.md` rule 1; `RUNBOOK-manual-build.md` §3 "Raise the floor to a + release with no golden" and the §4.2 correction resolved. +- **Architecture:** `09-update-architecture.md` §3 decisions 7 and 8, §6 and §6.1; `05-hub-architecture.md` + §5; `07-backup-architecture.md` §6; capability map "Update is GUARDED" row. +- **Register:** R-470, R-472, R-475 closed and compressed; R-477..R-480 opened; R-474 re-confirmed. +- **Evidence:** `documentation/audits/rulings-r472-r475-2026-09-13/` (live 01–10, red-proofs). +- `STATUS.md` items 15 and 16 rewritten as done; the cadence correction line fixed. `CONTEXT.md` updated. + +## Numbering note + +§3 already held a decision 6 from this morning, so the two new rulings are numbered 7 and 8 rather than 6 and 7. + +## Negative control + +An undeclared floor above the golden no longer reaches the hold: the form refuses it first (303 +`floor_needs_min_agent`, floor still 0.236.0, hub WARN). The hold itself for that shape is proven by +test C and its red-proof. ## Observations -1. **The golden cadence and the hub's floor rule contradict each other.** **FILED: R-472.** -2. **The glance catalog template crash-loops on a fresh install.** **FILED: R-473.** -3. **App removal with "delete backups" leaves the recovery unit and Tier-2 copy.** **FILED: R-474.** -4. **The update precondition is Tier-2-only.** **FILED: R-475.** -5. **The Mentések page dates a copy by its manifest, which lags the data.** **FILED: R-476.** +1. **The update's Tier-3 lookup takes its full 15 s bound and logs a misleading size WARN.** FILED: R-477 +2. **A leftover unit from a removed install counts as a reinstall's fresh Tier-1 copy.** FILED: R-478 +3. **A Tier-1 route back restores settings only for a bind-data app.** FILED: R-479 +4. **The card keeps the failure sentence after a successful restore.** FILED: R-480 +5. **Removal with `remove_backups` left units and prefs again.** FILED: R-474 diff --git a/STATUS.md b/STATUS.md index 4d885570..b74acd98 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,9 +1,13 @@ # STATUS — what works, what's broken, what's next +**Updated 2026-09-13 (fourth pass) — you decided both things. New releases reach the demo machines +by themselves again, and an app with any backup can now be updated. Both are live and proven on the +HP. Nothing needs you.** + **Updated 2026-09-13 (third pass) — the Update button now takes a backup first and tells the truth. It is live on both machines and I walked every case on the HP, including putting a broken update back from its backup. TWO THINGS NEED YOU: item 15 (the demo machines no longer get new releases by -themselves between golden images) and item 16 (apps with no second-drive copy cannot be updated).** +themselves between golden images) and item 16 (apps with no second-drive copy cannot be updated).** Both done the same afternoon. **Updated 2026-09-13 (second pass) — you decided both open items. The database engine now finishes its own conversion on the four MariaDB apps; the upgrade machine proved it and it landed on the HP @@ -252,31 +256,19 @@ nothing.* app came up, the regular backup copied the broken version into the app's local backup. It now leaves an app alone while it is updating. Live as 0.238.1. -15. **The demo machines no longer get new releases by themselves between golden images. One decision.** - This morning we agreed to bake the golden image weekly, on the promise that every release still - reaches both machines in about 20 seconds. **That promise was wrong.** The hub will not move a - machine to a release newer than the golden image, so between bakes I install releases on the demo - machines by hand. Today I did that three times. - - **Bake a golden image for every release again.** Releases reach the machines by themselves. The - cost is the weekly routine we just ended. - - **Let the hub move machines past the golden image when a release says it needs no newer agent.** - Releases reach the machines by themselves and goldens stay weekly. It needs one small hub change, - and first each release must reliably say which agent it needs (R-470). +15. **New releases reach the demo machines by themselves again.** You decided this, and it is done. + When I raise the floor for a release, I now also type which agent that release needs. The hub + then moves the machines past the golden image. If I leave that value out, the hub refuses to save. + Release 0.239.0 reached both machines this way in 15 seconds. Nothing needs you. - **My pick: the second.** **If you do nothing:** nothing breaks; I keep installing by hand and - saying so each time. - -16. **Apps with no copy on a second drive cannot be updated. One decision.** The update's safety rule - asks for the app's second-drive copy, as we decided on 2 September. On the HP, two apps (gokapi - and nextcloud) have no such copy, so their Update button refuses. A box with only one drive would - refuse every update. - - **Keep it.** Only apps with a second-drive copy can be updated. The refusal tells the customer how - to switch it on. - - **Also accept the backup on the same drive.** Updates work on one-drive boxes. That backup is lost - if that drive dies. - - **My pick: keep it for now,** and revisit when the first one-drive customer exists. **If you do - nothing:** those apps show the refusal and nothing else changes. +16. **Apps with no copy on a second drive can now be updated.** You decided this, and it is done. + The update now uses any backup: the second drive first, then the app's own copy, then the remote + copy. If none is younger than a day, it makes a fresh one first. If the new version fails, the page + names which backup to restore from. Proven on the HP: an app with only its own copy updated, a + broken update was held, and it came back from its own copy. Nothing needs you. + Four small things turned up and are written down: the remote check can take 15 seconds (R-477), + a leftover backup of a removed app counted as a fresh one (R-478), for some apps their own copy + holds only settings (R-479), and the card keeps the old failure sentence after a good restore (R-480). 8. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session, @@ -288,9 +280,9 @@ nothing.* In August I baked 25 goldens in 26 days, almost one per release, because the check trips on every release on purpose. From today: one golden a week, and always before a drill or a fresh install. **Correction, same day:** I wrote that every release would still reach both demo machines in about - 20 seconds between bakes. **That is wrong.** The hub will not move the machines past the image a new - machine starts from, so between bakes I install each release on the demo machines by hand. That is - item 15 below, and it is yours to decide. The check now reads a dated permission slip that runs out + 20 seconds between bakes. **That was wrong at first:** the hub would not move the machines past the + golden image. **It is true again since the afternoon:** the floor now carries the release when I type + which agent it needs (item 15). Release 0.239.0 arrived in 15 seconds. The check now reads a dated permission slip that runs out after at most 14 days; while it is valid the check warns instead of refusing, and when it runs out the check is red again until someone bakes or renews. A dated slip cannot be forgotten — it just expires. Today's golden (0.236.0) is baked, checked three ways, and live. diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 8b9a5e4b..4e6181f2 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -100,7 +100,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis |---|---|---|---|---| | Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only | | App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: `CAMPAIGN-3` proved `remove` removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets `[]` and a note.** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in `CAMPAIGN-3`; **remove WITH data: `audits/R442-2026-09-13/`**; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical | -| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up** | controller **v0.237.0 + v0.238.0 + v0.238.1** | **PROVEN-LIVE (2026-09-13)** — scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | **Tier-2-only precondition** — an app with no Tier-2 copy cannot be updated (R-475); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) | +| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) | | **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. | | **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. | | **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. | diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index 7e832174..607c759c 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -270,6 +270,12 @@ Recorded here so the encryption policy is not read as covering it. → **R-108** The tiers are **inputs to recovery**, not recovery routes. §7 and §8 say what they can actually do. +**The app update's safety precondition accepts ANY tier** (operator ruling 2026-09-13, controller +v0.239.0, R-475): the first fresh copy in the order Tier 2, Tier 1, Tier 3; with none, it backs up +first. Design: `09-update-architecture.md` §3 decision 8. **What a Tier-1 route back restores is only +what the unit holds** — for an app whose data is a bind mount that is the definition, not the data +(R-479). + ### 6.1 The four tiers, as configured on the live fleet | Tier | Location | Captures | Cadence (LIVE) | Retention (LIVE) | Encrypted | diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 3051dba7..1b102390 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -156,6 +156,28 @@ These are rulings, not proposals. Anything specced against a different assumptio --- +### 2026-09-13 (afternoon) — the floor carries a release without a golden, and every backup counts + +7. **A floor carries a release past the vouched golden when the release's MinAgent is declared with + it** (R-472; hub v0.112.0). Inside the golden the manifest's MinAgent governs, as before. Above it, + the MinAgent the operator declares with the floor — read from the release's CHANGELOG header, which + `minagent_header_gate.py` now guarantees — goes into the same per-box agent comparison. An + undeclared floor above the golden is still HELD, and both floor forms refuse to save one. **Why:** a + controller image is pulled by tag and needs no golden to be delivered; what the floor was missing + was only the agent requirement. This is what lets the weekly golden cadence (R-468) and per-release + delivery coexist. **Proven live:** both demo boxes self-updated 0.238.1 → 0.239.0 in 14 s and 15 s + from the save, the hub logging `SERVED … from declared` + (`audits/rulings-r472-r475-2026-09-13/03-declared-floor.txt`). +8. **Any backup tier lets an app update** (R-475; controller v0.239.0). The precondition takes the first + FRESH copy in the order second drive (Tier 2), the app's own recovery unit (Tier 1), off-site + (Tier 3, bounded; unreachable = absent). `update.backup_max_age` applies to whichever tier is chosen. + An app with nothing anywhere is backed up first; it is refused only when no backup can be taken + either. The hold names the tier and the date. **Tier 2 is required nowhere in the update path.** + This replaces decision 1's reading "the verified backup = the Tier-2 unit" — decision 1 itself (a + verified recent backup as a precondition, not a new copy) stands. **Proven live:** + `audits/rulings-r472-r475-2026-09-13/` 04 (nothing anywhere → backup first), 05 (Tier 1 alone), + 07 (a failed update held naming „saját meghajtó"), 08 (restored from „helyi", hold cleared). + ## 4. The vocabulary ruling — "rollback" is struck **App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually @@ -301,7 +323,7 @@ earlier feature is the failure mode to look for whenever a file changes meaning. | **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** | | **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** | | **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 | -| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13)** — §6.1 | +| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13); any backup tier since v0.239.0 (§3 decision 8)** — §6.1 | | **5** | **An upgrade test that runs again** — a harness that upgrades a real app with real data in it and asks the app for the data back. | **SHIPPED, `app-catalog/scripts/upgrade-test.py` (2026-09-06)** — 7 edges, 3 apps; see §4.1 and §10 | | **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 | | **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 | @@ -317,9 +339,9 @@ closed by construction: nothing reports an update complete on the compose exit c | # | phase | what happens | on failure | |---|---|---|---| -| 0 | refusals (409, before the intent is recorded) | held (R-439), busy (backup/restore/app-data op/quiesce), migration, already updating, deploying, memory (the deploy's `memoryVerdict`, releasing the app's own request), disk (**fixed 2 GB floor** — image size unknown without a registry), **no restorable Tier-2 copy** | nothing moves, nothing is recorded | -| 1 | `checking` | re-reads the precondition | nothing moves | -| 2 | `backing-up` — only when the proven copy is older than `update.backup_max_age` | `RunAppBackupNow`: this app's DB dump → volume dump → unit capture → Tier-2 copy | refused with the backup's own error; nothing moves | +| 0 | refusals (409, before the intent is recorded) | held (R-439), busy (backup/restore/app-data op/quiesce), migration, already updating, deploying, memory (the deploy's `memoryVerdict`, releasing the app's own request), disk (**fixed 2 GB floor** — image size unknown without a registry), **no copy on ANY tier and no backup can be taken now** (since v0.239.0; before it, no restorable Tier-2 copy) | nothing moves, nothing is recorded | +| 1 | `checking` | walks Tier 2 → Tier 1 → Tier 3 for the first copy younger than `update.backup_max_age` (v0.239.0) | nothing moves | +| 2 | `backing-up` — only when no tier holds a fresh copy | `RunAppBackupNow`: this app's DB dump → volume dump → unit capture (marked proven current) → Tier-2 copy, whose failure is a WARN since v0.239.0 | refused with the backup's own error; nothing moves | | 3 | `safety-dump` | `WriteUpdateSafetyDump` (R-361's undo copy) — **before the pin moves** | refused; nothing moves | | 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back | | 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) | @@ -337,7 +359,13 @@ SUCCESSFUL Tier-2 copy, not by the unit manifest's `created_at`**, and that was designed: a capture rewrites the manifest only when the app's DEFINITION changes, so on demo-hp bookstack's mirror held a 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aged by the manifest, a quiet app would be "stale" forever and a backup-first would not fix it. **The predicate is -Tier-2-only, as specified — an app with no Tier-2 copy cannot be updated (R-475).** +Tier-2-only, as specified — an app with no Tier-2 copy cannot be updated (R-475).** **SUPERSEDED in +v0.239.0 by §3 decision 8:** `backup.Manager.UpdateRestorePoints` walks all three tiers and the update +leans on the first fresh copy; `Tier2UnitRestorePoint` is still the Tier-2 half and still the page's +predicate for „Teljes visszaállítás". The same aging trap existed one tier down — a capture's checksum +skip leaves a quiet app's unit manifest untouched — so "back up first" now marks the captured unit +proven current. The hold stores `copy_tier` and names „második meghajtó" / „saját meghajtó" / +„távoli mentés"; a successful off-site restore now lifts an update hold too. **The hold** is `settings.RestoreHold` with `reason: update_failed` and `copy_date` — the SAME store and gate as R-379, so every start path that already honoured a restore hold honours this one. A successful @@ -361,7 +389,8 @@ the harness has proven it, is slice 6's. **The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0 -were hand-deployed to the demo guests. That contradiction is an operator decision. +were hand-deployed to the demo guests. **RESOLVED by §3 decision 7 (hub v0.112.0):** v0.239.0 reached +both demo boxes by the floor alone, with its MinAgent declared. **Found live, fixed in v0.238.1: the nightly legs must leave an app alone WHILE it is updating, not only once it is held.** In Scenario F the periodic unit capture ran at 10:17:09 — inside the 5-minute diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/01-hub-deploy.txt b/documentation/audits/rulings-r472-r475-2026-09-13/01-hub-deploy.txt new file mode 100644 index 00000000..4fa09fbc --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/01-hub-deploy.txt @@ -0,0 +1,45 @@ +=== hub 0.112.0 deploy — 2026-09-13T14:50:54Z === +sync=Synced health=Healthy revision=2f5d3af6f9f48f5508614e1b90c0ec098c881e5f op=Succeeded +image=gitea.dooplex.hu/admin/felhom-hub:0.112.0 +hub-687c88ccf8-lc54n 1/1 Running 0 88s 10.42.0.53 dooplex +--- hub log since start --- +2026/09/13 16:49:28 [INFO] felhom-hub 0.112.0 starting +2026/09/13 16:49:28 [INFO] Database opened at /data/hub.db +2026/09/13 16:49:28 [INFO] Default controller-version floor: 0.120.0 +2026/09/13 16:49:28 [INFO] Template fetcher started (every 1h) +2026/09/13 16:49:30 [INFO] Asset seed: all 210 files up-to-date +2026/09/13 16:49:30 [INFO] Asset manifest built: 211 files +2026/09/13 16:49:30 [INFO] App-email relay enabled (limit 30/min/customer, From domains [felhom.eu]) +2026/09/13 16:49:30 [INFO] Gitea artifact browser enabled (Day-0 version dropdowns) via http://gitea.gitea-system.svc.cluster.local:3000 +2026/09/13 16:49:30 [INFO] Offsite provisioning enabled (pool_box=611714, location=fsn1) +2026/09/13 16:49:30 [INFO] Offsite pool-box checker initialized: box=611714 fill warn=80% crit=90%, oversub warn=2.00x, unreachable after 3 consecutive failed reads, refresh 15m0s +2026/09/13 16:49:30 [INFO] Registry version checker started (every 6h) +2026/09/13 16:49:30 [INFO] WG peer-sync enabled (endpoint 167.233.158.164:22, user felhom-peersync) +2026/09/13 16:49:30 [INFO] PBS DR tenantsync enabled (endpoint 167.233.158.164:22, user felhom-peersync; WG-registration auto-provision hook armed) +2026/09/13 16:49:30 [INFO] PBS-DR box checker initialized: fill warn=80% crit=90%, unreachable after 3 consecutive failed reads, refresh 15m0s +2026/09/13 16:49:30 [INFO] agent-plane poke enabled (endpoint 167.233.158.164:22, user felhom-peersync; web + api admin seams armed) +2026/09/13 16:49:30 [INFO] PBS-DR self-heal reconciler started (interval 5m) +2026/09/13 16:49:30 [INFO] offsite credential self-heal reconciler started (interval 5m, debounce 2 reports) +2026/09/13 16:49:43 [INFO] Pruned 59 old report rows +2026/09/13 16:49:44 [INFO] Pruned 21 old event rows +2026/09/13 16:49:45 [INFO] Pruned 763 old app telemetry rows +2026/09/13 16:49:45 [INFO] Pruned 3 stale app issues +2026/09/13 16:49:45 [INFO] prune: next run at 2026-09-14 04:30 CEST (in 11h40m14s) +2026/09/13 16:49:45 [INFO] Staleness checker initialized: 2 ok, 0 stale, 1 down +2026/09/13 16:49:45 [INFO] Host staleness checker initialized: 2 ok, 0 stale, 0 down (stale/down left unseeded → first Check emits) +2026/09/13 16:49:45 [INFO] Host capability checker initialized: 2 ok, 0 degraded (degraded left unseeded → first Check emits) +2026/09/13 16:49:45 [INFO] Host leaf checker initialized: 2 host fingerprint(s) seeded +2026/09/13 16:49:45 [INFO] Host disk checker initialized: warn=90% crit=95%, 2 ok seeded, 0 already-breached left unseeded (first Check emits) +2026/09/13 16:49:45 [INFO] Storage fill checker initialized: warn=90% crit=95%, 5 ok seeded, 0 already-breached left unseeded, 2 root-backed excluded +2026/09/13 16:49:45 [INFO] Host mgmt-plane checker initialized: 0 host heal-state(s) seeded +2026/09/13 16:49:45 [INFO] Host OOB checker initialized: 0 host(s) seeded degraded +2026/09/13 16:49:45 [INFO] Offsite checker initialized: fill warn=90% crit=95%, stale after 48h0m0s, 3 ok-seeded +2026/09/13 16:49:45 [INFO] deadline-check: next run at 2026-09-14 05:00 CEST (in 12h10m14s) +2026/09/13 16:49:45 [INFO] Listening on :8080 +2026/09/13 16:50:03 [INFO] DR-recipe app-half stored for customer demo-felhom (v1) +2026/09/13 16:50:03 [INFO] Received report from demo-felhom (3268 bytes) +2026/09/13 16:50:03 [INFO] managed floor SERVED for demo-felhom: floor 0.236.0, agent requirement "0.129.0" from manifest (golden 0.236.0) +2026/09/13 16:50:04 [INFO] DR-recipe app-half stored for customer demo-hp (v1) +2026/09/13 16:50:04 [INFO] Received report from demo-hp (8019 bytes) +2026/09/13 16:50:04 [INFO] managed floor SERVED for demo-hp: floor 0.236.0, agent requirement "0.129.0" from manifest (golden 0.236.0) +2026/09/13 16:50:50 [INFO] Offsite pool-box refreshed: 0.3% full (2.7 GB of 1.00 TB), Σ shared quota 150 GB, oversub 0.15x diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/02-negative-control.txt b/documentation/audits/rulings-r472-r475-2026-09-13/02-negative-control.txt new file mode 100644 index 00000000..9ca71768 --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/02-negative-control.txt @@ -0,0 +1,17 @@ +=== baseline — 2026-09-13T15:18:28Z === +hub image: gitea.dooplex.hu/admin/felhom-hub:0.112.0 +DB floor: "0.236.0" declared field: name="min_agent" value="" +demo-hp (ssh hp) 9201: gitea.dooplex.hu/admin/felhom-controller:0.238.1 Up 5 hours (healthy) +demo-felhom (felhom-pve-lan) 9201: gitea.dooplex.hu/admin/felhom-controller:0.238.1 Up 5 hours (healthy) +VMID Status Lock Name +9201 running demo-felhom + +=== NEGATIVE CONTROL: floor 0.239.0 WITHOUT min_agent — 2026-09-13T15:18:34Z === +HTTP/1.1 303 See Other +Location: /configuration?flash=floor_needs_min_agent +DB floor after: "0.236.0" declared field after: name="min_agent" value="" +--- hub log, the refusal --- +2026/09/13 17:18:34 [WARN] floor 0.239.0 refused: it is above the vouched golden 0.236.0 and no MinAgent was declared (R-472) +--- the rendered flash (ASCII fragment + control) --- +1 +0 diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/03-declared-floor.txt b/documentation/audits/rulings-r472-r475-2026-09-13/03-declared-floor.txt new file mode 100644 index 00000000..c79b865a --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/03-declared-floor.txt @@ -0,0 +1,54 @@ +=== DECLARED FLOOR: 0.239.0 with min_agent 0.129.0 === +T0 POST at 2026-09-13T15:19:12Z +HTTP/1.1 303 See Other +Location: /configuration?flash=floor_set +DB floor + declared: id="global-floor-input" name="min_controller_version" value="0.239.0" | name="min_agent" value="0.129.0" placeholder="MinAgent from | +demo-hp on 0.239.0 at 2026-09-13T15:19:26Z (+14s): gitea.dooplex.hu/admin/felhom-controller:0.239.0 Up 5 seconds (healthy) +demo-felhom on 0.239.0 at 2026-09-13T15:19:27Z (+15s): gitea.dooplex.hu/admin/felhom-controller:0.239.0 Up 8 seconds (healthy) +--- hub log since T0 (managed floor) --- +2026/09/13 17:19:12 [INFO] Global controller-version floor set to "0.239.0" (declared MinAgent "0.129.0") +2026/09/13 17:19:14 [INFO] managed floor SERVED for demo-felhom: floor 0.239.0, agent requirement "0.129.0" from declared (golden 0.236.0) +2026/09/13 17:19:15 [INFO] managed floor SERVED for demo-hp: floor 0.239.0, agent requirement "0.129.0" from declared (golden 0.236.0) + +--- demo-hp: controller log (SetFloor / self-update / version) --- +2026/09/13 15:19:22 updater.go:96: [DEBUG] [selfupdate] VerifyStartup: checking update state in /opt/docker/felhom-controller/data +2026/09/13 15:19:22 updater.go:96: [DEBUG] [selfupdate] VerifyStartup: pending update found — target=0.239.0 previous=0.238.1 +2026/09/13 15:19:22 updater.go:809: [INFO] [selfupdate] Post-update startup: update successful (0.238.1 → 0.239.0) +2026/09/13 15:19:22 main.go:655: [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30) +2026/09/13 15:19:22 offsiteapply.go:151: [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply +2026/09/13 15:19:22 scheduler.go:102: [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s) +2026/09/13 15:19:22 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="selfupdate-check" interval=6h0m0s totalJobs=18 +2026/09/13 15:19:23 client.go:67: [DEBUG] [agentapi] agent version seen: 0.130.0 (was "") +2026/09/13 15:19:27 builder.go:40: [DEBUG] [report] BuildReport: starting — version=0.239.0, storagePaths=1 +2026/09/13 15:19:28 updater.go:96: [DEBUG] [selfupdate] SetFloor: floor "" → "0.239.0" +2026/09/13 15:19:28 updater.go:96: [DEBUG] [selfupdate] maybeAutoUpdate: current 0.239.0 >= floor 0.239.0 — no action +2026/09/13 15:19:32 offsiteapply.go:151: [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.239.0 (we are 0.239.0), no managed update running + +--- demo-hp: bootstrap service journal --- +Sep 13 15:19:20 demo-hp systemd[1]: felhom-controller-bootstrap.service: Deactivated successfully. +Sep 13 15:19:20 demo-hp systemd[1]: Stopped felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount). +Sep 13 15:19:20 demo-hp systemd[1]: Stopping felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)... +Sep 13 15:19:20 demo-hp systemd[1]: Starting felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)... +Sep 13 15:19:21 demo-hp felhom-controller-bootstrap.sh[1829428]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.239.0 from /etc/felhom-bootstrap/bootstrap.json (hostname=demo-hp) +Sep 13 15:19:21 demo-hp felhom-controller-bootstrap.sh[1829480]: d27aeecc7f869b5836ee7a6c2c566213df05fce8f0f97a1f6beca99661680e16 +Sep 13 15:19:21 demo-hp felhom-controller-bootstrap.sh[1829428]: [ctrl-bootstrap] controller started +Sep 13 15:19:21 demo-hp systemd[1]: Finished felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount). + +--- demo-felhom: controller log (SetFloor / self-update / version) --- +2026/09/13 15:19:20 [INFO] [selfupdate] Post-update startup: update successful (0.238.1 → 0.239.0) +2026/09/13 15:19:20 [INFO] Self-update enabled (check every 6h, auto-update: false, auto-update time: 04:30) +2026/09/13 15:19:20 [INFO] [scheduler] Registered periodic job: selfupdate-check (every 6h0m0s) +2026/09/13 15:19:20 [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply +2026/09/13 15:19:30 [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.239.0 (we are 0.239.0), no managed update running + +--- demo-felhom: bootstrap service journal --- +Sep 13 15:19:19 demo-felhom systemd[1]: felhom-controller-bootstrap.service: Deactivated successfully. +Sep 13 15:19:19 demo-felhom systemd[1]: Stopped felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount). +Sep 13 15:19:19 demo-felhom systemd[1]: Stopping felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)... +Sep 13 15:19:19 demo-felhom systemd[1]: Starting felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount)... +Sep 13 15:19:19 demo-felhom felhom-controller-bootstrap.sh[3492140]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.239.0 from /etc/felhom-bootstrap/bootstrap.json (hostname=demo-felhom) +Sep 13 15:19:19 demo-felhom felhom-controller-bootstrap.sh[3492189]: 30f28bf3afa5f7b665ef62ca3ae5df0b23a1610eb74385a834382e3eaf6b7805 +Sep 13 15:19:19 demo-felhom felhom-controller-bootstrap.sh[3492140]: [ctrl-bootstrap] controller started +Sep 13 15:19:19 demo-felhom systemd[1]: Finished felhom-controller-bootstrap.service - Felhom controller bootstrap (deploy the baked controller from the agent-populated config mount). + +SUMMARY T0=2026-09-13T15:19:12Z {'demo-hp': ('2026-09-13T15:19:26Z', 14), 'demo-felhom': ('2026-09-13T15:19:27Z', 15)} not done: [] diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/04-K-no-unit-anywhere.txt b/documentation/audits/rulings-r472-r475-2026-09-13/04-K-no-unit-anywhere.txt new file mode 100644 index 00000000..8a1a19dc --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/04-K-no-unit-anywhere.txt @@ -0,0 +1,58 @@ +=== Scenario K — no unit on ANY tier, a backup can be taken — 2026-09-13T15:25:41Z, controller 0.239.0 === +Fixture: throwaway actualbudget (deployed 15:24:21Z, Tier 2 switched off per app). A fresh deploy gets its +Tier-1 unit within seconds (attempt 04a), so the unit is REMOVED here to make 'no unit anywhere' true. +--- BEFORE removal (helper fixed: every tier location) --- +/mnt/sys_drive/felhom-data/backups/primary/actualbudget: + 2026-09-13T15:24:22.6041945430 1017 manifest.json + 2026-09-13T15:24:22.6041945430 291 compose/app.yaml + 2026-09-13T15:24:22.6041945430 2259 compose/.felhom.yml + 2026-09-13T15:24:22.6041945430 1412 compose/docker-compose.yml +/mnt/felhom-drives/hdd_1/backups/primary/actualbudget: ABSENT +/mnt/felhom-drives/hdd_1/backups/secondary/actualbudget: ABSENT +app_backup record (off-site per-app pref): {'enabled': False, 'cross_drive': {'enabled': False, 'method': 'rsync', 'destination_path': '', 'schedule': 'daily', 'user_disabled': True}} +--- remove the throwaway's own recovery unit, then POST the update with NO request in between --- +removed +ls: cannot access '/mnt/sys_drive/felhom-data/backups/primary/actualbudget': No such file or directory +--- POST /api/stacks/actualbudget/update at 2026-09-13T15:25:45Z --- +HTTP 202 +{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"} +end (2026-09-13T15:26:10Z, +24s): state=running updating=False phase=done err='' hold='' +--- the controller, verbatim --- +2026/09/13 15:25:45 update.go:336: [INFO] [stacks] update actualbudget: accepted — guarded update started +2026/09/13 15:25:45 update.go:763: [INFO] [stacks] update actualbudget: phase checking +2026/09/13 15:25:45 update_guard.go:173: [DEBUG] [backup] update precondition for actualbudget: no Tier-2 copy (nincs másodlagos fájlmásolat ehhez az alkalmazáshoz) +2026/09/13 15:26:00 update.go:452: [INFO] [stacks] update actualbudget: no copy younger than 24h0m0s on any tier (found: none) — backing up first +2026/09/13 15:26:00 update.go:763: [INFO] [stacks] update actualbudget: phase backing-up +2026/09/13 15:26:00 update_guard.go:287: [INFO] [backup] update pre-backup for actualbudget: starting (DB dump → volume dump → unit capture → Tier 2) +2026/09/13 15:26:00 manager.go:1199: [DEBUG] [stacks] StopStack actualbudget: current state=running deployed=true containers=1 +2026/09/13 15:26:01 manager.go:1117: [DEBUG] [stacks] StartStack actualbudget: current state=stopped deployed=true +2026/09/13 15:26:01 manager.go:1127: [DEBUG] [stacks] StartStack actualbudget: prepared 8 env vars for compose +2026/09/13 15:26:02 installed.go:397: [DEBUG] [stacks] installed-images actualbudget: unchanged (1 service(s)) — app.yaml not rewritten +2026/09/13 15:26:02 update_guard.go:337: [INFO] [backup] update pre-backup for actualbudget: volume dump OK +2026/09/13 15:26:02 update_guard.go:343: [INFO] [backup] update pre-backup for actualbudget: recovery unit captured (0 database dump(s)) +2026/09/13 15:26:02 tier2.go:293: [INFO] [backup] Tier 2 for actualbudget skipped — disabled by customer +2026/09/13 15:26:02 update_guard.go:352: [INFO] [backup] update pre-backup for actualbudget: complete in 1.423s +2026/09/13 15:26:02 update_guard.go:173: [DEBUG] [backup] update precondition for actualbudget: no Tier-2 copy (nincs másodlagos fájlmásolat ehhez az alkalmazáshoz) +2026/09/13 15:26:02 update.go:467: [INFO] [stacks] update actualbudget: precondition met after the backup — Tier 1 (own recovery unit) copy from 2026-09-13T15:26:02Z +2026/09/13 15:26:02 update.go:763: [INFO] [stacks] update actualbudget: phase safety-dump +2026/09/13 15:26:02 update_guard.go:403: [INFO] [backup] update safety dump for actualbudget: the app has no database — nothing to copy (no-op) +2026/09/13 15:26:02 update.go:482: [INFO] [stacks] update actualbudget: safety dump done (0 file(s)) [] +2026/09/13 15:26:02 update.go:763: [INFO] [stacks] update actualbudget: phase pinning +2026/09/13 15:26:02 pin.go:362: [INFO] [stacks] update actualbudget: pin advanced to the catalog's current definition (actualbudget=actualbudget/actual-server:26.7.0) +2026/09/13 15:26:02 update.go:763: [INFO] [stacks] update actualbudget: phase pulling +2026/09/13 15:26:03 update.go:763: [INFO] [stacks] update actualbudget: phase starting +2026/09/13 15:26:03 update.go:763: [INFO] [stacks] update actualbudget: phase verifying +2026/09/13 15:26:08 healthprobe.go:153: [DEBUG] Health probe actualbudget: HTTP GET :5006/ → 200 (4ms) +2026/09/13 15:26:08 update.go:558: [INFO] [stacks] update actualbudget: healthy after 5s (the app's health check passed) +2026/09/13 15:26:08 installed.go:397: [DEBUG] [stacks] installed-images actualbudget: unchanged (1 service(s)) — app.yaml not rewritten +2026/09/13 15:26:08 update.go:564: [INFO] [stacks] update actualbudget: DONE in 23s +--- AFTER: every tier location --- +/mnt/sys_drive/felhom-data/backups/primary/actualbudget: + 2026-09-13T15:26:02.2202745980 1059 manifest.json + 2026-09-13T15:26:02.2194343050 291 compose/app.yaml + 2026-09-13T15:26:02.2194343050 2259 compose/.felhom.yml + 2026-09-13T15:26:02.2194343050 1412 compose/docker-compose.yml + 2026-09-13T15:26:01.5313724220 74240 volume-dumps/actualbudget_actualbudget_data.tar +/mnt/felhom-drives/hdd_1/backups/primary/actualbudget: ABSENT +/mnt/felhom-drives/hdd_1/backups/secondary/actualbudget: ABSENT +=== K done 2026-09-13T15:26:12Z === diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/04a-attempt-deploy-already-made-a-tier1-unit.txt b/documentation/audits/rulings-r472-r475-2026-09-13/04a-attempt-deploy-already-made-a-tier1-unit.txt new file mode 100644 index 00000000..0f6908db --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/04a-attempt-deploy-already-made-a-tier1-unit.txt @@ -0,0 +1,84 @@ +=== Scenario K — no unit anywhere, a backup can be taken — 2026-09-13T15:24:19Z, controller 0.239.0 === +--- BEFORE: every tier's location --- +: + 2026-07-04T09:05:00.0000000000 132 .profile + 2026-08-18T10:54:08.4964149220 0 .bash_history + 2026-08-21T21:09:27.0884871020 2777 settings.json.drill-backup + 2026-07-04T09:05:00.0000000000 607 .bashrc +: + 2026-07-04T09:05:00.0000000000 132 .profile + 2026-08-18T10:54:08.4964149220 0 .bash_history + 2026-08-21T21:09:27.0884871020 2777 settings.json.drill-backup + 2026-07-04T09:05:00.0000000000 607 .bashrc +: + 2026-07-04T09:05:00.0000000000 132 .profile + 2026-08-18T10:54:08.4964149220 0 .bash_history + 2026-08-21T21:09:27.0884871020 2777 settings.json.drill-backup + 2026-07-04T09:05:00.0000000000 607 .bashrc +app_backup record: None +--- deploy actualbudget (throwaway) at 2026-09-13T15:24:21Z --- +HTTP 202 +{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"} +after deploy (2026-09-13T15:24:47Z): state=running updating=False phase=None err='' hold='' +--- Tier 2 OFF for actualbudget via POST /stacks/actualbudget/backup (form, enabled absent) at 2026-09-13T15:24:47Z --- +HTTP/2 303 +location: /stacks/actualbudget/backup?flash=A+2.+ment%C3%A9s+be%C3%A1ll%C3%ADt%C3%A1sa+elmentve. + +--- IMMEDIATELY BEFORE the update: still no unit? --- +: + 2026-07-04T09:05:00.0000000000 132 .profile + 2026-08-18T10:54:08.4964149220 0 .bash_history + 2026-08-21T21:09:27.0884871020 2777 settings.json.drill-backup + 2026-07-04T09:05:00.0000000000 607 .bashrc +: + 2026-07-04T09:05:00.0000000000 132 .profile + 2026-08-18T10:54:08.4964149220 0 .bash_history + 2026-08-21T21:09:27.0884871020 2777 settings.json.drill-backup + 2026-07-04T09:05:00.0000000000 607 .bashrc +: + 2026-07-04T09:05:00.0000000000 132 .profile + 2026-08-18T10:54:08.4964149220 0 .bash_history + 2026-08-21T21:09:27.0884871020 2777 settings.json.drill-backup + 2026-07-04T09:05:00.0000000000 607 .bashrc +--- POST /api/stacks/actualbudget/update at 2026-09-13T15:24:48Z --- +HTTP 202 +{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"} + + 0s phase safety-dump + + 3s phase done +end (2026-09-13T15:24:51Z, +3s): state=running updating=False phase=done err='' hold='' +--- the controller, verbatim (update + pre-backup + tier choice) --- +2026/09/13 15:24:48 update.go:336: [INFO] [stacks] update actualbudget: accepted — guarded update started +2026/09/13 15:24:48 update.go:763: [INFO] [stacks] update actualbudget: phase checking +2026/09/13 15:24:48 update_guard.go:173: [DEBUG] [backup] update precondition for actualbudget: no Tier-2 copy (nincs másodlagos fájlmásolat ehhez az alkalmazáshoz) +2026/09/13 15:24:48 update.go:450: [INFO] [stacks] update actualbudget: precondition met — Tier 1 (own recovery unit) copy from 2026-09-13T15:24:22Z (0s old, limit 24h0m0s) +2026/09/13 15:24:48 update.go:763: [INFO] [stacks] update actualbudget: phase safety-dump +2026/09/13 15:24:48 update_guard.go:403: [INFO] [backup] update safety dump for actualbudget: the app has no database — nothing to copy (no-op) +2026/09/13 15:24:48 update.go:482: [INFO] [stacks] update actualbudget: safety dump done (0 file(s)) [] +2026/09/13 15:24:48 update.go:763: [INFO] [stacks] update actualbudget: phase pinning +2026/09/13 15:24:48 pin.go:362: [INFO] [stacks] update actualbudget: pin advanced to the catalog's current definition (actualbudget=actualbudget/actual-server:26.7.0) +2026/09/13 15:24:48 update.go:763: [INFO] [stacks] update actualbudget: phase pulling +2026/09/13 15:24:50 update.go:763: [INFO] [stacks] update actualbudget: phase starting +2026/09/13 15:24:50 update.go:763: [INFO] [stacks] update actualbudget: phase verifying +2026/09/13 15:24:50 healthprobe.go:153: [DEBUG] Health probe actualbudget: HTTP GET :5006/ → 200 (3ms) +2026/09/13 15:24:50 update.go:558: [INFO] [stacks] update actualbudget: healthy after 0s (the app's health check passed) +2026/09/13 15:24:50 installed.go:397: [DEBUG] [stacks] installed-images actualbudget: unchanged (1 service(s)) — app.yaml not rewritten +2026/09/13 15:24:50 update.go:564: [INFO] [stacks] update actualbudget: DONE in 2s +--- AFTER: the unit the backup-first made --- +: + 2026-07-04T09:05:00.0000000000 132 .profile + 2026-08-18T10:54:08.4964149220 0 .bash_history + 2026-08-21T21:09:27.0884871020 2777 settings.json.drill-backup + 2026-07-04T09:05:00.0000000000 607 .bashrc +: + 2026-07-04T09:05:00.0000000000 132 .profile + 2026-08-18T10:54:08.4964149220 0 .bash_history + 2026-08-21T21:09:27.0884871020 2777 settings.json.drill-backup + 2026-07-04T09:05:00.0000000000 607 .bashrc +: + 2026-07-04T09:05:00.0000000000 132 .profile + 2026-08-18T10:54:08.4964149220 0 .bash_history + 2026-08-21T21:09:27.0884871020 2777 settings.json.drill-backup + 2026-07-04T09:05:00.0000000000 607 .bashrc +--- the card --- +state=running updating=False phase=done err='' hold='' +=== K done 2026-09-13T15:24:54Z === diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/05-H-gokapi-tier1.txt b/documentation/audits/rulings-r472-r475-2026-09-13/05-H-gokapi-tier1.txt new file mode 100644 index 00000000..203ce7c2 --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/05-H-gokapi-tier1.txt @@ -0,0 +1,52 @@ +=== Scenario H — gokapi, only its own recovery unit (Tier 1) — 2026-09-13T15:30:45Z, controller 0.239.0 === +--- BEFORE: every tier location (a unit left from 06:59Z, R-474); Tier 2 already switched off per app --- +/mnt/sys_drive/felhom-data/backups/primary/gokapi: + 2026-09-13T06:59:17.4202088100 1014 manifest.json + 2026-09-13T06:59:17.4202088100 394 compose/app.yaml + 2026-09-13T06:59:17.4202088100 2497 compose/.felhom.yml + 2026-09-13T06:59:17.4202088100 2993 compose/docker-compose.yml +/mnt/felhom-drives/hdd_1/backups/primary/gokapi: ABSENT +/mnt/felhom-drives/hdd_1/backups/secondary/gokapi: ABSENT +--- deploy gokapi (throwaway; GOKAPI_PASSWORD generated here, never printed or stored) at 2026-09-13T15:30:46Z --- +HTTP 202 +{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"} +after deploy (2026-09-13T15:31:07Z): state=running updating=False phase=None err='' hold='' +app_backup record: {'enabled': False, 'cross_drive': {'enabled': False, 'method': 'rsync', 'destination_path': '', 'schedule': 'daily', 'user_disabled': True}} +--- IMMEDIATELY BEFORE the update --- +/mnt/sys_drive/felhom-data/backups/primary/gokapi: + 2026-09-13T06:59:17.4202088100 1014 manifest.json + 2026-09-13T06:59:17.4202088100 394 compose/app.yaml + 2026-09-13T06:59:17.4202088100 2497 compose/.felhom.yml + 2026-09-13T06:59:17.4202088100 2993 compose/docker-compose.yml +/mnt/felhom-drives/hdd_1/backups/primary/gokapi: ABSENT +/mnt/felhom-drives/hdd_1/backups/secondary/gokapi: ABSENT +--- POST /api/stacks/gokapi/update at 2026-09-13T15:31:09Z --- +HTTP 202 +{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"} + + 0s phase safety-dump + + 3s phase done +end (2026-09-13T15:31:13Z, +3s): state=running updating=False phase=done err='' hold='' +--- the controller, verbatim --- +2026/09/13 15:31:10 update.go:336: [INFO] [stacks] update gokapi: accepted — guarded update started +2026/09/13 15:31:10 update.go:763: [INFO] [stacks] update gokapi: phase checking +2026/09/13 15:31:10 update_guard.go:173: [DEBUG] [backup] update precondition for gokapi: no Tier-2 copy (nincs másodlagos fájlmásolat ehhez az alkalmazáshoz) +2026/09/13 15:31:10 update.go:450: [INFO] [stacks] update gokapi: precondition met — Tier 1 (own recovery unit) copy from 2026-09-13T06:59:17Z (8h32m0s old, limit 24h0m0s) +2026/09/13 15:31:10 update.go:763: [INFO] [stacks] update gokapi: phase safety-dump +2026/09/13 15:31:10 update.go:482: [INFO] [stacks] update gokapi: safety dump done (0 file(s)) [] +2026/09/13 15:31:10 update.go:763: [INFO] [stacks] update gokapi: phase pinning +2026/09/13 15:31:10 pin.go:362: [INFO] [stacks] update gokapi: pin advanced to the catalog's current definition (gokapi=f0rc3/gokapi:v1.9.6) +2026/09/13 15:31:10 update.go:763: [INFO] [stacks] update gokapi: phase pulling +2026/09/13 15:31:11 update.go:763: [INFO] [stacks] update gokapi: phase starting +2026/09/13 15:31:11 update.go:763: [INFO] [stacks] update gokapi: phase verifying +2026/09/13 15:31:11 update.go:558: [INFO] [stacks] update gokapi: healthy after 0s (the app's health check passed) +2026/09/13 15:31:11 update.go:564: [INFO] [stacks] update gokapi: DONE in 2s +--- AFTER --- +/mnt/sys_drive/felhom-data/backups/primary/gokapi: + 2026-09-13T06:59:17.4202088100 1014 manifest.json + 2026-09-13T06:59:17.4202088100 394 compose/app.yaml + 2026-09-13T06:59:17.4202088100 2497 compose/.felhom.yml + 2026-09-13T06:59:17.4202088100 2993 compose/docker-compose.yml +/mnt/felhom-drives/hdd_1/backups/primary/gokapi: ABSENT +/mnt/felhom-drives/hdd_1/backups/secondary/gokapi: ABSENT +f0rc3/gokapi:v1.9.6 Up 29 seconds (healthy) +=== H done 2026-09-13T15:31:16Z === diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/05a-attempt-gokapi-deploy-refused-missing-password.txt b/documentation/audits/rulings-r472-r475-2026-09-13/05a-attempt-gokapi-deploy-refused-missing-password.txt new file mode 100644 index 00000000..3f29fcbc --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/05a-attempt-gokapi-deploy-refused-missing-password.txt @@ -0,0 +1,42 @@ +=== Scenario H — gokapi, only its own recovery unit (Tier 1) — 2026-09-13T15:26:36Z, controller 0.239.0 === +--- BEFORE: every tier location (a unit left from 06:59Z, R-474) --- +/mnt/sys_drive/felhom-data/backups/primary/gokapi: + 2026-09-13T06:59:17.4202088100 1014 manifest.json + 2026-09-13T06:59:17.4202088100 394 compose/app.yaml + 2026-09-13T06:59:17.4202088100 2497 compose/.felhom.yml + 2026-09-13T06:59:17.4202088100 2993 compose/docker-compose.yml +/mnt/felhom-drives/hdd_1/backups/primary/gokapi: ABSENT +/mnt/felhom-drives/hdd_1/backups/secondary/gokapi: ABSENT +--- deploy gokapi (throwaway) at 2026-09-13T15:26:37Z --- +HTTP 400 +{"ok":false,"error":"a(z) \"Admin jelszó\" mező kitöltése kötelező — használja a Generálás gombot vagy írjon be egy jelszót"} +after deploy (2026-09-13T15:29:59Z): state=not_deployed updating=False phase=None err='' hold='' +--- Tier 2 OFF for gokapi via POST /stacks/gokapi/backup (form, enabled absent) at 2026-09-13T15:29:59Z --- +HTTP/2 303 +location: /stacks/gokapi/backup?flash=A+2.+ment%C3%A9s+be%C3%A1ll%C3%ADt%C3%A1sa+elmentve. + +app_backup record: {'enabled': False, 'cross_drive': {'enabled': False, 'method': 'rsync', 'destination_path': '', 'schedule': 'daily', 'user_disabled': True}} +--- IMMEDIATELY BEFORE the update --- +/mnt/sys_drive/felhom-data/backups/primary/gokapi: + 2026-09-13T06:59:17.4202088100 1014 manifest.json + 2026-09-13T06:59:17.4202088100 394 compose/app.yaml + 2026-09-13T06:59:17.4202088100 2497 compose/.felhom.yml + 2026-09-13T06:59:17.4202088100 2993 compose/docker-compose.yml +/mnt/felhom-drives/hdd_1/backups/primary/gokapi: ABSENT +/mnt/felhom-drives/hdd_1/backups/secondary/gokapi: ABSENT +--- POST /api/stacks/gokapi/update at 2026-09-13T15:30:02Z --- +HTTP 409 +{"ok":false,"error":"Az alkalmazás nincs telepítve, ezért nem frissíthető."} +end (2026-09-13T15:30:05Z, +3s): state=not_deployed updating=False phase=None err='' hold='' +--- the controller, verbatim --- +2026/09/13 15:30:02 update.go:217: [ERROR] [stacks] update gokapi REFUSED (not_deployed): not deployed +--- AFTER --- +/mnt/sys_drive/felhom-data/backups/primary/gokapi: + 2026-09-13T06:59:17.4202088100 1014 manifest.json + 2026-09-13T06:59:17.4202088100 394 compose/app.yaml + 2026-09-13T06:59:17.4202088100 2497 compose/.felhom.yml + 2026-09-13T06:59:17.4202088100 2993 compose/docker-compose.yml +/mnt/felhom-drives/hdd_1/backups/primary/gokapi: ABSENT +/mnt/felhom-drives/hdd_1/backups/secondary/gokapi: ABSENT + +=== H done 2026-09-13T15:30:08Z === diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/06-F-setup-unit-rewritten.txt b/documentation/audits/rulings-r472-r475-2026-09-13/06-F-setup-unit-rewritten.txt new file mode 100644 index 00000000..32f28b9f --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/06-F-setup-unit-rewritten.txt @@ -0,0 +1,10 @@ +=== F setup — wait for the capture sweep to rewrite gokapi's Tier-1 unit for THIS install — 2026-09-13T15:31:51Z === +manifest mtime now: 2026-09-13 15:34:22.637661403 +0000 (+151s) +/mnt/sys_drive/felhom-data/backups/primary/gokapi: + 2026-09-13T15:34:22.6376614030 1014 manifest.json + 2026-09-13T15:34:22.6376614030 393 compose/app.yaml + 2026-09-13T15:34:22.6366613900 2497 compose/.felhom.yml + 2026-09-13T15:34:22.6366613900 2993 compose/docker-compose.yml +/mnt/felhom-drives/hdd_1/backups/primary/gokapi: ABSENT +/mnt/felhom-drives/hdd_1/backups/secondary/gokapi: ABSENT + SUBDOMAIN: gokapih5 diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/07-F-hold-names-own-drive.txt b/documentation/audits/rulings-r472-r475-2026-09-13/07-F-hold-names-own-drive.txt new file mode 100644 index 00000000..6e08df24 --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/07-F-hold-names-own-drive.txt @@ -0,0 +1,38 @@ +=== Scenario F-shape on Tier 1 — gokapi's new version never comes up — 2026-09-13T15:34:50Z, controller 0.239.0 === +Fixture: a BOX-LOCAL edit of gokapi's catalog-cache template (image -> alpine:3.20, which exits at once), +put back as soon as the job is past the pin. No catalog commit. +16: image: alpine:3.20 +--- POST /api/stacks/gokapi/update at 2026-09-13T15:34:51Z --- +HTTP 202 +{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"} + + 0s phase safety-dump + + 3s phase verifying + [fixture] template put back at 2026-09-13T15:34:55Z: 1 + +305s phase failed +end (2026-09-13T15:39:57Z, +305s): state=stopped updating=False phase=failed err='A(z) gokapi frissítése 2026-09-13 17:39-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 17:34.' hold='A(z) gokapi frissítése 2026-09-13 17:39-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 17:34.' +--- the hold sentence (Python string checks, no shell grep) --- +A(z) gokapi frissítése 2026-09-13 17:39-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 17:34. +names 'saját meghajtó': True | control 'második meghajtó': False | control 'távoli mentés': False +--- a held app refuses start --- +HTTP 409 +{"ok":false,"error":"A(z) gokapi frissítése 2026-09-13 17:39-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 17:34."} +--- the controller, verbatim --- +2026/09/13 15:34:51 update.go:336: [INFO] [stacks] update gokapi: accepted — guarded update started +2026/09/13 15:34:51 update.go:763: [INFO] [stacks] update gokapi: phase checking +2026/09/13 15:34:51 update_guard.go:173: [DEBUG] [backup] update precondition for gokapi: no Tier-2 copy (nincs másodlagos fájlmásolat ehhez az alkalmazáshoz) +2026/09/13 15:34:51 update.go:450: [INFO] [stacks] update gokapi: precondition met — Tier 1 (own recovery unit) copy from 2026-09-13T15:34:22Z (0s old, limit 24h0m0s) +2026/09/13 15:34:51 update.go:763: [INFO] [stacks] update gokapi: phase safety-dump +2026/09/13 15:34:52 update_guard.go:403: [INFO] [backup] update safety dump for gokapi: the app has no database — nothing to copy (no-op) +2026/09/13 15:34:52 update.go:482: [INFO] [stacks] update gokapi: safety dump done (0 file(s)) [] +2026/09/13 15:34:52 update.go:763: [INFO] [stacks] update gokapi: phase pinning +2026/09/13 15:34:52 pin.go:362: [INFO] [stacks] update gokapi: pin advanced to the catalog's current definition (gokapi=alpine:3.20) +2026/09/13 15:34:52 update.go:763: [INFO] [stacks] update gokapi: phase pulling +2026/09/13 15:34:54 update.go:763: [INFO] [stacks] update gokapi: phase starting +2026/09/13 15:34:54 update.go:763: [INFO] [stacks] update gokapi: phase verifying +2026/09/13 15:39:56 update.go:569: [ERROR] [stacks] update gokapi FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state restarting) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run) +2026/09/13 15:39:56 update_guard.go:468: [WARN] [backup] gokapi is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-13T15:34:22Z) +--- the hold record --- +{'stack': 'gokapi', 'at': '2026-09-13T15:39:56Z', 'reason': 'update_failed', 'copy_date': '2026-09-13T15:34:22Z', 'copy_tier': 1} +--- the template after the fixture --- +16: image: f0rc3/gokapi:v1.9.6 +=== F done 2026-09-13T15:40:00Z === diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/08-restore-from-helyi.txt b/documentation/audits/rulings-r472-r475-2026-09-13/08-restore-from-helyi.txt new file mode 100644 index 00000000..a533ad55 --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/08-restore-from-helyi.txt @@ -0,0 +1,56 @@ +=== The restore walk — from „helyi” through the restore page — 2026-09-13T15:40:16Z === +--- BEFORE: state=stopped updating=False phase=failed err='A(z) gokapi frissítése 2026-09-13 17:39-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 17:34.' hold='A(z) gokapi frissítése 2026-09-13 17:39-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 17:34.' +--- POST /backup/restore (form: stack_name=gokapi, snapshot_id=helyi — what the restore page submits) at 2026-09-13T15:40:16Z --- +HTTP/2 302 +location: /backups/restore?flash=Vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl. + +--- restore-status at the end --- +{ + "_raw": "404 page not found" +} +--- the app after the restore --- +state=starting updating=False phase=failed err='A(z) gokapi frissítése 2026-09-13 17:39-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 17:34.' hold='A(z) gokapi frissítése 2026-09-13 17:39-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 17:34.' +f0rc3/gokapi:v1.9.6 Up 4 seconds (health: starting) +pinned: {'gokapi': 'f0rc3/gokapi:v1.9.6'} | installed: None +--- the hold store --- +{'stack': 'gokapi', 'at': '2026-09-13T15:39:56Z', 'reason': 'update_failed', 'copy_date': '2026-09-13T15:34:22Z', 'copy_tier': 1} +--- the controller, verbatim --- +2026/09/13 15:40:16 auth.go:134: [DEBUG] [web] auth: valid session for POST /backup/restore +2026/09/13 15:40:16 server.go:393: [DEBUG] [web] ServeHTTP: POST /backup/restore from 172.18.0.3:33564 +2026/09/13 15:40:16 handlers.go:1480: [DEBUG] [web] backupRestoreHandler: stack=gokapi snapshot=helyi from 172.18.0.3:33564 +2026/09/13 15:40:16 handlers.go:1506: [WARN] [web] Restore requested (async): stack=gokapi, snapshot=helyi from 172.18.0.3:33564 +2026/09/13 15:40:16 restore_unit.go:313: [INFO] [backup] Restoring gokapi from recovery unit /mnt/sys_drive/felhom-data/backups/primary/gokapi: images=1, secrets recovered=1/1, data_keys=0 +2026/09/13 15:40:16 manager.go:1199: [DEBUG] [stacks] StopStack gokapi: current state=stopped deployed=true containers=0 +2026/09/13 15:40:16 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/restore-status +2026/09/13 15:40:16 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/restore-status from 172.18.0.3:33564 +2026/09/13 15:40:16 server.go:640: [WARN] [web] 404 Not Found: GET /backup/restore-status +2026/09/13 15:40:16 pin.go:93: [INFO] [stacks] pin gokapi: gokapi=f0rc3/gokapi:v1.9.6 +2026/09/13 15:40:16 manager.go:1117: [DEBUG] [stacks] StartStack gokapi: current state=stopped deployed=true +2026/09/13 15:40:16 manager.go:1127: [DEBUG] [stacks] StartStack gokapi: prepared 9 env vars for compose +2026/09/13 15:40:16 installed.go:408: [INFO] [stacks] installed-images gokapi: recorded 1 service(s) (gokapi=f0rc3/gokapi:v1.9.6 (sha256:ae9094a0ead8…)) +2026/09/13 15:40:19 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/restore-status +2026/09/13 15:40:19 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/restore-status from 172.18.0.3:33564 +2026/09/13 15:40:19 server.go:640: [WARN] [web] 404 Not Found: GET /backup/restore-status +2026/09/13 15:40:19 restore.go:212: [DEBUG] [backup] Post-restore health check: gokapi not yet running, waiting... +=== restore walk done 2026-09-13T15:40:22Z === + +=== SECOND READ — the first read above stopped after 3 s: it polled /backup/restore-status, which is 404, +=== so it saw the restore still running. This is the finished state — 2026-09-13T15:40:44Z === +GET /api/backup/restore-status -> HTTP 200 {"ok":true,"data":{"running":false,"op":"restore","stack":"gokapi","started_at":"2026-09-13T15:40:16.222440489Z","last":{"op":"restore","stack":"gokapi","ok":true,"message":"A(z) gokapi: a beállítások visszaálltak — az alkalmazás újraindult. FIGYELEM: ez a mentés csak a beállításokat tartalmazta, adatot nem. Az alkalmazás adatai NEM álltak vissza ebből a mentésből.","finished_at":"2026-09-13T15:40:24.982262899Z"},"last_recent":true}} +GET /backups/restore-status -> HTTP 404 404 page not found +GET /api/restore-status -> HTTP 404 {"ok":false,"error":"endpoint not found"} +--- the app --- +state=running updating=False phase=failed err='A(z) gokapi frissítése 2026-09-13 17:39-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 17:34.' hold='' +f0rc3/gokapi:v1.9.6 Up 29 seconds (healthy) +pinned: {'gokapi': 'f0rc3/gokapi:v1.9.6'} +--- the hold store --- +None +--- a start is no longer refused (the app is running; restart as the probe) --- +--- the controller, verbatim, restore end --- +2026/09/13 15:40:19 restore.go:212: [DEBUG] [backup] Post-restore health check: gokapi not yet running, waiting... +2026/09/13 15:40:24 restore.go:207: [DEBUG] [backup] Post-restore health check: gokapi is running +2026/09/13 15:40:24 restore_unit.go:391: [INFO] [backup] Restore-from-unit completed: gokapi — 0 volume(s) of 0 listed, 0 database(s) of 0 listed +2026/09/13 15:40:24 settings.go:1656: [INFO] [settings] restore hold CLEARED for gokapi +2026/09/13 15:40:24 update_guard.go:517: [INFO] [backup] gokapi: restore completed — the update hold (set 2026-09-13T15:39:56Z) is CLEARED +2026/09/13 15:40:24 handlers.go:1518: [INFO] [web] Restore completed (async): stack=gokapi in 8.759734524s (volumes 0/0, dbs 0/0) +=== second read done 2026-09-13T15:40:48Z === diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/09-card-after-restore.txt b/documentation/audits/rulings-r472-r475-2026-09-13/09-card-after-restore.txt new file mode 100644 index 00000000..dd7e9d5d --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/09-card-after-restore.txt @@ -0,0 +1,4 @@ +=== The card AFTER a successful restore — 2026-09-13T15:42:12Z === +API: state=running updating=False phase=failed err='A(z) gokapi frissítése 2026-09-13 17:39-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 17:34.' hold='' +GET /stacks (145958 bytes, read 15:41Z): 'leállítva marad' inside gokapi card region: True | control, actualbudget card: False +FINDING: the hold is cleared and the app runs, but update_phase=failed + update_error stay, so the card still says the app is stopped. diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/10-teardown.txt b/documentation/audits/rulings-r472-r475-2026-09-13/10-teardown.txt new file mode 100644 index 00000000..7d4bb904 --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/10-teardown.txt @@ -0,0 +1,83 @@ +=== Teardown — three layers — 2026-09-13T15:42:12Z === +--- layer 1: the two throwaway apps, through the product's own removal --- +gokapi HTTP 200 {"ok":true,"message":"Stack gokapi stop completed"} +gokapi HTTP 200 {"ok":true,"data":{"removed":"gokapi","volumes_removed":null,"hdd_paths_removed":[],"hdd_paths_preserved":[],"hdd_note":"Az alkalmazás nem tárolt saját adatot külső meghajtón, így ott nem volt mit törölni."},"message":"Stack gokapi removed"} +actualbudget HTTP 200 {"ok":true,"message":"Stack actualbudget stop completed"} +actualbudget HTTP 200 {"ok":true,"data":{"removed":"actualbudget","volumes_removed":null,"hdd_paths_removed":[],"hdd_paths_preserved":[],"hdd_note":"Az alkalmazás nem tárolt saját adatot külső meghajtón, így ott nem volt mit törölni."},"message":"Stack actualbudget removed"} +gokapi -> state=not_deployed updating=False phase=failed err='A(z) gokapi frissítése 2026-09-13 17:39-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 17:34.' hold='' + no containers +/mnt/sys_drive/felhom-data/backups/primary/gokapi: + 2026-09-13T15:34:22.6376614030 1014 manifest.json + 2026-09-13T15:34:22.6376614030 393 compose/app.yaml + 2026-09-13T15:34:22.6366613900 2497 compose/.felhom.yml + 2026-09-13T15:34:22.6366613900 2993 compose/docker-compose.yml +/mnt/felhom-drives/hdd_1/backups/primary/gokapi: ABSENT +/mnt/felhom-drives/hdd_1/backups/secondary/gokapi: ABSENT +actualbudget -> state=not_deployed updating=False phase=done err='' hold='' + no containers +/mnt/sys_drive/felhom-data/backups/primary/actualbudget: + 2026-09-13T15:26:02.2202745980 1059 manifest.json + 2026-09-13T15:26:02.2194343050 291 compose/app.yaml + 2026-09-13T15:26:02.2194343050 2259 compose/.felhom.yml + 2026-09-13T15:26:02.2194343050 1412 compose/docker-compose.yml + 2026-09-13T15:26:01.5313724220 74240 volume-dumps/actualbudget_actualbudget_data.tar +/mnt/felhom-drives/hdd_1/backups/primary/actualbudget: ABSENT +/mnt/felhom-drives/hdd_1/backups/secondary/actualbudget: ABSENT +app_backup prefs left: {'gokapi': {'enabled': False}, 'actualbudget': {'enabled': False}} +holds: None +--- layer 2: the box-local fixture --- +ls: cannot access '/root/gokapi-compose.orig': No such file or directory +16: image: f0rc3/gokapi:v1.9.6 +--- layer 3: fleet + hub, left as the release state (floor 0.239.0 with declared MinAgent 0.129.0 is the CORRECT floor now) --- +hp gitea.dooplex.hu/admin/felhom-controller:0.239.0 Up 23 minutes (healthy) +felhom-pve-lan gitea.dooplex.hu/admin/felhom-controller:0.239.0 Up 23 minutes (healthy) +standing apps: +2026/09/13 16:50:03 [INFO] managed floor SERVED for demo-felhom: floor 0.236.0, agent requirement "0.129.0" from manifest (golden 0.236.0) +2026/09/13 16:50:04 [INFO] managed floor SERVED for demo-hp: floor 0.236.0, agent requirement "0.129.0" from manifest (golden 0.236.0) +2026/09/13 17:19:14 [INFO] managed floor SERVED for demo-felhom: floor 0.239.0, agent requirement "0.129.0" from declared (golden 0.236.0) +2026/09/13 17:19:15 [INFO] managed floor SERVED for demo-hp: floor 0.239.0, agent requirement "0.129.0" from declared (golden 0.236.0) + +=== teardown done 2026-09-13T15:42:28Z === + +=== CORRECTION — the 'standing apps:' line above printed EMPTY because of a quoting slip in the teardown script (the tr argument). Re-read with a plain command — 2026-09-13T15:43:01Z === +bentopdf Up 11 hours (healthy) +bookstack Up 11 hours (healthy) +bookstack-db Up 8 hours (healthy) +calibre-web Up 11 hours (healthy) +cloudflared Up 11 days +docmost Up 11 hours (healthy) +docmost-postgres Up 11 hours (healthy) +docmost-redis Up 11 hours (healthy) +felhom-controller Up 23 minutes (healthy) +filebrowser Up 11 days (healthy) +kimai Up 11 hours (healthy) +kimai-db Up 11 hours (healthy) +opengist Up 11 hours (healthy) +privatebin Up 11 hours (healthy) +romm Up 11 hours (healthy) +romm-db Up 11 hours (healthy) +romm-redis Up 11 hours (healthy) +traefik Up 11 days + +=== CORRECTION 2 — the list above is MISSING paperless-ngx: the filter 'grep -v perl' (meant for a perl locale warning) also drops every line containing 'paperless'. Re-read filtering only the warning text — 2026-09-13T15:46:05Z === +bentopdf Up 11 hours (healthy) +bookstack Up 11 hours (healthy) +bookstack-db Up 8 hours (healthy) +calibre-web Up 11 hours (healthy) +cloudflared Up 11 days +docmost Up 11 hours (healthy) +docmost-postgres Up 11 hours (healthy) +docmost-redis Up 11 hours (healthy) +felhom-controller Up 26 minutes (healthy) +filebrowser Up 11 days (healthy) +kimai Up 11 hours (healthy) +kimai-db Up 11 hours (healthy) +opengist Up 11 hours (healthy) +paperless-postgres Up 11 hours (healthy) +paperless-redis Up 11 hours (healthy) +paperless-webserver Up 11 hours (healthy) +privatebin Up 11 hours (healthy) +romm Up 11 hours (healthy) +romm-db Up 11 hours (healthy) +romm-redis Up 11 hours (healthy) +traefik Up 11 days diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-F-gate-accepts-body-mention.txt b/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-F-gate-accepts-body-mention.txt new file mode 100644 index 00000000..8b70d795 --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-F-gate-accepts-body-mention.txt @@ -0,0 +1,17 @@ +F... +====================================================================== +FAIL: test_decoy_body_mention_is_not_the_line (__main__.MinAgentHeaderGateTest.test_decoy_body_mention_is_not_the_line) +R-421: the word in prose, a code span, or a non-bold form is the LABEL without the FACT. +---------------------------------------------------------------------- +Traceback (most recent call last): + File "/mnt/5_hdd/felhom.eu/git/felhom-controller/controller/scripts/test_minagent_header_gate.py", line 49, in test_decoy_body_mention_is_not_the_line + self.assertEqual(rc, 1, out) + ~~~~~~~~~~~~~~~~^^^^^^^^^^^^ +AssertionError: 0 != 1 : minagent header gate OK — v0.239.0 declares MinAgent 0.129.0 + + +---------------------------------------------------------------------- +Ran 4 tests in 0.124s + +FAILED (failures=1) +mutated-rc=1 diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-L-refuse-even-with-a-copy.txt b/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-L-refuse-even-with-a-copy.txt new file mode 100644 index 00000000..c48615f4 --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-L-refuse-even-with-a-copy.txt @@ -0,0 +1,11 @@ +MUTATION in /mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/stacks/update.go: + - if _, found, seen := g.RestorePoints(context.Background(), name, nil); !found { + + if _, _, seen := g.RestorePoints(context.Background(), name, nil); true { + +=== RUN TestR475_L_AnExistingCopyStillCarriesAnAppThatCannotBeBackedUp + update_tiers_test.go:118: a fresh off-site copy must carry the update even when no backup can be taken now, got "A(z) nextcloud nem frissíthető, mert nincs olyan biztonsági mentése, amelyből vissza lehetne állítani, és most új mentés sem készíthető róla. Ellenőrizd a Mentések oldalon, hogy az alkalmazás meghajtója elérhető-e — utána a frissítés elindítható." +--- FAIL: TestR475_L_AnExistingCopyStillCarriesAnAppThatCannotBeBackedUp (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.006s +FAIL +exit=1 diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-M-age-only-on-tier2.txt b/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-M-age-only-on-tier2.txt new file mode 100644 index 00000000..5a7acfbd --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-M-age-only-on-tier2.txt @@ -0,0 +1,23 @@ +MUTATION in /mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/stacks/update.go: + - return !p.ProvenAt.IsZero() && now.Sub(p.ProvenAt) <= maxAge + + return !p.ProvenAt.IsZero() && (p.Tier != UpdateTierSecondDrive || now.Sub(p.ProvenAt) <= maxAge) + +=== RUN TestR475_M_TheAgeRuleAppliesToTheChosenTier +=== RUN TestR475_M_TheAgeRuleAppliesToTheChosenTier/control:_a_stale_second-drive_copy_is_backed_up_first +=== RUN TestR475_M_TheAgeRuleAppliesToTheChosenTier/a_stale_own_unit_is_backed_up_first + update_tiers_test.go:150: backed up first = false, want true (calls [CanBackUp RestorePoints SafetyDump HoldAfterFailedUpdate]) + update_tiers_test.go:156: the chosen copy is 30h0m0s old — past the age limit +=== RUN TestR475_M_TheAgeRuleAppliesToTheChosenTier/a_stale_off-site_copy_is_backed_up_first + update_tiers_test.go:150: backed up first = false, want true (calls [CanBackUp RestorePoints SafetyDump HoldAfterFailedUpdate]) + update_tiers_test.go:153: the chosen copy is tier 3, want 1 (hold {Tier:3 ProvenAt:2026-09-12 04:00:00 +0000 UTC}, err "HELD-SENTENCE") + update_tiers_test.go:156: the chosen copy is 30h0m0s old — past the age limit +=== RUN TestR475_M_TheAgeRuleAppliesToTheChosenTier/a_stale_second_drive_does_not_block_a_fresh_own_unit +--- FAIL: TestR475_M_TheAgeRuleAppliesToTheChosenTier (0.02s) + --- PASS: TestR475_M_TheAgeRuleAppliesToTheChosenTier/control:_a_stale_second-drive_copy_is_backed_up_first (0.01s) + --- FAIL: TestR475_M_TheAgeRuleAppliesToTheChosenTier/a_stale_own_unit_is_backed_up_first (0.01s) + --- FAIL: TestR475_M_TheAgeRuleAppliesToTheChosenTier/a_stale_off-site_copy_is_backed_up_first (0.01s) + --- PASS: TestR475_M_TheAgeRuleAppliesToTheChosenTier/a_stale_second_drive_does_not_block_a_fresh_own_unit (0.01s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.032s +FAIL +exit=1 diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-offsite-clear-removed.txt b/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-offsite-clear-removed.txt new file mode 100644 index 00000000..0d6cda0f --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-offsite-clear-removed.txt @@ -0,0 +1,19 @@ +MUTATION in /mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/backup/offbox_reconstitute.go: + - m.clearUpdateHoldAfterRestore(stack) + + _ = stack + +=== RUN TestR475_OffsiteRestoreClearsAnUpdateHold +[INFO] [settings] No settings.json found, using defaults +[INFO] [settings] Settings saved +[INFO] [settings] Added storage path: /tmp/TestR475_OffsiteRestoreClearsAnUpdateHold1546124815/001 +[INFO] [settings] Settings saved +[WARN] [settings] restore hold SET for immich — the app stays stopped until it is cleared +[INFO] [settings] Settings saved +[WARN] [backup] immich is HELD STOPPED after a failed update (restore point: tier 3 "távoli mentés", 2026-09-13T11:00:00Z) +[INFO] [offbox] reconstituted immich from snapshot snap1: 3 file(s) placed, 0 volume(s) replayed, 0 DB dump(s) replayed, safety dump=., skewed=false + update_tiers_test.go:315: a successful off-site restore is the route back the hold names — the hold must be lifted +--- FAIL: TestR475_OffsiteRestoreClearsAnUpdateHold (3.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 3.008s +FAIL +exit=1 diff --git a/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-tail-no-fresh-mark.txt b/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-tail-no-fresh-mark.txt new file mode 100644 index 00000000..781efa10 --- /dev/null +++ b/documentation/audits/rulings-r472-r475-2026-09-13/rp-ctl-tail-no-fresh-mark.txt @@ -0,0 +1,12 @@ +MUTATION in /mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/backup/update_guard.go: + - os.Chtimes(RecoveryUnitManifestPath(nsRoot, stackName), now, now) + + os.Chtimes(RecoveryUnitManifestPath(nsRoot, stackName+".absent"), now, now) + +=== RUN TestR475_PreBackupTail_Tier2FailureIsAWarnAndTheOwnUnitIsFresh + update_tiers_test.go:241: the own unit must read as proven NOW (mtime 2026-09-12 11:08:10.733138645 +0200 CEST, now 2026-09-13 17:08:10.733144378 +0200 CEST m=+0.001178047) + update_tiers_test.go:249: a successful Tier-2 copy must not WARN; log = "[WARN] [backup] update pre-backup for gokapi: could not mark the recovery unit as proven current (chtimes /tmp/TestR475_PreBackupTail_Tier2FailureIsAWarnAndTheOwnUnitIsFresh2878992753/001/backups/primary/gokapi.absent/manifest.json: no such file or directory) — its own-unit copy may read older than it is\n" +--- FAIL: TestR475_PreBackupTail_Tier2FailureIsAWarnAndTheOwnUnitIsFresh (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.005s +FAIL +exit=1 diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index af5296d9..1ec3c127 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -261,3 +261,6 @@ Compressed here to title, shipping version, evidence, and the sentences that sta | **R-448** | **UPDATE ARC SLICE 4 — a guarded update: verified-backup precondition, abort path, truth at the moment of action.** Shipped controller **v0.237.0** (job) + **v0.238.0** (page) + **v0.238.1** (nightly legs skip an app mid-update, found live). Proven live on demo-hp 2026-09-13, scenarios A/B/E/F/H and the restore walk. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *the precondition is the existing verified backup, not a new copy* (ruling 2026-09-02) — `backup.Tier2UnitRestorePoint`, extracted from the backups page, not copied; *age a copy by its last SUCCESSFUL Tier-2 copy, never the manifest `created_at`* (measured: the manifest moves only on definition changes); *no automatic rollback — measured per-app, the route back is the restore*; *anything that writes a restore point skips an app that is held OR updating*. Open consequences: R-472, R-475, R-476. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` | | **R-443** | **The Update button reported success over an app it had just broken.** Closed by slice 4 (v0.237.0): 202 `accepted, not completed`; the outcome exists only as `update_phase` after health. Pinned by `TestR443_UpdateIsNeverReportedCompleteSynchronously`. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *a compose exit code is never a success signal* (spike §4: HTTP 200 over a crash loop). | **CLOSED 2026-09-13** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` | | **R-439** | **The restore hold was not honoured by the update path.** Closed by slice 4 (v0.237.0): `update` joined the router's hold check; live Scenario H refused start/restart/update and the boot sweep. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *a hold that only one path honours is not a hold* — the audit found the drive-return gate (restart + boot recreate) and the nightly volume dump ignoring any hold; all three fixed and red-proofed. | **CLOSED 2026-09-13** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` | +| **R-470** | **Four controller CHANGELOG headers (v0.233.0–v0.236.0) carried no `MinAgent:` line, while the vouch and now the declared floor read it from the header.** Closed 2026-09-13 (`felhom-controller` `f946b0d`): the four headers backfilled with `**MinAgent: 0.129.0** (unchanged)` — v0.232.0's value, proven unchanged (no commit under `internal/agentapi` since 2026-09-01; highest `featureMinAgent` 0.129.0) — and `controller/scripts/minagent_header_gate.py` (fast, blocking) refuses a newest header without the line; a prose or code-span mention does not count (decoy; red-proof F). | **CLOSED 2026-09-13 — GATED** | full text: `git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md` | +| **R-472** | **The golden cadence ruling and the hub's floor rule contradicted each other: a floor above the vouched golden delivered nothing.** Operator ruling 2026-09-13, hub **v0.112.0** (`f181efd`): a floor saved with the release's declared MinAgent is served above the golden under the same agent comparison; an undeclared one is still held and both forms refuse it (`floor_needs_min_agent`). Proven live: controller 0.239.0 reached demo-hp in 14 s and demo-felhom in 15 s from the save, hub `managed floor SERVED … from declared`. Evidence: `audits/rulings-r472-r475-2026-09-13/` 02, 03. **Reasoning kept:** *the manifest leads the floor inside the golden; above it, the release's own declared MinAgent does* (publish-train rule 1); a declaration binds to its exact floor. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md` | +| **R-475** | **The update precondition was Tier-2-only, so an app with no second-drive copy could not be updated.** Operator ruling 2026-09-13, controller **v0.239.0** (`b93c154`): the first fresh copy in the order Tier 2, Tier 1, Tier 3 (bounded, unreachable = absent + WARN); `backup_max_age` applies to the chosen tier; nothing anywhere → back up first; refused only when no copy and no backup can be taken; the hold names the tier; a Tier-2 failure in the pre-backup is a WARN. Proven live on demo-hp: nothing anywhere → backed up first (04), Tier 1 alone (05), held naming „saját meghajtó” (07), restored from „helyi” (08). Red-proof M (age only on Tier 2) fails. **Reasoning kept:** *first FRESH copy, not first copy* — a stale mirror must not force a backup while the own unit is minutes old. Follow-ups: R-477..R-480. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md` | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 2d000471..867ed8c2 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -697,15 +697,16 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** | | **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). | **READY — unblocked by R-448; rank P3-LOW; owner: CC (removal is a deliberate act)** | -| **R-470** | **[P3-LOW] Four consecutive controller CHANGELOG headers (v0.233.0 … v0.236.0) carry NO `MinAgent:` line, and `RUNBOOK-manual-build.md` §4.1 step 5 tells the vouch to read it from the header.** Measured 2026-09-13 while baking golden 0.236.0: `sed -n 1,349p CHANGELOG.md | grep -c MinAgent` → **0**; the newest statement is v0.232.0's `**MinAgent: 0.129.0** (unchanged)`. The vouch used 0.129.0 on the ground that no later entry declares a change and `internal/agentapi/features.go` gates per feature, not per release — but a reader following the runbook literally finds nothing to read, which is the R-233 shape (a document pointing at a string that is not there). **Fix:** either every release header carries the line (the v0.232.0 convention), or the runbook says where MinAgent actually lives when a header omits it. One of the two, not both. | **READY — rank P3-LOW; owner: CC** | | **R-471** | **[P3-LOW] `observations_gate.py` reads the FIRST `Observations` section of `REPORT.md` and nothing after it, so an appended second section with an unmarked item passes — and the decoy that proves it has been reading "LIVE HOLE" at HEAD, unnoticed, because the decoy suite is run by hand.** MEASURED 2026-09-13 on a clean worktree of `4b2e560`: `python3 scripts/test_gate_decoys.py` → `FAIL: observations/R-419: decoy PASSED - LIVE HOLE (rc=0)`; the gate's own output shows it scanned `## 11. Observations — noticed, documented, NOT acted on` (3 items, all marked) and never reached the appended `## Observations` block the decoy planted. Whether the cause is "first heading wins" or a heading-shape filter is NOT established — only that the planted unmarked item was not seen. **Consequence:** a report with two observation sections gets the second one unchecked. **Two fixes, both owed:** scan every section whose heading contains `Observations`, and put `test_gate_decoys.py` where something runs it (it is the instrument for R-421 and nothing in `repo_gates.py` invokes it). | **READY — rank P3-LOW; owner: CC** | -| **R-472** | **[P2-MEDIUM] THE GOLDEN CADENCE RULING AND THE HUB'S FLOOR RULE CONTRADICT EACH OTHER: a floor raise delivers NOTHING until a golden carries the release.** The 2026-09-13 ruling (R-468) rests on *"every release still raises the floor, so both demo boxes keep getting each release in ~20 s; only the golden moves to a cadence"*. MEASURED THE SAME DAY, releasing controller v0.237.0: `POST /configuration/global-floor` → `0.237.0` (303, re-read), and the hub logged for BOTH boxes `managed floor HELD for demo-hp: held: floor 0.237.0 is ABOVE the vouched golden 0.236.0, so its agent requirement is unknown — vouch a golden carrying the floor's controller (publish-train rule 1) (controller floor withheld)`; the box logged `SetFloor: floor "0.236.0" → ""`; neither box moved in 8 minutes. Publish-train rule 1 (hub `ResolveManagedFloor`) was built so a controller is never pushed past the agent it needs, and it reads that requirement from the vouched golden. **So under a weekly golden, releases between bakes do not reach the fleet by floor at all** — they need a hand deploy. v0.237.0 and v0.238.0 were hand-deployed to both demo guests by the skill's documented route and the floor was put back to 0.236.0 (a floor held on every box is a standing dashboard flag). Evidence `audits/slice4-2026-09-13/live/00-floor-0.237.0.txt`. **This is the operator's call, not CC's:** (a) bake per release again (the treadmill R-468 ended), or (b) let the floor carry a release above the golden when its CHANGELOG header states an unchanged `MinAgent` — which needs R-470's header line to be reliable first. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** | | **R-473** | **[P2-MEDIUM] The `glance` catalog template crash-loops on EVERY fresh install.** MEASURED 2026-09-13 on demo-hp (deployed as a throwaway for the slice-4 live test): `restarts=13 status=restarting`, the log repeating `parsing config: reading /app/config/glance.yml: open /app/config/glance.yml: no such file or directory`, the `glance_config` volume empty. The image does not create a default config and the template seeds none. The install page reports the deploy as started and the card then shows a restarting app. Not a version problem: the template's current tag v0.8.5 has the same config requirement (the catalog walk that pinned it never did a fresh install). **Fix shape:** seed a minimal `glance.yml` the way `gokapi`'s catalog entry seeds `config.json` (memory `gokapi-headless-setup`), then prove a fresh install lands healthy. Evidence `audits/slice4-2026-09-13/live/02-glance-abandoned.txt`. | **READY — rank P2-MEDIUM; owner: CC** | -| **R-474** | **[P3-LOW] "Remove app" with *also delete backups* deletes only the `db-dumps` directory, and reports `volumes_removed: null` over a volume it DID remove.** MEASURED 2026-09-13 removing the glance throwaway on controller v0.237.0 (`remove_hdd_data:true, remove_backups:true`): response `{"removed":"glance","volumes_removed":null,"hdd_paths_removed":[],…}`; afterwards `docker volume ls` shows no glance volume, while `backups/primary/glance/{compose,manifest.json}` and the stack dir's `applied-compose.yml` remain. The router passes ONLY `backup.AppDBDumpPath(nsRoot, name)` (router.go, beside a comment that disk-tier backup "moved to the host agent" — stale since Tier 2 returned to the controller). So the recovery unit and any Tier-2 copy survive a removal that promised to delete backups, and the answer names no volume. Same class as R-442: the removal's answer does not describe what happened. (Removal also refuses a crash-looping app as "still running — stop it first", which is correct and was observed.) **REPRODUCED a second time the same day** removing the uptime-kuma throwaway: `volumes_removed: null`, and `backups/primary/uptime-kuma/{compose,manifest.json,volume-dumps}` plus `backups/secondary/uptime-kuma/recovery-unit` survived `remove_backups:true` (residue cleared by hand; `audits/slice4-2026-09-13/` `live/15-teardown.txt`). | **READY — rank P3-LOW; owner: CC** | -| **R-475** | **[P2-MEDIUM] The update precondition is Tier-2-ONLY, so an app with no Tier-2 copy cannot be updated at all — even when its primary unit or off-site copy could restore it.** Slice 4 (v0.237.0) gates the Update button on `backup.Tier2UnitRestorePoint`, exactly as specified. MEASURED on demo-hp 2026-09-13: `gokapi` and `nextcloud` have NO `cross_drive` record, so both are refused with the no-backup sentence. The classes this reaches: a box with no second target (`no_target`), an app whose customer switched Tier 2 off, and a freshly installed app before its first Tier-2 run. **It is a finding about the specification, not a defect in the build:** the primary recovery unit (same drive) and the off-site unit (Tier 3) are both restorable routes the predicate does not consider — the driveless-app class R-356 names. Options: accept Tier 2 as the only update route (and say so in the refusal), or widen the predicate to the primary unit with an explicit "same drive" caveat. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules on the route set, CC implements** | +| **R-474** | **[P3-LOW] "Remove app" with *also delete backups* deletes only the `db-dumps` directory, and reports `volumes_removed: null` over a volume it DID remove.** MEASURED 2026-09-13 removing the glance throwaway on controller v0.237.0 (`remove_hdd_data:true, remove_backups:true`): response `{"removed":"glance","volumes_removed":null,"hdd_paths_removed":[],…}`; afterwards `docker volume ls` shows no glance volume, while `backups/primary/glance/{compose,manifest.json}` and the stack dir's `applied-compose.yml` remain. The router passes ONLY `backup.AppDBDumpPath(nsRoot, name)` (router.go, beside a comment that disk-tier backup "moved to the host agent" — stale since Tier 2 returned to the controller). So the recovery unit and any Tier-2 copy survive a removal that promised to delete backups, and the answer names no volume. Same class as R-442: the removal's answer does not describe what happened. (Removal also refuses a crash-looping app as "still running — stop it first", which is correct and was observed.) **REPRODUCED a second time the same day** removing the uptime-kuma throwaway: `volumes_removed: null`, and `backups/primary/uptime-kuma/{compose,manifest.json,volume-dumps}` plus `backups/secondary/uptime-kuma/recovery-unit` survived `remove_backups:true` (residue cleared by hand; `audits/slice4-2026-09-13/` `live/15-teardown.txt`). **REPRODUCED a third time 2026-09-13 (afternoon, v0.239.0)** removing the gokapi and actualbudget throwaways with `remove_backups:true`: both `backups/primary//` units survived (actualbudget's with its 74 KB volume tar), `volumes_removed: null` again, and both `app_backup` prefs remained (`10-teardown.txt`). It also fed R-478. | **READY — rank P3-LOW; owner: CC** | | **R-476** | **[P3-LOW] The Mentések page names a Tier-2 copy's date from the unit MANIFEST, which moves only when the app's DEFINITION changes — so it can undersell a fresh copy by a day or more.** MEASURED on demo-hp 2026-09-13: bookstack's mirror held `bookstack-mariadb.sql` written 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z; the Tier-2 run succeeded at 01:30Z. A capture rewrites the manifest only when checksums, dump NAMES or controller version change, and nightly dumps keep their names. R-403 chose the package date so a PRESERVED package is never shown as fresh — correct — but for a normal run it is the older, flattering-in-reverse date. The update (slice 4) deliberately ages the copy by the last successful copy instead (`Tier2RestorePoint.ProvenCopyTime`), so the two can name different dates. Fix shape: record the unit's DATA time (newest dump mtime) in the manifest, or name `LastSuccess` when the leg was not preserved. | **READY — rank P3-LOW; owner: CC** | +| **R-477** | **[P3-LOW] The update's Tier-3 lookup pays its full 15 s bound on demo-hp, and the bound shows up as a WARN about a DIFFERENT app.** MEASURED 2026-09-13 (controller v0.239.0, Scenario K): the `checking` phase ran 15:25:45 → 15:26:00, and the only line in the window was `[WARN] [offbox] inventory: size of kimai's newest snapshot unknown: offbox stats b4350226: signal: killed`. Cause: `UpdateRestorePoints` calls `OffsiteInventoryList`, which runs one `snapshots` AND one `stats` per app; the update needs only the snapshot times, and the 15 s context killed a `stats`. **Consequence:** every update with no fresh Tier 2/Tier 1 copy waits up to 15 s before backing up, and the operator log blames kimai for an update of actualbudget. **Fix shape:** a snapshots-only lookup for the update path (no sizes), so the bound is a guard and not the normal duration. `audits/rulings-r472-r475-2026-09-13/04-K-no-unit-anywhere.txt` | **READY — rank P3-LOW; owner: CC** | +| **R-478** | **[P3-LOW] A recovery unit LEFT BEHIND by a removed install counted as the fresh Tier-1 restore point of a reinstall.** MEASURED 2026-09-13 (v0.239.0, Scenario H): gokapi had been removed earlier; its unit (manifest 06:59:17Z) survived (R-474). A new gokapi was deployed at 15:30:46Z and updated at 15:31:10Z: `precondition met — Tier 1 (own recovery unit) copy from 2026-09-13T06:59:17Z (8h32m0s old)`. That unit belonged to the OLD install — a different `app.yaml` (subdomain, generated password). The capture sweep rewrote it only at 15:34:22Z. **Consequence:** in that window a failed update would name, and a restore would apply, the previous install's definition. **Fix shape:** R-474 deleting the unit on removal closes most of it; independently, a unit whose manifest predates the stack's current deploy should not count. `05-H-gokapi-tier1.txt`, `06-F-setup-unit-rewritten.txt` | **READY — rank P3-LOW; owner: CC** | +| **R-479** | **[P2-MEDIUM] For an app whose data is a bind mount, the Tier-1 unit holds settings only — so the route back a Tier-1 hold names restores the definition and NOT the data.** MEASURED 2026-09-13 restoring gokapi from „helyi” after a held update: `a beállítások visszaálltak … FIGYELEM: ez a mentés csak a beállításokat tartalmazta, adatot nem`. The restore message is honest; the HOLD sentence („Visszaállítható … saját meghajtó”) does not say it. The ruling (R-475) accepts any tier, in the order 2, 1, 3 — so for such an app a fresh Tier-1 unit is chosen ahead of an off-site snapshot that WOULD carry the data. **Consequence:** an update whose migration rewrote bind-mounted data has no data route back through the copy it named. **Decision-shaped:** either Tier 1 counts only for apps whose unit carries their data (DB dump / volume tar), or the hold sentence says "settings only" for that case. `07-F-hold-names-own-drive.txt`, `08-restore-from-helyi.txt` | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** | +| **R-480** | **[P2-MEDIUM] After a successful restore, the app card still shows the failed update's sentence — which says the running app „leállítva marad”.** MEASURED 2026-09-13 on demo-hp (v0.239.0): the restore cleared the hold (`restore hold CLEARED for gokapi`), gokapi ran healthy, `hold_reason` was empty — but `update_phase=failed` and `update_error` stayed, and `GET /stacks` rendered the sentence inside gokapi's card (control: absent from actualbudget's card). It survived the app's removal too (`not_deployed … phase=failed err=…`). Present since v0.238.0's page; this morning's restore walk did not check the card text. **Fix shape:** a successful restore (and a removal) clears the stack's last update outcome. `09-card-after-restore.txt`, `10-teardown.txt` | **READY — rank P2-MEDIUM; owner: CC** |