docs: rulings 7 and 8 shipped and proven live (R-470/R-472/R-475 CLOSED); R-477..R-480 opened
gates / gates (push) Successful in 21s

Hub v0.112.0 serves a floor above the golden with a declared MinAgent;
controller v0.239.0 reached both demo boxes by that floor in 14 s and 15 s
and updates on any backup tier. 09 §3 decisions 7 and 8, §6/§6.1; 07 §6
line; capability map row; STATUS items 15/16 done and the cadence line
corrected; CONTEXT; register: R-470/R-472/R-475 compressed to CLOSED-ITEMS
(full text at 2f5d3af), R-477..R-480 opened, R-474 reproduced a third
time. OPEN-ITEMS 431689 -> 432156 bytes, CLOSED-ITEMS 118051 -> 120598.
Evidence: documentation/audits/rulings-r472-r475-2026-09-13/.
This commit is contained in:
2026-09-13 17:54:23 +02:00
parent 2f5d3af6f9
commit 681c3d6a6d
25 changed files with 735 additions and 71 deletions
@@ -100,7 +100,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|---|---|---|---|---|
| Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only |
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: `CAMPAIGN-3` proved `remove` removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets `[]` and a note.** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in `CAMPAIGN-3`; **remove WITH data: `audits/R442-2026-09-13/`**; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up** | controller **v0.237.0 + v0.238.0 + v0.238.1** | **PROVEN-LIVE (2026-09-13)** — scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | **Tier-2-only precondition** — an app with no Tier-2 copy cannot be updated (R-475); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) |
| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) |
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
@@ -270,6 +270,12 @@ Recorded here so the encryption policy is not read as covering it. → **R-108**
The tiers are **inputs to recovery**, not recovery routes. §7 and §8 say what they can actually do.
**The app update's safety precondition accepts ANY tier** (operator ruling 2026-09-13, controller
v0.239.0, R-475): the first fresh copy in the order Tier 2, Tier 1, Tier 3; with none, it backs up
first. Design: `09-update-architecture.md` §3 decision 8. **What a Tier-1 route back restores is only
what the unit holds** — for an app whose data is a bind mount that is the definition, not the data
(R-479).
### 6.1 The four tiers, as configured on the live fleet
| Tier | Location | Captures | Cadence (LIVE) | Retention (LIVE) | Encrypted |
@@ -156,6 +156,28 @@ These are rulings, not proposals. Anything specced against a different assumptio
---
### 2026-09-13 (afternoon) — the floor carries a release without a golden, and every backup counts
7. **A floor carries a release past the vouched golden when the release's MinAgent is declared with
it** (R-472; hub v0.112.0). Inside the golden the manifest's MinAgent governs, as before. Above it,
the MinAgent the operator declares with the floor — read from the release's CHANGELOG header, which
`minagent_header_gate.py` now guarantees — goes into the same per-box agent comparison. An
undeclared floor above the golden is still HELD, and both floor forms refuse to save one. **Why:** a
controller image is pulled by tag and needs no golden to be delivered; what the floor was missing
was only the agent requirement. This is what lets the weekly golden cadence (R-468) and per-release
delivery coexist. **Proven live:** both demo boxes self-updated 0.238.1 → 0.239.0 in 14 s and 15 s
from the save, the hub logging `SERVED … from declared`
(`audits/rulings-r472-r475-2026-09-13/03-declared-floor.txt`).
8. **Any backup tier lets an app update** (R-475; controller v0.239.0). The precondition takes the first
FRESH copy in the order second drive (Tier 2), the app's own recovery unit (Tier 1), off-site
(Tier 3, bounded; unreachable = absent). `update.backup_max_age` applies to whichever tier is chosen.
An app with nothing anywhere is backed up first; it is refused only when no backup can be taken
either. The hold names the tier and the date. **Tier 2 is required nowhere in the update path.**
This replaces decision 1's reading "the verified backup = the Tier-2 unit" — decision 1 itself (a
verified recent backup as a precondition, not a new copy) stands. **Proven live:**
`audits/rulings-r472-r475-2026-09-13/` 04 (nothing anywhere → backup first), 05 (Tier 1 alone),
07 (a failed update held naming „saját meghajtó"), 08 (restored from „helyi", hold cleared).
## 4. The vocabulary ruling — "rollback" is struck
**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually
@@ -301,7 +323,7 @@ earlier feature is the failure mode to look for whenever a file changes meaning.
| **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** |
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** |
| **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 |
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13)** — §6.1 |
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13); any backup tier since v0.239.0 (§3 decision 8)** — §6.1 |
| **5** | **An upgrade test that runs again** — a harness that upgrades a real app with real data in it and asks the app for the data back. | **SHIPPED, `app-catalog/scripts/upgrade-test.py` (2026-09-06)** — 7 edges, 3 apps; see §4.1 and §10 |
| **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 |
| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 |
@@ -317,9 +339,9 @@ closed by construction: nothing reports an update complete on the compose exit c
| # | phase | what happens | on failure |
|---|---|---|---|
| 0 | refusals (409, before the intent is recorded) | held (R-439), busy (backup/restore/app-data op/quiesce), migration, already updating, deploying, memory (the deploy's `memoryVerdict`, releasing the app's own request), disk (**fixed 2 GB floor** — image size unknown without a registry), **no restorable Tier-2 copy** | nothing moves, nothing is recorded |
| 1 | `checking` | re-reads the precondition | nothing moves |
| 2 | `backing-up` — only when the proven copy is older than `update.backup_max_age` | `RunAppBackupNow`: this app's DB dump → volume dump → unit capture → Tier-2 copy | refused with the backup's own error; nothing moves |
| 0 | refusals (409, before the intent is recorded) | held (R-439), busy (backup/restore/app-data op/quiesce), migration, already updating, deploying, memory (the deploy's `memoryVerdict`, releasing the app's own request), disk (**fixed 2 GB floor** — image size unknown without a registry), **no copy on ANY tier and no backup can be taken now** (since v0.239.0; before it, no restorable Tier-2 copy) | nothing moves, nothing is recorded |
| 1 | `checking` | walks Tier 2 → Tier 1 → Tier 3 for the first copy younger than `update.backup_max_age` (v0.239.0) | nothing moves |
| 2 | `backing-up` — only when no tier holds a fresh copy | `RunAppBackupNow`: this app's DB dump → volume dump → unit capture (marked proven current) → Tier-2 copy, whose failure is a WARN since v0.239.0 | refused with the backup's own error; nothing moves |
| 3 | `safety-dump` | `WriteUpdateSafetyDump` (R-361's undo copy) — **before the pin moves** | refused; nothing moves |
| 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back |
| 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) |
@@ -337,7 +359,13 @@ SUCCESSFUL Tier-2 copy, not by the unit manifest's `created_at`**, and that was
designed: a capture rewrites the manifest only when the app's DEFINITION changes, so on demo-hp
bookstack's mirror held a 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aged by the
manifest, a quiet app would be "stale" forever and a backup-first would not fix it. **The predicate is
Tier-2-only, as specified — an app with no Tier-2 copy cannot be updated (R-475).**
Tier-2-only, as specified — an app with no Tier-2 copy cannot be updated (R-475).** **SUPERSEDED in
v0.239.0 by §3 decision 8:** `backup.Manager.UpdateRestorePoints` walks all three tiers and the update
leans on the first fresh copy; `Tier2UnitRestorePoint` is still the Tier-2 half and still the page's
predicate for „Teljes visszaállítás". The same aging trap existed one tier down — a capture's checksum
skip leaves a quiet app's unit manifest untouched — so "back up first" now marks the captured unit
proven current. The hold stores `copy_tier` and names „második meghajtó" / „saját meghajtó" /
„távoli mentés"; a successful off-site restore now lifts an update hold too.
**The hold** is `settings.RestoreHold` with `reason: update_failed` and `copy_date` — the SAME store and
gate as R-379, so every start path that already honoured a restore hold honours this one. A successful
@@ -361,7 +389,8 @@ the harness has proven it, is slice 6's.
**The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the
vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0
were hand-deployed to the demo guests. That contradiction is an operator decision.
were hand-deployed to the demo guests. **RESOLVED by §3 decision 7 (hub v0.112.0):** v0.239.0 reached
both demo boxes by the floor alone, with its MinAgent declared.
**Found live, fixed in v0.238.1: the nightly legs must leave an app alone WHILE it is updating, not
only once it is held.** In Scenario F the periodic unit capture ran at 10:17:09 — inside the 5-minute