slice 3 docs: the ruling, the shipped mechanism, and four rows closed
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
09-update-architecture.md gains the fourth dated operator ruling (2026-09-06, Option 1) and its section 5 is rewritten from a proposed shape into the shipped one: the pin, the stored definition, the render table, the four writers, the startup ordering, and the trap this slice set for slice 2 - the live compose file is now the frozen one, so a badge comparing against it would answer Naprakesz on exactly the apps that are behind. 02-controller-module-map.md said 'copy compose + .felhom.yml'. That stopped being true today, so it is corrected, and the two sections describing the old seam now carry a banner saying they describe v0.234.0 and below - kept because every box under v0.235.0 still behaves that way and because they are the measured account of why it changed. R-447, R-441, R-438 and R-455 closed and compressed into CLOSED-ITEMS; R-458 opened for the .felhom.yml asymmetry, with what would settle it by measurement. Live evidence: two real catalog pushes travelling the real 15-minute cycle, both reverted, the tree byte-identical afterwards. The restart that used to take 18.3 seconds and pull a new image now takes 0.1 seconds and pulls nothing.
This commit is contained in:
@@ -100,11 +100,12 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
|---|---|---|---|---|
|
||||
| Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only |
|
||||
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. What they do to app DATA is now measured too, and it is a separate row-worth of facts (below).** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove live in `CAMPAIGN-3`; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
|
||||
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller | **PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
|
||||
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
|
||||
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
|
||||
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
|
||||
| Protected infra stacks can't be stopped/removed from UI | controller | **PROVEN-LIVE** | `CAMPAIGN-nomercy` + `RERUN-p1p3` T-SEC-PROTECTED (refuse stop/remove, stay Up) | (Cited `CAMPAIGN-2` T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal) |
|
||||
| **What VERSION a box is running, and whether it is behind the catalog** | controller **v0.234.0** + catalog `69761cf` | **PROVEN-LIVE — both the record and the rendered badge.** | **`tests/VALIDATION-update-slice12-2026-09-02.md`** — on demo-hp 0.233.0, through a REAL production caller (`bootrecon → StartStack → compose up -d → recordInstalledImages`, no hand-set state): `bentopdf` recorded **1** service and `bookstack` recorded **2**, keyed by compose SERVICE name, and **all three digests match the ground truth read independently from the containers before anything was touched**. `catalog_since` reached the box on the normal 15-minute sync. bookstack's two encrypted secrets are byte-identical across the write. Badge evidence: same file §4. **v0.234.0 startup backfill, PROVEN LIVE 2026-09-03 on demo-hp:** the record was stripped from `privatebin` (1 service) and `romm` (3 services) to recreate the pre-0.233.0 shape, the controller restarted, and the backfill re-seeded **exactly** those two — every digest matching the ground truth read from the containers beforehand — while logging `2 app(s) recorded, 7 already had a record, 0 left unrecorded`. **All 9 deployed apps then carried „Naprakész" on `/stacks`** (ASCII fragments with a negative control at 0). On demo-felhom the operator's own case, OpenGist, now renders the badge. **Why the backfill exists at all: without it the label never reached an app that simply runs**, which the operator found the morning after v0.233.0. Unit side: `installed_test.go` + `updatebadge_test.go`, incl. a wiring test through a real `RestartStack`, an AST walk of all four call sites, and three companion red-proofs. | **THE UNEXERCISED LEG, NAMED: one badge STATE of four.** „Frissítés elérhető" WITHOUT an age needs an app whose `catalog_since` is absent, malformed or future-dated, and all 53 now carry a valid one — unit-tested only (`TestGroupF`). The other three are live: „Naprakész" ×2 on `/stacks` and on `/apps/bookstack`; **NO badge at all on `/apps/docmost`**, a deployed app with no record — absent is UNKNOWN and is not rendered as current; and „Frissítés elérhető — **52 napja**" on both surfaces, the age being real arithmetic on bentopdf's `catalog_since` 2026-07-12. **The behind state was staged by editing bentopdf's compose tag ONLY** — no container restarted, no `up -d` — then reverted byte-identically (`sha256` equal, `diff` empty); so the RENDER is measured and the syncer's own half stays measured separately in the spike. Searched with `grep -oF` ASCII fragments plus positive AND negative controls. **The Frissítés/Újraindítás/Leállítás buttons are unchanged on the live page in the behind state.** **Absent means UNKNOWN, never current** — a legacy `app.yaml` renders NOTHING, red-proved. **No version number is shown to the customer** and **no registry is queried**, so „Naprakész" CAN BE FALSE for the 23 floating pins (**R-446**). Nothing about updating changed: R-438, R-440, R-441, R-443 all stand. Reasoning: `architecture/09-update-architecture.md`; remaining slices R-447..R-452 |
|
||||
| **An app's VERSION is frozen to what the customer has; only a deliberate Update moves it, while template CORRECTIONS still arrive** | controller v0.235.0 | **PROVEN-LIVE (2026-09-06)** — by two REAL catalog pushes travelling the REAL 15-minute cycle, not a hand-edited file | **`tests/VALIDATION-update-slice3-2026-09-06.md`** — a non-image catalog change REACHED the pinned app (08:01:51Z) with the container untouched; an image change did NOT (08:20:29Z); and the restart afterwards took **0.1 s, did not recreate the container, and never pulled the new image**, against the spike's **18.3 s with a pull** for the identical sequence before. The Update button still moved the version (pin advanced 17 s BEFORE the pull completed) and the teardown update returned the container to the baseline digest byte for byte. The freeze holds in BOTH directions. Unit side: `internal/sync/render_test.go` (the whole render table, incl. self-healing in BOTH branches), `internal/stacks/pin_test.go` (adoption never guesses; the update advances the pin BEFORE the pull), `TestGroupG` (the badge reads the catalog, not the frozen file), `TestGroupH` (an AST walk of `cmd/controller/main.go` asserting the seam, adoption, and their ORDER against `syncer.Start()`). **Three companion red-proofs**, each run, failing, and reverted. | Operator ruling 2026-09-06, Option 1 (`architecture/09-update-architecture.md` §3.4). **Nothing was added to the thirteen `compose up -d` call sites** — most are repairs, and a repair that refuses to repair leaves an app down; they were made safe by removing the reason. `pinned_images` is INTENT, `installed_images` is an OBSERVATION — never fed from each other (R-166, one field over). **Known limitations, all recorded rather than fixed:** a frozen app is frozen WHOLE (§8.4); `.felhom.yml` keeps flowing, so a frozen app can get a probe for a newer version — false alarm, never data loss (**R-458**); and **the Update button is still unguarded** (R-448 is slice 4). Closes R-447, R-441, R-438 |
|
||||
| Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad `backup:` block degrades to legacy, loudly) | controller v0.132, catalog | **PROVEN-LIVE** | `CAMPAIGN-2` T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs | |
|
||||
| Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) | agent v0.94.0→**v0.95.0**, controller v0.169.0→v0.171.0→**v0.215.0**, hub v0.73.1 | **PROVEN-LIVE (healthy path + delivery + the severity wire).** **IMPLEMENTED, NOT proven-live: the Hiba-from-counters path** (v0.215.0) — it has never fired on real hardware, only against the committed fixture's values in unit tests (**R-332**) | **2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix):** the card on guest 9201 now shows BOTH real disks with **real verdicts + human model labels** — **„AirDisk 512GB SSD" → Rendben (34°C)** (the system SSD, via LVM/dm resolution) and **„TOSHIBA MQ04ABF100" → Rendben (30°C)** (the USB, via union-path SMART). `/disks` carries `smart.health=PASSED` + `model_name` for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (`SPIKE-smart-coverage-2026-07-25.md` had proven both disks answer `smartctl -a -j` PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. **Notification pipeline PROVEN-LIVE 2026-07-24** — a `disk_health_degraded` POST (the exact `notify.PushEvent` wire call) was **400-rejected by hub v0.73.0** and **200-accepted + „Operator email sent" by hub v0.73.1** | No new smartctl load; feature-detect by payload presence → **MinAgent floor unchanged**; no sudoers/`-d sat` change. **No global banner** (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin `local` on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + `model_name` capture. **2026-08-14 — a genuinely failing disk HAS now been seen, and it broke three assumptions** (`audits/DIAG-smart-passed-trap-2026-08-14.md` + two committed fixtures: raw `smartctl -a -j` and 406 `smartd` lines from ST3000VX010 S/N Z6A07P2G). **(1)** `smart_status.passed` is STRUCTURALLY incapable of failing on unreadable sectors — attrs 187/197/198 all carry `thresh: 0` and a normalized value floors at 1 — so the drive read PASSED at 352 pending sectors and 1001 uncorrectable reads. **(2)** The alert it did produce carried severity `"warn"`, which the hub coerces to `info` and never emails: **the counterfactual is ZERO emails about this drive** (R-328, fixed controller v0.215.0, and the `warning`-vs-`warn` pair proven side by side in `notification_log` on 2026-08-14 — `sent` vs no row at all). **(3)** The old check spoke once and forgot on restart, so between 8 and 352 sectors it emitted nothing. v0.215.0 adds the sustained/count/heat Hiba rules, persisted state and an hourly cadence. **The verdict half of that arm remains unit+red-proof covered only** — no live drive has reached Hiba from counters (R-332). **SMART history/trending (hub-side) PARKED** (ROADMAP R-73) |
|
||||
| App crashes → customer notified (one event per transition, no flapping spam) | controller v0.120, hub v0.48 | **IMPLEMENTED** | controller v0.120.0 (dead-app alerting, `app_start_failed`, one-event-per-transition red-proofs); `CAMPAIGN-3` F11 surfaced the gap | End-to-end crash→customer-email delivery never live-confirmed (6B deferred / 6C inconclusive: clean stop ≠ crash); anti-spam unit-proven |
|
||||
|
||||
Reference in New Issue
Block a user