night 2026-09-25 (in progress): part 7 shipped as controller v0.271.0 — decisions 31-33, evidence A/B/C/F, register (R-685..R-687; closed R-672 R-673 R-684 R-680 R-678)
gates / gates (push) Successful in 25s
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -111,7 +111,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **A crash-looping or out-of-memory app is STOPPED by the box, the household and operator are told, and Start gives one more try** (decision 28) | controller **v0.269.0**, hub **v0.123.0** | **PROVEN-LIVE (2026-09-24)** | `audits/night-2026-09-24/A3/21-gokapi-page-event.txt`, `A3/30-start-then-trip2.*`, `E/round-01/07/08.json` | gokapi stopped at 6 restarts in 10 min; Start lifted it, 9 fast restarts, stopped again with the support-informed sentence (hu + en); chaos hour: OOM storm stopped after a power cut (+185 s), a crash loop under a backup run (+116 s), a storm on a nearly full disk (+102 s). |
|
||||
| **A held app whose page says support is informed can only be removed KEEPING its data** (decision 27) | controller **v0.269.0** | **PROVEN-LIVE (2026-09-24)** | `audits/night-2026-09-24/A2/10-held-no-copy-keep-data.*` | the dialog reads `keep_data_only`; Remove with data or backups → 409 in both languages; the app stayed installed. |
|
||||
| **The box runs the exact TESTED image of a floating tag, and an installed app keeps its image until a guarded Update moves it** (`09` §6.4 part 6, box half) | controller **v0.269.1** (`stacks/digest.go`) | **PROVEN-LIVE (2026-09-24)** | `audits/night-2026-09-24/B/21-floating-tag-0.269.1.*`, `B/01-compose-accepts-digests.txt` | redis:7-alpine: a newer tested digest → the „Frissítés elérhető" badge, the sync left the running file alone (v0.269.0's sync did not — fixed), Update pulled exactly that digest. Compose accepts `tag@digest` for all 25 ladder apps (37 digests). |
|
||||
| **Automatic updates at night — the update leg** (`09` §6.4 part 7) | — | **SPIKED, NOT BUILT (2026-09-24)** | `audits/night-2026-09-24/C/31-night.json`, `09` §6.4.2 | a caller pressing only the public Update climbed 3 apps (1–3 steps) and set aside a failing one in 7 min 18 s; three gaps filed (R-678, R-679, R-680). |
|
||||
| **Automatic updates at night — the update leg** (`09` §6.4 part 7) | controller **v0.271.0** | **PROVEN-LIVE on scratch 9202 (six simulated nights, 2026-09-24/25) — the demo boxes' first real night: see `audits/DRILL-night-2026-09-25.md` Part D** | `audits/night-2026-09-25/C/`, `B/redproofs/` | after the off-site leg on every path; one step per app per night; `needs_person` never, `files_may_change` only with a whole copy; a failed step not re-pressed until the catalog re-tests it (R-680); a power cut mid-step resumed and finished; a controller kill mid-step put back; switch off = nothing pressed; data read back after every night. **Not claimed:** W+5h with steps left and a FAILING off-site leg live (unit only), the full-system gate waiting live (9202 has no agent), resume of the leg after a restart (R-686). |
|
||||
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
|
||||
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
|
||||
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
|
||||
|
||||
Reference in New Issue
Block a user