Slice 4 shipped (R-448/R-443/R-439 CLOSED, proven live); R-472..R-476; the floor-between-bakes claim corrected
gates / gates (push) Successful in 19s
gates / gates (push) Successful in 19s
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure. Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/). Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the runbook, STATUS, CONTEXT, R-468 and the gate docstring. Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
+20
@@ -14,6 +14,25 @@
|
||||
> language, one screen, no identifiers in the prose. Same subjects, different readers; merging them
|
||||
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
## The Update button is a guarded job, and the route back is the restore (2026-09-13, update arc slice 4 — R-448 / R-443 / R-439)
|
||||
|
||||
**Shipped** in controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (nightly-leg fix), **proven live**
|
||||
on demo-hp — `documentation/audits/slice4-2026-09-13/`, design `09-update-architecture.md` §6.1.
|
||||
|
||||
**Decisions a later session must not re-derive:**
|
||||
- **The precondition is the existing verified backup** (ruling 2026-09-02): `backup.Tier2UnitRestorePoint`,
|
||||
the predicate the destructive Tier-2 unit restore already uses — extracted, not copied. Tier-2-only by
|
||||
specification; that leaves apps with no Tier-2 copy un-updatable (R-475, operator decision).
|
||||
- **Age a copy by its last SUCCESSFUL Tier-2 copy, never by the unit manifest's `created_at`** — measured:
|
||||
the manifest moves only when the definition changes (R-476 is the page's side of this).
|
||||
- **No automatic rollback** — measured per-app in `SPIKE-upgrade-test-2026-09-06`. Health failure → stop +
|
||||
HOLD (`RestoreHold.Reason=update_failed`, same store and gate as R-379); the hold names the copy; a
|
||||
successful unit restore lifts it.
|
||||
- **Anything that writes a restore point skips an app that is held OR updating** (capture, Tier 2, volume
|
||||
dump) — the updating half was found live (v0.238.1).
|
||||
- **A release does not reach the fleet by floor between golden bakes** — the hub holds a floor above the
|
||||
vouched golden (R-472, operator decision). Slice 4's three releases were hand-deployed.
|
||||
|
||||
## Goldens move to a CADENCE, and the gate learns to read a dated waiver (2026-09-13, R-468 / R-242 / R-467)
|
||||
|
||||
**The ruling (operator, 2026-09-13):** *bake on a cadence — weekly, and always before any drill or
|
||||
@@ -22,6 +41,7 @@ design, and in August that produced **25 goldens in 26 days** plus thirteen decl
|
||||
bypasses (R-404/R-417). A guard bypassed that often teaches everyone to bypass it. Every release still
|
||||
raises the FLOOR, so the demo boxes keep getting each release in ~20 s; only the golden — which
|
||||
protects a fresh install and nothing else — moves to a cadence.
|
||||
**⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**
|
||||
|
||||
**The mechanism, because a rule without one is a wish (R-242 recurred the day after it was written):**
|
||||
`documentation/tests/golden-waiver.yml` — `issued`, `expires` (≤ 14 days, enforced by the gate's
|
||||
|
||||
@@ -1,273 +1,35 @@
|
||||
# REPORT — MariaDB finishes its own conversion, golden 0.236.0, and goldens move to a cadence (2026-09-13)
|
||||
# REPORT — update arc slice 4 documents, and a correction to this morning's golden-cadence page (2026-09-13)
|
||||
|
||||
*Overwritten each session. Nothing durable lives only here — every finding below has a register row.*
|
||||
*Overwritten each session. Nothing durable lives only here — every finding below has a register row.
|
||||
The implementation report is `felhom-controller/REPORT.md`.*
|
||||
|
||||
> **Everything the task asked for shipped, and every scenario A–H passed with the evidence quoted
|
||||
> below.** Four templates carry `MARIADB_AUTO_UPGRADE=1`; the harness saw the conversion RUN; the
|
||||
> change landed on demo-hp through the real 15-minute cycle and recreated nothing; an engine-major
|
||||
> gate holds every database engine inside its major until Slice 4; golden **0.236.0** is baked,
|
||||
> round-tripped, vouched and the floor raised; and `golden_currency_gate.py` reads a dated waiver.
|
||||
> **Three claims in the prompt turned out wrong or imprecise — named first, in §1.**
|
||||
> **The Update button is a guarded job, shipped as controller v0.237.0–v0.238.1 and proven live on
|
||||
> demo-hp.** This repo carries its documents: `09-update-architecture.md` §6.1, the register, the
|
||||
> capability map, ROADMAP, CONTEXT and STATUS. **It also carries a correction to what this repo said
|
||||
> this morning:** under the weekly golden cadence a release does NOT reach the fleet by floor.
|
||||
|
||||
## 1. Claims in the prompt that turned out wrong, named first
|
||||
## What changed here
|
||||
|
||||
1. **"a throwaway guest on demo-hp, disk on `/mnt/nvme-1tb` at its root."** `/mnt/nvme-1tb` does not
|
||||
exist on demo-hp; the 1 TB NVMe is mounted at **`/mnt/hdd_1`** (already recorded by the two
|
||||
September spikes and R-461). The guest's disk went on a dir storage at `/mnt/hdd_1`'s root, as
|
||||
those spikes did. `target-selection.md` still names the wrong path — that is R-461, not re-filed.
|
||||
2. **"the waiver is R-242's remaining half."** The gate's docstring names TWO things: the honest fix
|
||||
for a release nobody wants a golden for is "a recorded waiver, never a bypass" (built today), and
|
||||
R-242's *remaining half* is that **nothing gates the VOUCH** (unbuilt, unchanged, still open on
|
||||
R-242). The task text conflated them. The docstring, the row and this report keep them apart.
|
||||
3. **"Add the bake to the nightly-session checklist."** No document by that name exists in any repo
|
||||
(`grep -rln nightly` finds none under `documentation/runbooks/`). The step went into the two
|
||||
routines that do exist: `RUNBOOK-manual-build.md` §4.2 (the cadence) and this repo's
|
||||
`CLAUDE.md` end-of-session checklist. If a nightly checklist is created later, it points at §4.2.
|
||||
- `documentation/architecture/09-update-architecture.md` — slice 4 SHIPPED; §6.1 describes the sequence
|
||||
as shipped, the two knobs, the abort decision (not built, by measurement), the v0.238.1 finding and
|
||||
the live proof; §8.6 closed.
|
||||
- `documentation/backlog/OPEN-ITEMS.md` — **R-448, R-443, R-439 CLOSED** and compressed to
|
||||
`CLOSED-ITEMS.md`; **R-472..R-476 opened**; R-469 unblocked, not lifted. 214 open / 174 closed;
|
||||
431 689 bytes, from 439 828.
|
||||
- `documentation/architecture/00-capability-map.md` — "Update is GUARDED" row, PROVEN-LIVE, citing
|
||||
`audits/slice4-2026-09-13/`.
|
||||
- `documentation/backlog/ROADMAP.md` — slice 4 collapsed.
|
||||
- `documentation/audits/slice4-2026-09-13/` — live evidence, red-proof outputs, gate results.
|
||||
- `STATUS.md` — item 14 (plain words), items 15 and 16 (two operator decisions).
|
||||
- **The correction.** `RUNBOOK-manual-build.md` §4.2, `STATUS.md`, `CONTEXT.md`, R-468's row and the
|
||||
`golden_currency_gate.py` docstring said a release still reaches the demo boxes by floor in about
|
||||
20 seconds between bakes. **The hub holds a floor above the vouched golden**, so that is false. Each
|
||||
now says so and points at R-472. `scripts/CHANGELOG.md` records it.
|
||||
|
||||
Also imprecise, not wrong: the runbook's vouch step says to read `MinAgent` from the golden's
|
||||
controller CHANGELOG header, and **the last four headers carry no such line** — filed as **R-470**.
|
||||
## Observations
|
||||
|
||||
## 2. Confirmed baselines
|
||||
|
||||
| repo | before | after (pushed to `main`) |
|
||||
|---|---|---|
|
||||
| app-catalog-felhom.eu | `b7ef0c4a09d6` | `eec1228` templates → `bd32830` gate → `3525e35` CHANGELOG/REPORT |
|
||||
| felhom.eu | `4b2e5608c227` | `ae59c31` (documents + gate + evidence) → the compression commit on top |
|
||||
| felhom-controller | `155271672265` (v0.236.0) | **unchanged — no code.** Golden bake only. |
|
||||
|
||||
Highest R-id before: 467. Minted: **R-468** (waiver), **R-469** (engine-major rule expiry),
|
||||
**R-470** (MinAgent header line), **R-471** (observations decoy hole, pre-existing).
|
||||
|
||||
## 3. Scenario results, with the evidence quoted
|
||||
|
||||
Evidence directory: `documentation/audits/r459-close-2026-09-13/` (36 files); golden:
|
||||
`documentation/tests/golden-0.236.0-2026-09-13/`.
|
||||
|
||||
### A — the harness sees the conversion happen ✅
|
||||
|
||||
Throwaway LXC **9403** on demo-hp (Debian 13, Docker 29.8.0, Compose v5.5.1), harness copied from the
|
||||
local commit WITH the setting. `engine_state_after.bookstack-db.answer`, verbatim, on both edges:
|
||||
|
||||
```
|
||||
12.3.3-MariaDB | This installation of MariaDB is already upgraded to 12.3.3-MariaDB. There is no need to run mariadb-upgrade again. [exit=1]
|
||||
```
|
||||
|
||||
| edge | verdict | seed before / after | abort |
|
||||
|---|---|---|---|
|
||||
| E3 (app + engine 11.6→12.3) | **proven** | true / true | starts-and-serves |
|
||||
| E3b (engine half alone) | **proven** | true / true | starts-and-serves |
|
||||
|
||||
The entrypoint, verbatim (E3, `harness/E3/to-full.log`):
|
||||
|
||||
```
|
||||
[Note] [Entrypoint]: Backing up system database to system_mysql_backup_11.6.2-MariaDB.sql.zst
|
||||
[Note] [Entrypoint]: Backing up complete
|
||||
[Note] [Entrypoint]: Starting mariadb-upgrade
|
||||
Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!
|
||||
[Note] [Entrypoint]: Finished mariadb-upgrade
|
||||
```
|
||||
|
||||
`skipped due to $MARIADB_AUTO_UPGRADE`: **0** lines in both TO logs (it appeared on every start
|
||||
before this change — spike §4). Conversion 09:53:20 → 09:53:26 = **6 s**. The abort log prints
|
||||
`MariaDB upgrade not required` — the R-464 trap, quoted as an observation and not as soundness.
|
||||
|
||||
### B — the negative control still works ✅
|
||||
|
||||
C3 (`privatebin/pdo:2.0.5 → alpine:3.20`): **`failed`**, `healthy_after=false`, container
|
||||
`restarting`. Ran first, as the CLAUDE.md rule says.
|
||||
|
||||
### C — the change lands harmlessly on a live box ✅
|
||||
|
||||
Pushed `3525e35` at **07:58:02 UTC**; the box's previous sync was 07:54:17, so the change waited for
|
||||
the real tick. `live-9201/00-baseline-before-push.txt` → `01-after-sync.txt` → `02-restart.txt`.
|
||||
|
||||
The sync's own lines, verbatim:
|
||||
|
||||
```
|
||||
2026/09/13 08:09:17 sync.go:396: [INFO] [sync] Updated bookstack/docker-compose.yml
|
||||
2026/09/13 08:09:17 sync.go:409: [DEBUG] [sync] bookstack: stored definition refreshed with the delivered fix
|
||||
2026/09/13 08:09:17 sync.go:135: [INFO] [sync] Periodic sync: Sablonok frissítve — frissítve: bookstack, kimai, nextcloud, romm
|
||||
```
|
||||
|
||||
Live `docker-compose.yml` fragment from 9201 after the sync:
|
||||
|
||||
```
|
||||
bookstack-db:
|
||||
image: mariadb:12.3
|
||||
...
|
||||
- TZ=Europe/Budapest
|
||||
# MARIADB_AUTO_UPGRADE: on a MAJOR engine move the engine converts its own datadir (~7 s on a
|
||||
...
|
||||
- MARIADB_AUTO_UPGRADE=1
|
||||
```
|
||||
|
||||
**Not recreated by the sync:** both container IDs and `StartedAt` identical to the pre-push baseline
|
||||
(`bookstack` `26b555d7…` started 04:21:07Z; `bookstack-db` `b02f7a09…` started 04:21:01Z), `docker ps`
|
||||
`Up 4 hours (healthy)`. `applied-compose.yml` carries the setting too (fixes flow INTO the pin).
|
||||
|
||||
The one deliberate `POST /api/stacks/bookstack/restart` at 08:10:32 UTC → `{"ok":true,"message":"Stack
|
||||
bookstack restart completed"}`, healthy after 8 s. The engine's log after the restart, verbatim:
|
||||
|
||||
```
|
||||
2026-09-13 10:10:34+02:00 [Note] [Entrypoint]: MariaDB upgrade not required
|
||||
2026-09-13 10:10:35 0 [Note] mariadbd: ready for connections.
|
||||
```
|
||||
|
||||
and the engine asked directly (R-464): `12.3.2-MariaDB` / `This installation of MariaDB is already
|
||||
upgraded to 12.3.2-MariaDB.` `[exit=1]`; `MARIADB_AUTO_UPGRADE=1` present in the running container;
|
||||
`GET /login` → **200**. Nothing else on 9201 was touched; bentopdf stays.
|
||||
|
||||
### D — the rule exists and a gate knows it ✅
|
||||
|
||||
`app-catalog-felhom.eu/scripts/check-engine-major.py`, fourth row of `catalog_gates.py`, run by the
|
||||
pre-push hook with `--range=<remote sha>..<local sha>`. Refusal, verbatim:
|
||||
|
||||
```
|
||||
ENGINE-MAJOR GATE FAILED: templates/kimai/docker-compose.yml service kimai-db moves mariadb 11 -> 12 (mariadb:11.6 -> mariadb:12.3).
|
||||
RULE (app-catalog CLAUDE.md, operator ruling 2026-09-13): until the Update button takes a VERIFIED BACKUP as its precondition (Slice 4, felhom.eu OPEN-ITEMS.md R-448), no template may move a database-engine image across a MAJOR version.
|
||||
WHY: MariaDB sidecars now carry MARIADB_AUTO_UPGRADE=1 and WILL convert the customer's datadir on the next Update; PostgreSQL's image refuses to start on an older major's datadir (R-463). Either way this is a customer-data event with no backup in front of it.
|
||||
EXPIRY: this rule is removed DELIBERATELY when R-448 ships — the removal is its own register row, not a silent edit. Until then, keep the engine within its major.
|
||||
```
|
||||
|
||||
Pass, verbatim (this push's own range): `engine-major gate OK — no database engine crosses a major
|
||||
version (rule: CLAUDE.md, until Slice 4 / R-448 ships)`. Red-proof (`engine-major-gate/redproof-7-cases.txt`):
|
||||
`11.6→12.3` rc=1, `postgres:16→17` rc=1, `mariadb:lts` rc=2, `11.6→11.8` rc=0, and three decoys
|
||||
(comment + `serverVersion=` env, the app's own image, README) rc=0. **CI cannot run it** — the
|
||||
runner fetches at `--depth 1`; the runner announces the skip on a shallow clone (pinned by
|
||||
`test_catalog_gates.py`). Same gap as R-452, not re-filed.
|
||||
|
||||
### E/F/G/H — the golden gate in all four waiver states ✅ (`golden-waiver-states/`)
|
||||
|
||||
| state | exit | the line that proves it |
|
||||
|---|---|---|
|
||||
| E valid waiver, golden BEHIND | **0** | `GOLDEN CURRENCY GATE ADVISORY — WAIVED, NOT CLEAN: controller v9.9.9 is released and NO golden carries it (newest bake is 9.9.7; 2 release(s) behind: 9.9.8, 9.9.9).` … `Waived by R-468 until 2026-09-26 (13 day(s) left)` |
|
||||
| F expired waiver, behind | **1** | `The waiver at documentation/tests/golden-waiver.yml EXPIRED on 2026-09-12 (R-468). It ran out, as a dated waiver is meant to` |
|
||||
| G valid waiver, golden UNRECORDED | **1** | `A waiver exists (R-468, until 2026-09-26) and DOES NOT COVER THIS: it covers a golden that is behind the record, never one that is unrecorded (R-385).` |
|
||||
| H 15-day waiver | **2** | `the waiver … is MALFORMED — `expires:` is 15 days after `issued:` — the hard limit is 14.` |
|
||||
| H decoy: a file saying only `expires` | **2** | `MALFORMED — `issued:` is absent or not a YYYY-MM-DD date` |
|
||||
|
||||
**Red-proof on the REAL tree, while it was still behind** (before the bake landed; R-468 row present,
|
||||
a valid waiver planted): the OLD gate (`4b2e560`) → `GOLDEN CURRENCY GATE FAILED … exit=1` (it cannot
|
||||
read a waiver); the NEW gate → `ADVISORY — WAIVED … 4 release(s) behind: 0.233.0, 0.234.0, 0.235.0,
|
||||
0.236.0 … exit=0`. Tests: `scripts/test_golden_currency_gate.py` cases 5–15, all green
|
||||
(19 cases in the file now, 4 before).
|
||||
|
||||
**The first waiver, as committed** (`documentation/tests/golden-waiver.yml`, issued AFTER the bake):
|
||||
|
||||
```yaml
|
||||
issued: 2026-09-13
|
||||
expires: 2026-09-27
|
||||
reason: pre-customer development; goldens on a weekly cadence (operator ruling 2026-09-13)
|
||||
register_row: R-468
|
||||
```
|
||||
|
||||
Gate on the real tree now: `newest released controller : 0.236.0` / `newest golden baked : 0.236.0`
|
||||
/ `waiver : VALID until 2026-09-27 (14 day(s) left)` → **OK, exit 0**. Before the bake it read
|
||||
`FAILED … 4 release(s) behind`, exit 1 (`golden-waiver-states/00-real-tree-before-bake-no-waiver.txt`).
|
||||
|
||||
## 4. The golden (R-467) — version and vouch evidence
|
||||
|
||||
`GOLDEN_VERSION=0.236.0`, `GOLDEN_SHA256=58a3cc24c61dc2271f2cb508ef5a269af97005f386b671f18011258df5f958bf`,
|
||||
654 115 664 B. Three readers agreed (bake log, round trip hashed from the DOWNLOADED bytes, the hub's
|
||||
dropdown); `./etc/felhom-controller-image` out of the archive says `felhom-controller:0.236.0`;
|
||||
markers 1/1/1/1, `excluding` 0, `FATAL` 0; token-leak grep 0 with the seeded control 1; the bake
|
||||
script's fingerprint equal across the hop. Vouch: `golden_version 0.232.0 → 0.236.0`, `agent_version`
|
||||
0.130.0 and `min_agent` 0.129.0 unchanged, re-read from the page, R-120 banner absent; floor
|
||||
`0.232.0 → 0.236.0` re-read. **No box moved — both already ran 0.236.0 by hand**; the chain was last
|
||||
exercised 2026-09-01 and nothing about it changed. This bake carried four unbaked releases and is
|
||||
the last per-release one. **No `--no-verify` anywhere in this session.**
|
||||
|
||||
## 5. Files created / modified
|
||||
|
||||
**app-catalog-felhom.eu** (pushed `3525e35`): `templates/{bookstack,kimai,nextcloud,romm}/docker-compose.yml`,
|
||||
`scripts/check-engine-major.py` (new), `scripts/test_gate_decoys.py` (new), `scripts/catalog_gates.py`,
|
||||
`scripts/test_catalog_gates.py`, `.githooks/pre-push`, `CLAUDE.md`, `REUSE.md`, `CHANGELOG.md`, `REPORT.md`.
|
||||
|
||||
**felhom.eu**: `scripts/golden_currency_gate.py`, `scripts/test_golden_currency_gate.py`,
|
||||
`scripts/CHANGELOG.md`, `documentation/tests/golden-waiver.yml` (new),
|
||||
`documentation/tests/golden-0.236.0-2026-09-13/` (new, 12 files),
|
||||
`documentation/audits/r459-close-2026-09-13/` (new, 36 files),
|
||||
`documentation/runbooks/RUNBOOK-manual-build.md` (§4.2), `CLAUDE.md` (checklist step),
|
||||
`documentation/architecture/09-update-architecture.md` (§3 decisions 5 and 6, §8.7, engines section),
|
||||
`documentation/backlog/OPEN-ITEMS.md`, `documentation/backlog/CLOSED-ITEMS.md`, `CONTEXT.md`, `STATUS.md`, this file.
|
||||
|
||||
## 6. Tests
|
||||
|
||||
| suite | before | after |
|
||||
|---|---|---|
|
||||
| `app-catalog/scripts/test_catalog_gates.py` | 5 | **7**, green |
|
||||
| `app-catalog/scripts/test_gate_decoys.py` | — | **7 cases**, green |
|
||||
| `felhom.eu/scripts/test_golden_currency_gate.py` | 4 | **19 cases**, green |
|
||||
| `felhom.eu/scripts/repo_gates.py --fast` | 14 gates | 14 gates, **all OK** |
|
||||
| `felhom.eu/scripts/decoy_coverage_gate.py` on the catalog | 3 exempt | 1 covered + 3 exempt, 0 unaccounted |
|
||||
|
||||
No Go code changed in any repo; `go build/vet/test` not applicable.
|
||||
|
||||
`python3 scripts/unproven.py --summary`: 55 claims, walked 20 / partial 17 / built 14 / missing 4,
|
||||
**NOT WALKED 35 of 55 — unchanged** (nothing in this session walked a capability claim).
|
||||
|
||||
**CI, by `head_sha`:** app-catalog job **536** for `3525e35` → `completed` / `success`; felhom.eu run
|
||||
**537** (run_number 322) for `4727aaa` → `completed` / `success`, started 08:15:52Z. The previous run
|
||||
534 for `4b2e560` (the earlier session's `--no-verify` push) was `failure`, as that session reported —
|
||||
the bake made it green again.
|
||||
|
||||
## 7. Register — rows opened / closed, size
|
||||
|
||||
Closed: **R-459** (harness + live evidence), **R-467** (the bake). Narrowed: **R-242** (waiver half built;
|
||||
vouch half open). Opened: **R-468** (WATCHING — the waiver; renew ≤ 14 days or bake; retire at the
|
||||
first external install), **R-469** (BLOCKED on R-448 — remove the engine-major rule), **R-470**
|
||||
(READY — MinAgent header line), **R-471** (READY — observations decoy hole, pre-existing).
|
||||
Register size: **210 open rows / 169 closed before → 212 open / 171 closed after** (R-459 and R-467
|
||||
compressed to `CLOSED-ITEMS.md` in the second commit, citing `ae59c31` for the full text; nothing deleted).
|
||||
`OPEN-ITEMS.md`: 433 845 → 434988 bytes with the four new rows and two closures written out, → 429692 bytes after compression.
|
||||
|
||||
## 8. Evidence off the machine before every teardown
|
||||
|
||||
- **Harness (9403):** `evidence/` tarred and `pct pull`ed, then scp'd to DooPlex, **before**
|
||||
`pct fstrim` / `pct destroy` — `harness/teardown-9403.txt`.
|
||||
- **Bake (drill VM):** `bake.log` scp'd out and markers counted **before** `pct destroy 9100`, the
|
||||
`shred -u` and the `poweroff` — `05-teardown.txt`, leftovers in the VM: 0.
|
||||
- **Live (9201):** baseline captured **before** the push; after-sync and after-restart captured at
|
||||
the moment; nothing on 9201 was reverted.
|
||||
|
||||
## 9. Teardown — three layers
|
||||
|
||||
| layer | before | after |
|
||||
|---|---|---|
|
||||
| **1 — machines** | LXC 9403 on demo-hp (scratch dir storage `scratch-r459c` at `/mnt/hdd_1`); LXC 9100 inside the drill VM | 9403 **destroyed**, storage **removed**, template deleted, `pct list` shows only 9201; 9100 **destroyed**, drill VM **powered off**, `qemu` confirmed exited (`ps -eo comm`), disk reverted to the single `virgin` snapshot |
|
||||
| **2 — hosts** | demo-hp `local-lvm` **35.32 %**, `local` 20 899 060 KiB, `/mnt/hdd_1` 5 774 620 KiB | `local-lvm` **35.32 % — never touched**; `local` 20 908 828 KiB (+9.5 MB); `/mnt/hdd_1` **5 774 620 KiB — identical**; `pct fstrim 9403` returned 55.5 GiB; guest peak 3.3 GB (2.21 GB images). DooPlex `/mnt/5_hdd` 37 % before and after |
|
||||
| **3 — hub** | 2 enrolled hosts | **Checked, not asserted** (`harness/teardown-layer3-hub.txt`): `/hosts` lists exactly `demo-felhom-8363b5` and `demo-hp-bb76ea`; **this run created no customer, no appliance and no host record** — 9403 never enrolled, 9100 is the bake fixture. The hub's ONLY change is the vouch + floor, which is the deliverable. |
|
||||
|
||||
Helper files on demo-hp (`ctrl_pw`, `dh_user`, `dh_pat`, `upg.tgz`, `guest-setup.sh`, logs) were
|
||||
`shred -u`'d / removed; the Docker Hub login inside 9403 was logged out before the guest was destroyed.
|
||||
|
||||
## 10. NOT live-validated — stated plainly
|
||||
|
||||
- **The waiver in the "behind" state on the real tree, after this push:** it is dormant (the golden is
|
||||
current). Its real-tree behaviour was proven **before** the bake (§3 E, red-proof); after today it is
|
||||
exercised the next time a release ships without a bake — which is the intended cadence.
|
||||
- **The self-update chain for 0.236.0:** not re-exercised (both boxes were already there). Last proof
|
||||
2026-09-01.
|
||||
- **Kimai, Nextcloud, RomM under an actual engine major move:** the setting is inert until their pins
|
||||
move, and the engine-major rule forbids that until Slice 4. Only bookstack has a measurable edge.
|
||||
- **The engine-major gate in CI:** cannot run there (`--depth 1`); the hook is the enforcement.
|
||||
|
||||
## 11. Observations — noticed, documented, NOT acted on
|
||||
|
||||
1. **`felhom.eu/scripts/test_gate_decoys.py` reports a LIVE HOLE at HEAD, before this session:**
|
||||
`observations/R-419: decoy PASSED (rc=0)` — the gate scans the report's first `Observations`
|
||||
section and not an appended second one. Reproduced on a clean worktree of `4b2e560`. **FILED: R-471.**
|
||||
2. **Four consecutive controller CHANGELOG headers carry no `MinAgent:` line** while the vouch runbook
|
||||
says to read it from the header. **FILED: R-470.**
|
||||
3. **The stack restart recreated only the changed service** — bookstack-db got a new ID, the bookstack
|
||||
app container kept its 4-hour uptime — because the restart is compose up -d, which the update
|
||||
architecture §1.3 records as a chosen behaviour.
|
||||
**NOT-A-FINDING: documented design, and "restart completed" is accurate for a compose-level restart.**
|
||||
4. **`mariadb:12.3` resolved to 12.3.3 in the harness and 12.3.2 on 9201** — the floating-pin class.
|
||||
**NOT-A-FINDING: already R-446.**
|
||||
5. **`target-selection.md` names `/mnt/nvme-1tb`, which does not exist on demo-hp.** **NOT-A-FINDING:
|
||||
already R-461, and it says exactly this.**
|
||||
6. **The session ran with permission prompts disabled**, which the workspace `CLAUDE.md` says not to do
|
||||
on this host. Nothing outside the workspace, `~/build`, the drill directory and the two Tier-0
|
||||
boxes was touched; every host-side act is listed in §9. **NOT-A-FINDING: an operator setting, not
|
||||
a product defect; recorded so it is not hidden.**
|
||||
1. **The golden cadence and the hub's floor rule contradict each other.** **FILED: R-472.**
|
||||
2. **The glance catalog template crash-loops on a fresh install.** **FILED: R-473.**
|
||||
3. **App removal with "delete backups" leaves the recovery unit and Tier-2 copy.** **FILED: R-474.**
|
||||
4. **The update precondition is Tier-2-only.** **FILED: R-475.**
|
||||
5. **The Mentések page dates a copy by its manifest, which lags the data.** **FILED: R-476.**
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-13 (third pass) — the Update button now takes a backup first and tells the truth. It
|
||||
is live on both machines and I walked every case on the HP, including putting a broken update back
|
||||
from its backup. TWO THINGS NEED YOU: item 15 (the demo machines no longer get new releases by
|
||||
themselves between golden images) and item 16 (apps with no second-drive copy cannot be updated).**
|
||||
|
||||
**Updated 2026-09-13 (second pass) — you decided both open items. The database engine now finishes
|
||||
its own conversion on the four MariaDB apps; the upgrade machine proved it and it landed on the HP
|
||||
without a ripple. Goldens are now weekly and before any install, not per release, and the gate
|
||||
@@ -230,6 +235,49 @@ nothing.*
|
||||
|
||||
13. **"Delete my data too" now deletes the data — or tells you it could not.** Until today, when a customer removed an app and ticked the box, the box said it worked and left everything on the drive (128 MB of a Nextcloud on 2026-09-01). The cause: the removal asked one global setting for the drive, and no machine fills that setting in. Every other part of the box already asks the app itself where its data is. Now the removal does too. If the box cannot work out where the data is, it refuses and keeps the app, so you can try again — it never again reports success over data left behind. Proven on the HP with a throwaway Nextcloud: 63 MB the app wrote itself was gone after removal, and the answer listed it; the refusal was shown with the app still in place; an app with no drive data gets a plain "nothing to delete" note. Live on both machines (0.236.0). No standing app was touched. **If you do nothing:** nothing to do.
|
||||
|
||||
14. **The Update button now takes a backup first, and tells the truth.** Nothing needs you.
|
||||
Before today, pressing Update pulled the new version at once, checked nothing, and said
|
||||
"done" while the app could already be crashing. Now:
|
||||
- It refuses first if it must: the app is held, a backup is running, or memory or disk is short.
|
||||
It also refuses if the app has no backup it could be put back from.
|
||||
- If the backup is older than a day, it makes a fresh one first.
|
||||
- It waits for the app to actually be healthy before it says it worked. The button shows each step.
|
||||
- If the new version does not come up, the app is stopped and held. The page names the backup to
|
||||
restore it from. The box never puts the old version back by itself, because we measured that
|
||||
this works for some apps and breaks others.
|
||||
|
||||
**Proven on the HP with a throwaway app:** a real upgrade, a stale backup, a version that does not
|
||||
exist, a version that never starts, and the restore from the named backup back to the old version.
|
||||
**One problem showed up during the test and is fixed:** while an update was waiting to see if the
|
||||
app came up, the regular backup copied the broken version into the app's local backup. It now
|
||||
leaves an app alone while it is updating. Live as 0.238.1.
|
||||
|
||||
15. **The demo machines no longer get new releases by themselves between golden images. One decision.**
|
||||
This morning we agreed to bake the golden image weekly, on the promise that every release still
|
||||
reaches both machines in about 20 seconds. **That promise was wrong.** The hub will not move a
|
||||
machine to a release newer than the golden image, so between bakes I install releases on the demo
|
||||
machines by hand. Today I did that three times.
|
||||
- **Bake a golden image for every release again.** Releases reach the machines by themselves. The
|
||||
cost is the weekly routine we just ended.
|
||||
- **Let the hub move machines past the golden image when a release says it needs no newer agent.**
|
||||
Releases reach the machines by themselves and goldens stay weekly. It needs one small hub change,
|
||||
and first each release must reliably say which agent it needs (R-470).
|
||||
|
||||
**My pick: the second.** **If you do nothing:** nothing breaks; I keep installing by hand and
|
||||
saying so each time.
|
||||
|
||||
16. **Apps with no copy on a second drive cannot be updated. One decision.** The update's safety rule
|
||||
asks for the app's second-drive copy, as we decided on 2 September. On the HP, two apps (gokapi
|
||||
and nextcloud) have no such copy, so their Update button refuses. A box with only one drive would
|
||||
refuse every update.
|
||||
- **Keep it.** Only apps with a second-drive copy can be updated. The refusal tells the customer how
|
||||
to switch it on.
|
||||
- **Also accept the backup on the same drive.** Updates work on one-drive boxes. That backup is lost
|
||||
if that drive dies.
|
||||
|
||||
**My pick: keep it for now,** and revisit when the first one-drive customer exists. **If you do
|
||||
nothing:** those apps show the refusal and nothing else changes.
|
||||
|
||||
8. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
||||
as it misled one by an hour.
|
||||
@@ -239,8 +287,10 @@ nothing.*
|
||||
- **GOLDENS ARE NOW WEEKLY AND BEFORE ANY INSTALL, NOT PER RELEASE — AND THE GATE KNOWS. DECIDED 2026-09-13.**
|
||||
In August I baked 25 goldens in 26 days, almost one per release, because the check trips on every
|
||||
release on purpose. From today: one golden a week, and always before a drill or a fresh install.
|
||||
Every release still reaches both demo machines in about 20 seconds — only the image a **new**
|
||||
machine starts from moves to a cadence. The check now reads a dated permission slip that runs out
|
||||
**Correction, same day:** I wrote that every release would still reach both demo machines in about
|
||||
20 seconds between bakes. **That is wrong.** The hub will not move the machines past the image a new
|
||||
machine starts from, so between bakes I install each release on the demo machines by hand. That is
|
||||
item 15 below, and it is yours to decide. The check now reads a dated permission slip that runs out
|
||||
after at most 14 days; while it is valid the check warns instead of refusing, and when it runs out
|
||||
the check is red again until someone bakes or renews. A dated slip cannot be forgotten — it just
|
||||
expires. Today's golden (0.236.0) is baked, checked three ways, and live.
|
||||
|
||||
@@ -100,6 +100,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
|---|---|---|---|---|
|
||||
| Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only |
|
||||
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: `CAMPAIGN-3` proved `remove` removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets `[]` and a note.** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in `CAMPAIGN-3`; **remove WITH data: `audits/R442-2026-09-13/`**; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
|
||||
| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up** | controller **v0.237.0 + v0.238.0 + v0.238.1** | **PROVEN-LIVE (2026-09-13)** — scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | **Tier-2-only precondition** — an app with no Tier-2 copy cannot be updated (R-475); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) |
|
||||
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
|
||||
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
|
||||
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
|
||||
|
||||
@@ -301,11 +301,80 @@ earlier feature is the failure mode to look for whenever a file changes meaning.
|
||||
| **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** |
|
||||
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** |
|
||||
| **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 |
|
||||
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 |
|
||||
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13)** — §6.1 |
|
||||
| **5** | **An upgrade test that runs again** — a harness that upgrades a real app with real data in it and asks the app for the data back. | **SHIPPED, `app-catalog/scripts/upgrade-test.py` (2026-09-06)** — 7 edges, 3 apps; see §4.1 and §10 |
|
||||
| **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 |
|
||||
| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 |
|
||||
|
||||
### 6.1 Slice 4 as SHIPPED (controller v0.237.0 + v0.238.0, 2026-09-13)
|
||||
|
||||
`POST /api/stacks/{name}/update` is a guarded job. It answers **202** at once; the outcome exists only
|
||||
on `GET /api/stacks/{name}` (`updating`, `update_phase`, `update_phase_label`, `update_error`,
|
||||
`hold_reason`), and `update_phase=done` is written only after the app's health is known. **R-443 is
|
||||
closed by construction: nothing reports an update complete on the compose exit code.**
|
||||
|
||||
**The sequence, and the order is the design:**
|
||||
|
||||
| # | phase | what happens | on failure |
|
||||
|---|---|---|---|
|
||||
| 0 | refusals (409, before the intent is recorded) | held (R-439), busy (backup/restore/app-data op/quiesce), migration, already updating, deploying, memory (the deploy's `memoryVerdict`, releasing the app's own request), disk (**fixed 2 GB floor** — image size unknown without a registry), **no restorable Tier-2 copy** | nothing moves, nothing is recorded |
|
||||
| 1 | `checking` | re-reads the precondition | nothing moves |
|
||||
| 2 | `backing-up` — only when the proven copy is older than `update.backup_max_age` | `RunAppBackupNow`: this app's DB dump → volume dump → unit capture → Tier-2 copy | refused with the backup's own error; nothing moves |
|
||||
| 3 | `safety-dump` | `WriteUpdateSafetyDump` (R-361's undo copy) — **before the pin moves** | refused; nothing moves |
|
||||
| 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back |
|
||||
| 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) |
|
||||
| 6 | `starting` | `compose up -d --remove-orphans` | stop + HOLD |
|
||||
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe, or 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
|
||||
| 8 | `done` | installed images recorded, journal cleared | — |
|
||||
|
||||
**The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and
|
||||
`update.health_timeout` (default `5m`).
|
||||
|
||||
**The precondition is the existing verified backup, not a new copy** (§3 decision 1). It is
|
||||
`backup.Tier2UnitRestorePoint` — the SAME predicate that permits the destructive „Teljes
|
||||
visszaállítás", extracted from the backups page rather than copied. **The copy is aged by the last
|
||||
SUCCESSFUL Tier-2 copy, not by the unit manifest's `created_at`**, and that was measured before it was
|
||||
designed: a capture rewrites the manifest only when the app's DEFINITION changes, so on demo-hp
|
||||
bookstack's mirror held a 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aged by the
|
||||
manifest, a quiet app would be "stale" forever and a backup-first would not fix it. **The predicate is
|
||||
Tier-2-only, as specified — an app with no Tier-2 copy cannot be updated (R-475).**
|
||||
|
||||
**The hold** is `settings.RestoreHold` with `reason: update_failed` and `copy_date` — the SAME store and
|
||||
gate as R-379, so every start path that already honoured a restore hold honours this one. A successful
|
||||
unit restore lifts an update hold (only that kind). **Three unattended paths honoured no hold before
|
||||
v0.237.0 and now do:** the drive-return gate's restart and boot recreate, and the nightly volume dump
|
||||
(which ends in `StartStack`). The nightly capture and Tier-2 run skip a held app, so the restore point
|
||||
the hold text names is never overwritten.
|
||||
|
||||
**Crash safety is a journal**, `<data>/update-journal.json`, written before every phase. `RecoverUpdates`
|
||||
runs before the boot sweep: interrupted before the pin → dropped; while pinning/pulling → pin put back;
|
||||
after `up` → marked Updating (the boot sweep and the dead-app alarm leave it alone) and resumed by
|
||||
`ResumeInterruptedUpdates` once the backup side is wired, ending healthy or held.
|
||||
|
||||
**THE ABORT DECISION, restated so it is not reopened: NOT BUILT, BY MEASUREMENT.** Whether an old image
|
||||
starts on data a new one migrated is per-app (§4.1: PrivateBin yes, Docmost and Nextcloud no) and cannot
|
||||
be predicted. So the box never puts the old version back by itself. **The route back is the restore**,
|
||||
and slice 4's whole purpose is that the restore exists before anything moves. Per-app abort data, where
|
||||
the harness has proven it, is slice 6's.
|
||||
|
||||
**Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6.
|
||||
|
||||
**The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the
|
||||
vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0
|
||||
were hand-deployed to the demo guests. That contradiction is an operator decision.
|
||||
|
||||
**Found live, fixed in v0.238.1: the nightly legs must leave an app alone WHILE it is updating, not
|
||||
only once it is held.** In Scenario F the periodic unit capture ran at 10:17:09 — inside the 5-minute
|
||||
health wait, 53 s before the hold — and wrote the never-started definition into the app's PRIMARY unit.
|
||||
The Tier-2 mirror the hold names survived only because Tier 2 is daily. `backup.Manager.isHeld` now also
|
||||
answers true for an app a guarded update is moving (`SetUpdatingCheck`).
|
||||
|
||||
**Proven live on demo-hp, 2026-09-13**, with a throwaway uptime-kuma and real catalog tag changes (each
|
||||
reverted in the same phase): A (2.3.2→2.4.0, done after health), B (`backup_max_age: 2m` → backup first),
|
||||
E (non-existent tag → pin back, container untouched), F (`alpine:3.20` → held), H (three buttons and the
|
||||
boot sweep refuse the held app), and the restore walk (Mentések unit restore → back on 2.4.0, hold
|
||||
cleared). Live evidence: `audits/slice4-2026-09-13/`.
|
||||
|
||||
### The verdict record — the contract Slice 6 carries
|
||||
|
||||
Decided here rather than invented twice. The harness writes one of these per edge, beside its
|
||||
@@ -458,9 +527,11 @@ Version strings stay in the logs, the API and the hub.
|
||||
§5.4. So a frozen app can receive a health check written for a NEWER version and read as degraded.
|
||||
**The failure direction is a false alarm, never data loss**, and freezing `.felhom.yml` would break
|
||||
the update badge by withholding `catalog_since`. **R-458.**
|
||||
6. **The Update button is still unguarded.** It takes no backup, has no rollback, and can still
|
||||
attempt a multi-major jump the app will refuse (R-40). **Slice 3 did not change that and must not
|
||||
be read as having done so** — the precondition is slice 4 (R-448).
|
||||
6. ~~**The Update button is still unguarded.**~~ **CLOSED 2026-09-13 by slice 4 (v0.237.0, §6.1).** It
|
||||
refuses without a restorable, proven Tier-2 copy, backs up first when that copy is stale, takes a
|
||||
safety dump, and holds an app that does not come up. **What stays true:** it still has no automatic
|
||||
rollback (deliberately, §6.1) and can still attempt a multi-major jump the app will refuse (R-40) —
|
||||
that now ends HELD rather than crash-looping behind a green button.
|
||||
7. ~~**An engine major can be applied without its datadir upgrade, and nothing notices.**~~ **CLOSED
|
||||
2026-09-13 for MariaDB (R-459):** every `mariadb:` sidecar carries `MARIADB_AUTO_UPGRADE=1`, and the
|
||||
harness shows the conversion running on the E3/E3b edges (§3 decision 5). **What stays true:** the
|
||||
|
||||
@@ -0,0 +1,28 @@
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 10.311s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/agentapi 0.178s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.316s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/appbackup 0.020s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/appexport 0.971s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 316.772s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/backupwindow 0.003s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/bootrecon 0.007s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/bootstrap 14.023s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/channelhealth 0.004s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/config 0.005s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/fillwatch 0.007s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/infra 0.009s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/mailrelay 0.015s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/metrics 0.008s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/monitor 0.004s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/notify 0.778s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/offsiteapply 0.010s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.076s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/report 2.578s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/scheduler 0.154s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/selfupdate 0.022s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/settings 0.009s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/setup 0.009s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.540s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/sync 0.015s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/system 0.104s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 29.840s
|
||||
@@ -0,0 +1,28 @@
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 10.293s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/agentapi 0.190s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.298s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/appbackup 0.020s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/appexport 1.023s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 318.161s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/backupwindow 0.006s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/bootrecon 0.008s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/bootstrap 14.025s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/channelhealth 0.008s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/config 0.006s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/fillwatch 0.009s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/infra 0.009s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/mailrelay 0.015s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/metrics 0.008s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/monitor 0.006s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/notify 0.797s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/offsiteapply 0.012s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.079s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/report 2.578s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/scheduler 0.156s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/selfupdate 0.025s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/settings 0.013s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/setup 0.010s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.544s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/sync 0.019s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/system 0.094s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 29.766s
|
||||
@@ -0,0 +1,28 @@
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 10.306s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/agentapi 0.190s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.360s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/appbackup 0.022s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/appexport 1.127s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 317.940s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/backupwindow 0.006s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/bootrecon 0.008s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/bootstrap 14.027s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/channelhealth 0.005s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/config 0.006s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/fillwatch 0.009s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/infra 0.009s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/mailrelay 0.016s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/metrics 0.011s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/monitor 0.005s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/notify 0.817s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/offsiteapply 0.012s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.082s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/report 2.591s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/scheduler 0.156s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/selfupdate 0.027s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/settings 0.012s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/setup 0.010s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.655s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/sync 0.023s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/system 0.106s
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 29.955s
|
||||
@@ -0,0 +1,16 @@
|
||||
=== floor raise 2026-09-13T09:43:21Z ===
|
||||
POST http=303
|
||||
Location: /configuration?flash=floor_set
|
||||
name="min_controller_version" value="0.237.0"
|
||||
=== fleet 2026-09-13T09:51:44Z ===
|
||||
demo-hp: gitea.dooplex.hu/admin/felhom-controller:0.236.0 Up 3 hours (healthy)
|
||||
demo-felhom: gitea.dooplex.hu/admin/felhom-controller:0.236.0 Up 3 hours (healthy)
|
||||
=== the hub HELD the 0.237.0 floor on every box (hub log, verbatim) ===
|
||||
2026/09/13 11:43:21 [INFO] Global controller-version floor set to "0.237.0"
|
||||
2026/09/13 11:43:23 [INFO] managed floor HELD for demo-felhom: held: floor 0.237.0 is ABOVE the vouched golden 0.236.0, so its agent requirement is unknown — vouch a golden carrying the floor's controller (publish-train rule 1) (controller floor withheld)
|
||||
2026/09/13 11:43:24 [INFO] managed floor HELD for demo-hp: held: floor 0.237.0 is ABOVE the vouched golden 0.236.0, so its agent requirement is unknown — vouch a golden carrying the floor's controller (publish-train rule 1) (controller floor withheld)
|
||||
|
||||
=== floor put back to the vouched golden 0.236.0 2026-09-13T09:52:43Z ===
|
||||
POST http=303
|
||||
Location: /configuration?flash=floor_set
|
||||
name="min_controller_version" value="0.236.0"
|
||||
@@ -0,0 +1,4 @@
|
||||
=== hand-deploy 0.237.0 on felhom-pve-lan guest 9201 2026-09-13T09:52:50Z (skill route: pull, write the image file, restart the bootstrap unit) ===
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.237.0
|
||||
docker ps: gitea.dooplex.hu/admin/felhom-controller:0.237.0 Up 6 seconds (healthy)
|
||||
--- startup lines ---
|
||||
@@ -0,0 +1,6 @@
|
||||
=== hand-deploy 0.237.0 on hp guest 9201 2026-09-13T09:52:50Z (skill route: pull, write the image file, restart the bootstrap unit) ===
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.237.0
|
||||
docker ps: gitea.dooplex.hu/admin/felhom-controller:0.237.0 Up 7 seconds (healthy)
|
||||
--- startup lines ---
|
||||
2026/09/13 09:52:55 client.go:67: [DEBUG] [agentapi] agent version seen: 0.130.0 (was "")
|
||||
2026/09/13 09:52:59 builder.go:40: [DEBUG] [report] BuildReport: starting — version=0.237.0, storagePaths=1
|
||||
@@ -0,0 +1,29 @@
|
||||
=== glance abandoned as the test app 2026-09-13T09:59:15Z: its catalog template crash-loops on a fresh install (no glance.yml seeded) ===
|
||||
restarts=13 status=restarting
|
||||
parsing config: reading /app/config/glance.yml: open /app/config/glance.yml: no such file or directory
|
||||
parsing config: reading /app/config/glance.yml: open /app/config/glance.yml: no such file or directory
|
||||
--- removal through the feature (R-442), data + backups ---
|
||||
HTTP 409
|
||||
{"ok":false,"error":"stack \"glance\" is still running — stop it first before removing"}
|
||||
|
||||
--- after ---
|
||||
glance Restarting (1) 44 seconds ago
|
||||
glance_glance_config
|
||||
app.yaml
|
||||
applied-compose.yml
|
||||
docker-compose.yml
|
||||
compose
|
||||
manifest.json
|
||||
--- stop, then remove (data + backups) 2026-09-13T09:59:42Z ---
|
||||
HTTP 200
|
||||
{"ok":true,"message":"Stack glance stop completed"}
|
||||
|
||||
|
||||
HTTP 200
|
||||
{"ok":true,"data":{"removed":"glance","volumes_removed":null,"hdd_paths_removed":[],"hdd_paths_preserved":[],"hdd_note":"Az alkalmazás nem tárolt saját adatot külső meghajtón, így ott nem volt mit törölni."},"message":"Stack glance removed"}
|
||||
|
||||
|
||||
--- after ---
|
||||
volumes:
|
||||
stackdir: applied-compose.yml docker-compose.yml
|
||||
primary: compose manifest.json
|
||||
@@ -0,0 +1,32 @@
|
||||
=== L1 setup 2026-09-13T09:54:46Z: catalog a1f1c38 carries glance v0.8.4 ===
|
||||
--- sync via POST /api/sync (the dashboard's 'Sablonok frissítése' — the syncer's own entry point) ---
|
||||
HTTP 200
|
||||
{"ok":true,"data":{"ok":true,"updated":["glance"],"message":"Sablonok frissítve — frissítve: glance"},"message":"Sablonok frissítve — frissítve: glance"}
|
||||
|
||||
|
||||
image: glanceapp/glance:v0.8.4
|
||||
--- deploy glance (throwaway) ---
|
||||
HTTP 202
|
||||
{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
|
||||
|
||||
state after deploy:
|
||||
glanceapp/glance:v0.8.4 Restarting (1) 35 seconds ago
|
||||
Traceback (most recent call last):
|
||||
File "<string>", line 1, in <module>
|
||||
import json,sys; d=json.load(sys.stdin)["data"]; print("pinned_images:", (d.get("app_config") or {}).get("pinned_images"))
|
||||
~~~~~~~~~^^^^^^^^^^^
|
||||
File "/usr/lib/python3.13/json/__init__.py", line 293, in load
|
||||
return loads(fp.read(),
|
||||
cls=cls, object_hook=object_hook,
|
||||
parse_float=parse_float, parse_int=parse_int,
|
||||
parse_constant=parse_constant, object_pairs_hook=object_pairs_hook, **kw)
|
||||
File "/usr/lib/python3.13/json/__init__.py", line 346, in loads
|
||||
return _default_decoder.decode(s)
|
||||
~~~~~~~~~~~~~~~~~~~~~~~^^^
|
||||
File "/usr/lib/python3.13/json/decoder.py", line 345, in decode
|
||||
obj, end = self.raw_decode(s, idx=_w(s, 0).end())
|
||||
~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^
|
||||
File "/usr/lib/python3.13/json/decoder.py", line 363, in raw_decode
|
||||
raise JSONDecodeError("Expecting value", s, err.value) from None
|
||||
json.decoder.JSONDecodeError: Expecting value: line 2 column 1 (char 1)
|
||||
--- wait for the periodic capture to write glance's primary unit ---
|
||||
@@ -0,0 +1,30 @@
|
||||
=== L1 setup 2026-09-13T09:59:55Z: catalog 01c631d carries uptime-kuma 2.3.2 ===
|
||||
--- sync via POST /api/sync ---
|
||||
HTTP 200
|
||||
{"ok":true,"data":{"ok":true,"updated":["glance","uptime-kuma"],"message":"Sablonok frissítve — frissítve: glance, uptime-kuma"},"message":"Sablonok frissítve — frissítve: glance, uptime-kuma"}
|
||||
|
||||
|
||||
image: louislam/uptime-kuma:2.3.2
|
||||
--- deploy uptime-kuma (throwaway) ---
|
||||
HTTP 202
|
||||
{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
|
||||
|
||||
state after deploy: running (2026-09-13T10:00:57Z)
|
||||
louislam/uptime-kuma:2.3.2 Up 13 seconds (healthy)
|
||||
pinned_images: {'uptime-kuma': 'louislam/uptime-kuma:2.3.2'}
|
||||
--- wait for the periodic capture to write the primary unit ---
|
||||
total 8
|
||||
drwxr-xr-x 2 root root 4096 2026-09-13T10:02:54Z compose
|
||||
-rw-r--r-- 1 root root 1008 2026-09-13T10:02:54Z manifest.json
|
||||
--- Tier-2 via POST /api/backup/tier2 2026-09-13T10:03:03Z ---
|
||||
HTTP 200
|
||||
{"ok":true,"message":"2. mentés elindítva"}
|
||||
|
||||
|
||||
uptime-kuma cross_drive: 2026-09-13T10:03:07Z ok /mnt/felhom-drives/hdd_1
|
||||
total 8
|
||||
drwxr-xr-x 2 root root 4096 2026-09-13T10:02:54Z compose
|
||||
-rw-r--r-- 1 root root 1008 2026-09-13T10:02:54Z manifest.json
|
||||
--- the card's view ---
|
||||
state running catalog_images {'uptime-kuma': 'louislam/uptime-kuma:2.3.2'} installed {'uptime-kuma': {'ref': 'louislam/uptime-kuma:2.3.2', 'digest': 'sha256:9aeb4e51d038047f414309c77a1af553281ca535723cb88907d907269d0a908e', 'at': '2026-09-13T10:00:45Z'}}
|
||||
=== L1 done 2026-09-13T10:03:16Z ===
|
||||
@@ -0,0 +1,13 @@
|
||||
=== Scenario A catalog change 2026-09-13T10:07:28Z: revert 01c631d — uptime-kuma 2.3.2 -> 2.4.0 ===
|
||||
image: louislam/uptime-kuma:2.4.0
|
||||
catalog_since: "2026-07-18"
|
||||
all catalog gates OK
|
||||
pre-push [app-catalog-felhom.eu]: gates OK - push proceeding.
|
||||
catalog HEAD=28ce33b
|
||||
--- sync ---
|
||||
HTTP 200
|
||||
{"ok":true,"data":{"ok":true,"updated":["uptime-kuma"],"message":"Sablonok frissítve — frissítve: uptime-kuma"},"message":"Sablonok frissítve — frissítve: uptime-kuma"}
|
||||
|
||||
|
||||
image: louislam/uptime-kuma:2.3.2
|
||||
(the live file must STILL name 2.3.2 — the app is frozen until Update)
|
||||
@@ -0,0 +1,78 @@
|
||||
=== Scenario A — a real upgrade 2.3.2 -> 2.4.0 with a fresh proven copy — 2026-09-13T10:07:31Z ===
|
||||
--- BEFORE ---
|
||||
{
|
||||
"state": "running",
|
||||
"updating": false,
|
||||
"update_phase": null,
|
||||
"update_phase_label": null,
|
||||
"update_error": null,
|
||||
"hold_reason": null,
|
||||
"pinned_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.3.2"
|
||||
},
|
||||
"installed_images": {
|
||||
"uptime-kuma": {
|
||||
"ref": "louislam/uptime-kuma:2.3.2",
|
||||
"digest": "sha256:9aeb4e51d038047f414309c77a1af553281ca535723cb88907d907269d0a908e",
|
||||
"at": "2026-09-13T10:00:45Z"
|
||||
}
|
||||
},
|
||||
"catalog_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
|
||||
}
|
||||
}
|
||||
louislam/uptime-kuma:2.3.2 Up 6 minutes (healthy)
|
||||
--- POST /api/stacks/uptime-kuma/update at 2026-09-13T10:07:32Z ---
|
||||
HTTP 202
|
||||
{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"}
|
||||
|
||||
10:07:32 U safety-dump
|
||||
10:07:34 U pulling
|
||||
10:08:01 U starting
|
||||
10:08:16 U verifying
|
||||
10:08:20 - done
|
||||
--- AFTER 2026-09-13T10:08:20Z ---
|
||||
{
|
||||
"state": "running",
|
||||
"updating": false,
|
||||
"update_phase": "done",
|
||||
"update_phase_label": "Frissítve",
|
||||
"update_error": null,
|
||||
"hold_reason": null,
|
||||
"pinned_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
|
||||
},
|
||||
"installed_images": {
|
||||
"uptime-kuma": {
|
||||
"ref": "louislam/uptime-kuma:2.4.0",
|
||||
"digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985",
|
||||
"at": "2026-09-13T10:08:19Z"
|
||||
}
|
||||
},
|
||||
"catalog_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
|
||||
}
|
||||
}
|
||||
louislam/uptime-kuma:2.4.0 Up 7 seconds (healthy)
|
||||
--- the controller, verbatim (update + backup lines since 2026-09-13T10:07:32Z) ---
|
||||
2026/09/13 10:07:32 router.go:567: [INFO] [api] update requested for stack: uptime-kuma
|
||||
2026/09/13 10:07:32 update.go:278: [INFO] [stacks] update uptime-kuma: accepted — guarded update started
|
||||
2026/09/13 10:07:32 update.go:701: [INFO] [stacks] update uptime-kuma: phase checking
|
||||
2026/09/13 10:07:32 update.go:409: [INFO] [stacks] update uptime-kuma: precondition met — proven copy from 2026-09-13T10:03:07Z (4m0s old, limit 24h0m0s)
|
||||
2026/09/13 10:07:32 update.go:701: [INFO] [stacks] update uptime-kuma: phase safety-dump
|
||||
2026/09/13 10:07:32 update_guard.go:239: [INFO] [backup] update safety dump for uptime-kuma: the app has no database — nothing to copy (no-op)
|
||||
2026/09/13 10:07:32 update.go:423: [INFO] [stacks] update uptime-kuma: safety dump done (0 file(s)) []
|
||||
2026/09/13 10:07:32 update.go:701: [INFO] [stacks] update uptime-kuma: phase pinning
|
||||
2026/09/13 10:07:32 pin.go:362: [INFO] [stacks] update uptime-kuma: pin advanced to the catalog's current definition (uptime-kuma=louislam/uptime-kuma:2.4.0)
|
||||
2026/09/13 10:07:32 update.go:701: [INFO] [stacks] update uptime-kuma: phase pulling
|
||||
2026/09/13 10:08:01 update.go:701: [INFO] [stacks] update uptime-kuma: phase starting
|
||||
2026/09/13 10:08:14 update.go:701: [INFO] [stacks] update uptime-kuma: phase verifying
|
||||
2026/09/13 10:08:19 update.go:499: [INFO] [stacks] update uptime-kuma: healthy after 5s (the app's health check passed)
|
||||
2026/09/13 10:08:19 update.go:505: [INFO] [stacks] update uptime-kuma: DONE in 47s
|
||||
--- the stacks page card (badge / hold / phase fragments) ---
|
||||
1 Frissítés
|
||||
1 Naprakész
|
||||
1 stackAction(event, 'uptime-kuma', 'restart')
|
||||
1 stackAction(event, 'uptime-kuma', 'stop')
|
||||
1 stackAction(event, 'uptime-kuma', 'update')
|
||||
=== Scenario A — a real upgrade 2.3.2 -> 2.4.0 with a fresh proven copy done ===
|
||||
@@ -0,0 +1,6 @@
|
||||
=== Scenario B setup 2026-09-13T10:09:30Z: the operator knob update.backup_max_age set to 2m on demo-hp (a REAL config setting, not a test hook), reverted at the end of this phase ===
|
||||
042b71a3123d70da
|
||||
|
||||
update:
|
||||
backup_max_age: 2m
|
||||
controller: gitea.dooplex.hu/admin/felhom-controller:0.237.0 Up 16 seconds (healthy)
|
||||
@@ -0,0 +1,100 @@
|
||||
=== Scenario B — the proven copy is older than update.backup_max_age (2m): back up first — 2026-09-13T10:09:48Z ===
|
||||
--- BEFORE ---
|
||||
{
|
||||
"state": "unhealthy",
|
||||
"updating": false,
|
||||
"update_phase": null,
|
||||
"update_phase_label": null,
|
||||
"update_error": null,
|
||||
"hold_reason": null,
|
||||
"pinned_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
|
||||
},
|
||||
"installed_images": {
|
||||
"uptime-kuma": {
|
||||
"ref": "louislam/uptime-kuma:2.4.0",
|
||||
"digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985",
|
||||
"at": "2026-09-13T10:08:19Z"
|
||||
}
|
||||
},
|
||||
"catalog_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
|
||||
}
|
||||
}
|
||||
louislam/uptime-kuma:2.4.0 Up About a minute (healthy)
|
||||
--- POST /api/stacks/uptime-kuma/update at 2026-09-13T10:09:49Z ---
|
||||
HTTP 202
|
||||
{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"}
|
||||
|
||||
10:09:50 U backing-up
|
||||
10:09:52 U pulling
|
||||
10:09:54 U verifying
|
||||
10:09:58 - done
|
||||
--- AFTER 2026-09-13T10:09:58Z ---
|
||||
{
|
||||
"state": "running",
|
||||
"updating": false,
|
||||
"update_phase": "done",
|
||||
"update_phase_label": "Frissítve",
|
||||
"update_error": null,
|
||||
"hold_reason": null,
|
||||
"pinned_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
|
||||
},
|
||||
"installed_images": {
|
||||
"uptime-kuma": {
|
||||
"ref": "louislam/uptime-kuma:2.4.0",
|
||||
"digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985",
|
||||
"at": "2026-09-13T10:08:19Z"
|
||||
}
|
||||
},
|
||||
"catalog_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
|
||||
}
|
||||
}
|
||||
louislam/uptime-kuma:2.4.0 Up 8 seconds (healthy)
|
||||
--- the controller, verbatim (update + backup lines since 2026-09-13T10:09:49Z) ---
|
||||
2026/09/13 10:09:49 router.go:567: [INFO] [api] update requested for stack: uptime-kuma
|
||||
2026/09/13 10:09:49 update.go:278: [INFO] [stacks] update uptime-kuma: accepted — guarded update started
|
||||
2026/09/13 10:09:49 update.go:701: [INFO] [stacks] update uptime-kuma: phase checking
|
||||
2026/09/13 10:09:49 update.go:394: [INFO] [stacks] update uptime-kuma: the proven copy is 7m0s old (limit 2m0s) — backing up first
|
||||
2026/09/13 10:09:49 update.go:701: [INFO] [stacks] update uptime-kuma: phase backing-up
|
||||
2026/09/13 10:09:49 update_guard.go:143: [INFO] [backup] update pre-backup for uptime-kuma: starting (DB dump → volume dump → unit capture → Tier 2)
|
||||
2026/09/13 10:09:51 update_guard.go:192: [INFO] [backup] update pre-backup for uptime-kuma: volume dump OK
|
||||
2026/09/13 10:09:51 update_guard.go:198: [INFO] [backup] update pre-backup for uptime-kuma: recovery unit captured (0 database dump(s))
|
||||
2026/09/13 10:09:51 tier2.go:424: [INFO] [backup] Tier 2 copied uptime-kuma → /mnt/felhom-drives/hdd_1/backups/secondary/uptime-kuma (19.9 KB, 0 leg(s), 0s)
|
||||
2026/09/13 10:09:51 update_guard.go:214: [INFO] [backup] update pre-backup for uptime-kuma: complete in 1.395s
|
||||
2026/09/13 10:09:51 update.go:701: [INFO] [stacks] update uptime-kuma: phase safety-dump
|
||||
2026/09/13 10:09:51 update_guard.go:239: [INFO] [backup] update safety dump for uptime-kuma: the app has no database — nothing to copy (no-op)
|
||||
2026/09/13 10:09:51 update.go:423: [INFO] [stacks] update uptime-kuma: safety dump done (0 file(s)) []
|
||||
2026/09/13 10:09:51 update.go:701: [INFO] [stacks] update uptime-kuma: phase pinning
|
||||
2026/09/13 10:09:51 pin.go:362: [INFO] [stacks] update uptime-kuma: pin advanced to the catalog's current definition (uptime-kuma=louislam/uptime-kuma:2.4.0)
|
||||
2026/09/13 10:09:51 update.go:701: [INFO] [stacks] update uptime-kuma: phase pulling
|
||||
2026/09/13 10:09:52 update.go:701: [INFO] [stacks] update uptime-kuma: phase starting
|
||||
2026/09/13 10:09:53 update.go:701: [INFO] [stacks] update uptime-kuma: phase verifying
|
||||
2026/09/13 10:09:58 update.go:499: [INFO] [stacks] update uptime-kuma: healthy after 5s (the app's health check passed)
|
||||
2026/09/13 10:09:58 update.go:505: [INFO] [stacks] update uptime-kuma: DONE in 8s
|
||||
--- the stacks page card (badge / hold / phase fragments) ---
|
||||
1 Frissítés
|
||||
1 Naprakész
|
||||
1 stackAction(event, 'uptime-kuma', 'restart')
|
||||
1 stackAction(event, 'uptime-kuma', 'stop')
|
||||
1 stackAction(event, 'uptime-kuma', 'update')
|
||||
=== Scenario B — the proven copy is older than update.backup_max_age (2m): back up first done ===
|
||||
--- the copy after the backup-first leg ---
|
||||
cross_drive last_success 2026-09-13T10:09:51Z status ok
|
||||
/mnt/sys_drive/felhom-data/backups/primary/uptime-kuma/:
|
||||
total 12
|
||||
drwxr-xr-x 2 root root 4096 2026-09-13T10:09:51Z compose
|
||||
-rw-r--r-- 1 root root 1048 2026-09-13T10:09:51Z manifest.json
|
||||
drwxr-xr-x 2 root root 4096 2026-09-13T10:09:50Z volume-dumps
|
||||
|
||||
/mnt/sys_drive/felhom-data/backups/primary/uptime-kuma/volume-dumps/:
|
||||
total 4
|
||||
-rw-r--r-- 1 root root 3072 2026-09-13T10:09:50Z uptime-kuma_uptime_kuma_data.tar
|
||||
total 4
|
||||
-rw-r--r-- 1 root root 3072 2026-09-13T10:09:50Z uptime-kuma_uptime_kuma_data.tar
|
||||
=== revert the knob 2026-09-13T10:10:04Z ===
|
||||
042b71a3123d70da
|
||||
1
|
||||
controller: gitea.dooplex.hu/admin/felhom-controller:0.237.0 Up 16 seconds (healthy)
|
||||
@@ -0,0 +1,11 @@
|
||||
=== Scenario E setup 2026-09-13T10:11:01Z: catalog 29a8cbe names louislam/uptime-kuma:2.4.999 (does not resolve) ===
|
||||
--- BEFORE ---
|
||||
container: cb541381d93ef9a7a9d19f06c31e6b75d74cea70f3d11c0da8f9cd5a51cdd4fb louislam/uptime-kuma:2.4.0 started=2026-09-13T10:09:50.950051403Z
|
||||
live: image: louislam/uptime-kuma:2.4.0
|
||||
applied: image: louislam/uptime-kuma:2.4.0
|
||||
pre-update-copies: 0
|
||||
--- sync ---
|
||||
HTTP 200
|
||||
{"ok":true,"data":{"ok":true,"updated":["uptime-kuma"],"message":"Sablonok frissítve — frissítve: uptime-kuma"},"message":"Sablonok frissítve — frissítve: uptime-kuma"}
|
||||
|
||||
|
||||
@@ -0,0 +1,86 @@
|
||||
=== Scenario E — the pull fails (the catalog names a tag that does not exist) — 2026-09-13T10:11:03Z ===
|
||||
--- BEFORE ---
|
||||
{
|
||||
"state": "running",
|
||||
"updating": false,
|
||||
"update_phase": null,
|
||||
"update_phase_label": null,
|
||||
"update_error": null,
|
||||
"hold_reason": null,
|
||||
"pinned_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
|
||||
},
|
||||
"installed_images": {
|
||||
"uptime-kuma": {
|
||||
"ref": "louislam/uptime-kuma:2.4.0",
|
||||
"digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985",
|
||||
"at": "2026-09-13T10:08:19Z"
|
||||
}
|
||||
},
|
||||
"catalog_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.999"
|
||||
}
|
||||
}
|
||||
louislam/uptime-kuma:2.4.0 Up About a minute (healthy)
|
||||
--- POST /api/stacks/uptime-kuma/update at 2026-09-13T10:11:04Z ---
|
||||
HTTP 202
|
||||
{"ok":true,"data":{"accepted":true,"completed":false},"message":"FrissÃtés elindult – az állapot a kártyán követhetÅ‘"}
|
||||
|
||||
10:11:04 U safety-dump
|
||||
10:11:06 - failed
|
||||
--- AFTER 2026-09-13T10:11:06Z ---
|
||||
{
|
||||
"state": "running",
|
||||
"updating": false,
|
||||
"update_phase": "failed",
|
||||
"update_phase_label": "A frissÃtés nem sikerült",
|
||||
"update_error": "Az új verzió letöltése nem sikerült, ezért a frissÃtés elmaradt. Az alkalmazás a korábbi verzióval fut tovább.",
|
||||
"hold_reason": null,
|
||||
"pinned_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
|
||||
},
|
||||
"installed_images": {
|
||||
"uptime-kuma": {
|
||||
"ref": "louislam/uptime-kuma:2.4.0",
|
||||
"digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985",
|
||||
"at": "2026-09-13T10:08:19Z"
|
||||
}
|
||||
},
|
||||
"catalog_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.999"
|
||||
}
|
||||
}
|
||||
louislam/uptime-kuma:2.4.0 Up About a minute (healthy)
|
||||
--- the controller, verbatim (update + backup lines since 2026-09-13T10:11:04Z) ---
|
||||
2026/09/13 10:11:04 router.go:567: [INFO] [api] update requested for stack: uptime-kuma
|
||||
2026/09/13 10:11:04 update.go:278: [INFO] [stacks] update uptime-kuma: accepted — guarded update started
|
||||
2026/09/13 10:11:04 update.go:701: [INFO] [stacks] update uptime-kuma: phase checking
|
||||
2026/09/13 10:11:04 update.go:409: [INFO] [stacks] update uptime-kuma: precondition met — proven copy from 2026-09-13T10:09:51Z (1m0s old, limit 24h0m0s)
|
||||
2026/09/13 10:11:04 update.go:701: [INFO] [stacks] update uptime-kuma: phase safety-dump
|
||||
2026/09/13 10:11:04 update_guard.go:239: [INFO] [backup] update safety dump for uptime-kuma: the app has no database — nothing to copy (no-op)
|
||||
2026/09/13 10:11:04 update.go:423: [INFO] [stacks] update uptime-kuma: safety dump done (0 file(s)) []
|
||||
2026/09/13 10:11:04 update.go:701: [INFO] [stacks] update uptime-kuma: phase pinning
|
||||
2026/09/13 10:11:04 pin.go:362: [INFO] [stacks] update uptime-kuma: pin advanced to the catalog's current definition (uptime-kuma=louislam/uptime-kuma:2.4.999)
|
||||
2026/09/13 10:11:04 update.go:701: [INFO] [stacks] update uptime-kuma: phase pulling
|
||||
2026/09/13 10:11:06 update.go:551: [INFO] [stacks] update uptime-kuma: pin and definition PUT BACK to the pre-update version (uptime-kuma=louislam/uptime-kuma:2.4.0)
|
||||
2026/09/13 10:11:06 update.go:469: [ERROR] [stacks] update uptime-kuma: pull failed — pin and definition PUT BACK; the app was not touched. Docker said: exit code 1
|
||||
2026/09/13 10:11:06 update.go:373: [ERROR] [stacks] update uptime-kuma FAILED in phase pulling after 1.411s — nothing was moved: pull failed: exit code 1
|
||||
--- the stacks page card (badge / hold / phase fragments) ---
|
||||
1 FrissÃtés
|
||||
1 FrissÃtés gombot.">FrissÃtés elérhetÅ‘ â
|
||||
1 stackAction(event, 'uptime-kuma', 'restart')
|
||||
1 stackAction(event, 'uptime-kuma', 'stop')
|
||||
1 stackAction(event, 'uptime-kuma', 'update')
|
||||
=== Scenario E — the pull fails (the catalog names a tag that does not exist) done ===
|
||||
--- AFTER: the app untouched, the pin and definition put back ---
|
||||
container: cb541381d93ef9a7a9d19f06c31e6b75d74cea70f3d11c0da8f9cd5a51cdd4fb louislam/uptime-kuma:2.4.0 started=2026-09-13T10:09:50.950051403Z
|
||||
live: image: louislam/uptime-kuma:2.4.0
|
||||
applied: image: louislam/uptime-kuma:2.4.0
|
||||
pre-update-copies: 0
|
||||
--- the pull's own error, verbatim, from the log ---
|
||||
2026/09/13 10:11:06 manager.go:1394: [ERROR] [stacks] stderr: Image louislam/uptime-kuma:2.4.999 Pulling
|
||||
2026/09/13 10:11:06 update.go:551: [INFO] [stacks] update uptime-kuma: pin and definition PUT BACK to the pre-update version (uptime-kuma=louislam/uptime-kuma:2.4.0)
|
||||
2026/09/13 10:11:06 update.go:469: [ERROR] [stacks] update uptime-kuma: pull failed — pin and definition PUT BACK; the app was not touched. Docker said: exit code 1
|
||||
stderr: Image louislam/uptime-kuma:2.4.999 Pulling
|
||||
2026/09/13 10:11:06 update.go:373: [ERROR] [stacks] update uptime-kuma FAILED in phase pulling after 1.411s — nothing was moved: pull failed: exit code 1
|
||||
stderr: Image louislam/uptime-kuma:2.4.999 Pulling
|
||||
@@ -0,0 +1,3 @@
|
||||
=== hand-deploy 0.238.0 on felhom-pve-lan guest 9201 2026-09-13T10:12:05Z (the floor is held above the golden — R-472) ===
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.238.0
|
||||
docker ps: gitea.dooplex.hu/admin/felhom-controller:0.238.0 Up 6 seconds (healthy)
|
||||
@@ -0,0 +1,3 @@
|
||||
=== hand-deploy 0.238.0 on hp guest 9201 2026-09-13T10:12:05Z (the floor is held above the golden — R-472) ===
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.238.0
|
||||
docker ps: gitea.dooplex.hu/admin/felhom-controller:0.238.0 Up 7 seconds (healthy)
|
||||
@@ -0,0 +1,9 @@
|
||||
=== Scenario F setup 2026-09-13T10:12:56Z: catalog 6ce3f65 names alpine:3.20 for uptime-kuma; controller 0.238.0 ===
|
||||
--- units BEFORE ---
|
||||
primary-unit: image: louislam/uptime-kuma:2.4.0 manifest "created_at": "2026-09-13T10:12:09Z"
|
||||
tier2-mirror: image: louislam/uptime-kuma:2.4.0 manifest "created_at": "2026-09-13T10:09:51Z"
|
||||
--- sync ---
|
||||
HTTP 200
|
||||
{"ok":true,"data":{"ok":true,"message":"Sablonok naprakészek — nincs változás"},"message":"Sablonok naprakészek — nincs változás"}
|
||||
|
||||
|
||||
@@ -0,0 +1,96 @@
|
||||
=== Scenario F — the new version never becomes healthy (alpine:3.20 exits at once) — 2026-09-13T10:12:57Z ===
|
||||
--- BEFORE ---
|
||||
{
|
||||
"state": "running",
|
||||
"updating": false,
|
||||
"update_phase": null,
|
||||
"update_phase_label": null,
|
||||
"update_error": null,
|
||||
"hold_reason": null,
|
||||
"pinned_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
|
||||
},
|
||||
"installed_images": {
|
||||
"uptime-kuma": {
|
||||
"ref": "louislam/uptime-kuma:2.4.0",
|
||||
"digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985",
|
||||
"at": "2026-09-13T10:08:19Z"
|
||||
}
|
||||
},
|
||||
"catalog_images": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.999"
|
||||
}
|
||||
}
|
||||
louislam/uptime-kuma:2.4.0 Up 3 minutes (healthy)
|
||||
--- POST /api/stacks/uptime-kuma/update at 2026-09-13T10:12:58Z ---
|
||||
HTTP 202
|
||||
{"ok":true,"data":{"accepted":true,"completed":false},"message":"FrissÃtés elindult – az állapot a kártyán követhetÅ‘"}
|
||||
|
||||
10:12:59 U safety-dump
|
||||
10:13:01 U verifying
|
||||
10:18:04 - failed
|
||||
--- AFTER 2026-09-13T10:18:04Z ---
|
||||
{
|
||||
"state": "stopped",
|
||||
"updating": false,
|
||||
"update_phase": "failed",
|
||||
"update_phase_label": "A frissÃtés nem sikerült",
|
||||
"update_error": "A(z) uptime-kuma frissÃtése 2026-09-13 12:18-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállÃtva marad, hogy az adatai ne sérüljenek. VisszaállÃtható a(z) 2026-09-13 12:09-i biztonsági mentésbÅ‘l a Mentések oldalon.",
|
||||
"hold_reason": "A(z) uptime-kuma frissÃtése 2026-09-13 12:18-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállÃtva marad, hogy az adatai ne sérüljenek. VisszaállÃtható a(z) 2026-09-13 12:09-i biztonsági mentésbÅ‘l a Mentések oldalon.",
|
||||
"pinned_images": {
|
||||
"uptime-kuma": "alpine:3.20"
|
||||
},
|
||||
"installed_images": {
|
||||
"uptime-kuma": {
|
||||
"ref": "louislam/uptime-kuma:2.4.0",
|
||||
"digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985",
|
||||
"at": "2026-09-13T10:08:19Z"
|
||||
}
|
||||
},
|
||||
"catalog_images": {
|
||||
"uptime-kuma": "alpine:3.20"
|
||||
}
|
||||
}
|
||||
--- the controller, verbatim (update + backup lines since 2026-09-13T10:12:58Z) ---
|
||||
2026/09/13 10:12:59 router.go:567: [INFO] [api] update requested for stack: uptime-kuma
|
||||
2026/09/13 10:12:59 update.go:278: [INFO] [stacks] update uptime-kuma: accepted — guarded update started
|
||||
2026/09/13 10:12:59 update.go:701: [INFO] [stacks] update uptime-kuma: phase checking
|
||||
2026/09/13 10:12:59 update.go:409: [INFO] [stacks] update uptime-kuma: precondition met — proven copy from 2026-09-13T10:09:51Z (3m0s old, limit 24h0m0s)
|
||||
2026/09/13 10:12:59 update.go:701: [INFO] [stacks] update uptime-kuma: phase safety-dump
|
||||
2026/09/13 10:12:59 update_guard.go:239: [INFO] [backup] update safety dump for uptime-kuma: the app has no database — nothing to copy (no-op)
|
||||
2026/09/13 10:12:59 update.go:423: [INFO] [stacks] update uptime-kuma: safety dump done (0 file(s)) []
|
||||
2026/09/13 10:12:59 update.go:701: [INFO] [stacks] update uptime-kuma: phase pinning
|
||||
2026/09/13 10:12:59 pin.go:362: [INFO] [stacks] update uptime-kuma: pin advanced to the catalog's current definition (uptime-kuma=alpine:3.20)
|
||||
2026/09/13 10:12:59 update.go:701: [INFO] [stacks] update uptime-kuma: phase pulling
|
||||
2026/09/13 10:13:00 update.go:701: [INFO] [stacks] update uptime-kuma: phase starting
|
||||
2026/09/13 10:13:01 update.go:701: [INFO] [stacks] update uptime-kuma: phase verifying
|
||||
2026/09/13 10:18:02 update.go:510: [ERROR] [stacks] update uptime-kuma FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state restarting) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)
|
||||
2026/09/13 10:18:02 settings.go:1618: [WARN] [settings] restore hold SET for uptime-kuma — the app stays stopped until it is cleared
|
||||
2026/09/13 10:18:02 update_guard.go:294: [WARN] [backup] uptime-kuma is HELD STOPPED after a failed update (restore point: 2026-09-13T10:09:51Z)
|
||||
--- the stacks page card (badge / hold / phase fragments) ---
|
||||
1 FrissÃtés gombot.">FrissÃtés elérhetÅ‘ â
|
||||
1 data-held="true"
|
||||
=== Scenario F — the new version never becomes healthy (alpine:3.20 exits at once) done ===
|
||||
--- the hold record in settings.json ---
|
||||
{
|
||||
"uptime-kuma": {
|
||||
"stack": "uptime-kuma",
|
||||
"at": "2026-09-13T10:18:02Z",
|
||||
"reason": "update_failed",
|
||||
"copy_date": "2026-09-13T10:09:51Z"
|
||||
}
|
||||
}
|
||||
--- units AFTER (did anything overwrite the restore point during the 5-minute verify window?) ---
|
||||
primary-unit: image: alpine:3.20 manifest "created_at": "2026-09-13T10:17:09Z"
|
||||
tier2-mirror: image: louislam/uptime-kuma:2.4.0 manifest "created_at": "2026-09-13T10:09:51Z"
|
||||
--- containers ---
|
||||
(none listed = removed by the hold's stop)
|
||||
--- app_info page fragments (ASCII only; controls included) ---
|
||||
2 Mentések
|
||||
2 Visszaáll
|
||||
1 data-held="true"
|
||||
3 friss
|
||||
2 href="/backups/apps"
|
||||
3 leáll
|
||||
--- controls: fragments that must be ABSENT on a held card ---
|
||||
lifecycle buttons on the uptime-kuma card: 0
|
||||
@@ -0,0 +1,4 @@
|
||||
2026-09-13T10:16:27Z
|
||||
primary-unit: image: louislam/uptime-kuma:2.4.0 "created_at": "2026-09-13T10:12:09Z"
|
||||
tier2-mirror: image: louislam/uptime-kuma:2.4.0 "created_at": "2026-09-13T10:09:51Z"
|
||||
alpine:3.20 Restarting (0) 41 seconds ago
|
||||
@@ -0,0 +1,24 @@
|
||||
=== Scenario H — a held app cannot be started from any path — 2026-09-13T10:18:35Z ===
|
||||
--- POST /api/stacks/uptime-kuma/start ---
|
||||
HTTP 409
|
||||
{"ok":false,"error":"A(z) uptime-kuma frissítése 2026-09-13 12:18-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a(z) 2026-09-13 12:09-i biztonsági mentésből a Mentések oldalon."}
|
||||
|
||||
--- POST /api/stacks/uptime-kuma/restart ---
|
||||
HTTP 409
|
||||
{"ok":false,"error":"A(z) uptime-kuma frissítése 2026-09-13 12:18-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a(z) 2026-09-13 12:09-i biztonsági mentésből a Mentések oldalon."}
|
||||
|
||||
--- POST /api/stacks/uptime-kuma/update ---
|
||||
HTTP 409
|
||||
{"ok":false,"error":"A(z) uptime-kuma frissítése 2026-09-13 12:18-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a(z) 2026-09-13 12:09-i biztonsági mentésből a Mentések oldalon."}
|
||||
|
||||
--- container after the three presses (must still be absent) ---
|
||||
(end)
|
||||
--- controller restart at 2026-09-13T10:18:37Z: the boot reconciler gets its pass ---
|
||||
--- the boot reconciler, the app-stop recovery and the update recovery, verbatim ---
|
||||
2026/09/13 10:19:04 bootrecon.go:236: [INFO] [bootrecon] "uptime-kuma" is a boot orphan by intent but is HELD (held after a failed update (2026-09-13T10:18:02Z) — restore it from its backup to start it) — NOT starting it; whatever is holding it owns its recovery
|
||||
2026/09/13 10:19:04 bootrecon.go:251: [INFO] [bootrecon] Boot reconciliation: nothing to start — 1 app(s) held (absent drive / quiesce / app-data operation): [uptime-kuma]
|
||||
--- container after the controller restart + boot sweep (must still be absent) ---
|
||||
(end)
|
||||
--- the API after the restart ---
|
||||
state stopped | hold_reason: A(z) uptime-kuma frissítése 2026-09-13 12:18-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a(z) 2026-09-13 12:09-i biztonsági mentésből a Mentések oldalon.
|
||||
=== Scenario H done ===
|
||||
@@ -0,0 +1,37 @@
|
||||
=== the restore walk — 2026-09-13T10:25:02Z ===
|
||||
--- POST /backup/tier2/unit-restore (form, as the page submits it) ---
|
||||
HTTP/2 302
|
||||
location: /backups/apps?flash=Teljes+vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl.
|
||||
--- restore-status at the end ---
|
||||
{
|
||||
"running": false,
|
||||
"op": "tier2-unit-restore",
|
||||
"stack": "uptime-kuma",
|
||||
"started_at": "2026-09-13T10:25:02.394608073Z",
|
||||
"last": {
|
||||
"op": "tier2-unit-restore",
|
||||
"stack": "uptime-kuma",
|
||||
"ok": true,
|
||||
"message": "A(z) uptime-kuma: 1 adatkötet visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-09-13 12:09).",
|
||||
"finished_at": "2026-09-13T10:25:11.504258982Z"
|
||||
},
|
||||
"last_recent": true
|
||||
}
|
||||
--- the app after the restore ---
|
||||
container: louislam/uptime-kuma:2.4.0 Up 10 seconds (healthy)
|
||||
state running | pinned {'uptime-kuma': 'louislam/uptime-kuma:2.4.0'} | hold_reason None | installed {'uptime-kuma': 'louislam/uptime-kuma:2.4.0'}
|
||||
--- the hold store ---
|
||||
restore_holds: None
|
||||
--- the controller, verbatim ---
|
||||
2026/09/13 10:25:02 handlers.go:1881: [WARN] [web] Tier-2 UNIT restore requested (async, OVERWRITES live data): stack=uptime-kuma from 172.18.0.3:42310
|
||||
2026/09/13 10:25:02 tier2_restore.go:197: [WARN] [backup] Tier-2 UNIT restore for uptime-kuma from the secondary mirror /mnt/felhom-drives/hdd_1/backups/secondary/uptime-kuma/recovery-unit — this OVERWRITES live app data
|
||||
2026/09/13 10:25:11 restore_unit.go:391: [INFO] [backup] Restore-from-unit completed: uptime-kuma — 1 volume(s) of 1 listed, 0 database(s) of 0 listed
|
||||
2026/09/13 10:25:11 update_guard.go:324: [INFO] [backup] uptime-kuma: restore completed — the update hold (set 2026-09-13T10:18:02Z) is CLEARED
|
||||
2026/09/13 10:25:11 handlers.go:1891: [INFO] [web] Tier-2 unit restore completed (async): stack=uptime-kuma in 9.109539298s (volumes 1/1, dbs 0/0)
|
||||
--- the card ---
|
||||
1 Frissítés
|
||||
1 Naprakész
|
||||
1 stackAction(event, 'uptime-kuma', 'restart')
|
||||
1 stackAction(event, 'uptime-kuma', 'stop')
|
||||
1 stackAction(event, 'uptime-kuma', 'update')
|
||||
=== restore walk done ===
|
||||
@@ -0,0 +1,26 @@
|
||||
=== teardown layer 1: uptime-kuma removed through the feature (R-442) 2026-09-13T10:25:53Z ===
|
||||
HTTP 200
|
||||
{"ok":true,"message":"Stack uptime-kuma stop completed"}
|
||||
|
||||
|
||||
HTTP 200
|
||||
{"ok":true,"data":{"removed":"uptime-kuma","volumes_removed":null,"hdd_paths_removed":[],"hdd_paths_preserved":[],"hdd_note":"Az alkalmazás nem tárolt saját adatot külső meghajtón, így ott nem volt mit törölni."},"message":"Stack uptime-kuma removed"}
|
||||
|
||||
|
||||
--- what the removal left (R-474) ---
|
||||
containers: 0
|
||||
volumes:
|
||||
uptime-kuma stackdir: applied-compose.yml docker-compose.yml
|
||||
uptime-kuma primary: compose manifest.json volume-dumps
|
||||
uptime-kuma secondary: recovery-unit
|
||||
glance stackdir: applied-compose.yml docker-compose.yml
|
||||
glance primary: compose manifest.json
|
||||
glance secondary:
|
||||
images: louislam/uptime-kuma:2.4.0 louislam/uptime-kuma:2.3.2 alpine:3.20 glanceapp/glance:v0.8.4
|
||||
--- residue removed BY HAND (throwaway apps only, named paths, no prune) ---
|
||||
8
|
||||
after: 0 dirs
|
||||
0
|
||||
journal-absent
|
||||
--- standing apps untouched ---
|
||||
bentopdf|Up 6 hours (healthy) bookstack-db|Up 2 hours (healthy) bookstack|Up 6 hours (healthy) calibre-web|Up 6 hours (healthy) cloudflared|Up 11 days docmost-postgres|Up 6 hours (healthy) docmost-redis|Up 6 hours (healthy) docmost|Up 6 hours (healthy) filebrowser|Up 11 days (healthy) kimai-db|Up 6 hours (healthy) kimai|Up 6 hours (healthy) opengist|Up 6 hours (healthy) paperless-postgres|Up 6 hours (healthy) paperless-redis|Up 6 hours (healthy) paperless-webserver|Up 6 hours (healthy) privatebin|Up 6 hours (healthy) romm-db|Up 6 hours (healthy) romm-redis|Up 6 hours (healthy) romm|Up 6 hours (healthy) traefik|Up 11 days
|
||||
@@ -0,0 +1,6 @@
|
||||
=== teardown layer 3: the hub 2026-09-13T10:26:04Z ===
|
||||
host: demo-felhom-8363b5
|
||||
host: demo-hp-bb76ea
|
||||
customer configs: 0
|
||||
floor: name="min_controller_version" value="0.236.0"
|
||||
(this run created no host, no customer and no appliance record; two app-deployed events from the throwaways reached the hub as ordinary events; the floor is back at the vouched golden)
|
||||
@@ -0,0 +1,3 @@
|
||||
=== hand-deploy 0.238.1 on felhom-pve-lan guest 9201 2026-09-13T10:26:31Z (floor held above the golden — R-472) ===
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.238.1
|
||||
docker ps: gitea.dooplex.hu/admin/felhom-controller:0.238.1 Up 6 seconds (healthy)
|
||||
@@ -0,0 +1,3 @@
|
||||
=== hand-deploy 0.238.1 on hp guest 9201 2026-09-13T10:26:31Z (floor held above the golden — R-472) ===
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.238.1
|
||||
docker ps: gitea.dooplex.hu/admin/felhom-controller:0.238.1 Up 7 seconds (healthy)
|
||||
@@ -0,0 +1,11 @@
|
||||
MUTATION in internal/stacks/update.go:
|
||||
- healthy, detail := m.updateHealth(ctx, name, timeout)
|
||||
+ healthy, detail := true, "RED-PROOF: no health wait"; _ = timeout
|
||||
|
||||
=== RUN TestSlice4_A_SuccessIsDeclaredOnlyAfterHealth
|
||||
update_test.go:180: the health wait was never reached; state: updating=false phase=done err=""
|
||||
--- FAIL: TestSlice4_A_SuccessIsDeclaredOnlyAfterHealth (5.01s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 5.017s
|
||||
FAIL
|
||||
exit=1
|
||||
@@ -0,0 +1,15 @@
|
||||
MUTATION in internal/stacks/update.go:
|
||||
- rp, err := g.RestorePoint(name)
|
||||
if err != nil || !rp.Restorable || !rp.Proven {
|
||||
return m.refuseUpdate(
|
||||
+ rp, err := g.RestorePoint(name)
|
||||
if err != nil || !rp.Proven {
|
||||
return m.refuseUpdate(
|
||||
|
||||
=== RUN TestSlice4_C_NoRestorableCopyRefusesBeforeAnythingMoves
|
||||
update_test.go:295: an app with no restorable copy must be REFUSED (no_backup), got <nil>
|
||||
--- FAIL: TestSlice4_C_NoRestorableCopyRefusesBeforeAnythingMoves (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
|
||||
FAIL
|
||||
exit=1
|
||||
@@ -0,0 +1,11 @@
|
||||
MUTATION in internal/stacks/update.go:
|
||||
- } else if err := g.HoldAfterFailedUpdate(name, m.now(), provenAt); err != nil {
|
||||
+ } else if err := error(nil); err != nil {
|
||||
|
||||
=== RUN TestSlice4_F_HealthFailureHoldsTheAppAndKeepsTheNewPin
|
||||
update_test.go:409: an app that did not come up must be HELD
|
||||
--- FAIL: TestSlice4_F_HealthFailureHoldsTheAppAndKeepsTheNewPin (0.01s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.012s
|
||||
FAIL
|
||||
exit=1
|
||||
@@ -0,0 +1,11 @@
|
||||
MUTATION in internal/api/router.go:
|
||||
- if action == "start" || action == "restart" || action == "update" {
|
||||
+ if action == "start" || action == "restart" {
|
||||
|
||||
=== RUN TestR439_UpdateOfAHeldAppIsRefused
|
||||
slice4_update_test.go:117: a HELD app's update must be refused with the hold's own sentence: code=409 ok=false error="A(z) app nem frissíthető, mert nincs olyan biztonsági mentése, amelyből vissza lehetne állítani. Kapcsold be a 2. mentést az alkalmazás mentési beállításainál a Mentések oldalon, és várd meg az első sikeres másolatot — utána a frissítés elindítható."
|
||||
--- FAIL: TestR439_UpdateOfAHeldAppIsRefused (0.06s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.072s
|
||||
FAIL
|
||||
exit=1
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
MUTATION in internal/backup/backup.go:
|
||||
- if m.isHeld(stack.Name) {
|
||||
m.logger.Printf("[WARN] [backup] Skipping volume dump
|
||||
+ if false && m.isHeld(stack.Name) {
|
||||
m.logger.Printf("[WARN] [backup] Skipping volume dump
|
||||
|
||||
=== RUN TestSlice4_NightlyLegsLeaveAHeldAppAlone
|
||||
slice4_update_guard_test.go:145: the nightly volume dump touched a HELD app (dumped=[held free] stopped=[])
|
||||
--- FAIL: TestSlice4_NightlyLegsLeaveAHeldAppAlone (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.006s
|
||||
FAIL
|
||||
exit=1
|
||||
@@ -0,0 +1,13 @@
|
||||
MUTATION in internal/web/intermediary.go:
|
||||
- if s.appHeld(name) {
|
||||
continue
|
||||
+ if false && s.appHeld(name) {
|
||||
continue
|
||||
|
||||
=== RUN TestSlice4_DriveReturnGateSkipsAHeldApp
|
||||
slice4_update_test.go:70: the drive-return gate tried to START a held app: runtime error: invalid memory address or nil pointer dereference
|
||||
--- FAIL: TestSlice4_DriveReturnGateSkipsAHeldApp (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.009s
|
||||
FAIL
|
||||
exit=1
|
||||
+16
@@ -0,0 +1,16 @@
|
||||
MUTATION in internal/web/templates/stacks.html:
|
||||
- {{if .Updating}}
|
||||
<span class="tag tag-progress"
|
||||
+ {{if and .Updating (not (isOperational .State))}}
|
||||
<span class="tag tag-progress"
|
||||
|
||||
=== RUN TestSlice4_Page_UpdatingCardShowsThePhaseAndNoLifecycleButton
|
||||
slice4_page_test.go:34: an updating card must show its phase label
|
||||
slice4_page_test.go:38: an updating card must offer no lifecycle button, found "stackAction(event, 'bookstack', 'update')"
|
||||
slice4_page_test.go:38: an updating card must offer no lifecycle button, found "stackAction(event, 'bookstack', 'restart')"
|
||||
slice4_page_test.go:38: an updating card must offer no lifecycle button, found "stackAction(event, 'bookstack', 'stop')"
|
||||
--- FAIL: TestSlice4_Page_UpdatingCardShowsThePhaseAndNoLifecycleButton (0.01s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.020s
|
||||
FAIL
|
||||
exit=1
|
||||
+14
@@ -0,0 +1,14 @@
|
||||
MUTATION in internal/web/templates/stacks.html:
|
||||
- {{else if .HoldReason}}
|
||||
+ {{else if and .HoldReason (not (isOperational .State))}}
|
||||
|
||||
=== RUN TestSlice4_Page_HeldCardShowsTheHoldAndTheWayBack
|
||||
slice4_page_test.go:49: a held card must show the hold sentence and link Mentések
|
||||
slice4_page_test.go:53: a held card must offer nothing that starts or updates it, found "stackAction(event, 'bookstack', 'update')"
|
||||
slice4_page_test.go:53: a held card must offer nothing that starts or updates it, found "stackAction(event, 'bookstack', 'restart')"
|
||||
slice4_page_test.go:53: a held card must offer nothing that starts or updates it, found "stackAction(event, 'bookstack', 'stop')"
|
||||
--- FAIL: TestSlice4_Page_HeldCardShowsTheHoldAndTheWayBack (0.01s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.022s
|
||||
FAIL
|
||||
exit=1
|
||||
+11
@@ -0,0 +1,11 @@
|
||||
MUTATION in internal/backup/update_guard.go:
|
||||
- return m.updatingCheck != nil && m.updatingCheck(stackName)
|
||||
+ return false
|
||||
|
||||
=== RUN TestSlice4_NightlyLegsLeaveAnAppMidUpdateAlone
|
||||
slice4_update_guard_test.go:188: a nightly leg touched an app MID-UPDATE (dumped=[updating free] stopped=[] captured=[updating free] mirrored=[updating free])
|
||||
--- FAIL: TestSlice4_NightlyLegsLeaveAnAppMidUpdateAlone (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.006s
|
||||
FAIL
|
||||
exit=1
|
||||
@@ -258,3 +258,6 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-403** | **A poorer copy deleted a richer one: an EMPTY recovery unit on the primary drive was mirrored over a COMPLETE copy on the second drive, with `--delete`.** Shipped in controller **v0.230.0**. **MEASURED before it was fixed** — on the shipped v0.229.0, on demo-hp: 120 082 104 B (4 database dumps + 3 volume tars) -> 7 036 B (none of either) in one nightly run, recorded as a success. Evidence: `audits/DRILL-r403-tier2-delete-2026-08-31/`. **Reasoning kept:** *hollowness is a MANIFEST question, never a size question* - a unit with a fat compose capture and no dumps is the dangerous shape and a 360-byte unit for a tiny app is healthy; absent or unparseable manifest counts as hollow, fail closed. *The guard fences ONE shape and not shrinking* - `07` §8 row 5's derived-copy rebuild is a DESIGN DECISION, `--delete` stays, the data legs are untouched, and only source-hollow-over-destination-complete is refused (§8.2 records the exception beside the rule so nobody 'fixes' it back). *The rehydrate happens INSIDE the restore* - the hollow manifest was written two seconds later by the 5-minute capture job, so any follow-up job races it; and *the capture is deliberately NOT guarded*, because a capture describing an empty drive as empty is correct and guarding it would make the manifest lie. *A warning that fires on everything costs the same as the comforting lie it replaces* - the first draft flagged 'package older than the run', which is true of every healthy app, and four healthy apps on the box would have been warned. | **CLOSED 2026-08-31 - controller v0.230.0, PROVEN-LIVE both ways** (the loss reproduced on v0.229.0, then the same state preserved on v0.230.0 with all 7 files sha256-identical) | full text: `git show 66156c619fd2:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-459** | **The skipped MariaDB conversion is STABLE but never self-resolving; converting costs 7 s and keeps the abort — RULED YES and shipped 2026-09-13: `MARIADB_AUTO_UPGRADE=1` on `bookstack-db`, `kimai-db`, `nextcloud-db`, `romm-db`** (catalog `eec1228`/`bd32830`/`3525e35`; no image moved, `catalog_since` untouched). Measured `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md`; proven `audits/r459-close-2026-09-13/` — harness E3/E3b `proven` with `engine_state_after` = `already upgraded to 12.3.3-MariaDB [exit=1]`, `skipped due to $MARIADB_AUTO_UPGRADE` 0 lines, C3 still `failed`; landed on demo-hp via the real 15-min cycle with both container IDs unchanged, one deliberate restart → `MariaDB upgrade not required`, `/login` 200. **Reasoning kept:** *ask the engine, not the log* — the entrypoint prints `MariaDB upgrade not required` on an unsupported downgrade too (R-464); `mariadb-upgrade --check-if-upgrade-is-needed` exit 0 = needed, 1 = not. *Not established, unchanged:* whether any MariaDB feature misbehaves on an UNCONVERTED datadir. *The precaution that keeps the setting inert until Slice 4:* the engine-major rule + gate, removal tracked as R-469. | **CLOSED 2026-09-13 — shipped in the catalog, PROVEN by harness and live** | full text: `git show ae59c31:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-467** | **Controller v0.236.0 owed a golden — PAID 2026-09-13:** golden `0.236.0` baked (`GOLDEN_SHA256=58a3cc24…958bf`, 654 115 664 B), round-tripped from the DOWNLOADED bytes, `./etc/felhom-controller-image` says `felhom-controller:0.236.0`, hub dropdown agreed, three-field vouch re-read (`0.236.0` / `agent 0.130.0` / `min_agent 0.129.0`, not the R-216 shape), floor raised 0.232.0 → 0.236.0. Evidence `documentation/tests/golden-0.236.0-2026-09-13/`. **Reasoning kept:** *this bake carried FOUR unbaked releases and is the LAST per-release bake* — goldens are weekly and before any install from today (R-468); *the MinAgent line was missing from four headers* (R-470). | **CLOSED 2026-09-13 — baked, vouched, floor raised** | full text: `git show ae59c31:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-448** | **UPDATE ARC SLICE 4 — a guarded update: verified-backup precondition, abort path, truth at the moment of action.** Shipped controller **v0.237.0** (job) + **v0.238.0** (page) + **v0.238.1** (nightly legs skip an app mid-update, found live). Proven live on demo-hp 2026-09-13, scenarios A/B/E/F/H and the restore walk. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *the precondition is the existing verified backup, not a new copy* (ruling 2026-09-02) — `backup.Tier2UnitRestorePoint`, extracted from the backups page, not copied; *age a copy by its last SUCCESSFUL Tier-2 copy, never the manifest `created_at`* (measured: the manifest moves only on definition changes); *no automatic rollback — measured per-app, the route back is the restore*; *anything that writes a restore point skips an app that is held OR updating*. Open consequences: R-472, R-475, R-476. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-443** | **The Update button reported success over an app it had just broken.** Closed by slice 4 (v0.237.0): 202 `accepted, not completed`; the outcome exists only as `update_phase` after health. Pinned by `TestR443_UpdateIsNeverReportedCompleteSynchronously`. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *a compose exit code is never a success signal* (spike §4: HTTP 200 over a crash loop). | **CLOSED 2026-09-13** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-439** | **The restore hold was not honoured by the update path.** Closed by slice 4 (v0.237.0): `update` joined the router's hold check; live Scenario H refused start/restart/update and the boot sweep. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *a hold that only one path honours is not a hold* — the audit found the drive-return gate (restart + boot recreate) and the nightly volume dump ignoring any hold; all three fixed and red-proofed. | **CLOSED 2026-09-13** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
@@ -675,13 +675,10 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** |
|
||||
| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. **STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner.** `docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic` reads: *"Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH."* So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses `rclone:`, measured with a control). **AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all.** That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. **Corroboration, unlooked for:** that page's table of port-23 commands matches, item for item, the `help` output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. | **OPEN — ask the vendor before building anything** |
|
||||
| **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** |
|
||||
| **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). **RE-READ AND CONFIRMED 2026-09-01; the severity argument SURVIVES but its stated reason was imprecise and is corrected here.** The task's reason was *"the UI only renders Frissites when the app is operational, and a held app is stopped"*. Half right: `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded` as operational too**, and this was OBSERVED live — the green `Frissites` button rendered over a crash-looping app during the spike's Phase 3b. **So the button is hidden specifically because a held app is `StateStopped`, not because broken apps hide it.** LOW stands; the reason must be stated precisely or the next reader will widen it. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-443** | **[P2-MEDIUM] The Update button reports SUCCESS over an app it has just broken, and the truth arrives 5m16s later by a different road.** MEASURED on demo-hp 2026-09-01: `POST /api/stacks/bentopdf/update` against an image that pulls cleanly and then fails to run returned **HTTP 200 `{"ok":true,"message":"Stack bentopdf update completed"}`** and logged `Stack bentopdf updated successfully (took 3.5s)`, while the container went to `status=restarting RestartCount=9`. **The controller's own post-start line told the truth (`manager.go:1403 ... alpine:3.20 restarting`) — but it runs AFTER the API has already answered.** This is this repo's own `up -d` exits 0 on a crash-loop invariant surfacing at the customer's most consequential button. **What the customer's page then said:** badge **`Ujraindites...`**, `Restarting (0) 15 seconds ago`, and the full green button row — because `isOperationalState` counts `StateRestarting` as operational (see R-439). *"Restarting"* reads as transient, not as failure, and nothing says the update caused it. **THE HONEST OTHER HALF, and it must travel with this row: the customer IS told.** `app_start_failed` fired at 18:05:59Z with severity `warning` (inside the hub's exact vocabulary, so it really delivers) — 5m16s after the update, from `crashLoopAfter = 5 * time.Minute`, a threshold whose own comment argues it well. **So this is NOT the silent-dead-app class; it is a TRUTHFULNESS-AT-THE-MOMENT-OF-ACTION problem.** Also recorded: on a pull FAILURE the product behaves correctly — HTTP 500, and `compose up -d` resolves images before touching a container, so the running app survives (measured twice). Owner: **CC to propose, VIKTOR to rule on whether Update should wait and verify.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules** |
|
||||
| **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** |
|
||||
| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-448** | **[P2-MEDIUM] UPDATE ARC SLICE 4 — a guarded update: a verified backup as a precondition, an abort path, and the truth at the moment of action.** Three parts, each already evidenced. (a) **The precondition is a VERIFIED RECENT BACKUP, not a new copy** (operator ruling 2026-09-02); the guest-snapshot alternative must be SPIKED before anything is designed around it. Today's safety machinery is DATABASE-ONLY (`Manager.writeSafetyDump`, `internal/backup/offbox_reconstitute.go:207`) and the file half was never priced (spike §6). (b) **The abort path, never a "rollback"** — spike §7 proved the word is wrong: once a migration has run, the old image refuses to start on the migrated data. The two available shapes are ABORT (before anything migrated) and RESTORE FROM A COPY (after). (c) **Truth at the moment of action** — this subsumes **R-443**: the Update button returns HTTP 200 over an app it has just broken and the alarm arrives 5m16s later. `architecture/09-update-architecture.md` §3, §4 | **READY — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules on (c)** |
|
||||
| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-452** | **[P3-LOW] Nothing enforces `catalog_since`, so the one number the update badge shows can silently under-report.** `app-catalog-felhom.eu` `CLAUDE.md` now states the rule — any commit that changes an `image:` line must set that app's `catalog_since` to the same day — and all 53 apps were backfilled from git history on 2026-09-02 (`69761cf`). **A rule with no instrument is a wish; that is this project's most-repeated finding and this row exists so it is not repeated silently.** A stale `catalog_since` makes „Frissítés elérhető — N napja" under-report N, which is the single number the badge exists to give. **WHY IT WAS NOT BUILT IN THE SAME SESSION, stated rather than implied:** the gate would have to diff an `image:` line against the PARENT commit, and `catalog_gates.py` runs under a runner that fetches at `--depth 1` — there is no parent to diff against. The gate therefore needs a deeper fetch, which is a change to the CI shape and not to a script. **This is the R-421 class in advance: an enumerated gap becomes a row in the same session it is enumerated.** `architecture/09-update-architecture.md` §8.2 | **READY — rank P3-LOW; owner: CC** |
|
||||
@@ -698,12 +695,18 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-465** | **[P3-LOW] `cfg.Paths.HDDPath` — the global that R-442 proved is set on NO box (demo-hp AND demo-felhom: 0 `hdd_path`, 0 `FELHOM_PATHS_*`) — still has SIX readers, each reading an always-empty value:** `report/builder.go:69`, `monitor/healthcheck.go:35`, `api/router.go` (system-info), `web/server.go:740`, `cmd/controller/main.go` (auto-discovery seed + metrics HDD path). Removal was silently inert for months on the very same read. Whether any of these is inert the same way — a report field that is always empty, a health check that never fires, a metric never collected — is a one-hour audit: for each reader, name the POSITIVE observable that must appear when it works and check it on the box. R-442 §5 said "do not delete it here"; this row is the audit it deferred. Owner: **CC.** `felhom-controller/REPORT.md` (v0.236.0, Observations 1) | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-466** | **[P3-LOW] Removing an app with „Mentési adatok törlése" ticked leaves its recovery unit's `compose/` + `manifest.json` on the drive.** The router passes only `backup.AppDBDumpPath(nsRoot, name)` to `RemoveStack`, so `backups/primary/<app>/db-dumps/` goes and the unit root keeps `compose/` (the app's `app.yaml` with the portable secret class at 0600) and `manifest.json`. MEASURED 2026-09-13 on demo-hp after the R-442 Scenario A removal — `audits/R442-2026-09-13/teardown-and-log.txt`, residue check: `backups/primary/nextcloud` still listed with the app gone; removed by hand at teardown. A customer who asked for the backups to go is left with the app's definition and a manifest. **Decide:** the button means the WHOLE unit (pass the unit root, under the same `backups/`-prefix guard) or stays db-dumps-only (then the modal must say so). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
| **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver.yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** |
|
||||
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. | **BLOCKED — on R-448; rank P3-LOW; owner: CC** |
|
||||
| **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** |
|
||||
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). | **READY — unblocked by R-448; rank P3-LOW; owner: CC (removal is a deliberate act)** |
|
||||
| **R-470** | **[P3-LOW] Four consecutive controller CHANGELOG headers (v0.233.0 … v0.236.0) carry NO `MinAgent:` line, and `RUNBOOK-manual-build.md` §4.1 step 5 tells the vouch to read it from the header.** Measured 2026-09-13 while baking golden 0.236.0: `sed -n 1,349p CHANGELOG.md | grep -c MinAgent` → **0**; the newest statement is v0.232.0's `**MinAgent: 0.129.0** (unchanged)`. The vouch used 0.129.0 on the ground that no later entry declares a change and `internal/agentapi/features.go` gates per feature, not per release — but a reader following the runbook literally finds nothing to read, which is the R-233 shape (a document pointing at a string that is not there). **Fix:** either every release header carries the line (the v0.232.0 convention), or the runbook says where MinAgent actually lives when a header omits it. One of the two, not both. | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
| **R-471** | **[P3-LOW] `observations_gate.py` reads the FIRST `Observations` section of `REPORT.md` and nothing after it, so an appended second section with an unmarked item passes — and the decoy that proves it has been reading "LIVE HOLE" at HEAD, unnoticed, because the decoy suite is run by hand.** MEASURED 2026-09-13 on a clean worktree of `4b2e560`: `python3 scripts/test_gate_decoys.py` → `FAIL: observations/R-419: decoy PASSED - LIVE HOLE (rc=0)`; the gate's own output shows it scanned `## 11. Observations — noticed, documented, NOT acted on` (3 items, all marked) and never reached the appended `## Observations` block the decoy planted. Whether the cause is "first heading wins" or a heading-shape filter is NOT established — only that the planted unmarked item was not seen. **Consequence:** a report with two observation sections gets the second one unchecked. **Two fixes, both owed:** scan every section whose heading contains `Observations`, and put `test_gate_decoys.py` where something runs it (it is the instrument for R-421 and nothing in `repo_gates.py` invokes it). | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
| **R-472** | **[P2-MEDIUM] THE GOLDEN CADENCE RULING AND THE HUB'S FLOOR RULE CONTRADICT EACH OTHER: a floor raise delivers NOTHING until a golden carries the release.** The 2026-09-13 ruling (R-468) rests on *"every release still raises the floor, so both demo boxes keep getting each release in ~20 s; only the golden moves to a cadence"*. MEASURED THE SAME DAY, releasing controller v0.237.0: `POST /configuration/global-floor` → `0.237.0` (303, re-read), and the hub logged for BOTH boxes `managed floor HELD for demo-hp: held: floor 0.237.0 is ABOVE the vouched golden 0.236.0, so its agent requirement is unknown — vouch a golden carrying the floor's controller (publish-train rule 1) (controller floor withheld)`; the box logged `SetFloor: floor "0.236.0" → ""`; neither box moved in 8 minutes. Publish-train rule 1 (hub `ResolveManagedFloor`) was built so a controller is never pushed past the agent it needs, and it reads that requirement from the vouched golden. **So under a weekly golden, releases between bakes do not reach the fleet by floor at all** — they need a hand deploy. v0.237.0 and v0.238.0 were hand-deployed to both demo guests by the skill's documented route and the floor was put back to 0.236.0 (a floor held on every box is a standing dashboard flag). Evidence `audits/slice4-2026-09-13/live/00-floor-0.237.0.txt`. **This is the operator's call, not CC's:** (a) bake per release again (the treadmill R-468 ended), or (b) let the floor carry a release above the golden when its CHANGELOG header states an unchanged `MinAgent` — which needs R-470's header line to be reliable first. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-473** | **[P2-MEDIUM] The `glance` catalog template crash-loops on EVERY fresh install.** MEASURED 2026-09-13 on demo-hp (deployed as a throwaway for the slice-4 live test): `restarts=13 status=restarting`, the log repeating `parsing config: reading /app/config/glance.yml: open /app/config/glance.yml: no such file or directory`, the `glance_config` volume empty. The image does not create a default config and the template seeds none. The install page reports the deploy as started and the card then shows a restarting app. Not a version problem: the template's current tag v0.8.5 has the same config requirement (the catalog walk that pinned it never did a fresh install). **Fix shape:** seed a minimal `glance.yml` the way `gokapi`'s catalog entry seeds `config.json` (memory `gokapi-headless-setup`), then prove a fresh install lands healthy. Evidence `audits/slice4-2026-09-13/live/02-glance-abandoned.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-474** | **[P3-LOW] "Remove app" with *also delete backups* deletes only the `db-dumps` directory, and reports `volumes_removed: null` over a volume it DID remove.** MEASURED 2026-09-13 removing the glance throwaway on controller v0.237.0 (`remove_hdd_data:true, remove_backups:true`): response `{"removed":"glance","volumes_removed":null,"hdd_paths_removed":[],…}`; afterwards `docker volume ls` shows no glance volume, while `backups/primary/glance/{compose,manifest.json}` and the stack dir's `applied-compose.yml` remain. The router passes ONLY `backup.AppDBDumpPath(nsRoot, name)` (router.go, beside a comment that disk-tier backup "moved to the host agent" — stale since Tier 2 returned to the controller). So the recovery unit and any Tier-2 copy survive a removal that promised to delete backups, and the answer names no volume. Same class as R-442: the removal's answer does not describe what happened. (Removal also refuses a crash-looping app as "still running — stop it first", which is correct and was observed.) **REPRODUCED a second time the same day** removing the uptime-kuma throwaway: `volumes_removed: null`, and `backups/primary/uptime-kuma/{compose,manifest.json,volume-dumps}` plus `backups/secondary/uptime-kuma/recovery-unit` survived `remove_backups:true` (residue cleared by hand; `audits/slice4-2026-09-13/` `live/15-teardown.txt`). | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-475** | **[P2-MEDIUM] The update precondition is Tier-2-ONLY, so an app with no Tier-2 copy cannot be updated at all — even when its primary unit or off-site copy could restore it.** Slice 4 (v0.237.0) gates the Update button on `backup.Tier2UnitRestorePoint`, exactly as specified. MEASURED on demo-hp 2026-09-13: `gokapi` and `nextcloud` have NO `cross_drive` record, so both are refused with the no-backup sentence. The classes this reaches: a box with no second target (`no_target`), an app whose customer switched Tier 2 off, and a freshly installed app before its first Tier-2 run. **It is a finding about the specification, not a defect in the build:** the primary recovery unit (same drive) and the off-site unit (Tier 3) are both restorable routes the predicate does not consider — the driveless-app class R-356 names. Options: accept Tier 2 as the only update route (and say so in the refusal), or widen the predicate to the primary unit with an explicit "same drive" caveat. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules on the route set, CC implements** |
|
||||
| **R-476** | **[P3-LOW] The Mentések page names a Tier-2 copy's date from the unit MANIFEST, which moves only when the app's DEFINITION changes — so it can undersell a fresh copy by a day or more.** MEASURED on demo-hp 2026-09-13: bookstack's mirror held `bookstack-mariadb.sql` written 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z; the Tier-2 run succeeded at 01:30Z. A capture rewrites the manifest only when checksums, dump NAMES or controller version change, and nightly dumps keep their names. R-403 chose the package date so a PRESERVED package is never shown as fresh — correct — but for a normal run it is the older, flattering-in-reverse date. The update (slice 4) deliberately ages the copy by the last successful copy instead (`Tier2RestorePoint.ProvenCopyTime`), so the two can name different dates. Fix shape: record the unit's DATA time (newest dump mtime) in the manifest, or name `LastSuccess` when the leg was not preserved. | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
Clearing a row means the check was DONE and its result recorded in that R-row —
|
||||
|
||||
@@ -52,7 +52,7 @@
|
||||
|
||||
| ID | Item | Size | Status | Notes |
|
||||
|----|------|------|--------|-------|
|
||||
| **UPDATE-ARC** | **The app-update arc — seven slices, from "nobody knows what any box runs" to "an update is a decision the box can take safely."** Opened after `audits/SPIKE-app-update-2026-09-01.md` measured what an update actually does. | L | **slices 1, 1b, 2, 3 & 5 SHIPPED (controller v0.233.0 / v0.234.0 / v0.235.0; harness `app-catalog/scripts/upgrade-test.py`); slices 4, 6, 7 open** | **The reasoning has a home and it is the point of the exercise: `architecture/09-update-architecture.md`** — a LIVING document, updated by every slice in the same session, created because its ABSENCE was a finding (R-438: the mechanism was chosen deliberately and written down nowhere). **Capability-map rows this flips:** the App-lifecycle row (`00-capability-map.md`, "start/stop/restart/update/logs/remove/redeploy") and its 2026-09-01 sibling ("what restart and update do to a deployed app whose compose file the catalog already moved") — both recorded that the ACTIONS work and said nothing about VERSIONS. Slices 1 and 2 add the version half. **The findings live in the register, per the ONE REGISTER ruling:** R-438, R-440, R-441, R-443 (existing) and R-446..R-452 (this session) — one row per remaining slice, each with a rank and an owner, so the arc is visible in the register and not only in a task file. **Slice 3 SHIPPED 2026-09-06 (controller v0.235.0) on the operator's ruling — Option 1, *freeze the version, keep the fixes flowing*.** It was BLOCKED until then, deliberately: changing what the syncer does to a deployed app reverses a decision, not a bug. R-447, R-441 and R-438 are CLOSED by it; **Slice 5 SHIPPED 2026-09-06 as a spike** — `scripts/upgrade-test.py`, 7 edges over 3 apps, proven by a RED negative control, closing R-449 and opening R-459/R-460/R-462. **Its headline changes an assumption this arc was carrying: whether an update can be undone is a property of the individual APP, not of updates** — docmost refuses, privatebin does not — so any design assuming one answer for all 53 is designing against a checked-and-false fact. **R-459 was measured 2026-09-06 and narrowed to a decision, not a defect:** the skipped MariaDB conversion is stable but never self-resolving, and correcting it costs 7 s and does **not** cost the ability to abort — so the trade that row was expected to produce does not exist. Opened R-463 (the PostgreSQL analogue, which fails in the OPPOSITE direction — it refuses to start, and 8 templates sit on `postgres:16-alpine`) and R-464 (an entrypoint line that says `upgrade not required` on an unsupported downgrade). **R-448 (slice 4, the guarded update) is now the head of the arc** and is where the backup precondition goes — slice 3 did NOT make the Update button safer and must not be read as having done so. |
|
||||
| **UPDATE-ARC** | **The app-update arc — seven slices, from "nobody knows what any box runs" to "an update is a decision the box can take safely."** Opened after `audits/SPIKE-app-update-2026-09-01.md` measured what an update actually does. | L | **slices 1, 1b, 2, 3 & 5 SHIPPED (controller v0.233.0 / v0.234.0 / v0.235.0; harness `app-catalog/scripts/upgrade-test.py`); slices 4, 6, 7 open** **2026-09-13 — SLICE 4 COLLAPSED: SHIPPED** (controller v0.237.0/v0.238.0/v0.238.1, R-448 CLOSED, proven live — `audits/slice4-2026-09-13/`). Remaining: slice 6 (R-450), slice 7 (R-451); R-469 unblocked. | **The reasoning has a home and it is the point of the exercise: `architecture/09-update-architecture.md`** — a LIVING document, updated by every slice in the same session, created because its ABSENCE was a finding (R-438: the mechanism was chosen deliberately and written down nowhere). **Capability-map rows this flips:** the App-lifecycle row (`00-capability-map.md`, "start/stop/restart/update/logs/remove/redeploy") and its 2026-09-01 sibling ("what restart and update do to a deployed app whose compose file the catalog already moved") — both recorded that the ACTIONS work and said nothing about VERSIONS. Slices 1 and 2 add the version half. **The findings live in the register, per the ONE REGISTER ruling:** R-438, R-440, R-441, R-443 (existing) and R-446..R-452 (this session) — one row per remaining slice, each with a rank and an owner, so the arc is visible in the register and not only in a task file. **Slice 3 SHIPPED 2026-09-06 (controller v0.235.0) on the operator's ruling — Option 1, *freeze the version, keep the fixes flowing*.** It was BLOCKED until then, deliberately: changing what the syncer does to a deployed app reverses a decision, not a bug. R-447, R-441 and R-438 are CLOSED by it; **Slice 5 SHIPPED 2026-09-06 as a spike** — `scripts/upgrade-test.py`, 7 edges over 3 apps, proven by a RED negative control, closing R-449 and opening R-459/R-460/R-462. **Its headline changes an assumption this arc was carrying: whether an update can be undone is a property of the individual APP, not of updates** — docmost refuses, privatebin does not — so any design assuming one answer for all 53 is designing against a checked-and-false fact. **R-459 was measured 2026-09-06 and narrowed to a decision, not a defect:** the skipped MariaDB conversion is stable but never self-resolving, and correcting it costs 7 s and does **not** cost the ability to abort — so the trade that row was expected to produce does not exist. Opened R-463 (the PostgreSQL analogue, which fails in the OPPOSITE direction — it refuses to start, and 8 templates sit on `postgres:16-alpine`) and R-464 (an entrypoint line that says `upgrade not required` on an unsupported downgrade). **R-448 (slice 4, the guarded update) is now the head of the arc** and is where the backup precondition goes — slice 3 did NOT make the Update button safer and must not be read as having done so. |
|
||||
| R-6 | **Spike: LAN service discovery from the guest** — SSDP multicast (UDP 1900, DLNA), WSD (Windows discovery), mDNS; host-network vs macvlan; is the customer LXC LAN-bridged in appliance deployments? | M | **spiked (2026-07-18)** | **VERDICT: appliance guest IS LAN-bridged (own DHCP lease on the household /24); multicast discovery works ONLY in the guest netns — guest-direct or Docker `--network host` (SSDP/mDNS/WSD all PASS both ways); the default docker bridge is categorically DEAF to LAN multicast (WSD/mDNS RX FAIL, unicast-publish PASS). Real samba+wsdd on host-net → Windows 11 ProbeMatch + FELHOM-SPIKE renders in Explorer + 445 + authenticated SMB round-trip all PASS; real SSDP `MediaServer:1` advert reaches both LAN clients. → R-7 SMB stack MUST be host-network LAN-bound; R-8 Jellyfin-DLNA plausible if host-network. Caveat: `vmbr0 multicast_snooping=1` worked only because the household router is a live querier — customer LANs w/ snooping+no-querier, and Peti's BYO bridge, are UNTESTED gaps.** **S4b (human leg, the sharpest finding): wsdd makes the box VISIBLE but the Explorer double-click FAILS `0x80070035` — WSD gives no name resolution; the flat `\\FELHOM-SPIKE` resolved by no path. Adding `nmbd` (NetBIOS) fixed it live (flat name resolves + mounts). → R-7 needs smbd+wsdd+nmbd (+avahi/.local for modern clients), not wsdd alone.** Doc: `audits/SPIKE-lan-discovery-2026-07-18.md`. |
|
||||
| R-7b | **Share backup EXECUTION** — put share data into the live tier-2 + offsite runs (the design fork reported by R-7 slice 1) | M | **SHIPPED (controller v0.145.0, 2026-07-18)** | **Viktor's ruling: Model B′ — a SIBLING shares source.** New, additive job/leg code reusing the proven primitives (tier-2 mirror seam, restic wrappers, soft-quota/enlargement gate, status recorders) while leaving **every per-app engine path byte-identical** — NOT a synthetic recovery unit (breaks on multi-drive shares, wraps 1 KB of JSON in dump machinery) and NOT engine-loop surgery. The B′ invariant is enforced by test in both tiers, red-proofed. Tier 2 → `RunSharesTier2` (legs grouped by SOURCE drive → `backups/secondary/_shares/<driveKey>/<share>`, payload at `_payload/`, layout marker LAST). Tier 3 → `runOffboxSharesLeg`: ONE extra `restic backup --tag felhom-offbox --tag _shares` placed after the app loop and BEFORE retention, so `forget --group-by host,tags` covers the new group with no flag change; a quota-blocked push degrades to the **manifest only, never to nothing**. Restore → „Megosztások" on `/backups/restore`: scratch, then a missing-only merge whose every destination is PREFIX-ASSERTED against live storage roots, definitions merged existing-wins, then `ReconcileSamba`, then the credential. The **payload** (`_shares-manifest.json` + a best-effort secret-bearing `passdb.tar`) is what makes DR return files + configuration + password rather than loose bytes. **Fold-in: samba joins the liveness set** — `EffectiveProtected` adds the CONTAINER `felhom-samba` exactly while sharing is on. **FULLY PROVEN-LIVE on demo (2026-07-18), all four legs.** (1) tier-2: real `/api/backup/tier2` trigger → `_shares` tree + marker + payload on the cross-drive target, mirrored file md5-identical, payload 0600 preserved. (2) offsite: Viktor's manual run 12:18:16Z → snapshot **`e0b9d723`** (tags `felhom-offbox,_shares`) with the payload dir + both share folders; a second run via the „Távoli mentés" button → **`4e2b15ec`**, containing `_shares-manifest.json` (418 B) AND `passdb.tar` (855 040 B), both 0600, share files with uid 1000 preserved. (3) restore round-trip: probe file + the `dokumentumok` DEFINITION deleted via the real endpoints, then „Megosztások" restore + place → `1 file(s), 1 definition(s) re-added, 1 kept, 0 refused, credential=true`; probe back md5-identical, the two pre-existing files NOT overwritten (missing-only proven on live data), definition back with its ORIGINAL flags and created_at, `smb.conf` re-rendered, `filmek` untouched. (4) liveness: samba stopped → `health_critical` pushed and hub-accepted (200) → self-healed. Remaining human leg: SMB positive auth with the real household password (never persisted by design). **Correction:** an earlier revision of this row and of the ship REPORT wrongly claimed the demo box had no offsite target — the verification read a guessed settings key (`offbox_target`) instead of the real one (`offbox`); root cause dissected in REPORT §7b. Findings: the reserved-name assumption was FALSE (`nbNameRe` accepted „_shares" as a share name — now refused); the alert/e-mail pipeline needed NO change and adds no new event type. Docs: `controller/sharing.md`; ship report `felhom-controller/REPORT.md`. |
|
||||
| R-8 | DLNA (**gate input now exists — R-6 spiked 2026-07-18: SSDP reaches LAN clients from host-net**): validate Jellyfin's built-in DLNA server first; only add minidlna to the catalog if Jellyfin-DLNA fails | S | idea (unblocked) | Don't add catalog weight before proving the cheap path. **R-6 confirmed the cheap path is physically viable — Jellyfin DLNA must run host-network (same multicast constraint as R-7)** |
|
||||
|
||||
@@ -210,9 +210,10 @@ The full 0.188.0 run, with the observables: `documentation/audits/tester-gate-go
|
||||
**What was measured before the ruling:** 25 goldens in 26 days in August, almost one per release,
|
||||
because `golden_currency_gate.py` trips on every release by design and the only honest ways past it
|
||||
were a bake or a declared `--no-verify` (thirteen of those by 2026-09-01 — R-404/R-417). **The
|
||||
ruling:** bake on a cadence. Every release still raises the FLOOR (§4.1 step 5), so both demo boxes
|
||||
keep getting each release in ~20 s; only the golden — which protects a **fresh install** and nothing
|
||||
else — moves to a cadence.
|
||||
ruling:** bake on a cadence. The ruling assumed every release would still raise the FLOOR (§4.1 step
|
||||
5) and reach both demo boxes in ~20 s, with only the golden moving to a cadence. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.** A
|
||||
release between bakes is hand-deployed per the `felhom-build-deploy` skill, and the floor is left at
|
||||
the vouched golden.
|
||||
|
||||
**The cadence, and it is a step in a routine, not a memory:**
|
||||
|
||||
|
||||
@@ -1,3 +1,11 @@
|
||||
## golden_currency_gate.py docstring corrected: the floor does NOT carry a release between bakes (2026-09-13, R-472) — NOT A RELEASE
|
||||
|
||||
Docstring only; the gate's behaviour is unchanged. This morning's entry stated that under the weekly
|
||||
golden cadence "every release still raises the FLOOR, so the fleet keeps getting each release in
|
||||
~20 s". Measured the same day releasing controller v0.237.0: the hub HOLDS any floor above the vouched
|
||||
golden (publish-train rule 1) and neither demo box moved. The same false sentence was corrected in
|
||||
`RUNBOOK-manual-build.md` §4.2, `STATUS.md`, `CONTEXT.md` and R-468's row. R-472 carries the decision.
|
||||
|
||||
## the golden waiver — goldens on a cadence, not per release (2026-09-13, R-468 / R-242) — NOT A RELEASE
|
||||
|
||||
**No product code, no version bump, no image.** A scripts change is not a release.
|
||||
|
||||
@@ -81,9 +81,10 @@ What happened between 2026-08-07 and 2026-09-01: **25 goldens in 26 days**, almo
|
||||
because this gate trips on every release by design (see WHY VERSION AND NOT BEHAVIOUR) and the only
|
||||
honest ways past it were a bake or a `--no-verify`. Thirteen bypasses were counted by 2026-09-01
|
||||
(R-404/R-417). The operator ruled on 2026-09-13: **bake on a cadence — weekly, and always before any
|
||||
drill or fresh install — not per release.** Every release still raises the FLOOR, so the fleet keeps
|
||||
getting each release in ~20 s; only the golden, which protects a fresh install and nothing else,
|
||||
moves to a cadence.
|
||||
drill or fresh install — not per release.** The ruling assumed every release would still raise the
|
||||
FLOOR and reach the fleet in ~20 s. CORRECTED THE SAME DAY (R-472): the hub HOLDS a floor above the
|
||||
vouched golden, so between bakes a release reaches the demo boxes only by hand-deploy. This gate's
|
||||
behaviour is unaffected — it reads the bake record, not the fleet.
|
||||
|
||||
The mechanism is a small tracked file, `documentation/tests/golden-waiver.yml`:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user