v0.234.0: a known limitation written on 2026-09-02 was a defect by the next morning
gates / gates (push) Successful in 18s

The operator looked at demo-felhom and found OpenGist - up 15 hours, running
exactly the catalog pin, showing no badge at all. 09-update-architecture.md had
recorded that as an accepted limitation the day before: 'the fleet view fills in
gradually'. On a quiet box gradually means never, and a feature that fills itself
in on an event nobody triggers is, on the quiet installations, not shipped. That
limitation row is now struck with the reason kept.

The living document gains slice 1b, the two admission rules of the backfill (it
never overwrites, and it refuses to seed a partial observation because the badge
reads a service-count mismatch as BEHIND), and the note that the same field having
two writers with two different admission rules is deliberate.

Live evidence added: all nine apps already had records by the time 0.234.0 was
ready, so the natural fleet state could no longer exercise the new code - said
plainly rather than papered over. The pre-0.233.0 shape was recreated on demo-hp
by stripping two records; the backfill re-seeded exactly those two with digests
matching independently-read ground truth and left the other seven alone.

The refusal half was deliberately NOT staged live: it needs a degraded app, and
manufacturing one risks the false-customer-email class that already cost 61 mails
(R-330). Unit-tested with a red-proof, and recorded as unproven-live.

R-457: a test that hardcodes a date and asserts an age derived from it is green
only on the day it is written. Mine was, and it went red overnight. Six other
files carry both a date literal and time.Now() - named as candidates, not accused.
This commit is contained in:
2026-09-03 12:01:44 +02:00
parent 7941b0c159
commit bc47dd4ef9
6 changed files with 112 additions and 12 deletions
+8 -2
View File
@@ -1,4 +1,4 @@
# REPORT — update arc slices 1 & 2: documents, register, roadmap, capability map (2026-09-02)
# REPORT — update arc slices 1, 1b & 2: documents, register, roadmap, capability map (2026-09-03)
*Overwritten each session. Nothing durable lives only here.*
@@ -29,7 +29,7 @@
## Register
**194 rows before, 205 after.** Nothing closed, and that is stated rather than implied: R-438 and
**194 rows before, 206 after** (205 on 2026-09-02, plus R-457 on 2026-09-03). Nothing closed, and that is stated rather than implied: R-438 and
R-440 are **amended and stay OPEN** — the mechanism is now documented, not changed — so nothing moved
to `CLOSED-ITEMS.md` and that file is untouched.
@@ -46,6 +46,7 @@ to `CLOSED-ITEMS.md` and that file is untouched.
| R-454 | five `gofmt`-unclean `internal/web` test files, and no gate notices | READY, P3-LOW |
| R-455 | **DooPlex has no Docker Hub login — the ceiling now blocks BUILDS, not just gates.** The mirror workaround is **MEASURED byte-identical** to Hub, so it is a sound fallback, not a hope | **WAITING-ON-OPERATOR**, P2-MEDIUM |
| R-456 | a partly-dead stack is not a boot orphan, and that rule is written down nowhere | READY, P3-LOW |
| R-457 | **a test that hardcodes a date and asserts an age is green only on the day it is written** — one proven, six candidate files named | READY, P3-LOW |
## Live validation
@@ -66,4 +67,9 @@ double quotes from a single-quoted value.** The operator caught it in one line.
## Sibling repos
- `felhom-controller` **v0.233.0** — `8025304acc0a`, deployed to demo-hp and verified healthy.
- `felhom-controller` **v0.234.0** — `38d28b5b624c`, the startup backfill. Deployed to **both** demo
boxes and verified: 2 apps seeded on demo-hp with digests matching ground truth, 7 left untouched;
all 9 then badged. **Its cause is worth keeping: `09-update-architecture.md` §8.3 was written as an
accepted limitation on 2026-09-02 and was a defect by the next morning** — the operator found an app
that simply ran and therefore showed nothing.
- `app-catalog-felhom.eu` — `69761cf91bfc` (backfill) + `8220f8d` (REPORT).
+23 -5
View File
@@ -1,9 +1,12 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-02 — the box now writes down which version of each app it is running, and shows one
**Updated 2026-09-03 — you spotted that OpenGist had no label. You were right, and it was a real gap:
the label only appeared on apps something had restarted. Fixed and live (0.234.0). Every app on both
machines now carries one. NOTHING IS WAITING ON YOU.**
**Earlier 2026-09-02 — the box now writes down which version of each app it is running, and shows one
small label saying whether it is up to date: „Naprakész" or „Frissítés elérhető — 52 napja". No version
numbers, and nothing about updating changed. Both halves are proven on the real machine, labels read off
the real pages. NOTHING NEW IS WAITING ON YOU.**
numbers, and nothing about updating changed.**
**Earlier 2026-09-01 (third pass) — the backup work is FINISHED for beta, and I have written down
where it stops. One alarm that was telling you something untrue is fixed and live (hub 0.111.1).
@@ -116,8 +119,23 @@ nothing.*
Full measurement, with the controls and the quoted output:
`felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md`.
9. **Nothing here needs you. I got in on my own — after getting it wrong first, which is worth one
line.**
9. **Nothing here needs you. Two things I got wrong, both now fixed.**
**You found the first one.** OpenGist showed no label. That was not a misunderstanding — the label
only appeared on an app that something had restarted, so an app that simply runs showed **nothing,
possibly for months**. On a quiet machine that is every app, which is exactly the machine we most
want to be able to look at. **Now the box reads what every app is on when it starts up, and writes
it down.** It only looks — it starts nothing and changes no app. Live on both machines: all nine
apps on the HP now carry a label, and so does OpenGist.
**One thing to expect, so it does not look broken:** the label says **„Naprakész"** when an app is
on the newest version, with **no number at all**. A number only appears when the app is *behind* —
and then it is how long the newer version has been waiting, not how old the running one is. That
was your ruling and I think it is still the right one.
**The second one was mine and smaller:** a test I wrote pinned a date and an age, so it passed the
day I wrote it and failed the next morning. Fixed, and I have written down that the same trap may
sit in six other test files — named, not accused; someone has to read them.
The box now writes down which version of each app it is really running, and shows the customer one
small label: **„Naprakész"** or **„Frissítés elérhető — 52 napja"**. No version numbers — a
@@ -104,7 +104,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
| Protected infra stacks can't be stopped/removed from UI | controller | **PROVEN-LIVE** | `CAMPAIGN-nomercy` + `RERUN-p1p3` T-SEC-PROTECTED (refuse stop/remove, stay Up) | (Cited `CAMPAIGN-2` T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal) |
| **What VERSION a box is running, and whether it is behind the catalog** | controller v0.233.0 + catalog `69761cf` | **PROVEN-LIVE — both the record and the rendered badge.** | **`tests/VALIDATION-update-slice12-2026-09-02.md`** — on demo-hp 0.233.0, through a REAL production caller (`bootrecon → StartStack → compose up -d → recordInstalledImages`, no hand-set state): `bentopdf` recorded **1** service and `bookstack` recorded **2**, keyed by compose SERVICE name, and **all three digests match the ground truth read independently from the containers before anything was touched**. `catalog_since` reached the box on the normal 15-minute sync. bookstack's two encrypted secrets are byte-identical across the write. Badge evidence: same file §4. Unit side: `installed_test.go` + `updatebadge_test.go`, incl. a wiring test through a real `RestartStack`, an AST walk of all four call sites, and three companion red-proofs. | **THE UNEXERCISED LEG, NAMED: one badge STATE of four.** „Frissítés elérhető" WITHOUT an age needs an app whose `catalog_since` is absent, malformed or future-dated, and all 53 now carry a valid one — unit-tested only (`TestGroupF`). The other three are live: „Naprakész" ×2 on `/stacks` and on `/apps/bookstack`; **NO badge at all on `/apps/docmost`**, a deployed app with no record — absent is UNKNOWN and is not rendered as current; and „Frissítés elérhető — **52 napja**" on both surfaces, the age being real arithmetic on bentopdf's `catalog_since` 2026-07-12. **The behind state was staged by editing bentopdf's compose tag ONLY** — no container restarted, no `up -d` — then reverted byte-identically (`sha256` equal, `diff` empty); so the RENDER is measured and the syncer's own half stays measured separately in the spike. Searched with `grep -oF` ASCII fragments plus positive AND negative controls. **The Frissítés/Újraindítás/Leállítás buttons are unchanged on the live page in the behind state.** **Absent means UNKNOWN, never current** — a legacy `app.yaml` renders NOTHING, red-proved. **No version number is shown to the customer** and **no registry is queried**, so „Naprakész" CAN BE FALSE for the 23 floating pins (**R-446**). Nothing about updating changed: R-438, R-440, R-441, R-443 all stand. Reasoning: `architecture/09-update-architecture.md`; remaining slices R-447..R-452 |
| **What VERSION a box is running, and whether it is behind the catalog** | controller **v0.234.0** + catalog `69761cf` | **PROVEN-LIVE — both the record and the rendered badge.** | **`tests/VALIDATION-update-slice12-2026-09-02.md`** — on demo-hp 0.233.0, through a REAL production caller (`bootrecon → StartStack → compose up -d → recordInstalledImages`, no hand-set state): `bentopdf` recorded **1** service and `bookstack` recorded **2**, keyed by compose SERVICE name, and **all three digests match the ground truth read independently from the containers before anything was touched**. `catalog_since` reached the box on the normal 15-minute sync. bookstack's two encrypted secrets are byte-identical across the write. Badge evidence: same file §4. **v0.234.0 startup backfill, PROVEN LIVE 2026-09-03 on demo-hp:** the record was stripped from `privatebin` (1 service) and `romm` (3 services) to recreate the pre-0.233.0 shape, the controller restarted, and the backfill re-seeded **exactly** those two — every digest matching the ground truth read from the containers beforehand — while logging `2 app(s) recorded, 7 already had a record, 0 left unrecorded`. **All 9 deployed apps then carried „Naprakész" on `/stacks`** (ASCII fragments with a negative control at 0). On demo-felhom the operator's own case, OpenGist, now renders the badge. **Why the backfill exists at all: without it the label never reached an app that simply runs**, which the operator found the morning after v0.233.0. Unit side: `installed_test.go` + `updatebadge_test.go`, incl. a wiring test through a real `RestartStack`, an AST walk of all four call sites, and three companion red-proofs. | **THE UNEXERCISED LEG, NAMED: one badge STATE of four.** „Frissítés elérhető" WITHOUT an age needs an app whose `catalog_since` is absent, malformed or future-dated, and all 53 now carry a valid one — unit-tested only (`TestGroupF`). The other three are live: „Naprakész" ×2 on `/stacks` and on `/apps/bookstack`; **NO badge at all on `/apps/docmost`**, a deployed app with no record — absent is UNKNOWN and is not rendered as current; and „Frissítés elérhető — **52 napja**" on both surfaces, the age being real arithmetic on bentopdf's `catalog_since` 2026-07-12. **The behind state was staged by editing bentopdf's compose tag ONLY** — no container restarted, no `up -d` — then reverted byte-identically (`sha256` equal, `diff` empty); so the RENDER is measured and the syncer's own half stays measured separately in the spike. Searched with `grep -oF` ASCII fragments plus positive AND negative controls. **The Frissítés/Újraindítás/Leállítás buttons are unchanged on the live page in the behind state.** **Absent means UNKNOWN, never current** — a legacy `app.yaml` renders NOTHING, red-proved. **No version number is shown to the customer** and **no registry is queried**, so „Naprakész" CAN BE FALSE for the 23 floating pins (**R-446**). Nothing about updating changed: R-438, R-440, R-441, R-443 all stand. Reasoning: `architecture/09-update-architecture.md`; remaining slices R-447..R-452 |
| Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad `backup:` block degrades to legacy, loudly) | controller v0.132, catalog | **PROVEN-LIVE** | `CAMPAIGN-2` T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs | |
| Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) | agent v0.94.0→**v0.95.0**, controller v0.169.0→v0.171.0→**v0.215.0**, hub v0.73.1 | **PROVEN-LIVE (healthy path + delivery + the severity wire).** **IMPLEMENTED, NOT proven-live: the Hiba-from-counters path** (v0.215.0) — it has never fired on real hardware, only against the committed fixture's values in unit tests (**R-332**) | **2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix):** the card on guest 9201 now shows BOTH real disks with **real verdicts + human model labels** — **„AirDisk 512GB SSD" → Rendben (34°C)** (the system SSD, via LVM/dm resolution) and **„TOSHIBA MQ04ABF100" → Rendben (30°C)** (the USB, via union-path SMART). `/disks` carries `smart.health=PASSED` + `model_name` for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (`SPIKE-smart-coverage-2026-07-25.md` had proven both disks answer `smartctl -a -j` PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. **Notification pipeline PROVEN-LIVE 2026-07-24** — a `disk_health_degraded` POST (the exact `notify.PushEvent` wire call) was **400-rejected by hub v0.73.0** and **200-accepted + „Operator email sent" by hub v0.73.1** | No new smartctl load; feature-detect by payload presence → **MinAgent floor unchanged**; no sudoers/`-d sat` change. **No global banner** (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin `local` on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + `model_name` capture. **2026-08-14 — a genuinely failing disk HAS now been seen, and it broke three assumptions** (`audits/DIAG-smart-passed-trap-2026-08-14.md` + two committed fixtures: raw `smartctl -a -j` and 406 `smartd` lines from ST3000VX010 S/N Z6A07P2G). **(1)** `smart_status.passed` is STRUCTURALLY incapable of failing on unreadable sectors — attrs 187/197/198 all carry `thresh: 0` and a normalized value floors at 1 — so the drive read PASSED at 352 pending sectors and 1001 uncorrectable reads. **(2)** The alert it did produce carried severity `"warn"`, which the hub coerces to `info` and never emails: **the counterfactual is ZERO emails about this drive** (R-328, fixed controller v0.215.0, and the `warning`-vs-`warn` pair proven side by side in `notification_log` on 2026-08-14 — `sent` vs no row at all). **(3)** The old check spoke once and forgot on restart, so between 8 and 352 sectors it emitted nothing. v0.215.0 adds the sustained/count/heat Hiba rules, persisted state and an hourly cadence. **The verdict half of that arm remains unit+red-proof covered only** — no live drive has reached Hiba from counters (R-332). **SMART history/trending (hub-side) PARKED** (ROADMAP R-73) |
| App crashes → customer notified (one event per transition, no flapping spam) | controller v0.120, hub v0.48 | **IMPLEMENTED** | controller v0.120.0 (dead-app alerting, `app_start_failed`, one-event-per-transition red-proofs); `CAMPAIGN-3` F11 surfaced the gap | End-to-end crash→customer-email delivery never live-confirmed (6B deferred / 6C inconclusive: clean stop ≠ crash); anti-spam unit-proven |
@@ -139,6 +139,7 @@ prerequisite for judging how urgent it is.
| # | slice | status |
|---|---|---|
| **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** |
| **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** |
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** |
| **3** | **The compose file becomes DERIVED** — stop the syncer overwriting a deployed app's file; the pin in `app.yaml` wins. Needs the operator's ruling on R-438 first. | OPEN — R-447 |
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 |
@@ -180,6 +181,22 @@ Three rules, each with its reason:
- **It is NOT called from `StartStackServices`** — the R-47 DB-only restore window would overwrite a
complete record with a partial one.
**Seeded at startup (v0.234.0).** `Manager.BackfillInstalledImages` runs once at boot, beside the
desired-state backfill and before the boot reconciler, and records what every deployed app is ALREADY
on. It only READS containers. **This was not a refinement — without it the feature did not reach a
quiet box at all:** see §8.3, which was written as a known limitation on 2026-09-02 and was a defect
by the next morning.
Two admission rules, and the second is the design:
- **It never overwrites an existing record.** The bring-up paths own updates; this fills gaps only.
- **It refuses to seed a PARTIAL observation.** §7.2's comparison reads a service-count mismatch as
BEHIND, so a degraded or crash-looping app seeded from its visible containers would render
„Frissítés elérhető" over an app that is perfectly current. The bring-up paths may write a partial
because they follow a SUCCESSFUL `up -d`, where a gap is real news and is logged; a backfill meets a
box in whatever state it is in. **Same field, two writers, two different admission rules — that is
deliberate and must not be "made consistent".**
**Nothing reads it to take a decision.** Slice 2 reads it to render a label.
### 7.2 The label (slice 2)
@@ -209,10 +226,15 @@ Version strings stay in the logs, the API and the hub.
2. **Nothing enforces `catalog_since`.** A commit that moves an `image:` line and forgets the date
under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent
commit to diff against, so the gate needs a deeper fetch — **R-452**.
3. **The record only appears after the next lifecycle action.** An app that is running and untouched
keeps a legacy `app.yaml` and therefore no badge, until someone restarts, updates or redeploys it.
That is correct — the alternative is inventing a record from the file §1.2 says has already moved —
but it means the fleet view fills in gradually rather than at upgrade.
3. ~~**The record only appears after the next lifecycle action.**~~ **CLOSED in v0.234.0, and the way
it closed is worth keeping.** This was written on 2026-09-02 as an accepted limitation — *"the
fleet view fills in gradually"*. The operator looked at demo-felhom the next morning and found
OpenGist, up 15 hours, running exactly the catalog pin, showing **nothing at all**. **On a quiet
box "gradually" means "never", and a feature that fills itself in on an event nobody triggers is,
on the quiet installations, not shipped.** `BackfillInstalledImages` now seeds the absences at
startup by reading containers (§7.1). **The residue that stays:** the seed happens at controller
START, so a box between upgrade and its next restart still shows nothing — bounded by one restart
rather than unbounded.
4. **The hub does not record image tags at all.** Its report's container payload carries name, state,
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
change; it is not derivable from what is already reported.
+1
View File
@@ -694,6 +694,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** |
| **R-455** | **[P2-MEDIUM] DooPlex has no Docker Hub login, and the unauthenticated ceiling now blocks BUILDS, not just gates.** MEASURED 2026-09-02: `build.sh 0.233.0 --push` failed at `[internal] load metadata for docker.io/library/debian:bookworm-slim: 429 Too Many Requests`, with `golang:1.24-bookworm` cancelled behind it. Neither base image was in the local store, so there was no fallback. The same window also left the catalog's `image-resolvable` gate INCONCLUSIVE (6 of 65 lookups throttled). **This is R-41's known note — *"the full sweep is still OWED — DooPlex is not logged in to Docker Hub"* — arriving with a bigger bill: it is no longer only a gate that cannot reach a verdict, it is a release that cannot be built.** **WORKAROUND USED AND RECORDED RATHER THAN BURIED:** both base images were pulled from Google's official Docker Hub mirror (`mirror.gcr.io/library/...`) and retagged locally, after which `build.sh` ran unmodified; the two digests are written into `felhom-controller/REPORT.md` §5 so the identity check against Hub is one command when the window clears. **THE IDENTITY CLAIM IS NOW MEASURED, not left as an IOU:** the throttle cleared 40 minutes later and `docker pull docker.io/library/debian:bookworm-slim` and `…/golang:1.24-bookworm` both answered **`Status: Image is up to date`** — Docker Hub's own manifest resolved to the images already local, i.e. the ones the mirror supplied and v0.233.0 was built from; and the `docker manifest inspect` bodies are identical between the two registries (`959bc47a76ff713a…` debian, `de3a17b36657e232…` golang). **So the shipped image is byte-for-byte what a Docker-Hub build would have produced, and the mirror is a sound source — which is what makes option (b) below a real option rather than a hope.** **The fix is a credential, and it is a decision:** a Docker Hub account for DooPlex (free tier lifts the ceiling ~6x for authenticated pulls), or an explicit standing ruling that the mirror is the sanctioned source and `build.sh` should name it. **A build that depends on an anonymous third-party quota is not a build you can run when you need to.** Owner: **VIKTOR decides, CC executes.** | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR decides, CC executes** |
| **R-456** | **[P3-LOW] A partly-dead stack is not a boot orphan, and that is written down nowhere.** MEASURED 2026-09-02 on demo-hp while validating v0.233.0: `docker rm -f bookstack` (leaving `bookstack-db` running) then a controller restart produced `Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]` — **bookstack was NOT selected**, although the app container was gone and `desired_state: running` was recorded. Removing `bookstack-db` as well made the whole stack orphaned and the very next pass repaired it in 6.3 s. **So `bootrecon.isBootOrphan` requires the stack as a WHOLE to be down; one live member is enough to make it invisible to the reconciler.** **NOT called a defect, and the reason is part of the row:** `StateDegraded` IS in `IsDownState`, and the crash-loop/dead-app alarm path (`classifyRunStates`) does count a degraded stack as down — so the customer IS told; it is the automatic REPAIR that does not fire, and there may be a good reason (repairing half a stack while its DB is live is not obviously safe). **What is certain is that nobody has written the rule down**, so the next session re-derives it the same way this one did — by watching a reconciliation not happen, which is an absent observable and the weakest possible evidence. Either state the rule in `02-controller-module-map.md` with a test pinning it, or change it. Owner: **CC.** `tests/VALIDATION-update-slice12-2026-09-02.md` §2.2 | **READY — rank P3-LOW; owner: CC** |
| **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
@@ -215,6 +215,59 @@ The badge is wired to nothing.
future-dated, and all 53 catalog apps now carry a valid one. Covered by `TestGroupF`, which walks
absent, blank, `tegnap`, `18/07/2026`, `2026-13-45` and a future date.
## 4b. v0.234.0 — the startup backfill, PROVEN LIVE 2026-09-03
**Why it exists.** §4 of this file, written 2026-09-02, recorded *"the record only appears after the
next lifecycle action"* as an accepted limitation. The operator looked at **demo-felhom** the next
morning and found **OpenGist** — up 15 hours, running exactly the catalog pin, showing **no badge at
all**. On a quiet box "fills in gradually" means "never".
### The staging, and why it was needed
By the time v0.234.0 was ready, **all nine deployed apps on demo-hp and the one on demo-felhom already
had records** — the overnight backup cycle had restarted them and v0.233.0's recorder had fired on
every one (opengist's record is stamped `2026-09-03T00:31:11Z`). **So the natural fleet state could no
longer exercise the new code**, and saying that plainly matters more than a green log line.
The pre-0.233.0 shape was therefore recreated on **demo-hp** (Tier 0): the `installed_images:` block
was deleted from `privatebin` (1 service) and `romm` (**3** services) — a record, never data — and the
controller restarted. Ground truth was read from the containers first, independently.
### The result
```
09:59:45 installed.go:519: [INFO] installed-images backfill: privatebin recorded 1 service(s) (privatebin=privatebin/pdo:2.0.5 (sha256:8a2cac16eff6…))
09:59:45 installed.go:519: [INFO] installed-images backfill: romm recorded 3 service(s) (romm=rommapp/romm:5.0.0 (sha256:91f6611eca5a…), romm-db=mariadb:11.4 (sha256:4f1d8d202fcf…), romm-redis=redis:7-alpine (sha256:ff02b58f971e…))
09:59:45 installed.go:525: [INFO] installed-images backfill: 2 app(s) recorded, 7 already had a record, 0 left unrecorded
```
- **Exactly the two stripped apps were seeded**, and **the other seven were not touched** — the
never-overwrite rule, observed rather than asserted.
- **Every digest matches the ground truth** read from the containers beforehand:
`privatebin/pdo@sha256:8a2cac16eff6caed4dc622e7b3ebd0c0ffcb28dfc0eba0b9c335636552eadac1`,
`rommapp/romm@sha256:91f6611eca5a4dafc4f4a1d72a1ed7dd66a11375d939f28410dc1d1de0b80b1b`,
`mariadb@sha256:4f1d8d202fcf7bcb3902f63af09f9c1a050c2922a89652f22abaec0d4f015e83`.
- The re-seeded records are identical to the ones stripped, apart from `at` — which is correct: it
records when THIS observation was first made.
- **`/stacks` on demo-hp then carried „Naprakész" ×9** — one for every deployed app — with the
negative control `zzz-never-present` at **0** and `napja` at **0** (nothing is behind on that box).
- On **demo-felhom** the operator's own case renders:
`<span class="tag tag-ok" …>Naprakész</span>` on `/stacks?filter=running` **and** on `/apps/opengist`.
### What was NOT proven live, and why the choice was made
**The refusal half — that a partial observation is not seeded — was not staged on a live box.** Doing
so means stopping one container of a multi-service app to make it degraded, and a degraded app is a
dead-app alarm candidate; **this project has already paid for 61 false customer e-mails from exactly
that class (R-330), and manufacturing one to demonstrate a guard is a bad trade.** It is covered by
`TestGroupG_BackfillRefusesAPartialObservation` **with a companion red-proof** — removing the guard
makes it fail with *"backfilled 1, want 0"*. Stated as unproven-live rather than left unmentioned.
### Teardown
The two `app.yaml` backups taken on the box were removed. **Nothing was provisioned** and no app was
started, stopped or upgraded at any point.
## 5. End state — nothing left broken, nothing provisioned
```