R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s

The gate failed only on `released > baked`, so it could catch a forgotten bake
and nothing else. A golden AHEAD of the record passed silently - and that is
how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG
heading still read v0.221.0, with every gate green. Reproduced on the real
history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0.

The gate now asks whether the version being shipped is WRITTEN DOWN: the baked
version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG.
Membership rather than `baked > released` deliberately - a comparison against
the newest heading alone goes green the moment any later entry is written,
leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2)
preserved; every refusal names a reason and a route.

Red-proofed both directions: old gate/old record exit 0, new gate/old record
exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0.

08-alarm-ladder.md is new, and its absence was itself the finding: no document
owned "when does a broken app raise an alarm?". The rules lived as comments in
four packages, each locally correct, with the ordering between them legible only
by reading one function top to bottom - which is how R-384 survived review.

R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed
closed. R-386 filed OPEN: a single-container app stopped out of band raises no
alarm, and a comment claims the opposite - measured live, 9 scans, 0 events,
against a positive control from the same box 17 minutes earlier. Not fixed here.

Golden 0.222.0 baked and published; vouching is the operator's act.
This commit is contained in:
2026-08-23 07:59:52 +02:00
parent 1eb64bec51
commit 55274d5ef3
39 changed files with 4986 additions and 65 deletions
+83 -48
View File
@@ -1,62 +1,97 @@
# REPORT — R-361 and the two loose ends v0.220.2 left (2026-08-22 → 23)
# REPORT — felhom.eu: the golden-currency gate could not see an unrecorded golden (R-385)
Companion to `felhom-controller` **v0.221.0 → v0.221.1**. Full record:
`documentation/audits/DRILL-r361-2026-08-22/`.
**Session 2026-08-23.** Companion to `felhom-controller` v0.222.0 (R-384, R-383) — see that repo's
`REPORT.md` for the controller work and the full live walk.
## What this repo carried
## What was wrong here
- **`documentation/architecture/07-backup-architecture.md`** — a dated **[FACT]** on R-361 (the
comment that asserted an invariant the code did not have, and what it cost), and a **[DESIGN]** on
the `db_dumps` decision **including the trap it created**: a stable list lets
`CaptureRecoveryUnit`'s already-current early return fire, so anything that must happen on every
capture has to sit above that check.
- **`documentation/architecture/00-capability-map.md`** — the **negative** from Part 3, recorded so it
is not re-derived: a HELD app does **not** raise the dead-app alarm, measured on the shipped build,
and the reading that said it would was wrong and why.
- **`STATUS.md`** — the outcome in plain words; the deciding section says what happens if nothing is
done.
- **`documentation/tests/golden-0.221.1-2026-08-23/`** — the golden bake.
- **Register** — R-361 closed and compressed; **R-383** and **R-384** opened.
`scripts/golden_currency_gate.py` asked ONE question — *is the golden BEHIND the record?* — and
therefore could only ever catch a forgotten bake. **It said nothing when the golden was AHEAD of the
record**, and that direction is not harmless: a golden ahead of every CHANGELOG heading was built
from something never written down.
## Part 3 — a measurement that cancelled a Part, and that is a good outcome
That is not hypothetical. Controller **0.221.1** was built, baked **and vouched** on 2026-08-23 while
the newest heading in the controller CHANGELOG still read `v0.221.0`. Measured on the real history,
with the old gate:
The runbook's reading was that a held app alarms as a dead app. **It does not.** On the shipped
v0.220.2 a hold was created deliberately; `docmost` aggregated to `unhealthy`; `IsDownState` is
`{stopped, exited, degraded}`; the dead-app heartbeat read **`0 currently down`** at scans 600 and
620 with the scans demonstrably running over it. **Part 2 was dropped in full** and **no register row
was opened**, exactly as the runbook directs.
```
newest released controller : 0.221.0 (## v0.221.0 — taking the undo copy destroyed …)
newest golden baked : 0.221.1 (documentation/tests/golden-0.221.1-2026-08-23)
golden currency gate OK …
EXIT=0
```
**The positive control took three attempts, and that is the second finding.** Two live attempts
failed to produce a lasting down state at all — `privatebin` went `stopped` (whitelisted by design)
and `bookstack` went `degraded` then `unhealthy`. An absent alarm from a detector never shown working
proves nothing, so the control was moved to the layer the detector lives in: `classifyRunStates` is a
pure function, and it raises the banner for `degraded`/`exited` while staying silent for the states
measured live.
Every gate was green while the fleet ran a version the record did not name.
## Findings opened
## The fix, and why it is membership and not a comparison
- **R-383 (MEDIUM)** — the double-failure message tells the customer *"a korábbi állapot mentése
megvan"* while naming the very file whose absence caused the failure. Observed on **both** v0.220.2
and v0.221.1. **R-361's own class** — a sentence asserting a property the code does not check.
- **R-384 (MEDIUM)** — an app whose **database** has died reads `unhealthy` and raises no dead-app
banner and no customer e-mail, because `aggregateState` checks `unhealthy > 0` before the
mixed-case degraded branch. Not invisible everywhere (the health report counts it), but it does not
alarm.
The gate now asks **"is the version we are shipping WRITTEN DOWN?"** — the baked version must have its
own `## vX.Y.Z` heading **anywhere** in the controller CHANGELOG, not merely at the top (an entry may
legitimately be overtaken by later ones; what may never happen is that it is absent).
**Register size: `OPEN-ITEMS.md` 325 236 → 327 266 bytes; `CLOSED-ITEMS.md` 66 777 → 68 464.**
R-361's full text: `git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md`.
**Membership, not `baked > released`, deliberately:** a comparison against the newest heading alone
goes green the moment ANY later entry is written — which would have left 0.221.1 permanently
unrecorded and the gate permanently silent about it.
## The golden was baked here, and why
Preserved unchanged: **INCONCLUSIVE (exit 2)** for an absent clone or an unparseable CHANGELOG — *not
knowing is never a pass, and never a conviction*. Every refusal names a reason **and** a route,
including what to do if a bake was a throwaway that must never be delivered.
`golden-currency` refused the docs push: 0.221.1 released with no golden. **Not circular** — a golden
needs the controller image, already pushed, not this commit — so the gate was satisfied by doing the
work rather than bypassed. **No push in this session used `--no-verify`.**
## Red-proofs — both directions, against the real history
Golden **0.221.1**, sha256 `1c8bf6cf08cadabeca6331f38360d905e867c235067cd10c2716915b6e6df089`,
656 966 079 B, round-trip verified, all markers hit, both negative controls at zero, both token-leak
greps proved able to convict before their zeros were accepted.
| Run | Gate | CHANGELOG | Golden baked | Exit | |
|---|---|---|---|---|---|
| `gate-01` | **old** | v0.221.0 | 0.221.1 | **0** | the blindness, reproduced |
| `gate-02` | **new** | v0.221.0 | 0.221.1 | **1** | convicted |
| `gate-03` | new | v0.221.1 | 0.221.1 | **0** | Part 0's heading makes it pass |
| `gate-04` | new | absent clone | — | **2** | INCONCLUSIVE preserved |
| `gate-05` | new | v0.222.0 | 0.222.0 | **0** | post-bake |
## Operator follow-up
Transcripts: `documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt`.
**Vouch** the golden — a **three-field** save: `golden_version` **0.221.1**, `agent_version`
**0.130.0**, `min_agent` **0.129.0**. **Then** raise the floor to **0.221.1**, last, in its own save.
## Files changed
| File | Change |
|---|---|
| `scripts/golden_currency_gate.py` | the unrecorded-golden conviction; `newest_released` → `released_versions` returning the whole set; docstring records the second blindness |
| `documentation/architecture/08-alarm-ladder.md` | **NEW.** The alarm ladder as a dated [DESIGN] |
| `documentation/architecture/00-capability-map.md` | R-384 marked closed with its live evidence; points at the new doc |
| `documentation/backlog/OPEN-ITEMS.md` | R-383/R-384 removed (closed); **R-385** (closed) and **R-386** (open) filed |
| `documentation/backlog/CLOSED-ITEMS.md` | R-383 + R-384 compressed, each keeping its rules and naming `git show 1eb64bec5183:…` for the full text |
| `documentation/tests/golden-0.222.0-2026-08-23/` | **NEW.** Bake evidence + log + the vouching instructions |
| `STATUS.md` | the 0.222.0 vouch replaces the (now completed) 0.221.1 one; R-386 added in plain words |
| `documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/` | **NEW.** The drill record and 29 evidence files |
## The alarm ladder had no owning document — that absence is a finding
Nothing in `documentation/architecture/` owned the question *"when does a customer's app being broken
raise an alarm?"* The rules lived as comments across four packages, each locally correct, with the
ordering between them legible only by reading `aggregateState` top to bottom. **That is precisely how
R-384 survived review**, and three separate defects in this ladder (R-51, C9-F2, R-384) were each
found on live hardware rather than by reading. `08-alarm-ladder.md` now owns it.
## Golden
**Baked and PUBLISHED: 0.222.0.** `GOLDEN_SHA256 = 19f5904f5379…`, `upload OK (HTTP 201)`, round-trip
`HTTP 206` from the package URL, all acceptance markers counted.
**VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
**Deviation recorded:** the bake runbook's §4.1 is missing a `pveam update`. On the `virgin` snapshot
the template index is stale, so the listed template cannot be downloaded and the failure presents as
`400 Parameter verification failed. template: no such template` rather than as a stale index.
## Register size
| File | Before | After |
|---|---|---|
| `OPEN-ITEMS.md` | 327,266 B | **328,325 B** |
| `CLOSED-ITEMS.md` | 68,464 B | **71,441 B** |
OPEN grew ~1 KB despite two closures, because R-386 is a substantial new finding. Recorded rather
than smoothed over.
## Hub numbers as read at session start (live, `GET /configuration`)
`golden_version` **0.221.1** · `agent_version` **0.130.0** · `min_agent` **0.129.0** ·
controller floor **0.221.1**. The task expected 0.220.2/0.220.2; the operator had already vouched.
**The hub was READ ONLY this session** — nothing was written to it.
+21 -9
View File
@@ -1,7 +1,7 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-08-22 — the off-site restore now works for all 53 apps, not 13. It is released and
NOT yet delivered: two steps below are yours.**
**Updated 2026-08-23 — an app whose database dies now raises an alarm. It did not before, and the
watcher said "nothing is down" the whole time. Released and NOT yet delivered: step 1 is yours.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
@@ -12,14 +12,26 @@ NOT yet delivered: two steps below are yours.**
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **Vouch the golden carrying controller 0.221.1** — Hub → Configuration → Day-0 artifacts.
1. **Vouch the golden carrying controller 0.222.0** — Hub → Configuration → Day-0 artifacts.
**It is already baked, published and round-trip verified**
(`documentation/tests/golden-0.221.1-2026-08-23/`); only the vouch is left, and only you can do it.
**It is a THREE-field save:** `golden_version` → **0.221.1**, `agent_version` → **0.130.0**,
`min_agent` → **0.129.0**. **Then** raise the floor to **0.221.1**, last, in its own save.
**If you do nothing:** the fleet stays on 0.219.0, so a failed database restore still leaves an app
broken with an unusable copy — the thing today's release fixes reaches nobody. New machines still
receive 0.219.0. The build system stays red about it and will mail you on every push.
(`documentation/tests/golden-0.222.0-2026-08-23/`); only the vouch is left, and only you can do it.
**It is a THREE-field save:** `golden_version` → **0.222.0**, `agent_version` → **0.130.0**,
`min_agent` → **0.129.0**. **Then** raise the floor to **0.222.0**, last, in its own save.
**If you do nothing:** the fleet stays on 0.221.1, where an app whose database has died reports
nothing at all — no banner, no event — and the watcher keeps printing "0 currently down". New
machines still receive 0.221.1.
*(Thank you — the 0.221.1 vouch from earlier today has landed; the hub reads golden 0.221.1 and
floor 0.221.1. Nothing is owed on that one.)*
2. **An app that is stopped from outside still reports nothing** (R-386) — and this one I found today
and deliberately did **not** fix. If a single-container app is stopped by hand on the machine
rather than through the product, nothing is said: no banner, no e-mail, no operator event. I
measured it: nine checks ran over four minutes and every one stayed silent. A comment in our own
code claims the opposite, which is why nobody noticed. **The reason I stopped rather than fixed it:**
from the outside this looks exactly like a customer pressing Stop, and the obvious fix would start
alarming every time somebody legitimately stops their own app. That trade is a decision, not a
patch. **If you do nothing:** it stays as it is — this is not a new fault, it has always been so;
it is newly *known*.
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
@@ -67,6 +67,23 @@ the banner, pinned by `TestClassifyRunStates_PositiveControl_ADownStackDoesAlarm
measurement DID expose is **R-384**: an app whose database has died is `unhealthy` too, and is
likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt`.
> **R-384 CLOSED in controller v0.222.0 (2026-08-23), proven live.** The defect was the ORDER of two
> questions, not the `unhealthy` exclusion: `aggregateState` now asks *"is a supervised member dead?"*
> **before** the `unhealthy`/`starting`/`restarting` returns, and "some members are up" counts any
> member not in the down bucket rather than `running` alone. `IsDownState` is byte-identical.
> **Measured on `demo-hp` 2026-08-23** with the same fixture that read `0 currently down` the day
> before: `bookstack-db` stopped 05:30:07Z → `app_start_failed` fired at **05:30:14Z**, the banner read
> *„Telepített alkalmazás nem fut: BookStack (degraded)"*, the stack read `state=degraded` **while its
> front end was `unhealthy`**, and the heartbeat printed **`1 currently down`** against the previous
> day's `0`. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`.
>
> **The HELD-app half of the paragraph above is now also covered** — a held app keeps its database
> container, so it is the same shape and reaches the same `degraded` verdict.
>
> **The alarm ladder that decides all of this now has an owning document:** see
> `08-alarm-ladder.md` (written 2026-08-23 — before that date no document owned it, and that absence
> is why the ordering defect was legible only from source).
**WHAT IS STILL NOT CLAIMED:** the FAILURE path is where this class is weak, not the success path — a corrupt dump leaves the customer with an emptied or partially-applied database and an undo copy **no product action can apply** (**R-379**), and on MariaDB it does so behind an app that reports `health=healthy` (**R-380**). The success story is proven; the recovery-from-a-bad-restore story is not.
**NARROWED 2026-08-21 by the backup-truth drill — kept as history; both halves have since closed, see the entry immediately above.** Proven that night on `demo-hp` with planted, hash-recorded files: the **declared-userdata** leg of a **drive-declaring** app does come back byte-identical (`calibre-web`, 5/5 including two Hungarian accented filenames). **Two legs of the same story do NOT:** (a) the off-site restore has **no named-volume leg at all**, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (**R-354**); (b) for the **40 of 53** apps that declare no data drive the off-site restore **refuses outright**, saying a running app „nincs telepítve" (**R-356**). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is **unproven for that class and disproven for the volume leg generally**. The escrow/key half of this row is untouched by that and still stands. Evidence: `audits/DRILL-backup-truth-2026-08-21/evidence/` and `REPORT.md` (2026-08-21). |
@@ -0,0 +1,141 @@
# 08 — The app-down alarm ladder
**Written 2026-08-23, with controller v0.222.0 (R-384).**
**The absence is the finding.** Until this file existed, no document owned the question *"when does a
customer's app being broken raise an alarm?"* The rules were spread across four packages as comments,
each locally correct, and the ordering between them was legible only by reading
`aggregateState` top to bottom. That is exactly how R-384 survived: every individual rule was right,
and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each
found on live hardware rather than by review, and each is a case where a reader could not see the
whole ladder at once.
Everything below is **[DESIGN]** — deliberate, with the reason recorded — unless marked otherwise.
---
## 1. The two questions, and their order
Two different questions get asked about a multi-container app, and **the order between them is
load-bearing**:
1. **Is a SUPERVISED member of this app dead?** — a container Docker's restart policy says should be
running, that is not.
2. **Is a RUNNING member failing its healthcheck?**
**Question 1 is asked FIRST.** [DESIGN, R-384, v0.222.0]
Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose
database exits goes `unhealthy` seconds later *because it cannot reach that database*. So the symptom
the dead database causes was the thing that suppressed the alarm for it. Measured on `demo-hp`
2026-08-22 — `bookstack-db` stopped at 21:27:01 and the watcher reported `0 currently down`
throughout.
**"Some members are up" means any member NOT in the down bucket** — `running`, `unhealthy`,
`starting` or `restarting`. [DESIGN, R-384] The earlier guard was `running > 0`, counting only
`StateRunning`, which made the supervised test unreachable in precisely the case it was written for:
an unhealthy survivor beside a dead database counted as nothing being up.
---
## 2. Where each decision is made
| Decision | Where | Notes |
|---|---|---|
| container → stack aggregate state | `internal/stacks/manager.go` `aggregateState` | the ladder in §3 |
| is a down member supervised? | `internal/stacks/manager.go` `supervisedPolicy` | `no`/`on-failure` benign; everything else, **including unknown**, supervised |
| which states mean "down" | `internal/stacks/manager.go` `IsDownState` | `{stopped, exited, degraded}` |
| stack state → "this app is down" | `cmd/controller/main.go` `classifyRunStates` | **the single derivation point**; all three suppressions live here |
| sustained restarting → down | `internal/stacks/manager.go` `CrashLooping` | 5-minute threshold |
| quiesce suppression | `internal/quiesce/suppress.go` | cycle-keyed, 180 s grace |
| boot repair | `internal/bootrecon/bootrecon.go` | consumes `IsDownState` |
---
## 3. The aggregation ladder, in order
`aggregateState(containers, policyOf)` — **priority: degraded > unhealthy/starting > restarting >
all-running > stopped.**
1. no containers → `not_deployed`
2. **any DOWN member is supervised, and any member is up → `degraded`** ← R-384 put this first
3. any `unhealthy` → `unhealthy`
4. any `starting` → `starting`
5. any `restarting` → `restarting`
6. all running → `running`
7. all down → `stopped`
8. mix, every down member benign → `running`
Step 2's `policyOf` is consulted **only** for the down members, and only when something is up. `nil`
is allowed; every down member then reads as supervised.
---
## 4. Which states alarm, and which deliberately do not
`IsDownState` = `{stopped, exited, degraded}`.
| State | Down? | Why |
|---|---|---|
| `stopped`, `exited` | **yes** | not running, will not recover alone |
| `degraded` | **yes** | [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert |
| `unhealthy` | **NO** | [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. **R-384 did not change this** — it asks a prior question instead |
| `restarting` | **NO**, until sustained | [DESIGN, C9-F2] `restarting` is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after **5 min** (`crashLoopAfter`) |
| `starting`, `deploying` | no | mid-start |
| `paused` | no | a deliberate user action |
| `unknown` | no | [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read |
**Note the two fail directions are deliberately opposite.** `IsDownState` fails OPEN on `unknown`
(ambiguous *state* → do not alarm). `supervisedPolicy` fails CLOSED on unknown (we already KNOW a
member is dead; only the excuse is missing). Both are recorded at their sites.
---
## 5. The three suppressions, all at `classifyRunStates`
| Suppression | Rule | Expires? |
|---|---|---|
| **deliberate user stop** | `StateStopped` is not down **unless** the quiesce loop reports it failed to restart that stack | n/a — lifted by `failedRestart` |
| **quiesce cycle** | a stack this backup cycle stopped is exempt | **yes**, 180 s after unquiescing |
| **boot grace** | no evaluation for 90 s after controller start | **yes** |
**None of them latch.** [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false
alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite
direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its
loss.
The quiesce suppression is **cycle-keyed, not state-keyed** — an app caught mid-restart is
`starting`/`unhealthy`, not `stopped`, so no state test can see it. It is therefore **state-blind**,
which is why R-384 moving a stack from `unhealthy` to `degraded` cannot weaken it.
---
## 6. The alarm itself
Edge-triggered: `app_start_failed`, one event per transition into down, **not** per scan. Verified
live 2026-08-23 — one event across 22 scans.
- **Operator/hub event + dashboard banner.** `app_start_failed` is **not** in
`settings.DefaultEnabledEvents`, so by default it does **not** e-mail the customer.
- **F-OBS heartbeat**, every 20 scans (~10 min), at `[INFO]`:
`[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down`.
This line exists because an absent alarm and a stopped detector look identical in a log.
> ⚠ **R-329, OPEN and it bites here.** `app_start_failed` is pushed with severity **`warn`**, which is
> **not** in the hub's vocabulary (`{info, warning, error, critical}`) and is silently coerced to
> `info` — which e-mails nobody, while the POST still returns 200. Observed again on 2026-08-23:
> `PushEvent: type=app_start_failed severity=warn`. R-384 makes this event actually fire, so the
> severity bug now matters more than it did while the event was unreachable.
---
## 7. Known gap, filed not fixed
> **R-386 (filed 2026-08-23, OPEN).** An all-down stack aggregates to `stopped` — `StateExited` is
> folded into the same counter and never survives aggregation. `classifyRunStates` then whitelists
> `stopped` as a deliberate user stop. So a **single-container app stopped out of band raises no
> alarm at all**, which directly contradicts the comment at `cmd/controller/main.go`: *"An out-of-band
> `docker compose stop` leaves the containers present → StateExited → still alerts."*
> **Measured on `demo-hp` 2026-08-23:** `privatebin` stopped out of band, 9 dead-app scans over 4+
> minutes, `state=stopped`, **zero events and zero banner lines** — against a positive control from
> the same box 17 minutes earlier. Not fixed in v0.222.0 deliberately; it is a separate decision.
@@ -0,0 +1,96 @@
# DRILL — R-384: an app whose database dies raised no alarm (2026-08-23)
**Controller v0.221.1 → v0.222.0. Live leg on `demo-hp` (Tier 0, disposable), guest 9201.**
**UNATTENDED.** Method: endpoint-level — no browser exists on DooPlex, so every read is either the
exact endpoint the UI calls or the controller's own log. Guest clock is UTC.
## Verdict
| Part | Outcome |
|---|---|
| Part 0 — the record | ✅ `v0.221.1` given its own heading, pushed ALONE (`da75603`) |
| Part 1 — the blind gate | ✅ fixed; both directions red-proofed against the real history |
| Part 2 — R-384 | ✅ shipped v0.222.0, **proven live** |
| Part 3 — R-383 | ✅ shipped v0.222.0 |
| §4 — the measurement | ⚠ **reproduced. Filed as R-386. NOT fixed — that was the instruction.** |
| Live walk step 4 (Scenario E) | **DROPPED** — consequence of the halt; drop-list item (3) |
## The one-line result
The same fixture that printed **`0 currently down`** on 2026-08-22 printed **`1 currently down`** on
2026-08-23, with 8 apps evaluated both times:
```
2026/08/22 21:13:48 [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down
2026/08/23 05:37:44 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down
```
## What was actually wrong
**The ORDER of two questions**, not the `unhealthy` exclusion. "Is a supervised member dead?" and "is
a running member failing its healthcheck?" are different questions, and the second was answering the
first — because a dying database drags its own front end `unhealthy`, **the symptom the fault causes
was what suppressed the alarm for it.** `IsDownState` was not touched; no state was minted.
The fix has **two halves and either alone leaves the defect standing**: the hoist, and widening "some
members are up" from `running > 0` to *any member not in the down bucket*. The old guard made the
R-51 block unreachable in precisely the case R-51 was written for.
## Evidence index (`evidence/`)
| File | What it shows |
|---|---|
| `gate-01-old-gate-old-changelog.txt` | the blindness: old gate, real history, **exit 0** |
| `gate-02-new-gate-old-changelog.txt` | new gate on the same history, **exit 1**, naming the fix |
| `gate-03-new-gate-new-changelog.txt` | with Part 0's heading, **exit 0** |
| `gate-04-inconclusive.txt` | INCONCLUSIVE (**exit 2**) preserved |
| `gate-05-new-gate-post-bake.txt` | v0.222.0 + golden 0.222.0, **exit 0** |
| `redproof-R384-1-order.txt` | mutation: hoist reverted → `"unhealthy", want "degraded"` |
| `redproof-R384-2-upguard.txt` | mutation: `up` narrowed → all three survivor shapes convict |
| `redproof-R384-3-classifier.txt` | mutation: classifier ignores `degraded` → empty banner |
| `redproof-R383-undo-phrase.txt` | mutation: unconditional claim → prints the false sentence |
| `live-00-the-old-heartbeat-2026-08-22.txt` | yesterday's `0 currently down`, carried in for contrast |
| `live-01`…`live-06` | Scenario A: pre-state, stop, log, decisive read, heartbeat, banner |
| `live-08-golden-bake-markers.txt` | the bake's acceptance markers, each counted |
| `live-09-…-full-controller-log.txt` | 1808 lines, pulled off **before** the app was restarted |
| `live-10`…`live-13` | Scenario B: unhealthy with nothing dead, no alarm, banner cleared |
| `live-14`…`live-16` | Scenario D: full stop→start cycle, **0 alarms across 9 scans** |
| `live-17`, `live-18` | §4: `privatebin` out-of-band stop, 9 scans, **0 events, 0 banner** |
| `live-19-…-full-log.txt` | 2139 lines covering Scenario D and §4 |
## §4 — what it found, in plain words
`aggregateState` folds `StateExited` into the `stopped` counter, so an all-down stack returns
`StateStopped` and **`StateExited` never survives aggregation** — that is the path the task suspected
and could not find in source. `classifyRunStates` then whitelists `StateStopped` as a deliberate user
stop. So this comment in `cmd/controller/main.go` is **false**:
> *"(An out-of-band `docker compose stop` leaves the containers present → StateExited → still alerts,
> which is correct: out-of-band tampering IS reportable.)"*
Measured: `privatebin` (1 container, `unless-stopped`) stopped 05:47:35Z; at 05:51:53Z it read
`state=stopped` with 9 dead-app scans behind it, **zero events and zero banner lines**.
**The absence is trustworthy because the detector was shown alive first** (standing rule 3):
`app_start_failed` fired for BookStack at 05:30:14Z on the same box 17 minutes earlier.
**Scoped honestly:** a genuine crash under `unless-stopped` is restarted by Docker and surfaces as
`restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit
out-of-band stop of a stack with no surviving member.
## Two things noticed that are NOT this drill's work
1. **R-329 moved from unreachable to load-bearing.** `app_start_failed` ships severity `warn`, which
is not in the hub's vocabulary and coerces silently to `info` — e-mailing nobody, POST still 200.
Observed again today. While the event never fired, this was harmless; it no longer is.
2. **The golden-bake runbook is missing `pveam update`.** On the `virgin` snapshot the template index
is stale, so `pveam available` offers `13.1-2` and downloading it fails with
`400 Parameter verification failed. template: no such template` — a confusing 400 rather than a
legible "your index is old". Recorded in `documentation/tests/golden-0.222.0-2026-08-23/README.md`.
## Teardown
Nothing was provisioned. All apps restored and confirmed healthy (`bookstack` + `bookstack-db`,
`docmost` ×3, `privatebin`); planted data untouched; no app rebuilt or restored. **Hub-side: nothing
to discard — the hub was READ ONLY this session** (`GET /configuration`, `GET /events`); no appliance
registered, no config written, no artifact manifest changed.
@@ -0,0 +1,5 @@
### OLD gate + OLD changelog (v0.221.0 top, golden 0.221.1 baked) ###
newest released controller : 0.221.0 (## v0.221.0 — taking the undo copy destroyed the app's own database backup (2026-08-22, R-)
newest golden baked : 0.221.1 (documentation/tests/golden-0.221.1-2026-08-23)
golden currency gate OK — the newest released controller has a golden (NOTE: this checks the BAKE, not the vouch — see the module docstring)
EXIT=0
@@ -0,0 +1,9 @@
### NEW gate + OLD changelog (v0.221.0 top, golden 0.221.1 baked) -> must FAIL ###
newest released controller : 0.221.0 (## v0.221.0 — taking the undo copy destroyed the app's own database backup (2026-08-22, R-)
newest golden baked : 0.221.1 (documentation/tests/golden-0.221.1-2026-08-23)
GOLDEN CURRENCY GATE FAILED: golden 0.221.1 is baked but UNRECORDED — the controller CHANGELOG has no '## v0.221.1' heading.
The newest heading is 0.221.0. A golden ahead of the record was built from a version nobody wrote down, so no one can read what the fleet is running.
Fix: give v0.221.1 its own '## v0.221.1 — <what changed>' heading in felhom-controller/CHANGELOG.md, above the entries it supersedes. If its fix is currently described inside another version's entry, MOVE that text — do not duplicate it, and do not delete the reasoning.
If this bake was a throwaway that must never be delivered, delete its documentation/tests/golden-<VER>-<DATE>/ directory — never leave it to read as shipped.
EXIT=1
@@ -0,0 +1,5 @@
### NEW gate + FIXED changelog (v0.221.1 heading present) -> must PASS ###
newest released controller : 0.221.1 (## v0.221.1 — the undo-copy prune stopped running because another fix made its guard reach)
newest golden baked : 0.221.1 (documentation/tests/golden-0.221.1-2026-08-23)
golden currency gate OK — the newest released controller has a golden (NOTE: this checks the BAKE, not the vouch — see the module docstring)
EXIT=0
@@ -0,0 +1,3 @@
### INCONCLUSIVE preserved: unreadable CHANGELOG path ###
GOLDEN CURRENCY GATE INCONCLUSIVE: controller clone not found at /nonexistent/CHANGELOG.md
EXIT=2
@@ -0,0 +1,5 @@
### NEW gate, post-bake: CHANGELOG v0.222.0 + golden 0.222.0 -> must PASS ###
newest released controller : 0.222.0 (## v0.222.0 — an app whose database dies raised no alarm, because the wrong question answe)
newest golden baked : 0.222.0 (documentation/tests/golden-0.222.0-2026-08-23)
golden currency gate OK — the newest released controller has a golden (NOTE: this checks the BAKE, not the vouch — see the module docstring)
EXIT=0
@@ -0,0 +1,18 @@
=== PART 3 MEASUREMENT — v0.220.2, no code change
measured at: 21:13:51Z (hold created 21:11:19Z)
elapsed: ~2m20s = 5 deadapp scans at 30s cadence (heartbeat lines confirm 9 scans in the log window)
-- 3a. WHICH RUN STATE did docmost aggregate to?
bookstack state=running deployed=True
docmost state=unhealthy deployed=True
privatebin state=running deployed=True
-- 3b. IS THE DATABASE STILL UP?
docmost-postgres | Up 2 minutes (healthy)
-- 3c. DID A CUSTOMER-FACING EVENT FIRE for docmost?
2026/08/22 21:11:19 notifier.go:234: [INFO] Event pushed: backup_run_failures (error) — App "docmost" is HELD STOPPED: its off-site database restore failed AND the rollback to the customer's own pre-restore copy also failed. The app will not start from any path until the hold is cleared. Replay error: importing postgres dump for docmost: postgres import into docmost-postgres failed: exit status 3. Rollback error: a visszavonáshoz szükséges mentés nem található (pre-restore-20260822T211114Z-docmost-postgres.sql): stat /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T211114Z-docmost-postgres.sql: no such file or directory
(none above = no event fired)
-- 3d. DEAD-APP BANNER STATE:
2026/08/22 21:13:48 main.go:1730: [INFO] [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down
@@ -0,0 +1,11 @@
=== SCENARIO A — a database dies behind a healthy-looking app (v0.222.0) ===
controller version:
gitea.dooplex.hu/admin/felhom-controller:0.222.0 Up 3 minutes (healthy)
-- PRE-STATE (guest UTC) --
2026-08-23T05:30:02Z
bookstack Up 3 hours (healthy)
bookstack-db Up 3 hours (healthy)
/bookstack restart=unless-stopped
/bookstack-db restart=unless-stopped
@@ -0,0 +1,6 @@
-- STOPPING bookstack-db OUT OF BAND --
2026-08-23T05:30:07Z
bookstack-db
2026-08-23T05:30:07Z
bookstack Up 3 hours (healthy)
bookstack-db Exited (0) Less than a second ago
@@ -0,0 +1,18 @@
-- controller log since the stop (05:30:00Z) --
2026/08/23 05:30:04 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m10s ago, effective interval 5m0s, healthy=true
2026/08/23 05:30:14 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
2026/08/23 05:30:14 manager.go:703: [DEBUG] [stacks] restart-policy of down member "bookstack-db" = "unless-stopped"
2026/08/23 05:30:14 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m20s ago, effective interval 5m0s, healthy=true
2026/08/23 05:30:14 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warn url=https://hub.felhom.eu/api/v1/event
2026/08/23 05:30:14 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
2026/08/23 05:30:14 notifier.go:234: [INFO] Event pushed: app_start_failed (warn) — Telepített alkalmazás nem fut: BookStack
2026/08/23 05:30:24 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m30s ago, effective interval 5m0s, healthy=true
2026/08/23 05:30:34 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m40s ago, effective interval 5m0s, healthy=true
2026/08/23 05:30:44 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
2026/08/23 05:30:44 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m50s ago, effective interval 5m0s, healthy=true
2026/08/23 05:30:44 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "bookstack" deployed=true composePath=/opt/docker/stacks/bookstack/docker-compose.yml
2026/08/23 05:30:54 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 4m0s ago, effective interval 5m0s, healthy=true
2026/08/23 05:31:04 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 4m10s ago, effective interval 5m0s, healthy=true
2026/08/23 05:31:14 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 4m20s ago, effective interval 5m0s, healthy=true
2026/08/23 05:31:14 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
2026/08/23 05:31:24 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 4m30s ago, effective interval 5m0s, healthy=true
@@ -0,0 +1,13 @@
-- DECISIVE READ: front end UNHEALTHY, database EXITED --
2026-08-23T05:32:03Z
bookstack Up 3 hours (unhealthy)
bookstack-db Exited (0) About a minute ago
bookstack state=degraded deployed=True
calibre-web state=running deployed=True
docmost state=running deployed=True
kimai state=running deployed=True
opengist state=running deployed=True
paperless-ngx state=running deployed=True
privatebin state=running deployed=True
romm state=running deployed=True
@@ -0,0 +1,2 @@
-- waiting for the F-OBS heartbeat (every 20 scans = ~10 min) --
2026/08/23 05:37:44 main.go:1730: [INFO] [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down
@@ -0,0 +1,9 @@
-- the alert surface: /api/alerts + the launcher page --
### /api/alerts
http=404 bytes=42
### /launcher
http=200 bytes=42201
<span class="alert-message">Telepített alkalmazás nem fut: BookStack (degraded)</span>
### /dashboard
http=200 bytes=56437
<span class="alert-message">Telepített alkalmazás nem fut: BookStack (degraded)</span>
@@ -0,0 +1,2 @@
-- hub-side: the event as the hub STORED it (R-329 severity check) --
http=404
@@ -0,0 +1,13 @@
=== GOLDEN 0.222.0 BAKE — acceptance markers ===
docker OK (overlay2 : 1
including mount point: 2
upload OK (HTTP 201) : 1
excluding (must be 0): 0
FATAL (must be 0): 0
--- the marker lines ---
docker OK (overlay2; data-root /var/lib/docker)
INFO: including mount point rootfs ('/') in backup
INFO: including mount point mp0 ('/var/lib/felhom') in backup
[golden] upload OK (HTTP 201)
GOLDEN_VERSION=0.222.0
GOLDEN_SHA256=19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037
@@ -0,0 +1,7 @@
=== SCENARIO B — unhealthy with NOTHING dead (the flapping case) ===
-- restarting bookstack-db; the front end stays unhealthy for a while with nothing down --
2026-08-23T05:40:01Z
bookstack-db
2026-08-23T05:40:09Z
bookstack Up 3 hours (unhealthy)
bookstack-db Up 8 seconds (healthy)
@@ -0,0 +1,6 @@
-- SCENARIO B decisive read: front end unhealthy, database UP, nothing dead --
2026-08-23T05:40:20Z
bookstack Up 3 hours (healthy)
bookstack-db Up 18 seconds (healthy)
bookstack state=unhealthy
@@ -0,0 +1,7 @@
-- SCENARIO B: any NEW alarm during the recovery window? --
app_start_failed events since 05:40:00Z: 0
--- what the aggregate read during the window ---
2026/08/23 05:40:04 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m0s ago, effective interval 5m0s, healthy=true
2026/08/23 05:40:14 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m10s ago, effective interval 5m0s, healthy=true
2026/08/23 05:40:24 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m20s ago, effective interval 5m0s, healthy=true
2026/08/23 05:40:34 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m30s ago, effective interval 5m0s, healthy=true
@@ -0,0 +1,6 @@
-- banner cleared + bookstack healthy again --
2026-08-23T05:40:53Z
bookstack Up 3 hours (healthy)
bookstack-db Up 52 seconds (healthy)
banner lines: 0
bookstack state=running
@@ -0,0 +1,13 @@
=== SCENARIO D — a full stop -> start cycle through the PRODUCTION endpoint ===
method: POST /api/stacks/docmost/{stop,start} — the exact call the launcher's buttons make
subject: docmost (3 containers: docmost, docmost-postgres, docmost-redis)
csrf len=64
BASELINE alarms before: 0
T0=2026-08-23T05:42:15Z --- STOP ---
{"ok":true,"message":"Stack docmost stop completed"}
http=200
after stop: 2026-08-23T05:42:28Z
--- START ---
{"ok":true,"message":"Stack docmost start completed"}
http=200
@@ -0,0 +1,11 @@
-- SCENARIO D: watching a FULL cycle settle (5 min past the start) --
05:42:47 docmost:Up 7 seconds (health: starting) docmost-postgres:Up 18 seconds (healthy) docmost-redis:Up 18 seconds (healthy)
05:43:12 docmost:Up 33 seconds (healthy) docmost-postgres:Up 43 seconds (healthy) docmost-redis:Up 43 seconds (healthy)
05:43:37 docmost:Up 58 seconds (healthy) docmost-postgres:Up About a minute (healthy) docmost-redis:Up About a minute (healthy)
05:44:02 docmost:Up About a minute (healthy) docmost-postgres:Up About a minute (healthy) docmost-redis:Up About a minute (healthy)
05:44:27 docmost:Up About a minute (healthy) docmost-postgres:Up About a minute (healthy) docmost-redis:Up About a minute (healthy)
05:44:52 docmost:Up 2 minutes (healthy) docmost-postgres:Up 2 minutes (healthy) docmost-redis:Up 2 minutes (healthy)
05:45:17 docmost:Up 2 minutes (healthy) docmost-postgres:Up 2 minutes (healthy) docmost-redis:Up 2 minutes (healthy)
05:45:42 docmost:Up 3 minutes (healthy) docmost-postgres:Up 3 minutes (healthy) docmost-redis:Up 3 minutes (healthy)
05:46:07 docmost:Up 3 minutes (healthy) docmost-postgres:Up 3 minutes (healthy) docmost-redis:Up 3 minutes (healthy)
05:46:32 docmost:Up 3 minutes (healthy) docmost-postgres:Up 4 minutes (healthy) docmost-redis:Up 4 minutes (healthy)
@@ -0,0 +1,8 @@
-- SCENARIO D VERDICT: alarms across the whole cycle (T0=05:42:15Z) --
app_start_failed events since T0 : 0
deadapp scans in the window : 9
supervised-down path entered : 1
--- any docmost down-state reading? ---
2026/08/23 05:42:34 manager.go:703: [DEBUG] [stacks] restart-policy of down member "docmost" = "unless-stopped"
(no lines above = none)
@@ -0,0 +1,11 @@
=== §4 MEASUREMENT — a SINGLE-container app crashes out of band ===
POSITIVE CONTROL, established on this box today: app_start_failed fired at 05:30:14Z for
BookStack (see live-09). The detector demonstrably works here and now, so an absence below
is a real absence, not a dead detector.
subject: privatebin (1 container, restart policy below). Stopped OUT OF BAND — the controller
did not do it, so no Deploying flag and no quiesce key.
privatebin restart=unless-stopped
T0=2026-08-23T05:47:35Z
2026-08-23T05:47:40Z
privatebin Exited (0) 5 seconds ago
@@ -0,0 +1,12 @@
-- §4: waiting 4 minutes, well past every grace window (deadapp scan 30s, quiesce grace 180s) --
2026-08-23T05:51:53Z
privatebin Exited (0) 4 minutes ago
-- the aggregate state the controller reads --
privatebin state=stopped deployed=True
-- DID ANY ALARM FIRE? --
app_start_failed since T0 : 0
deadapp scans since T0 : 9
-- banner? --
banner lines: 0
@@ -0,0 +1,18 @@
### RED-PROOF R-383 — mutation: undoCopyPhrase reverted to the unconditional pre-fix claim ###
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist (0.00s)
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist/absent_—_must_NOT_claim_it_exists,_must_still_name_where_it_should_be (0.00s)
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-mariadb.sql" does not contain "NEM találjuk"
r383_undo_phrase_test.go:97: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-mariadb.sql" contains "mentése megvan" — it asserts a file that is not on disk
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist/zero-length_—_counts_as_missing (0.00s)
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-empty.sql" does not contain "NEM találjuk"
r383_undo_phrase_test.go:97: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-empty.sql" contains "mentése megvan" — it asserts a file that is not on disk
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist/partial_—_both_halves_named,_neither_hidden (0.00s)
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-postgres.sql" does not contain "RÉSZBEN"
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-postgres.sql" does not contain "HIÁNYZIK"
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-postgres.sql" does not contain "pre-restore-20260823T120000Z-app-mariadb.sql"
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist/no_undo_was_ever_written_—_said_plainly,_not_silently (0.00s)
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: ." does not contain "nem készült"
r383_undo_phrase_test.go:97: phrase "a korábbi állapot mentése megvan: ." contains "megvan" — it asserts a file that is not on disk
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.008s
FAIL
@@ -0,0 +1,9 @@
### RED-PROOF R-384 #1 — the ORDER (mutation: hoist moved back below `unhealthy > 0`) ###
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst (0.00s)
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst/unhealthy_survivor_—_the_bookstack_case_measured_live (0.00s)
degraded_test.go:165: survivor "unhealthy" beside a dead SUPERVISED member: aggregateState = "unhealthy", want "degraded" — a dead database must not hide behind it
--- FAIL: TestR384_WiresTheDeadDatabaseThroughTheRealPath (0.00s)
degraded_test.go:458: bookstack state = "unhealthy", want "degraded" — a dead database must not hide behind its own unhealthy front end (measured live 2026-08-22: this read "unhealthy" and nothing alarmed)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
FAIL
@@ -0,0 +1,13 @@
### RED-PROOF R-384 #2 — the GUARD (mutation: up = running only, the old `running > 0`) ###
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst (0.00s)
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst/unhealthy_survivor_—_the_bookstack_case_measured_live (0.00s)
degraded_test.go:165: survivor "unhealthy" beside a dead SUPERVISED member: aggregateState = "unhealthy", want "degraded" — a dead database must not hide behind it
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst/starting_survivor (0.00s)
degraded_test.go:165: survivor "starting" beside a dead SUPERVISED member: aggregateState = "starting", want "degraded" — a dead database must not hide behind it
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst/restarting_survivor (0.00s)
degraded_test.go:165: survivor "restarting" beside a dead SUPERVISED member: aggregateState = "restarting", want "degraded" — a dead database must not hide behind it
--- FAIL: TestR384_WiresTheDeadDatabaseThroughTheRealPath (0.00s)
degraded_test.go:458: bookstack state = "unhealthy", want "degraded" — a dead database must not hide behind its own unhealthy front end (measured live 2026-08-22: this read "unhealthy" and nothing alarmed)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
FAIL
@@ -0,0 +1,6 @@
### RED-PROOF R-384 #3 — the CONSEQUENCE layer (mutation: classifier ignores StateDegraded) ###
--- FAIL: TestR384_ADeadDatabaseBehindAnUnhealthyAppAlarms (0.00s)
r384_dead_db_alarm_test.go:29: dead-app banner = [], want exactly one entry for bookstack
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.008s
FAIL
+2
View File
@@ -160,3 +160,5 @@
| **R-370** | **PROCESS: between 2026-08-19 and 2026-08-22 the reviewing side called a documented architectural decision a defect, in four places, because it read the register and live source and never `documentation/architecture/`.** Evidence: `documentation/architecture/`. **Reasoning kept:** R-352 (re-framed), R-369 The record is corrected in place with the framing marked rather than deleted, per the standing rule that a document which quietly changes its mind teaches nobody. | **CLOSED — corrected 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
| **R-96** | **Two standing rules were agreed in chat and never committed** Evidence: `documentation/runbooks/workspace-CLAUDE.md:48-70`. **Reasoning kept:** **Two standing rules were agreed in chat and never committed** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-27, size XS, roadmap state `idea — found 2026-07-27`.** Moved verbatim; nothing added or reinterpreted. | **CLOSED — migrated from ROADMAP 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
| **R-107** | **No offsite action unpacks the named-volume tars Tier-3 captures on every run.** Shipped in v0.218.0. | **CLOSED — migrated from ROADMAP 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
| **R-383** | **The double-failure message told the customer their previous state was saved, and named a file that was not there.** Shipped in controller v0.222.0. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`. **Reasoning kept:** *One of the two ways a rollback fails is that the undo copy is missing — so the sentence was most likely to be false in exactly the case it was printed.* **Do NOT simply drop the filename:** an operator needs it, and R-351's lesson is that a refusal naming nothing forces someone to remember what the product already knows — so the absent case still names WHERE the file should have been. **A zero-length dump counts as MISSING**, because a 0-byte file restores nothing and calling it present is the same false reassurance one step smaller. The check is `os.Stat` and deliberately not an integrity test: this runs at the end of a failed restore on a machine that may be unwell, and presence is the honest claim available there. | **CLOSED — SHIPPED** (controller v0.222.0, 2026-08-23; `undoCopyPhrase`, four cases, plus an AST seam test that the message is still wired to the builder) | full text: `git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md` |
| **R-384** | **An app whose DATABASE had died raised no dead-app alarm — the wrong question answered first.** Shipped in controller v0.222.0. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`. **Reasoning kept:** *The defect was the ORDER of two questions, not the `unhealthy` exclusion.* "Is a SUPERVISED member dead?" and "is a RUNNING member failing its healthcheck?" are different questions, and the second was answering the first — a dying database drags its own front end `unhealthy`, so the symptom the fault causes was what suppressed the alarm for it. **`IsDownState` is byte-identical and `unhealthy` stays excluded** — an unhealthy container is RUNNING, and folding it in reintroduces the flapping that exclusion exists to stop; **no new state was minted**, `StateDegraded` already means this. **Two things had to move and either alone leaves the defect standing:** the hoist, AND widening "some members are up" from `running > 0` to *any member not in the down bucket* — the old guard made the R-51 block unreachable in precisely the case it was written for. **The register's own suggested fix was WRONG and is recorded as such:** it proposed a sustained-`unhealthy` threshold on the `crashLoopAfter` model; the actual defect needed no threshold at all. **PROVEN LIVE the only way it can be** — the same fixture that printed `0 currently down` on 2026-08-22 printed **`1 currently down`** on 2026-08-23, with `app_start_failed` 7 s after the stop and the banner reading *„…nem fut: BookStack (degraded)"*. Scenario D measured **0 alarms across 9 scans** through a full stop→start cycle. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.222.0, 2026-08-23) | full text: `git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md` |
+2 -2
View File
@@ -136,8 +136,8 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
| ID | What | State |
|---|---|---|
| **R-383** | **The double-failure message tells the customer their previous state was saved, and names a file that is not there.** The sentence ends *"a korábbi állapot mentése megvan: <file>"* — "the backup of the previous state EXISTS" — built from the path `writeSafetyDump` returned, WITHOUT asking whether it is still on disk. But one of the two ways a rollback can fail is that the undo copy is missing or unreadable, and in exactly that case the sentence is FALSE. **Measured live twice, on v0.220.2 (2026-08-22 21:11:19) and again on v0.221.1 (22:20:32):** rollback failed with `stat …pre-restore-…sql: no such file or directory`, and the customer message named that same file as existing. 369 bytes, `offbox_reconstitute.go` (the double-failure branch). **This is R-361's own class** — a sentence asserting a property the code does not check — one surface over. | **OPEN — MEDIUM** | — | Say what is true: name the undo copy only when it is verifiably on disk, and say plainly when it is not. **Do not simply drop the filename** — an operator needs it, and R-351's lesson was that a refusal which names nothing forces someone to remember what the product already knew. Evidence: `audits/DRILL-r361-2026-08-22/evidence/03-observed-false-sentence.txt`, `audits/DRILL-r361-2026-08-22/evidence/16-part4-message.txt`. | CC |
| **R-384** | **An app whose DATABASE has died raises no dead-app alarm — `unhealthy` masks the mixed state.** `aggregateState` (`internal/stacks/manager.go`) checks `if unhealthy > 0 → StateUnhealthy` BEFORE the mixed-case degraded branch, and `IsDownState` (`manager.go:54`) is `{stopped, exited, degraded}` — `unhealthy` is absent. So a multi-container app whose database container dies goes `degraded` for a moment and then `unhealthy` as its own healthcheck fails, and stops being a fault. **Measured live 2026-08-22:** `bookstack-db` stopped out-of-band at 21:27:01; `bookstack` read `unhealthy`; the dead-app heartbeat reported **`0 currently down`** across the whole window (scans 600 and 620), with 8 apps evaluated. **NOT invisible everywhere** — the health report counts it (`cr.Unhealthy++`, `internal/report/builder.go:251`) and that reaches the hub — but it raises no banner and no customer e-mail. **This is the F-CRIT-1 class the `classifyRunStates` comment says was closed:** it WAS closed for `StateStopped`+failedRestart, and `unhealthy` was never in scope. Found while building a positive control for a different question. | **OPEN — MEDIUM** | — | Decide whether a SUSTAINED `unhealthy` is a fault (it is not a brief one — that is why it is excluded), on the `crashLoopAfter` model: a threshold above the deploy/health windows rather than a state test. **Do not simply add `unhealthy` to `IsDownState`** — it has other callers and would alarm on every deploy, which is the over-correction F-A1 nearly cost. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt`. | CC |
| **R-385** | **A controller was built, baked AND vouched with no CHANGELOG entry of its own, and every gate stayed green.** Controller **0.221.1** shipped on 2026-08-23 while the newest heading in `felhom-controller/CHANGELOG.md` still read `v0.221.0` — the prune-ordering fix (commit `810b18a`) had been written INSIDE the v0.221.0 entry instead of getting its own. The image was never in question; the RECORD was, and the fleet ran a version the record did not name. **`scripts/golden_currency_gate.py` could not catch it by construction:** it failed only on `released > baked`, so a golden AHEAD of the record passed silently. Measured on the real history: `newest released 0.221.0 / newest golden baked 0.221.1 → OK, exit 0`. | **CLOSED — 2026-08-23** | — | **Both halves fixed, both directions red-proofed.** The record: `v0.221.1` has its own heading carrying the MOVED (not duplicated, not deleted) reasoning — commit `da75603`, pushed alone before anything else. The gate now asks *"is the baked version WRITTEN DOWN?"* — the baked version must have its own `## vX.Y.Z` heading **anywhere** in the CHANGELOG. **Membership, not `baked > released`, deliberately:** a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded and the gate permanently silent about it. INCONCLUSIVE (exit 2) preserved. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt` — old gate/old record `exit 0`, new gate/old record `exit 1`, new gate/fixed record `exit 0`. | CC |
| **R-386** | **A single-container app stopped OUT OF BAND raises no alarm at all — and a comment states the opposite as settled fact.** `aggregateState` folds `StateExited` into the `stopped` counter, so an all-down stack returns `StateStopped` and **`StateExited` never survives aggregation**. `classifyRunStates` then whitelists `StateStopped` as a deliberate user stop unless the quiesce loop reports a failed restart. The comment at `cmd/controller/main.go` says: *"(An out-of-band `docker compose stop` leaves the containers present → StateExited → still alerts, which is correct: out-of-band tampering IS reportable.)"* — **measured FALSE.** The neighbouring I2 claim (*"a CRASHING app never comes to rest at `stopped` — faults surface as StateExited"*) is false in the same way. **Measured live on `demo-hp` 2026-08-23 (controller v0.222.0):** `privatebin` (1 container, `unless-stopped`) stopped out of band at 05:47:35Z; at 05:51:53Z it read `state=stopped`, **9 dead-app scans had run, and there were ZERO `app_start_failed` events and ZERO banner lines.** **The absence is trustworthy — positive control from the same box 17 minutes earlier:** `app_start_failed` fired for BookStack at 05:30:14Z, so the detector demonstrably works there. **SCOPE, stated so it is not overclaimed:** a genuine crash under `unless-stopped` is RESTARTED by Docker and surfaces as `restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit out-of-band stop of a stack with no surviving member. **This is case #10 of "a comment asserting an invariant the code does not provide".** Found by §4 of the R-384 task, which asked for a measurement and explicitly forbade a fix in that session. | **OPEN — MEDIUM** | — | Decide whether an out-of-band stop is distinguishable from a customer stop at all — they are byte-identical on the Docker side, exactly as invariant I1 says, so the answer is probably NOT a state test but a recorded intent (`DesiredStateOf` already exists and `bootrecon` already consumes it). **Do NOT simply un-whitelist `StateStopped`** — that re-alarms every genuine customer stop, which is the over-correction F-CRIT-1's fix was careful to avoid. **And fix the comment either way:** it is load-bearing and it is false. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/live-17-sec4-stop.txt`, `live-18-sec4-verdict.txt`, `live-19-scenarioD-sec4-full-log.txt`. | CC |
| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor |
| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor |
| **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor |
@@ -0,0 +1,44 @@
# Golden bake — controller 0.222.0 (2026-08-23)
**Baked and PUBLISHED by Claude Code. NOT vouched — vouching is the operator's act.**
| Field | Value |
|---|---|
| `GOLDEN_VERSION` | `0.222.0` |
| `GOLDEN_SHA256` | `19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037` |
| Controller image | `gitea.dooplex.hu/admin/felhom-controller:0.222.0` |
| MinAgent | `0.129.0` (unchanged) |
| Package URL | `https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.222.0/golden.tar.zst` |
| Archive size | 655,194,776 B (624 MB) |
| LXC template | `debian-13-standard_13.6-1_amd64.tar.zst` |
## Acceptance markers (RUNBOOK-manual-build.md §4.1), each counted from `bake.log`
| Marker | Required | Observed |
|---|---|---|
| `docker OK (overlay2` | ≥1 | **1** — `docker OK (overlay2; data-root /var/lib/docker)` |
| `including mount point` (rootfs + mp0) | 2 | **2** |
| `upload OK (HTTP 201)` | 1 | **1** |
| `excluding` | 0 | **0** |
| `FATAL` | 0 | **0** |
Round-trip check on the published package: `HTTP 206` on a ranged GET, so the bytes are fetchable at
the URL Day-0 will use.
## One deviation from the runbook, recorded
§4.1 step 2 says to list the current Debian template because "the exact point release rots". On the
`virgin` snapshot the **`pveam` index is itself stale**: `pveam available` offered
`debian-13-standard_13.1-2_amd64.tar.zst`, and downloading it failed with
`400 Parameter verification failed. template: no such template`. **`pveam update` first**, then the
list reads `13.6-1` and the download succeeds. The runbook does not say to run `pveam update`; that
is the step that was missing, and it presents as a confusing 400 rather than as a stale index.
## Vouching — the OPERATOR's step, not done here
Hub → Configuration → Day-0 artifacts:
- `golden_version` → `0.222.0`
- `golden_sha256` → `19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037`
- `min_agent` → `0.129.0` (unchanged)
- then, **last and in its own save**, the global controller floor → `0.222.0`.
@@ -0,0 +1,327 @@
[golden] build-golden.sh v3.0.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.222.0
[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) …
Logical volume "vm-9100-disk-0" created.
Logical volume pve/vm-9100-disk-0 changed.
Creating filesystem with 8388608 4k blocks and 2097152 inodes
Filesystem UUID: d6274bb1-fd60-4a03-89d3-b7abfd683aa9
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
4096000, 7962624
Logical volume "vm-9100-disk-1" created.
Logical volume pve/vm-9100-disk-1 changed.
Creating filesystem with 6291456 4k blocks and 1572864 inodes
Filesystem UUID: b2064916-d0a7-475a-859e-36f1f3158ab8
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst'
Total bytes read: 553512960 (528MiB, 121MiB/s)
Detected container architecture: amd64
Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ...
done: SHA256:hVvFtsj2mnuHVaTWL4MhGSxyILPGznY0wzDSi5+g20o root@felhom-golden
Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ...
done: SHA256:9Pzut6x5TUrA9DlDV86rMDQha+u9s1PqfTJgj0wWYnI root@felhom-golden
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
done: SHA256:AN0aOUwJ2hzq6EBxLXIucTm5CKGlhNxVyTG89dezqDE root@felhom-golden
[golden] starting + installing Docker (official repo, trixie channel) …
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = (unset),
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to the standard locale ("C").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = (unset),
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to the standard locale ("C").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …
[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds …
[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) …
Unable to find image 'hello-world:latest' locally
latest: Pulling from library/hello-world
4f55086f7dd0: Pulling fs layer
4f55086f7dd0: Verifying Checksum
4f55086f7dd0: Download complete
4f55086f7dd0: Pull complete
Digest: sha256:5dd0d3e6e255913fc30f90b9f2b1d359cc2cbdb48090cc4b65f1676e203243cc
Status: Downloaded newer image for hello-world:latest
docker OK (overlay2; data-root /var/lib/docker)
/var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4
/mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4
both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576
[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.222.0 (no registry cred at deploy) …
WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'.
Configure a credential helper to remove this warning. See
https://docs.docker.com/go/credential-store/
0.222.0: Pulling from admin/felhom-controller
039e6f9f9752: Pulling fs layer
0094c3ac0914: Pulling fs layer
deca1dac7403: Pulling fs layer
11c19a33d1b8: Pulling fs layer
5dd7add0958f: Pulling fs layer
521bc476eadb: Pulling fs layer
11c19a33d1b8: Waiting
5dd7add0958f: Waiting
521bc476eadb: Waiting
deca1dac7403: Verifying Checksum
deca1dac7403: Download complete
11c19a33d1b8: Verifying Checksum
11c19a33d1b8: Download complete
5dd7add0958f: Verifying Checksum
5dd7add0958f: Download complete
521bc476eadb: Verifying Checksum
521bc476eadb: Download complete
0094c3ac0914: Verifying Checksum
0094c3ac0914: Download complete
039e6f9f9752: Verifying Checksum
039e6f9f9752: Download complete
039e6f9f9752: Pull complete
0094c3ac0914: Pull complete
deca1dac7403: Pull complete
11c19a33d1b8: Pull complete
5dd7add0958f: Pull complete
521bc476eadb: Pull complete
Digest: sha256:07320cd3ac46abd9e8404eb6a06b667930f3062e3a4e0814e1f9e9d33c227b73
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.222.0
gitea.dooplex.hu/admin/felhom-controller:0.222.0
[golden] asking the controller which infra images it manages …
[golden] baking infra images (4): traefik:v3.6.7 cloudflare/cloudflared:2026.6.0 gtstef/filebrowser:1.3.3-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 …
v3.6.7: Pulling from library/traefik
589002ba0eae: Pulling fs layer
ef63511ea6cc: Pulling fs layer
0738e5cb835e: Pulling fs layer
3e6813f70c64: Pulling fs layer
3e6813f70c64: Waiting
589002ba0eae: Verifying Checksum
589002ba0eae: Download complete
ef63511ea6cc: Verifying Checksum
ef63511ea6cc: Download complete
3e6813f70c64: Verifying Checksum
3e6813f70c64: Download complete
0738e5cb835e: Verifying Checksum
0738e5cb835e: Download complete
589002ba0eae: Pull complete
ef63511ea6cc: Pull complete
0738e5cb835e: Pull complete
3e6813f70c64: Pull complete
Digest: sha256:a9890c898f379c1905ee5b28342f6b408dc863f08db2dab20e46c267d1ff463a
Status: Downloaded newer image for traefik:v3.6.7
docker.io/library/traefik:v3.6.7
2026.6.0: Pulling from cloudflare/cloudflared
47de5dd0b812: Pulling fs layer
c172f21841df: Pulling fs layer
99515e7b4d35: Pulling fs layer
99ba982a9142: Pulling fs layer
d6b1b89eccac: Pulling fs layer
2780920e5dbf: Pulling fs layer
7c12895b777b: Pulling fs layer
3214acf345c0: Pulling fs layer
52630fc75a18: Pulling fs layer
dd64bf2dd177: Pulling fs layer
b839dfae01f6: Pulling fs layer
ebddc55facdc: Pulling fs layer
bdfd7f7e5bf6: Pulling fs layer
2d4d7adf6272: Pulling fs layer
40008157d8d2: Pulling fs layer
bd8962e29291: Pulling fs layer
cac2ae0193cb: Pulling fs layer
74d1dac84ecc: Pulling fs layer
dd64bf2dd177: Waiting
b839dfae01f6: Waiting
ebddc55facdc: Waiting
bdfd7f7e5bf6: Waiting
2d4d7adf6272: Waiting
40008157d8d2: Waiting
bd8962e29291: Waiting
cac2ae0193cb: Waiting
74d1dac84ecc: Waiting
2780920e5dbf: Waiting
7c12895b777b: Waiting
3214acf345c0: Waiting
52630fc75a18: Waiting
99ba982a9142: Waiting
d6b1b89eccac: Waiting
47de5dd0b812: Download complete
c172f21841df: Verifying Checksum
c172f21841df: Download complete
99ba982a9142: Verifying Checksum
99ba982a9142: Download complete
d6b1b89eccac: Verifying Checksum
d6b1b89eccac: Download complete
99515e7b4d35: Verifying Checksum
99515e7b4d35: Download complete
47de5dd0b812: Pull complete
2780920e5dbf: Verifying Checksum
2780920e5dbf: Download complete
7c12895b777b: Verifying Checksum
7c12895b777b: Download complete
3214acf345c0: Verifying Checksum
3214acf345c0: Download complete
52630fc75a18: Verifying Checksum
52630fc75a18: Download complete
c172f21841df: Pull complete
dd64bf2dd177: Verifying Checksum
dd64bf2dd177: Download complete
b839dfae01f6: Download complete
ebddc55facdc: Verifying Checksum
ebddc55facdc: Download complete
bdfd7f7e5bf6: Verifying Checksum
bdfd7f7e5bf6: Download complete
2d4d7adf6272: Verifying Checksum
2d4d7adf6272: Download complete
bd8962e29291: Verifying Checksum
bd8962e29291: Download complete
cac2ae0193cb: Verifying Checksum
cac2ae0193cb: Download complete
40008157d8d2: Verifying Checksum
40008157d8d2: Download complete
99515e7b4d35: Pull complete
74d1dac84ecc: Verifying Checksum
74d1dac84ecc: Download complete
99ba982a9142: Pull complete
d6b1b89eccac: Pull complete
2780920e5dbf: Pull complete
7c12895b777b: Pull complete
3214acf345c0: Pull complete
52630fc75a18: Pull complete
dd64bf2dd177: Pull complete
b839dfae01f6: Pull complete
ebddc55facdc: Pull complete
bdfd7f7e5bf6: Pull complete
2d4d7adf6272: Pull complete
40008157d8d2: Pull complete
bd8962e29291: Pull complete
cac2ae0193cb: Pull complete
74d1dac84ecc: Pull complete
Digest: sha256:ba461b8aa9c042156dbd39c38657fe7431bafa063220eab8d5330a523863da9f
Status: Downloaded newer image for cloudflare/cloudflared:2026.6.0
docker.io/cloudflare/cloudflared:2026.6.0
1.3.3-stable: Pulling from gtstef/filebrowser
6a0ac1617861: Pulling fs layer
ef8806083e82: Pulling fs layer
b74107c861c7: Pulling fs layer
adc935def003: Pulling fs layer
4f4fb700ef54: Pulling fs layer
18695ccc900a: Pulling fs layer
45d119d5c397: Pulling fs layer
dac52db4fc51: Pulling fs layer
6d598f86b2f2: Pulling fs layer
8aa349c8396c: Pulling fs layer
dac52db4fc51: Waiting
6d598f86b2f2: Waiting
8aa349c8396c: Waiting
adc935def003: Waiting
4f4fb700ef54: Waiting
18695ccc900a: Waiting
45d119d5c397: Waiting
b74107c861c7: Verifying Checksum
b74107c861c7: Download complete
6a0ac1617861: Verifying Checksum
6a0ac1617861: Download complete
ef8806083e82: Verifying Checksum
ef8806083e82: Download complete
adc935def003: Verifying Checksum
adc935def003: Download complete
4f4fb700ef54: Verifying Checksum
4f4fb700ef54: Download complete
45d119d5c397: Verifying Checksum
45d119d5c397: Download complete
dac52db4fc51: Verifying Checksum
dac52db4fc51: Download complete
6d598f86b2f2: Verifying Checksum
6d598f86b2f2: Download complete
18695ccc900a: Verifying Checksum
18695ccc900a: Download complete
6a0ac1617861: Pull complete
8aa349c8396c: Verifying Checksum
8aa349c8396c: Download complete
ef8806083e82: Pull complete
b74107c861c7: Pull complete
adc935def003: Pull complete
4f4fb700ef54: Pull complete
18695ccc900a: Pull complete
45d119d5c397: Pull complete
dac52db4fc51: Pull complete
6d598f86b2f2: Pull complete
8aa349c8396c: Pull complete
Digest: sha256:eb3733681db8757412632c61a99ad656f0d94ed6781bb2ea114b4d70babab78c
Status: Downloaded newer image for gtstef/filebrowser:1.3.3-stable
docker.io/gtstef/filebrowser:1.3.3-stable
1.1.0: Pulling from admin/felhom-samba
897d797d2723: Pulling fs layer
3051591aa250: Pulling fs layer
ce57a3f93416: Pulling fs layer
fb94eeec2fe1: Pulling fs layer
fb94eeec2fe1: Waiting
ce57a3f93416: Verifying Checksum
ce57a3f93416: Download complete
fb94eeec2fe1: Verifying Checksum
fb94eeec2fe1: Download complete
897d797d2723: Verifying Checksum
897d797d2723: Download complete
3051591aa250: Verifying Checksum
3051591aa250: Download complete
897d797d2723: Pull complete
3051591aa250: Pull complete
ce57a3f93416: Pull complete
fb94eeec2fe1: Pull complete
Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0
gitea.dooplex.hu/admin/felhom-samba:1.1.0
[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'.
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'.
[golden] baking the first-boot SSH host-key regeneration unit (F3) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'.
[golden] identity-clean + minimize …
[golden] stop + archive …
INFO: including mount point rootfs ('/') in backup
INFO: including mount point mp0 ('/var/lib/felhom') in backup
INFO: archive file size: 624MB
INFO: Finished Backup of VM 9100 (00:00:40)
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_08_23-07_32_56.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive)
[golden] publishing golden (655194776 bytes, sha256 19f5904f53792684…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.222.0/golden.tar.zst
[golden] pre-delete existing: HTTP 404 (404/204 expected)
[golden] upload OK (HTTP 201)
GOLDEN_VERSION=0.222.0
GOLDEN_SHA256=19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037
[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.222.0 / 19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037
[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)
+60 -6
View File
@@ -50,6 +50,28 @@ recorded waiver in the register, never a habit of bypassing.
FAIL-CLOSED, BUT HONEST ABOUT NOT KNOWING. An absent controller clone, or a CHANGELOG whose top
header cannot be parsed, exits **2 (INCONCLUSIVE)** — never 0. The runner reports 2 distinctly for
exactly this reason: an undetermined result is not a pass, and it is not a conviction either.
── THE SECOND BLINDNESS, R-385 (added 2026-08-23) ────────────────────────────────────────────
Until this change the gate asked ONE question — *is the golden BEHIND the record?* — and so it could
only ever catch a forgotten bake. It said nothing when the golden was **AHEAD** of the record, and
that is not a harmless direction: a golden ahead of every CHANGELOG heading was built from something
**never written down**.
That is not a hypothetical. Controller **0.221.1** was built, baked AND vouched on 2026-08-23 while
the newest heading in the controller CHANGELOG still read v0.221.0 — the fix had been written inside
the v0.221.0 entry instead of getting its own. Every gate was green throughout, including this one,
measured: `newest released 0.221.0 / newest golden baked 0.221.1 → OK`. The fleet ran a version the
record did not name.
So the test is no longer "behind?" but "**is the version we are shipping WRITTEN DOWN?**". The gate
now looks for the baked version's own `## vX.Y.Z` heading anywhere in the CHANGELOG — not merely at
the top, because an entry may legitimately be overtaken by later ones; what may never happen is that
it is absent. An unrecorded golden is convicted (exit 1) exactly like a stale one.
**Why membership and not `baked > released`.** A comparison against the newest heading alone would go
green again the moment ANY later entry was written, leaving 0.221.1 permanently unrecorded and the
gate permanently silent about it. Membership cannot be satisfied by an unrelated later release.
"""
import os
import re
@@ -67,16 +89,30 @@ RELEASED_RE = re.compile(r"^##\s+v(\d+)\.(\d+)\.(\d+)\b")
EVIDENCE_RE = re.compile(r"^golden-(\d+)\.(\d+)\.(\d+)-\d{4}-\d{2}-\d{2}$")
def newest_released():
"""(tuple, str) of the newest controller release, or (None, reason)."""
def released_versions():
"""(newest_tuple, note, set_of_all_tuples) of the controller releases, or (None, reason, set()).
R-385: the whole set is returned, not only the newest. The newest answers "is the golden behind?";
membership answers "is the version we are shipping written down at all?" — and only the second
question could have caught 0.221.1, whose heading was missing while a NEWER heading existed.
"""
if not os.path.isfile(CONTROLLER_CHANGELOG):
return None, "controller clone not found at %s" % CONTROLLER_CHANGELOG
return None, "controller clone not found at %s" % CONTROLLER_CHANGELOG, set()
newest = None
note = ""
every = set()
with open(CONTROLLER_CHANGELOG, encoding="utf-8") as fh:
for line in fh:
m = RELEASED_RE.match(line)
if m:
return tuple(int(g) for g in m.groups()), line.strip()[:90]
return None, "no '## vX.Y.Z' header found in %s" % CONTROLLER_CHANGELOG
v = tuple(int(g) for g in m.groups())
every.add(v)
if newest is None:
# newest-first by convention: the FIRST heading is the newest release.
newest, note = v, line.strip()[:90]
if newest is None:
return None, "no '## vX.Y.Z' header found in %s" % CONTROLLER_CHANGELOG, set()
return newest, note, every
def newest_baked():
@@ -99,7 +135,7 @@ def vstr(v):
def main():
released, rel_note = newest_released()
released, rel_note, every_released = released_versions()
if released is None:
print("GOLDEN CURRENCY GATE INCONCLUSIVE: %s" % rel_note)
sys.exit(2)
@@ -113,6 +149,24 @@ def main():
print(" newest released controller : %s (%s)" % (vstr(released), rel_note))
print(" newest golden baked : %s (documentation/tests/%s)" % (vstr(baked), bake_note))
# R-385 — UNRECORDED, checked before "behind". A golden whose version has no heading of its own
# was built from something never written down, and that is a different (worse) fault than a
# forgotten bake: there is nothing to read to find out what the fleet is running.
if baked not in every_released:
print("")
print("GOLDEN CURRENCY GATE FAILED: golden %s is baked but UNRECORDED — the controller "
"CHANGELOG has no '## v%s' heading." % (vstr(baked), vstr(baked)))
print("The newest heading is %s. A golden ahead of the record was built from a version "
"nobody wrote down, so no one can read what the fleet is running." % vstr(released))
print("Fix: give v%s its own '## v%s — <what changed>' heading in "
"felhom-controller/CHANGELOG.md, above the entries it supersedes. If its fix is "
"currently described inside another version's entry, MOVE that text — do not "
"duplicate it, and do not delete the reasoning." % (vstr(baked), vstr(baked)))
print("If this bake was a throwaway that must never be delivered, delete its "
"documentation/tests/golden-<VER>-<DATE>/ directory — never leave it to read as "
"shipped.")
sys.exit(1)
if released > baked:
print("")
print("GOLDEN CURRENCY GATE FAILED: controller v%s is released and NO golden carries it "