R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
The gate failed only on `released > baked`, so it could catch a forgotten bake and nothing else. A golden AHEAD of the record passed silently - and that is how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG heading still read v0.221.0, with every gate green. Reproduced on the real history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0. The gate now asks whether the version being shipped is WRITTEN DOWN: the baked version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG. Membership rather than `baked > released` deliberately - a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2) preserved; every refusal names a reason and a route. Red-proofed both directions: old gate/old record exit 0, new gate/old record exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0. 08-alarm-ladder.md is new, and its absence was itself the finding: no document owned "when does a broken app raise an alarm?". The rules lived as comments in four packages, each locally correct, with the ordering between them legible only by reading one function top to bottom - which is how R-384 survived review. R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed closed. R-386 filed OPEN: a single-container app stopped out of band raises no alarm, and a comment claims the opposite - measured live, 9 scans, 0 events, against a positive control from the same box 17 minutes earlier. Not fixed here. Golden 0.222.0 baked and published; vouching is the operator's act.
This commit is contained in:
@@ -1,62 +1,97 @@
|
||||
# REPORT — R-361 and the two loose ends v0.220.2 left (2026-08-22 → 23)
|
||||
# REPORT — felhom.eu: the golden-currency gate could not see an unrecorded golden (R-385)
|
||||
|
||||
Companion to `felhom-controller` **v0.221.0 → v0.221.1**. Full record:
|
||||
`documentation/audits/DRILL-r361-2026-08-22/`.
|
||||
**Session 2026-08-23.** Companion to `felhom-controller` v0.222.0 (R-384, R-383) — see that repo's
|
||||
`REPORT.md` for the controller work and the full live walk.
|
||||
|
||||
## What this repo carried
|
||||
## What was wrong here
|
||||
|
||||
- **`documentation/architecture/07-backup-architecture.md`** — a dated **[FACT]** on R-361 (the
|
||||
comment that asserted an invariant the code did not have, and what it cost), and a **[DESIGN]** on
|
||||
the `db_dumps` decision **including the trap it created**: a stable list lets
|
||||
`CaptureRecoveryUnit`'s already-current early return fire, so anything that must happen on every
|
||||
capture has to sit above that check.
|
||||
- **`documentation/architecture/00-capability-map.md`** — the **negative** from Part 3, recorded so it
|
||||
is not re-derived: a HELD app does **not** raise the dead-app alarm, measured on the shipped build,
|
||||
and the reading that said it would was wrong and why.
|
||||
- **`STATUS.md`** — the outcome in plain words; the deciding section says what happens if nothing is
|
||||
done.
|
||||
- **`documentation/tests/golden-0.221.1-2026-08-23/`** — the golden bake.
|
||||
- **Register** — R-361 closed and compressed; **R-383** and **R-384** opened.
|
||||
`scripts/golden_currency_gate.py` asked ONE question — *is the golden BEHIND the record?* — and
|
||||
therefore could only ever catch a forgotten bake. **It said nothing when the golden was AHEAD of the
|
||||
record**, and that direction is not harmless: a golden ahead of every CHANGELOG heading was built
|
||||
from something never written down.
|
||||
|
||||
## Part 3 — a measurement that cancelled a Part, and that is a good outcome
|
||||
That is not hypothetical. Controller **0.221.1** was built, baked **and vouched** on 2026-08-23 while
|
||||
the newest heading in the controller CHANGELOG still read `v0.221.0`. Measured on the real history,
|
||||
with the old gate:
|
||||
|
||||
The runbook's reading was that a held app alarms as a dead app. **It does not.** On the shipped
|
||||
v0.220.2 a hold was created deliberately; `docmost` aggregated to `unhealthy`; `IsDownState` is
|
||||
`{stopped, exited, degraded}`; the dead-app heartbeat read **`0 currently down`** at scans 600 and
|
||||
620 with the scans demonstrably running over it. **Part 2 was dropped in full** and **no register row
|
||||
was opened**, exactly as the runbook directs.
|
||||
```
|
||||
newest released controller : 0.221.0 (## v0.221.0 — taking the undo copy destroyed …)
|
||||
newest golden baked : 0.221.1 (documentation/tests/golden-0.221.1-2026-08-23)
|
||||
golden currency gate OK …
|
||||
EXIT=0
|
||||
```
|
||||
|
||||
**The positive control took three attempts, and that is the second finding.** Two live attempts
|
||||
failed to produce a lasting down state at all — `privatebin` went `stopped` (whitelisted by design)
|
||||
and `bookstack` went `degraded` then `unhealthy`. An absent alarm from a detector never shown working
|
||||
proves nothing, so the control was moved to the layer the detector lives in: `classifyRunStates` is a
|
||||
pure function, and it raises the banner for `degraded`/`exited` while staying silent for the states
|
||||
measured live.
|
||||
Every gate was green while the fleet ran a version the record did not name.
|
||||
|
||||
## Findings opened
|
||||
## The fix, and why it is membership and not a comparison
|
||||
|
||||
- **R-383 (MEDIUM)** — the double-failure message tells the customer *"a korábbi állapot mentése
|
||||
megvan"* while naming the very file whose absence caused the failure. Observed on **both** v0.220.2
|
||||
and v0.221.1. **R-361's own class** — a sentence asserting a property the code does not check.
|
||||
- **R-384 (MEDIUM)** — an app whose **database** has died reads `unhealthy` and raises no dead-app
|
||||
banner and no customer e-mail, because `aggregateState` checks `unhealthy > 0` before the
|
||||
mixed-case degraded branch. Not invisible everywhere (the health report counts it), but it does not
|
||||
alarm.
|
||||
The gate now asks **"is the version we are shipping WRITTEN DOWN?"** — the baked version must have its
|
||||
own `## vX.Y.Z` heading **anywhere** in the controller CHANGELOG, not merely at the top (an entry may
|
||||
legitimately be overtaken by later ones; what may never happen is that it is absent).
|
||||
|
||||
**Register size: `OPEN-ITEMS.md` 325 236 → 327 266 bytes; `CLOSED-ITEMS.md` 66 777 → 68 464.**
|
||||
R-361's full text: `git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md`.
|
||||
**Membership, not `baked > released`, deliberately:** a comparison against the newest heading alone
|
||||
goes green the moment ANY later entry is written — which would have left 0.221.1 permanently
|
||||
unrecorded and the gate permanently silent about it.
|
||||
|
||||
## The golden was baked here, and why
|
||||
Preserved unchanged: **INCONCLUSIVE (exit 2)** for an absent clone or an unparseable CHANGELOG — *not
|
||||
knowing is never a pass, and never a conviction*. Every refusal names a reason **and** a route,
|
||||
including what to do if a bake was a throwaway that must never be delivered.
|
||||
|
||||
`golden-currency` refused the docs push: 0.221.1 released with no golden. **Not circular** — a golden
|
||||
needs the controller image, already pushed, not this commit — so the gate was satisfied by doing the
|
||||
work rather than bypassed. **No push in this session used `--no-verify`.**
|
||||
## Red-proofs — both directions, against the real history
|
||||
|
||||
Golden **0.221.1**, sha256 `1c8bf6cf08cadabeca6331f38360d905e867c235067cd10c2716915b6e6df089`,
|
||||
656 966 079 B, round-trip verified, all markers hit, both negative controls at zero, both token-leak
|
||||
greps proved able to convict before their zeros were accepted.
|
||||
| Run | Gate | CHANGELOG | Golden baked | Exit | |
|
||||
|---|---|---|---|---|---|
|
||||
| `gate-01` | **old** | v0.221.0 | 0.221.1 | **0** | the blindness, reproduced |
|
||||
| `gate-02` | **new** | v0.221.0 | 0.221.1 | **1** | convicted |
|
||||
| `gate-03` | new | v0.221.1 | 0.221.1 | **0** | Part 0's heading makes it pass |
|
||||
| `gate-04` | new | absent clone | — | **2** | INCONCLUSIVE preserved |
|
||||
| `gate-05` | new | v0.222.0 | 0.222.0 | **0** | post-bake |
|
||||
|
||||
## Operator follow-up
|
||||
Transcripts: `documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt`.
|
||||
|
||||
**Vouch** the golden — a **three-field** save: `golden_version` **0.221.1**, `agent_version`
|
||||
**0.130.0**, `min_agent` **0.129.0**. **Then** raise the floor to **0.221.1**, last, in its own save.
|
||||
## Files changed
|
||||
|
||||
| File | Change |
|
||||
|---|---|
|
||||
| `scripts/golden_currency_gate.py` | the unrecorded-golden conviction; `newest_released` → `released_versions` returning the whole set; docstring records the second blindness |
|
||||
| `documentation/architecture/08-alarm-ladder.md` | **NEW.** The alarm ladder as a dated [DESIGN] |
|
||||
| `documentation/architecture/00-capability-map.md` | R-384 marked closed with its live evidence; points at the new doc |
|
||||
| `documentation/backlog/OPEN-ITEMS.md` | R-383/R-384 removed (closed); **R-385** (closed) and **R-386** (open) filed |
|
||||
| `documentation/backlog/CLOSED-ITEMS.md` | R-383 + R-384 compressed, each keeping its rules and naming `git show 1eb64bec5183:…` for the full text |
|
||||
| `documentation/tests/golden-0.222.0-2026-08-23/` | **NEW.** Bake evidence + log + the vouching instructions |
|
||||
| `STATUS.md` | the 0.222.0 vouch replaces the (now completed) 0.221.1 one; R-386 added in plain words |
|
||||
| `documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/` | **NEW.** The drill record and 29 evidence files |
|
||||
|
||||
## The alarm ladder had no owning document — that absence is a finding
|
||||
|
||||
Nothing in `documentation/architecture/` owned the question *"when does a customer's app being broken
|
||||
raise an alarm?"* The rules lived as comments across four packages, each locally correct, with the
|
||||
ordering between them legible only by reading `aggregateState` top to bottom. **That is precisely how
|
||||
R-384 survived review**, and three separate defects in this ladder (R-51, C9-F2, R-384) were each
|
||||
found on live hardware rather than by reading. `08-alarm-ladder.md` now owns it.
|
||||
|
||||
## Golden
|
||||
|
||||
**Baked and PUBLISHED: 0.222.0.** `GOLDEN_SHA256 = 19f5904f5379…`, `upload OK (HTTP 201)`, round-trip
|
||||
`HTTP 206` from the package URL, all acceptance markers counted.
|
||||
**VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
|
||||
|
||||
**Deviation recorded:** the bake runbook's §4.1 is missing a `pveam update`. On the `virgin` snapshot
|
||||
the template index is stale, so the listed template cannot be downloaded and the failure presents as
|
||||
`400 Parameter verification failed. template: no such template` rather than as a stale index.
|
||||
|
||||
## Register size
|
||||
|
||||
| File | Before | After |
|
||||
|---|---|---|
|
||||
| `OPEN-ITEMS.md` | 327,266 B | **328,325 B** |
|
||||
| `CLOSED-ITEMS.md` | 68,464 B | **71,441 B** |
|
||||
|
||||
OPEN grew ~1 KB despite two closures, because R-386 is a substantial new finding. Recorded rather
|
||||
than smoothed over.
|
||||
|
||||
## Hub numbers as read at session start (live, `GET /configuration`)
|
||||
|
||||
`golden_version` **0.221.1** · `agent_version` **0.130.0** · `min_agent` **0.129.0** ·
|
||||
controller floor **0.221.1**. The task expected 0.220.2/0.220.2; the operator had already vouched.
|
||||
**The hub was READ ONLY this session** — nothing was written to it.
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-08-22 — the off-site restore now works for all 53 apps, not 13. It is released and
|
||||
NOT yet delivered: two steps below are yours.**
|
||||
**Updated 2026-08-23 — an app whose database dies now raises an alarm. It did not before, and the
|
||||
watcher said "nothing is down" the whole time. Released and NOT yet delivered: step 1 is yours.**
|
||||
|
||||
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
||||
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
||||
@@ -12,14 +12,26 @@ NOT yet delivered: two steps below are yours.**
|
||||
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
||||
nothing.*
|
||||
|
||||
1. **Vouch the golden carrying controller 0.221.1** — Hub → Configuration → Day-0 artifacts.
|
||||
1. **Vouch the golden carrying controller 0.222.0** — Hub → Configuration → Day-0 artifacts.
|
||||
**It is already baked, published and round-trip verified**
|
||||
(`documentation/tests/golden-0.221.1-2026-08-23/`); only the vouch is left, and only you can do it.
|
||||
**It is a THREE-field save:** `golden_version` → **0.221.1**, `agent_version` → **0.130.0**,
|
||||
`min_agent` → **0.129.0**. **Then** raise the floor to **0.221.1**, last, in its own save.
|
||||
**If you do nothing:** the fleet stays on 0.219.0, so a failed database restore still leaves an app
|
||||
broken with an unusable copy — the thing today's release fixes reaches nobody. New machines still
|
||||
receive 0.219.0. The build system stays red about it and will mail you on every push.
|
||||
(`documentation/tests/golden-0.222.0-2026-08-23/`); only the vouch is left, and only you can do it.
|
||||
**It is a THREE-field save:** `golden_version` → **0.222.0**, `agent_version` → **0.130.0**,
|
||||
`min_agent` → **0.129.0**. **Then** raise the floor to **0.222.0**, last, in its own save.
|
||||
**If you do nothing:** the fleet stays on 0.221.1, where an app whose database has died reports
|
||||
nothing at all — no banner, no event — and the watcher keeps printing "0 currently down". New
|
||||
machines still receive 0.221.1.
|
||||
*(Thank you — the 0.221.1 vouch from earlier today has landed; the hub reads golden 0.221.1 and
|
||||
floor 0.221.1. Nothing is owed on that one.)*
|
||||
|
||||
2. **An app that is stopped from outside still reports nothing** (R-386) — and this one I found today
|
||||
and deliberately did **not** fix. If a single-container app is stopped by hand on the machine
|
||||
rather than through the product, nothing is said: no banner, no e-mail, no operator event. I
|
||||
measured it: nine checks ran over four minutes and every one stayed silent. A comment in our own
|
||||
code claims the opposite, which is why nobody noticed. **The reason I stopped rather than fixed it:**
|
||||
from the outside this looks exactly like a customer pressing Stop, and the obvious fix would start
|
||||
alarming every time somebody legitimately stops their own app. That trade is a decision, not a
|
||||
patch. **If you do nothing:** it stays as it is — this is not a new fault, it has always been so;
|
||||
it is newly *known*.
|
||||
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||
|
||||
@@ -67,6 +67,23 @@ the banner, pinned by `TestClassifyRunStates_PositiveControl_ADownStackDoesAlarm
|
||||
measurement DID expose is **R-384**: an app whose database has died is `unhealthy` too, and is
|
||||
likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt`.
|
||||
|
||||
> **R-384 CLOSED in controller v0.222.0 (2026-08-23), proven live.** The defect was the ORDER of two
|
||||
> questions, not the `unhealthy` exclusion: `aggregateState` now asks *"is a supervised member dead?"*
|
||||
> **before** the `unhealthy`/`starting`/`restarting` returns, and "some members are up" counts any
|
||||
> member not in the down bucket rather than `running` alone. `IsDownState` is byte-identical.
|
||||
> **Measured on `demo-hp` 2026-08-23** with the same fixture that read `0 currently down` the day
|
||||
> before: `bookstack-db` stopped 05:30:07Z → `app_start_failed` fired at **05:30:14Z**, the banner read
|
||||
> *„Telepített alkalmazás nem fut: BookStack (degraded)"*, the stack read `state=degraded` **while its
|
||||
> front end was `unhealthy`**, and the heartbeat printed **`1 currently down`** against the previous
|
||||
> day's `0`. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`.
|
||||
>
|
||||
> **The HELD-app half of the paragraph above is now also covered** — a held app keeps its database
|
||||
> container, so it is the same shape and reaches the same `degraded` verdict.
|
||||
>
|
||||
> **The alarm ladder that decides all of this now has an owning document:** see
|
||||
> `08-alarm-ladder.md` (written 2026-08-23 — before that date no document owned it, and that absence
|
||||
> is why the ordering defect was legible only from source).
|
||||
|
||||
**WHAT IS STILL NOT CLAIMED:** the FAILURE path is where this class is weak, not the success path — a corrupt dump leaves the customer with an emptied or partially-applied database and an undo copy **no product action can apply** (**R-379**), and on MariaDB it does so behind an app that reports `health=healthy` (**R-380**). The success story is proven; the recovery-from-a-bad-restore story is not.
|
||||
|
||||
**NARROWED 2026-08-21 by the backup-truth drill — kept as history; both halves have since closed, see the entry immediately above.** Proven that night on `demo-hp` with planted, hash-recorded files: the **declared-userdata** leg of a **drive-declaring** app does come back byte-identical (`calibre-web`, 5/5 including two Hungarian accented filenames). **Two legs of the same story do NOT:** (a) the off-site restore has **no named-volume leg at all**, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (**R-354**); (b) for the **40 of 53** apps that declare no data drive the off-site restore **refuses outright**, saying a running app „nincs telepítve" (**R-356**). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is **unproven for that class and disproven for the volume leg generally**. The escrow/key half of this row is untouched by that and still stands. Evidence: `audits/DRILL-backup-truth-2026-08-21/evidence/` and `REPORT.md` (2026-08-21). |
|
||||
|
||||
@@ -0,0 +1,141 @@
|
||||
# 08 — The app-down alarm ladder
|
||||
|
||||
**Written 2026-08-23, with controller v0.222.0 (R-384).**
|
||||
|
||||
**The absence is the finding.** Until this file existed, no document owned the question *"when does a
|
||||
customer's app being broken raise an alarm?"* The rules were spread across four packages as comments,
|
||||
each locally correct, and the ordering between them was legible only by reading
|
||||
`aggregateState` top to bottom. That is exactly how R-384 survived: every individual rule was right,
|
||||
and the composition was wrong. Three separate defects in this ladder (R-51, C9-F2, R-384) were each
|
||||
found on live hardware rather than by review, and each is a case where a reader could not see the
|
||||
whole ladder at once.
|
||||
|
||||
Everything below is **[DESIGN]** — deliberate, with the reason recorded — unless marked otherwise.
|
||||
|
||||
---
|
||||
|
||||
## 1. The two questions, and their order
|
||||
|
||||
Two different questions get asked about a multi-container app, and **the order between them is
|
||||
load-bearing**:
|
||||
|
||||
1. **Is a SUPERVISED member of this app dead?** — a container Docker's restart policy says should be
|
||||
running, that is not.
|
||||
2. **Is a RUNNING member failing its healthcheck?**
|
||||
|
||||
**Question 1 is asked FIRST.** [DESIGN, R-384, v0.222.0]
|
||||
|
||||
Until v0.222.0 it was asked second, and the consequence was not subtle: a two-container app whose
|
||||
database exits goes `unhealthy` seconds later *because it cannot reach that database*. So the symptom
|
||||
the dead database causes was the thing that suppressed the alarm for it. Measured on `demo-hp`
|
||||
2026-08-22 — `bookstack-db` stopped at 21:27:01 and the watcher reported `0 currently down`
|
||||
throughout.
|
||||
|
||||
**"Some members are up" means any member NOT in the down bucket** — `running`, `unhealthy`,
|
||||
`starting` or `restarting`. [DESIGN, R-384] The earlier guard was `running > 0`, counting only
|
||||
`StateRunning`, which made the supervised test unreachable in precisely the case it was written for:
|
||||
an unhealthy survivor beside a dead database counted as nothing being up.
|
||||
|
||||
---
|
||||
|
||||
## 2. Where each decision is made
|
||||
|
||||
| Decision | Where | Notes |
|
||||
|---|---|---|
|
||||
| container → stack aggregate state | `internal/stacks/manager.go` `aggregateState` | the ladder in §3 |
|
||||
| is a down member supervised? | `internal/stacks/manager.go` `supervisedPolicy` | `no`/`on-failure` benign; everything else, **including unknown**, supervised |
|
||||
| which states mean "down" | `internal/stacks/manager.go` `IsDownState` | `{stopped, exited, degraded}` |
|
||||
| stack state → "this app is down" | `cmd/controller/main.go` `classifyRunStates` | **the single derivation point**; all three suppressions live here |
|
||||
| sustained restarting → down | `internal/stacks/manager.go` `CrashLooping` | 5-minute threshold |
|
||||
| quiesce suppression | `internal/quiesce/suppress.go` | cycle-keyed, 180 s grace |
|
||||
| boot repair | `internal/bootrecon/bootrecon.go` | consumes `IsDownState` |
|
||||
|
||||
---
|
||||
|
||||
## 3. The aggregation ladder, in order
|
||||
|
||||
`aggregateState(containers, policyOf)` — **priority: degraded > unhealthy/starting > restarting >
|
||||
all-running > stopped.**
|
||||
|
||||
1. no containers → `not_deployed`
|
||||
2. **any DOWN member is supervised, and any member is up → `degraded`** ← R-384 put this first
|
||||
3. any `unhealthy` → `unhealthy`
|
||||
4. any `starting` → `starting`
|
||||
5. any `restarting` → `restarting`
|
||||
6. all running → `running`
|
||||
7. all down → `stopped`
|
||||
8. mix, every down member benign → `running`
|
||||
|
||||
Step 2's `policyOf` is consulted **only** for the down members, and only when something is up. `nil`
|
||||
is allowed; every down member then reads as supervised.
|
||||
|
||||
---
|
||||
|
||||
## 4. Which states alarm, and which deliberately do not
|
||||
|
||||
`IsDownState` = `{stopped, exited, degraded}`.
|
||||
|
||||
| State | Down? | Why |
|
||||
|---|---|---|
|
||||
| `stopped`, `exited` | **yes** | not running, will not recover alone |
|
||||
| `degraded` | **yes** | [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert |
|
||||
| `unhealthy` | **NO** | [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. **R-384 did not change this** — it asks a prior question instead |
|
||||
| `restarting` | **NO**, until sustained | [DESIGN, C9-F2] `restarting` is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after **5 min** (`crashLoopAfter`) |
|
||||
| `starting`, `deploying` | no | mid-start |
|
||||
| `paused` | no | a deliberate user action |
|
||||
| `unknown` | no | [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read |
|
||||
|
||||
**Note the two fail directions are deliberately opposite.** `IsDownState` fails OPEN on `unknown`
|
||||
(ambiguous *state* → do not alarm). `supervisedPolicy` fails CLOSED on unknown (we already KNOW a
|
||||
member is dead; only the excuse is missing). Both are recorded at their sites.
|
||||
|
||||
---
|
||||
|
||||
## 5. The three suppressions, all at `classifyRunStates`
|
||||
|
||||
| Suppression | Rule | Expires? |
|
||||
|---|---|---|
|
||||
| **deliberate user stop** | `StateStopped` is not down **unless** the quiesce loop reports it failed to restart that stack | n/a — lifted by `failedRestart` |
|
||||
| **quiesce cycle** | a stack this backup cycle stopped is exempt | **yes**, 180 s after unquiescing |
|
||||
| **boot grace** | no evaluation for 90 s after controller start | **yes** |
|
||||
|
||||
**None of them latch.** [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false
|
||||
alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite
|
||||
direction. Every window expires; the cost is a bounded DELAY in reporting a real failure, never its
|
||||
loss.
|
||||
|
||||
The quiesce suppression is **cycle-keyed, not state-keyed** — an app caught mid-restart is
|
||||
`starting`/`unhealthy`, not `stopped`, so no state test can see it. It is therefore **state-blind**,
|
||||
which is why R-384 moving a stack from `unhealthy` to `degraded` cannot weaken it.
|
||||
|
||||
---
|
||||
|
||||
## 6. The alarm itself
|
||||
|
||||
Edge-triggered: `app_start_failed`, one event per transition into down, **not** per scan. Verified
|
||||
live 2026-08-23 — one event across 22 scans.
|
||||
|
||||
- **Operator/hub event + dashboard banner.** `app_start_failed` is **not** in
|
||||
`settings.DefaultEnabledEvents`, so by default it does **not** e-mail the customer.
|
||||
- **F-OBS heartbeat**, every 20 scans (~10 min), at `[INFO]`:
|
||||
`[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down`.
|
||||
This line exists because an absent alarm and a stopped detector look identical in a log.
|
||||
|
||||
> ⚠ **R-329, OPEN and it bites here.** `app_start_failed` is pushed with severity **`warn`**, which is
|
||||
> **not** in the hub's vocabulary (`{info, warning, error, critical}`) and is silently coerced to
|
||||
> `info` — which e-mails nobody, while the POST still returns 200. Observed again on 2026-08-23:
|
||||
> `PushEvent: type=app_start_failed severity=warn`. R-384 makes this event actually fire, so the
|
||||
> severity bug now matters more than it did while the event was unreachable.
|
||||
|
||||
---
|
||||
|
||||
## 7. Known gap, filed not fixed
|
||||
|
||||
> **R-386 (filed 2026-08-23, OPEN).** An all-down stack aggregates to `stopped` — `StateExited` is
|
||||
> folded into the same counter and never survives aggregation. `classifyRunStates` then whitelists
|
||||
> `stopped` as a deliberate user stop. So a **single-container app stopped out of band raises no
|
||||
> alarm at all**, which directly contradicts the comment at `cmd/controller/main.go`: *"An out-of-band
|
||||
> `docker compose stop` leaves the containers present → StateExited → still alerts."*
|
||||
> **Measured on `demo-hp` 2026-08-23:** `privatebin` stopped out of band, 9 dead-app scans over 4+
|
||||
> minutes, `state=stopped`, **zero events and zero banner lines** — against a positive control from
|
||||
> the same box 17 minutes earlier. Not fixed in v0.222.0 deliberately; it is a separate decision.
|
||||
@@ -0,0 +1,96 @@
|
||||
# DRILL — R-384: an app whose database dies raised no alarm (2026-08-23)
|
||||
|
||||
**Controller v0.221.1 → v0.222.0. Live leg on `demo-hp` (Tier 0, disposable), guest 9201.**
|
||||
**UNATTENDED.** Method: endpoint-level — no browser exists on DooPlex, so every read is either the
|
||||
exact endpoint the UI calls or the controller's own log. Guest clock is UTC.
|
||||
|
||||
## Verdict
|
||||
|
||||
| Part | Outcome |
|
||||
|---|---|
|
||||
| Part 0 — the record | ✅ `v0.221.1` given its own heading, pushed ALONE (`da75603`) |
|
||||
| Part 1 — the blind gate | ✅ fixed; both directions red-proofed against the real history |
|
||||
| Part 2 — R-384 | ✅ shipped v0.222.0, **proven live** |
|
||||
| Part 3 — R-383 | ✅ shipped v0.222.0 |
|
||||
| §4 — the measurement | ⚠ **reproduced. Filed as R-386. NOT fixed — that was the instruction.** |
|
||||
| Live walk step 4 (Scenario E) | **DROPPED** — consequence of the halt; drop-list item (3) |
|
||||
|
||||
## The one-line result
|
||||
|
||||
The same fixture that printed **`0 currently down`** on 2026-08-22 printed **`1 currently down`** on
|
||||
2026-08-23, with 8 apps evaluated both times:
|
||||
|
||||
```
|
||||
2026/08/22 21:13:48 [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down
|
||||
2026/08/23 05:37:44 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down
|
||||
```
|
||||
|
||||
## What was actually wrong
|
||||
|
||||
**The ORDER of two questions**, not the `unhealthy` exclusion. "Is a supervised member dead?" and "is
|
||||
a running member failing its healthcheck?" are different questions, and the second was answering the
|
||||
first — because a dying database drags its own front end `unhealthy`, **the symptom the fault causes
|
||||
was what suppressed the alarm for it.** `IsDownState` was not touched; no state was minted.
|
||||
|
||||
The fix has **two halves and either alone leaves the defect standing**: the hoist, and widening "some
|
||||
members are up" from `running > 0` to *any member not in the down bucket*. The old guard made the
|
||||
R-51 block unreachable in precisely the case R-51 was written for.
|
||||
|
||||
## Evidence index (`evidence/`)
|
||||
|
||||
| File | What it shows |
|
||||
|---|---|
|
||||
| `gate-01-old-gate-old-changelog.txt` | the blindness: old gate, real history, **exit 0** |
|
||||
| `gate-02-new-gate-old-changelog.txt` | new gate on the same history, **exit 1**, naming the fix |
|
||||
| `gate-03-new-gate-new-changelog.txt` | with Part 0's heading, **exit 0** |
|
||||
| `gate-04-inconclusive.txt` | INCONCLUSIVE (**exit 2**) preserved |
|
||||
| `gate-05-new-gate-post-bake.txt` | v0.222.0 + golden 0.222.0, **exit 0** |
|
||||
| `redproof-R384-1-order.txt` | mutation: hoist reverted → `"unhealthy", want "degraded"` |
|
||||
| `redproof-R384-2-upguard.txt` | mutation: `up` narrowed → all three survivor shapes convict |
|
||||
| `redproof-R384-3-classifier.txt` | mutation: classifier ignores `degraded` → empty banner |
|
||||
| `redproof-R383-undo-phrase.txt` | mutation: unconditional claim → prints the false sentence |
|
||||
| `live-00-the-old-heartbeat-2026-08-22.txt` | yesterday's `0 currently down`, carried in for contrast |
|
||||
| `live-01`…`live-06` | Scenario A: pre-state, stop, log, decisive read, heartbeat, banner |
|
||||
| `live-08-golden-bake-markers.txt` | the bake's acceptance markers, each counted |
|
||||
| `live-09-…-full-controller-log.txt` | 1808 lines, pulled off **before** the app was restarted |
|
||||
| `live-10`…`live-13` | Scenario B: unhealthy with nothing dead, no alarm, banner cleared |
|
||||
| `live-14`…`live-16` | Scenario D: full stop→start cycle, **0 alarms across 9 scans** |
|
||||
| `live-17`, `live-18` | §4: `privatebin` out-of-band stop, 9 scans, **0 events, 0 banner** |
|
||||
| `live-19-…-full-log.txt` | 2139 lines covering Scenario D and §4 |
|
||||
|
||||
## §4 — what it found, in plain words
|
||||
|
||||
`aggregateState` folds `StateExited` into the `stopped` counter, so an all-down stack returns
|
||||
`StateStopped` and **`StateExited` never survives aggregation** — that is the path the task suspected
|
||||
and could not find in source. `classifyRunStates` then whitelists `StateStopped` as a deliberate user
|
||||
stop. So this comment in `cmd/controller/main.go` is **false**:
|
||||
|
||||
> *"(An out-of-band `docker compose stop` leaves the containers present → StateExited → still alerts,
|
||||
> which is correct: out-of-band tampering IS reportable.)"*
|
||||
|
||||
Measured: `privatebin` (1 container, `unless-stopped`) stopped 05:47:35Z; at 05:51:53Z it read
|
||||
`state=stopped` with 9 dead-app scans behind it, **zero events and zero banner lines**.
|
||||
|
||||
**The absence is trustworthy because the detector was shown alive first** (standing rule 3):
|
||||
`app_start_failed` fired for BookStack at 05:30:14Z on the same box 17 minutes earlier.
|
||||
|
||||
**Scoped honestly:** a genuine crash under `unless-stopped` is restarted by Docker and surfaces as
|
||||
`restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit
|
||||
out-of-band stop of a stack with no surviving member.
|
||||
|
||||
## Two things noticed that are NOT this drill's work
|
||||
|
||||
1. **R-329 moved from unreachable to load-bearing.** `app_start_failed` ships severity `warn`, which
|
||||
is not in the hub's vocabulary and coerces silently to `info` — e-mailing nobody, POST still 200.
|
||||
Observed again today. While the event never fired, this was harmless; it no longer is.
|
||||
2. **The golden-bake runbook is missing `pveam update`.** On the `virgin` snapshot the template index
|
||||
is stale, so `pveam available` offers `13.1-2` and downloading it fails with
|
||||
`400 Parameter verification failed. template: no such template` — a confusing 400 rather than a
|
||||
legible "your index is old". Recorded in `documentation/tests/golden-0.222.0-2026-08-23/README.md`.
|
||||
|
||||
## Teardown
|
||||
|
||||
Nothing was provisioned. All apps restored and confirmed healthy (`bookstack` + `bookstack-db`,
|
||||
`docmost` ×3, `privatebin`); planted data untouched; no app rebuilt or restored. **Hub-side: nothing
|
||||
to discard — the hub was READ ONLY this session** (`GET /configuration`, `GET /events`); no appliance
|
||||
registered, no config written, no artifact manifest changed.
|
||||
+5
@@ -0,0 +1,5 @@
|
||||
### OLD gate + OLD changelog (v0.221.0 top, golden 0.221.1 baked) ###
|
||||
newest released controller : 0.221.0 (## v0.221.0 — taking the undo copy destroyed the app's own database backup (2026-08-22, R-)
|
||||
newest golden baked : 0.221.1 (documentation/tests/golden-0.221.1-2026-08-23)
|
||||
golden currency gate OK — the newest released controller has a golden (NOTE: this checks the BAKE, not the vouch — see the module docstring)
|
||||
EXIT=0
|
||||
+9
@@ -0,0 +1,9 @@
|
||||
### NEW gate + OLD changelog (v0.221.0 top, golden 0.221.1 baked) -> must FAIL ###
|
||||
newest released controller : 0.221.0 (## v0.221.0 — taking the undo copy destroyed the app's own database backup (2026-08-22, R-)
|
||||
newest golden baked : 0.221.1 (documentation/tests/golden-0.221.1-2026-08-23)
|
||||
|
||||
GOLDEN CURRENCY GATE FAILED: golden 0.221.1 is baked but UNRECORDED — the controller CHANGELOG has no '## v0.221.1' heading.
|
||||
The newest heading is 0.221.0. A golden ahead of the record was built from a version nobody wrote down, so no one can read what the fleet is running.
|
||||
Fix: give v0.221.1 its own '## v0.221.1 — <what changed>' heading in felhom-controller/CHANGELOG.md, above the entries it supersedes. If its fix is currently described inside another version's entry, MOVE that text — do not duplicate it, and do not delete the reasoning.
|
||||
If this bake was a throwaway that must never be delivered, delete its documentation/tests/golden-<VER>-<DATE>/ directory — never leave it to read as shipped.
|
||||
EXIT=1
|
||||
+5
@@ -0,0 +1,5 @@
|
||||
### NEW gate + FIXED changelog (v0.221.1 heading present) -> must PASS ###
|
||||
newest released controller : 0.221.1 (## v0.221.1 — the undo-copy prune stopped running because another fix made its guard reach)
|
||||
newest golden baked : 0.221.1 (documentation/tests/golden-0.221.1-2026-08-23)
|
||||
golden currency gate OK — the newest released controller has a golden (NOTE: this checks the BAKE, not the vouch — see the module docstring)
|
||||
EXIT=0
|
||||
+3
@@ -0,0 +1,3 @@
|
||||
### INCONCLUSIVE preserved: unreadable CHANGELOG path ###
|
||||
GOLDEN CURRENCY GATE INCONCLUSIVE: controller clone not found at /nonexistent/CHANGELOG.md
|
||||
EXIT=2
|
||||
+5
@@ -0,0 +1,5 @@
|
||||
### NEW gate, post-bake: CHANGELOG v0.222.0 + golden 0.222.0 -> must PASS ###
|
||||
newest released controller : 0.222.0 (## v0.222.0 — an app whose database dies raised no alarm, because the wrong question answe)
|
||||
newest golden baked : 0.222.0 (documentation/tests/golden-0.222.0-2026-08-23)
|
||||
golden currency gate OK — the newest released controller has a golden (NOTE: this checks the BAKE, not the vouch — see the module docstring)
|
||||
EXIT=0
|
||||
+18
@@ -0,0 +1,18 @@
|
||||
=== PART 3 MEASUREMENT — v0.220.2, no code change
|
||||
measured at: 21:13:51Z (hold created 21:11:19Z)
|
||||
elapsed: ~2m20s = 5 deadapp scans at 30s cadence (heartbeat lines confirm 9 scans in the log window)
|
||||
|
||||
-- 3a. WHICH RUN STATE did docmost aggregate to?
|
||||
bookstack state=running deployed=True
|
||||
docmost state=unhealthy deployed=True
|
||||
privatebin state=running deployed=True
|
||||
|
||||
-- 3b. IS THE DATABASE STILL UP?
|
||||
docmost-postgres | Up 2 minutes (healthy)
|
||||
|
||||
-- 3c. DID A CUSTOMER-FACING EVENT FIRE for docmost?
|
||||
2026/08/22 21:11:19 notifier.go:234: [INFO] Event pushed: backup_run_failures (error) — App "docmost" is HELD STOPPED: its off-site database restore failed AND the rollback to the customer's own pre-restore copy also failed. The app will not start from any path until the hold is cleared. Replay error: importing postgres dump for docmost: postgres import into docmost-postgres failed: exit status 3. Rollback error: a visszavonáshoz szükséges mentés nem található (pre-restore-20260822T211114Z-docmost-postgres.sql): stat /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T211114Z-docmost-postgres.sql: no such file or directory
|
||||
(none above = no event fired)
|
||||
|
||||
-- 3d. DEAD-APP BANNER STATE:
|
||||
2026/08/22 21:13:48 main.go:1730: [INFO] [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down
|
||||
+11
@@ -0,0 +1,11 @@
|
||||
=== SCENARIO A — a database dies behind a healthy-looking app (v0.222.0) ===
|
||||
controller version:
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.222.0 Up 3 minutes (healthy)
|
||||
|
||||
-- PRE-STATE (guest UTC) --
|
||||
2026-08-23T05:30:02Z
|
||||
bookstack Up 3 hours (healthy)
|
||||
bookstack-db Up 3 hours (healthy)
|
||||
|
||||
/bookstack restart=unless-stopped
|
||||
/bookstack-db restart=unless-stopped
|
||||
+6
@@ -0,0 +1,6 @@
|
||||
-- STOPPING bookstack-db OUT OF BAND --
|
||||
2026-08-23T05:30:07Z
|
||||
bookstack-db
|
||||
2026-08-23T05:30:07Z
|
||||
bookstack Up 3 hours (healthy)
|
||||
bookstack-db Exited (0) Less than a second ago
|
||||
+18
@@ -0,0 +1,18 @@
|
||||
-- controller log since the stop (05:30:00Z) --
|
||||
2026/08/23 05:30:04 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m10s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:30:14 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
|
||||
2026/08/23 05:30:14 manager.go:703: [DEBUG] [stacks] restart-policy of down member "bookstack-db" = "unless-stopped"
|
||||
2026/08/23 05:30:14 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m20s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:30:14 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warn url=https://hub.felhom.eu/api/v1/event
|
||||
2026/08/23 05:30:14 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
|
||||
2026/08/23 05:30:14 notifier.go:234: [INFO] Event pushed: app_start_failed (warn) — Telepített alkalmazás nem fut: BookStack
|
||||
2026/08/23 05:30:24 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m30s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:30:34 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m40s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:30:44 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
|
||||
2026/08/23 05:30:44 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m50s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:30:44 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "bookstack" deployed=true composePath=/opt/docker/stacks/bookstack/docker-compose.yml
|
||||
2026/08/23 05:30:54 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 4m0s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:31:04 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 4m10s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:31:14 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 4m20s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:31:14 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
|
||||
2026/08/23 05:31:24 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 4m30s ago, effective interval 5m0s, healthy=true
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
-- DECISIVE READ: front end UNHEALTHY, database EXITED --
|
||||
2026-08-23T05:32:03Z
|
||||
bookstack Up 3 hours (unhealthy)
|
||||
bookstack-db Exited (0) About a minute ago
|
||||
|
||||
bookstack state=degraded deployed=True
|
||||
calibre-web state=running deployed=True
|
||||
docmost state=running deployed=True
|
||||
kimai state=running deployed=True
|
||||
opengist state=running deployed=True
|
||||
paperless-ngx state=running deployed=True
|
||||
privatebin state=running deployed=True
|
||||
romm state=running deployed=True
|
||||
+2
@@ -0,0 +1,2 @@
|
||||
-- waiting for the F-OBS heartbeat (every 20 scans = ~10 min) --
|
||||
2026/08/23 05:37:44 main.go:1730: [INFO] [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down
|
||||
+9
@@ -0,0 +1,9 @@
|
||||
-- the alert surface: /api/alerts + the launcher page --
|
||||
### /api/alerts
|
||||
http=404 bytes=42
|
||||
### /launcher
|
||||
http=200 bytes=42201
|
||||
<span class="alert-message">Telepített alkalmazás nem fut: BookStack (degraded)</span>
|
||||
### /dashboard
|
||||
http=200 bytes=56437
|
||||
<span class="alert-message">Telepített alkalmazás nem fut: BookStack (degraded)</span>
|
||||
+2
@@ -0,0 +1,2 @@
|
||||
-- hub-side: the event as the hub STORED it (R-329 severity check) --
|
||||
http=404
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
=== GOLDEN 0.222.0 BAKE — acceptance markers ===
|
||||
docker OK (overlay2 : 1
|
||||
including mount point: 2
|
||||
upload OK (HTTP 201) : 1
|
||||
excluding (must be 0): 0
|
||||
FATAL (must be 0): 0
|
||||
--- the marker lines ---
|
||||
docker OK (overlay2; data-root /var/lib/docker)
|
||||
INFO: including mount point rootfs ('/') in backup
|
||||
INFO: including mount point mp0 ('/var/lib/felhom') in backup
|
||||
[golden] upload OK (HTTP 201)
|
||||
GOLDEN_VERSION=0.222.0
|
||||
GOLDEN_SHA256=19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037
|
||||
+1808
File diff suppressed because it is too large
Load Diff
+7
@@ -0,0 +1,7 @@
|
||||
=== SCENARIO B — unhealthy with NOTHING dead (the flapping case) ===
|
||||
-- restarting bookstack-db; the front end stays unhealthy for a while with nothing down --
|
||||
2026-08-23T05:40:01Z
|
||||
bookstack-db
|
||||
2026-08-23T05:40:09Z
|
||||
bookstack Up 3 hours (unhealthy)
|
||||
bookstack-db Up 8 seconds (healthy)
|
||||
+6
@@ -0,0 +1,6 @@
|
||||
-- SCENARIO B decisive read: front end unhealthy, database UP, nothing dead --
|
||||
2026-08-23T05:40:20Z
|
||||
bookstack Up 3 hours (healthy)
|
||||
bookstack-db Up 18 seconds (healthy)
|
||||
|
||||
bookstack state=unhealthy
|
||||
+7
@@ -0,0 +1,7 @@
|
||||
-- SCENARIO B: any NEW alarm during the recovery window? --
|
||||
app_start_failed events since 05:40:00Z: 0
|
||||
--- what the aggregate read during the window ---
|
||||
2026/08/23 05:40:04 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m0s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:40:14 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m10s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:40:24 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m20s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:40:34 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m30s ago, effective interval 5m0s, healthy=true
|
||||
+6
@@ -0,0 +1,6 @@
|
||||
-- banner cleared + bookstack healthy again --
|
||||
2026-08-23T05:40:53Z
|
||||
bookstack Up 3 hours (healthy)
|
||||
bookstack-db Up 52 seconds (healthy)
|
||||
banner lines: 0
|
||||
bookstack state=running
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
=== SCENARIO D — a full stop -> start cycle through the PRODUCTION endpoint ===
|
||||
method: POST /api/stacks/docmost/{stop,start} — the exact call the launcher's buttons make
|
||||
subject: docmost (3 containers: docmost, docmost-postgres, docmost-redis)
|
||||
csrf len=64
|
||||
BASELINE alarms before: 0
|
||||
T0=2026-08-23T05:42:15Z --- STOP ---
|
||||
{"ok":true,"message":"Stack docmost stop completed"}
|
||||
http=200
|
||||
|
||||
after stop: 2026-08-23T05:42:28Z
|
||||
--- START ---
|
||||
{"ok":true,"message":"Stack docmost start completed"}
|
||||
http=200
|
||||
+11
@@ -0,0 +1,11 @@
|
||||
-- SCENARIO D: watching a FULL cycle settle (5 min past the start) --
|
||||
05:42:47 docmost:Up 7 seconds (health: starting) docmost-postgres:Up 18 seconds (healthy) docmost-redis:Up 18 seconds (healthy)
|
||||
05:43:12 docmost:Up 33 seconds (healthy) docmost-postgres:Up 43 seconds (healthy) docmost-redis:Up 43 seconds (healthy)
|
||||
05:43:37 docmost:Up 58 seconds (healthy) docmost-postgres:Up About a minute (healthy) docmost-redis:Up About a minute (healthy)
|
||||
05:44:02 docmost:Up About a minute (healthy) docmost-postgres:Up About a minute (healthy) docmost-redis:Up About a minute (healthy)
|
||||
05:44:27 docmost:Up About a minute (healthy) docmost-postgres:Up About a minute (healthy) docmost-redis:Up About a minute (healthy)
|
||||
05:44:52 docmost:Up 2 minutes (healthy) docmost-postgres:Up 2 minutes (healthy) docmost-redis:Up 2 minutes (healthy)
|
||||
05:45:17 docmost:Up 2 minutes (healthy) docmost-postgres:Up 2 minutes (healthy) docmost-redis:Up 2 minutes (healthy)
|
||||
05:45:42 docmost:Up 3 minutes (healthy) docmost-postgres:Up 3 minutes (healthy) docmost-redis:Up 3 minutes (healthy)
|
||||
05:46:07 docmost:Up 3 minutes (healthy) docmost-postgres:Up 3 minutes (healthy) docmost-redis:Up 3 minutes (healthy)
|
||||
05:46:32 docmost:Up 3 minutes (healthy) docmost-postgres:Up 4 minutes (healthy) docmost-redis:Up 4 minutes (healthy)
|
||||
+8
@@ -0,0 +1,8 @@
|
||||
-- SCENARIO D VERDICT: alarms across the whole cycle (T0=05:42:15Z) --
|
||||
app_start_failed events since T0 : 0
|
||||
deadapp scans in the window : 9
|
||||
supervised-down path entered : 1
|
||||
|
||||
--- any docmost down-state reading? ---
|
||||
2026/08/23 05:42:34 manager.go:703: [DEBUG] [stacks] restart-policy of down member "docmost" = "unless-stopped"
|
||||
(no lines above = none)
|
||||
+11
@@ -0,0 +1,11 @@
|
||||
=== §4 MEASUREMENT — a SINGLE-container app crashes out of band ===
|
||||
POSITIVE CONTROL, established on this box today: app_start_failed fired at 05:30:14Z for
|
||||
BookStack (see live-09). The detector demonstrably works here and now, so an absence below
|
||||
is a real absence, not a dead detector.
|
||||
|
||||
subject: privatebin (1 container, restart policy below). Stopped OUT OF BAND — the controller
|
||||
did not do it, so no Deploying flag and no quiesce key.
|
||||
privatebin restart=unless-stopped
|
||||
T0=2026-08-23T05:47:35Z
|
||||
2026-08-23T05:47:40Z
|
||||
privatebin Exited (0) 5 seconds ago
|
||||
+12
@@ -0,0 +1,12 @@
|
||||
-- §4: waiting 4 minutes, well past every grace window (deadapp scan 30s, quiesce grace 180s) --
|
||||
2026-08-23T05:51:53Z
|
||||
privatebin Exited (0) 4 minutes ago
|
||||
|
||||
-- the aggregate state the controller reads --
|
||||
privatebin state=stopped deployed=True
|
||||
|
||||
-- DID ANY ALARM FIRE? --
|
||||
app_start_failed since T0 : 0
|
||||
deadapp scans since T0 : 9
|
||||
-- banner? --
|
||||
banner lines: 0
|
||||
+2139
File diff suppressed because it is too large
Load Diff
+18
@@ -0,0 +1,18 @@
|
||||
### RED-PROOF R-383 — mutation: undoCopyPhrase reverted to the unconditional pre-fix claim ###
|
||||
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist (0.00s)
|
||||
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist/absent_—_must_NOT_claim_it_exists,_must_still_name_where_it_should_be (0.00s)
|
||||
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-mariadb.sql" does not contain "NEM találjuk"
|
||||
r383_undo_phrase_test.go:97: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-mariadb.sql" contains "mentése megvan" — it asserts a file that is not on disk
|
||||
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist/zero-length_—_counts_as_missing (0.00s)
|
||||
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-empty.sql" does not contain "NEM találjuk"
|
||||
r383_undo_phrase_test.go:97: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-empty.sql" contains "mentése megvan" — it asserts a file that is not on disk
|
||||
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist/partial_—_both_halves_named,_neither_hidden (0.00s)
|
||||
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-postgres.sql" does not contain "RÉSZBEN"
|
||||
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-postgres.sql" does not contain "HIÁNYZIK"
|
||||
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-postgres.sql" does not contain "pre-restore-20260823T120000Z-app-mariadb.sql"
|
||||
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist/no_undo_was_ever_written_—_said_plainly,_not_silently (0.00s)
|
||||
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: ." does not contain "nem készült"
|
||||
r383_undo_phrase_test.go:97: phrase "a korábbi állapot mentése megvan: ." contains "megvan" — it asserts a file that is not on disk
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.008s
|
||||
FAIL
|
||||
+9
@@ -0,0 +1,9 @@
|
||||
### RED-PROOF R-384 #1 — the ORDER (mutation: hoist moved back below `unhealthy > 0`) ###
|
||||
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst (0.00s)
|
||||
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst/unhealthy_survivor_—_the_bookstack_case_measured_live (0.00s)
|
||||
degraded_test.go:165: survivor "unhealthy" beside a dead SUPERVISED member: aggregateState = "unhealthy", want "degraded" — a dead database must not hide behind it
|
||||
--- FAIL: TestR384_WiresTheDeadDatabaseThroughTheRealPath (0.00s)
|
||||
degraded_test.go:458: bookstack state = "unhealthy", want "degraded" — a dead database must not hide behind its own unhealthy front end (measured live 2026-08-22: this read "unhealthy" and nothing alarmed)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
|
||||
FAIL
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
### RED-PROOF R-384 #2 — the GUARD (mutation: up = running only, the old `running > 0`) ###
|
||||
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst (0.00s)
|
||||
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst/unhealthy_survivor_—_the_bookstack_case_measured_live (0.00s)
|
||||
degraded_test.go:165: survivor "unhealthy" beside a dead SUPERVISED member: aggregateState = "unhealthy", want "degraded" — a dead database must not hide behind it
|
||||
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst/starting_survivor (0.00s)
|
||||
degraded_test.go:165: survivor "starting" beside a dead SUPERVISED member: aggregateState = "starting", want "degraded" — a dead database must not hide behind it
|
||||
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst/restarting_survivor (0.00s)
|
||||
degraded_test.go:165: survivor "restarting" beside a dead SUPERVISED member: aggregateState = "restarting", want "degraded" — a dead database must not hide behind it
|
||||
--- FAIL: TestR384_WiresTheDeadDatabaseThroughTheRealPath (0.00s)
|
||||
degraded_test.go:458: bookstack state = "unhealthy", want "degraded" — a dead database must not hide behind its own unhealthy front end (measured live 2026-08-22: this read "unhealthy" and nothing alarmed)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
|
||||
FAIL
|
||||
+6
@@ -0,0 +1,6 @@
|
||||
### RED-PROOF R-384 #3 — the CONSEQUENCE layer (mutation: classifier ignores StateDegraded) ###
|
||||
--- FAIL: TestR384_ADeadDatabaseBehindAnUnhealthyAppAlarms (0.00s)
|
||||
r384_dead_db_alarm_test.go:29: dead-app banner = [], want exactly one entry for bookstack
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.008s
|
||||
FAIL
|
||||
@@ -160,3 +160,5 @@
|
||||
| **R-370** | **PROCESS: between 2026-08-19 and 2026-08-22 the reviewing side called a documented architectural decision a defect, in four places, because it read the register and live source and never `documentation/architecture/`.** Evidence: `documentation/architecture/`. **Reasoning kept:** R-352 (re-framed), R-369 The record is corrected in place with the framing marked rather than deleted, per the standing rule that a document which quietly changes its mind teaches nobody. | **CLOSED — corrected 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-96** | **Two standing rules were agreed in chat and never committed** Evidence: `documentation/runbooks/workspace-CLAUDE.md:48-70`. **Reasoning kept:** **Two standing rules were agreed in chat and never committed** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-27, size XS, roadmap state `idea — found 2026-07-27`.** Moved verbatim; nothing added or reinterpreted. | **CLOSED — migrated from ROADMAP 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-107** | **No offsite action unpacks the named-volume tars Tier-3 captures on every run.** Shipped in v0.218.0. | **CLOSED — migrated from ROADMAP 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-383** | **The double-failure message told the customer their previous state was saved, and named a file that was not there.** Shipped in controller v0.222.0. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`. **Reasoning kept:** *One of the two ways a rollback fails is that the undo copy is missing — so the sentence was most likely to be false in exactly the case it was printed.* **Do NOT simply drop the filename:** an operator needs it, and R-351's lesson is that a refusal naming nothing forces someone to remember what the product already knows — so the absent case still names WHERE the file should have been. **A zero-length dump counts as MISSING**, because a 0-byte file restores nothing and calling it present is the same false reassurance one step smaller. The check is `os.Stat` and deliberately not an integrity test: this runs at the end of a failed restore on a machine that may be unwell, and presence is the honest claim available there. | **CLOSED — SHIPPED** (controller v0.222.0, 2026-08-23; `undoCopyPhrase`, four cases, plus an AST seam test that the message is still wired to the builder) | full text: `git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-384** | **An app whose DATABASE had died raised no dead-app alarm — the wrong question answered first.** Shipped in controller v0.222.0. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`. **Reasoning kept:** *The defect was the ORDER of two questions, not the `unhealthy` exclusion.* "Is a SUPERVISED member dead?" and "is a RUNNING member failing its healthcheck?" are different questions, and the second was answering the first — a dying database drags its own front end `unhealthy`, so the symptom the fault causes was what suppressed the alarm for it. **`IsDownState` is byte-identical and `unhealthy` stays excluded** — an unhealthy container is RUNNING, and folding it in reintroduces the flapping that exclusion exists to stop; **no new state was minted**, `StateDegraded` already means this. **Two things had to move and either alone leaves the defect standing:** the hoist, AND widening "some members are up" from `running > 0` to *any member not in the down bucket* — the old guard made the R-51 block unreachable in precisely the case it was written for. **The register's own suggested fix was WRONG and is recorded as such:** it proposed a sustained-`unhealthy` threshold on the `crashLoopAfter` model; the actual defect needed no threshold at all. **PROVEN LIVE the only way it can be** — the same fixture that printed `0 currently down` on 2026-08-22 printed **`1 currently down`** on 2026-08-23, with `app_start_failed` 7 s after the stop and the banner reading *„…nem fut: BookStack (degraded)"*. Scenario D measured **0 alarms across 9 scans** through a full stop→start cycle. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.222.0, 2026-08-23) | full text: `git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
@@ -136,8 +136,8 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
|
||||
|
||||
| ID | What | State |
|
||||
|---|---|---|
|
||||
| **R-383** | **The double-failure message tells the customer their previous state was saved, and names a file that is not there.** The sentence ends *"a korábbi állapot mentése megvan: <file>"* — "the backup of the previous state EXISTS" — built from the path `writeSafetyDump` returned, WITHOUT asking whether it is still on disk. But one of the two ways a rollback can fail is that the undo copy is missing or unreadable, and in exactly that case the sentence is FALSE. **Measured live twice, on v0.220.2 (2026-08-22 21:11:19) and again on v0.221.1 (22:20:32):** rollback failed with `stat …pre-restore-…sql: no such file or directory`, and the customer message named that same file as existing. 369 bytes, `offbox_reconstitute.go` (the double-failure branch). **This is R-361's own class** — a sentence asserting a property the code does not check — one surface over. | **OPEN — MEDIUM** | — | Say what is true: name the undo copy only when it is verifiably on disk, and say plainly when it is not. **Do not simply drop the filename** — an operator needs it, and R-351's lesson was that a refusal which names nothing forces someone to remember what the product already knew. Evidence: `audits/DRILL-r361-2026-08-22/evidence/03-observed-false-sentence.txt`, `audits/DRILL-r361-2026-08-22/evidence/16-part4-message.txt`. | CC |
|
||||
| **R-384** | **An app whose DATABASE has died raises no dead-app alarm — `unhealthy` masks the mixed state.** `aggregateState` (`internal/stacks/manager.go`) checks `if unhealthy > 0 → StateUnhealthy` BEFORE the mixed-case degraded branch, and `IsDownState` (`manager.go:54`) is `{stopped, exited, degraded}` — `unhealthy` is absent. So a multi-container app whose database container dies goes `degraded` for a moment and then `unhealthy` as its own healthcheck fails, and stops being a fault. **Measured live 2026-08-22:** `bookstack-db` stopped out-of-band at 21:27:01; `bookstack` read `unhealthy`; the dead-app heartbeat reported **`0 currently down`** across the whole window (scans 600 and 620), with 8 apps evaluated. **NOT invisible everywhere** — the health report counts it (`cr.Unhealthy++`, `internal/report/builder.go:251`) and that reaches the hub — but it raises no banner and no customer e-mail. **This is the F-CRIT-1 class the `classifyRunStates` comment says was closed:** it WAS closed for `StateStopped`+failedRestart, and `unhealthy` was never in scope. Found while building a positive control for a different question. | **OPEN — MEDIUM** | — | Decide whether a SUSTAINED `unhealthy` is a fault (it is not a brief one — that is why it is excluded), on the `crashLoopAfter` model: a threshold above the deploy/health windows rather than a state test. **Do not simply add `unhealthy` to `IsDownState`** — it has other callers and would alarm on every deploy, which is the over-correction F-A1 nearly cost. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt`. | CC |
|
||||
| **R-385** | **A controller was built, baked AND vouched with no CHANGELOG entry of its own, and every gate stayed green.** Controller **0.221.1** shipped on 2026-08-23 while the newest heading in `felhom-controller/CHANGELOG.md` still read `v0.221.0` — the prune-ordering fix (commit `810b18a`) had been written INSIDE the v0.221.0 entry instead of getting its own. The image was never in question; the RECORD was, and the fleet ran a version the record did not name. **`scripts/golden_currency_gate.py` could not catch it by construction:** it failed only on `released > baked`, so a golden AHEAD of the record passed silently. Measured on the real history: `newest released 0.221.0 / newest golden baked 0.221.1 → OK, exit 0`. | **CLOSED — 2026-08-23** | — | **Both halves fixed, both directions red-proofed.** The record: `v0.221.1` has its own heading carrying the MOVED (not duplicated, not deleted) reasoning — commit `da75603`, pushed alone before anything else. The gate now asks *"is the baked version WRITTEN DOWN?"* — the baked version must have its own `## vX.Y.Z` heading **anywhere** in the CHANGELOG. **Membership, not `baked > released`, deliberately:** a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded and the gate permanently silent about it. INCONCLUSIVE (exit 2) preserved. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt` — old gate/old record `exit 0`, new gate/old record `exit 1`, new gate/fixed record `exit 0`. | CC |
|
||||
| **R-386** | **A single-container app stopped OUT OF BAND raises no alarm at all — and a comment states the opposite as settled fact.** `aggregateState` folds `StateExited` into the `stopped` counter, so an all-down stack returns `StateStopped` and **`StateExited` never survives aggregation**. `classifyRunStates` then whitelists `StateStopped` as a deliberate user stop unless the quiesce loop reports a failed restart. The comment at `cmd/controller/main.go` says: *"(An out-of-band `docker compose stop` leaves the containers present → StateExited → still alerts, which is correct: out-of-band tampering IS reportable.)"* — **measured FALSE.** The neighbouring I2 claim (*"a CRASHING app never comes to rest at `stopped` — faults surface as StateExited"*) is false in the same way. **Measured live on `demo-hp` 2026-08-23 (controller v0.222.0):** `privatebin` (1 container, `unless-stopped`) stopped out of band at 05:47:35Z; at 05:51:53Z it read `state=stopped`, **9 dead-app scans had run, and there were ZERO `app_start_failed` events and ZERO banner lines.** **The absence is trustworthy — positive control from the same box 17 minutes earlier:** `app_start_failed` fired for BookStack at 05:30:14Z, so the detector demonstrably works there. **SCOPE, stated so it is not overclaimed:** a genuine crash under `unless-stopped` is RESTARTED by Docker and surfaces as `restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit out-of-band stop of a stack with no surviving member. **This is case #10 of "a comment asserting an invariant the code does not provide".** Found by §4 of the R-384 task, which asked for a measurement and explicitly forbade a fix in that session. | **OPEN — MEDIUM** | — | Decide whether an out-of-band stop is distinguishable from a customer stop at all — they are byte-identical on the Docker side, exactly as invariant I1 says, so the answer is probably NOT a state test but a recorded intent (`DesiredStateOf` already exists and `bootrecon` already consumes it). **Do NOT simply un-whitelist `StateStopped`** — that re-alarms every genuine customer stop, which is the over-correction F-CRIT-1's fix was careful to avoid. **And fix the comment either way:** it is load-bearing and it is false. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/live-17-sec4-stop.txt`, `live-18-sec4-verdict.txt`, `live-19-scenarioD-sec4-full-log.txt`. | CC |
|
||||
| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor |
|
||||
| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor |
|
||||
| **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor |
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
# Golden bake — controller 0.222.0 (2026-08-23)
|
||||
|
||||
**Baked and PUBLISHED by Claude Code. NOT vouched — vouching is the operator's act.**
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| `GOLDEN_VERSION` | `0.222.0` |
|
||||
| `GOLDEN_SHA256` | `19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037` |
|
||||
| Controller image | `gitea.dooplex.hu/admin/felhom-controller:0.222.0` |
|
||||
| MinAgent | `0.129.0` (unchanged) |
|
||||
| Package URL | `https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.222.0/golden.tar.zst` |
|
||||
| Archive size | 655,194,776 B (624 MB) |
|
||||
| LXC template | `debian-13-standard_13.6-1_amd64.tar.zst` |
|
||||
|
||||
## Acceptance markers (RUNBOOK-manual-build.md §4.1), each counted from `bake.log`
|
||||
|
||||
| Marker | Required | Observed |
|
||||
|---|---|---|
|
||||
| `docker OK (overlay2` | ≥1 | **1** — `docker OK (overlay2; data-root /var/lib/docker)` |
|
||||
| `including mount point` (rootfs + mp0) | 2 | **2** |
|
||||
| `upload OK (HTTP 201)` | 1 | **1** |
|
||||
| `excluding` | 0 | **0** |
|
||||
| `FATAL` | 0 | **0** |
|
||||
|
||||
Round-trip check on the published package: `HTTP 206` on a ranged GET, so the bytes are fetchable at
|
||||
the URL Day-0 will use.
|
||||
|
||||
## One deviation from the runbook, recorded
|
||||
|
||||
§4.1 step 2 says to list the current Debian template because "the exact point release rots". On the
|
||||
`virgin` snapshot the **`pveam` index is itself stale**: `pveam available` offered
|
||||
`debian-13-standard_13.1-2_amd64.tar.zst`, and downloading it failed with
|
||||
`400 Parameter verification failed. template: no such template`. **`pveam update` first**, then the
|
||||
list reads `13.6-1` and the download succeeds. The runbook does not say to run `pveam update`; that
|
||||
is the step that was missing, and it presents as a confusing 400 rather than as a stale index.
|
||||
|
||||
## Vouching — the OPERATOR's step, not done here
|
||||
|
||||
Hub → Configuration → Day-0 artifacts:
|
||||
|
||||
- `golden_version` → `0.222.0`
|
||||
- `golden_sha256` → `19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037`
|
||||
- `min_agent` → `0.129.0` (unchanged)
|
||||
- then, **last and in its own save**, the global controller floor → `0.222.0`.
|
||||
@@ -0,0 +1,327 @@
|
||||
[golden] build-golden.sh v3.0.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.222.0
|
||||
[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) …
|
||||
Logical volume "vm-9100-disk-0" created.
|
||||
Logical volume pve/vm-9100-disk-0 changed.
|
||||
Creating filesystem with 8388608 4k blocks and 2097152 inodes
|
||||
Filesystem UUID: d6274bb1-fd60-4a03-89d3-b7abfd683aa9
|
||||
Superblock backups stored on blocks:
|
||||
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||
4096000, 7962624
|
||||
Logical volume "vm-9100-disk-1" created.
|
||||
Logical volume pve/vm-9100-disk-1 changed.
|
||||
Creating filesystem with 6291456 4k blocks and 1572864 inodes
|
||||
Filesystem UUID: b2064916-d0a7-475a-859e-36f1f3158ab8
|
||||
Superblock backups stored on blocks:
|
||||
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||
extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst'
|
||||
Total bytes read: 553512960 (528MiB, 121MiB/s)
|
||||
Detected container architecture: amd64
|
||||
Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ...
|
||||
done: SHA256:hVvFtsj2mnuHVaTWL4MhGSxyILPGznY0wzDSi5+g20o root@felhom-golden
|
||||
Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ...
|
||||
done: SHA256:9Pzut6x5TUrA9DlDV86rMDQha+u9s1PqfTJgj0wWYnI root@felhom-golden
|
||||
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
|
||||
done: SHA256:AN0aOUwJ2hzq6EBxLXIucTm5CKGlhNxVyTG89dezqDE root@felhom-golden
|
||||
[golden] starting + installing Docker (official repo, trixie channel) …
|
||||
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = (unset),
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to the standard locale ("C").
|
||||
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
|
||||
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = (unset),
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to the standard locale ("C").
|
||||
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
|
||||
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||
[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …
|
||||
[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds …
|
||||
[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) …
|
||||
Unable to find image 'hello-world:latest' locally
|
||||
latest: Pulling from library/hello-world
|
||||
4f55086f7dd0: Pulling fs layer
|
||||
4f55086f7dd0: Verifying Checksum
|
||||
4f55086f7dd0: Download complete
|
||||
4f55086f7dd0: Pull complete
|
||||
Digest: sha256:5dd0d3e6e255913fc30f90b9f2b1d359cc2cbdb48090cc4b65f1676e203243cc
|
||||
Status: Downloaded newer image for hello-world:latest
|
||||
docker OK (overlay2; data-root /var/lib/docker)
|
||||
/var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4
|
||||
/mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4
|
||||
both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576
|
||||
[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.222.0 (no registry cred at deploy) …
|
||||
|
||||
WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'.
|
||||
Configure a credential helper to remove this warning. See
|
||||
https://docs.docker.com/go/credential-store/
|
||||
|
||||
0.222.0: Pulling from admin/felhom-controller
|
||||
039e6f9f9752: Pulling fs layer
|
||||
0094c3ac0914: Pulling fs layer
|
||||
deca1dac7403: Pulling fs layer
|
||||
11c19a33d1b8: Pulling fs layer
|
||||
5dd7add0958f: Pulling fs layer
|
||||
521bc476eadb: Pulling fs layer
|
||||
11c19a33d1b8: Waiting
|
||||
5dd7add0958f: Waiting
|
||||
521bc476eadb: Waiting
|
||||
deca1dac7403: Verifying Checksum
|
||||
deca1dac7403: Download complete
|
||||
11c19a33d1b8: Verifying Checksum
|
||||
11c19a33d1b8: Download complete
|
||||
5dd7add0958f: Verifying Checksum
|
||||
5dd7add0958f: Download complete
|
||||
521bc476eadb: Verifying Checksum
|
||||
521bc476eadb: Download complete
|
||||
0094c3ac0914: Verifying Checksum
|
||||
0094c3ac0914: Download complete
|
||||
039e6f9f9752: Verifying Checksum
|
||||
039e6f9f9752: Download complete
|
||||
039e6f9f9752: Pull complete
|
||||
0094c3ac0914: Pull complete
|
||||
deca1dac7403: Pull complete
|
||||
11c19a33d1b8: Pull complete
|
||||
5dd7add0958f: Pull complete
|
||||
521bc476eadb: Pull complete
|
||||
Digest: sha256:07320cd3ac46abd9e8404eb6a06b667930f3062e3a4e0814e1f9e9d33c227b73
|
||||
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.222.0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.222.0
|
||||
[golden] asking the controller which infra images it manages …
|
||||
[golden] baking infra images (4): traefik:v3.6.7 cloudflare/cloudflared:2026.6.0 gtstef/filebrowser:1.3.3-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 …
|
||||
v3.6.7: Pulling from library/traefik
|
||||
589002ba0eae: Pulling fs layer
|
||||
ef63511ea6cc: Pulling fs layer
|
||||
0738e5cb835e: Pulling fs layer
|
||||
3e6813f70c64: Pulling fs layer
|
||||
3e6813f70c64: Waiting
|
||||
589002ba0eae: Verifying Checksum
|
||||
589002ba0eae: Download complete
|
||||
ef63511ea6cc: Verifying Checksum
|
||||
ef63511ea6cc: Download complete
|
||||
3e6813f70c64: Verifying Checksum
|
||||
3e6813f70c64: Download complete
|
||||
0738e5cb835e: Verifying Checksum
|
||||
0738e5cb835e: Download complete
|
||||
589002ba0eae: Pull complete
|
||||
ef63511ea6cc: Pull complete
|
||||
0738e5cb835e: Pull complete
|
||||
3e6813f70c64: Pull complete
|
||||
Digest: sha256:a9890c898f379c1905ee5b28342f6b408dc863f08db2dab20e46c267d1ff463a
|
||||
Status: Downloaded newer image for traefik:v3.6.7
|
||||
docker.io/library/traefik:v3.6.7
|
||||
2026.6.0: Pulling from cloudflare/cloudflared
|
||||
47de5dd0b812: Pulling fs layer
|
||||
c172f21841df: Pulling fs layer
|
||||
99515e7b4d35: Pulling fs layer
|
||||
99ba982a9142: Pulling fs layer
|
||||
d6b1b89eccac: Pulling fs layer
|
||||
2780920e5dbf: Pulling fs layer
|
||||
7c12895b777b: Pulling fs layer
|
||||
3214acf345c0: Pulling fs layer
|
||||
52630fc75a18: Pulling fs layer
|
||||
dd64bf2dd177: Pulling fs layer
|
||||
b839dfae01f6: Pulling fs layer
|
||||
ebddc55facdc: Pulling fs layer
|
||||
bdfd7f7e5bf6: Pulling fs layer
|
||||
2d4d7adf6272: Pulling fs layer
|
||||
40008157d8d2: Pulling fs layer
|
||||
bd8962e29291: Pulling fs layer
|
||||
cac2ae0193cb: Pulling fs layer
|
||||
74d1dac84ecc: Pulling fs layer
|
||||
dd64bf2dd177: Waiting
|
||||
b839dfae01f6: Waiting
|
||||
ebddc55facdc: Waiting
|
||||
bdfd7f7e5bf6: Waiting
|
||||
2d4d7adf6272: Waiting
|
||||
40008157d8d2: Waiting
|
||||
bd8962e29291: Waiting
|
||||
cac2ae0193cb: Waiting
|
||||
74d1dac84ecc: Waiting
|
||||
2780920e5dbf: Waiting
|
||||
7c12895b777b: Waiting
|
||||
3214acf345c0: Waiting
|
||||
52630fc75a18: Waiting
|
||||
99ba982a9142: Waiting
|
||||
d6b1b89eccac: Waiting
|
||||
47de5dd0b812: Download complete
|
||||
c172f21841df: Verifying Checksum
|
||||
c172f21841df: Download complete
|
||||
99ba982a9142: Verifying Checksum
|
||||
99ba982a9142: Download complete
|
||||
d6b1b89eccac: Verifying Checksum
|
||||
d6b1b89eccac: Download complete
|
||||
99515e7b4d35: Verifying Checksum
|
||||
99515e7b4d35: Download complete
|
||||
47de5dd0b812: Pull complete
|
||||
2780920e5dbf: Verifying Checksum
|
||||
2780920e5dbf: Download complete
|
||||
7c12895b777b: Verifying Checksum
|
||||
7c12895b777b: Download complete
|
||||
3214acf345c0: Verifying Checksum
|
||||
3214acf345c0: Download complete
|
||||
52630fc75a18: Verifying Checksum
|
||||
52630fc75a18: Download complete
|
||||
c172f21841df: Pull complete
|
||||
dd64bf2dd177: Verifying Checksum
|
||||
dd64bf2dd177: Download complete
|
||||
b839dfae01f6: Download complete
|
||||
ebddc55facdc: Verifying Checksum
|
||||
ebddc55facdc: Download complete
|
||||
bdfd7f7e5bf6: Verifying Checksum
|
||||
bdfd7f7e5bf6: Download complete
|
||||
2d4d7adf6272: Verifying Checksum
|
||||
2d4d7adf6272: Download complete
|
||||
bd8962e29291: Verifying Checksum
|
||||
bd8962e29291: Download complete
|
||||
cac2ae0193cb: Verifying Checksum
|
||||
cac2ae0193cb: Download complete
|
||||
40008157d8d2: Verifying Checksum
|
||||
40008157d8d2: Download complete
|
||||
99515e7b4d35: Pull complete
|
||||
74d1dac84ecc: Verifying Checksum
|
||||
74d1dac84ecc: Download complete
|
||||
99ba982a9142: Pull complete
|
||||
d6b1b89eccac: Pull complete
|
||||
2780920e5dbf: Pull complete
|
||||
7c12895b777b: Pull complete
|
||||
3214acf345c0: Pull complete
|
||||
52630fc75a18: Pull complete
|
||||
dd64bf2dd177: Pull complete
|
||||
b839dfae01f6: Pull complete
|
||||
ebddc55facdc: Pull complete
|
||||
bdfd7f7e5bf6: Pull complete
|
||||
2d4d7adf6272: Pull complete
|
||||
40008157d8d2: Pull complete
|
||||
bd8962e29291: Pull complete
|
||||
cac2ae0193cb: Pull complete
|
||||
74d1dac84ecc: Pull complete
|
||||
Digest: sha256:ba461b8aa9c042156dbd39c38657fe7431bafa063220eab8d5330a523863da9f
|
||||
Status: Downloaded newer image for cloudflare/cloudflared:2026.6.0
|
||||
docker.io/cloudflare/cloudflared:2026.6.0
|
||||
1.3.3-stable: Pulling from gtstef/filebrowser
|
||||
6a0ac1617861: Pulling fs layer
|
||||
ef8806083e82: Pulling fs layer
|
||||
b74107c861c7: Pulling fs layer
|
||||
adc935def003: Pulling fs layer
|
||||
4f4fb700ef54: Pulling fs layer
|
||||
18695ccc900a: Pulling fs layer
|
||||
45d119d5c397: Pulling fs layer
|
||||
dac52db4fc51: Pulling fs layer
|
||||
6d598f86b2f2: Pulling fs layer
|
||||
8aa349c8396c: Pulling fs layer
|
||||
dac52db4fc51: Waiting
|
||||
6d598f86b2f2: Waiting
|
||||
8aa349c8396c: Waiting
|
||||
adc935def003: Waiting
|
||||
4f4fb700ef54: Waiting
|
||||
18695ccc900a: Waiting
|
||||
45d119d5c397: Waiting
|
||||
b74107c861c7: Verifying Checksum
|
||||
b74107c861c7: Download complete
|
||||
6a0ac1617861: Verifying Checksum
|
||||
6a0ac1617861: Download complete
|
||||
ef8806083e82: Verifying Checksum
|
||||
ef8806083e82: Download complete
|
||||
adc935def003: Verifying Checksum
|
||||
adc935def003: Download complete
|
||||
4f4fb700ef54: Verifying Checksum
|
||||
4f4fb700ef54: Download complete
|
||||
45d119d5c397: Verifying Checksum
|
||||
45d119d5c397: Download complete
|
||||
dac52db4fc51: Verifying Checksum
|
||||
dac52db4fc51: Download complete
|
||||
6d598f86b2f2: Verifying Checksum
|
||||
6d598f86b2f2: Download complete
|
||||
18695ccc900a: Verifying Checksum
|
||||
18695ccc900a: Download complete
|
||||
6a0ac1617861: Pull complete
|
||||
8aa349c8396c: Verifying Checksum
|
||||
8aa349c8396c: Download complete
|
||||
ef8806083e82: Pull complete
|
||||
b74107c861c7: Pull complete
|
||||
adc935def003: Pull complete
|
||||
4f4fb700ef54: Pull complete
|
||||
18695ccc900a: Pull complete
|
||||
45d119d5c397: Pull complete
|
||||
dac52db4fc51: Pull complete
|
||||
6d598f86b2f2: Pull complete
|
||||
8aa349c8396c: Pull complete
|
||||
Digest: sha256:eb3733681db8757412632c61a99ad656f0d94ed6781bb2ea114b4d70babab78c
|
||||
Status: Downloaded newer image for gtstef/filebrowser:1.3.3-stable
|
||||
docker.io/gtstef/filebrowser:1.3.3-stable
|
||||
1.1.0: Pulling from admin/felhom-samba
|
||||
897d797d2723: Pulling fs layer
|
||||
3051591aa250: Pulling fs layer
|
||||
ce57a3f93416: Pulling fs layer
|
||||
fb94eeec2fe1: Pulling fs layer
|
||||
fb94eeec2fe1: Waiting
|
||||
ce57a3f93416: Verifying Checksum
|
||||
ce57a3f93416: Download complete
|
||||
fb94eeec2fe1: Verifying Checksum
|
||||
fb94eeec2fe1: Download complete
|
||||
897d797d2723: Verifying Checksum
|
||||
897d797d2723: Download complete
|
||||
3051591aa250: Verifying Checksum
|
||||
3051591aa250: Download complete
|
||||
897d797d2723: Pull complete
|
||||
3051591aa250: Pull complete
|
||||
ce57a3f93416: Pull complete
|
||||
fb94eeec2fe1: Pull complete
|
||||
Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10
|
||||
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0
|
||||
gitea.dooplex.hu/admin/felhom-samba:1.1.0
|
||||
[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) …
|
||||
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'.
|
||||
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
|
||||
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'.
|
||||
[golden] baking the first-boot SSH host-key regeneration unit (F3) …
|
||||
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'.
|
||||
[golden] identity-clean + minimize …
|
||||
[golden] stop + archive …
|
||||
INFO: including mount point rootfs ('/') in backup
|
||||
INFO: including mount point mp0 ('/var/lib/felhom') in backup
|
||||
INFO: archive file size: 624MB
|
||||
INFO: Finished Backup of VM 9100 (00:00:40)
|
||||
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_08_23-07_32_56.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive)
|
||||
[golden] publishing golden (655194776 bytes, sha256 19f5904f53792684…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.222.0/golden.tar.zst
|
||||
[golden] pre-delete existing: HTTP 404 (404/204 expected)
|
||||
[golden] upload OK (HTTP 201)
|
||||
GOLDEN_VERSION=0.222.0
|
||||
GOLDEN_SHA256=19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037
|
||||
[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.222.0 / 19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037
|
||||
[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)
|
||||
@@ -50,6 +50,28 @@ recorded waiver in the register, never a habit of bypassing.
|
||||
FAIL-CLOSED, BUT HONEST ABOUT NOT KNOWING. An absent controller clone, or a CHANGELOG whose top
|
||||
header cannot be parsed, exits **2 (INCONCLUSIVE)** — never 0. The runner reports 2 distinctly for
|
||||
exactly this reason: an undetermined result is not a pass, and it is not a conviction either.
|
||||
|
||||
── THE SECOND BLINDNESS, R-385 (added 2026-08-23) ────────────────────────────────────────────
|
||||
|
||||
Until this change the gate asked ONE question — *is the golden BEHIND the record?* — and so it could
|
||||
only ever catch a forgotten bake. It said nothing when the golden was **AHEAD** of the record, and
|
||||
that is not a harmless direction: a golden ahead of every CHANGELOG heading was built from something
|
||||
**never written down**.
|
||||
|
||||
That is not a hypothetical. Controller **0.221.1** was built, baked AND vouched on 2026-08-23 while
|
||||
the newest heading in the controller CHANGELOG still read v0.221.0 — the fix had been written inside
|
||||
the v0.221.0 entry instead of getting its own. Every gate was green throughout, including this one,
|
||||
measured: `newest released 0.221.0 / newest golden baked 0.221.1 → OK`. The fleet ran a version the
|
||||
record did not name.
|
||||
|
||||
So the test is no longer "behind?" but "**is the version we are shipping WRITTEN DOWN?**". The gate
|
||||
now looks for the baked version's own `## vX.Y.Z` heading anywhere in the CHANGELOG — not merely at
|
||||
the top, because an entry may legitimately be overtaken by later ones; what may never happen is that
|
||||
it is absent. An unrecorded golden is convicted (exit 1) exactly like a stale one.
|
||||
|
||||
**Why membership and not `baked > released`.** A comparison against the newest heading alone would go
|
||||
green again the moment ANY later entry was written, leaving 0.221.1 permanently unrecorded and the
|
||||
gate permanently silent about it. Membership cannot be satisfied by an unrelated later release.
|
||||
"""
|
||||
import os
|
||||
import re
|
||||
@@ -67,16 +89,30 @@ RELEASED_RE = re.compile(r"^##\s+v(\d+)\.(\d+)\.(\d+)\b")
|
||||
EVIDENCE_RE = re.compile(r"^golden-(\d+)\.(\d+)\.(\d+)-\d{4}-\d{2}-\d{2}$")
|
||||
|
||||
|
||||
def newest_released():
|
||||
"""(tuple, str) of the newest controller release, or (None, reason)."""
|
||||
def released_versions():
|
||||
"""(newest_tuple, note, set_of_all_tuples) of the controller releases, or (None, reason, set()).
|
||||
|
||||
R-385: the whole set is returned, not only the newest. The newest answers "is the golden behind?";
|
||||
membership answers "is the version we are shipping written down at all?" — and only the second
|
||||
question could have caught 0.221.1, whose heading was missing while a NEWER heading existed.
|
||||
"""
|
||||
if not os.path.isfile(CONTROLLER_CHANGELOG):
|
||||
return None, "controller clone not found at %s" % CONTROLLER_CHANGELOG
|
||||
return None, "controller clone not found at %s" % CONTROLLER_CHANGELOG, set()
|
||||
newest = None
|
||||
note = ""
|
||||
every = set()
|
||||
with open(CONTROLLER_CHANGELOG, encoding="utf-8") as fh:
|
||||
for line in fh:
|
||||
m = RELEASED_RE.match(line)
|
||||
if m:
|
||||
return tuple(int(g) for g in m.groups()), line.strip()[:90]
|
||||
return None, "no '## vX.Y.Z' header found in %s" % CONTROLLER_CHANGELOG
|
||||
v = tuple(int(g) for g in m.groups())
|
||||
every.add(v)
|
||||
if newest is None:
|
||||
# newest-first by convention: the FIRST heading is the newest release.
|
||||
newest, note = v, line.strip()[:90]
|
||||
if newest is None:
|
||||
return None, "no '## vX.Y.Z' header found in %s" % CONTROLLER_CHANGELOG, set()
|
||||
return newest, note, every
|
||||
|
||||
|
||||
def newest_baked():
|
||||
@@ -99,7 +135,7 @@ def vstr(v):
|
||||
|
||||
|
||||
def main():
|
||||
released, rel_note = newest_released()
|
||||
released, rel_note, every_released = released_versions()
|
||||
if released is None:
|
||||
print("GOLDEN CURRENCY GATE INCONCLUSIVE: %s" % rel_note)
|
||||
sys.exit(2)
|
||||
@@ -113,6 +149,24 @@ def main():
|
||||
print(" newest released controller : %s (%s)" % (vstr(released), rel_note))
|
||||
print(" newest golden baked : %s (documentation/tests/%s)" % (vstr(baked), bake_note))
|
||||
|
||||
# R-385 — UNRECORDED, checked before "behind". A golden whose version has no heading of its own
|
||||
# was built from something never written down, and that is a different (worse) fault than a
|
||||
# forgotten bake: there is nothing to read to find out what the fleet is running.
|
||||
if baked not in every_released:
|
||||
print("")
|
||||
print("GOLDEN CURRENCY GATE FAILED: golden %s is baked but UNRECORDED — the controller "
|
||||
"CHANGELOG has no '## v%s' heading." % (vstr(baked), vstr(baked)))
|
||||
print("The newest heading is %s. A golden ahead of the record was built from a version "
|
||||
"nobody wrote down, so no one can read what the fleet is running." % vstr(released))
|
||||
print("Fix: give v%s its own '## v%s — <what changed>' heading in "
|
||||
"felhom-controller/CHANGELOG.md, above the entries it supersedes. If its fix is "
|
||||
"currently described inside another version's entry, MOVE that text — do not "
|
||||
"duplicate it, and do not delete the reasoning." % (vstr(baked), vstr(baked)))
|
||||
print("If this bake was a throwaway that must never be delivered, delete its "
|
||||
"documentation/tests/golden-<VER>-<DATE>/ directory — never leave it to read as "
|
||||
"shipped.")
|
||||
sys.exit(1)
|
||||
|
||||
if released > baked:
|
||||
print("")
|
||||
print("GOLDEN CURRENCY GATE FAILED: controller v%s is released and NO golden carries it "
|
||||
|
||||
Reference in New Issue
Block a user