diff --git a/REPORT.md b/REPORT.md index 13058648..796c6453 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,77 +1,62 @@ -# REPORT — R-379/R-380/R-381/R-382: the undo copy goes back (2026-08-22) +# REPORT — R-361 and the two loose ends v0.220.2 left (2026-08-22 → 23) -Companion to `felhom-controller` **v0.220.0 → v0.220.1 → v0.220.2**. Full record: -`documentation/audits/DRILL-r379-rollback-2026-08-22/`. +Companion to `felhom-controller` **v0.221.0 → v0.221.1**. Full record: +`documentation/audits/DRILL-r361-2026-08-22/`. -## What shipped +## What this repo carried -**R-379 and R-380 were one failure with one fix.** Both ended with a half-restored database; the only -difference was whether it looked broken. When the replay fails, the product now re-applies the -customer's own pre-restore copy — the same `ImportDump` call a person ran by hand yesterday to recover -both apps — and the app comes back with a message saying **both** that the restore failed and that the -data is as it was. +- **`documentation/architecture/07-backup-architecture.md`** — a dated **[FACT]** on R-361 (the + comment that asserted an invariant the code did not have, and what it cost), and a **[DESIGN]** on + the `db_dumps` decision **including the trap it created**: a stable list lets + `CaptureRecoveryUnit`'s already-current early return fire, so anything that must happen on every + capture has to sit above that check. +- **`documentation/architecture/00-capability-map.md`** — the **negative** from Part 3, recorded so it + is not re-derived: a HELD app does **not** raise the dead-app alarm, measured on the shipped build, + and the reading that said it would was wrong and why. +- **`STATUS.md`** — the outcome in plain words; the deciding section says what happens if nothing is + done. +- **`documentation/tests/golden-0.221.1-2026-08-23/`** — the golden bake. +- **Register** — R-361 closed and compressed; **R-383** and **R-384** opened. -**The whole undo set, matched on the run's own stamp**, never on the `pre-restore-` prefix and never -just the first file. **When the rollback also fails the app is held stopped** — the operator's ruling — -with every start path refusing it, the app-stop marker ended so nothing auto-restarts it, and the row -red rather than green. +## Part 3 — a measurement that cancelled a Part, and that is a good outcome -**R-381:** the failure message stopped pasting engine output (407→257 bytes on Postgres; the MariaDB -one had been 615 bytes with rows out of the customer's own database). The full text now reaches the -operator log, which never had it. -**R-382:** the summary log prints the volume count it already held. -**Also:** undo copies resolve to their own app and are capped at 3. +The runbook's reading was that a held app alarms as a dead app. **It does not.** On the shipped +v0.220.2 a hold was created deliberately; `docmost` aggregated to `unhealthy`; `IsDownState` is +`{stopped, exited, degraded}`; the dead-app heartbeat read **`0 currently down`** at scans 600 and +620 with the scans demonstrably running over it. **Part 2 was dropped in full** and **no register row +was opened**, exactly as the runbook directs. -## Documents updated here +**The positive control took three attempts, and that is the second finding.** Two live attempts +failed to produce a lasting down state at all — `privatebin` went `stopped` (whitelisted by design) +and `bookstack` went `degraded` then `unhealthy`. An absent alarm from a detector never shown working +proves nothing, so the control was moved to the layer the detector lives in: `classifyRunStates` is a +pure function, and it raises the banner for `degraded`/`exited` while staying silent for the states +measured live. -- `documentation/architecture/07-backup-architecture.md` §6.3 — a dated **[DESIGN]** paragraph on the - failure ladder **replay → rollback → hold**, and why an engine flag does not close it. -- `STATUS.md` — the outcome in plain words; the deciding section says what happens if nothing is done. -- `documentation/backlog/` — R-379…R-382 compressed into `CLOSED-ITEMS.md`, each keeping its title, - shipping version, evidence path and every sentence that states a rule. Full text: - `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md`. +## Findings opened -**Register size:** `OPEN-ITEMS.md` **330 683 → 325 236 bytes**; `CLOSED-ITEMS.md` **63 507 → 66 777**. +- **R-383 (MEDIUM)** — the double-failure message tells the customer *"a korábbi állapot mentése + megvan"* while naming the very file whose absence caused the failure. Observed on **both** v0.220.2 + and v0.221.1. **R-361's own class** — a sentence asserting a property the code does not check. +- **R-384 (MEDIUM)** — an app whose **database** has died reads `unhealthy` and raises no dead-app + banner and no customer e-mail, because `aggregateState` checks `unhealthy > 0` before the + mixed-case degraded branch. Not invisible everywhere (the health report counts it), but it does not + alarm. -## The live walk found two defects in the fix itself +**Register size: `OPEN-ITEMS.md` 325 236 → 327 266 bytes; `CLOSED-ITEMS.md` 66 777 → 68 464.** +R-361's full text: `git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md`. -**Both are recorded because the walk, not the tests, caught them.** +## The golden was baked here, and why -1. **The rollback used a dead container (fixed v0.220.1).** The DB-only start re-creates the DB - container, so the id captured at dump time is dead by rollback time. Measured: captured - `9adbc14f9af6`, re-created `309795897b82`, rollback timed out after 30 s — **the app was held for an - infrastructure reason while its data was recoverable.** No unit test could see it: they all inject - the import seam and never look at container identity. -2. **The operator route did not take effect (fixed v0.220.2).** `--clear-restore-hold` runs as a second - process; it cleared the file and the running controller went on refusing. Found by using it. +`golden-currency` refused the docs push: 0.221.1 released with no golden. **Not circular** — a golden +needs the controller image, already pushed, not this commit — so the gate was satisfied by doing the +work rather than bypassed. **No push in this session used `--no-verify`.** -## A red-proof that PASSED - -Of nine mutations, **one did not fail its test** and is reported rather than omitted: the R-381 -behavioural test injected below `ImportDump`, so a leak reintroduced inside `ImportDump` was invisible -to it. A guard now sits at that layer and the mutation convicts. - -## What did not reproduce - -The undo copies rendering as app rows on the customer's backup page. The live page was read **before** -any change: zero `pre-restore` strings while four such files sat on disk, with a positive control -showing 8 real rows. The phantom name was real as a map **key**, never a row. Fixed as a naming defect; -their visibility is unchanged and deliberate. - -## The golden was baked in this session, and why - -The `golden-currency` gate refused this docs push: 0.220.2 was released with no golden. **That block -is not circular** — a golden needs the controller image, which was already pushed, not this commit — -so the gate was satisfied by doing the work it asked for rather than bypassed with `--no-verify`. -**No push in this session used `--no-verify`.** - -Golden **0.220.2**, sha256 `cb439418c7005ce01bcb8126bb6385688c2f408c3c4649f6740001bc2c864ed5`, -657 271 965 B, round-trip verified. All five markers hit, both negative controls at zero, both -token-leak greps proved able to convict before their zeros were accepted. Record: -`documentation/tests/golden-0.220.2-2026-08-22/`. +Golden **0.221.1**, sha256 `1c8bf6cf08cadabeca6331f38360d905e867c235067cd10c2716915b6e6df089`, +656 966 079 B, round-trip verified, all markers hit, both negative controls at zero, both token-leak +greps proved able to convict before their zeros were accepted. ## Operator follow-up -**Vouch** the golden — Hub → Configuration → Day-0 artifacts, a **three-field** save: -`golden_version` **0.220.2**, `agent_version` **0.130.0**, `min_agent` **0.129.0**. **Then** raise the -floor to **0.220.2**, last, in its own save. +**Vouch** the golden — a **three-field** save: `golden_version` **0.221.1**, `agent_version` +**0.130.0**, `min_agent` **0.129.0**. **Then** raise the floor to **0.221.1**, last, in its own save. diff --git a/STATUS.md b/STATUS.md index abbbd513..75b3c397 100644 --- a/STATUS.md +++ b/STATUS.md @@ -12,11 +12,11 @@ NOT yet delivered: two steps below are yours.** *This section is allowed to be longer than one screen, and each item says what happens if you do nothing.* -1. **Vouch the golden carrying controller 0.220.2** — Hub → Configuration → Day-0 artifacts. +1. **Vouch the golden carrying controller 0.221.1** — Hub → Configuration → Day-0 artifacts. **It is already baked, published and round-trip verified** - (`documentation/tests/golden-0.220.2-2026-08-22/`); only the vouch is left, and only you can do it. - **It is a THREE-field save:** `golden_version` → **0.220.2**, `agent_version` → **0.130.0**, - `min_agent` → **0.129.0**. **Then** raise the floor to **0.220.2**, last, in its own save. + (`documentation/tests/golden-0.221.1-2026-08-23/`); only the vouch is left, and only you can do it. + **It is a THREE-field save:** `golden_version` → **0.221.1**, `agent_version` → **0.130.0**, + `min_agent` → **0.129.0**. **Then** raise the floor to **0.221.1**, last, in its own save. **If you do nothing:** the fleet stays on 0.219.0, so a failed database restore still leaves an app broken with an unusable copy — the thing today's release fixes reaches nobody. New machines still receive 0.219.0. The build system stays red about it and will mail you on every push. @@ -50,6 +50,14 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an ## Shipped +- **Taking the safety copy no longer destroys the app's own backup** (R-361, controller 0.221.1, + proven on `demo-hp`). Before every restore the machine saves a copy of your live database. To do + that it called the ordinary backup routine — **which always writes to the app's normal backup + filename first** — so the app's real backup was overwritten and then renamed away. Until the next + nightly run the app had **no database backup of its own**, and a local recovery in that window would + have told you the app never had a database. A comment in the code said this could not happen; it + could, and had been happening for four months. Proven fixed the only way it can be: the app's own + backup file is now **byte-identical before and after a restore**, on both database types. - **A failed database restore now puts your data back by itself** (R-379/R-380, controller 0.220.2, proven on `demo-hp`). Until today, if a restore of an app's database went wrong, the machine had already taken a copy of your live database — a good copy — and **nothing in the product could put it diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 433114bf..4b0074ec 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -55,7 +55,19 @@ | **The offsite repository password can be RECOVERED from the sealed escrow with the customer's recovery code** | hub v0.94.0, agent v0.125.0, controller v0.195.0 | **PROVEN-LIVE (2026-08-04)** | On demo-felhom, through the real endpoints end to end: the box fetched its own sealed blob from the hub with its own per-host credential (hub log: *escrow blob SERVED … 572 opaque bytes, self_scope=true*), the agent unsealed it with the customer's recovery code, and the extracted repository password's sha256 was **byte-identical** to the one on disk — `c60c8bc737a6…`, which is ALSO the hash the hub had independently stored, so three sources agree. Five minutes earlier the same path with a WRONG code failed closed at age's KDF with nothing written, which proves links 6 and 7 ran independently of the success. R was searched for afterwards and found in 0 log lines and 0 files, with a positive control confirming the search would have found it. Evidence: per-repo CHANGELOGs; `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 links 6–8 | **WHAT THIS ROW DOES NOT CLAIM, stated because the previous over-claim here was struck out four hours earlier.** It covers the KEY, not the DATA. **R-199 BACK-POINTER (omitted when this row was written): the recovery chain's link inventory and the per-link status live in `audits/RECON-offsite-dr-chain-2026-08-04.md` §3; links 1–8 are walked, 9–11 are not.** **Re-confirmed 2026-08-04 evening by an attempt to prove the DATA half:** the R-201 drill was prepared on demo-hp and **halted before the wipe** — the sentinel file was not in the off-site snapshot (R-203), so the wipe would have destroyed it and proven nothing. **No file has still ever been restored from an off-site backup after a wipe** (`audits/DRILL-r201-offsite-recovery-2026-08-04.md`). The install half of the chain (controller v0.196.0 `--recover-offsite-install`, R-200) is likewise unit-proven only — it has never run against a live recovery. A recovered password has never been **installed** (the diagnostic compares and refuses to write, by design), no existing repository has ever been **reopened** under one, and **no file has ever been restored** from an off-site history via a recovered key. Links 9–11 of the chain are open (R-200's remaining half, R-201). The proof also used a box whose local key still exists — the rebuilt-box case, where there is nothing to compare against, is exactly what the drill covers and it has not run || DR tier by default: PBS + WireGuard base infra on every install, hub-controlled activation | installer v1.15, agent v0.86, hub v0.51 | **PROVEN-LIVE** (2026-07-21) | `DRILL-day0-take2-2026-07-12` §2 (WG enabled both modes, PBS-DR descriptor auto-provisioned ~1s after WG registration, zero operator steps); ships installer v1.15/agent v0.86/hub v0.51 | Live only on demo/drill fleet. (Cited spike was slice-0 mechanics — shipped nothing; corrected.) **⚠ The candidate upgrade to PROVEN-LIVE is WITHDRAWN — the 2026-07-18 rehearsal produced a live counter-example (R-39).** On the reborn N100 the descriptor auto-provisioned and the agent reported `converged state=applied` (16:45:53), yet **the storage is dead**: `pvesm status` → `felhom-pbs: error fetching datastores - 401 Unauthorized` / `inactive`, and a direct probe with the stored credential returns **401 on every endpoint including `/version`** while the WG transport is healthy (handshake 9 s, 27.9 ms RTT) — i.e. authentication failure, not ACL scope. Root cause in the evidence: **the hub minted a SECOND token secret at 16:47:52, two minutes after the agent had applied the first, and `consumed_at` is still NULL**; the converged state machine will not re-apply, and the agent's 15-minute verify loop **cannot even read the credential to notice** (`open /etc/pve/priv/storage/felhom-pbs.pw: permission denied` — non-root agent reading a file it writes through a root wrapper). A tier that reports `applied` while silently unable to authenticate is exactly the shape that must not carry a PROVEN-LIVE badge. See `tests/VALIDATION-n100-rehearsal-2026-07-18.md` F2 and `pbs-dr-state.txt`. **agent v0.89.0 closes the F4 non-default-storage-id gap (R-22) — PROVEN-LIVE 2026-07-17:** the reconcile self-grants the ACL through the root wrapper on a pre-check 403 instead of dead-locking. Reproduced F4 on the demo (marker moved aside = reinstall fresh-state + felhom-offsite ACLs revoked) → next reconcile tick `pbsdr: pre-check 403 … self-granting … (R-22)` → `converged state=adopted` in ~3 s, ACLs self-restored, `pvesm status felhom-offsite`=active, zero operator action. No more one-shot `pveum` grant **2026-07-21 — the R-39 fleet fix SHIPPED (hub v0.68.0 + agent v0.91.2), closing the self-heal chain end to end.** The three defects that let a box be `applied` and dead simultaneously are each addressed: the hub stamps a monotonic `secret_generation` into the descriptor so a credential re-key finally MOVES the content hash the agent re-applies on; the wrapper gains a narrow `read` verb so the non-root agent can read the credential it writes (it never could — `/etc/pve/priv` is 0700 root:www-data, which made the verify loop blind by construction); and `pbs.ProbeAuth` turns a 401 into a loud `auth_failed` that the existing `pbsdrheal` damper escalates to a fresh mint. Plus a consumed_at honesty gauge for the disagreement no single tier can see (box says `applied`, hub's staged secret never consumed). Proven live on felhom-pve: the agent read its credential through the wrapper (`rc=0`) and probed successfully (`credential probe OK storage=felhom-pbs`). **STOP-2 RAN 2026-07-21 AND THE CHAIN CLOSED — 13 SECONDS, operator click to converged.** The operator pressed **Re-issue PBS credentials**; the identical click on 2026-07-18 did nothing at all. Full chain (hub UTC / host CEST = UTC+2): `08:39:31Z` hub mints a fresh secret, **generation 0 → 1**, and the descriptor gains `"secret_generation": 1` — with `token_id` and `fingerprint` **byte-identical**, i.e. exactly the re-key shape that used to be invisible → `10:39:34` the agent READS its credential through the wrapper (leg b — the read that was impossible until v0.91.0) → `10:39:38` **`ERROR pbsdr: the DR endpoint REJECTED this box's credential — the tier is applied and DEAD` `previous_state=applied`** (leg c: the exact R-39 failure state, detected out loud for the first time ever) → `10:39:45` **`one-time token secret consumed`** `secret_len=36` (leg a: **NO short-circuit** — this is the line that never appeared on 2026-07-18) → `10:39:45` `felhom-pbs-apply reconcile` (the set-only wrapper, no `--server`) → `10:39:47` **`pbsdr: converged state=applied`**. Corroboration: the agent marker hash moved to `afbb3b41…` (it was byte-identical to the pre-reissue marker in the failure); `consumed_at` stamped `08:39:45Z`; the on-disk secret's mtime moved `2026-07-18 20:28:52` → `2026-07-21 10:39:45`; a live probe with the NEW credential returns **200**; three consecutive hub reports trace the whole state machine `applied → auth_failed → applied`; and **zero** `pbsdr_selfheal` escalations fired — the box healed through the descriptor path before the damper was ever needed, with exactly ONE mint and ONE consume and no `consumed-failed.json`. **Row upgraded to PROVEN-LIVE (2026-07-21).** Evidence: `felhom-agent/REPORT.md` (2026-07-21). | -| **A customer's file survives a machine rebuild and comes back — the whole off-site story, end to end** | hub v0.94.0, agent v0.125.0, controller v0.197.0 | **PROVEN-LIVE (2026-08-04 night drill)** | demo-hp's controller data volume was destroyed and the sentinel deleted from disk. The customer's recovery code then produced `8a9e33aa4da6…` — **byte-identical** to the pre-wipe on-disk key and to the hub's independent record — and installed cleanly on the bare box. `identity_blob` was unchanged across the wipe (572 B, `updated_at` still 11:11:37). Evidence: `audits/DRILL-r201-night-run-2026-08-04.md` §2 | **WHAT IT DOES AND DOES NOT CLAIM.** PROVEN: after a real rebuild the key recovers byte-identical, the EXISTING repository opens (3 snapshots, 42 026 B — the pre-wipe size; not a fresh history), and the customer restore flow returns the file byte-identical. **ALL FOUR MANUAL INTERVENTIONS ARE CLOSED, AND THE CUSTOMER IS NOW OFFERED THE RECOVERY (2026-08-05). ONE QUALIFIER REMAINS AND IT IS NOT THE OLD ONE — state it precisely.** The 2026-08-04 drill needed four interventions between "the key is recoverable" and "the file is back", none of them in any design document (R-204). **Items 1–3 closed in controller v0.198.0 + hub v0.95.0** (the reset code works first time; a Re-issue no longer marks a healthy escrow stale; a unit restore states it returned the app's definition and database and NOT the customer's files). **Item 4 closed in controller v0.199.0 + hub v0.96.0** (a rebuilt box DECLARES `offsite.state=needs_credential` and `internal/offsiteheal` re-arms the stored credential before minting). **AND THE "needs someone who knows to look" HALF IS NOW GONE TOO (controller v0.200.0, R-193):** a full-page recovery screen takes over the landing pages while the hub holds a sealed package this box cannot open, explains that nobody can replace a lost recovery code, takes the code, opens the repository and lists what is in it — proven live on demo-felhom, which is genuinely in that state (`/launcher` → 302 `/recovery`; a wrong code refused with nothing written, three times, no lockout; the code found in no file, no container log and no debug ring, with a planted-copy positive control that first exposed a mis-aimed sweep). **WHAT THE QUALIFIER IS NOW, and it is narrower:** (1) **the final unlock has never been driven with a CORRECT code through the page** — no recovery code was kept for demo-felhom's orphaned history and demo-hp's is operator-held out of band, so the live run exercised the whole chain (handler → agent → hub fetch → age KDF) and stopped at the unseal; the install and listing rest on tests. (2) **putting files back in place is deliberately NOT part of this** — restore stays per-app, and the step after the listing is R-213. (3) **the journey has still not been re-walked end to end since these fixes** — the closures are proven individually, not as one uninterrupted run. That re-walk is one more drill and it is what is owed. Scope: `demo-hp`, a controller-data-volume rebuild — **NOT** a total host loss, and **NOT** a guest reprovision. **⚠ THE RE-WALK RAN ON 2026-08-05 (CAMPAIGN-11) AND IT FAILED. THIS ROW STAYS AS IT IS — it records a key recovering and a file returning, which remains true — but the JOURNEY row below is the honest verdict and this row must not be read as covering it.** The campaign built a throwaway appliance from the published ISO, destroyed it, and walked the customer's route: **the data half PASSED** (three sentinels byte-identical, including a 12 MB binary and a non-ASCII Hungarian filename, out of the pre-wipe snapshot in 16 s) and **the journey half FAILED** — four operator interventions, three needing root on the appliance, and the first of them was the machine telling the customer their perfectly correct recovery code was wrong (R-216). **R-198's retention is NO LONGER unit-proven — it was PROVEN IN PRODUCTION on 2026-08-05** (CAMPAIGN-11 Phase 3): the first supersession since the fix retained `identity_blob` at **572 B, byte-length exact**, carrying the OLD key `626e424670248db3` while the current row moved to the newly minted `e11a6c542b73477a`; `offsite_repo_key_changed` and `escrow_superseded` both fired at the instant of supersession, and the next off-site run REFUSED as `orphaned` rather than starting a fresh history. Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` **RE-WIDENED 2026-08-22, and the narrowing below is now HISTORY — read both, in order.** Both defects named in the 2026-08-21 narrowing are closed and proven live: **R-354** (controller v0.218.0, `volReplay`) and **R-356** (controller v0.219.0). The **40-class end-to-end story is now WALKED**, including the hardest ten of it: `audits/DRILL-r356-hot-only-restore-2026-08-22/` proved a driveless app with NO database (`privatebin`: planted, off-sited, deleted, restored, **15/15 files byte-identical**, two Hungarian accented names), and `audits/DRILL-r356b-driveless-db-restore-2026-08-22/` proved a driveless app **WITH** a database on both engines — `docmost` (Postgres 16) and `bookstack` (MariaDB 12.3), each planted through the app's own interface, destroyed for real, and returned with accented names byte-identical. **Five legs that had never run in any combination all ran and all succeeded:** the undo copy, DB-service identification, the volume replay, the DB-only start window, and the dump replay on top. **A second claim was walked at the same time:** R-164's F17 ordering — *the logical dump wins over the volume tar's copy of the same database* — was recorded only for the LOCAL path (`restore_unit.go:262-266`) and is now measured on the **off-site** path too, by a three-way discriminator (volume tar `ORIGINAL-VALUE-A`, altered dump `ALTERED-VALUE-B`, live `LIVE-VALUE-C3`; result **`ALTERED-VALUE-B`**). **WHAT IS STILL NOT CLAIMED:** the FAILURE path is where this class is weak, not the success path — a corrupt dump leaves the customer with an emptied or partially-applied database and an undo copy **no product action can apply** (**R-379**), and on MariaDB it does so behind an app that reports `health=healthy` (**R-380**). The success story is proven; the recovery-from-a-bad-restore story is not. +| **A customer's file survives a machine rebuild and comes back — the whole off-site story, end to end** | hub v0.94.0, agent v0.125.0, controller v0.197.0 | **PROVEN-LIVE (2026-08-04 night drill)** | demo-hp's controller data volume was destroyed and the sentinel deleted from disk. The customer's recovery code then produced `8a9e33aa4da6…` — **byte-identical** to the pre-wipe on-disk key and to the hub's independent record — and installed cleanly on the bare box. `identity_blob` was unchanged across the wipe (572 B, `updated_at` still 11:11:37). Evidence: `audits/DRILL-r201-night-run-2026-08-04.md` §2 | **WHAT IT DOES AND DOES NOT CLAIM.** PROVEN: after a real rebuild the key recovers byte-identical, the EXISTING repository opens (3 snapshots, 42 026 B — the pre-wipe size; not a fresh history), and the customer restore flow returns the file byte-identical. **ALL FOUR MANUAL INTERVENTIONS ARE CLOSED, AND THE CUSTOMER IS NOW OFFERED THE RECOVERY (2026-08-05). ONE QUALIFIER REMAINS AND IT IS NOT THE OLD ONE — state it precisely.** The 2026-08-04 drill needed four interventions between "the key is recoverable" and "the file is back", none of them in any design document (R-204). **Items 1–3 closed in controller v0.198.0 + hub v0.95.0** (the reset code works first time; a Re-issue no longer marks a healthy escrow stale; a unit restore states it returned the app's definition and database and NOT the customer's files). **Item 4 closed in controller v0.199.0 + hub v0.96.0** (a rebuilt box DECLARES `offsite.state=needs_credential` and `internal/offsiteheal` re-arms the stored credential before minting). **AND THE "needs someone who knows to look" HALF IS NOW GONE TOO (controller v0.200.0, R-193):** a full-page recovery screen takes over the landing pages while the hub holds a sealed package this box cannot open, explains that nobody can replace a lost recovery code, takes the code, opens the repository and lists what is in it — proven live on demo-felhom, which is genuinely in that state (`/launcher` → 302 `/recovery`; a wrong code refused with nothing written, three times, no lockout; the code found in no file, no container log and no debug ring, with a planted-copy positive control that first exposed a mis-aimed sweep). **WHAT THE QUALIFIER IS NOW, and it is narrower:** (1) **the final unlock has never been driven with a CORRECT code through the page** — no recovery code was kept for demo-felhom's orphaned history and demo-hp's is operator-held out of band, so the live run exercised the whole chain (handler → agent → hub fetch → age KDF) and stopped at the unseal; the install and listing rest on tests. (2) **putting files back in place is deliberately NOT part of this** — restore stays per-app, and the step after the listing is R-213. (3) **the journey has still not been re-walked end to end since these fixes** — the closures are proven individually, not as one uninterrupted run. That re-walk is one more drill and it is what is owed. Scope: `demo-hp`, a controller-data-volume rebuild — **NOT** a total host loss, and **NOT** a guest reprovision. **⚠ THE RE-WALK RAN ON 2026-08-05 (CAMPAIGN-11) AND IT FAILED. THIS ROW STAYS AS IT IS — it records a key recovering and a file returning, which remains true — but the JOURNEY row below is the honest verdict and this row must not be read as covering it.** The campaign built a throwaway appliance from the published ISO, destroyed it, and walked the customer's route: **the data half PASSED** (three sentinels byte-identical, including a 12 MB binary and a non-ASCII Hungarian filename, out of the pre-wipe snapshot in 16 s) and **the journey half FAILED** — four operator interventions, three needing root on the appliance, and the first of them was the machine telling the customer their perfectly correct recovery code was wrong (R-216). **R-198's retention is NO LONGER unit-proven — it was PROVEN IN PRODUCTION on 2026-08-05** (CAMPAIGN-11 Phase 3): the first supersession since the fix retained `identity_blob` at **572 B, byte-length exact**, carrying the OLD key `626e424670248db3` while the current row moved to the newly minted `e11a6c542b73477a`; `offsite_repo_key_changed` and `escrow_superseded` both fired at the instant of supersession, and the next off-site run REFUSED as `orphaned` rather than starting a fresh history. Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` **RE-WIDENED 2026-08-22, and the narrowing below is now HISTORY — read both, in order.** Both defects named in the 2026-08-21 narrowing are closed and proven live: **R-354** (controller v0.218.0, `volReplay`) and **R-356** (controller v0.219.0). The **40-class end-to-end story is now WALKED**, including the hardest ten of it: `audits/DRILL-r356-hot-only-restore-2026-08-22/` proved a driveless app with NO database (`privatebin`: planted, off-sited, deleted, restored, **15/15 files byte-identical**, two Hungarian accented names), and `audits/DRILL-r356b-driveless-db-restore-2026-08-22/` proved a driveless app **WITH** a database on both engines — `docmost` (Postgres 16) and `bookstack` (MariaDB 12.3), each planted through the app's own interface, destroyed for real, and returned with accented names byte-identical. **Five legs that had never run in any combination all ran and all succeeded:** the undo copy, DB-service identification, the volume replay, the DB-only start window, and the dump replay on top. **A second claim was walked at the same time:** R-164's F17 ordering — *the logical dump wins over the volume tar's copy of the same database* — was recorded only for the LOCAL path (`restore_unit.go:262-266`) and is now measured on the **off-site** path too, by a three-way discriminator (volume tar `ORIGINAL-VALUE-A`, altered dump `ALTERED-VALUE-B`, live `LIVE-VALUE-C3`; result **`ALTERED-VALUE-B`**). **MEASURED 2026-08-22, AND THE READING THAT PROMPTED IT WAS WRONG — recorded so nobody re-derives it.** +It was read from source that a HELD app would raise the dead-app banner and a customer e-mail, because +it keeps one container (its database) and so is not `StateStopped`. **It does not.** A held app +aggregates to `unhealthy`, and `aggregateState` checks `unhealthy > 0` before the mixed-case degraded +branch while `IsDownState` excludes `unhealthy` entirely. Measured on `demo-hp` on the shipped +v0.220.2: hold created 21:11:19Z, dead-app scans every 30 s ran over it, and the heartbeat reported +`0 currently down` throughout. **No suppression was built, because there was nothing to suppress.** +The detector itself is sound — `classifyRunStates` is pure and a `degraded`/`exited` stack does raise +the banner, pinned by `TestClassifyRunStates_PositiveControl_ADownStackDoesAlarm`. What the same +measurement DID expose is **R-384**: an app whose database has died is `unhealthy` too, and is +likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt`. + +**WHAT IS STILL NOT CLAIMED:** the FAILURE path is where this class is weak, not the success path — a corrupt dump leaves the customer with an emptied or partially-applied database and an undo copy **no product action can apply** (**R-379**), and on MariaDB it does so behind an app that reports `health=healthy` (**R-380**). The success story is proven; the recovery-from-a-bad-restore story is not. **NARROWED 2026-08-21 by the backup-truth drill — kept as history; both halves have since closed, see the entry immediately above.** Proven that night on `demo-hp` with planted, hash-recorded files: the **declared-userdata** leg of a **drive-declaring** app does come back byte-identical (`calibre-web`, 5/5 including two Hungarian accented filenames). **Two legs of the same story do NOT:** (a) the off-site restore has **no named-volume leg at all**, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (**R-354**); (b) for the **40 of 53** apps that declare no data drive the off-site restore **refuses outright**, saying a running app „nincs telepítve" (**R-356**). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is **unproven for that class and disproven for the volume leg generally**. The escrow/key half of this row is untouched by that and still stands. Evidence: `audits/DRILL-backup-truth-2026-08-21/evidence/` and `REPORT.md` (2026-08-21). | | **The customer's own UNAIDED recovery journey, end to end** | controller v0.206.0, hub v0.98.0, agent v0.127.0 | **PROVEN-LIVE (2026-08-07, the fifth walk) — an unaided customer journey on a real installation, SCOPED: see what it does not claim** | A throwaway appliance was built from the published ISO on demo-hp, claimed, given three sentinels, escrowed, backed up off-site, then destroyed — guest purged and both enrolled drives wiped. The rebuild and the customer's route were then walked with **no command line inside the guest**. **Data: PASS.** All three sentinels byte-identical (`beb9175d…`, `7c8cb0ad…`, `e012e76f…`), including a 12 MB binary and `C11-őrszem-ékezetes-árvíztűrő.txt` — the Hungarian filename survived disk → restic → SFTP → Storage Box → restore → disk intact. Restored in **16 s** out of the pre-wipe snapshot `f3d9cd67`, non-destructively. **Journey: FAIL.** Four interventions, three of them root on the appliance: R-216 (a correct code called wrong), R-218 (the store never opened; the box stopped asking), R-219 (the promised listing structurally could not render), R-220 (the app could not be redeployed — its drives were unenrollable). Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` | **WHAT IT DOES NOT CLAIM, and what the v0.201.0/v0.97.1 fixes do and do not change.** Six of the nine findings are fixed and each is red-proofed (R-216, R-218, R-219, R-217, R-222, R-215) — but **fixes are not a journey.** This row goes green only when the walk is repeated end to end and completes with no operator intervention. **Three findings are deliberately still open and each is a live blocker for some flow:** **R-214** (the console never stops showing a stale pairing code), **R-220** (after a rebuild the customer's drives cannot be re-enrolled — the deploy refuses and the wizard's list is empty, so no app can be redeployed onto its own data drive), **R-221** (a rebuilt box cannot run the escrow ceremony at all, because the pbsdr convergence marker survives on the host while the seeded `escrow.pbs_storage_id` is rewritten out of `agent.json`). **And R-216's fix does not by itself make a NEW box work:** with the hub guard corrected but the Day-0 manifest still vouching agent 0.120.0, a new box is **HELD, not served** — it stops being lied to, which is better, but the feature does not work for it until agent 0.125.0 is vouched. **⚠ CAMPAIGN-11 PHASES 2 + 4 RAN OVERNIGHT 2026-08-05/06 AND THIS ROW STAYS **FAIL**.** **What they add:** eleven injected faults against the recovery journey, each with a positive control, and a full unattended scheduled cycle. **The backup promise strengthened** — the off-site tier ran itself at 04:15 on a box rebuilt twice and set aside hours earlier (`snapshot_count` 1 → 2), all five daily jobs fired exactly once, and nothing on the must-not list fired. **R-217's and R-215's fixes were proven live under exactly their faults**, the set-aside proved it does not delete (12 535 KB byte-exact at the far end), and a wrong code was refused three times with nothing written and no lockout. **What they do NOT add: any progress on the JOURNEY.** These faults are **not a re-walk** — no customer route was walked end to end — and four new findings say the journey got no better where it matters: **R-224** (a hub outage and a stopped agent are BOTH reported to the customer as a bad recovery code, in 0.056 s and 0.030 s — no unseal attempted; the agent's own `err` distinguishes them and it is discarded at the HTTP boundary, and R-216's gate answers `source=version` so it cannot see reachability — **Phase 1's headline defect relocated from the version channel to the transport**), **R-226** (M1, the only message telling a customer to check their typing, is unreachable on any box that has re-escrowed), **R-225** (an unread store renders `0 pillanatkép · 0 GB` above a card saying it holds backups — 1 snapshot and 12 535 KB were really there), **R-228** (the set-aside history is recorded in `orphaned_renamed_to` and shown by nobody — `OrphanedRenamedTo` has zero references in any template or handler). **§4.1 is now MEASURED rather than deduced** — the floor IS served, from the box's own rendered `GetFloor()` and a cold-started controller's settle-gate line. **§4.2's positive half is still NOT measured** and needs a rebuild. Evidence: `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`, `tests/campaign11-evidence-2026-08-05/journal-phase24.md` **2026-08-06 — THE FIVE PHASE-2 FINDINGS ARE FIXED AND THIS ROW STILL SAYS FAIL.** controller **v0.202.0** + agent **v0.126.0** close R-224, R-226, R-225, R-227, R-228. **What changed:** the unlock path now classifies WHY it failed, from the value and never the text, and the message that mentions typing is reachable from exactly one class — a real refusal — with **unknown defaulting to neutral**. Proven live on the venue with the same wrong code and only the hub's reachability changed: `400 → 502 → 400`. **What it does NOT change: this row.** Fixes are not a journey; nothing here walked a customer end to end, and three of the blockers that made Phase 1 fail are still open (**R-220** in particular is worked around BY HAND on the venue — without that unmount no app can be deployed on a rebuilt box). **What still needs proving:** a re-walk with no operator intervention, and specifically the customer-facing messages end-to-end — they were NOT re-driven, because `/recovery` correctly retires itself once the old data is set aside and restoring that state would be the reconfiguration the fix task forbade. Evidence: `felhom-controller/REPORT.md` **⚠ THE RE-WALK RAN 2026-08-06 (R-201, attended) AND THIS ROW STAYS FAIL — but it is a nearer no.** A brand-new appliance was built from the published ISO, given three sentinels, escrowed, backed up off-site, destroyed (guest purged AND both drives wiped), reinstalled and walked. **THE DATA: PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also identical**, restored in **23 s** out of the pre-destruction snapshot through the customer's own flow. **THE JOURNEY: FAIL — TWO dead ends against Phase 1's four.** (1) **R-218's CONSUME half**: the hub re-staged the credential and said "the box re-consumes on its next cycle"; a full cycle ran (positive control) and it did not — no customer-reachable action fetches it, and the only lever is a command line INSIDE THE GUEST. (2) **R-220**: the drives are still unenrollable after a rebuild, needing a Proxmox-host unmount, without which no app can be redeployed and the restore page stays empty. **The unaided RTO is therefore STILL UNDEFINED**; the attended figure was 30 m 13 s and must not be quoted as the customer number. **What passed and is new:** the recovery screen appeared WITHOUT being sought, answered all three of its questions with a seal date matching the hub exactly, the emailed reset code worked first try, and the unlock was a real 1.528 s unseal that placed the key. **⚠ AND IT DOES NOT CLAIM DELIVERY:** the fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes — which were installed BY HAND. Fleet delivery needs a golden carrying the fixed controller and a vouched agent, and nothing was vouched. Evidence: `tests/rewalk-r201-2026-08-06/journal.md`   **⚠ 2026-08-06 (later the same day) — THE TWO DEAD ENDS ARE CLOSED AND VOUCHED, AND THIS ROW STILL SAYS FAIL.** controller **v0.203.0** + agent **v0.127.0** close **R-218's consume half** and **R-220**, and golden **0.203.0** / agent **0.127.0** / `min_agent` **0.127.0** are now VOUCHED — so the delivery gap the re-walk recorded (“the fresh install landed on versions neither carrying the fixes”) is gone. **Both were then proven live on a genuinely rebuilt box** (Part 4 venue, VM 323): the reinstall landed on agent 0.127.0 with **no downgrade and no hand upgrade** (the previous re-walk's reinstall downgraded 0.126.0→0.125.0), `/disks/candidates` returned **both drives** after the guest was purged with the raw mounts still on the surviving host — before the fix, two empty lists — and both **re-attached through the customer endpoint**; the off-site credential was re-staged by `offsiteheal` at 13:24:57Z and **collected by the box on a tick, unaided**. **Why the row is still FAIL:** the walk did not finish. It stopped at the restore surface — **R-237**, the list was keyed on apps that are currently INSTALLED and currently TOGGLED ON for future backups, so a rebuilt box was shown nothing to restore while the repository held its snapshots. Fixed in **controller v0.204.0** (the store is now the source of the list) and proven live on that same venue with the toggle switched OFF, but **no sentinel was restored in that walk**, so the data half is unproven in EITHER direction for this venue. **R-238 was RECLASSIFIED, not fixed as filed:** `mode=full` without `confirm=1` is step 1 of a deliberate two-step that starts no job by design; the endpoint driver did not carry `?full_prep=` forward. The operator drove the same restore to completion in a browser. Its real residue — that step logging nothing at all, including when the headroom gate refuses — is fixed. **R-236 was WITHDRAWN: not a defect.** **This row goes green only when a walk completes end to end with no operator intervention, restoring a byte-identical sentinel.** Evidence: `tests/part4-rewalk-2026-08-06/journal.md`, `tests/teardown-2026-08-06.md`, `felhom-controller/REPORT.md` **⚠ THE FINAL WALK RAN OVERNIGHT 2026-08-06/07 AND THIS ROW STAYS FAIL — the data half passed again, the journey half failed further from the line than before.** A brand-new appliance was built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed and rebuilt. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `f5c53b03`, including a 12 MB binary and an accented Hungarian filename **whose name BYTES are identical too**. **THE JOURNEY: FAIL, and this time the customer has NO route at all** — `/` lands on the launcher with no recovery pointer, `/recovery` **302s away**, and the remote-backup page offers to **CREATE a new recovery code**, which would orphan the very history the customer's code protects. **There is no field anywhere to enter the code they hold**, and the operator's documented remedy refuses too. Cause (**R-241**): the credential self-heal — proven working unaided in the same walk, six hours earlier, as its best result — writes a FRESH repository password, which moves the box out of `OffsiteRecoveryOffer()`'s (a) pristine case while (b) orphan-detected is unreachable because runs are blocked by `escrow_state: pending`. **R-218's shape one level up.** **What DID pass, and is new:** the whole credential chain ran end to end with zero human action on an unclaimed box — declare, `offsiteheal` re-stages after two reports, the box's 5-minute retry collects it, tier applied — the **first live sighting** of that success line, settling R-218's consume half and R-236's withdrawal. Also new: **R-239** — a fresh install lands on controller **0.203.0** while 0.205.0 is released, so R-234 and R-237 are written but **not delivered**, proven from the customer's side (T2, T3). **This row goes green only when a walk completes with no operator intervention AND a byte-identical sentinel.** Evidence: `tests/finalwalk-r201-2026-08-07/journal.md` **⚠ 2026-08-07 — R-241 WAS DIAGNOSED BY A READ-ONLY SPIKE ON THE STANDING VENUE, AND THE DIAGNOSIS REVERSES THE FIX. This row is unchanged: still FAIL, and no code was written.** The question was whether R-241 is a screen-predicate defect or a minting defect. **It is a MINTING defect.** The screen was telling the truth — there genuinely was nothing recoverable under the key the box held, because the box minted that key itself over the top of a sealed package it already knew the hub was holding. **Three measurements, from the venue rather than from the earlier report:** (1) `WriteOffboxSecrets` (`offbox.go:411`) mints on ONE input — does the file exist — while `OffsiteRecoveryOffer()` and `needsOffsiteCredential()`, both in the same file, consult `GetHubEscrowIdentityPresent()`; the same fact is available on three paths and used on two. (2) That flag was the **precondition of the chain that reached the minting**: the retry job logs only when the declaration is live, and the venue logged `credential retry: … (the box still declares a need; retrying)` at **02:48:03Z** — thirty minutes and six ticks before the mint at 03:18:06Z. (3) **The box computed the answer and discarded it**: at **03:28:03Z**, thirty-five minutes before the customer looked, `EscrowAutoConfirmer.Reconcile` (`escrow_confirm.go:154`) logged `the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…)`. It is recomputed every report cycle and never persisted. **And the hub explicitly disclaims doing this** — `offsiteheal`'s package doc: *"credential automatic, key customer-present — the ruling this session implements and must not quietly widen."* **Shape (b) is structurally unreachable on this box** (the escrow gate at `offbox.go:743` sits upstream of `ensureOffboxRepo`, the only producer of `RepoState="orphaned"`; positive control: the scheduler was alive, 241 `agent-channel-health` ticks, and `offbox-backup` is a `sched.Daily` leg whose slot fell before the destruction). **Two by-products:** **R-243** — a box in this state silently stops backing up off-site and **no alarm fires** (`isStale` requires `escrowed`, `offsite_delivery_stuck` skips the `applied` shape, `backup_failed` needs a run that never happens); and the trap in the obvious fix — `ResetOrphanedRepo` clears the orphan flag without a ceremony, so a hash-mismatch discriminator alone would re-offer the screen forever to a customer who declined the old data. **Q7:** the „Helyreállítási kód létrehozása" button does NOT destroy the data — R-198's retention copies the `identity_blob` — but it converts a self-service recovery into one needing an unbuilt read path (R-199), and it re-enables the screen while invalidating the code that screen accepts. Evidence: `audits/SPIKE-r241-recovery-offer-2026-08-07.md` **⚠ 2026-08-07 (later) — R-241 IS FIXED (controller v0.206.0 + hub v0.98.0) AND THIS ROW STILL SAYS FAIL.** Three changes, following the spike's ruling rather than the obvious reading: the box **no longer mints a repository key while the hub holds a sealed package** (a conjunction, so a first-time box is untouched; the refusal is a HOLDING state that still writes the transport, declared as `offsite.state=awaiting_recovery_key` and shown inert to every existing hub reader); the **hub-vs-local key comparison that was computed every cycle and discarded is now persisted** and drives the offer as **shape (c)**; and **abandoning starts a 14-day countdown** whose terminal step removes the set-aside store and the sealed package **together**, so the offer ends because the state is right rather than because a flag suppresses it. Surface: the full page appears **once per ENTRY into the offered state**, three dismissal levers with three scopes, and **none removes the entry point**. Q7's trap is closed — „Helyreállítási kód létrehozása" is **unavailable** while a recovery is outstanding. **WHY THE ROW STAYS FAIL: these are fixes, not a walk.** Nothing here walked a customer end to end, and this row goes green only when one completes with **no operator intervention AND a byte-identical sentinel**. Two of Phase 1's blockers are also still open (**R-214**, **R-202**), and **R-240** is untouched. Evidence: `felhom-controller/CHANGELOG.md` v0.206.0, `hub/CHANGELOG.md` v0.98.0. **DELIVERY, separately: R-239 is CLOSED 2026-08-07** — golden **0.205.0** baked, published, round-trip verified (`./etc/felhom-controller-image` read OUT of the downloaded archive) and **VOUCHED**, so a fresh install now lands on 0.205.0 carrying R-234 and R-237. In the event it was a ONE-field change: `agent_version` and `min_agent` both stayed 0.127.0. Evidence: `tests/golden-0.205.0-2026-08-07/`. **The row still says FAIL** — delivery is not a journey, and R-241 is diagnosed, not fixed. **✅ 2026-08-07 — THE FIFTH WALK PASSED, BOTH HALVES, AND THIS ROW TURNS.** A brand-new appliance (VM **325**, customer `walk5`) was installed from the published ISO on `demo-hp`, claimed, given three sentinels, escrowed, backed up off-site, then **destroyed on purpose** — guest purged, both data drives wiped — and rebuilt through the documented day-0 path. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `5b0f20f7` (`11eb7fb2…`, `6b504d1e…`, `0baaf402…`), including a 12 MB binary and `WALK5-őrszem-ékezetes-árvíztűrő.txt` **whose name BYTES are identical too**, read back with `os.listdir` on a bytes path so no decode round trip could launder a `U+FFFD`. **THE JOURNEY: PASS — ZERO guest command lines were needed to progress**, against three on the previous walk. **RTO 71.7 s** from login to an open store (12.44 s of it the unseal; ~22 s a harness retry). The recovery screen **appeared without being sought** (`/` → `/launcher` → `/recovery`), answered all three of its questions, and its sealed-at timestamp matched `host_escrow.created_at` exactly. **WHAT MADE THE DIFFERENCE — R-241's mint guard, exercised live for the first time:** at **14:58:52Z**, unaided and before anyone logged in, the rebuilt box collected its re-staged credential, configured the transport and **refused to mint a repository password** over the sealed package — `NOT minting a repository password: the hub holds a sealed recovery package for this box`. Sampled every 20 s from T0: **no key at any moment**, with the scheduler proven alive throughout. At the equivalent moment the previous walk minted one and lost the journey silently. **AND DELIVERY IS PART OF THE PASS:** the fresh install landed on **controller 0.206.0 + agent 0.127.0 — the vouched set, no hand upgrade**, both times, so the box under test is the box a customer receives. **WHAT THIS ROW STILL DOES NOT CLAIM.** (1) **Putting files back in place is built and worked here** — `reconstitute` placed 6 files — but **only after two obstacles the customer must guess past**: **R-252** (the restore refuses with „nincs elérhető adatmeghajtó" because a rebuild loses the drive *registration*, and nothing on the recovery path says to re-attach) and **R-253** (the restore refuses because the app is not installed, on a page that says three lines above that the restore reinstalls it). Both were cleared **from the dashboard with no shell** — which is why the journey passes — but neither is signposted, so *unaided* here means *possible without a shell*, not *obvious*. (2) **Shape (c) did NOT fire positively.** With the mint guard holding there is no local key, so the offer comes from **shape (a)**; shape (c) was measured in Phase A in its **negative** half (hub hash == local hash, correctly silent). The mint guard is proven positively; the discriminator only negatively. (3) **R-214, R-202 and R-240 are untouched.** (4) The venue is **torn down** (2026-08-08) — full-schema census 168 rows → 67, of which 37 are audit-by-design and **30 are R-244**, predicted before the run rather than found after; every layer verified absent against a surviving positive control, and **16.64 GiB** returned. `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. Evidence: `tests/walk5-r201-2026-08-07/journal.md` | diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index f46f5166..266b13f1 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -352,6 +352,29 @@ above stands exactly as written: the secondary unit mirror is still read by noth where the primary drive is lost — and in exactly that case the primary unit is gone while this mirror survives on the second drive, unreachable by any customer action. +**[FACT] 2026-08-23 — taking the undo copy used to DESTROY the app's own database backup (R-361).** +`writeSafetyDump` called `DumpOne` into the app's own unit directory and renamed the result to +`pre-restore-*` afterwards. `DumpOne` writes `-.sql` — the app's canonical dump, the +name the replay loop matches exactly — so **every safety dump overwrote the app's real backup and then +moved it away**, leaving the app with no database backup of its own until the next nightly run. A local +restore-from-unit in that window finds no `.sql` and tells the customer the app never had a database. +The comment beside it asserted the rename meant it *"can never overwrite the app's real dump"*; it was +false as written and stood for four months. **Measured before the fix on `demo-hp` 2026-08-22:** +`docmost` and `bookstack` each held only `pre-restore-*` files and no canonical dump. + +Fixed in controller **v0.221.0**: `DumpOneTo` takes an explicit final path and derives its own `.tmp` +from it, so neither the destination nor the scratch file can collide with a nightly dump running +beside it. **Proven the only way it can be** — the canonical dump's sha256, unchanged across a +restore, on both engines. + +**[DESIGN] 2026-08-23 — `db_dumps` lists the app's OWN dumps, not the undo copies.** They are local +material for a restore that went wrong, not part of the app's recovery set: nothing reads them from +the manifest, and three per app were being pushed off-site permanently for no recovery value. The +files are neither deleted nor hidden — their visibility is a separate recorded decision and it stands. +**One consequence, recorded because it bit within minutes:** a stable `db_dumps` lets +`CaptureRecoveryUnit`'s already-current early return fire, so anything that must happen on every +capture — such as bounding the undo copies — has to sit ABOVE that check, not after it. + **[DESIGN] 2026-08-22 — the failure ladder of a database restore: replay → rollback → hold.** Recorded here rather than only in a closed register row, because a decision that survives only inside a closed work item is a decision nobody will find. diff --git a/documentation/audits/DRILL-r361-2026-08-22/README.md b/documentation/audits/DRILL-r361-2026-08-22/README.md new file mode 100644 index 00000000..1abce466 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/README.md @@ -0,0 +1,103 @@ +# DRILL — R-361, and the two loose ends v0.220.2 left (2026-08-22 → 23) + +**Subject:** `demo-hp` (Tier 0), guest 9201. Controller **v0.220.2 → v0.221.0 → v0.221.1**. +**Method:** endpoint-level, the exact endpoints the UI's forms post to. No browser on DooPlex. +**Subjects FOUND, not rebuilt:** `docmost` (Postgres 16) and `bookstack` (MariaDB 12.3), plus +`privatebin`, `kimai`, `calibre-web`, `opengist`, `paperless-ngx`, `romm`. + +## R-361 — the whole thing, in one comparison + +The canonical dump's sha256, before and after a restore: + +| app | before | after | +|---|---|---| +| `docmost` | `5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed` | **identical** | +| `bookstack` | `7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b` | **identical** | + +**Before the fix, on the same box:** neither app had a canonical dump at all — only `pre-restore-*` +copies. Every safety dump had overwritten the app's real backup and then moved it away. + +**The dangerous lookalike, named:** a test asserting "the pre-restore file exists" passes just as well +when the app's own backup was destroyed. Only the canonical dump's **bytes** convict. + +## Part 3 — the measurement that cancelled Part 2 + +The runbook's reading was that a HELD app raises a dead-app banner and a customer e-mail. **It does +not.** Measured on the shipped v0.220.2, hold created 21:11:19Z, scans every 30 s running over it: +`docmost` aggregated to **`unhealthy`**, `IsDownState` is `{stopped, exited, degraded}`, and the +heartbeat read **`0 currently down`** at scans 600 and 620. **Part 2 was dropped in full** — 2.1, 2.2 +and 2.3 — because 2.2 and 2.3 existed only to make 2.1 safe and complete. + +**Why the reading was wrong:** its first half was right (a held app is not `StateStopped`); its second +half assumed the remaining state would be a fault. `aggregateState` checks `unhealthy > 0` **before** +the mixed-case degraded branch, and `unhealthy` is deliberately not a down state. + +**The positive control, and two live attempts that failed.** `docker stop privatebin` → `stopped`, +whitelisted by design. `docker stop bookstack-db` → `degraded` for a moment, then `unhealthy`. Neither +put a stack in a lasting down state, and an absent alarm from a detector never shown working proves +nothing. The control that works is at the layer the detector lives in: `classifyRunStates` is pure, +and it raises the banner for `degraded`/`exited` while staying silent for the states measured live. + +## Part 4 — the double-failure path, on what actually ships + +**The trigger, chosen deliberately:** the undo copy is **lost between being written and being needed** +— a drive that goes away, a filesystem that goes read-only, an external cleanup. It is realistic, it +is one of the only two ways a rollback can fail, and the product must not assume the file it wrote +twenty seconds ago is still there. + +Verified on 0.221.1: replay failed → rollback failed → **app held**; the customer's start button +refused with a reason and a route; a **full controller restart** left it stopped (`Recover()` and the +boot sweep both honoured the hold); the operator route listed it, cleared it, and after the restart +the command names, the app started. + +**The customer message, verbatim, 369 bytes:** + +> A teljes visszaállítás sikertelen: a(z) docmost adatbázisának visszaállítása sikertelen, és a korábbi +> állapot visszatöltése sem sikerült. Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai +> ne sérüljenek tovább. Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: +> pre-restore-20260822T222028Z-docmost-postgres.sql + +Hex in `16-part4-message.txt`. **No engine output** — R-381 holds. **But the last clause is false**, +and that is now **R-383**: it says the previous state's backup exists while naming the very file whose +absence caused the failure. + +**An accident worth keeping:** the first trigger attempt removed the `.tmp` instead of the final file +— a side-effect of R-361's own fix, which now writes `.tmp`. The safety dump failed and the +**fail-closed refusal fired with the app untouched**: *"a visszaállítás nem indult el"*. Not the case +being tested, but a free confirmation that no-undo-means-no-restore still holds. + +## Findings + +| id | what | +|---|---| +| **R-361** | CLOSED — shipped v0.221.0/.1, proven by the sha256 comparison above | +| **R-383** | NEW — the double-failure message names an undo copy that is not there | +| **R-384** | NEW — an app whose database has died reads `unhealthy` and raises no dead-app alarm | +| *held-app alarm* | **no row opened** — it does not fire; the negative is in the capability map | + +## Two defects found in this session's own work + +1. **A red-proof passed, twice over.** The behavioural tests inject the dump seam, so a mutation + *inside* `DumpOneTo` was invisible to them; and Part 1.3 initially had no test at all. Guards were + added at the layer each defect lives in, and both mutations then convicted. +2. **One change made another unreachable.** Excluding the undo copies from `db_dumps` made that list + stable, which let `CaptureRecoveryUnit`'s already-current early return fire — and the prune sat + after it. **Four copies on disk against a cap of three, counted on the box minutes later.** Fixed in + v0.221.1; the cap now holds at 3 on both apps, verified live. + +**And the lost-update window recorded in v0.220.2 was observed, not just reasoned:** clearing a hold +and restarting in one breath lost the clear. Doing it with a verification between the steps held. + +## Teardown — three layers + +1. **Nothing was provisioned.** Every app was already on the box; all 20 containers healthy at the end. + `docmost` still holds its 4 pages and 1 user; both canonical dumps present. +2. **No storage added.** The undo copies are capped at 3 per app and the cap was verified live. +3. **No hub-side record created.** No customer, no appliance. + +**No app is left held** — the hold created for Part 4 was cleared and the app restarted. + +## Deliberately left open + +R-102 (the Tier-2 unit mirror read by nothing) and R-359 (no readability check on the off-site store). +Untouched here. diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/00-baselines.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/00-baselines.txt new file mode 100644 index 00000000..229029ac --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/00-baselines.txt @@ -0,0 +1,27 @@ +Baselines re-verified at session start, 2026-08-22 (Phase 1, before any change). + +repo HEAD clean in-sync +felhom-controller 2024ed99826717bf765803e159c074909b28bf83 yes yes (v0.220.2) +felhom-agent 40d857b52711dd8c9c88bdd21ffcaad819a33a84 yes yes (v0.130.0) +felhom.eu a8caa0fdde7c678bb88b5b25727044695bbdeaa9 yes yes +app-catalog 459766cb16395fd1d1a66282f5cc6da59ead5924 yes yes + +Hub (/configuration, Basic auth, ClusterIP): golden_version 0.220.2, agent_version 0.130.0, +min_agent 0.129.0, global controller floor 0.220.2 — all as the runbook states. + +Register ceiling: R-382. + +SUBJECTS FOUND (not rebuilt): docmost + docmost-postgres + docmost-redis, bookstack + bookstack-db, +all healthy, up ~5 hours. Also on the box: privatebin, kimai, calibre-web, opengist, paperless-ngx, +romm. + +R-361 CONFIRMED LIVE BEFORE ANY CHANGE — the canonical dumps are absent on BOTH subjects: + docmost/db-dumps/ : only 3x pre-restore-* (no docmost-postgres.sql) + bookstack/db-dumps/ : only 3x pre-restore-* (no bookstack-mariadb.sql) +Mechanism at source: appbackup/dbdump.go:246 builds `-.sql`, :247-248 derive tmp and +final from it, and writeSafetyDump (offbox_reconstitute.go:186-196) calls DumpOne into the app's own +unit dir and only THEN renames. The comment at offbox_reconstitute.go:190-191 says the rename means it +"can never overwrite the app's real dump" — false as written. + +Dead-app watcher cadence: sched.Every("deadapp-check", 30s) at cmd/controller/main.go:726, with a 90s +boot grace. classifyRunStates at :2150. IsDownState (manager.go:54) includes StateDegraded. diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/01-part3-hold-created.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/01-part3-hold-created.txt new file mode 100644 index 00000000..b09c6c05 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/01-part3-hold-created.txt @@ -0,0 +1,7 @@ +=== the trigger: +TRIGGER: removed the freshly-written undo copy pre-restore-20260822T211114Z-docmost-postgres.sql +=== the legs: +2026/08/22 21:11:14 offbox_reconstitute.go:200: [INFO] [offbox] docmost: pre-restore safety dump written → pre-restore-20260822T211114Z-docmost-postgres.sql (138.0 KB) +2026/08/22 21:11:18 offbox_reconstitute.go:690: [ERROR] [offbox] docmost: database replay failed, rolling back to the pre-restore state: importing postgres dump for docmost: postgres import into docmost-postgres failed: exit status 3 +2026/08/22 21:11:19 offbox_reconstitute.go:695: [ERROR] [offbox] docmost: ROLLBACK ALSO FAILED (a visszavonáshoz szükséges mentés nem található (pre-restore-20260822T211114Z-docmost-postgres.sql): stat /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T211114Z-docmost-postgres.sql: no such file or directory) — holding the app stopped; replay error was: importing postgres dump for docmost: postgres import into docmost-postgres failed: exit status 3 +2026/08/22 21:11:19 offbox_handlers.go:451: [ERROR] [web] off-box reconstitute docmost (async): a(z) docmost adatbázisának visszaállítása sikertelen, és a korábbi állapot visszatöltése sem sikerült. Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai ne sérüljenek tovább. Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: pre-restore-20260822T211114Z-docmost-postgres.sql diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/02-part3-T0.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/02-part3-T0.txt new file mode 100644 index 00000000..3cf6d414 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/02-part3-T0.txt @@ -0,0 +1,7 @@ +=== T+0 immediately after the hold + now: 21:11:37Z + containers: +docmost-postgres | Up 20 seconds (healthy) + (docmost app + redis gone; the DATABASE is still up — Part 2.3's subject) + hold on disk: + ['docmost'] diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/03-observed-false-sentence.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/03-observed-false-sentence.txt new file mode 100644 index 00000000..be9305ae --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/03-observed-false-sentence.txt @@ -0,0 +1,23 @@ +OBSERVED 2026-08-22 21:11:19, on the shipped v0.220.2, while creating the hold for Part 3. + +The double-failure message ends: + + "... Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: + pre-restore-20260822T211114Z-docmost-postgres.sql" + +"a korábbi állapot mentése megvan" = "the backup of the previous state EXISTS". + +IN THIS CASE IT DOES NOT. The rollback failed precisely BECAUSE that file was gone: + + ROLLBACK ALSO FAILED (a visszavonáshoz szükséges mentés nem található + (pre-restore-20260822T211114Z-docmost-postgres.sql): stat ...: no such file or directory) + +So the customer is told, in the same sentence, that we could not restore their previous state AND +that a copy of it exists — naming a file that does not. The sentence is built from `safety`, the path +writeSafetyDump returned, without asking whether it is still there. + +This is the SAME CLASS as R-361 itself: a sentence asserting a property the code does not check. +It is not the trigger's artefact — ANY rollback failure whose cause is a missing or unreadable undo +copy produces it, and that is one of the two ways a rollback can fail. + +NOT ACTED ON in this session's Part 1/2 scope. Filed as a register row. diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/04-part3-measurement.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/04-part3-measurement.txt new file mode 100644 index 00000000..7b960a9c --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/04-part3-measurement.txt @@ -0,0 +1,18 @@ +=== PART 3 MEASUREMENT — v0.220.2, no code change + measured at: 21:13:51Z (hold created 21:11:19Z) + elapsed: ~2m20s = 5 deadapp scans at 30s cadence (heartbeat lines confirm 9 scans in the log window) + +-- 3a. WHICH RUN STATE did docmost aggregate to? + bookstack state=running deployed=True + docmost state=unhealthy deployed=True + privatebin state=running deployed=True + +-- 3b. IS THE DATABASE STILL UP? +docmost-postgres | Up 2 minutes (healthy) + +-- 3c. DID A CUSTOMER-FACING EVENT FIRE for docmost? + 2026/08/22 21:11:19 notifier.go:234: [INFO] Event pushed: backup_run_failures (error) — App "docmost" is HELD STOPPED: its off-site database restore failed AND the rollback to the customer's own pre-restore copy also failed. The app will not start from any path until the hold is cleared. Replay error: importing postgres dump for docmost: postgres import into docmost-postgres failed: exit status 3. Rollback error: a visszavonáshoz szükséges mentés nem található (pre-restore-20260822T211114Z-docmost-postgres.sql): stat /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T211114Z-docmost-postgres.sql: no such file or directory + (none above = no event fired) + +-- 3d. DEAD-APP BANNER STATE: + 2026/08/22 21:13:48 main.go:1730: [INFO] [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/05-part3-positive-control.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/05-part3-positive-control.txt new file mode 100644 index 00000000..960c0b25 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/05-part3-positive-control.txt @@ -0,0 +1,5 @@ +=== PART 3 step 4 — POSITIVE CONTROL: stop a different app OUT-OF-BAND so it is a genuine fault + (docker stop leaves the container present -> StateExited, which the code says still alerts; + a compose-down user stop would be StateStopped, which is whitelisted by design) + privatebin stopped out-of-band + at: 21:14:18Z diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt new file mode 100644 index 00000000..8a6174c7 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt @@ -0,0 +1,50 @@ +PART 3 — THE MEASUREMENT, AND THE DECISION IT PRODUCED +Measured on demo-hp, controller v0.220.2 as shipped, BEFORE any code change. + +THE READING (prompt §3): "a held app has one container still up — its database — so it does not +aggregate to StateStopped; it should aggregate to a state the down-predicate treats as a fault", and +therefore a held app gets a dead-app banner and a customer email. + +THE MEASUREMENT: IT DID NOT REPRODUCE. + + hold created 21:11:19Z (deliberate trigger, see 01) + waited 2m20s, then longer; deadapp-check runs every 30s (main.go:726) and the + scheduler log shows 9 executions in the window — the scans DID run over it + docmost's run state `unhealthy` <- NOT StateStopped, and NOT a down state + IsDownState {stopped, exited, degraded} (stacks/manager.go:54) — `unhealthy` absent + dead-app count `0 currently down`, 8 apps evaluated, at 21:13:48 and 21:23:48 + customer event NONE. The only event was this session's own operator-tier + `backup_run_failures`, which is the hold's intended notification. + database container docmost-postgres STILL UP (Part 2.3's subject — that half was right) + +WHY THE READING WAS WRONG. The first half was right: a held app is not StateStopped. The second half +assumed the remaining state would be one the predicate treats as a fault. It is not. +`aggregateState` (stacks/manager.go:~810) checks `if unhealthy > 0 → StateUnhealthy` FIRST, before the +mixed-case degraded branch. A held app's surviving database plus its failing app healthcheck put the +stack in `unhealthy`, and `unhealthy` is deliberately excluded from IsDownState. + +DECISION: PART 2 IS DROPPED IN FULL — 2.1 (suppression), 2.2 (replacement banner) and 2.3 (stopping +the database). 2.2 and 2.3 existed only to make 2.1 safe and complete; with nothing to suppress there +is nothing to replace, and stopping the database would be a behaviour change with no defect behind it. +No register row is opened for the held-app alarm, per the runbook. The negative is recorded in the +capability map instead. + +POSITIVE CONTROL — and the two live attempts that FAILED, recorded because they are the finding. + attempt 1: `docker stop privatebin` (single-container stack) -> stack read `stopped`, which is + WHITELISTED as a deliberate user stop. No alarm, correctly by design. + attempt 2: `docker stop bookstack-db` (multi-container) -> `degraded` for a moment, then `unhealthy` + once the app's own healthcheck failed. No alarm. + So NEITHER live attempt put a stack in a lasting down state, and an absent alarm from a detector + never shown working proves nothing. + THE CONTROL THAT WORKS is at the layer the detector lives in — `classifyRunStates` is a pure + function: TestClassifyRunStates_PositiveControl_ADownStackDoesAlarm feeds it `degraded` and + `exited` and both raise the banner and a Down run state; the companion test feeds it the states + MEASURED LIVE (`unhealthy`, `stopped`) and both are silent. The detector works; these states do not + reach it. + +ADJACENT FINDING, measured rather than reasoned, and OUT OF THIS TASK'S SCOPE: + `bookstack` sat with its DATABASE CONTAINER EXITED and its app `unhealthy` for several minutes and + the dead-app watcher reported `0 currently down` throughout. An app whose database has died raises + no dead-app banner and no customer email, because `unhealthy` masks the mixed state. It is not + wholly invisible — the health report counts it (`cr.Unhealthy++`, report/builder.go:251) and that + reaches the hub — but it does not alarm. Filed as a register row; NOT acted on here. diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/07-hold-clear-lost-update.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/07-hold-clear-lost-update.txt new file mode 100644 index 00000000..96d9d1d8 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/07-hold-clear-lost-update.txt @@ -0,0 +1,25 @@ +=== step 1: clear (full output, nothing filtered) + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), +[INFO] [settings] Loaded settings from /opt/docker/felhom-controller/data/settings.json +[INFO] [settings] restore hold CLEARED for docmost +[INFO] [settings] Settings saved +restore hold cleared for docmost in settings.json. + NOW RESTART THE CONTROLLER, or it will keep refusing to start the app: + systemctl restart felhom-controller-bootstrap.service + Check the app's data first — the undo copies are in its unit's db-dumps dir. +=== step 2: verify the FILE immediately, before any restart + restore_holds = None +=== step 3: restart, then verify the file AGAIN (the lost-update check) + after restart restore_holds = None diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/08-scenarioA-before.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/08-scenarioA-before.txt new file mode 100644 index 00000000..105e8f3b --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/08-scenarioA-before.txt @@ -0,0 +1,3 @@ +=== SCENARIO A — the canonical dump's sha256 BEFORE the restore + 5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/docmost-postgres.sql + 7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b /mnt/sys_drive/felhom-data/backups/primary/bookstack/db-dumps/bookstack-mariadb.sql diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/09-scenarioA-after-docmost.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/09-scenarioA-after-docmost.txt new file mode 100644 index 00000000..732169c3 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/09-scenarioA-after-docmost.txt @@ -0,0 +1,18 @@ +=== the restore that just ran: +2026/08/22 21:54:32 offbox_reconstitute.go:207: [INFO] [offbox] docmost: pre-restore safety dump written → pre-restore-20260822T215432Z-docmost-postgres.sql (138.0 KB) +2026/08/22 21:55:01 offbox_reconstitute.go:732: [INFO] [offbox] reconstituted docmost from snapshot 750b7b4d: 0 file(s) placed, 3 volume(s) replayed, 1 DB dump(s) replayed, safety dump=pre-restore-20260822T215432Z-docmost-postgres.sql, skewed=false +2026/08/22 21:55:01 offbox_handlers.go:455: [INFO] [web] off-box reconstitute docmost completed (async): files=0 dbs=1 snapshot=750b7b4d + +=== SCENARIO A — THE CANONICAL DUMP AFTER THE RESTORE + 5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/docmost-postgres.sql + BEFORE was: 5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed + +=== and the undo copy exists beside it, under its own name: + total 708 + drwxr-xr-x 2 root root 4096 Aug 22 21:54 . + drwxr-xr-x 5 root root 4096 Aug 22 21:44 .. + -rw-r--r-- 1 root root 141363 Aug 22 21:43 docmost-postgres.sql + -rw-r--r-- 1 root root 141363 Aug 22 16:05 pre-restore-20260822T160544Z-docmost-postgres.sql + -rw-r--r-- 1 root root 141363 Aug 22 16:23 pre-restore-20260822T162347Z-docmost-postgres.sql + -rw-r--r-- 1 root root 141363 Aug 22 16:27 pre-restore-20260822T162708Z-docmost-postgres.sql + -rw-r--r-- 1 root root 141363 Aug 22 21:54 pre-restore-20260822T215432Z-docmost-postgres.sql diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/10-scenarioA-after-bookstack.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/10-scenarioA-after-bookstack.txt new file mode 100644 index 00000000..685970cb --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/10-scenarioA-after-bookstack.txt @@ -0,0 +1,4 @@ +=== SCENARIO A — bookstack (MariaDB) +2026/08/22 21:56:07 offbox_reconstitute.go:732: [INFO] [offbox] reconstituted bookstack from snapshot 245932e4: 0 file(s) placed, 2 volume(s) replayed, 1 DB dump(s) replayed, safety dump=pre-restore-20260822T215540Z-bookstack-mariadb.sql, skewed=false + BEFORE: 7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b + AFTER : 7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b /mnt/sys_drive/felhom-data/backups/primary/bookstack/db-dumps/bookstack-mariadb.sql diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/11-manifest.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/11-manifest.txt new file mode 100644 index 00000000..5f38df34 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/11-manifest.txt @@ -0,0 +1,9 @@ +=== the MANIFEST after these restores (Part 1.3 — must list only the app's own dump) +-- docmost + db_dumps : ['docmost-postgres.sql'] + vol_dumps : ['docmost_docmost_postgres_data.tar', 'docmost_docmost_redis_data.tar', 'docmost_docmost_storage.tar'] + undo copies ON DISK (positive control for the exclusion): 4 +-- bookstack + db_dumps : ['bookstack-mariadb.sql'] + vol_dumps : ['bookstack_bookstack_config.tar', 'bookstack_bookstack_db_data.tar'] + undo copies ON DISK (positive control for the exclusion): 4 diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/12-prune-cap-live.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/12-prune-cap-live.txt new file mode 100644 index 00000000..61a37ff3 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/12-prune-cap-live.txt @@ -0,0 +1,3 @@ +=== the cap, live on 0.221.1 (was 4 with a cap of 3 before the follow-on fix) + docmost undo copies: 3 + bookstack undo copies: 3 diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/13-scenarioE.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/13-scenarioE.txt new file mode 100644 index 00000000..ec4e8537 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/13-scenarioE.txt @@ -0,0 +1,9 @@ +=== SCENARIO E — a genuinely dead app must still alarm, byte-identical to v0.220.2 +Part 2 was DROPPED, so no suppression was added. The proof is that the classifier is untouched: + +-- git diff of classifyRunStates and its caller, v0.220.2 (2024ed99) -> HEAD: + (empty above = main.go was not changed at all this session) + +-- the only cmd/controller change is a NEW TEST FILE: + .../cmd/controller/r361_classifier_control_test.go | 55 ++++++++++++++++++++++ + 1 file changed, 55 insertions(+) diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/14-part4-hold.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/14-part4-hold.txt new file mode 100644 index 00000000..abb2177e --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/14-part4-hold.txt @@ -0,0 +1,9 @@ +=== PART 4 — double failure on the SHIPPED 0.221.1 +TRIGGER (deliberate, realistic): the undo copy is LOST between being written and being needed — + a drive that goes away, a filesystem that goes read-only, or an external cleanup. The product + must not assume the file it wrote twenty seconds ago is still there. + TRIGGER: removed the freshly-written undo copy pre-restore-20260822T221623Z-docmost-postgres.sql.tmp +=== app state — must be NOT running: +docmost | Up 12 minutes (healthy) +docmost-postgres | Up 12 minutes (healthy) +docmost-redis | Up 12 minutes (healthy) diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/15-part4-refusals.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/15-part4-refusals.txt new file mode 100644 index 00000000..e7f546e5 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/15-part4-refusals.txt @@ -0,0 +1,7 @@ +=== PART 4 on 0.221.1 — every start path +-- app state (must be NOT running): +docmost-postgres | Up 23 seconds (healthy) +-- the CUSTOMER's start button: + {"ok":false,"error":"a(z) docmost adatainak visszaállítása 2026-08-22 22:17-kor megszakadt, és a korábbi állapotot sem sikerült visszatölteni. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek tovább. Vedd fel velünk a kapcsolatot"} +-- a FULL controller restart, then Recover() and the boot sweep: +docmost-postgres | Up 37 seconds (healthy) diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/16-part4-message.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/16-part4-message.txt new file mode 100644 index 00000000..43372c72 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/16-part4-message.txt @@ -0,0 +1,15 @@ +=== CUSTOMER MESSAGE, VERBATIM — double failure, controller 0.221.1, read from the page === +A teljes visszaállítás sikertelen: a(z) docmost adatbázisának visszaállítása sikertelen, és a korábbi állapot visszatöltése sem sikerült. Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai ne sérüljenek tovább. Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: pre-restore-20260822T222028Z-docmost-postgres.sql + +=== UTF-8 hex === +412074656c6a657320766973737a61c3a16c6cc3ad74c3a1732073696b657274656c656e3a2061287a2920646f636d6f7374206164617462c3a17a6973c3a16e616b20766973737a61c3a16c6cc3ad74c3a173612073696b657274656c656e2c20c3a9732061206b6f72c3a162626920c3a16c6c61706f7420766973737a6174c3b66c74c3a973652073656d2073696b6572c3bc6c742e20417a20616c6b616c6d617ac3a173742062697a746f6e73c3a16769206f6b62c3b36c204c45c3814c4cc38d545641206861677974756b2c20686f677920617a20616461746169206e652073c3a972c3bc6c6a656e656b20746f76c3a162622e20566564642066656c2076656cc3bc6e6b2061206b617063736f6c61746f7420e280942061206b6f72c3a162626920c3a16c6c61706f74206d656e74c3a97365206d656776616e3a207072652d726573746f72652d3230323630383232543232323032385a2d646f636d6f73742d706f7374677265732e73716c + +=== byte length: 369 + +=== R-381 must hold — engine-output leak check: + ERROR: present: False + LINE 1: present: False + COPY public. present: False + exit status present: False + psql present: False + INSERT INTO present: False diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/17-part4-operator-route.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/17-part4-operator-route.txt new file mode 100644 index 00000000..3aaf5838 --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/17-part4-operator-route.txt @@ -0,0 +1,12 @@ +=== PART 4 — the operator route, end to end on 0.221.1 +-- list: + docmost held since 2026-08-22T22:20:32Z + replay error : importing postgres dump for docmost: postgres import into docmost-postgres failed: exit status 3 + rollback err : a visszavonáshoz szükséges mentés nem található (pre-restore-20260822T222028Z-docmost-postgres.sql): stat /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T222028Z-docmost-postgres.sql: no such file or directory +-- clear: + restore hold cleared for docmost in settings.json. + NOW RESTART THE CONTROLLER, or it will keep refusing to start the app: + systemctl restart felhom-controller-bootstrap.service + Check the app's data first — the undo copies are in its unit's db-dumps dir. +-- after the restart the command names, the app starts: + {"ok":true,"message":"Stack docmost start completed"} diff --git a/documentation/audits/DRILL-r361-2026-08-22/evidence/18-teardown.txt b/documentation/audits/DRILL-r361-2026-08-22/evidence/18-teardown.txt new file mode 100644 index 00000000..b3b00f9c --- /dev/null +++ b/documentation/audits/DRILL-r361-2026-08-22/evidence/18-teardown.txt @@ -0,0 +1,29 @@ +=== TEARDOWN CHECK — 2026-08-23 +-- holds in force (must be none): + None +-- every app on the box: +bookstack | Up 17 minutes (healthy) +bookstack-db | Up 17 minutes (healthy) +calibre-web | Up 17 minutes (healthy) +cloudflared | Up 30 hours +docmost | Up 15 seconds (healthy) +docmost-postgres | Up About a minute (healthy) +docmost-redis | Up 26 seconds (healthy) +felhom-controller | Up 37 seconds (healthy) +filebrowser | Up 30 hours (healthy) +kimai | Up 17 minutes (healthy) +kimai-db | Up 17 minutes (healthy) +opengist | Up 17 minutes (healthy) +paperless-postgres | Up 17 minutes (healthy) +paperless-redis | Up 17 minutes (healthy) +paperless-webserver | Up 16 minutes (healthy) +privatebin | Up 16 minutes (healthy) +romm | Up 16 minutes (healthy) +romm-db | Up 16 minutes (healthy) +romm-redis | Up 16 minutes (healthy) +traefik | Up 30 hours +-- docmost's data survived the whole session: + 4 pages, 1 users +-- canonical dumps present on both subjects: + docmost docmost-postgres.sql + bookstack bookstack-mariadb.sql diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 60330692..7c57324c 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,7 @@ --- +| **R-361** | **The pre-restore safety dump overwrote the app's own DB dump, and the comment beside it said it could not.** Shipped in controller v0.221.0 (+v0.221.1). Evidence: `audits/DRILL-r361-2026-08-22/evidence/`. **Reasoning kept:** *`DumpOne` writes `-.sql` — the app's canonical dump, the name the replay loop matches EXACTLY — so nothing else may ever be written to it.* The fix is a DESTINATION, not a rename: `DumpOneTo` takes the final path and derives its own `.tmp` from it, so neither the destination nor the scratch file can collide with a nightly dump running beside it. **`DumpOne`'s signature did not move** — it has callers outside this concern. **The manifest no longer lists the undo copies:** every consumer of `Manifest.DBDumps` was grepped and named — three, all inside `recovery_unit.go`, none reading it for recovery. **AND THAT CHANGE MADE ANOTHER UNREACHABLE:** a stable `db_dumps` let `CaptureRecoveryUnit`'s already-current early return fire, and the undo-copy prune sat after it — four copies on disk against a cap of three, counted live. The prune now runs ABOVE the check; it is housekeeping on the dump directory and is independent of whether the manifest needs rewriting. **PROVEN LIVE the only way it can be:** the canonical dump's sha256, unchanged across a restore — `docmost` `5d35678349bb…`, `bookstack` `7837aa5de295…`, both byte-identical before and after. A test asserting merely that the undo copy exists passes just as well when the app's backup was destroyed. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.221.1, 2026-08-23) | full text: `git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md` | | **R-379** | **The pre-restore undo copy was valid, was named to the customer, and no product action could apply it.** Shipped in controller v0.220.0 (+v0.220.1, v0.220.2). Evidence: `audits/DRILL-r379-rollback-2026-08-22/evidence/`. **Reasoning kept:** *R-379 and R-380 were ONE failure with ONE fix — both ended with a half-restored database and the only difference was whether it looked broken.* **The undo set is matched on THE RUN'S OWN STAMP, never on the `pre-restore-` prefix** (four copies coexisted on one app in one afternoon; a prefix match replays an arbitrary older state) **and never just the first file** (a two-database app would have had one restored and the other left half-written). **The rollback RE-DISCOVERS the container** — the undo file is stable, the container is not: the DB-only start re-creates it, and v0.220.0's own first live run held an app for 30 s of `waitDBReady` against a dead id while its data was recoverable. **No unit test saw that: they all inject the import seam and never look at container identity.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.1, 2026-08-22; docmost and bookstack both rolled back to byte-identical prior state) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` | | **R-380** | **A failed MariaDB replay left a partially-applied database behind an app reporting `health=healthy`.** Shipped in controller v0.220.0. Evidence: `audits/DRILL-r379-rollback-2026-08-22/evidence/13-step2-verify.txt`. **Reasoning kept:** **no engine flag closes this** — `--single-transaction` was added to the Postgres import and does make it all-or-nothing, but **MariaDB's DDL is not transactional**, so a partial apply there is unavoidable at the engine. The flag is a belt; the rollback is the fix, and this row must not be read as saying otherwise. Proven live: `bookstack`'s `migrations` table back at **102 rows**, the exact cell the defect was measured in. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` | | **R-381** | **The restore-failure message pasted raw engine stderr — including rows out of the customer's own database — into the Hungarian customer surface.** Shipped in controller v0.220.0. **Reasoning kept:** the full engine text now goes to the operator log, **which never had it before — the diagnostic was ADDED, not removed**. Measured: 407 bytes (Postgres) and 615 (MariaDB, whose middle was an `INSERT INTO migrations VALUES (…)` listing); now 257 bytes with no engine tokens. **A red-proof for this PASSED and the test was hollow**: it injected below `ImportDump`, so a leak reintroduced inside `ImportDump` could not fail it. The guard now sits at that layer. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index e3d7ef6a..d2db78ad 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -136,6 +136,8 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour | ID | What | State | |---|---|---| +| **R-383** | **The double-failure message tells the customer their previous state was saved, and names a file that is not there.** The sentence ends *"a korábbi állapot mentése megvan: "* — "the backup of the previous state EXISTS" — built from the path `writeSafetyDump` returned, WITHOUT asking whether it is still on disk. But one of the two ways a rollback can fail is that the undo copy is missing or unreadable, and in exactly that case the sentence is FALSE. **Measured live twice, on v0.220.2 (2026-08-22 21:11:19) and again on v0.221.1 (22:20:32):** rollback failed with `stat …pre-restore-…sql: no such file or directory`, and the customer message named that same file as existing. 369 bytes, `offbox_reconstitute.go` (the double-failure branch). **This is R-361's own class** — a sentence asserting a property the code does not check — one surface over. | **OPEN — MEDIUM** | — | Say what is true: name the undo copy only when it is verifiably on disk, and say plainly when it is not. **Do not simply drop the filename** — an operator needs it, and R-351's lesson was that a refusal which names nothing forces someone to remember what the product already knew. Evidence: `audits/DRILL-r361-2026-08-22/evidence/03-observed-false-sentence.txt`, `audits/DRILL-r361-2026-08-22/evidence/16-part4-message.txt`. | CC | +| **R-384** | **An app whose DATABASE has died raises no dead-app alarm — `unhealthy` masks the mixed state.** `aggregateState` (`internal/stacks/manager.go`) checks `if unhealthy > 0 → StateUnhealthy` BEFORE the mixed-case degraded branch, and `IsDownState` (`manager.go:54`) is `{stopped, exited, degraded}` — `unhealthy` is absent. So a multi-container app whose database container dies goes `degraded` for a moment and then `unhealthy` as its own healthcheck fails, and stops being a fault. **Measured live 2026-08-22:** `bookstack-db` stopped out-of-band at 21:27:01; `bookstack` read `unhealthy`; the dead-app heartbeat reported **`0 currently down`** across the whole window (scans 600 and 620), with 8 apps evaluated. **NOT invisible everywhere** — the health report counts it (`cr.Unhealthy++`, `internal/report/builder.go:251`) and that reaches the hub — but it raises no banner and no customer e-mail. **This is the F-CRIT-1 class the `classifyRunStates` comment says was closed:** it WAS closed for `StateStopped`+failedRestart, and `unhealthy` was never in scope. Found while building a positive control for a different question. | **OPEN — MEDIUM** | — | Decide whether a SUSTAINED `unhealthy` is a fault (it is not a brief one — that is why it is excluded), on the `crashLoopAfter` model: a threshold above the deploy/health windows rather than a state test. **Do not simply add `unhealthy` to `IsDownState`** — it has other callers and would alarm on every deploy, which is the over-correction F-A1 nearly cost. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt`. | CC | | **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor | | **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | | **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor | @@ -529,7 +531,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-358** | **A FAILED scratch restore leaves a partial copy that the product then offers as a full restore source — and the destructive restore runs from it and reports success.** `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only whether the directory exists and is non-empty; its comment defers completeness to `PlaceOffsiteRestore`, which stats top-level placements, not files. Proven 2026-08-21 22:54-22:56 against a deliberately corrupted store: the restore failed honestly (`ciphertext verification failed`, 54 files, 15 of 16 originals), the wizard then offered „Teljes visszaállítás indítása", and pressing it reported `ok=true`. **The failure is detected and then forgotten.** | **OPEN — MEDIUM** | — | Record the failure against the scratch and refuse to place from it until it is re-prepared. | CC | | **R-359** | **The off-site restic store is never verified by anything, ever.** The complete set of restic verbs in the controller is `restore, snapshots, backup, unlock, stats, init, forget, prune, cat` — **no `check`**. The agent's `RestoreTest` is PBS-tier only. Established 2026-08-21 by deliberately corrupting one pack: `restic check` catches it immediately („ciphertext verification failed", „Fatal: repository contains errors"), and the product only meets the damage when a customer is already trying to recover. | **OPEN — MEDIUM** | — | A periodic `restic check` (structure) with an occasional `--read-data`, reported like any other backup verdict. Note PBS already has verify jobs; this is the tier that does not. | CC | | **R-360** | **The verification-copy delete gates on the concurrency flag, which the verification restore never holds — so the copy is deletable for the whole restore, and the handler's own comment claims the opposite.** `offboxVerifyCopyDeleteHandler` (`web/offbox_handlers.go:502`) reads `backupMgr.IsRunning()`; `RestoreOffboxScratch` (`offbox_restore.go:211`) **never calls `acquireRunning`**. R-351b moved all seven restore handlers onto `restoreOpBlocked()` (both flags) and left this one behind. Demonstrated 2026-08-21 22:35 with the flags read immediately before and after: `display=True offbox-restore kimai / concurrency=False` on both sides, and the delete succeeded. The guard does not compare app names, so the same call naming the restoring app removes the directory the restore is writing into. **Observed in two consecutive reports and filed neither time; filed now.** | **OPEN — MEDIUM** | — | `restoreOpBlocked()`, and a test that asserts the CONSEQUENCE — the copy survives a delete attempt mid-restore. | CC | -| **R-361** | **The pre-restore safety dump overwrites the app's own DB dump, and the comment beside it says it cannot.** `writeSafetyDump` calls `DumpOne`, which writes the canonical `-.sql` (`appbackup/dbdump.go:200-202`) — i.e. the unit's real dump — and only THEN renames it to `pre-restore-*`. The comment at `offbox_reconstitute.go:147-148` states the rename means it "can never overwrite the app's real dump". Proven 2026-08-21: `romm`'s `db-dumps/` held `romm-mariadb.sql` (62 270 B) at 22:59 and held ONLY `pre-restore-20260821T210246Z-romm-mariadb.sql` after one reconstitute. Until the next backup run the local restore-from-unit finds no `.sql` and reports the app has no database. The `pre-restore-*` file is also enumerated into the manifest's `db_dumps` and shipped off-site. | **OPEN — MEDIUM** | — | Dump to the safety name directly, or to a temp name. Pin the invariant the comment already asserts. | CC | | **R-362** | **A data drive detached mid-restore is reported as „permission denied".** Observed 2026-08-21 23:15: the guest-visible bind was unmounted 4 s into a scratch restore; the restore correctly failed and wrote nothing to the wrong place, but said „A visszaállítás sikertelen: restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied". The controller has a drive-state concept (`IsDisconnected`, used by both backup legs) and the restore path never consults it. **A correct refusal that misdescribes why sends the reader at a permissions problem that does not exist.** Creditable in the same test: the agent re-bound the drive 5 s later, unaided. | **OPEN — MEDIUM** | — | Consult drive state when a restore path operation fails on ENOENT/EACCES and name the drive. | CC | | **R-363** | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** `sched.Daily("fill-watch", "03:30", …)` (`cmd/controller/main.go:1092`) plus one startup check. Proven 2026-08-21 23:17: the 69 GB filesystem carrying the Docker data-root, the system namespace and ALL 40-class app data was filled to 99% / 1.2 GiB free; the backup reserve refused `kimai` per app and the hub received `recovery_unit_capture_failed` (error) naming the filesystem, **and the fill watcher said nothing at all**. The package comment says it "warns the CUSTOMER that a filesystem is filling, BEFORE anything fails"; at a daily cadence it frequently cannot. | **OPEN — MEDIUM** | — | The reserve already computes the same numbers every run. Let the watcher share that reading rather than owning a separate daily one. | CC | | **R-364** | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC | diff --git a/documentation/tests/golden-0.221.1-2026-08-23/README.md b/documentation/tests/golden-0.221.1-2026-08-23/README.md new file mode 100644 index 00000000..7277846c --- /dev/null +++ b/documentation/tests/golden-0.221.1-2026-08-23/README.md @@ -0,0 +1,65 @@ +# Golden bake — 0.221.1 (2026-08-23) + +Baked in the drill VM on DooPlex per `documentation/runbooks/RUNBOOK-manual-build.md` §4.0/§4.1, +carrying controller **v0.221.1** (R-361 — the safety dump no longer destroys the app's own database +backup, plus the follow-on that keeps the undo-copy cap reachable). + +| | | +|---|---| +| `GOLDEN_VERSION` | **0.221.1** | +| `GOLDEN_SHA256` | **1c8bf6cf08cadabeca6331f38360d905e867c235067cd10c2716915b6e6df089** | +| package | `https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.221.1/golden.tar.zst` | +| size | 656 966 079 B | +| controller image | `gitea.dooplex.hu/admin/felhom-controller:0.221.1` | +| template | `debian-13-standard_13.6-1_amd64.tar.zst` (**listed fresh**, checksum verified on download) | +| `build-golden.sh` | v3.0.0, from `felhom-agent` @ `40d857b52711` | +| `MinAgent` | **0.129.0** (from the controller CHANGELOG header — unchanged) | + +## Why this bake happened in this session + +The `golden-currency` gate refused the docs push: 0.221.1 was released with no golden carrying it. +**That block is not circular** — a golden needs the controller image, which was already built and +pushed, not the docs commit — so the gate was satisfied by doing the work it asked for rather than +bypassed. **No push in this session used `--no-verify`.** + +## Pass markers — each checked, with the negative controls + +``` +docker OK (overlay2 : 1 -> " docker OK (overlay2; data-root /var/lib/docker)" +including mount point : 2 -> rootfs ('/') and mp0 ('/var/lib/felhom') [there is no mp1] +upload OK (HTTP 201) : 1 -> pre-delete returned HTTP 404 (404/204 expected) +excluding : 0 <- negative control +FATAL : 0 <- negative control +``` + +**The 404 pre-gate was controlled before it was believed:** the same URL shape for **0.220.2 returned +HTTP 200** in the same minute, so the 404 on 0.221.1 means absent, not a wrong URL. + +## Verified by ROUND TRIP + +Downloaded again — **HTTP 200, 656 966 079 bytes** — and the sha256 recomputed: `1c8bf6cf08ca…` on +both sides. What a machine receives is byte-identical to what was baked. + +## Token hygiene + +Copied **file → file**, read by a runner script inside the VM. +`systemctl show golden-bake -p Environment -p ExecStart | grep -c -F "$(cat /root/.gitea-token)"` → **0**, +and that grep was **proved able to convict first** (planted copy → **1**, copy shredded). The same +control was run on the **committed** `bake.log`: **0**, positive control **1**. + +## Teardown + +`pct destroy 9100 --purge`; token, runner, build script and in-VM log `shred -u`'d **after** +`bake.log` was copied out — all four confirmed absent; `poweroff`; waited for qemu to exit; +`qemu-img snapshot -a virgin`. + +## NOT vouched + +The hub's Day-0 artifact manifest was **not** changed and the floor was **not** raised. Both are the +operator's, and it is a **three-field** save: + +| field | value | +|---|---| +| `golden_version` | **0.221.1** | +| `agent_version` | **0.130.0** | +| `min_agent` | **0.129.0** | diff --git a/documentation/tests/golden-0.221.1-2026-08-23/bake.log b/documentation/tests/golden-0.221.1-2026-08-23/bake.log new file mode 100644 index 00000000..3a9807cf --- /dev/null +++ b/documentation/tests/golden-0.221.1-2026-08-23/bake.log @@ -0,0 +1,326 @@ +[golden] build-golden.sh v3.0.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.221.1 +[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) … + Logical volume "vm-9100-disk-0" created. + Logical volume pve/vm-9100-disk-0 changed. +Creating filesystem with 8388608 4k blocks and 2097152 inodes +Filesystem UUID: e16a42b3-de76-4000-91e6-caf9f09113ac +Superblock backups stored on blocks: + 32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208, + 4096000, 7962624 + Logical volume "vm-9100-disk-1" created. + Logical volume pve/vm-9100-disk-1 changed. +Creating filesystem with 6291456 4k blocks and 1572864 inodes +Filesystem UUID: 9da5bfbf-649a-43a1-8dd6-a60a09c0a484 +Superblock backups stored on blocks: + 32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208, +extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst' +Total bytes read: 553512960 (528MiB, 119MiB/s) +Detected container architecture: amd64 +Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ... +done: SHA256:XFxzbJ3tcRoO0GHqqMVUbizdKJ+Uo51heqZbzwWXIRo root@felhom-golden +Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ... +done: SHA256:xV5uo5d4abiH7PncNubY/MNlJRVVt6+O58UAR7IPxsE root@felhom-golden +Creating SSH host key 'ssh_host_rsa_key' - this may take some time ... +done: SHA256:M4mkSWVcp0gQcNWTbZtdJ5mQ2S4eJoyfG0FLsBwN/MM root@felhom-golden +[golden] starting + installing Docker (official repo, trixie channel) … +apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct! +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = (unset), + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to the standard locale ("C"). +locale: Cannot set LC_CTYPE to default locale: No such file or directory +locale: Cannot set LC_MESSAGES to default locale: No such file or directory +locale: Cannot set LC_ALL to default locale: No such file or directory +apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct! +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = (unset), + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to the standard locale ("C"). +locale: Cannot set LC_CTYPE to default locale: No such file or directory +locale: Cannot set LC_MESSAGES to default locale: No such file or directory +locale: Cannot set LC_ALL to default locale: No such file or directory +[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation … +[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds … +[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) … +Unable to find image 'hello-world:latest' locally +latest: Pulling from library/hello-world +4f55086f7dd0: Pulling fs layer +4f55086f7dd0: Verifying Checksum +4f55086f7dd0: Download complete +4f55086f7dd0: Pull complete +Digest: sha256:5dd0d3e6e255913fc30f90b9f2b1d359cc2cbdb48090cc4b65f1676e203243cc +Status: Downloaded newer image for hello-world:latest + docker OK (overlay2; data-root /var/lib/docker) + /var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4 + /mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4 + both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576 +[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.221.1 (no registry cred at deploy) … + +WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'. +Configure a credential helper to remove this warning. See +https://docs.docker.com/go/credential-store/ + +0.221.1: Pulling from admin/felhom-controller +039e6f9f9752: Pulling fs layer +0094c3ac0914: Pulling fs layer +deca1dac7403: Pulling fs layer +11c19a33d1b8: Pulling fs layer +449fade02d9e: Pulling fs layer +6eec5b6da5a5: Pulling fs layer +11c19a33d1b8: Waiting +449fade02d9e: Waiting +6eec5b6da5a5: Waiting +deca1dac7403: Verifying Checksum +deca1dac7403: Download complete +11c19a33d1b8: Verifying Checksum +11c19a33d1b8: Download complete +449fade02d9e: Verifying Checksum +449fade02d9e: Download complete +6eec5b6da5a5: Verifying Checksum +6eec5b6da5a5: Download complete +0094c3ac0914: Verifying Checksum +0094c3ac0914: Download complete +039e6f9f9752: Verifying Checksum +039e6f9f9752: Download complete +039e6f9f9752: Pull complete +0094c3ac0914: Pull complete +deca1dac7403: Pull complete +11c19a33d1b8: Pull complete +449fade02d9e: Pull complete +6eec5b6da5a5: Pull complete +Digest: sha256:9dab837d1b538b283b88dae0716f035d9fc54a0743f19ff55e4f4f79bd6efa95 +Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.221.1 +gitea.dooplex.hu/admin/felhom-controller:0.221.1 +[golden] asking the controller which infra images it manages … +[golden] baking infra images (4): traefik:v3.6.7 cloudflare/cloudflared:2026.6.0 gtstef/filebrowser:1.3.3-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 … +v3.6.7: Pulling from library/traefik +589002ba0eae: Pulling fs layer +ef63511ea6cc: Pulling fs layer +0738e5cb835e: Pulling fs layer +3e6813f70c64: Pulling fs layer +3e6813f70c64: Waiting +589002ba0eae: Verifying Checksum +589002ba0eae: Download complete +ef63511ea6cc: Verifying Checksum +ef63511ea6cc: Download complete +3e6813f70c64: Verifying Checksum +3e6813f70c64: Download complete +589002ba0eae: Pull complete +0738e5cb835e: Verifying Checksum +0738e5cb835e: Download complete +ef63511ea6cc: Pull complete +0738e5cb835e: Pull complete +3e6813f70c64: Pull complete +Digest: sha256:a9890c898f379c1905ee5b28342f6b408dc863f08db2dab20e46c267d1ff463a +Status: Downloaded newer image for traefik:v3.6.7 +docker.io/library/traefik:v3.6.7 +2026.6.0: Pulling from cloudflare/cloudflared +47de5dd0b812: Pulling fs layer +c172f21841df: Pulling fs layer +99515e7b4d35: Pulling fs layer +99ba982a9142: Pulling fs layer +d6b1b89eccac: Pulling fs layer +2780920e5dbf: Pulling fs layer +7c12895b777b: Pulling fs layer +3214acf345c0: Pulling fs layer +52630fc75a18: Pulling fs layer +dd64bf2dd177: Pulling fs layer +b839dfae01f6: Pulling fs layer +ebddc55facdc: Pulling fs layer +bdfd7f7e5bf6: Pulling fs layer +2d4d7adf6272: Pulling fs layer +40008157d8d2: Pulling fs layer +bd8962e29291: Pulling fs layer +cac2ae0193cb: Pulling fs layer +74d1dac84ecc: Pulling fs layer +3214acf345c0: Waiting +52630fc75a18: Waiting +dd64bf2dd177: Waiting +b839dfae01f6: Waiting +ebddc55facdc: Waiting +bdfd7f7e5bf6: Waiting +2d4d7adf6272: Waiting +40008157d8d2: Waiting +bd8962e29291: Waiting +cac2ae0193cb: Waiting +74d1dac84ecc: Waiting +99ba982a9142: Waiting +d6b1b89eccac: Waiting +2780920e5dbf: Waiting +7c12895b777b: Waiting +47de5dd0b812: Download complete +c172f21841df: Verifying Checksum +c172f21841df: Download complete +99515e7b4d35: Verifying Checksum +99515e7b4d35: Download complete +99ba982a9142: Verifying Checksum +99ba982a9142: Download complete +d6b1b89eccac: Verifying Checksum +d6b1b89eccac: Download complete +2780920e5dbf: Download complete +47de5dd0b812: Pull complete +7c12895b777b: Verifying Checksum +7c12895b777b: Download complete +3214acf345c0: Verifying Checksum +3214acf345c0: Download complete +52630fc75a18: Verifying Checksum +52630fc75a18: Download complete +dd64bf2dd177: Download complete +b839dfae01f6: Verifying Checksum +b839dfae01f6: Download complete +ebddc55facdc: Verifying Checksum +ebddc55facdc: Download complete +c172f21841df: Pull complete +bdfd7f7e5bf6: Verifying Checksum +bdfd7f7e5bf6: Download complete +2d4d7adf6272: Verifying Checksum +2d4d7adf6272: Download complete +bd8962e29291: Verifying Checksum +bd8962e29291: Download complete +40008157d8d2: Verifying Checksum +40008157d8d2: Download complete +cac2ae0193cb: Verifying Checksum +cac2ae0193cb: Download complete +74d1dac84ecc: Verifying Checksum +74d1dac84ecc: Download complete +99515e7b4d35: Pull complete +99ba982a9142: Pull complete +d6b1b89eccac: Pull complete +2780920e5dbf: Pull complete +7c12895b777b: Pull complete +3214acf345c0: Pull complete +52630fc75a18: Pull complete +dd64bf2dd177: Pull complete +b839dfae01f6: Pull complete +ebddc55facdc: Pull complete +bdfd7f7e5bf6: Pull complete +2d4d7adf6272: Pull complete +40008157d8d2: Pull complete +bd8962e29291: Pull complete +cac2ae0193cb: Pull complete +74d1dac84ecc: Pull complete +Digest: sha256:ba461b8aa9c042156dbd39c38657fe7431bafa063220eab8d5330a523863da9f +Status: Downloaded newer image for cloudflare/cloudflared:2026.6.0 +docker.io/cloudflare/cloudflared:2026.6.0 +1.3.3-stable: Pulling from gtstef/filebrowser +6a0ac1617861: Pulling fs layer +ef8806083e82: Pulling fs layer +b74107c861c7: Pulling fs layer +adc935def003: Pulling fs layer +4f4fb700ef54: Pulling fs layer +18695ccc900a: Pulling fs layer +45d119d5c397: Pulling fs layer +dac52db4fc51: Pulling fs layer +6d598f86b2f2: Pulling fs layer +8aa349c8396c: Pulling fs layer +adc935def003: Waiting +4f4fb700ef54: Waiting +18695ccc900a: Waiting +45d119d5c397: Waiting +dac52db4fc51: Waiting +6d598f86b2f2: Waiting +8aa349c8396c: Waiting +b74107c861c7: Verifying Checksum +b74107c861c7: Download complete +6a0ac1617861: Verifying Checksum +6a0ac1617861: Download complete +adc935def003: Verifying Checksum +adc935def003: Download complete +4f4fb700ef54: Verifying Checksum +4f4fb700ef54: Download complete +45d119d5c397: Verifying Checksum +45d119d5c397: Download complete +dac52db4fc51: Verifying Checksum +dac52db4fc51: Download complete +6a0ac1617861: Pull complete +ef8806083e82: Verifying Checksum +ef8806083e82: Download complete +18695ccc900a: Verifying Checksum +18695ccc900a: Download complete +6d598f86b2f2: Verifying Checksum +6d598f86b2f2: Download complete +8aa349c8396c: Verifying Checksum +8aa349c8396c: Download complete +ef8806083e82: Pull complete +b74107c861c7: Pull complete +adc935def003: Pull complete +4f4fb700ef54: Pull complete +18695ccc900a: Pull complete +45d119d5c397: Pull complete +dac52db4fc51: Pull complete +6d598f86b2f2: Pull complete +8aa349c8396c: Pull complete +Digest: sha256:eb3733681db8757412632c61a99ad656f0d94ed6781bb2ea114b4d70babab78c +Status: Downloaded newer image for gtstef/filebrowser:1.3.3-stable +docker.io/gtstef/filebrowser:1.3.3-stable +1.1.0: Pulling from admin/felhom-samba +897d797d2723: Pulling fs layer +3051591aa250: Pulling fs layer +ce57a3f93416: Pulling fs layer +fb94eeec2fe1: Pulling fs layer +fb94eeec2fe1: Waiting +ce57a3f93416: Verifying Checksum +ce57a3f93416: Download complete +fb94eeec2fe1: Verifying Checksum +fb94eeec2fe1: Download complete +897d797d2723: Verifying Checksum +897d797d2723: Download complete +3051591aa250: Verifying Checksum +3051591aa250: Download complete +897d797d2723: Pull complete +3051591aa250: Pull complete +ce57a3f93416: Pull complete +fb94eeec2fe1: Pull complete +Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10 +Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0 +gitea.dooplex.hu/admin/felhom-samba:1.1.0 +[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'. +[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'. +[golden] baking the first-boot SSH host-key regeneration unit (F3) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'. +[golden] identity-clean + minimize … +[golden] stop + archive … +INFO: including mount point rootfs ('/') in backup +INFO: including mount point mp0 ('/var/lib/felhom') in backup +INFO: archive file size: 626MB +INFO: Finished Backup of VM 9100 (00:00:32) +[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_08_23-00_29_27.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive) +[golden] publishing golden (656966079 bytes, sha256 1c8bf6cf08cadabe…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.221.1/golden.tar.zst +[golden] pre-delete existing: HTTP 404 (404/204 expected) +[golden] upload OK (HTTP 201) +GOLDEN_VERSION=0.221.1 +GOLDEN_SHA256=1c8bf6cf08cadabeca6331f38360d905e867c235067cd10c2716915b6e6df089 +[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.221.1 / 1c8bf6cf08cadabeca6331f38360d905e867c235067cd10c2716915b6e6df089 +[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)