R-379/R-380 docs: the failure ladder, the drill record, register housekeeping
gates / gates (push) Successful in 17s

07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay ->
rollback -> hold, including why no engine flag closes it: --single-transaction
makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the
fix and the flag is a belt.

Drill record for the live walk, including the TWO defects the walk found in the
fix itself (a rollback into a re-created container; an operator route that
cleared the file while the running controller kept refusing) and the ONE
red-proof that PASSED, which is reported rather than omitted.

R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes.

STATUS.md restates the outcome and names the next operator step.
This commit is contained in:
2026-08-22 18:31:23 +02:00
parent 4e488321bf
commit a8caa0fdde
23 changed files with 744 additions and 75 deletions
+4
View File
@@ -26,6 +26,10 @@
---
| **R-379** | **The pre-restore undo copy was valid, was named to the customer, and no product action could apply it.** Shipped in controller v0.220.0 (+v0.220.1, v0.220.2). Evidence: `audits/DRILL-r379-rollback-2026-08-22/evidence/`. **Reasoning kept:** *R-379 and R-380 were ONE failure with ONE fix — both ended with a half-restored database and the only difference was whether it looked broken.* **The undo set is matched on THE RUN'S OWN STAMP, never on the `pre-restore-` prefix** (four copies coexisted on one app in one afternoon; a prefix match replays an arbitrary older state) **and never just the first file** (a two-database app would have had one restored and the other left half-written). **The rollback RE-DISCOVERS the container** — the undo file is stable, the container is not: the DB-only start re-creates it, and v0.220.0's own first live run held an app for 30 s of `waitDBReady` against a dead id while its data was recoverable. **No unit test saw that: they all inject the import seam and never look at container identity.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.1, 2026-08-22; docmost and bookstack both rolled back to byte-identical prior state) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
| **R-380** | **A failed MariaDB replay left a partially-applied database behind an app reporting `health=healthy`.** Shipped in controller v0.220.0. Evidence: `audits/DRILL-r379-rollback-2026-08-22/evidence/13-step2-verify.txt`. **Reasoning kept:** **no engine flag closes this** — `--single-transaction` was added to the Postgres import and does make it all-or-nothing, but **MariaDB's DDL is not transactional**, so a partial apply there is unavoidable at the engine. The flag is a belt; the rollback is the fix, and this row must not be read as saying otherwise. Proven live: `bookstack`'s `migrations` table back at **102 rows**, the exact cell the defect was measured in. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
| **R-381** | **The restore-failure message pasted raw engine stderr — including rows out of the customer's own database — into the Hungarian customer surface.** Shipped in controller v0.220.0. **Reasoning kept:** the full engine text now goes to the operator log, **which never had it before — the diagnostic was ADDED, not removed**. Measured: 407 bytes (Postgres) and 615 (MariaDB, whose middle was an `INSERT INTO migrations VALUES (…)` listing); now 257 bytes with no engine tokens. **A red-proof for this PASSED and the test was hollow**: it injected below `ImportDump`, so a leak reintroduced inside `ImportDump` could not fail it. The guard now sits at that layer. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
| **R-382** | **The reconstitution's summary log line omitted the volume count it already held.** Shipped in controller v0.220.0. Proven live: `0 file(s) placed, 3 volume(s) replayed, 1 DB dump(s) replayed`. | **CLOSED — SHIPPED** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
| **R-356** | **The off-site restore refused every app that has no data drive — it asked "does this app have an HDD path?" to answer "is this app installed?", and for 40 of 53 catalogue apps the honest answer to the first is permanently no.** Shipped in controller v0.219.0. Evidence: `audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/`. **Reasoning kept:** *the restore destination is resolved by the SAME rule as the capture destination — the drive if the app has one, the system data path otherwise (`Manager.GetAppDrivePath`, one expression). The refusal that protects a drive app from being restored onto the wrong disk applies to apps that HAVE a drive to get wrong.* **An app with no drive is not misconfigured** — `01-topology-and-trust.md` §8 carries the `[DESIGN]` marker; between 19 and 22 August that design was called a defect four times. **Deployment is asked of `ListDeployedStacks()` and FAILS CLOSED on a nil provider:** "cannot tell" must not become "go ahead" when the caller's next act is a write. **Two different failures get two different sentences** — installed-but-no-resolvable-data-root has its own refusal and its own route; widening `nincs telepítve` to cover it would send a customer to reinstall a running app and hide the real fault. **Measured, and load-bearing: 53 templates, 13 `needs_hdd: true`, 40 `false`** (catalogue @ `459766cb1639`). **The capture side's raw `GetStackHDDPath` is FENCED and was not changed** — capture resolves an app's declared `userdata`/`import` file legs against that value, and a system-data fallback there would write a snapshot claiming to hold files it does not. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.219.0, 2026-08-22; `privatebin` on `demo-hp`: planted, backed up, deleted, restored, 15/15 files byte-identical including two Hungarian accented names) | full text: `git show e18668f9e19f:documentation/backlog/OPEN-ITEMS.md` |
| **R-216** | **A correct recovery code was reported to the customer as wrong.** Shipped in 0.120.0, v0.125.0. | **SHIPPED** (controller v0.201.0 + hub v0.97.0/0.97.1) — **but see R-223**: the feature does not work on a NEW box until the manifest vouches agent 0.125.0. Until then such a box is correctly HELD, not lied to | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
| **R-218** | **Succeeding at recovery stopped the box asking for what it still needed.** Shipped in v0.203.0. Evidence: `documentation/tests/part4-rewalk-2026-08-06/journal.md`. | **CLOSED 2026-08-06 — controller v0.203.0, proven live.** *(State corrected 2026-08-06: this field read REOPENED while the body below already recorded the fix shipped and proven. The history of the over-claim is kept deliberately — it is why the row is worded as it is.)* **The over-claim, as it stood: the fix covered the DECLARATION half only.** Measured on the R-201 re-walk: the box declared, and **`offsiteheal` re-staged the secret at 11:44:57** saying *"the box re-consumes on its next cycle"* — **the next cycle came and went** (`host-report` 11:55:46, `Received report` 11:55:54, a full cycle **with a positive control that it ran**) **and the credential was still not consumed.** 23 minutes after the re-stage the box's last off-site-apply attempt was still the pre-re-stage one. A census of the customer-reachable actions on `/backups/remote` (`config`, `reset`, `run`, `toggle`) found **none that fetches a staged credential**, and the only lever is `systemctl restart felhom-controller-bootstrap.service` **inside the guest** — which worked in **18 s** (Campaign 11 measured 17), confirming nothing was wrong with the credential, the target or the key: **the only thing missing is anything at all to trigger a retry.** **This is the FIRST of the two dead ends that keep the recovery journey failing** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
-4
View File
@@ -136,10 +136,6 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
| ID | What | State |
|---|---|---|
| **R-379** | **The pre-restore undo copy is taken, is valid, is named to the customer — and NOTHING IN THE PRODUCT CAN APPLY IT.** When an off-site DB replay fails, `reimportDBDumpsFrom` returns and the refusal names the safety dump by filename (`offbox_reconstitute.go:436-441`). That filename appears ONLY inside the error string: there is no button, no list entry, no route. `preRestoreDumpPrefix` ("pre-restore-") is deliberately SKIPPED at three sites so these files are never offered as a restore source — `internal/backup/restore_unit.go:125`, `internal/backup/offbox_reconstitute.go:489`, `internal/backup/offbox_reconstitute.go:539`. **PROVEN LIVE 2026-08-22 on `demo-hp`, both engines.** Postgres (`docmost`): after a truncated dump the live database held 43 tables and **0 rows** in `pages`, `users` and `spaces`, and the app crash-looped. The undo copy (141 363 B, 43 COPY blocks, 4 page rows, 1 user row, the accented title present) was applied BY HAND and restored the exact prior state. MariaDB (`bookstack`): same, `migrations` 0 -> 102 rows. **So the data is recoverable — by us, by hand, over a support conversation. The customer has a filename.** | **OPEN — HIGH** | — | Offer the undo copy as a restore source on the app's restore page when one exists, or state in the message that recovery needs support and how to ask. The skip at the three sites is correct for *normal* listing — the gap is that there is no deliberate second surface. **Do NOT widen the three skips**: they exist so a safety dump is never mistaken for the app's own backup (that confusion is R-361's neighbourhood). | CC |
| **R-380** | **A failed MariaDB replay leaves a PARTIALLY APPLIED database behind an app that reports HEALTHY — Postgres fails visibly, MariaDB does not.** `ImportDump` gives the Postgres branch `-v ON_ERROR_STOP=1`; the MariaDB branch is a plain `mariadb -u root -p<pw> <db>` with no equivalent (`internal/appbackup/dbdump.go:670-692`). Both DO surface the failure — H3's predicted 'quiet success' did NOT occur — but the STATE they leave differs, and that is the defect. **Measured 2026-08-22 on `demo-hp` with the same truncation on both engines.** Postgres: everything emptied, app crash-loops, `Restarting (1)` — visibly broken. MariaDB: the dump's DROP/CREATE/INSERT runs table by table, so tables it reached are rebuilt, tables it never reached keep their ORIGINAL data, and the table it died inside is left EMPTY. Result on `bookstack`: `entities` 1 (intact), `users` 2 (intact), **`migrations` 0 rows (wiped)** — the schema-version ledger — while `docker inspect` reported **`health=healthy running=true restarts=0`** and the app served HTTP. An empty `migrations` table means BookStack believes no migration has ever run; the next upgrade would re-run all 102 against an existing schema. **Nothing signals ongoing damage.** | **OPEN — HIGH** | — | Make a failed replay leave a KNOWN state rather than a partial one: wrap the MariaDB import so a failure is atomic, or re-apply the undo copy automatically on import failure (which needs R-379 first), or at minimum mark the app unhealthy so the dashboard stops saying it is fine. **The MariaDB client's default IS to abort on error — that was measured, not assumed — so this is not a missing flag; it is the absence of a transaction boundary.** | CC |
| **R-381** | **The restore-failure message pastes raw database-engine stderr — including the customer's own database rows — into a Hungarian customer-facing surface.** `ImportDump` truncates stderr to 300 chars and wraps it verbatim (`internal/appbackup/dbdump.go:700-706`); `offbox_reconstitute.go:436-441` wraps that again; the flash renders it in an `alert alert-error` block. **Measured verbatim 2026-08-22.** Postgres, **407 bytes**, of which ~250 are untranslated English psql output with a caret diagram and `exit status 3`. MariaDB, **615 bytes**, whose middle is an `INSERT INTO \`migrations\` VALUES (1,&#39;2014_10_12_000000_create_users_table&#39;,1),(2,...` listing — i.e. **actual table contents, HTML-escaped, shown to the customer**. On a real app that statement could be any row the dump died inside. | **OPEN — MEDIUM** | — | Keep the engine text in the operator log where it belongs; give the customer the reason, the undo copy and the route. Same class as **R-79** (English on customer surfaces) and **R-257** (internal state names in customer copy), but a distinct producer and with a content-disclosure dimension neither has: this one can print rows. | CC |
| **R-382** | **The reconstitution's summary log line omits the volume count it already computed.** `offbox_reconstitute.go:452` logs `%d file(s) placed, %d DB dump(s) replayed, safety dump=%s, skewed=%v` — `res.VolumesReplayed` is set at line 412 and never printed. Measured 2026-08-22: `docmost` logged `0 file(s) placed, 1 DB dump(s) replayed` on a run that replayed **3** volumes including the entire 52 MB Postgres data directory; `bookstack` logged the same shape on a run that replayed 2 including a 161 MB one. The customer-facing flash DOES name the volumes („0 fájl és 1 adatkötet visszaállítva") — so the operator log is less informative than the customer message. This directly obstructed answering the 2026-08-22 drill's Q2 from the log and forced a planted discriminator instead. | **OPEN — LOW** | — | Add `%d volume(s) replayed` to the line. One format string. | CC |
| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor |
| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor |
| **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor |