DRILL R-356b: the off-site restore for a driveless app that HAS a database
gates / gates (push) Successful in 16s
gates / gates (push) Successful in 16s
A drill, not an implementation. No code, no version bump, no CHANGELOG entry. Ten of the forty driveless apps carry a database; I re-measured that count and got 10. For those ten the restore is a five-leg operation that never ran at all until this week, because R-356 refused before any of it started. Walked end to end on demo-hp for both engines - docmost (Postgres 16) and bookstack (MariaDB 12.3) - each deployed for the drill, planted through the app's own interface, destroyed for real, restored through the endpoint the UI posts to. Q1 does it complete: YES. All five legs ran and succeeded, 32s / 25s. Accented names byte-identical both directions. Q2 which leg won: the SQL DUMP. Three-way discriminator returned the altered dump's value. This confirms R-164's F17 ordering on the OFF-SITE path; R-164 only ever cited the local one. Scratch-only mutation; store proved unmutated. Q3 does a failure tell the truth: partly, and two defects. Filed R-379 (HIGH, the undo copy is valid, named, and unappliable by any product action - proven by applying it by hand on both engines), R-380 (HIGH, a failed MariaDB replay leaves a partial database behind an app reporting healthy, where Postgres crash-loops visibly), R-381 (MEDIUM, the failure message pastes engine stderr including customer table rows into the Hungarian surface), R-382 (LOW, the summary log omits the volume count it already has). H1, H2 and H4 did NOT fire and that is recorded. H3 fired in a shape nobody predicted: not a quiet success, but a loud error over a silent inconsistency. R-361 reproduced independently on a second app. restic check: no errors, 29 snapshots. A flaw in the drill's own planting - a double-escaped accented title - was caught by reading stored bytes as hex, recorded, and re-measured in Phase 1b. Register 325236 -> 330683 bytes. Nothing dropped. Teardown: two apps retained with reason, no pvesm before-snapshot taken (said plainly), no hub-side record created.
This commit is contained in:
@@ -136,6 +136,10 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
|
||||
|
||||
| ID | What | State |
|
||||
|---|---|---|
|
||||
| **R-379** | **The pre-restore undo copy is taken, is valid, is named to the customer — and NOTHING IN THE PRODUCT CAN APPLY IT.** When an off-site DB replay fails, `reimportDBDumpsFrom` returns and the refusal names the safety dump by filename (`offbox_reconstitute.go:436-441`). That filename appears ONLY inside the error string: there is no button, no list entry, no route. `preRestoreDumpPrefix` ("pre-restore-") is deliberately SKIPPED at three sites so these files are never offered as a restore source — `internal/backup/restore_unit.go:125`, `internal/backup/offbox_reconstitute.go:489`, `internal/backup/offbox_reconstitute.go:539`. **PROVEN LIVE 2026-08-22 on `demo-hp`, both engines.** Postgres (`docmost`): after a truncated dump the live database held 43 tables and **0 rows** in `pages`, `users` and `spaces`, and the app crash-looped. The undo copy (141 363 B, 43 COPY blocks, 4 page rows, 1 user row, the accented title present) was applied BY HAND and restored the exact prior state. MariaDB (`bookstack`): same, `migrations` 0 -> 102 rows. **So the data is recoverable — by us, by hand, over a support conversation. The customer has a filename.** | **OPEN — HIGH** | — | Offer the undo copy as a restore source on the app's restore page when one exists, or state in the message that recovery needs support and how to ask. The skip at the three sites is correct for *normal* listing — the gap is that there is no deliberate second surface. **Do NOT widen the three skips**: they exist so a safety dump is never mistaken for the app's own backup (that confusion is R-361's neighbourhood). | CC |
|
||||
| **R-380** | **A failed MariaDB replay leaves a PARTIALLY APPLIED database behind an app that reports HEALTHY — Postgres fails visibly, MariaDB does not.** `ImportDump` gives the Postgres branch `-v ON_ERROR_STOP=1`; the MariaDB branch is a plain `mariadb -u root -p<pw> <db>` with no equivalent (`internal/appbackup/dbdump.go:670-692`). Both DO surface the failure — H3's predicted 'quiet success' did NOT occur — but the STATE they leave differs, and that is the defect. **Measured 2026-08-22 on `demo-hp` with the same truncation on both engines.** Postgres: everything emptied, app crash-loops, `Restarting (1)` — visibly broken. MariaDB: the dump's DROP/CREATE/INSERT runs table by table, so tables it reached are rebuilt, tables it never reached keep their ORIGINAL data, and the table it died inside is left EMPTY. Result on `bookstack`: `entities` 1 (intact), `users` 2 (intact), **`migrations` 0 rows (wiped)** — the schema-version ledger — while `docker inspect` reported **`health=healthy running=true restarts=0`** and the app served HTTP. An empty `migrations` table means BookStack believes no migration has ever run; the next upgrade would re-run all 102 against an existing schema. **Nothing signals ongoing damage.** | **OPEN — HIGH** | — | Make a failed replay leave a KNOWN state rather than a partial one: wrap the MariaDB import so a failure is atomic, or re-apply the undo copy automatically on import failure (which needs R-379 first), or at minimum mark the app unhealthy so the dashboard stops saying it is fine. **The MariaDB client's default IS to abort on error — that was measured, not assumed — so this is not a missing flag; it is the absence of a transaction boundary.** | CC |
|
||||
| **R-381** | **The restore-failure message pastes raw database-engine stderr — including the customer's own database rows — into a Hungarian customer-facing surface.** `ImportDump` truncates stderr to 300 chars and wraps it verbatim (`internal/appbackup/dbdump.go:700-706`); `offbox_reconstitute.go:436-441` wraps that again; the flash renders it in an `alert alert-error` block. **Measured verbatim 2026-08-22.** Postgres, **407 bytes**, of which ~250 are untranslated English psql output with a caret diagram and `exit status 3`. MariaDB, **615 bytes**, whose middle is an `INSERT INTO \`migrations\` VALUES (1,'2014_10_12_000000_create_users_table',1),(2,...` listing — i.e. **actual table contents, HTML-escaped, shown to the customer**. On a real app that statement could be any row the dump died inside. | **OPEN — MEDIUM** | — | Keep the engine text in the operator log where it belongs; give the customer the reason, the undo copy and the route. Same class as **R-79** (English on customer surfaces) and **R-257** (internal state names in customer copy), but a distinct producer and with a content-disclosure dimension neither has: this one can print rows. | CC |
|
||||
| **R-382** | **The reconstitution's summary log line omits the volume count it already computed.** `offbox_reconstitute.go:452` logs `%d file(s) placed, %d DB dump(s) replayed, safety dump=%s, skewed=%v` — `res.VolumesReplayed` is set at line 412 and never printed. Measured 2026-08-22: `docmost` logged `0 file(s) placed, 1 DB dump(s) replayed` on a run that replayed **3** volumes including the entire 52 MB Postgres data directory; `bookstack` logged the same shape on a run that replayed 2 including a 161 MB one. The customer-facing flash DOES name the volumes („0 fájl és 1 adatkötet visszaállítva") — so the operator log is less informative than the customer message. This directly obstructed answering the 2026-08-22 drill's Q2 from the log and forced a planted discriminator instead. | **OPEN — LOW** | — | Add `%d volume(s) replayed` to the line. One format string. | CC |
|
||||
| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor |
|
||||
| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor |
|
||||
| **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor |
|
||||
|
||||
Reference in New Issue
Block a user