Files
felhom.eu/REPORT.md
T
admin a8caa0fdde
gates / gates (push) Successful in 17s
R-379/R-380 docs: the failure ladder, the drill record, register housekeeping
07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay ->
rollback -> hold, including why no engine flag closes it: --single-transaction
makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the
fix and the flag is a belt.

Drill record for the live walk, including the TWO defects the walk found in the
fix itself (a rollback into a re-created container; an operator route that
cleared the file while the running controller kept refusing) and the ONE
red-proof that PASSED, which is reported rather than omitted.

R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes.

STATUS.md restates the outcome and names the next operator step.
2026-08-22 18:40:22 +02:00

4.4 KiB

REPORT — R-379/R-380/R-381/R-382: the undo copy goes back (2026-08-22)

Companion to felhom-controller v0.220.0 → v0.220.1 → v0.220.2. Full record: documentation/audits/DRILL-r379-rollback-2026-08-22/.

What shipped

R-379 and R-380 were one failure with one fix. Both ended with a half-restored database; the only difference was whether it looked broken. When the replay fails, the product now re-applies the customer's own pre-restore copy — the same ImportDump call a person ran by hand yesterday to recover both apps — and the app comes back with a message saying both that the restore failed and that the data is as it was.

The whole undo set, matched on the run's own stamp, never on the pre-restore- prefix and never just the first file. When the rollback also fails the app is held stopped — the operator's ruling — with every start path refusing it, the app-stop marker ended so nothing auto-restarts it, and the row red rather than green.

R-381: the failure message stopped pasting engine output (407→257 bytes on Postgres; the MariaDB one had been 615 bytes with rows out of the customer's own database). The full text now reaches the operator log, which never had it. R-382: the summary log prints the volume count it already held. Also: undo copies resolve to their own app and are capped at 3.

Documents updated here

  • documentation/architecture/07-backup-architecture.md §6.3 — a dated [DESIGN] paragraph on the failure ladder replay → rollback → hold, and why an engine flag does not close it.
  • STATUS.md — the outcome in plain words; the deciding section says what happens if nothing is done.
  • documentation/backlog/ — R-379…R-382 compressed into CLOSED-ITEMS.md, each keeping its title, shipping version, evidence path and every sentence that states a rule. Full text: git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md.

Register size: OPEN-ITEMS.md 330 683 → 325 236 bytes; CLOSED-ITEMS.md 63 507 → 66 777.

The live walk found two defects in the fix itself

Both are recorded because the walk, not the tests, caught them.

  1. The rollback used a dead container (fixed v0.220.1). The DB-only start re-creates the DB container, so the id captured at dump time is dead by rollback time. Measured: captured 9adbc14f9af6, re-created 309795897b82, rollback timed out after 30 s — the app was held for an infrastructure reason while its data was recoverable. No unit test could see it: they all inject the import seam and never look at container identity.
  2. The operator route did not take effect (fixed v0.220.2). --clear-restore-hold runs as a second process; it cleared the file and the running controller went on refusing. Found by using it.

A red-proof that PASSED

Of nine mutations, one did not fail its test and is reported rather than omitted: the R-381 behavioural test injected below ImportDump, so a leak reintroduced inside ImportDump was invisible to it. A guard now sits at that layer and the mutation convicts.

What did not reproduce

The undo copies rendering as app rows on the customer's backup page. The live page was read before any change: zero pre-restore strings while four such files sat on disk, with a positive control showing 8 real rows. The phantom name was real as a map key, never a row. Fixed as a naming defect; their visibility is unchanged and deliberate.

The golden was baked in this session, and why

The golden-currency gate refused this docs push: 0.220.2 was released with no golden. That block is not circular — a golden needs the controller image, which was already pushed, not this commit — so the gate was satisfied by doing the work it asked for rather than bypassed with --no-verify. No push in this session used --no-verify.

Golden 0.220.2, sha256 cb439418c7005ce01bcb8126bb6385688c2f408c3c4649f6740001bc2c864ed5, 657 271 965 B, round-trip verified. All five markers hit, both negative controls at zero, both token-leak greps proved able to convict before their zeros were accepted. Record: documentation/tests/golden-0.220.2-2026-08-22/.

Operator follow-up

Vouch the golden — Hub → Configuration → Day-0 artifacts, a three-field save: golden_version 0.220.2, agent_version 0.130.0, min_agent 0.129.0. Then raise the floor to 0.220.2, last, in its own save.