17 KiB
REPORT — R-403: a poorer copy must never delete a richer one
Controller v0.230.0 · 2026-08-31 · MinAgent 0.129.0 (unchanged)
2. PART 1'S RESULT, FIRST — the loss is REAL and was reproduced before anything was built
On the shipped v0.229.0, on demo-hp, app docmost. The hollow primary was produced through the
R-102 restore path, exactly as the 2026-08-31 observation was — not hand-crafted.
BEFORE AFTER (one Tier-2 run)
db-dumps : 4 files db-dumps : 0 files
volume-dumps : 3 files volume-dumps : 0 files
unit size : 120 082 104 bytes unit size : 7 036 bytes
9f676376…a092a28 db-dumps/docmost-postgres.sql GONE
73917ba6…fe15ef1 db-dumps/pre-restore-20260822T162347Z-…sql GONE
4c134c2c…4935949 db-dumps/pre-restore-20260822T162708Z-…sql GONE
13e5a864…422af25 db-dumps/pre-restore-20260822T215432Z-…sql GONE
f46a2fc3…4c8e3b1 volume-dumps/docmost_docmost_postgres_data.tar GONE
a8df17c4…1cfa1a73 volume-dumps/docmost_docmost_redis_data.tar GONE
88f21f49…5f2ba751d volume-dumps/docmost_docmost_storage.tar GONE
The run's own line: [backup] Tier 2 copied docmost → …/backups/secondary/docmost (14.9 KB, 0 leg(s), 0s) — recorded as a success.
VERDICT: LOSS CONFIRMED. The code reading filed yesterday was right, and it is now a measurement.
The hollow primary manifest that armed it: created_at 2026-08-31T11:32:47Z, db_dumps: [],
volume_dumps: null, written by the 5-minute backup-cache job ~140 s after the restore.
Evidence: felhom.eu/documentation/audits/DRILL-r403-tier2-delete-2026-08-31/ — phase1a (before),
phase1b (the hollow primary), phase1c (the loss), phase1d (repair).
1. Confirmed baselines
Re-checked before the first edit. No drift.
| Repo | main @ start |
expected | version |
|---|---|---|---|
felhom-controller |
fed272e62adf3c72a2c61f7f2f1f0a5e21164c71 |
same | v0.229.0 → v0.230.0 |
felhom.eu |
83ff9e8e3856fa822bfa1f80fc305783964aef68 |
same | — (docs + one script) |
felhom-agent |
not touched | — | v0.130.0 unchanged |
git status --porcelain empty in both; disk 37% / 51%; demo-hp was running the shipped 0.229.0,
which is what Part 1 required.
3. Files created / modified
felhom-controller
| File | Change |
|---|---|
controller/internal/backup/r403_hollow.go |
new — unitCarriesData / unitIsHollow, the ONE predicate |
controller/internal/backup/tier2.go |
the unit-leg precondition; tier2UnitPreservedWarning; unitPackageDate; recordTier2SuccessWithUnit (the 5-arg form kept as a thin caller) |
controller/internal/backup/tier2_shares.go |
one comment — shares have no unit leg |
controller/internal/backup/tier2_restore.go |
rehydratePrimaryUnit, countUnitFiles; Tier2Coverage.{UnitPackageDate,UnitLegPreserved}; UnitRestoreDate() |
controller/internal/backup/backup.go |
the unitRehydrate seam |
controller/internal/settings/settings.go |
CrossDriveBackup.{UnitLegSkipped,UnitPackageDate} |
controller/internal/web/handlers.go |
tier2UnitStaleClause, tier2UnitStaleNoticeFmt, tier2UnitConfirmWithStaleness (2-arg form kept as a thin caller); the row's Tier2UnitStaleNotice; tier2UnitSourceMsg now names the package date |
controller/internal/web/templates/backups_apps.html |
the stale notice on the card |
controller/internal/backup/r403_{hollow,mirror_guard,rehydrate}_test.go, controller/internal/web/r403_surface_test.go |
new — Groups A–D |
CHANGELOG.md, CONTEXT.md, controller/README.md, REPORT.md |
documentation |
felhom.eu — scripts/read_credential.py + scripts/test_read_credential.py (new, Part 4);
documentation/architecture/07-backup-architecture.md (§8 row 5 note + new §8.2);
documentation/architecture/00-capability-map.md; documentation/backlog/{OPEN,CLOSED}-ITEMS.md;
STATUS.md; documentation/audits/DRILL-r403-tier2-delete-2026-08-31/ (new, 9 files).
Not modified: felhom-agent, app-catalog-felhom.eu, rsyncMirror, the data legs, the R-165/R-181
capture floor, the off-site path, golden_currency_gate.py (that is R-404 and it is Viktor's).
4. Commits pushed to main
| Repo | Commit | What |
|---|---|---|
felhom.eu |
66156c619fd2bceff5af6213f11a0836183f6880 |
drill evidence + read_credential.py |
felhom-controller |
2358e561b741b3f08e254b28b29d8b6906dd52a0 |
the guard, the honesty, the rehydrate, Groups A–D |
felhom-controller |
5429d651ee882ce3354216aa44f4669dce15e30b |
the over-eager stale flag, caught by the live run |
felhom-controller |
b48a7fa326dbe5505d1568ac55c66488e8de125c |
the outcome names the package date |
felhom-controller |
1cfdde968fab0d2cd07013830f532f70ecc786ab |
CHANGELOG / CONTEXT / README / REPORT |
felhom.eu |
dddcc808be95d1c89b276b4d791491bad3c96bba |
architecture §8.2, capability map, register, STATUS — pushed with --no-verify, declared |
5. Per-test results and the named red-proofs
Full suite: go build ./... && go vet ./... && go test ./... — all packages ok
(internal/backup 313 s). controller_gates.py — all 13 OK. test_read_credential.py — OK.
| Group | Tests | Result |
|---|---|---|
| A | TestR403_{ManifestWithNoDumpsIsHollow, ManifestWithVolumeDumpsOnlyIsNotHollow, ManifestWithDBDumpsOnlyIsNotHollow, AbsentManifestIsHollow, UnparseableManifestIsHollow, SizeIsNeverConsulted, AHealthyAppIsNeverCalledStale} |
PASS |
| B | TestR403_{HollowSourceOverCompleteDestIsSkipped, CompleteSourceStillMirrors, HollowOverHollowStillMirrors, FirstCopyStillMirrors, OtherLegsStillRunWhenTheUnitLegIsSkipped, SkipIsRecordedForTheSurface, DataLegShrinkIsUnaffected, GuardUsesTheSharedPredicate} |
PASS |
| C | TestR403_{HollowPrimaryIsRefilledFromTheMirror, CompletePrimaryIsLeftByteIdentical, FailedRestoreDoesNotWriteAPackage, RehydrateHappensBeforeTheCallReturns, RehydrateFailureDoesNotFailTheRestore} |
PASS |
| D | TestR403_{SkippedUnitLegIsNotRenderedAsFresh, UnitRestoreOfferNamesTheOlderPackageDate, TheOrdinaryConfirmIsUnchanged, OutcomeNamesThePackageDateNotTheRunDate} |
PASS |
| E | test_r404_credential_length_mismatch_fails_loudly, test_the_value_is_never_printed |
PASS |
| E2 | the existing suite unmodified — no existing test was edited (both changed signatures kept their old form as thin callers) | PASS |
Red-proofs — each mutated, run, observed failing, reverted:
| # | Mutation | Observed failure |
|---|---|---|
| A6 | predicate → dirSizeBytes > 1024 |
FAIL: "a 400346-byte unit listing NO dumps was called data-bearing" + "a 360-byte unit listing a volume tar was called hollow" |
| B1 | the guard removed | FAIL: "the destination unit CHANGED", all three files "was DELETED from the copy", and "the mirror seam WAS called for the unit leg" |
| B6 | a general never-shrink rule (refuse any leg whose destination exists) | FAIL: "a data leg stopped shrinking — the guard is TOO WIDE and is fencing a design decision" |
| C2 | the only-when-hollow condition dropped | FAIL: "the rehydrate ran 1 time(s) over a COMPLETE primary" |
| E1 | the quote assertion removed | FAIL ×3 by name: one-sided strip, trailing-only quote, mismatched pair |
| (extra) | reinstate the package-older-than-the-run comparison | FAIL: "a healthy app … was flagged as preserved/stale" |
B6's first mutation was wrong and is recorded rather than quietly re-done. It skipped the data legs only when the unit leg was skipped, and B6's fixture has a COMPLETE source, so the mutation never reached it —
OtherLegsStillRunfailed instead. Re-done as a true never-shrink rule, which is what B6 actually guards, and then it failed correctly.
6. Test count
Go: 1632 → 1656 (+24). Python: +2 (test_read_credential.py).
7. Deployed version
$ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"
gitea.dooplex.hu/admin/felhom-controller:0.230.0 Up (healthy)
The fleet is on 0.229.0 — WHICH CARRIES THE DEFECT. demo-hp was updated by hand; demo-felhom
is still on 0.229.0. A golden carrying 0.230.0 is owed (R-242, Viktor's), and this time the day-0
ground that justified the previous six bypasses does not apply: R-403 is a defect in the nightly
Tier-2 copy, which a newly installed box starts running on its first night.
8. The live evidence — Scenarios B, D, E
Scenario B — the same state, on the fixed build. The WARN, verbatim:
[WARN] [backup] Tier 2 docmost: unit leg SKIPPED — the recovery unit on the source drive lists no
database dumps and no volume tars, while the existing copy at
/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit does. The copy was PRESERVED rather
than replaced with an empty one (R-403). The other legs continue.
[INFO] [backup] Tier 2 copied docmost → …/secondary/docmost (14.9 KB, 0 leg(s), 0s)
[unit leg SKIPPED — existing package preserved, R-403]
db-dumps: 4 volume-dumps: 3 size: 120082104 before and after, and all seven sha256 values
identical. On v0.229.0 the same state left 0 files.
Scenario D — the surfaces, per row, with the other seven apps as the control:
app notice FIGYELEM package date in the confirm
bookstack False False 2026-08-31 14:03
calibre-web False False 2026-08-31 14:03
docmost True True 2026-08-31 11:43 <- the preserved package
kimai False False 2026-08-31 14:03
opengist False False 2026-08-31 14:03
paperless-ngx False False 2026-08-31 14:03
privatebin False False 2026-08-31 14:03
romm False False 2026-08-31 14:03
Only the skipped app carries the notice, and its confirm names the package's date (11:43) while
every other row names its freshly-mirrored one (14:03). ASCII fragments (adatcsomagja, FIGYELEM)
with the seven other rows as the negative control.
Scenario E — the rehydrate. Immediately after the call returned, with nothing waited for:
BEFORE created_at: 2026-08-31T12:08:49Z db_dumps: [] volume_dumps: None
AFTER created_at: 2026-08-31T09:43:41Z db_dumps: ['docmost-postgres.sql']
volume_dumps: ['…postgres_data.tar','…redis_data.tar','…storage.tar']
volume tars on the app drive: 3 db dumps: 4
[INFO] [backup] docmost: primary unit refilled from the secondary mirror (R-403) —
3 volume tar(s), 4 database dump(s) now on the app's own drive
And after waiting out 3 backup-cache cycles (330 s) — the job that wrote the hollow manifest in
the first place — the primary is still a real package: created_at 2026-08-31T12:28:32Z, 1 db dump,
3 volume tars. The hollow state is gone, and the capture is describing reality.
9. NOT live-validated — explicit
- Scenario C3/C4 live (hollow→hollow, and a data leg shrinking) — unit-tested only. The live box had no app in either state and manufacturing one would have meant breaking a second app's backup.
- The rehydrate's failure path (
unitRehydratereturning an error) — unit-tested only; no way to make a realrsyncfail on that box without damaging something. - A real second-drive failure. Everything here was proven by making a package hollow, never by
removing a disk.
07§8 row 4 remains PARTIAL for that reason and did not move. demo-felhomwas deliberately untouched; the fix is proven on one machine.- Rendering — endpoint level, as always here: the markup is proven, the browser dialog is not.
10. Rows moved
07§8 row 5 — status unchanged; a pointer added to the new §8.2, which states the derived-copy rule is intact and names the single exception, so a future reader does not "fix" the skip back.07§8.2 — new section, with the measurement, the four-case table, and the reason the data legs are not guarded.00-capability-map.mdTier-2 row — no status change, stated explicitly rather than left ambiguous. R-403 removes a way the route could be DESTROYED between uses; it does not change what the route can be relied on for.07§8 rows 3b and 4 — unchanged, and that is deliberate: 3b is PROVEN on what the route does, which R-403 does not alter.
11. Teardown
- Machines provisioned: none. Hub records created: none.
- The drilled app:
docmostis running and healthy, and both copies are complete and byte-identical — primary and secondary each 3 tars + 4 dumps, 120 082 104 B, sha256 matching. The final Tier-2 run mirrored normally (114.5 MB, 0 leg(s), 0 skips), which also proves Scenario C1 live. The surface shows no stale notice on any app. - On-box artefacts: the safety-net copy of the unit (deliberately placed OUTSIDE every backup tree, because yesterday's set-aside was swallowed by a directory the product re-created), the endpoint driver, the password and session files and every phase script — all removed or shredded.
- All 8 apps on the box report
healthy.
12. Register
| Row | Action |
|---|---|
| R-403 | CLOSED — controller v0.230.0, proven live both ways. Compressed into CLOSED-ITEMS.md naming git show 66156c619fd2:…/OPEN-ITEMS.md for the original |
| R-404 | FILED and deliberately NOT acted on — a decision for Viktor on whether a documents-only push should be subject to the golden-currency gate. Both sides stated, plus what happens if he does nothing. The gate was not changed. |
| R-242 | appended — seventh conviction, and the first where the day-0 ground does NOT apply |
Register size: OPEN-ITEMS.md 594 → 593 lines (R-403 out, R-404 in); CLOSED-ITEMS.md 239 →
240; ROADMAP.md unchanged (it carries no R-403 row; one_register_gate.py green).
13. Observations
-
FILED: R-404 — six correct bypasses of one gate is a habit, not a guard. Filed as a decision, not built. Detailed above.
-
FILED: R-242 — the seventh conviction, and the ground that justified the other six has expired. The
felhom.eupush usedgit push --no-verify, declared. Unlike the previous six, this release does bite a day-0 box: R-403 is a defect in the nightly Tier-2 copy, which a new machine starts running on its first night. -
NOT-A-FINDING: my own live validation found a defect my unit tests did not, and the shape is worth naming. The first draft flagged "the package is older than the run" by comparing dates — true of every healthy app, because a unit is always captured shortly before the run that mirrors it. Four healthy apps on the box would have been warned. It is not a register row because it was found and fixed inside this task, but it is recorded in
CONTEXT.mdas a shape: a warning that fires on everything costs the same as the comforting lie it replaces. -
NOT-A-FINDING: my first Scenario-D control was broken and produced a false alarm. I scanned a fixed 9000-character window from each app's name, which spilled into the next app's row, so kimai appeared to carry docmost's warning. Re-done by splitting on the real row container. The broken control is what surfaced observation 3, so it is recorded rather than quietly replaced.
-
NOT-A-FINDING:
rsyncis not installed in guest 9201 — it lives inside the controller container. My first repair script shelled out to it withset -uo pipefail(no-e) and silently did nothing; the log says so at the top ofphase1d-repair.log. Same trap as yesterday'sdocker volume rmin a different disguise: an unchecked exit code that looks like success. -
NOT-A-FINDING: one
--no-verifywas used unnecessarily. The evidence/script push at66156c6was pushed with--no-verifybefore I had checked whether the gate was green — it was (the immediately following no-op push printedgates OK). Harmless, and recorded because a bypass that was not needed is exactly the habit R-404 is about. -
NOT-A-FINDING:
recordTier2Successandtier2UnitConfirmMsgkept their old signatures as thin callers. Both needed new arguments, and both had existing callers including tests. §9/E2 forbids editing an existing test, so each gained a…WithUnit/…WithStalenesscore with the old form as the thin caller — the ONE-implementation-two-callers pattern this repo already uses. No existing test was touched. -
NOT-A-FINDING: the drill's session expired mid-run and a POST silently did nothing. After the 0.230.0 restart the recorded
felhom_sessionwas dead;ctl.sh postprinted no status line and no Tier-2 ran. Caught because the secondary was unchanged when it should have been evaluated. Recorded in the evidence asphase3-scenarioB-first-attempt-session-expired.lograther than deleted.
14. Final verification
felhom-controller/controller$ go build ./... && go vet ./... && go test ./... → all ok
felhom-controller$ python3 controller/scripts/controller_gates.py → all 13 gates OK
felhom.eu$ python3 scripts/test_read_credential.py → OK
felhom.eu$ python3 scripts/repo_gates.py → 11 OK, golden-currency FAILED (declared, §13.2)