Files
felhom.eu/documentation/audits/DRILL-r403-tier2-delete-2026-08-31/phase1d-repair.log
T
admin 66156c619f
gates / gates (push) Successful in 16s
R-403 drill evidence + the credential reader that ends a three-time mistake
The drill: the loss reproduced on the shipped v0.229.0 before anything was built. 120 082 104 B ->
7 036 B in one Tier-2 run, recorded as a success. Phases 1a (before), 1b (the hollow primary,
produced through the R-102 restore path exactly as the 2026-08-31 observation was), 1c (the loss),
1d (repair).

scripts/read_credential.py is Part 4's rider, and it exists because a note did not work three times:
2026-07-20 a Failed login was diagnosed as a stale password and written into memory; 2026-08-31 the
same misreading recurred and was caught; 2026-08-31, hours later, it recurred AGAIN and rewrote a
live box's password hash. Between them the project already had a memory file stating the rule, a
worked recipe in it, and a session report describing the mistake. The rule now lives in the code
path: one matching quote pair is unwrapped, the result is REFUSED if it still carries a quote, and
--expect-length gives the caller a second opinion. The value goes file->file at 0600 and stdout gets
only its length. test_read_credential.py asserts each refusal by its reason, with a positive control
before believing the not-in-stdout result.

Red-proof E1: remove the final quote assertion -> three cases fail by name.
2026-08-31 14:02:26 +02:00

27 lines
1.8 KiB
Plaintext

######## R-403 PHASE 1d — repair the box before building the fix ########
The first repair attempt used `rsync -a --delete` INSIDE the guest and silently did nothing:
rsync in the guest: NOT-INSTALLED
rsync lives in the CONTROLLER CONTAINER, not in guest 9201 — which is why Tier-2 (which shells out
from inside the container) works while a guest-side script does not. The script ran with
`set -uo pipefail` and no `-e`, so a missing binary continued as if it had succeeded. Recorded rather
than quietly re-run: an unchecked exit code that looks like success is the same trap Phase 4 of the
R-102 drill hit yesterday, in a different disguise.
Repaired with `cp -a` from the safety net:
primary created_at: 2026-08-31T09:43:41Z
db_dumps : ['docmost-postgres.sql']
volume_dumps: ['docmost_docmost_postgres_data.tar','docmost_docmost_redis_data.tar','docmost_docmost_storage.tar']
Secondary rebuilt by a real Tier-2 run through POST /api/backup/tier2:
db-dumps: 4 volume-dumps: 3 size: 120082104 bytes (identical to the phase-1a BEFORE state)
f46a2fc3aa9a7ae2502d83b1c6ef27e503102f5ba71a0c6559246d9674c8e3b1 volume-dumps/docmost_docmost_postgres_data.tar
a8df17c444e41f54762e122ce1be998315c969015a1580bb8abdc7211cfa1a73 volume-dumps/docmost_docmost_redis_data.tar
88f21f491d0766aa7a1fc9eba5866e5fffd7a72fa640c55f7bccf575f2ba751d volume-dumps/docmost_docmost_storage.tar
9f676376f759733f5b62e590e4a2b31dddd66ff49990df3394332b790a092a28 db-dumps/docmost-postgres.sql
docmost / docmost-redis / docmost-postgres: all healthy
The safety net at /mnt/sys_drive/felhom-data/r403-safekeeping/docmost-unit was intact throughout and
is what made the repair possible. It was deliberately placed OUTSIDE every backup tree, because
yesterday's set-aside was placed inside backups/primary/ and was swallowed by a directory the product
re-created underneath it.