DRILL closeout: the scheduled cycle agrees, the abandonment sweep watched firing, R-366
gates / gates (push) Successful in 16s

The box's own 02:30 / 03:30 / 04:15 cycle ran unattended and agrees with the manual one
on every measure. The one that matters: R-355 is not an artefact of manual triggering —
the scheduled run again wrote paperless-ngx's PostgreSQL dump into a directory for a
stack that does not exist, and again left the app's own unit recording db_dumps: null.

Part 4.2's terminal deletion has now been observed. At 05:10 the sweep removed exactly
the recorded set-aside store and left the live repository untouched; at 05:13 the hub
dropped the sealed package that protected it and said so (event 3025). Both halves went
together, three minutes apart, and the controller cleared its own state. demo-felhom's
two preserved fixtures were verified untouched throughout.

R-366 (HIGH) filed, found incidentally: the 21 August reinstall orphaned demo-hp's PBS
whole-guest archives as well as its restic repo, so a rebuilt box loses BOTH off-premises
tiers at once. The restore-test caught it and named the key mismatch precisely; it is
merely called "a failed restore test" rather than "your older backups are unreadable".

demo-hp is left HEALTHY, not broken. Ceiling R-353 -> R-366.
This commit is contained in:
2026-08-22 05:25:30 +02:00
parent d895d9f7dd
commit 7064596c2e
4 changed files with 249 additions and 6 deletions
+154 -6
View File
@@ -245,9 +245,33 @@ leads, and the log line alone says only „backup OK".
---
## 5. THE TWO CYCLES COMPARED
## 5. THE TWO CYCLES COMPARED — AND THEY AGREE
*(filled in after the scheduled run — see §11)*
**By hand:** local cycle 22:12, off-site 22:17 (after enabling the per-app switches the rebuild had
silently cleared). **By the clock:** the box's own `db-dump` at 02:30 CEST, cross-drive at 03:30,
off-site at 04:15.
| | manual | scheduled | agree? |
|---|---|---|---|
| units refreshed | all 6 apps | all 6 apps, `00:30Z` | **yes** |
| `privatebin` volume tar | 1 055 744 B | 1 055 744 B | **yes** |
| `opengist` volume tar | 181 248 B | 181 248 B | **yes** |
| `calibre-web` config tar | 1 422 848 B | 368 640 B † | **yes, and explained** |
| `paperless-ngx` `db_dumps` | **absent** | **absent** | **yes — the defect reproduces on the scheduled path** |
| the orphan `…/primary/paperless/db-dumps/` | written | **written again, 294 936 B, `00:30Z`** | **yes** |
| planted files on the drive | 5/5 identical | **5/5 identical**, plus `POST-SNAPSHOT.txt` | **yes** |
| off-site run | ok, 27 → snapshots | ok, `02:17:12Z`, **2m8s**, 42.8 MB, **27 snapshots** | **yes** |
† the tar shrank because the drill had by then removed the planted 1 MB from that volume — the
expected value, not a discrepancy.
**Two independent observations, no disagreement.** The one that matters: **R-355 is not an artefact of
my manual triggering.** The box, unattended, on its own schedule, wrote paperless-ngx's PostgreSQL
dump into a directory for a stack that does not exist and left the app's own unit recording
`db_dumps: null`.
The scheduled cross-drive (Tier-2) leg also completed for all six apps at 01:30Z — `crossdrive_completed`
events 3019–3024.
---
@@ -347,7 +371,7 @@ means in all three cases.
## 8. RANKED REGISTER ROWS OPENED
Ceiling was **R-353**; it **moved to R-365**.
Ceiling was **R-353**; it **moved to R-366**.
| id | rank | what |
|---|---|---|
@@ -363,6 +387,7 @@ Ceiling was **R-353**; it **moved to R-365**.
| **R-363** | **MEDIUM** | The fill watcher runs **once a day (03:30)** plus at startup. A filesystem that fills at 03:31 is unannounced for ~24 h — while the backup reserve is already refusing apps. |
| **R-364** | **LOW** | Accented-text search is a silently-transforming instrument; discipline has failed at least three times. Guard proposed, not built. **(Part 5.2)** |
| **R-365** | **LOW** | An **overdue** abandonment countdown renders its past due-date in the future tense („…2026-08-20 napján véglegesen töröljük" shown on 2026-08-21). |
| **R-366** | **HIGH** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives too** — the box cannot read its own pre-reinstall backups (`manifest's key 3f:4f:65:c0… does not match provided key dd:d1:d8:53…`, hub event 3016, filed unprompted at 21:59). The PBS-tier analogue of R-193: a rebuilt box loses **both** off-premises tiers at once. The restore-test caught it precisely; the gap is that it is called *a failed restore test* rather than *your older whole-guest backups are unreadable*. **Found incidentally — nobody was looking for it.** |
**Confirmed still live, not re-filed:** **R-329** — `app_start_failed` emits severity `"warn"`
(`internal/notify/notifier.go:546`), outside the hub's vocabulary, so it coerces to `info` and emails
@@ -440,12 +465,135 @@ over a unit that *did* have a data leg.
---
## 11. SCHEDULED CYCLE AND THE COUNTDOWN
## 11. PART 4.2 — THE ABANDONMENT COUNTDOWN, WATCHED FIRING
*(filled in as they fire — 02:30 local backup, 04:15 off-site, 05:10 abandonment sweep, all CEST)*
**The terminal deletion has now been observed.** It did exactly what it claims, including the half
nobody had seen.
- **State created** 23:08–23:09 CEST. A realistic set-aside store was built at
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, the countdown written overdue
(started 2026-08-07, due 2026-08-20), and the product's own CLI confirmed it:
`abandonment countdown RUNNING … deleted on: 2026-08-20 … days left: 0`.
- **05:10 CEST — it fired.** `/home` on the storage box is stamped `03:10Z`; the set-aside store is
**gone**; **`/home/felhom-repo`, the live repository, is untouched** (mtime still 4 Aug).
- **05:13 CEST — the hub half.** Event **3025 `offsite_abandon_purged`**: *„Az ügyfél korábbi távoli
mentései és a hozzájuk tartozó megőrzött helyreállítási csomag is törölve (1 csomag). Az ügyfél
döntése alapján, a 14 napos türelmi idő lejárta után."* And demo-hp's superseded escrow row **is
gone from the hub** — it held one before.
- **The controller closed itself out.** Every `abandon_*` field has been removed from `settings.json`
and `--abandon-status` reports *"no abandonment countdown is running on this box"*.
**So the two-phase commit's central promise — *"it removes BOTH halves or neither"* — is confirmed
live for the first time:** the ciphertext and the sealed package that protects it went together, three
minutes apart, and the state that remembered the operation cleaned itself up. **What it removes:**
exactly the recorded set-aside path. **What survives:** the live repository, the live escrow, and the
box's current recovery path.
**Judged:** the completion message is accurate and in plain Hungarian. The only wrong note is while
the countdown is *overdue but not yet swept* — the card then states a past date in the future tense
(**R-365**).
**No countdown is left running anywhere.** `demo-hp` cleared itself; `demo-felhom` reports none.
---
## 11b. WHAT THE NIGHT FOUND THAT NOBODY WAS LOOKING FOR
At 21:59 the box filed, unprompted, hub event 3016:
> `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could
> not be restored+booted … wrong key — manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided
> key dd:d1:d8:53:44:62:5e:0b`
That archive predates the 21 August reinstall by three days. **The rebuild orphaned the PBS whole-guest
archives exactly as it orphaned the restic repository** — so a rebuilt box loses *both* off-premises
tiers at once. The restore-test mechanism deserves credit: it caught it and named the key mismatch
precisely. The gap is what it is *called* — "a restore test failed" reads as a flaky verification, not
as "every whole-guest backup taken before the reinstall is unreadable on this machine". Filed **R-366
(HIGH)**.
---
## 12. MACHINE STATES AT THE END
*(final state recorded in §13 after the scheduled runs)*
**`demo-hp` — HEALTHY, not broken. Nothing needs bringing back.** At 05:28 CEST: all 15 containers up
and healthy (`felhom-controller` 0.217.0, `traefik`, `cloudflared`, `filebrowser`, plus the six drill
apps); `/` 4%, `/var/lib/felhom` 13%, the data drive 1%; **no filler files left**, **no app-stop marker**,
**no operation in flight**, **no countdown running**, and `restic check` over its off-site repository
reports **no errors were found**. Its off-site backup, silent since 9 August, is working again and ran
on its own schedule at 04:15.
**Deliberately left in place, each with a reason** (retained, not forgotten):
| left behind | why |
|---|---|
| `opengist`, `privatebin`, `calibre-web` deployed with the planted `DRILL-2026-08-21` fixture | the reproduction for **R-354** and **R-356**; the hashes in §2 make the fix verifiable without rebuilding the case |
| `paperless-ngx` deployed | the **only** reproduction of **R-355**, and it regenerates the orphan directory on every nightly cycle |
| `romm` deployed | the only app on the box where the safety dump actually works — the fixture for **R-361** and the control for R-355 |
| `kimai` deployed | a second correct-derivation DB app, the negative control for R-355 |
| five verification copies under `backups/offsite-restore/` | harmless, and they are the **R-358** fixture |
| `romm`'s `pre-restore-…sql` in its unit | the physical evidence for **R-361** |
| DooPlex's public key in `demo-hp:/root/.ssh/authorized_keys`, and the `hp` alias in `~/.ssh/config` | the box had lost the key and its tailnet address is dead; without it the next session must go through break-glass again |
**Removed / restored during the drill:** both corrupted restic packs (repo verified clean), both
`fallocate` fillers, the set-aside store (by the sweep, as intended), and demo-hp's superseded escrow
row (by the sweep's hub half, as intended).
**`demo-hp`'s tailnet address `100.76.96.79` is still unreachable** and was not repaired — the box is
reachable on the LAN at `192.168.0.104`. That is the one thing about it that is worse than it should
be, and it predates tonight.
**`demo-felhom` — untouched and healthy**, controller 0.217.0, its own off-site ran at 02:15 with
`last_status: ok`, both fixtures verified present (§10), no countdown running.
**`drill-r50`** — not used. Still `DOWN` in the hub, agent 0.129.0, as it was.
---
## 13. TEARDOWN — ALL FOUR LAYERS, STATED
| layer | state |
|---|---|
| **Off-site (storage box)** | The set-aside store I created was **deleted by the product's own sweep**, as designed. The two corrupted packs were **restored byte-identical** and `restic check` passes. Nothing else was written. `restic forget`/`prune` were never invoked by me. |
| **PVE host `demo-hp`** | `/root/settings.json.drill-backup` **retained** (the pre-drill controller settings, in case the abandonment edit needs reverting). DooPlex's SSH key **retained**, with the reason above. `/tmp/plant.tar` left; harmless. |
| **Guest 9201 / controller** | Six apps and the planted fixtures **retained with reasons** (table above). Controller state is clean: no marker, no countdown, no in-flight op. |
| **Hub — stated explicitly** | **No customer record and no host record was created, so none needs deleting.** I used the existing `demo-hp` customer and host throughout. The only hub-side *removal* was demo-hp's superseded escrow row, done by the abandonment sweep itself and reported as event 3025. Events 3002–3025 were generated as a normal consequence of the work and are left as the record. **Nothing is owed here and nothing is blocked.** |
| **DooPlex (this machine)** | The hub DB copies (which contain every host's break-glass secret) are in the session scratchpad only, never in a committed file, and are shredded in the closing step. |
---
## 14. OBSERVATIONS — noticed, not acted on
- **The hub does not display the controller version it is told.** `demo-hp`'s host page shows the
guest's Controller column as **„—"** hours after receiving `controller_updated: 0.216.0 → 0.217.0`.
Two hub surfaces, one blind. Not filed — `REPORT-hub-blindness.md` already exists and this may be
part of it.
- **`demo-felhom` moved to 0.217.0 without a `controller_updated` event** (started 19:31Z, 17 minutes
before the floor was saved). Consistent with a by-hand deployment during the golden bake, not a
defect — recorded so a later reader does not mistake it for one.
- **`restic --latest N` is per-group, not a total.** It briefly read as "18 snapshots became 10" and
would have been reported as data loss. Caught by listing all of them and by an independent
`restic stats` (26.44 MiB / 588 blobs). **An unpaginated listing is not a total** — the rule earned
its place again.
- **`find -newermt` is the wrong probe for a restore**: restic preserves the snapshot's mtimes, so a
freshly restored tree looks old. It briefly read as "nothing was restored".
- **`sftp -b` aborts on the first failing line**, so a batch listing 256 shard directories returned 7
packs and looked like a total. Prefixing each line with `-` fixed it. Same family as the two above.
---
## 15. WHAT WAS DROPPED, PLAINLY
- **Nothing in Parts 0, 2, 3, 4 or 5 was skipped.** Every Part 4 item was attempted and reached a
verdict.
- **The fill-watch alarm was not watched firing on its own schedule.** I proved its cadence from
source (daily 03:30) and observed it fire from the startup path (`disk_critical`, 21:10:58Z, correct
Hungarian copy naming the drive and the free space). Holding a filesystem at 99% for six hours would
have sat across the scheduled backup and the abandonment sweep, and I judged the scheduled cycle —
which the task asks for explicitly — worth more than a second sighting of an alarm whose trigger I
had already read and seen work.
- **I did not drive the abandonment through the real orphan→reset path**, because that path
move-asides the *live* repository, which would have destroyed demo-hp's whole off-site history
including the 9 August snapshots this drill exists to read. The sweep's own code path was exercised
in full on a real store; the deviation is only in how the state was created.
- **`peti-felhom` was not contacted**, per the fence.