DRILL closeout: the scheduled cycle agrees, the abandonment sweep watched firing, R-366
gates / gates (push) Successful in 16s

The box's own 02:30 / 03:30 / 04:15 cycle ran unattended and agrees with the manual one
on every measure. The one that matters: R-355 is not an artefact of manual triggering —
the scheduled run again wrote paperless-ngx's PostgreSQL dump into a directory for a
stack that does not exist, and again left the app's own unit recording db_dumps: null.

Part 4.2's terminal deletion has now been observed. At 05:10 the sweep removed exactly
the recorded set-aside store and left the live repository untouched; at 05:13 the hub
dropped the sealed package that protected it and said so (event 3025). Both halves went
together, three minutes apart, and the controller cleared its own state. demo-felhom's
two preserved fixtures were verified untouched throughout.

R-366 (HIGH) filed, found incidentally: the 21 August reinstall orphaned demo-hp's PBS
whole-guest archives as well as its restic repo, so a rebuilt box loses BOTH off-premises
tiers at once. The restore-test caught it and named the key mismatch precisely; it is
merely called "a failed restore test" rather than "your older backups are unreadable".

demo-hp is left HEALTHY, not broken. Ceiling R-353 -> R-366.
This commit is contained in:
2026-08-22 05:25:30 +02:00
parent d895d9f7dd
commit 7064596c2e
4 changed files with 249 additions and 6 deletions
@@ -0,0 +1,23 @@
PART 4.2 — the abandonment sweep, watched firing. 2026-08-22.
State created 2026-08-21 23:08-23:09 CEST:
set-aside store u629488-sub3:/home/felhom-repo-superseded-drill-20260821 (config/data/index/snapshots)
countdown started 2026-08-07, due 2026-08-20 (written into settings.json, controller restarted)
product CLI "abandonment countdown RUNNING … deleted on: 2026-08-20 … days left: 0"
05:10 CEST — the daily `offsite-abandon-sweep` job fired.
/home on the storage box now stamped 03:10Z
/home/felhom-repo-superseded-drill-20260821 -> GONE
/home/felhom-repo (the LIVE repository) -> present, mtime still Aug 4 ** survived **
05:13 CEST — the hub half. Event 3025 `offsite_abandon_purged`:
"Az ügyfél korábbi távoli mentései és a hozzájuk tartozó megőrzött helyreállítási csomag is
törölve (1 csomag). Az ügyfél döntése alapján, a 14 napos türelmi idő lejárta után."
demo-hp's row in host_escrow_superseded: REMOVED (there was exactly one).
demo-felhom's rows 4, 11, 12: ALL PRESENT — the protected fixtures were not touched.
Controller closed itself out: every abandon_* field removed from settings.json;
`--abandon-status` reports "no abandonment countdown is running on this box".
VERDICT: the two-phase commit's promise — "it removes BOTH halves or neither" — is confirmed
live for the first time. It removed exactly the recorded set-aside path and nothing else.
@@ -0,0 +1,71 @@
[2026-08-21 23:34:08 CEST] collector v2 started (epoch waits)
[2026-08-22 02:38:09 CEST] === after the 02:30 SCHEDULED local cycle
[2026-08-22 02:38:09 CEST] --- units:
## /mnt/sys_drive/felhom-data/backups/primary/kimai
2235 2026-08-21 21:00 compose/.felhom.yml
488 2026-08-21 21:00 compose/app.yaml
2195 2026-08-21 21:00 compose/docker-compose.yml
48217 2026-08-22 00:30 db-dumps/kimai-mariadb.sql
1275 2026-08-21 21:00 manifest.json
160331776 2026-08-22 00:30 volume-dumps/kimai_kimai_db_data.tar
52845056 2026-08-22 00:30 volume-dumps/kimai_kimai_var.tar
## /mnt/sys_drive/felhom-data/backups/primary/opengist
1750 2026-08-21 21:00 compose/.felhom.yml
287 2026-08-21 21:00 compose/app.yaml
1260 2026-08-21 21:00 compose/docker-compose.yml
1121 2026-08-21 21:00 manifest.json
181248 2026-08-22 00:30 volume-dumps/opengist_opengist_data.tar
## /mnt/sys_drive/felhom-data/backups/primary/paperless
294936 2026-08-22 00:30 db-dumps/paperless-postgres.sql
## /mnt/sys_drive/felhom-data/backups/primary/privatebin
1723 2026-08-21 21:00 compose/.felhom.yml
288 2026-08-21 21:00 compose/app.yaml
1223 2026-08-21 21:00 compose/docker-compose.yml
1118 2026-08-21 21:00 manifest.json
1055744 2026-08-22 00:30 volume-dumps/privatebin_privatebin_data.tar
## /mnt/felhom-drives/hdd_1/backups/primary/calibre-web
3035 2026-08-21 21:00 compose/.felhom.yml
327 2026-08-21 21:00 compose/app.yaml
2122 2026-08-21 21:00 compose/docker-compose.yml
1165 2026-08-21 21:00 manifest.json
368640 2026-08-22 00:30 volume-dumps/calibre-web_calibre_web_config.tar
## /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx
5971 2026-08-21 21:00 compose/.felhom.yml
664 2026-08-21 21:00 compose/app.yaml
5802 2026-08-21 21:00 compose/docker-compose.yml
1462 2026-08-21 21:00 manifest.json
231424 2026-08-22 00:30 volume-dumps/paperless-ngx_paperless_data.tar
71417344 2026-08-22 00:30 volume-dumps/paperless-ngx_paperless_postgres_data.tar
116736 2026-08-22 00:30 volume-dumps/paperless-ngx_paperless_redis_data.tar
## /mnt/felhom-drives/hdd_1/backups/primary/romm
5520 2026-08-21 21:20 compose/.felhom.yml
576 2026-08-21 21:20 compose/app.yaml
4399 2026-08-21 21:20 compose/docker-compose.yml
62270 2026-08-21 21:02 db-dumps/pre-restore-20260821T210246Z-romm-mariadb.sql
62270 2026-08-22 00:30 db-dumps/romm-mariadb.sql
1466 2026-08-21 21:20 manifest.json
2560 2026-08-22 00:30 volume-dumps/romm_romm_config.tar
160247296 2026-08-22 00:30 volume-dumps/romm_romm_db_data.tar
14677504 2026-08-22 00:30 volume-dumps/romm_romm_redis_data.tar
[2026-08-22 02:38:10 CEST] --- planted files still byte-identical (calibre-web books):
eee5880e304b27c13d205aac9989e901a5f2393be7ecf076ee2c3a0195a2358c DRILL-2026-08-21/POST-SNAPSHOT.txt
0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991 DRILL-2026-08-21/SENTINEL.txt
725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a DRILL-2026-08-21/binary-1mb.bin
a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4 DRILL-2026-08-21/nested/őszibarack.md
07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1 DRILL-2026-08-21/plain.txt
0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea DRILL-2026-08-21/árvíztűrő-tükörfúrógép.txt
[2026-08-22 02:38:12 CEST] --- scheduled db-dump log:
grep: /opt/docker/felhom-controller/data/debug-ring.log: No such file or directory
[2026-08-22 04:30:13 CEST] === after the 04:15 SCHEDULED off-site run
{"last_duration":"2m8s","last_error":"","last_run":"2026-08-22T02:17:12Z","orphaned":false,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elapsed_sec":0,"phase":""},"repo_size_human":"42.8 MB","snapshots":27,"status":"ok"}
[2026-08-22 04:30:13 CEST] --- offsite log:
grep: /opt/docker/felhom-controller/data/debug-ring.log: No such file or directory
[2026-08-22 05:20:15 CEST] === after the 05:10 abandonment sweep
[2026-08-22 05:20:15 CEST] --- sweep log:
grep: /opt/docker/felhom-controller/data/debug-ring.log: No such file or directory
[2026-08-22 05:20:16 CEST] --- product CLI:
[INFO] [settings] Loaded settings from /opt/docker/felhom-controller/data/settings.json
no abandonment countdown is running on this box
[2026-08-22 05:20:17 CEST] --- settings abandon fields:
[2026-08-22 05:20:18 CEST] collector finished
+1
View File
@@ -665,6 +665,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-363** | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** `sched.Daily("fill-watch", "03:30", …)` (`cmd/controller/main.go:1092`) plus one startup check. Proven 2026-08-21 23:17: the 69 GB filesystem carrying the Docker data-root, the system namespace and ALL 40-class app data was filled to 99% / 1.2 GiB free; the backup reserve refused `kimai` per app and the hub received `recovery_unit_capture_failed` (error) naming the filesystem, **and the fill watcher said nothing at all**. The package comment says it "warns the CUSTOMER that a filesystem is filling, BEFORE anything fails"; at a daily cadence it frequently cannot. | **OPEN — MEDIUM** | — | The reserve already computes the same numbers every run. Let the watcher share that reading rather than owning a separate daily one. | CC |
| **R-364** | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC |
| **R-365** | **An overdue abandonment countdown renders its past due-date in the future tense.** With the terminal step due and the daily sweep not yet run, the card reads „A kérésed szerint a korábbi távoli mentéseidet **2026-08-20** napján véglegesen töröljük" — on 2026-08-21. The window is up to ~29 h in production (due moment → next 05:10 sweep). | **OPEN — LOW** | — | Say "due, will run at the next daily sweep" once the date has passed. | CC |
| **R-366** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC |
| **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** Both box checkers (`OffsiteBoxChecker` over the Hetzner API, `PBSDRBoxChecker` over ep0's `usage` op) held their last snapshot and returned quietly on a failed fetch. That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. During the 2026-08-18 ep0 incident the hub said nothing for the entire outage; the only mails came from the boxes' own backup failures, and **only because the WEEKLY offsite run happened to fall inside the window**. Two days earlier, nothing would have fired at all | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | — | **PROVEN-LIVE still owed.** No real or constructed outage has exercised the emit path end to end, and one cannot be manufactured without making ep0 or the Hetzner API unreachable — ep0 is Tier 2 protected, so that is forbidden. The honest route is a constructed outage against a scratch hub instance with the tenantsync client pointed at a blackholed address. **Do not close this row on the unit tests** | CC |
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |