DRILL closeout: the scheduled cycle agrees, the abandonment sweep watched firing, R-366
gates / gates (push) Successful in 16s
gates / gates (push) Successful in 16s
The box's own 02:30 / 03:30 / 04:15 cycle ran unattended and agrees with the manual one on every measure. The one that matters: R-355 is not an artefact of manual triggering — the scheduled run again wrote paperless-ngx's PostgreSQL dump into a directory for a stack that does not exist, and again left the app's own unit recording db_dumps: null. Part 4.2's terminal deletion has now been observed. At 05:10 the sweep removed exactly the recorded set-aside store and left the live repository untouched; at 05:13 the hub dropped the sealed package that protected it and said so (event 3025). Both halves went together, three minutes apart, and the controller cleared its own state. demo-felhom's two preserved fixtures were verified untouched throughout. R-366 (HIGH) filed, found incidentally: the 21 August reinstall orphaned demo-hp's PBS whole-guest archives as well as its restic repo, so a rebuilt box loses BOTH off-premises tiers at once. The restore-test caught it and named the key mismatch precisely; it is merely called "a failed restore test" rather than "your older backups are unreadable". demo-hp is left HEALTHY, not broken. Ceiling R-353 -> R-366.
This commit is contained in:
@@ -245,9 +245,33 @@ leads, and the log line alone says only „backup OK".
|
||||
|
||||
---
|
||||
|
||||
## 5. THE TWO CYCLES COMPARED
|
||||
## 5. THE TWO CYCLES COMPARED — AND THEY AGREE
|
||||
|
||||
*(filled in after the scheduled run — see §11)*
|
||||
**By hand:** local cycle 22:12, off-site 22:17 (after enabling the per-app switches the rebuild had
|
||||
silently cleared). **By the clock:** the box's own `db-dump` at 02:30 CEST, cross-drive at 03:30,
|
||||
off-site at 04:15.
|
||||
|
||||
| | manual | scheduled | agree? |
|
||||
|---|---|---|---|
|
||||
| units refreshed | all 6 apps | all 6 apps, `00:30Z` | **yes** |
|
||||
| `privatebin` volume tar | 1 055 744 B | 1 055 744 B | **yes** |
|
||||
| `opengist` volume tar | 181 248 B | 181 248 B | **yes** |
|
||||
| `calibre-web` config tar | 1 422 848 B | 368 640 B † | **yes, and explained** |
|
||||
| `paperless-ngx` `db_dumps` | **absent** | **absent** | **yes — the defect reproduces on the scheduled path** |
|
||||
| the orphan `…/primary/paperless/db-dumps/` | written | **written again, 294 936 B, `00:30Z`** | **yes** |
|
||||
| planted files on the drive | 5/5 identical | **5/5 identical**, plus `POST-SNAPSHOT.txt` | **yes** |
|
||||
| off-site run | ok, 27 → snapshots | ok, `02:17:12Z`, **2m8s**, 42.8 MB, **27 snapshots** | **yes** |
|
||||
|
||||
† the tar shrank because the drill had by then removed the planted 1 MB from that volume — the
|
||||
expected value, not a discrepancy.
|
||||
|
||||
**Two independent observations, no disagreement.** The one that matters: **R-355 is not an artefact of
|
||||
my manual triggering.** The box, unattended, on its own schedule, wrote paperless-ngx's PostgreSQL
|
||||
dump into a directory for a stack that does not exist and left the app's own unit recording
|
||||
`db_dumps: null`.
|
||||
|
||||
The scheduled cross-drive (Tier-2) leg also completed for all six apps at 01:30Z — `crossdrive_completed`
|
||||
events 3019–3024.
|
||||
|
||||
---
|
||||
|
||||
@@ -347,7 +371,7 @@ means in all three cases.
|
||||
|
||||
## 8. RANKED REGISTER ROWS OPENED
|
||||
|
||||
Ceiling was **R-353**; it **moved to R-365**.
|
||||
Ceiling was **R-353**; it **moved to R-366**.
|
||||
|
||||
| id | rank | what |
|
||||
|---|---|---|
|
||||
@@ -363,6 +387,7 @@ Ceiling was **R-353**; it **moved to R-365**.
|
||||
| **R-363** | **MEDIUM** | The fill watcher runs **once a day (03:30)** plus at startup. A filesystem that fills at 03:31 is unannounced for ~24 h — while the backup reserve is already refusing apps. |
|
||||
| **R-364** | **LOW** | Accented-text search is a silently-transforming instrument; discipline has failed at least three times. Guard proposed, not built. **(Part 5.2)** |
|
||||
| **R-365** | **LOW** | An **overdue** abandonment countdown renders its past due-date in the future tense („…2026-08-20 napján véglegesen töröljük" shown on 2026-08-21). |
|
||||
| **R-366** | **HIGH** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives too** — the box cannot read its own pre-reinstall backups (`manifest's key 3f:4f:65:c0… does not match provided key dd:d1:d8:53…`, hub event 3016, filed unprompted at 21:59). The PBS-tier analogue of R-193: a rebuilt box loses **both** off-premises tiers at once. The restore-test caught it precisely; the gap is that it is called *a failed restore test* rather than *your older whole-guest backups are unreadable*. **Found incidentally — nobody was looking for it.** |
|
||||
|
||||
**Confirmed still live, not re-filed:** **R-329** — `app_start_failed` emits severity `"warn"`
|
||||
(`internal/notify/notifier.go:546`), outside the hub's vocabulary, so it coerces to `info` and emails
|
||||
@@ -440,12 +465,135 @@ over a unit that *did* have a data leg.
|
||||
|
||||
---
|
||||
|
||||
## 11. SCHEDULED CYCLE AND THE COUNTDOWN
|
||||
## 11. PART 4.2 — THE ABANDONMENT COUNTDOWN, WATCHED FIRING
|
||||
|
||||
*(filled in as they fire — 02:30 local backup, 04:15 off-site, 05:10 abandonment sweep, all CEST)*
|
||||
**The terminal deletion has now been observed.** It did exactly what it claims, including the half
|
||||
nobody had seen.
|
||||
|
||||
- **State created** 23:08–23:09 CEST. A realistic set-aside store was built at
|
||||
`u629488-sub3:/home/felhom-repo-superseded-drill-20260821`, the countdown written overdue
|
||||
(started 2026-08-07, due 2026-08-20), and the product's own CLI confirmed it:
|
||||
`abandonment countdown RUNNING … deleted on: 2026-08-20 … days left: 0`.
|
||||
- **05:10 CEST — it fired.** `/home` on the storage box is stamped `03:10Z`; the set-aside store is
|
||||
**gone**; **`/home/felhom-repo`, the live repository, is untouched** (mtime still 4 Aug).
|
||||
- **05:13 CEST — the hub half.** Event **3025 `offsite_abandon_purged`**: *„Az ügyfél korábbi távoli
|
||||
mentései és a hozzájuk tartozó megőrzött helyreállítási csomag is törölve (1 csomag). Az ügyfél
|
||||
döntése alapján, a 14 napos türelmi idő lejárta után."* And demo-hp's superseded escrow row **is
|
||||
gone from the hub** — it held one before.
|
||||
- **The controller closed itself out.** Every `abandon_*` field has been removed from `settings.json`
|
||||
and `--abandon-status` reports *"no abandonment countdown is running on this box"*.
|
||||
|
||||
**So the two-phase commit's central promise — *"it removes BOTH halves or neither"* — is confirmed
|
||||
live for the first time:** the ciphertext and the sealed package that protects it went together, three
|
||||
minutes apart, and the state that remembered the operation cleaned itself up. **What it removes:**
|
||||
exactly the recorded set-aside path. **What survives:** the live repository, the live escrow, and the
|
||||
box's current recovery path.
|
||||
|
||||
**Judged:** the completion message is accurate and in plain Hungarian. The only wrong note is while
|
||||
the countdown is *overdue but not yet swept* — the card then states a past date in the future tense
|
||||
(**R-365**).
|
||||
|
||||
**No countdown is left running anywhere.** `demo-hp` cleared itself; `demo-felhom` reports none.
|
||||
|
||||
---
|
||||
|
||||
## 11b. WHAT THE NIGHT FOUND THAT NOBODY WAS LOOKING FOR
|
||||
|
||||
At 21:59 the box filed, unprompted, hub event 3016:
|
||||
|
||||
> `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could
|
||||
> not be restored+booted … wrong key — manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided
|
||||
> key dd:d1:d8:53:44:62:5e:0b`
|
||||
|
||||
That archive predates the 21 August reinstall by three days. **The rebuild orphaned the PBS whole-guest
|
||||
archives exactly as it orphaned the restic repository** — so a rebuilt box loses *both* off-premises
|
||||
tiers at once. The restore-test mechanism deserves credit: it caught it and named the key mismatch
|
||||
precisely. The gap is what it is *called* — "a restore test failed" reads as a flaky verification, not
|
||||
as "every whole-guest backup taken before the reinstall is unreadable on this machine". Filed **R-366
|
||||
(HIGH)**.
|
||||
|
||||
---
|
||||
|
||||
## 12. MACHINE STATES AT THE END
|
||||
|
||||
*(final state recorded in §13 after the scheduled runs)*
|
||||
**`demo-hp` — HEALTHY, not broken. Nothing needs bringing back.** At 05:28 CEST: all 15 containers up
|
||||
and healthy (`felhom-controller` 0.217.0, `traefik`, `cloudflared`, `filebrowser`, plus the six drill
|
||||
apps); `/` 4%, `/var/lib/felhom` 13%, the data drive 1%; **no filler files left**, **no app-stop marker**,
|
||||
**no operation in flight**, **no countdown running**, and `restic check` over its off-site repository
|
||||
reports **no errors were found**. Its off-site backup, silent since 9 August, is working again and ran
|
||||
on its own schedule at 04:15.
|
||||
|
||||
**Deliberately left in place, each with a reason** (retained, not forgotten):
|
||||
|
||||
| left behind | why |
|
||||
|---|---|
|
||||
| `opengist`, `privatebin`, `calibre-web` deployed with the planted `DRILL-2026-08-21` fixture | the reproduction for **R-354** and **R-356**; the hashes in §2 make the fix verifiable without rebuilding the case |
|
||||
| `paperless-ngx` deployed | the **only** reproduction of **R-355**, and it regenerates the orphan directory on every nightly cycle |
|
||||
| `romm` deployed | the only app on the box where the safety dump actually works — the fixture for **R-361** and the control for R-355 |
|
||||
| `kimai` deployed | a second correct-derivation DB app, the negative control for R-355 |
|
||||
| five verification copies under `backups/offsite-restore/` | harmless, and they are the **R-358** fixture |
|
||||
| `romm`'s `pre-restore-…sql` in its unit | the physical evidence for **R-361** |
|
||||
| DooPlex's public key in `demo-hp:/root/.ssh/authorized_keys`, and the `hp` alias in `~/.ssh/config` | the box had lost the key and its tailnet address is dead; without it the next session must go through break-glass again |
|
||||
|
||||
**Removed / restored during the drill:** both corrupted restic packs (repo verified clean), both
|
||||
`fallocate` fillers, the set-aside store (by the sweep, as intended), and demo-hp's superseded escrow
|
||||
row (by the sweep's hub half, as intended).
|
||||
|
||||
**`demo-hp`'s tailnet address `100.76.96.79` is still unreachable** and was not repaired — the box is
|
||||
reachable on the LAN at `192.168.0.104`. That is the one thing about it that is worse than it should
|
||||
be, and it predates tonight.
|
||||
|
||||
**`demo-felhom` — untouched and healthy**, controller 0.217.0, its own off-site ran at 02:15 with
|
||||
`last_status: ok`, both fixtures verified present (§10), no countdown running.
|
||||
|
||||
**`drill-r50`** — not used. Still `DOWN` in the hub, agent 0.129.0, as it was.
|
||||
|
||||
---
|
||||
|
||||
## 13. TEARDOWN — ALL FOUR LAYERS, STATED
|
||||
|
||||
| layer | state |
|
||||
|---|---|
|
||||
| **Off-site (storage box)** | The set-aside store I created was **deleted by the product's own sweep**, as designed. The two corrupted packs were **restored byte-identical** and `restic check` passes. Nothing else was written. `restic forget`/`prune` were never invoked by me. |
|
||||
| **PVE host `demo-hp`** | `/root/settings.json.drill-backup` **retained** (the pre-drill controller settings, in case the abandonment edit needs reverting). DooPlex's SSH key **retained**, with the reason above. `/tmp/plant.tar` left; harmless. |
|
||||
| **Guest 9201 / controller** | Six apps and the planted fixtures **retained with reasons** (table above). Controller state is clean: no marker, no countdown, no in-flight op. |
|
||||
| **Hub — stated explicitly** | **No customer record and no host record was created, so none needs deleting.** I used the existing `demo-hp` customer and host throughout. The only hub-side *removal* was demo-hp's superseded escrow row, done by the abandonment sweep itself and reported as event 3025. Events 3002–3025 were generated as a normal consequence of the work and are left as the record. **Nothing is owed here and nothing is blocked.** |
|
||||
| **DooPlex (this machine)** | The hub DB copies (which contain every host's break-glass secret) are in the session scratchpad only, never in a committed file, and are shredded in the closing step. |
|
||||
|
||||
---
|
||||
|
||||
## 14. OBSERVATIONS — noticed, not acted on
|
||||
|
||||
- **The hub does not display the controller version it is told.** `demo-hp`'s host page shows the
|
||||
guest's Controller column as **„—"** hours after receiving `controller_updated: 0.216.0 → 0.217.0`.
|
||||
Two hub surfaces, one blind. Not filed — `REPORT-hub-blindness.md` already exists and this may be
|
||||
part of it.
|
||||
- **`demo-felhom` moved to 0.217.0 without a `controller_updated` event** (started 19:31Z, 17 minutes
|
||||
before the floor was saved). Consistent with a by-hand deployment during the golden bake, not a
|
||||
defect — recorded so a later reader does not mistake it for one.
|
||||
- **`restic --latest N` is per-group, not a total.** It briefly read as "18 snapshots became 10" and
|
||||
would have been reported as data loss. Caught by listing all of them and by an independent
|
||||
`restic stats` (26.44 MiB / 588 blobs). **An unpaginated listing is not a total** — the rule earned
|
||||
its place again.
|
||||
- **`find -newermt` is the wrong probe for a restore**: restic preserves the snapshot's mtimes, so a
|
||||
freshly restored tree looks old. It briefly read as "nothing was restored".
|
||||
- **`sftp -b` aborts on the first failing line**, so a batch listing 256 shard directories returned 7
|
||||
packs and looked like a total. Prefixing each line with `-` fixed it. Same family as the two above.
|
||||
|
||||
---
|
||||
|
||||
## 15. WHAT WAS DROPPED, PLAINLY
|
||||
|
||||
- **Nothing in Parts 0, 2, 3, 4 or 5 was skipped.** Every Part 4 item was attempted and reached a
|
||||
verdict.
|
||||
- **The fill-watch alarm was not watched firing on its own schedule.** I proved its cadence from
|
||||
source (daily 03:30) and observed it fire from the startup path (`disk_critical`, 21:10:58Z, correct
|
||||
Hungarian copy naming the drive and the free space). Holding a filesystem at 99% for six hours would
|
||||
have sat across the scheduled backup and the abandonment sweep, and I judged the scheduled cycle —
|
||||
which the task asks for explicitly — worth more than a second sighting of an alarm whose trigger I
|
||||
had already read and seen work.
|
||||
- **I did not drive the abandonment through the real orphan→reset path**, because that path
|
||||
move-asides the *live* repository, which would have destroyed demo-hp's whole off-site history
|
||||
including the 9 August snapshots this drill exists to read. The sweep's own code path was exercised
|
||||
in full on a real store; the deviation is only in how the state was created.
|
||||
- **`peti-felhom` was not contacted**, per the fence.
|
||||
|
||||
+23
@@ -0,0 +1,23 @@
|
||||
PART 4.2 — the abandonment sweep, watched firing. 2026-08-22.
|
||||
|
||||
State created 2026-08-21 23:08-23:09 CEST:
|
||||
set-aside store u629488-sub3:/home/felhom-repo-superseded-drill-20260821 (config/data/index/snapshots)
|
||||
countdown started 2026-08-07, due 2026-08-20 (written into settings.json, controller restarted)
|
||||
product CLI "abandonment countdown RUNNING … deleted on: 2026-08-20 … days left: 0"
|
||||
|
||||
05:10 CEST — the daily `offsite-abandon-sweep` job fired.
|
||||
/home on the storage box now stamped 03:10Z
|
||||
/home/felhom-repo-superseded-drill-20260821 -> GONE
|
||||
/home/felhom-repo (the LIVE repository) -> present, mtime still Aug 4 ** survived **
|
||||
|
||||
05:13 CEST — the hub half. Event 3025 `offsite_abandon_purged`:
|
||||
"Az ügyfél korábbi távoli mentései és a hozzájuk tartozó megőrzött helyreállítási csomag is
|
||||
törölve (1 csomag). Az ügyfél döntése alapján, a 14 napos türelmi idő lejárta után."
|
||||
demo-hp's row in host_escrow_superseded: REMOVED (there was exactly one).
|
||||
demo-felhom's rows 4, 11, 12: ALL PRESENT — the protected fixtures were not touched.
|
||||
|
||||
Controller closed itself out: every abandon_* field removed from settings.json;
|
||||
`--abandon-status` reports "no abandonment countdown is running on this box".
|
||||
|
||||
VERDICT: the two-phase commit's promise — "it removes BOTH halves or neither" — is confirmed
|
||||
live for the first time. It removed exactly the recorded set-aside path and nothing else.
|
||||
+71
@@ -0,0 +1,71 @@
|
||||
[2026-08-21 23:34:08 CEST] collector v2 started (epoch waits)
|
||||
[2026-08-22 02:38:09 CEST] === after the 02:30 SCHEDULED local cycle
|
||||
[2026-08-22 02:38:09 CEST] --- units:
|
||||
## /mnt/sys_drive/felhom-data/backups/primary/kimai
|
||||
2235 2026-08-21 21:00 compose/.felhom.yml
|
||||
488 2026-08-21 21:00 compose/app.yaml
|
||||
2195 2026-08-21 21:00 compose/docker-compose.yml
|
||||
48217 2026-08-22 00:30 db-dumps/kimai-mariadb.sql
|
||||
1275 2026-08-21 21:00 manifest.json
|
||||
160331776 2026-08-22 00:30 volume-dumps/kimai_kimai_db_data.tar
|
||||
52845056 2026-08-22 00:30 volume-dumps/kimai_kimai_var.tar
|
||||
## /mnt/sys_drive/felhom-data/backups/primary/opengist
|
||||
1750 2026-08-21 21:00 compose/.felhom.yml
|
||||
287 2026-08-21 21:00 compose/app.yaml
|
||||
1260 2026-08-21 21:00 compose/docker-compose.yml
|
||||
1121 2026-08-21 21:00 manifest.json
|
||||
181248 2026-08-22 00:30 volume-dumps/opengist_opengist_data.tar
|
||||
## /mnt/sys_drive/felhom-data/backups/primary/paperless
|
||||
294936 2026-08-22 00:30 db-dumps/paperless-postgres.sql
|
||||
## /mnt/sys_drive/felhom-data/backups/primary/privatebin
|
||||
1723 2026-08-21 21:00 compose/.felhom.yml
|
||||
288 2026-08-21 21:00 compose/app.yaml
|
||||
1223 2026-08-21 21:00 compose/docker-compose.yml
|
||||
1118 2026-08-21 21:00 manifest.json
|
||||
1055744 2026-08-22 00:30 volume-dumps/privatebin_privatebin_data.tar
|
||||
## /mnt/felhom-drives/hdd_1/backups/primary/calibre-web
|
||||
3035 2026-08-21 21:00 compose/.felhom.yml
|
||||
327 2026-08-21 21:00 compose/app.yaml
|
||||
2122 2026-08-21 21:00 compose/docker-compose.yml
|
||||
1165 2026-08-21 21:00 manifest.json
|
||||
368640 2026-08-22 00:30 volume-dumps/calibre-web_calibre_web_config.tar
|
||||
## /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx
|
||||
5971 2026-08-21 21:00 compose/.felhom.yml
|
||||
664 2026-08-21 21:00 compose/app.yaml
|
||||
5802 2026-08-21 21:00 compose/docker-compose.yml
|
||||
1462 2026-08-21 21:00 manifest.json
|
||||
231424 2026-08-22 00:30 volume-dumps/paperless-ngx_paperless_data.tar
|
||||
71417344 2026-08-22 00:30 volume-dumps/paperless-ngx_paperless_postgres_data.tar
|
||||
116736 2026-08-22 00:30 volume-dumps/paperless-ngx_paperless_redis_data.tar
|
||||
## /mnt/felhom-drives/hdd_1/backups/primary/romm
|
||||
5520 2026-08-21 21:20 compose/.felhom.yml
|
||||
576 2026-08-21 21:20 compose/app.yaml
|
||||
4399 2026-08-21 21:20 compose/docker-compose.yml
|
||||
62270 2026-08-21 21:02 db-dumps/pre-restore-20260821T210246Z-romm-mariadb.sql
|
||||
62270 2026-08-22 00:30 db-dumps/romm-mariadb.sql
|
||||
1466 2026-08-21 21:20 manifest.json
|
||||
2560 2026-08-22 00:30 volume-dumps/romm_romm_config.tar
|
||||
160247296 2026-08-22 00:30 volume-dumps/romm_romm_db_data.tar
|
||||
14677504 2026-08-22 00:30 volume-dumps/romm_romm_redis_data.tar
|
||||
[2026-08-22 02:38:10 CEST] --- planted files still byte-identical (calibre-web books):
|
||||
eee5880e304b27c13d205aac9989e901a5f2393be7ecf076ee2c3a0195a2358c DRILL-2026-08-21/POST-SNAPSHOT.txt
|
||||
0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991 DRILL-2026-08-21/SENTINEL.txt
|
||||
725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a DRILL-2026-08-21/binary-1mb.bin
|
||||
a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4 DRILL-2026-08-21/nested/őszibarack.md
|
||||
07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1 DRILL-2026-08-21/plain.txt
|
||||
0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea DRILL-2026-08-21/árvíztűrő-tükörfúrógép.txt
|
||||
[2026-08-22 02:38:12 CEST] --- scheduled db-dump log:
|
||||
grep: /opt/docker/felhom-controller/data/debug-ring.log: No such file or directory
|
||||
[2026-08-22 04:30:13 CEST] === after the 04:15 SCHEDULED off-site run
|
||||
{"last_duration":"2m8s","last_error":"","last_run":"2026-08-22T02:17:12Z","orphaned":false,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elapsed_sec":0,"phase":""},"repo_size_human":"42.8 MB","snapshots":27,"status":"ok"}
|
||||
|
||||
[2026-08-22 04:30:13 CEST] --- offsite log:
|
||||
grep: /opt/docker/felhom-controller/data/debug-ring.log: No such file or directory
|
||||
[2026-08-22 05:20:15 CEST] === after the 05:10 abandonment sweep
|
||||
[2026-08-22 05:20:15 CEST] --- sweep log:
|
||||
grep: /opt/docker/felhom-controller/data/debug-ring.log: No such file or directory
|
||||
[2026-08-22 05:20:16 CEST] --- product CLI:
|
||||
[INFO] [settings] Loaded settings from /opt/docker/felhom-controller/data/settings.json
|
||||
no abandonment countdown is running on this box
|
||||
[2026-08-22 05:20:17 CEST] --- settings abandon fields:
|
||||
[2026-08-22 05:20:18 CEST] collector finished
|
||||
@@ -665,6 +665,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-363** | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** `sched.Daily("fill-watch", "03:30", …)` (`cmd/controller/main.go:1092`) plus one startup check. Proven 2026-08-21 23:17: the 69 GB filesystem carrying the Docker data-root, the system namespace and ALL 40-class app data was filled to 99% / 1.2 GiB free; the backup reserve refused `kimai` per app and the hub received `recovery_unit_capture_failed` (error) naming the filesystem, **and the fill watcher said nothing at all**. The package comment says it "warns the CUSTOMER that a filesystem is filling, BEFORE anything fails"; at a daily cadence it frequently cannot. | **OPEN — MEDIUM** | — | The reserve already computes the same numbers every run. Let the watcher share that reading rather than owning a separate daily one. | CC |
|
||||
| **R-364** | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC |
|
||||
| **R-365** | **An overdue abandonment countdown renders its past due-date in the future tense.** With the terminal step due and the daily sweep not yet run, the card reads „A kérésed szerint a korábbi távoli mentéseidet **2026-08-20** napján véglegesen töröljük" — on 2026-08-21. The window is up to ~29 h in production (due moment → next 05:10 sweep). | **OPEN — LOW** | — | Say "due, will run at the next daily sweep" once the date has passed. | CC |
|
||||
| **R-366** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC |
|
||||
|
||||
| **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** Both box checkers (`OffsiteBoxChecker` over the Hetzner API, `PBSDRBoxChecker` over ep0's `usage` op) held their last snapshot and returned quietly on a failed fetch. That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. During the 2026-08-18 ep0 incident the hub said nothing for the entire outage; the only mails came from the boxes' own backup failures, and **only because the WEEKLY offsite run happened to fall inside the window**. Two days earlier, nothing would have fired at all | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | — | **PROVEN-LIVE still owed.** No real or constructed outage has exercised the emit path end to end, and one cannot be manufactured without making ep0 or the Hetzner API unreachable — ep0 is Tier 2 protected, so that is forbidden. The honest route is a constructed outage against a scratch hub instance with the tenantsync client pointed at a blackholed address. **Do not close this row on the unit tests** | CC |
|
||||
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |
|
||||
|
||||
Reference in New Issue
Block a user