DRILL 2026-08-21: the off-site restore never replays named volumes (R-354..R-365)
gates / gates (push) Successful in 16s
gates / gates (push) Successful in 16s
Diagnostic only — no code changed, no version bumped, nothing deployed.
The verdict is a mixture. The unit and the off-site snapshot HOLD the data, proven
by identity in both storage classes including two Hungarian accented filenames. The
loss is in the last leg: ReconstituteFromOffsite skips every isUnit placement and the
volume tars live inside the unit, so the off-site full restore has no named-volume
leg at all — while the local restore-from-unit does, and returned the same tar
byte-identical minutes later.
Twelve rows opened, ceiling R-353 -> R-365. Three HIGH:
R-354 off-site restore never replays volume dumps
R-355 paperless-ngx's Postgres is dumped under a non-existent stack, so its unit
has no DB dump, no safety dump is taken, and the customer is told it has none
R-356 the off-site restore refuses for all 40 no-drive apps saying the running app
"is not installed", with a remedy those apps make impossible
R-353's instruction (2) is satisfied and annotated: the 40-class DOES reach the
off-site tier. Its instruction (1) stands and is now larger. R-329 confirmed still
live and now the only bad-severity emit fleet-wide.
Evidence: documentation/audits/DRILL-backup-truth-2026-08-21/evidence/
This commit is contained in:
@@ -652,7 +652,19 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:<HUB_PW>@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes |
|
||||
| **R-351** | **The restore never read back where the backup said the data lived, and a second press started a second restore.** Two findings, one session, both shipped. **(a) The blindness.** Every recovery-unit `manifest.json` has carried `drive` and `namespace_root` since schema 1 (`controller/internal/backup/recovery_unit.go:48-49`), written at capture from the app's own placement. `grep -rE '\.Drive\b\|\.NamespaceRoot\b' --include=*.go` found **no non-test reader anywhere** — the reconstitution opened the manifest (`offbox_reconstitute.go:235`) purely for the coherence stamp and resolved its destination from the LIVE app instead. **A restore into a destination different from the recorded one therefore succeeded silently, under a green message.** **(b) The second press.** All seven restore handlers gated on `backupMgr.IsRunning()` — the CONCURRENCY flag, acquired *inside* the goroutine (`offbox_reconstitute.go:180`) **after** the handler returned. Established with a test before any change: both the reconstitute and place handlers answered „…elindult" and **overwrote the first restore's op/stack**. The wizard had read the correct flag since v0.154.0 and said so in a comment; the handlers were never moved over. **(c)** The banner gated its terminal result on a page-local `sawRunning`, so a restore that finished before the page opened — the 8.666 s OpenGist restore — was shown to nobody. | **CLOSED 2026-08-21** — controller | — | **Shipped:** `backup/offbox_placement.go` (`CheckPlacement`, `PlacementMismatchMessage`, `RecordedUnitForStack`); mismatch **named and refused** before the safety dump, with `ack_placement` as a **separate** field from `confirm=1`; the not-installed refusal names the recorded drive; deploy page **prefills the address and folder from the app's own backup**; `Server.restoreOpBlocked()` reads BOTH flags; `RestoreOpStatus.LastRecent` + `RestoreResultWindow` moved to `internal/backup` as ONE expression for two surfaces. Red-proofs: B with **both** guards removed **was seen starting a restore with no drive attached** (no error, full 3.00 s run into `/tmp/mutant-destination`); C, E and A each returned their wrong outcome; D forced on broke 8 ordinary reconstitute tests, proving reachability both ways. | CC |
|
||||
| **R-352** | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes |
|
||||
| **R-353** | **A restore reported success having returned configuration and no data — and no screen could have told the customer.** `demo-hp`, 2026-08-21, OpenGist. The off-site reconstitution refused at 16:37:14 (not installed); the person reinstalled and ran the local unit restore, which reported `Restore-from-unit completed: opengist in 8.666896042s`. **The unit it restored from contains `manifest.json` + `compose/{app.yaml,.felhom.yml,docker-compose.yml}` and NOTHING else — `volume_dumps: None`, `db_dumps: None`** — and the off-site snapshot was **182.3 KB**. So the restore returned the app's configuration; there was no data leg in the unit to return, and the outcome said only that it had completed. **A warning beside a success is read as a success, and an unknown must never be drawn as healthy.** **Compounding, and recorded as UNKNOWN rather than fine:** whether the 40-class reaches the off-site tier at all has **not been observed** — `runVolumeDumps` (`backup/backup.go:607+`) covers them on paper, but every unit on the box reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly dump run had happened on a one-hour-old box. | **OPEN — NEXT SESSION'S FIRST ITEM** | — | **Two things, in order. (1)** A restore whose unit carries no `db_dumps` and no `volume_dumps` must **say so in its outcome** — „a mentés csak a beállításokat tartalmazta, adatot nem" — instead of reporting a bare completion. The verdict must consult what was actually placed, not merely that the operation ended. **(2)** Then *prove* the off-site coverage of a named-volume app by running a dump cycle and reading the resulting manifest, rather than inferring it from the gate order. Do not close (1) on the strength of (2) being likely. | CC |
|
||||
| **R-353** | **A restore reported success having returned configuration and no data — and no screen could have told the customer.** `demo-hp`, 2026-08-21, OpenGist. The off-site reconstitution refused at 16:37:14 (not installed); the person reinstalled and ran the local unit restore, which reported `Restore-from-unit completed: opengist in 8.666896042s`. **The unit it restored from contains `manifest.json` + `compose/{app.yaml,.felhom.yml,docker-compose.yml}` and NOTHING else — `volume_dumps: None`, `db_dumps: None`** — and the off-site snapshot was **182.3 KB**. So the restore returned the app's configuration; there was no data leg in the unit to return, and the outcome said only that it had completed. **A warning beside a success is read as a success, and an unknown must never be drawn as healthy.** **Compounding, and recorded as UNKNOWN rather than fine:** whether the 40-class reaches the off-site tier at all has **not been observed** — `runVolumeDumps` (`backup/backup.go:607+`) covers them on paper, but every unit on the box reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly dump run had happened on a one-hour-old box. | **OPEN — NEXT SESSION'S FIRST ITEM** | — | **Two things, in order. (1)** A restore whose unit carries no `db_dumps` and no `volume_dumps` must **say so in its outcome** — „a mentés csak a beállításokat tartalmazta, adatot nem" — instead of reporting a bare completion. The verdict must consult what was actually placed, not merely that the operation ended. **(2)** Then *prove* the off-site coverage of a named-volume app by running a dump cycle and reading the resulting manifest, rather than inferring it from the gate order. Do not close (1) on the strength of (2) being likely. **(2) IS NOW SATISFIED — drill 2026-08-21.** A dump cycle was run and the manifests read: `privatebin volume_dumps=[privatebin_privatebin_data.tar]`, `opengist volume_dumps=[opengist_opengist_data.tar]`, `kimai volume_dumps=[kimai_kimai_db_data.tar, kimai_kimai_var.tar]` — the 40-class DOES reach the off-site tier, and PrivateBin's planted 1 MB came back byte-identical from its off-site snapshot into the checking folder. **(1) stands and is now strictly larger than when written:** R-354 shows the bare completion is also reported over a unit that DID carry a data leg, because the off-site restore never replays volume dumps at all. | CC |
|
||||
| **R-354** | **The off-site full restore has NO named-volume leg — the tar is in the unit, in the snapshot and in the checking folder, and is never replayed.** Proven live on `demo-hp` 2026-08-21 22:23 with planted files. `calibre-web`'s `calibre_web_config` tar (1 422 848 B) was present at every stage and the restore returned 5 declared user files and **not the volume**, under „5 fájl visszaállítva … az alkalmazás újraindult". Cause: `ReconstituteFromOffsite` skips every placement flagged `isUnit` (`controller/internal/backup/offbox_reconstitute.go:341-346`) and the volume tars live INSIDE the unit; a grep for a volume-restore call across the whole off-site path returns nothing. The **local** restore does have one (`restore.go:99 restoreDockerVolumes`) — proven the same night by returning PrivateBin's planted 1 MB byte-identical from the same tar. **For the 13 drive-declaring apps the lost leg is the app's own configuration; for the 40 no-drive apps it is the entire dataset.** | **OPEN — HIGH** | — | Restore the unit's `volume-dumps/*.tar` from the SCRATCH unit on the off-site path, as `RestoreFromRecoveryUnit` already does from the live unit. Assert the CONSEQUENCE (planted bytes come back), not the mechanism. | CC |
|
||||
| **R-355** | **`paperless-ngx`'s PostgreSQL is dumped into a directory for a stack that does not exist, so its unit has never contained a database dump — and the destructive restore therefore takes no safety dump and tells the customer the app has no database.** `deriveStackName("paperless-postgres", known)` (`controller/internal/appbackup/dbdump.go:770-798`) strips the `postgres` suffix to `paperless`, finds it is NOT a known stack, finds no known stack is a prefix of the container name, and then **returns the unresolved candidate anyway** — no warning, no refusal. Observed live 2026-08-21: the dump (284 617 B, 72 tables, valid) landed in `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/` on the SYSTEM drive while the app's unit sits on `/mnt/felhom-drives/hdd_1` recording `"db_dumps": null`. The orphan directory is outside the app's off-site capture set, so the only copy of that dump is on the machine it protects. `writeSafetyDump` filters on the same wrong name, so `hasDB` is false: **no undo is taken and the fail-closed refusal cannot fire** — verified, `find /mnt -name "pre-restore-*"` empty before AND after a destructive restore. Outcome said „0 fájl visszaállítva … Ennek az alkalmazásnak nincs adatbázisa." **A catalogue-wide sweep of every DB-bearing template shows this is the ONLY affected app (1 of 53).** | **OPEN — HIGH** | — | Two candidates, both two-repo: rename the container to `paperless-ngx-postgres`, or make an unresolved candidate a loud skip rather than a silent fallback. The second is the one that generalises. Pin with a test that a container whose name resolves to no known stack is never dumped silently. | CC |
|
||||
| **R-356** | **The off-site restore refuses for all 40 no-drive apps, says the app "is not installed" when it is running, and then gives an instruction those apps make impossible.** `ReconstituteFromOffsite` refuses when `GetStackHDDPath(stack)` is empty (`offbox_reconstitute.go:208-227`); for a 40-class app that is ALWAYS empty, because they are offered no storage field at deploy time (R-352's own measurement). Observed 2026-08-21 22:21 on `privatebin` while it was `deployed=true, state=running, healthy`: „a(z) privatebin nincs telepítve, ezért nincs hová visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra az alkalmazást ugyanerre a helyre…". **The predicate is "has an HDD path"; the sentence says "is not installed"; for this class they are different things**, and the remedy offered cannot be carried out. This is also what the 2026-08-21 afternoon OpenGist journey hit before falling back to the local restore (R-353). | **OPEN — HIGH** | — | Separate the two questions. A 40-class app has a destination — the system data path — and the restore already knows it. | CC |
|
||||
| **R-357** | **The DESTRUCTIVE restore has no free-space gate; the three that exist are all on non-destructive paths.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the gates sit at `offbox_restore.go:231` (scratch restore), `:297` (prepare-full) and `:423` (place-to-live). Proven 2026-08-21 23:11 with 300 KB free and 1 MB to write: it stopped `paperless-ngx`, failed halfway (`rsync … No space left on device (28)`), left the data directory holding **2 of 5** planted entries, and restarted the app. The message is honest but is raw rsync output. | **OPEN — MEDIUM** | — | Same gate, same wording as `:297`, before the stop. | CC |
|
||||
| **R-358** | **A FAILED scratch restore leaves a partial copy that the product then offers as a full restore source — and the destructive restore runs from it and reports success.** `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only whether the directory exists and is non-empty; its comment defers completeness to `PlaceOffsiteRestore`, which stats top-level placements, not files. Proven 2026-08-21 22:54-22:56 against a deliberately corrupted store: the restore failed honestly (`ciphertext verification failed`, 54 files, 15 of 16 originals), the wizard then offered „Teljes visszaállítás indítása", and pressing it reported `ok=true`. **The failure is detected and then forgotten.** | **OPEN — MEDIUM** | — | Record the failure against the scratch and refuse to place from it until it is re-prepared. | CC |
|
||||
| **R-359** | **The off-site restic store is never verified by anything, ever.** The complete set of restic verbs in the controller is `restore, snapshots, backup, unlock, stats, init, forget, prune, cat` — **no `check`**. The agent's `RestoreTest` is PBS-tier only. Established 2026-08-21 by deliberately corrupting one pack: `restic check` catches it immediately („ciphertext verification failed", „Fatal: repository contains errors"), and the product only meets the damage when a customer is already trying to recover. | **OPEN — MEDIUM** | — | A periodic `restic check` (structure) with an occasional `--read-data`, reported like any other backup verdict. Note PBS already has verify jobs; this is the tier that does not. | CC |
|
||||
| **R-360** | **The verification-copy delete gates on the concurrency flag, which the verification restore never holds — so the copy is deletable for the whole restore, and the handler's own comment claims the opposite.** `offboxVerifyCopyDeleteHandler` (`web/offbox_handlers.go:502`) reads `backupMgr.IsRunning()`; `RestoreOffboxScratch` (`offbox_restore.go:211`) **never calls `acquireRunning`**. R-351b moved all seven restore handlers onto `restoreOpBlocked()` (both flags) and left this one behind. Demonstrated 2026-08-21 22:35 with the flags read immediately before and after: `display=True offbox-restore kimai / concurrency=False` on both sides, and the delete succeeded. The guard does not compare app names, so the same call naming the restoring app removes the directory the restore is writing into. **Observed in two consecutive reports and filed neither time; filed now.** | **OPEN — MEDIUM** | — | `restoreOpBlocked()`, and a test that asserts the CONSEQUENCE — the copy survives a delete attempt mid-restore. | CC |
|
||||
| **R-361** | **The pre-restore safety dump overwrites the app's own DB dump, and the comment beside it says it cannot.** `writeSafetyDump` calls `DumpOne`, which writes the canonical `<stack>-<dbtype>.sql` (`appbackup/dbdump.go:200-202`) — i.e. the unit's real dump — and only THEN renames it to `pre-restore-*`. The comment at `offbox_reconstitute.go:147-148` states the rename means it "can never overwrite the app's real dump". Proven 2026-08-21: `romm`'s `db-dumps/` held `romm-mariadb.sql` (62 270 B) at 22:59 and held ONLY `pre-restore-20260821T210246Z-romm-mariadb.sql` after one reconstitute. Until the next backup run the local restore-from-unit finds no `.sql` and reports the app has no database. The `pre-restore-*` file is also enumerated into the manifest's `db_dumps` and shipped off-site. | **OPEN — MEDIUM** | — | Dump to the safety name directly, or to a temp name. Pin the invariant the comment already asserts. | CC |
|
||||
| **R-362** | **A data drive detached mid-restore is reported as „permission denied".** Observed 2026-08-21 23:15: the guest-visible bind was unmounted 4 s into a scratch restore; the restore correctly failed and wrote nothing to the wrong place, but said „A visszaállítás sikertelen: restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied". The controller has a drive-state concept (`IsDisconnected`, used by both backup legs) and the restore path never consults it. **A correct refusal that misdescribes why sends the reader at a permissions problem that does not exist.** Creditable in the same test: the agent re-bound the drive 5 s later, unaided. | **OPEN — MEDIUM** | — | Consult drive state when a restore path operation fails on ENOENT/EACCES and name the drive. | CC |
|
||||
| **R-363** | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** `sched.Daily("fill-watch", "03:30", …)` (`cmd/controller/main.go:1092`) plus one startup check. Proven 2026-08-21 23:17: the 69 GB filesystem carrying the Docker data-root, the system namespace and ALL 40-class app data was filled to 99% / 1.2 GiB free; the backup reserve refused `kimai` per app and the hub received `recovery_unit_capture_failed` (error) naming the filesystem, **and the fill watcher said nothing at all**. The package comment says it "warns the CUSTOMER that a filesystem is filling, BEFORE anything fails"; at a daily cadence it frequently cannot. | **OPEN — MEDIUM** | — | The reserve already computes the same numbers every run. Let the watcher share that reading rather than owning a separate daily one. | CC |
|
||||
| **R-364** | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC |
|
||||
| **R-365** | **An overdue abandonment countdown renders its past due-date in the future tense.** With the terminal step due and the daily sweep not yet run, the card reads „A kérésed szerint a korábbi távoli mentéseidet **2026-08-20** napján véglegesen töröljük" — on 2026-08-21. The window is up to ~29 h in production (due moment → next 05:10 sweep). | **OPEN — LOW** | — | Say "due, will run at the next daily sweep" once the date has passed. | CC |
|
||||
|
||||
| **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** Both box checkers (`OffsiteBoxChecker` over the Hetzner API, `PBSDRBoxChecker` over ep0's `usage` op) held their last snapshot and returned quietly on a failed fetch. That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. During the 2026-08-18 ep0 incident the hub said nothing for the entire outage; the only mails came from the boxes' own backup failures, and **only because the WEEKLY offsite run happened to fall inside the window**. Two days earlier, nothing would have fired at all | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | — | **PROVEN-LIVE still owed.** No real or constructed outage has exercised the emit path end to end, and one cannot be manufactured without making ep0 or the Hetzner API unreachable — ep0 is Tier 2 protected, so that is forbidden. The honest route is a constructed outage against a scratch hub instance with the tenantsync client pointed at a blackholed address. **Do not close this row on the unit tests** | CC |
|
||||
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |
|
||||
|
||||
Reference in New Issue
Block a user