soak phase 5: R-403 guard PROVEN live; R-412 CORRECTED down after measuring the mechanism
gates / gates (push) Failing after 18s
gates / gates (push) Failing after 18s
R-403 MIRROR GUARD - PASS, forced after the natural test evaporated. privatebin was injected hollow at 23:34 to meet the 03:30 mirror; the 02:30 db-dump re-made its tar, so by 03:30 the primary was complete and the guard had nothing to refuse. Forced instead on calibre-web through the real Tier-2 path: the guard fired and named itself - "unit leg SKIPPED ... The copy was PRESERVED rather than replaced with an empty one (R-403). The other legs continue." Secondary byte-identical, 23 files, tar sha d7e7f422. R-412 CORRECTED, AND I OVERSTATED IT WHEN I FILED IT. The first wording claimed the hollow unit sits in the store for a whole cycle because the volume-dump leg runs only on the backup schedule. That is WRONG. The 04:15 off-site run has its OWN pre-push dump leg - "Stopping calibre-web for safe volume dump", "Volume dump: ... -> 877.5 KB" - so a unit that is hollow when a run starts is REPAIRED before it is pushed. Measured twice: opengist and calibre-web both went in hollow and came out complete, and the snapshot pulled back from the store (6fee3b5a) holds the volume tar and all 17 userdata files. What remains real is narrower: the one hollow snapshot that DID reach the store was created when the unit was destroyed INSIDE a run that had already completed that app's dump leg. The race is real and was observed, and "backed up opengist (... 0 mandatory path(s))" is a success line over a backup holding none of the app's data either way. Severity HIGH -> LOW, with the correction stated in the row rather than quietly rewritten. Phase 6 interim: the observer is clean so far - db-dump 674ms, tier2-backup 3ms (a no-op, cause to be established not assumed), zero ERROR/WARN since 23:00. Also recorded: two of Phase 5's four injections were NOT performed, with the reasons established rather than asserted - there is no endpoint that reaches SetDisconnected and a hand-set flag would be reverted by the live monitor before 04:15; and a corrupted manifest provably never reaches the store because the capture rewrites it first.
This commit is contained in:
@@ -0,0 +1,12 @@
|
||||
=== MY ERROR: the 03:15 guard was [ $H -ge 0315 ] with H=2334 -> true immediately.
|
||||
The hollow primary was injected at 23:34, not 03:15. Consequence recorded, not hidden:
|
||||
the hollow window is ~4h instead of ~15min, so the 04:15 off-site run WILL push a hollow
|
||||
privatebin snapshot. That makes Phase 5 richer (the whole R-412 chain plays out across the
|
||||
real cycle unattended) but it was NOT the plan.
|
||||
|
||||
=== has the 5-min capture rebuilt privatebin HOLLOW? ===
|
||||
primary : 4 files, 4234 bytes
|
||||
secondary: 6 files, 2123007 bytes
|
||||
secondary tar sha: c3ea1bae0731bcc3082d6c94
|
||||
primary manifest: db_dumps=[] volume_dumps=None
|
||||
HOLLOW: True
|
||||
@@ -0,0 +1,12 @@
|
||||
=== Phase 5 machinery check at 23:40 ===
|
||||
demo-hp recorder bytes : 4123669
|
||||
driver log lines : 33
|
||||
light load flag : on
|
||||
--- last driver lines ---
|
||||
(it cannot conjure a volume tar - that leg runs on the backup schedule). ===
|
||||
primary removed
|
||||
23:34:14 load: restore kimai -> [http=302 wall=0.012033s]
|
||||
|
||||
=== observer (demo-felhom) - MUST be untouched since 22:41 CEST ===
|
||||
recorder bytes: 94456
|
||||
last line: 2026-08-31T21:40:28.365022808Z 2026/08/31 21:40:28 [INFO] [stacks] Status refresh: 4 conta
|
||||
+41
@@ -0,0 +1,41 @@
|
||||
# Phase 5 — two of the four injections were NOT performed. The reasons, established not assumed.
|
||||
|
||||
## 1. "before offbox-backup 04:15: mark one app's drive disconnected" — NOT PERFORMED
|
||||
|
||||
There is no clean route on this box.
|
||||
|
||||
- `SetDisconnected` (`settings.go:1639`) is called by the **storage monitor** when a drive actually
|
||||
goes away. There is **no endpoint** that reaches it — `grep` over `server.go` and `router.go` for a
|
||||
disconnect route returns nothing.
|
||||
- Hand-editing `storage_paths[].disconnected` in `settings.json` is a **hand-set state** of exactly
|
||||
the shape the F9 lesson forbids, and the live monitor would revert it within a cycle anyway, so the
|
||||
injection would very likely not survive to 04:15 — an injection that silently un-injects is worse
|
||||
than none, because the run would then be read as a passing test of a fault that was not present.
|
||||
- `SetDecommissioned` is reached only through `finalizeDecommission`, a multi-step **migrate-then-
|
||||
decommission** workflow that genuinely moves data. That is a different and destructive operation,
|
||||
not a simulation of a missing drive.
|
||||
- Physically unmounting `hdd_1` was considered and rejected: three apps hold live data there, and a
|
||||
busy mount that refuses to unmount leaves a wedged mount — the exact hazard `CLAUDE.md` warns
|
||||
cannot be recovered without a reboot.
|
||||
|
||||
**Instead**, 04:15 is observed against the fault **already** injected: `privatebin`'s hollow primary.
|
||||
That is a real fault, it exercises R-412's chain through the real nightly run, and it needs no
|
||||
hand-set state.
|
||||
|
||||
## 2. "before proof 05:30: corrupt another app's manifest" — NOT PERFORMED, and Phase 4 is why
|
||||
|
||||
Phase 4 established by measurement that **a corrupted primary manifest never reaches the store**: the
|
||||
off-site run's own pre-push capture rewrites it. Verified by restoring snapshot `5dcee6e3` and reading
|
||||
its manifest — valid, not the corruption planted minutes earlier.
|
||||
|
||||
The proof judges the **snapshot**, not the live primary. So there is no route by which corrupting a
|
||||
manifest now can change what the 05:30 proof sees. Doing it anyway would produce a green result that
|
||||
proves nothing about the guard — the "instrument that cannot fail" shape.
|
||||
|
||||
The behaviour is covered by `TestR87_UnparseableManifestFails` (fail-closed).
|
||||
|
||||
## What IS injected for 05:30
|
||||
|
||||
**Stopping an app** — real, one command, and it does reach the proof's world: a stopped app still
|
||||
gets captured (measured tonight), so the question is whether the rotation and the verdicts stay
|
||||
correct with one app down. That will be done just before 05:30.
|
||||
+17
@@ -0,0 +1,17 @@
|
||||
=== THE DECISIVE MEASUREMENT: did the hollow primary overwrite the good secondary? ===
|
||||
primary : 5 files, 2122928 bytes
|
||||
primary manifest: db_dumps=[] volume_dumps=['privatebin_privatebin_data.tar'] HOLLOW=False
|
||||
secondary: 6 files, 2122929 bytes
|
||||
(was 6 files, 2123007 bytes before injection)
|
||||
secondary tar sha: c3ea1bae0731bcc3082d6c94
|
||||
(was c3ea1bae0731bcc3082d6c94)
|
||||
|
||||
=== did the R-403 guard say anything? ===
|
||||
RunHealthProbes: skipping privatebin — last check 3m50s ago, effective interval 5m0s, healthy=true
|
||||
RunHealthProbes: collected 0 targets (8 skipped not due, 1 skipped no container)
|
||||
RunHealthProbes: skipping privatebin — last check 4m0s ago, effective interval 5m0s, healthy=true
|
||||
RunHealthProbes: collected 0 targets (8 skipped not due, 1 skipped no container)
|
||||
RunHealthProbes: skipping privatebin — last check 4m10s ago, effective interval 5m0s, healthy=true
|
||||
RunHealthProbes: collected 0 targets (8 skipped not due, 1 skipped no container)
|
||||
RunHealthProbes: skipping privatebin — last check 4m20s ago, effective interval 5m0s, healthy=true
|
||||
RunHealthProbes: collected 0 targets (8 skipped not due, 1 skipped no container)
|
||||
+527
@@ -0,0 +1,527 @@
|
||||
=== every job that RAN since the mark (23:41), in order ===
|
||||
2026-08-31T21:41:14 agent-channel-health
|
||||
2026-08-31T21:42:14 stack-scan
|
||||
2026-08-31T21:42:14 agent-channel-health
|
||||
2026-08-31T21:43:14 agent-channel-health
|
||||
2026-08-31T21:44:14 stack-scan
|
||||
2026-08-31T21:44:14 agent-channel-health
|
||||
2026-08-31T21:45:14 offsite-credential-retry
|
||||
2026-08-31T21:45:14 backup-cache
|
||||
2026-08-31T21:45:14 system-health
|
||||
2026-08-31T21:45:14 agent-channel-health
|
||||
2026-08-31T21:46:14 stack-scan
|
||||
2026-08-31T21:46:14 agent-channel-health
|
||||
2026-08-31T21:47:14 agent-channel-health
|
||||
2026-08-31T21:48:14 stack-scan
|
||||
2026-08-31T21:48:14 agent-channel-health
|
||||
2026-08-31T21:49:14 agent-channel-health
|
||||
2026-08-31T21:50:14 backup-cache
|
||||
2026-08-31T21:50:14 system-health
|
||||
2026-08-31T21:50:14 stack-scan
|
||||
2026-08-31T21:50:14 offsite-credential-retry
|
||||
2026-08-31T21:50:14 agent-channel-health
|
||||
2026-08-31T21:51:14 agent-channel-health
|
||||
2026-08-31T21:52:14 stack-scan
|
||||
2026-08-31T21:52:14 agent-channel-health
|
||||
2026-08-31T21:53:14 agent-channel-health
|
||||
2026-08-31T21:54:14 stack-scan
|
||||
2026-08-31T21:54:14 agent-channel-health
|
||||
2026-08-31T21:55:14 backup-cache
|
||||
2026-08-31T21:55:14 offsite-credential-retry
|
||||
2026-08-31T21:55:14 hub-report
|
||||
2026-08-31T21:55:14 system-health
|
||||
2026-08-31T21:55:14 agent-channel-health
|
||||
2026-08-31T21:56:14 stack-scan
|
||||
2026-08-31T21:56:14 agent-channel-health
|
||||
2026-08-31T21:57:14 agent-channel-health
|
||||
2026-08-31T21:58:14 stack-scan
|
||||
2026-08-31T21:58:14 agent-channel-health
|
||||
2026-08-31T21:59:14 agent-channel-health
|
||||
2026-08-31T22:00:14 stack-scan
|
||||
2026-08-31T22:00:14 backup-cache
|
||||
2026-08-31T22:00:14 offsite-credential-retry
|
||||
2026-08-31T22:00:14 system-health
|
||||
2026-08-31T22:00:14 agent-channel-health
|
||||
2026-08-31T22:01:14 agent-channel-health
|
||||
2026-08-31T22:02:14 stack-scan
|
||||
2026-08-31T22:02:14 agent-channel-health
|
||||
2026-08-31T22:03:14 agent-channel-health
|
||||
2026-08-31T22:04:14 stack-scan
|
||||
2026-08-31T22:04:14 agent-channel-health
|
||||
2026-08-31T22:05:14 offsite-credential-retry
|
||||
2026-08-31T22:05:14 backup-cache
|
||||
2026-08-31T22:05:14 system-health
|
||||
2026-08-31T22:05:14 agent-channel-health
|
||||
2026-08-31T22:06:14 stack-scan
|
||||
2026-08-31T22:06:14 agent-channel-health
|
||||
2026-08-31T22:07:14 agent-channel-health
|
||||
2026-08-31T22:08:14 stack-scan
|
||||
2026-08-31T22:08:14 agent-channel-health
|
||||
2026-08-31T22:09:14 agent-channel-health
|
||||
2026-08-31T22:10:14 stack-scan
|
||||
2026-08-31T22:10:14 hub-report
|
||||
2026-08-31T22:10:14 system-health
|
||||
2026-08-31T22:10:14 backup-cache
|
||||
2026-08-31T22:10:14 offsite-credential-retry
|
||||
2026-08-31T22:10:14 agent-channel-health
|
||||
2026-08-31T22:10:14 disk-health-check
|
||||
2026-08-31T22:11:14 agent-channel-health
|
||||
2026-08-31T22:12:14 stack-scan
|
||||
2026-08-31T22:12:14 agent-channel-health
|
||||
2026-08-31T22:13:14 agent-channel-health
|
||||
2026-08-31T22:14:14 stack-scan
|
||||
2026-08-31T22:14:14 agent-channel-health
|
||||
2026-08-31T22:15:14 offsite-credential-retry
|
||||
2026-08-31T22:15:14 system-health
|
||||
2026-08-31T22:15:14 backup-cache
|
||||
2026-08-31T22:15:14 agent-channel-health
|
||||
2026-08-31T22:16:14 stack-scan
|
||||
2026-08-31T22:16:14 agent-channel-health
|
||||
2026-08-31T22:17:14 agent-channel-health
|
||||
2026-08-31T22:18:14 stack-scan
|
||||
2026-08-31T22:18:14 agent-channel-health
|
||||
2026-08-31T22:19:14 agent-channel-health
|
||||
2026-08-31T22:20:14 system-health
|
||||
2026-08-31T22:20:14 backup-cache
|
||||
2026-08-31T22:20:14 offsite-credential-retry
|
||||
2026-08-31T22:20:14 stack-scan
|
||||
2026-08-31T22:20:14 agent-channel-health
|
||||
2026-08-31T22:21:14 agent-channel-health
|
||||
2026-08-31T22:22:14 stack-scan
|
||||
2026-08-31T22:22:14 agent-channel-health
|
||||
2026-08-31T22:23:14 agent-channel-health
|
||||
2026-08-31T22:24:14 stack-scan
|
||||
2026-08-31T22:24:14 agent-channel-health
|
||||
2026-08-31T22:25:14 system-health
|
||||
2026-08-31T22:25:14 offsite-credential-retry
|
||||
2026-08-31T22:25:14 hub-report
|
||||
2026-08-31T22:25:14 backup-cache
|
||||
2026-08-31T22:25:14 agent-channel-health
|
||||
2026-08-31T22:26:14 stack-scan
|
||||
2026-08-31T22:26:14 agent-channel-health
|
||||
2026-08-31T22:27:14 agent-channel-health
|
||||
2026-08-31T22:28:14 stack-scan
|
||||
2026-08-31T22:28:14 agent-channel-health
|
||||
2026-08-31T22:29:14 agent-channel-health
|
||||
2026-08-31T22:30:14 system-health
|
||||
2026-08-31T22:30:14 offsite-credential-retry
|
||||
2026-08-31T22:30:14 backup-cache
|
||||
2026-08-31T22:30:14 stack-scan
|
||||
2026-08-31T22:30:14 agent-channel-health
|
||||
2026-08-31T22:31:14 agent-channel-health
|
||||
2026-08-31T22:32:14 stack-scan
|
||||
2026-08-31T22:32:14 agent-channel-health
|
||||
2026-08-31T22:33:14 agent-channel-health
|
||||
2026-08-31T22:34:14 stack-scan
|
||||
2026-08-31T22:34:14 agent-channel-health
|
||||
2026-08-31T22:35:14 system-health
|
||||
2026-08-31T22:35:14 backup-cache
|
||||
2026-08-31T22:35:14 offsite-credential-retry
|
||||
2026-08-31T22:35:14 agent-channel-health
|
||||
2026-08-31T22:36:14 stack-scan
|
||||
2026-08-31T22:36:14 agent-channel-health
|
||||
2026-08-31T22:37:14 agent-channel-health
|
||||
2026-08-31T22:38:14 stack-scan
|
||||
2026-08-31T22:38:14 agent-channel-health
|
||||
2026-08-31T22:39:14 agent-channel-health
|
||||
2026-08-31T22:40:14 offsite-credential-retry
|
||||
2026-08-31T22:40:14 hub-report
|
||||
2026-08-31T22:40:14 backup-cache
|
||||
2026-08-31T22:40:14 system-health
|
||||
2026-08-31T22:40:14 stack-scan
|
||||
2026-08-31T22:40:14 agent-channel-health
|
||||
2026-08-31T22:41:14 agent-channel-health
|
||||
2026-08-31T22:42:14 stack-scan
|
||||
2026-08-31T22:42:14 agent-channel-health
|
||||
2026-08-31T22:43:14 agent-channel-health
|
||||
2026-08-31T22:44:14 stack-scan
|
||||
2026-08-31T22:44:14 agent-channel-health
|
||||
2026-08-31T22:45:14 offsite-credential-retry
|
||||
2026-08-31T22:45:14 backup-cache
|
||||
2026-08-31T22:45:14 system-health
|
||||
2026-08-31T22:45:14 agent-channel-health
|
||||
2026-08-31T22:46:14 stack-scan
|
||||
2026-08-31T22:46:14 agent-channel-health
|
||||
2026-08-31T22:47:14 agent-channel-health
|
||||
2026-08-31T22:48:14 stack-scan
|
||||
2026-08-31T22:48:14 agent-channel-health
|
||||
2026-08-31T22:49:14 agent-channel-health
|
||||
2026-08-31T22:50:14 offsite-credential-retry
|
||||
2026-08-31T22:50:14 backup-cache
|
||||
2026-08-31T22:50:14 stack-scan
|
||||
2026-08-31T22:50:14 system-health
|
||||
2026-08-31T22:50:14 agent-channel-health
|
||||
2026-08-31T22:51:14 agent-channel-health
|
||||
2026-08-31T22:52:14 stack-scan
|
||||
2026-08-31T22:52:14 agent-channel-health
|
||||
2026-08-31T22:53:14 agent-channel-health
|
||||
2026-08-31T22:54:14 stack-scan
|
||||
2026-08-31T22:54:14 agent-channel-health
|
||||
2026-08-31T22:55:14 hub-report
|
||||
2026-08-31T22:55:14 offsite-credential-retry
|
||||
2026-08-31T22:55:14 backup-cache
|
||||
2026-08-31T22:55:14 system-health
|
||||
2026-08-31T22:55:14 agent-channel-health
|
||||
2026-08-31T22:56:14 stack-scan
|
||||
2026-08-31T22:56:14 agent-channel-health
|
||||
2026-08-31T22:57:14 agent-channel-health
|
||||
2026-08-31T22:58:14 stack-scan
|
||||
2026-08-31T22:58:14 agent-channel-health
|
||||
2026-08-31T22:59:14 agent-channel-health
|
||||
2026-08-31T23:00:14 system-health
|
||||
2026-08-31T23:00:14 offsite-credential-retry
|
||||
2026-08-31T23:00:14 stack-scan
|
||||
2026-08-31T23:00:14 backup-cache
|
||||
2026-08-31T23:00:14 agent-channel-health
|
||||
2026-08-31T23:01:14 agent-channel-health
|
||||
2026-08-31T23:02:14 stack-scan
|
||||
2026-08-31T23:02:14 agent-channel-health
|
||||
2026-08-31T23:03:14 agent-channel-health
|
||||
2026-08-31T23:04:14 stack-scan
|
||||
2026-08-31T23:04:14 agent-channel-health
|
||||
2026-08-31T23:05:14 backup-cache
|
||||
2026-08-31T23:05:14 system-health
|
||||
2026-08-31T23:05:14 offsite-credential-retry
|
||||
2026-08-31T23:05:14 agent-channel-health
|
||||
2026-08-31T23:06:14 stack-scan
|
||||
2026-08-31T23:06:14 agent-channel-health
|
||||
2026-08-31T23:07:14 agent-channel-health
|
||||
2026-08-31T23:08:14 stack-scan
|
||||
2026-08-31T23:08:14 agent-channel-health
|
||||
2026-08-31T23:09:14 agent-channel-health
|
||||
2026-08-31T23:10:14 hub-report
|
||||
2026-08-31T23:10:14 system-health
|
||||
2026-08-31T23:10:14 offsite-credential-retry
|
||||
2026-08-31T23:10:14 backup-cache
|
||||
2026-08-31T23:10:14 stack-scan
|
||||
2026-08-31T23:10:14 disk-health-check
|
||||
2026-08-31T23:10:14 agent-channel-health
|
||||
2026-08-31T23:11:14 agent-channel-health
|
||||
2026-08-31T23:12:14 stack-scan
|
||||
2026-08-31T23:12:14 agent-channel-health
|
||||
2026-08-31T23:13:14 agent-channel-health
|
||||
2026-08-31T23:14:14 stack-scan
|
||||
2026-08-31T23:14:14 agent-channel-health
|
||||
2026-08-31T23:15:14 offsite-credential-retry
|
||||
2026-08-31T23:15:14 system-health
|
||||
2026-08-31T23:15:14 backup-cache
|
||||
2026-08-31T23:15:14 agent-channel-health
|
||||
2026-08-31T23:16:14 stack-scan
|
||||
2026-08-31T23:16:14 agent-channel-health
|
||||
2026-08-31T23:17:14 agent-channel-health
|
||||
2026-08-31T23:18:14 stack-scan
|
||||
2026-08-31T23:18:14 agent-channel-health
|
||||
2026-08-31T23:19:14 agent-channel-health
|
||||
2026-08-31T23:20:14 stack-scan
|
||||
2026-08-31T23:20:14 backup-cache
|
||||
2026-08-31T23:20:14 system-health
|
||||
2026-08-31T23:20:14 offsite-credential-retry
|
||||
2026-08-31T23:20:14 agent-channel-health
|
||||
2026-08-31T23:21:14 agent-channel-health
|
||||
2026-08-31T23:22:14 stack-scan
|
||||
2026-08-31T23:22:14 agent-channel-health
|
||||
2026-08-31T23:23:14 agent-channel-health
|
||||
2026-08-31T23:24:14 stack-scan
|
||||
2026-08-31T23:24:14 agent-channel-health
|
||||
2026-08-31T23:25:14 backup-cache
|
||||
2026-08-31T23:25:14 hub-report
|
||||
2026-08-31T23:25:14 system-health
|
||||
2026-08-31T23:25:14 offsite-credential-retry
|
||||
2026-08-31T23:25:14 agent-channel-health
|
||||
2026-08-31T23:26:14 stack-scan
|
||||
2026-08-31T23:26:14 agent-channel-health
|
||||
2026-08-31T23:27:14 agent-channel-health
|
||||
2026-08-31T23:28:14 stack-scan
|
||||
2026-08-31T23:28:14 agent-channel-health
|
||||
2026-08-31T23:29:14 agent-channel-health
|
||||
2026-08-31T23:30:14 system-health
|
||||
2026-08-31T23:30:14 stack-scan
|
||||
2026-08-31T23:30:14 offsite-credential-retry
|
||||
2026-08-31T23:30:14 backup-cache
|
||||
2026-08-31T23:30:14 agent-channel-health
|
||||
2026-08-31T23:31:14 agent-channel-health
|
||||
2026-08-31T23:32:14 stack-scan
|
||||
2026-08-31T23:32:14 agent-channel-health
|
||||
2026-08-31T23:33:14 agent-channel-health
|
||||
2026-08-31T23:34:14 stack-scan
|
||||
2026-08-31T23:34:14 agent-channel-health
|
||||
2026-08-31T23:35:14 offsite-credential-retry
|
||||
2026-08-31T23:35:14 system-health
|
||||
2026-08-31T23:35:14 backup-cache
|
||||
2026-08-31T23:35:14 agent-channel-health
|
||||
2026-08-31T23:36:14 stack-scan
|
||||
2026-08-31T23:36:14 agent-channel-health
|
||||
2026-08-31T23:37:14 agent-channel-health
|
||||
2026-08-31T23:38:14 stack-scan
|
||||
2026-08-31T23:38:14 agent-channel-health
|
||||
2026-08-31T23:39:14 agent-channel-health
|
||||
2026-08-31T23:40:14 stack-scan
|
||||
2026-08-31T23:40:14 offsite-credential-retry
|
||||
2026-08-31T23:40:14 hub-report
|
||||
2026-08-31T23:40:14 system-health
|
||||
2026-08-31T23:40:14 backup-cache
|
||||
2026-08-31T23:40:14 agent-channel-health
|
||||
2026-08-31T23:41:14 agent-channel-health
|
||||
2026-08-31T23:42:14 stack-scan
|
||||
2026-08-31T23:42:14 agent-channel-health
|
||||
2026-08-31T23:43:14 agent-channel-health
|
||||
2026-08-31T23:44:14 stack-scan
|
||||
2026-08-31T23:44:14 agent-channel-health
|
||||
2026-08-31T23:45:14 system-health
|
||||
2026-08-31T23:45:14 backup-cache
|
||||
2026-08-31T23:45:14 offsite-credential-retry
|
||||
2026-08-31T23:45:14 agent-channel-health
|
||||
2026-08-31T23:46:14 stack-scan
|
||||
2026-08-31T23:46:14 agent-channel-health
|
||||
2026-08-31T23:47:14 agent-channel-health
|
||||
2026-08-31T23:48:14 stack-scan
|
||||
2026-08-31T23:48:14 agent-channel-health
|
||||
2026-08-31T23:49:14 agent-channel-health
|
||||
2026-08-31T23:50:14 offsite-credential-retry
|
||||
2026-08-31T23:50:14 stack-scan
|
||||
2026-08-31T23:50:14 system-health
|
||||
2026-08-31T23:50:14 backup-cache
|
||||
2026-08-31T23:50:14 agent-channel-health
|
||||
2026-08-31T23:51:14 agent-channel-health
|
||||
2026-08-31T23:52:14 stack-scan
|
||||
2026-08-31T23:52:14 agent-channel-health
|
||||
2026-08-31T23:53:14 agent-channel-health
|
||||
2026-08-31T23:54:14 stack-scan
|
||||
2026-08-31T23:54:14 agent-channel-health
|
||||
2026-08-31T23:55:14 backup-cache
|
||||
2026-08-31T23:55:14 offsite-credential-retry
|
||||
2026-08-31T23:55:14 hub-report
|
||||
2026-08-31T23:55:14 system-health
|
||||
2026-08-31T23:55:14 agent-channel-health
|
||||
2026-08-31T23:56:14 stack-scan
|
||||
2026-08-31T23:56:14 agent-channel-health
|
||||
2026-08-31T23:57:14 agent-channel-health
|
||||
2026-08-31T23:58:14 stack-scan
|
||||
2026-08-31T23:58:14 agent-channel-health
|
||||
2026-08-31T23:59:14 agent-channel-health
|
||||
2026-09-01T00:00:14 backup-cache
|
||||
2026-09-01T00:00:14 stack-scan
|
||||
2026-09-01T00:00:14 system-health
|
||||
2026-09-01T00:00:14 offsite-credential-retry
|
||||
2026-09-01T00:00:14 agent-channel-health
|
||||
2026-09-01T00:01:14 agent-channel-health
|
||||
2026-09-01T00:02:14 stack-scan
|
||||
2026-09-01T00:02:14 agent-channel-health
|
||||
2026-09-01T00:03:14 agent-channel-health
|
||||
2026-09-01T00:04:14 stack-scan
|
||||
2026-09-01T00:04:14 agent-channel-health
|
||||
2026-09-01T00:05:14 backup-cache
|
||||
2026-09-01T00:05:14 offsite-credential-retry
|
||||
2026-09-01T00:05:14 system-health
|
||||
2026-09-01T00:05:14 agent-channel-health
|
||||
2026-09-01T00:06:14 stack-scan
|
||||
2026-09-01T00:06:14 agent-channel-health
|
||||
2026-09-01T00:07:14 agent-channel-health
|
||||
2026-09-01T00:08:14 stack-scan
|
||||
2026-09-01T00:08:14 agent-channel-health
|
||||
2026-09-01T00:09:14 agent-channel-health
|
||||
2026-09-01T00:10:14 backup-cache
|
||||
2026-09-01T00:10:14 offsite-credential-retry
|
||||
2026-09-01T00:10:14 hub-report
|
||||
2026-09-01T00:10:14 stack-scan
|
||||
2026-09-01T00:10:14 system-health
|
||||
2026-09-01T00:10:14 disk-health-check
|
||||
2026-09-01T00:10:14 agent-channel-health
|
||||
2026-09-01T00:11:14 agent-channel-health
|
||||
2026-09-01T00:12:14 stack-scan
|
||||
2026-09-01T00:12:14 agent-channel-health
|
||||
2026-09-01T00:13:14 agent-channel-health
|
||||
2026-09-01T00:14:14 stack-scan
|
||||
2026-09-01T00:14:14 agent-channel-health
|
||||
2026-09-01T00:15:14 backup-cache
|
||||
2026-09-01T00:15:14 offsite-credential-retry
|
||||
2026-09-01T00:15:14 system-health
|
||||
2026-09-01T00:15:14 agent-channel-health
|
||||
2026-09-01T00:16:14 stack-scan
|
||||
2026-09-01T00:16:14 agent-channel-health
|
||||
2026-09-01T00:17:14 agent-channel-health
|
||||
2026-09-01T00:18:14 stack-scan
|
||||
2026-09-01T00:18:14 agent-channel-health
|
||||
2026-09-01T00:19:14 agent-channel-health
|
||||
2026-09-01T00:20:14 system-health
|
||||
2026-09-01T00:20:14 stack-scan
|
||||
2026-09-01T00:20:14 offsite-credential-retry
|
||||
2026-09-01T00:20:14 backup-cache
|
||||
2026-09-01T00:20:14 agent-channel-health
|
||||
2026-09-01T00:21:14 agent-channel-health
|
||||
2026-09-01T00:22:14 stack-scan
|
||||
2026-09-01T00:22:14 agent-channel-health
|
||||
2026-09-01T00:23:14 agent-channel-health
|
||||
2026-09-01T00:24:14 stack-scan
|
||||
2026-09-01T00:24:14 agent-channel-health
|
||||
2026-09-01T00:25:14 hub-report
|
||||
2026-09-01T00:25:14 offsite-credential-retry
|
||||
2026-09-01T00:25:14 system-health
|
||||
2026-09-01T00:25:14 backup-cache
|
||||
2026-09-01T00:25:14 agent-channel-health
|
||||
2026-09-01T00:26:14 stack-scan
|
||||
2026-09-01T00:26:14 agent-channel-health
|
||||
2026-09-01T00:27:14 agent-channel-health
|
||||
2026-09-01T00:28:14 stack-scan
|
||||
2026-09-01T00:28:14 agent-channel-health
|
||||
2026-09-01T00:29:14 agent-channel-health
|
||||
2026-09-01T00:30:00 db-dump
|
||||
2026-09-01T00:30:14 offsite-credential-retry
|
||||
2026-09-01T00:30:14 system-health
|
||||
2026-09-01T00:30:14 backup-cache
|
||||
2026-09-01T00:30:14 stack-scan
|
||||
2026-09-01T00:30:14 agent-channel-health
|
||||
2026-09-01T00:31:14 agent-channel-health
|
||||
2026-09-01T00:32:14 stack-scan
|
||||
2026-09-01T00:32:14 agent-channel-health
|
||||
2026-09-01T00:33:14 agent-channel-health
|
||||
2026-09-01T00:34:14 stack-scan
|
||||
2026-09-01T00:34:14 agent-channel-health
|
||||
2026-09-01T00:35:14 system-health
|
||||
2026-09-01T00:35:14 offsite-credential-retry
|
||||
2026-09-01T00:35:14 backup-cache
|
||||
2026-09-01T00:35:14 agent-channel-health
|
||||
2026-09-01T00:36:14 stack-scan
|
||||
2026-09-01T00:36:14 agent-channel-health
|
||||
2026-09-01T00:37:14 agent-channel-health
|
||||
2026-09-01T00:38:14 stack-scan
|
||||
2026-09-01T00:38:14 agent-channel-health
|
||||
2026-09-01T00:39:14 agent-channel-health
|
||||
2026-09-01T00:40:14 stack-scan
|
||||
2026-09-01T00:40:14 system-health
|
||||
2026-09-01T00:40:14 offsite-credential-retry
|
||||
2026-09-01T00:40:14 backup-cache
|
||||
2026-09-01T00:40:14 hub-report
|
||||
2026-09-01T00:40:14 agent-channel-health
|
||||
2026-09-01T00:41:14 agent-channel-health
|
||||
2026-09-01T00:42:14 stack-scan
|
||||
2026-09-01T00:42:14 agent-channel-health
|
||||
2026-09-01T00:43:14 agent-channel-health
|
||||
2026-09-01T00:44:14 stack-scan
|
||||
2026-09-01T00:44:14 agent-channel-health
|
||||
2026-09-01T00:45:14 offsite-credential-retry
|
||||
2026-09-01T00:45:14 system-health
|
||||
2026-09-01T00:45:14 backup-cache
|
||||
2026-09-01T00:45:14 agent-channel-health
|
||||
2026-09-01T00:46:14 stack-scan
|
||||
2026-09-01T00:46:14 agent-channel-health
|
||||
2026-09-01T00:47:14 agent-channel-health
|
||||
2026-09-01T00:48:14 stack-scan
|
||||
2026-09-01T00:48:14 agent-channel-health
|
||||
2026-09-01T00:49:14 agent-channel-health
|
||||
2026-09-01T00:50:14 backup-cache
|
||||
2026-09-01T00:50:14 system-health
|
||||
2026-09-01T00:50:14 stack-scan
|
||||
2026-09-01T00:50:14 offsite-credential-retry
|
||||
2026-09-01T00:50:14 agent-channel-health
|
||||
2026-09-01T00:51:14 agent-channel-health
|
||||
2026-09-01T00:52:14 stack-scan
|
||||
2026-09-01T00:52:14 agent-channel-health
|
||||
2026-09-01T00:53:14 agent-channel-health
|
||||
2026-09-01T00:54:14 stack-scan
|
||||
2026-09-01T00:54:14 agent-channel-health
|
||||
2026-09-01T00:55:14 backup-cache
|
||||
2026-09-01T00:55:14 system-health
|
||||
2026-09-01T00:55:14 hub-report
|
||||
2026-09-01T00:55:14 offsite-credential-retry
|
||||
2026-09-01T00:55:14 agent-channel-health
|
||||
2026-09-01T00:56:14 stack-scan
|
||||
2026-09-01T00:56:14 agent-channel-health
|
||||
2026-09-01T00:57:14 agent-channel-health
|
||||
2026-09-01T00:58:14 stack-scan
|
||||
2026-09-01T00:58:14 agent-channel-health
|
||||
2026-09-01T00:59:14 agent-channel-health
|
||||
2026-09-01T01:00:14 backup-cache
|
||||
2026-09-01T01:00:14 offsite-credential-retry
|
||||
2026-09-01T01:00:14 system-health
|
||||
2026-09-01T01:00:14 stack-scan
|
||||
2026-09-01T01:00:14 agent-channel-health
|
||||
2026-09-01T01:01:14 agent-channel-health
|
||||
2026-09-01T01:02:14 stack-scan
|
||||
2026-09-01T01:02:14 agent-channel-health
|
||||
2026-09-01T01:03:14 agent-channel-health
|
||||
2026-09-01T01:04:14 stack-scan
|
||||
2026-09-01T01:04:14 agent-channel-health
|
||||
2026-09-01T01:05:14 backup-cache
|
||||
2026-09-01T01:05:14 offsite-credential-retry
|
||||
2026-09-01T01:05:14 system-health
|
||||
2026-09-01T01:05:14 agent-channel-health
|
||||
2026-09-01T01:06:14 stack-scan
|
||||
2026-09-01T01:06:14 agent-channel-health
|
||||
2026-09-01T01:07:14 agent-channel-health
|
||||
2026-09-01T01:08:14 stack-scan
|
||||
2026-09-01T01:08:14 agent-channel-health
|
||||
2026-09-01T01:09:14 agent-channel-health
|
||||
2026-09-01T01:10:14 hub-report
|
||||
2026-09-01T01:10:14 backup-cache
|
||||
2026-09-01T01:10:14 system-health
|
||||
2026-09-01T01:10:14 offsite-credential-retry
|
||||
2026-09-01T01:10:14 geo-verify
|
||||
2026-09-01T01:10:14 stack-scan
|
||||
2026-09-01T01:10:14 selfupdate-check
|
||||
2026-09-01T01:10:14 disk-health-check
|
||||
2026-09-01T01:10:14 agent-channel-health
|
||||
2026-09-01T01:11:14 agent-channel-health
|
||||
2026-09-01T01:12:14 stack-scan
|
||||
2026-09-01T01:12:14 agent-channel-health
|
||||
2026-09-01T01:13:14 agent-channel-health
|
||||
2026-09-01T01:14:14 stack-scan
|
||||
2026-09-01T01:14:14 agent-channel-health
|
||||
2026-09-01T01:15:14 backup-cache
|
||||
2026-09-01T01:15:14 system-health
|
||||
2026-09-01T01:15:14 offsite-credential-retry
|
||||
2026-09-01T01:15:14 agent-channel-health
|
||||
2026-09-01T01:16:14 stack-scan
|
||||
2026-09-01T01:16:14 agent-channel-health
|
||||
2026-09-01T01:17:14 agent-channel-health
|
||||
2026-09-01T01:18:14 stack-scan
|
||||
2026-09-01T01:18:14 agent-channel-health
|
||||
2026-09-01T01:19:14 agent-channel-health
|
||||
2026-09-01T01:20:14 backup-cache
|
||||
2026-09-01T01:20:14 stack-scan
|
||||
2026-09-01T01:20:14 offsite-credential-retry
|
||||
2026-09-01T01:20:14 system-health
|
||||
2026-09-01T01:20:14 agent-channel-health
|
||||
2026-09-01T01:21:14 agent-channel-health
|
||||
2026-09-01T01:22:14 stack-scan
|
||||
2026-09-01T01:22:14 agent-channel-health
|
||||
2026-09-01T01:23:14 agent-channel-health
|
||||
2026-09-01T01:24:14 stack-scan
|
||||
2026-09-01T01:24:14 agent-channel-health
|
||||
2026-09-01T01:25:14 system-health
|
||||
2026-09-01T01:25:14 backup-cache
|
||||
2026-09-01T01:25:14 offsite-credential-retry
|
||||
2026-09-01T01:25:14 hub-report
|
||||
2026-09-01T01:25:14 agent-channel-health
|
||||
2026-09-01T01:26:14 stack-scan
|
||||
2026-09-01T01:26:14 agent-channel-health
|
||||
2026-09-01T01:27:14 agent-channel-health
|
||||
2026-09-01T01:28:14 stack-scan
|
||||
2026-09-01T01:28:14 agent-channel-health
|
||||
2026-09-01T01:29:14 agent-channel-health
|
||||
2026-09-01T01:30:00 fill-watch
|
||||
2026-09-01T01:30:00 tier2-backup
|
||||
2026-09-01T01:30:14 backup-cache
|
||||
2026-09-01T01:30:14 stack-scan
|
||||
2026-09-01T01:30:14 offsite-credential-retry
|
||||
2026-09-01T01:30:14 system-health
|
||||
2026-09-01T01:30:14 agent-channel-health
|
||||
2026-09-01T01:31:14 agent-channel-health
|
||||
2026-09-01T01:32:14 stack-scan
|
||||
2026-09-01T01:32:14 agent-channel-health
|
||||
|
||||
=== when did privatebin regain its volume tar? ===
|
||||
2026-08-31T21:42:14 ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml
|
||||
2026-08-31T21:44:14 ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml
|
||||
2026-08-31T21:45:14 groupStacksByDrive: /mnt/sys_drive → [bentopdf, bookstack, docmost, kimai, opengist, privatebin]
|
||||
2026-08-31T21:45:14.638876337Z 33651ab96aed privatebin privatebin privatebin/pdo:2.0.5
|
||||
2026-08-31T21:45:14 DiscoverDatabases: skipping container privatebin (image=privatebin/pdo:2.0.5, not a database)
|
||||
2026-08-31T21:45:14 ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/privatebin/docker-compose.yml
|
||||
2026-08-31T21:46:14 ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml
|
||||
2026-08-31T21:48:14 ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml
|
||||
2026-08-31T21:50:14 ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml
|
||||
2026-08-31T21:50:14 groupStacksByDrive: /mnt/sys_drive → [bentopdf, bookstack, docmost, kimai, opengist, privatebin]
|
||||
2026-08-31T21:50:14.680542869Z 33651ab96aed privatebin privatebin privatebin/pdo:2.0.5
|
||||
2026-08-31T21:50:14 DiscoverDatabases: skipping container privatebin (image=privatebin/pdo:2.0.5, not a database)
|
||||
2026-08-31T21:50:14 ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/privatebin/docker-compose.yml
|
||||
2026-08-31T21:52:14 ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml
|
||||
+16
@@ -0,0 +1,16 @@
|
||||
BEFORE primary=5 files, 905207 B secondary=23 files, 5808704 B
|
||||
sec tar sha: d7e7f422a34c02455525
|
||||
--- inject: remove the primary unit; the 5-min capture rebuilds it hollow ---
|
||||
hollow after ~165s
|
||||
AFTER-INJECT primary=4 files, 6527 B
|
||||
--- fire the Tier-2 mirror. The guard must REFUSE the unit leg. ---
|
||||
AFTER-MIRROR secondary=23 files, 5808704 B
|
||||
sec tar sha: d7e7f422a34c02455525
|
||||
RESULT: the good copy SURVIVED
|
||||
--- what the run said about calibre-web ---
|
||||
groupStacksByDrive: /mnt/felhom-drives/hdd_1 → [calibre-web, paperless-ngx, romm]
|
||||
Recovery unit captured for calibre-web → /mnt/felhom-drives/hdd_1/backups/primary/calibre-web (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0)
|
||||
Tier 2 calibre-web: unit leg SKIPPED — the recovery unit on the source drive lists no database dumps and no volume tars, while the existing copy at /mnt/sys_drive/felhom-data/backups/secondary/calibre-web/recovery-unit does. The copy was PRESERVED rather than replaced with an empty one (R-403). The other legs continue.
|
||||
[unit leg SKIPPED — existing package preserved, R-403]
|
||||
debug cross-drive run for calibre-web completed
|
||||
Event pushed: crossdrive_completed (info) — Másodlagos mentés elkészült: calibre-web
|
||||
@@ -0,0 +1,44 @@
|
||||
# Phase 5 row 1 — the R-403 mirror guard. VERDICT: **PASS** (forced, after the natural test was lost)
|
||||
|
||||
## The natural test was lost, and the reason is itself the finding
|
||||
|
||||
`privatebin` was injected hollow at 23:34 to meet the 03:30 `tier2-backup`. **It healed at 02:30**,
|
||||
an hour before the mirror ran, because the **`db-dump` job re-creates the volume tars**. By 03:30 the
|
||||
primary was complete (`volume_dumps: ['privatebin_privatebin_data.tar']`, 2 122 928 B) and the mirror
|
||||
correctly copied a sound unit. **The guard was never exercised.**
|
||||
|
||||
That fixes the timing in R-412: the hollow window closes at the next **02:30 db-dump**, not at the
|
||||
next off-site backup — so a unit that loses its tar just after 02:30 stays hollow for nearly 24 h,
|
||||
with the 04:15 off-site push inside that window.
|
||||
|
||||
## So it was forced, on an HDD app, through the real Tier-2 path
|
||||
|
||||
`calibre-web` (an HDD app, so `crossDriveTargets` includes it) injected hollow — primary
|
||||
905 207 B → 6 527 B — then the real Tier-2 mirror fired.
|
||||
|
||||
```
|
||||
BEFORE secondary = 23 files, 5 808 704 B, tar sha d7e7f422a34c02455525
|
||||
AFTER secondary = 23 files, 5 808 704 B, tar sha d7e7f422a34c02455525
|
||||
```
|
||||
|
||||
**The guard fired and named itself:**
|
||||
|
||||
```
|
||||
Tier 2 calibre-web: unit leg SKIPPED — the recovery unit on the source drive lists no database
|
||||
dumps and no volume tars, while the existing copy at /mnt/sys_drive/felhom-data/backups/secondary/
|
||||
calibre-web/recovery-unit does. The copy was PRESERVED rather than replaced with an empty one
|
||||
(R-403). The other legs continue.
|
||||
[unit leg SKIPPED — existing package preserved, R-403]
|
||||
```
|
||||
|
||||
The other legs continued and the run completed. **The good copy survived byte-identical.**
|
||||
|
||||
## One thing the runbook asked that is only half true
|
||||
|
||||
The runbook's acceptance was *"the run does not report a plain success"*. **The log does not** — it
|
||||
states the skip and the reason. **The customer EVENT does**: `crossdrive_completed (info) — Másodlagos
|
||||
mentés elkészült: calibre-web`, with no mention of the skipped leg. It is `info` severity, which
|
||||
`severityNotifies` drops, so nobody is mailed a false success — but an operator surface listing events
|
||||
would show a plain completion over a run that deliberately skipped a leg. Recorded as an observation,
|
||||
not filed: the event is not delivered, and widening it is the kind of coarse-vs-per-app decision
|
||||
08 §6.2 fences.
|
||||
@@ -0,0 +1,22 @@
|
||||
mark2 = 62484 at 03:36
|
||||
offbox-backup ran at 04:19
|
||||
calibre-web/docker-compose.yml: hash match, skipped
|
||||
calibre-web/.felhom.yml: hash match, skipped
|
||||
logscanner: scanned calibre-web: errors=0 warnings=0 issues=0 (took 22ms)
|
||||
off-box restore bookstack completed (full=false, async)
|
||||
calibre-web/docker-compose.yml: hash match, skipped
|
||||
calibre-web/.felhom.yml: hash match, skipped
|
||||
logscanner: scanned calibre-web: errors=0 warnings=0 issues=0 (took 21ms)
|
||||
calibre-web/docker-compose.yml: hash match, skipped
|
||||
calibre-web/.felhom.yml: hash match, skipped
|
||||
logscanner: scanned calibre-web: errors=0 warnings=0 issues=0 (took 24ms)
|
||||
off-box restore docmost completed (full=false, async)
|
||||
backup run started (9 app(s) toggled)
|
||||
Stopping calibre-web for safe volume dump
|
||||
StopStack calibre-web: current state=running deployed=true containers=1
|
||||
Stopping stack: calibre-web
|
||||
Running: docker compose down (in /opt/docker/stacks/calibre-web)
|
||||
Recovery unit captured for calibre-web → /mnt/felhom-drives/hdd_1/backups/primary/calibre-web (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0)
|
||||
Stack calibre-web stopped successfully (took 4.2s)
|
||||
Dumping volume calibre-web_calibre_web_config for calibre-web
|
||||
Volume dump: calibre-web/calibre-web_calibre_web_config → 877.5 KB
|
||||
+28
@@ -0,0 +1,28 @@
|
||||
=== calibre-web's primary AFTER the 04:15 run ===
|
||||
5 files, 905207 B
|
||||
db_dumps=[] volume_dumps=['calibre-web_calibre_web_config.tar'] HOLLOW=False
|
||||
|
||||
=== what reached the STORE - restore the newest snapshot and look inside ===
|
||||
newest calibre-web snapshot: 6fee3b5a
|
||||
/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/manifest.json
|
||||
/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/app.yaml
|
||||
/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/.felhom.yml
|
||||
/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/docker-compose.yml
|
||||
/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/volume-dumps/calibre-web_calibre_web_config.tar
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/nested/őszibarack.md
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/SENTINEL.txt
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/POST-SNAPSHOT.txt
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/binary-1mb.bin
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/árvíztűrő-tükörfúrógép.txt
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/plain.txt
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/binary-3mb.bin
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/nested/őszibarack.md
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/árvíztűrő-tükörfúrógép.txt
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/plain.txt
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/metadata.db
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/metadata.db-wal
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-SENTINEL.txt
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/SENTINEL.txt
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/book.bin
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/Örkény-egyperces-r356.txt
|
||||
/mnt/felhom-drives/hdd_1/userdata/media/books/metadata.db-shm
|
||||
@@ -0,0 +1,21 @@
|
||||
# Phase 6 interim note (taken at 03:37, mid-cycle — the full verdict comes after 06:00)
|
||||
|
||||
The analyser was dry-run at 03:37 so it would not be exercised for the first time at 06:05.
|
||||
It works, and it already shows two things worth carrying into the final verdict.
|
||||
|
||||
## 1. `tier2-backup` on the observer completed in **3 ms**
|
||||
|
||||
```
|
||||
2026-09-01 01:30:00 Running job: tier2-backup
|
||||
2026-09-01 01:30:00 Job tier2-backup completed (took 3ms)
|
||||
```
|
||||
|
||||
On `demo-hp` the same job took **~6 s** and copied 9 apps. 3 ms is a no-op. The likely reason is that
|
||||
`demo-felhom` has no second drive to mirror to — **to be established at Phase 6, not assumed.** If
|
||||
that is right it is correct behaviour on a single-drive box; if it is not, it is a finding.
|
||||
|
||||
## 2. Nothing has alarmed on the untouched box
|
||||
|
||||
`db-dump` completed in 674 ms and pushed `db_dump_completed (info)` — the only event so far. **Zero
|
||||
ERROR or WARN lines** in the whole capture since 23:00. That is the single most valuable line Phase 6
|
||||
can produce, and it is on track.
|
||||
@@ -595,7 +595,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-409** | **Nothing in the product can vouch for the bytes of a restored recovery unit — the only hash record covers 0.002 % of it.** MEASURED on demo-hp 2026-08-31 against kimai's restored unit: `manifest.json`'s `checksums` object carries sha256 for `.felhom.yml` (2 235 B), `app.yaml` (488 B) and `docker-compose.yml` (2 195 B) — **4 918 bytes of a 213 231 242-byte unit**. The database dump (48 217 B) and the two named-volume tars (160 331 776 B + 52 845 056 B) — the recoverable data, 99.998 % of the bytes — have no recorded hash anywhere. **And nothing else supplies one:** restic 0.14.0's `restore --verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving one-byte corruption of a 160 MB tar passed clean, red-proofed), and `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps and **no content hash**. **So "the restore produced correct files" is currently unanswerable by any automated means.** **What is NOT claimed here:** `restic check --read-data-subset=100%` already proves the STORE's packs, and the config files that ARE hashed are the ones a wrong-content failure would be hardest to spot in. | **OPEN — MEDIUM** | R-87, R-361 | Cheapest fix, and it is already half-built: extend the capture's `checksums` to cover `db_dumps` and `volume_dumps` — R-361 already computes a canonical dump sha256 to prove itself, so the value exists at capture time. Then a restore-test has a real reference and R-87's narrow version becomes a content check rather than a completeness one. Evidence: `audits/SPIKE-restic-restore-test-2026-08-31.md` §Q2, §Q3. | CC |
|
||||
| **R-410** | **`golden_currency_gate.py` is satisfied by a DIRECTORY NAME — `mkdir documentation/tests/golden-<VER>-<DATE>` turns it green with no bake behind it.** `EVIDENCE_RE = ^golden-(\d+)\.(\d+)\.(\d+)-\d{4}-\d{2}-\d{2}$` matched against `os.listdir` (`scripts/golden_currency_gate.py:89,123`) — it never opens the directory, never looks for a log, never checks a sha, and never asks Gitea whether the package exists. Noticed 2026-08-31 while the 0.230.0 bake turned it from red to green: **I created the directory before the bake finished, and the gate would have passed at that moment.** **What the gate DOES say honestly:** its own success line already reads *"this checks the BAKE, not the vouch"* — so the vouch hole is declared. **This one is not:** nothing tells a reader the bake check is a filename check. **Class:** an instrument that cannot distinguish the thing from a label for the thing — the same shape as R-378's whole-field match and R-233's un-matchable grep, aimed this time at the release gate. **Exposure today is low** because the runbook produces a real evidence directory and two bakes in a row have; it is the NEXT hurried session that pays. | **OPEN — LOW** | R-242 | Make it read something the bake alone can produce: the `GOLDEN_SHA256=` line in the directory's `bake.log`, or a HEAD against the Gitea package URL for that version. Prefer the log — it keeps the gate offline and `--fast`. Ship a red-proof: an empty `golden-9.9.9-2026-01-01/` directory must FAIL. | CC |
|
||||
| **R-411** | **A background job DELETES the lock of a live customer restore, and logs it as a crash that did not happen.** MEASURED on demo-hp 2026-08-31 during the overnight soak, through the product's own endpoints — this is R-408's consequence, which until tonight had only been reasoned about. **The chain, every step observed:** (1) a customer full-restore runs `OffboxRestorePrepareFull` → `restic stats`, and **`restic stats` TAKES A REPOSITORY LOCK** (clean-room test: nothing else running, 4x stats, sampler reads `locks=1`); (2) the restore holds `opRunning` but **NOT `acquireRunning`** (R-408), so the integrity check is not blocked and runs concurrently; (3) the check meets that lock, and `resticStep` escalates to **`unlock --remove-all`** — caught by the argv sampler at **20:50:51 with `restore 3c11059b --target …` and `unlock --remove-all` in the SAME sample**; (4) the log says *"cleared a stale exclusive lock left by a previous crash (single-writer repo)"* — **there was no crash**, and `resticStep` cannot know there was, because it fires on ANY `repository is already locked`. **THE CUSTOMER-FACING CONSEQUENCE WAS CONTAINED, and that is R-359's guard working:** the check returned `ok:false` in 6.7 s and was classified **Unreachable, NOT damage** — *"the check could not run to a verdict (other) — NOT reported as damage"* — so no `backup_integrity_failed` and no customer mail. Due-ness was not advanced either, so it retries. **What is NOT contained:** a live operation's lock is deleted by a background job; the single-writer premise `resticStep`'s own comment rests on is false in this pairing; and that night's integrity check silently did not verify the store, with only a WARN. **The opposite direction is FENCED and was measured too:** five restores fired into a running check at 5/15/25/35/40 s offsets were ALL refused by `restoreOpBlocked` (`offbox_handlers.go:358`), zero restic invoked — so the hazard is reachable only restore-FIRST. | **OPEN — MEDIUM** | R-408, R-407, R-359 | Decide ONE way, and R-408 is the same decision: either `RestoreOffboxScratch` (and the full-restore preparation) takes `acquireRunning`, or `resticStep`'s escalation stops claiming a crash it cannot verify and refuses instead of removing. **Pin whichever is chosen with a test that reproduces this pairing** — a unit test on `resticStep` alone cannot see it. Evidence: `audits/DRILL-soak-2026-08-31/phase1-lock-collision/`. | CC |
|
||||
| **R-412** | **A lost recovery unit is rebuilt WITHOUT its volume dumps, and the off-site backup then ships that hollow unit and reports success.** OBSERVED on demo-hp 2026-08-31 during the soak, produced by the product with no construction — **this is the natural instance of R-403's shape that yesterday's R-87 session could not produce and had to hand-build.** **The chain, every step in the log:** (1) `opengist`'s primary unit was removed (a restore, a crash mid-capture or a remount does the same — R-403's own named causes); (2) the 5-minute capture rebuilt it — *"Recovery unit captured for opengist"* — with compose and manifest but **no volume tar**, manifest reading `db_dumps: []`, `volume_dumps: None` (the key absent entirely), **185 664 B → 4 382 B**; (3) the off-site backup pushed it and logged *"backed up opengist … 0 mandatory path(s)"* — a SUCCESS line over a backup containing none of the app's data; (4) the newest off-site snapshot `35ba9fe7` is now hollow. **THE CAPTURE IS NOT WRONG IN ISOLATION** — a capture describing an empty tree as empty is correct, and R-403 deliberately refused to guard it for that reason. **What is wrong is that nothing between the capture and the off-site push notices that a unit which HAD dumps yesterday has none today**, and the run reports success. The volume-dump leg runs on the backup schedule, not on capture, so the window is a whole cycle wide. | **OPEN — HIGH** | R-403, R-87, R-413 | Decide where the notice belongs: the capture (which R-403 fenced off), the off-site run's own pre-push phase (which already re-captures and could compare against what it is about to replace), or a shrink-detector on the primary. **Do NOT simply guard the capture** — 08 §8.2 records why that would make the manifest lie. Evidence: `audits/DRILL-soak-2026-08-31/phase2-guard-interactions/`. | CC |
|
||||
| **R-412** | **A recovery unit lost DURING an off-site run — after its own dump leg, before its push — is shipped hollow and the run reports success.** **CORRECTED 2026-09-01 04:22, and the first wording of this row OVERSTATED it.** As first filed it claimed the hollow unit sat in the store for a whole cycle because "the volume-dump leg runs on the backup schedule, not on capture". **That is wrong, and measuring it overnight is what showed it:** the off-site run has its OWN pre-push dump leg — *"Stopping calibre-web for safe volume dump"*, *"Volume dump: calibre-web/calibre-web_calibre_web_config -> 877.5 KB"* — so a unit that is hollow when a run starts is **REPAIRED before it is pushed**. Proven twice: `opengist` (2026-08-31 21:0x) and `calibre-web` (2026-09-01 04:15) both went in hollow and came out complete, and the snapshot pulled back from the store (`6fee3b5a`) holds the volume tar and all 17 userdata files. **WHAT REMAINS REAL, and it is narrower:** the one hollow snapshot that DID reach the store (`35ba9fe7`, opengist) was created when the unit was destroyed **inside** a run that had already completed opengist's dump leg — so the push shipped what the capture had just rebuilt empty, and logged *"backed up opengist (… 0 mandatory path(s))"*, **a success line over a backup holding none of the app's data**. That race is real, it was observed, and the success wording is wrong either way. **The R-403 mirror guard holds throughout** — proven live: *"unit leg SKIPPED … The copy was PRESERVED rather than replaced with an empty one"*, secondary byte-identical. | **OPEN — LOW (was HIGH; the correction is the reason)** | R-403, R-87, R-413 | Two separable things. (1) The success line: a per-app push that carried no dumps and no tars should not read as a plain success — that is a wording fix in the run's own reporting, not a new guard. (2) The race: decide whether the push should re-read the unit it is about to send, or whether the window is small enough to accept. **Do NOT guard the capture** (08 §8.2). Evidence: `audits/DRILL-soak-2026-08-31/phase2-guard-interactions/` and `phase5-mutated-cycle/09-what-reached-the-store.txt`. | CC |
|
||||
| **R-413** | **R-87's proof caught a naturally-produced hollow snapshot, end to end, unattended — the validation yesterday's session could only do with a declared hand-built fixture.** 2026-08-31 soak, demo-hp. After R-412's chain left `opengist`'s newest off-site snapshot hollow, the nightly proof rotated to it and returned **`verdict:"fail"`, `reason:"volumes_expected_none_captured"`, missing `opengist_data`**, logged *"READABLE AND EMPTY — the store is not damaged; the backup does not contain this app's data"*, and pushed **one** `offsite_proof_empty` at severity `error`. The four apps ahead of it in the rotation all passed, so the discrimination is real and not a constant fail. **This is recorded as a row rather than only as a report line because it upgrades a claim:** the capability map's R-87 row cites a CONSTRUCTED failing case; it can now cite a natural one. | **CLOSED 2026-08-31 — the claim it upgrades is recorded** | R-87, R-412 | Nothing to build. When the capability map is next touched, cite this instead of the constructed case. | CC |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
|
||||
Reference in New Issue
Block a user