Files
felhom.eu/documentation/audits/CAMPAIGN-10-two-storage-soak-2026-07-31.md
T
admin 5f35aa0346 Campaign 10: M-band RTO measured — RTO ~= 40s + 26.9s/GB, and a capacity ceiling that matters more
The S figures (66 MB -> 42.0s, two passes agreeing to 0.6s) had a spread tight
enough to prove fixed work dominates, which is exactly why they said nothing about
M. Second point taken 327x larger, same app, same method: clock from restore
request to the app serving the correct discriminator.

rallly's postgres volume grown 66 MB -> 21.1 GB (200k rows, STORAGE EXTERNAL so
TOAST cannot compress it into a fake number). Two reps:

  rep 1  backup 406.4s  unit 41149 MB  RTO 624.5s  discriminator correct
  rep 2  backup 387.2s  unit 41133 MB  RTO 591.8s  discriminator correct

327x the data cost 14.5x the time - strongly sub-linear:
  RTO ~= 40s + 26.9 s/GB      backup ~= 29s + 17.4 s/GB
  10 GB -> 5.2 min   20 GB -> 9.6 min (measured 10.1)   100 GB -> 46 min
The fixed ~40s dominates below ~1.5 GB, which IS the S band and explains its tight
clustering.

The more consequential result is capacity. A DB-backed app's recovery unit is
1.90x its data (volume tar PLUS SQL dump): 21.1 GB produced a 40.2 GB unit. The
default appliance ships /mnt/sys_drive at 20 GB, so the largest app that can hold
a local Tier-1/2 recovery unit on a default box is about 10 GB - and that fills the
volume. The M band does not fit on a default box at all; this test only reached
21 GB because sys_drive was first grown 20G -> 70G with the same operation the
product performs via SysDataGrowGB. A tier-sizing decision, not a defect, but it
is invisible until an app crosses it.

Caveats stated in the doc: two points define a line but do not test linearity; the
1.90x is DB-app-specific and a file-only app should be nearer 1.0x (inferred, not
measured); synthetic incompressible data; one app, one box.
2026-08-02 08:31:51 +02:00

550 lines
33 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CAMPAIGN 10 — two-storage adversarial soak (2026-07-31 → 08-01)
**Phase A COMPLETE and gated (2026-07-31). Phase B run in three passes (2026-08-01/02): run 1
27 cycles, run 2a 10 cycles (stopped — harness defect), run 2b **39 cycles**. Together **66 clean
cycles + 39 with the full atom set**, 1 461 invariant checks.**
**Findings: R-156 (papra), R-157 (bootrecon, two mechanisms).** One structural scope constraint —
Tier 3 cannot be isolated (§3). **Three suspicions investigated and DISPROVED** (§5.2, §5.3), and
**two violations traced to the harness, not the product** (§4d). Monotonic-growth sampling over
13.5 h found **no leak** (§6b).
Evidence: `../tests/campaign10-evidence-2026-07-31/``state/phaseA-journal.md` (step-by-step),
`isolation/pbs-denial.txt`, `isolation/restic-denial.txt`, `r156-papra-volume.txt`.
---
## 1. Venue and baselines
| | |
|---|---|
| Host | `demo-hp` (HP t740), **Tier 0**, the designated drill/build host |
| VM | **311 `c10-appliance`** — q35/OVMF (`pre-enrolled-keys=0`), 4 cores, 8 GB, `cpu=host` |
| Disks | `scsi0` 200 G system · `scsi1` 50 G · `scsi2` 50 G, qcow2 on `c10-scratch` |
| Storage | **`c10-scratch`**, `dir` at **`/mnt/nvme-1tb` — the mount ROOT**, `content=images` |
| Box | `c10.felhom.eu` @ `192.168.0.105` (demo-hp's vmbr0, `192.168.0.87/24`) |
| Console | no browser on DooPlex → `qm monitor screendump` → PNG, read visually; `sendkey` for input |
| Untouched | `drill-r50` (VM 300), guest 9201 on both demo boxes, `local-lvm`, both backup targets, ep0, the Storage Box |
**The `exactMount` choice, made deliberately (the brief asks which).** `c10-scratch` sits at the mount
**root**, not a subdirectory. A dir storage at a subdirectory sets `exactMount=false`
`reachable=false``StorageStateDisconnected` (`felhom-agent/internal/storage/observe.go:336`), which
would have emitted `storage_disconnected` for demo-hp for the entire run — **the exact signal I1/I2
exist to discriminate**. It coexists with `felhom-backup` on the same path, which was not modified.
### Baselines — every value read fresh
| What | Value | How |
|---|---|---|
| controller `main` | **v0.188.0** | CHANGELOG top; `HEAD`=`origin/main`=`4115e88`, clean |
| **golden's baked controller** | **0.188.0** (`7353d8be…`) | hub `/configuration`, selected option |
| agent published + vouched | **0.119.0** (`e37aca82…`) | hub `/configuration`, selected option |
| hub | **0.86.0** | live deploy image tag == CHANGELOG top |
| ISO | **1.26.1**, `f3cc86d5…`, 1 705 322 496 B | live round trip from `iso.felhom.eu` |
| `felhom.eu` HEAD | `e9a74a0`, clean | `git rev-parse` |
| managed floor | v0.156.0 (DB override) | hub `/configuration` |
**Golden vs `main`: not behind — both 0.188.0.** The deliberate decision the brief asks for is
therefore vacuous this run; the campaign tests what is shipping *and* what is current.
**Two of the brief's assumptions were stale; reading fresh caught both.**
1. The brief was written against ISO **v1.25.0**. `iso.felhom.eu` began serving **1.26.1** at 18:37 Z
the same day, ~90 minutes before this run. Operator chose 1.26.1 — the current published artifact.
2. The brief and three docs say demo-hp has **no baked SSH key**; `ssh -o BatchMode=yes demo-hp`
authenticated **by key, first try**. Already **R-129**; **not re-filed**.
---
## 2. The isolation gate — §2
| Target | Isolation | Proof |
|---|---|---|
| **PBS** | dedicated datastore `felhom-c10` on DooPlex + auth-id `c10@pbs!box` scoped to it alone | **CAPTURED**`isolation/pbs-denial.txt` |
| **restic** | scratch subaccount `u629488-sub4`, own chroot home `felhom-campaign10` | **CAPTURED**`isolation/restic-denial.txt` |
| **hub** | the campaign's own record `c10-soak`; no other customer touched | **CAPTURED** — §2.3 |
**Both denials carry a POSITIVE CONTROL, and the first attempt failed it.** The initial PBS capture
returned four tidy 403s — including on the campaign's *own* datastore. All four were worthless: PBS
API tokens carry privilege separation, so a token's rights are the intersection of the user's ACL and
the token's, and user `c10@pbs` had none. Granting the scoped role to the user as well turned the
control green and left the denials standing. Without the control this would have shipped as a clean
isolation proof that proved only that the credential was broken.
## POSITIVE CONTROL — same token, its OWN datastore
GET /api2/json/admin/datastore/felhom-c10/snapshots {"data":[]} HTTP 200
## DENIAL 1 — production datastore felhom-spike
permission check failed - missing Datastore.Audit|Datastore.Backup on /datastore/felhom-spike 403
## DENIAL 3 — DELETE namespace demo-felhom (destructive)
permission check failed - missing Datastore.Modify on /datastore/felhom-spike 403
restic, same shape — the credential writes and reads back in its own home, and every production path
is outside the chroot:
POSITIVE CONTROL put/get/rm c10-canary.txt in own home → succeeds, content verified
DENIAL cd /home/felhom-peti-felhom → stat remote: No such file or directory
DENIAL ls .. → remote readdir("/home/../"): Permission denied
**2.3 — hub.** The run created exactly one record, `c10-soak`, and bound exactly one appliance
(uuid `55ab3776-…`) to it. Existing customers (`demo-felhom`, `demo-hp`, `peti-felhom`, `drill-r50`,
`sess-f`) were read but never written. Verified from the hub's own log: every `c10-soak` line is the
campaign's, and no line names another customer.
---
## 3. Scope constraint — Tier 3 cannot be isolated, so it was not run
The brief requires the campaign's PBS to be *a dedicated datastore and namespace on DooPlex*. That is
unreachable, for two reasons that are each an **already-recorded deliberate position**:
1. **Offsite hard-requires the DR tier**`hub/internal/web/configs.go:1300` refuses outright:
*"Offsite backup requires the DR tier — enable it first (the escrow ceremony depends on the PBS
key)"*. Closed by policy 2026-07-12 (drill F-6, DR-tier-by-default decisions 3/4).
2. **The DR tier only ever provisions on ep0** — peer allocation and endpoint sync use the lowest
`endpoint_id` only; per-endpoint allocation is an explicitly deferred arc (`hub/README.md:260`).
So the only two configurations are *DR tier on* → campaign PBS traffic lands on **ep0** (Tier 2, the
offsite of last resort, RAM-constrained per R-90, fenced by §3), or *DR tier off* → **no Tier 3 at
all**. Chosen: **DR tier OFF, offsite OFF** — the only option §3 permits.
Consequence, stated up front rather than discovered in the results: the campaign touches **neither ep0
nor the Storage Box at any point** — stronger isolation than asked for, obtained by not running the
tier. **Lost: every Tier-3 atom, I8 entirely, and the Tier-3 RTO/RPO rows.** Unaffected: the
two-drive core (R-112/113/114/116/117), D5, and I1I7, I9I11.
**No `R-n` minted** — the register was grepped first and both halves are already recorded. This is a
campaign constraint, not a defect.
---
## 4. Phase A gates
| Gate | Result | Positive observable |
|---|---|---|
| **A1** install from the published ISO | **PASS** | `ii felhom-bootstrap 1.26.1`; unit `enabled`, fired at first boot; registered unclaimed with pairing code `FD6-A6S`, MACs, hardware and 3 SSH host keys; **4.2 GB actually written**, system booted on its own |
| **A2** claim | **PASS** | discriminator flipped `dashboard not yet claimed``authentication required`; hub: `[claim] customer c10-soak CLAIMED its dashboard` |
| **A3** enrol both drives, apps across them | **PASS** | both registered via the real endpoint; backup target healthy; 4 apps healthy |
| **A4** discriminators | **PASS** | all four apps seed and read back the same generation |
| **A5** budget + watchdog | **PASS** (built for Phase B) | disk floor 120 G on `/mnt/nvme-1tb` + VM-liveness watchdog; heartbeat `status.txt` per phase; `journal.tsv` per check. Neither fired: the run ended on its deadline with 871 G free |
**A1 — 1.26.1 cannot auto-install, verified against the bytes** (not inferred from its manifest):
`auto-installer-mode.toml` 0 hits, `answer.toml` 0, `proxmox-auto-install` 0, exactly 2 interactive
menuentries. Driven blind through the **Terminal UI** (the graphical entry is undrivable — monitor
mouse events do not move the guest cursor). Three traps hit and handled:
- GRUB's default is the graphical entry; the first `down` landed after the 15 s countdown. Reset and
sent it inside the window in one round trip.
- **The guest keymap was Hungarian while `sendkey` emits US scancodes.** Every symbol afterwards —
notably `@`, which is AltGr+V on a Hungarian layout — would have been silently mangled, *including
the 24-char root password*, locking the box out with no way back in. Switched to U.S. English
before typing anything; confirmed positively when `c10@felhom.eu` rendered a correct `@` and `.`.
- The recorded boot-order trap: `--boot` set in its **own** `qm set` after the disks existed and
verified from `qm config`; flipped to `order=scsi0` post-install and verified again.
Only `scsi0` was attached during the install, so the installer offered exactly one target and could
not reach a data drive. The 50 G drives were attached afterwards.
**A2** — the installer's root password stops working at day-0; it is replaced by a hub-vaulted
managed credential (`vaulted break-glass recovery credential … secret 32 chars`), retrieved via
`POST /hosts/c10-soak-8a405e/reveal-recovery-credential`. R-119 confirmed live again: the claim code
is emailed-only and the operator relayed it — the single human step.
**A3** — both drives enrolled through **the real endpoint** (`POST /api/storage/attach`), which does
format → assign → guest-attach → register-stable-path in order:
| Drive | Device | Stable path | Role |
|---|---|---|---|
| `adatok` | `/dev/sdb` | `/mnt/felhom-drives/adatok` | app data, `set_default` |
| `mentes` | `/dev/sdc` | `/mnt/felhom-drives/mentes` | **backup target** (PVE storage `felhom-backup` at `/mnt/mentes`) |
A false start worth recording: hand-driving `format` + `assign` first left the drives **mounted but
unregistered** — invisible to both the raw-candidate scan and `/disks`, so the real enroll endpoint
could not resolve a UUID for them (`resolveEnrollUUID` sees neither list). Backed out (eject leaves
the raw `/mnt/<name>` mount by design, `disks.go:606`, so the mounts also had to be undone) and re-ran
the real flow clean. **The lesson is the brief's own:** use the real endpoint, or manufacture a state
the product never produces.
Backup-target offer flow (R-112/R-114/E-2) exercised as the customer would: assign returned
`restart_required: true`, and the state only flipped after the agent restart it asked for —
before: {"degraded":true, "target":"local", "offer_path":"/mnt/felhom-drives/adatok", "message":"A rendszermentés jelenleg ugyanazon a lemezen van…"}
after: {"degraded":false,"known":true,"label":"Mentes","target":"felhom-backup"}
That healthy read is the **I5/I6 baseline**: everything healthy ⇒ no degraded banner, target reads
healthy, zero absent events.
**Apps — four, healthy, covering both sides of D5's split:**
| App | Data | Secret classes | Why chosen |
|---|---|---|---|
| **rallly** | postgres (`rallly-postgres`) | 2 × `type: secret` (`SECRET_PASSWORD` hex:32 labelled *Titkosítási kulcs*, `DB_PASSWORD`) | the DB app; travelling class |
| **homebox** | sqlite volume | 1 × `secret`, **`data_key: true`** | data-encrypting key |
| **papra** | see **R-156** | 1 × `secret`, **`data_key: true`** | data-encrypting key |
| **grafana** | volume | 1 × **`type: password`** | **the withheld class — I10's negative half** |
4 travelling `secret` fields and 1 `password` field that must never travel, so I10 can assert both
directions rather than only the positive one.
**A4 — discriminators work.** Every app carries a generation marker; a full seed/read cycle returns
the seeded generation from all four. rallly's is a row in a `cc_proof` table read over **the path
`DATABASE_URL` actually names** (`docker run --network container:rallly … psql -h rallly-postgres`),
never `127.0.0.1` inside the postgres container — the trusted-socket path that produced D5's §1.2
false pass.
---
## 4b. Phase B — the soak
Three passes. Run 2a was stopped deliberately at cycle 10 when two violations turned out to be
**harness** defects (§4d); the harness was fixed and red-proofed, and run 2b restarted from cycle 1
with the full atom set.
| Pass | Cycles | Atoms | Violations |
|---|---|---|---|
| **run 1** | 27 | 6 families | **0** |
| **run 2a** | 10 | 13 families | 2 — both **harness**, §4d |
| **run 2b** | **39** | 13 families | 9 (see below) |
**Run 2b — 875 invariant rows over 39 cycles, 13.5 h, ended on its own deadline.** No watchdog halt,
no atom exception, no I11 breach.
| Check | PASS | VIOL | Note |
|---|---|---|---|
| **I7** | **38** | **0** | + 1 `I7-SKIP` — the fixed harness declining to judge, not a false green |
| I1 / I1-pair | 37 / 36 | 2 / 3 | one 3-cycle window, **self-recovered** — §5.3 |
| I2 / I2-pair | 39 / 39 | 0 / 0 | |
| I3 / I3-abort | 78 / 4 | 0 | |
| I4 / I4-abort / -recover | 78 / 4 / 4 | 0 | |
| I5 / I6 | 39 / 39 | 0 | |
| I10 | **195** | 0 | 5 secret-class fields × 39 |
| I11 | 39 | 0 | |
| I1-under-load / -recover | 5 / 5 | **0** | target pulled **during** a running backup |
| BACKUP / KILL-CTRL | 82 / 39 | 0 | |
| KILL-AGENT / REBOOT / REBOOT-VM | 5 / 5 / 5 | 0 | |
| CONCURRENCY / FILL-DRIVE | 5 / 5 | 0 | |
| HARD-RESET | 2 | **3** | **R-157** |
| REDEPLOY | 38 | 1 | harness: 5-min wait too short under load; the app returned just after |
**The full atom set ran this time.** The six families run 1 skipped are all present: filesystem
aborted in place (R-117 Q7), kill-agent-mid-backup, hard-reset-mid-write, reboot-VM, both concurrency
atoms, and fill-drive-near-full. The run-1 flaw where `reboot` was appended *after* the shuffle — so it
never interleaved with a detach — was fixed; heavy atoms are now permuted in.
**Two results worth naming.** `I1-under-load` passed 5/5: pulling the backup target **while a backup
was running** still produced `backup_target_absent` and a clean recovery. And `I4-abort` / `I3-abort`
passed 4/4 — R-117's Q7 case, the one its spike called *"the worse half"*, holds (details in
`../tests/campaign10-evidence-2026-07-31/r117-q7-abort-in-place.md`).
## 4d. Two violations that were the HARNESS, not the product
Recorded because **a check that fails for the wrong reason is as corrosive as one that passes for the
wrong reason**, and this arc already has six of the latter. Full write-up:
`../tests/campaign10-evidence-2026-07-31/run2a-violations-were-harness.md`.
Run 2a cycle 10 reported `HARD-RESET canaries_intact=False` and, worse, **`I7` claiming a stale
restore** (`want=C10-C010-A got=C10-C009-A, restore_ok=True`). The `cc_proof` table settled it: its
highest row was `C10-C009-A`**there was no `C010-A` row at all**, so the seed never landed, because
the hard reset earlier in the same cycle had left `rallly` `Exited`. `atom_restore_verify` called
`seed()` without checking its return; `atom_hard_reset_mid_write` read canaries out of app containers
that were still starting.
**A real stale restore would have produced a byte-identical journal row**, which is why this justified
stopping a running 10-cycle pass rather than annotating it. Fixed with `apps_ready()`, `seed(verify=True)`
(read-back), and an `I7-SKIP` verdict so a skipped check is never silently a green one — red-proofed
both directions before restarting. Run 2b's single `I7-SKIP` is that fix working.
## 4c. Primitives proven by hand before automation
| Step | Observed |
|---|---|
| backup | `POST /api/backup/run``Mentés elindítva`, settles in ~2334 s |
| restore | `POST /backup/restore` (form) is **async**; polled on `/api/backup/restore-status`, never `IsRunning()` |
| I7, first proof | seeded A → backup → seeded B → restored A; rallly returned **A** while the other three still read **B** |
| I1 | `backup_target_absent (error) — A rendszermentés meghajtója nem érhető el: Mentes` |
| I2 | `storage_disconnected (error) — Meghajtó váratlanul leválasztva: Adatok` |
| I6 degraded copy | *"A rendszermentés meghajtója nem érhető el — amíg vissza nem csatlakoztatod, a teljes rendszermentés nem készül el."* |
| I10 | rallly `DB_PASSWORD`+`SECRET_PASSWORD`, homebox `HBOX_AUTH_API_KEY_PEPPER`, papra `AUTH_SECRET` present; grafana `GF_SECURITY_ADMIN_PASSWORD` **absent**; unit mode `0600` |
## 5. Findings
### R-156 — papra's data is neither persisted nor backed up, and it reports healthy
**New. Register grepped first:** papra appears in R-41 (never-deployable / healthcheck) and R-127
(`data_key` misclassification); neither covers this. *(Also corrected en route: a first grep suggested
R-164 was taken — it is a mis-citation of controller **v0.164.0**, already flagged in
`REPORT-record-correction-2026-07-29.md:141`. Highest real number is R-155, so this is R-156.)*
The template mounts `papra_data:/app/data`, but the app persists to **`/app/app-data/db/db.sqlite`**.
The mounted volume is therefore empty, and the real database lives in the container's writable layer.
Evidence (`r156-papra-volume.txt`), all four legs measured:
1. volume papra_papra_data -> /app/data
2. uid=999(nonroot); drwxr-xr-x 2 root root /app/data
touch: cannot touch '/app/data/.w': Permission denied
3. /app/data contains only the canary this campaign wrote AS ROOT
4. /app/app-data/db/db.sqlite 475136 bytes owned nonroot <-- the real DB, in no volume
5. health status: "healthy"
**Consequences.** (a) papra's documents do not survive `docker compose down`/redeploy/host restore —
the writable layer is not persisted. (b) `DumpAppVolumes` dumps the volume, so papra is **backed up
as an empty directory**: its Tier-1/Tier-2 backup is real, verifiable, and contains nothing. (c) The
healthcheck only probes the HTTP port, so none of this surfaces — the app is green throughout.
This is the *presence is not success* family from `CLAUDE.md`: a backup exists, a healthcheck passes,
and neither means the data is there. It is also a second instance of the pattern R-41 names — an app
that deploys and looks fine while being fundamentally broken.
**Two-repo shape**, if fixed: the volume target belongs in `app-catalog-felhom.eu`, and the general
defence is a catalog gate asserting each template's mounted volume is a path the app actually writes.
**Not fixed here** — the fences forbid it, and a fix mid-run would prove a version that did not exist
when the run started.
### R-157 — bootrecon's start-ONCE sweep acts on an unsettled snapshot; the boot orphan survives
**New. Register grepped: R-52 is the feature, nothing covers its timing.** **INTERMITTENT — 3 of 6
hard resets (~50%)** across two runs, with an identical signature each time. (An earlier draft called
this deterministic because the failing cycle numbers matched; they matched only because the runner's
RNG is seeded. Intermittency is what a race predicts and a wrong predicate does not.) Full evidence:
`../tests/campaign10-evidence-2026-07-31/r157-bootrecon-start-once-race.md`.
A `qm reset` mid-backup brought everything back **except the app half of the DB-backed stack**:
`rallly` left `Exited (255)` (`oom=false`, `restarts=0`, its own log ending `✓ Ready` — it died
healthy) while `rallly-postgres` returned healthy.
20:28:13 Status refresh: 8 containers across 55 stacks <-- docker ps -a shows NINE
20:28:18 [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)
20:28:25 Status refresh: 7 containers … 8 containers <-- still churning AFTER the sweep
20:39:14 [deadapp] 20 scans since boot, 5 deployed app(s) evaluated, 1 currently down
The predicate is not at fault — once settled the controller reports `rallly state=degraded
containers=2`, and `IsDownState` includes `StateDegraded` (`manager.go:55`), so
`len(Containers)>0 && IsDownState(State)` holds. **The snapshot was**: bootrecon fires as
`go runBootReconcile(...)` (`cmd/controller/main.go:236`) ~5 s after start, while docker is still
restoring containers, and is start-once by design so it never re-checks.
**Consequence: the app stays down indefinitely.** Detection is perfect and recovery never happens —
exactly R-52's original shape, *an alarm with no recovery* (F5: two apps Exited for ~18 h). In run 2a
the app only returned because a later campaign atom redeployed it. **Not fixed** — the fences forbid
it, and the settle-condition fix belongs with a test that pins the *consequence*.
### 5.2 — Investigated and DISPROVED: `/api/disks` reports `state=attached` for an absent drive
Recorded because the *chase* is the useful artefact, not because anything is broken.
With `adatok` physically removed, `/api/disks` kept reporting `state: "attached"` for it across three
polls (~90 s) — while `/dev/sdb` was gone, `/mnt/adatok` was absent from `/proc/mounts`, and the guest
bind returned `Input/output error`. The other fields were right (`backing_device: ""`,
`guest_attached: false`, `bound_under_parent: false`).
It looked like the R-116/R-117 family, and there is a real consumer:
`internal/web/intermediary.go:230``present[d.MountPath] |= d.State == "attached"`. A presence gate
reading an absent drive as present is precisely I3's failure mode.
**It is inert, and the gate is correct.** `planDriveGates` only gates paths under
`StableParentDir+"/"` (`/mnt/felhom-drives/…`). The registered path is the **stable** one, whose
presence is computed from `d.BoundUnderParent` — observed **false** — not from `State`. The
`State`-derived entry is keyed on the **raw** `/mnt/adatok`, which is not a registered stable path, so
the loop skips it. The observable end-to-end behaviour agrees: the storage page showed
*"Meghajtó leválasztva: Adatok"* with a disconnect timestamp and a reconnect action, i.e. the gate fired.
The stale `state` field comes from the **registry** row surviving in the union after the
storage-observation row drops out (a raw enrolled drive is not a PVE storage), carrying its configured
`MountPath`. **No `R-n` minted** — no consumer is misled, and the R-116 comment's claim that this
"cannot make the gate read an absent drive as PRESENT" held under test.
### 5.3 — Investigated and DISPROVED: the `backup_target_*` pair going silent
Five I1/I1-pair violations in a ~10-minute window looked like R-116's unmatchable pair returning.
It is not. Full write-up: `../tests/campaign10-evidence-2026-07-31/RESOLVED-backup-target-pair-went-silent.md`.
The violations **cluster at cycles 31-33 and nowhere else** in 39 cycles; c34-c39 are clean, so it
recovered with no intervention. On the quiesced box afterwards, one slow detach with 4 minutes either
side produced a perfect pair (`backup_target_absent``backup_target_restored``health_recovered`).
And the alarming false-healthy I sampled alongside it (`mentes bound_under_parent=False` while
`degraded: false`) **does not survive quiescence** — three minutes after the run both drives read
`bound=True` and the target read healthy. I had been reading the two halves at different instants of
a detach. **No `R-n`.**
### 5.4 — Investigated and DISPROVED: the hub's `SQLITE_BUSY` event drops
[ERROR] Failed to save event from c10-soak: database is locked (5) (SQLITE_BUSY)
7 in 24 h, **including one for the real customer `demo-felhom`**, so not a campaign artefact. The hub
returns HTTP 500 on the failed save (`hub/internal/api/handler.go:1715`) and dispatches notifications
only *after* a successful save — so a genuinely lost event would be a lost alarm.
**But the controller retries 3×, and zero events exhausted their attempts** across the whole run
(`grep -c "Event push failed after 3 attempts"` = 0). The 07:04:39 drop landed 3 s later at 07:04:42,
same event. **Nothing was lost. No `R-n`.** One cosmetic note kept: `[ERROR] Failed to save event`
reads like data loss and is not.
### Not filed
- **demo-hp SSH key** — already R-129.
- **Tier 3 isolation** — §3; two recorded deliberate positions, not a defect.
- **`/api/disks` `state`** — §5.2, disproved.
- **`backup_target_*` silence** — §5.3, transient and self-recovered.
- **hub `SQLITE_BUSY` drops** — §5.4, absorbed by the controller's retry.
- **F-CRIT-1 (`StateStopped` assumed deliberate)** — confirmed live again via R-157 mechanism B, but
already recorded in `CLAUDE.md`'s false-invariant table; **not re-filed**.
---
## 6. RTO/RPO — measured
**66 restores across run 1 and run 2b. Every one returned the correct discriminator.**
| Tier | App | Data volume | n | min | median | p90 | max | mean |
|---|---|---|---|---|---|---|---|---|
| Tier 1 (local) | rallly (postgres) | 66 MB volume / 67 MB unit | 28 (run 1) | 38.8 s | **42.0 s** | 42.5 s | 44.3 s | 42.0 s |
| Tier 1 (local) | rallly (postgres) | 66 MB | **38 (run 2b)** | 29.1 s | **41.4 s** | 42.1 s | 49.6 s | 40.7 s |
The two independent passes agree to **0.6 s on the median**, and run 2b's wider spread (29.149.6 s
against 38.844.3 s) is the cost of measuring under hard resets, agent kills and a near-full drive.
### M band — measured 2026-08-02
The S figures could not speak for M, so a second point was taken **327× larger, same app, same
method**. Full evidence: `../tests/campaign10-evidence-2026-07-31/rto-m-band-2026-08-02.md`.
| rep | volume | backup | recovery unit | **RTO to correct data** | discriminator |
|---|---|---|---|---|---|
| 1 | 21 616 MB | 406.4 s | 41 149 MB | **624.5 s** | correct |
| 2 | 21 615 MB | 387.2 s | 41 133 MB | **591.8 s** | correct |
| **mean** | **21.1 GB** | **396.8 s** | **40.2 GB (1.90×)** | **608.1 s (10.1 min)** | **2/2** |
**327× the data cost 14.5× the time** — strongly sub-linear, so:
> **RTO ≈ 40 s + 26.9 s/GB** · **backup ≈ 29 s + 17.4 s/GB**
>
> 1 GB → 67 s · 10 GB → 5.2 min · **20 GB → 9.6 min** (measured 10.1) · 50 GB → 23 min · 100 GB → 46 min
**The fixed ~40 s dominates below ~1.5 GB** — that is the S band, and it is why the S numbers were so
tightly clustered. Above it, RTO is essentially linear in data.
### The capacity ceiling — the more consequential result
**A DB-backed app's recovery unit is 1.90× its data** (volume tar **plus** SQL dump): 21.1 GB produced
a **40.2 GB** unit. The default appliance ships `/mnt/sys_drive` at **20 GB** (`mp1 … size=20G`;
`SysDataGrowGB` is a per-customer grow, 0 for this customer). Therefore:
> **on a default box the largest app that can hold a local Tier-1/2 recovery unit is ≈ 10 GB, and that
> fills the volume completely — realistically ≈ 8 GB. The M band does not fit on a default box at
> all.** This test only reached 21 GB because `/mnt/sys_drive` was first grown to 70 G.
A sizing decision for the S/M/L tiers rather than a defect — but it is the constraint that actually
bites, and it is invisible until an app crosses it.
**Caveats, stated not buried:** two points define a line but do not test linearity (no ~5 GB point was
taken); the 1.90× is DB-app-specific and a file-only app should be nearer 1.0× — **inferred, not
measured**; the data is deliberately incompressible synthetic, so well-compressing real data moves
both numbers; one app, one shape, one box.
**RPO is still not measured.** Backups were driven on demand, not on the schedule.
---
## 6b. Monotonic growth — the half the invariants cannot see
I1I11 are *correctness* checks: they answer "is the system telling the truth this cycle", and every
one of them can pass while something climbs. **9 457 samples of 19 metrics over 13.5 h**, following
Campaign 2's `controller_rss.tsv` precedent (`state/growth.tsv`, analyser in `runner/`).
**Leak-class — nothing grew.** Across 39 cycles containing dozens of redeploys, controller and agent
kills, guest reboots, VM reboots and hard resets:
| Metric | start → end | verdict |
|---|---|---|
| controller RSS | 105.1 → 31.8 MB (min 0 / max 132 across restarts) | no leak |
| agent RSS | 19 008 → 19 632 KB (+2.07/cycle) | no leak |
| controller fd / agent fd | 14 → 11 / 14 → 13 | no leak |
| docker volumes / images / containers | 7 → 7 / 13 → 13 / 9 → 9 | **no orphans from repeated redeploys** |
| controller restarts | 0 | — |
**Expected accumulation — all bounded and explicable:**
| Metric | start → end | note |
|---|---|---|
| **agent journal** | **193.8 → 463.4 MB** (+6.5/cycle, ~20 MB/h) | **the only curve with real slope** |
| VM qcow2 on the host | 17 439 → 24 903 MB (+7.5 GB) | matches host free space falling 7.5 GB |
| backup target `mentes` | 623 → 6 424 MB | tracks `VZDUMP_N` 3 → 11 |
| guest root / docker logs | 5 667 → 5 940 MB / 4 → 9 MB | |
Two things to watch rather than findings: the **agent journal** at ~20 MB/h (journald caps it, so it
is bounded, but it is a lot of logging); and **whole-guest dumps accumulating 3 → 11** on a 50 GB
target, which at soak rates would fill it in days — a real box does not back up this often, so this is
an artefact of the campaign's cadence rather than a defect.
**Answering the question directly: after 13.5 h and 39 destructive cycles, nothing leaks.** That is a
genuine positive result and the half a targeted session structurally cannot produce.
---
## 7. I1I11 across cycles
Full tables in §4b. Aggregated over run 1 (27 cycles) and run 2b (39 cycles):
- **I7 — 66 restores, 66 correct discriminators, 0 stale, 0 empty.** The campaign's headline.
- **I2, I3, I4, I5, I6, I10, I11 — zero violations in either pass**, including the abort-in-place
variants and 195 + 135 secret-class assertions covering both sides of D5's split.
- **I1 — 2 violations in 39 cycles**, one self-recovering 3-cycle window under cycling far faster than
a real box sees (§5.3, disproved as a defect).
- **I8 — never checked.** Tier 3 could not be isolated (§3).
- **I9 — not automated.** Cited from the tester-gate run on this same controller 0.188.0, not
re-claimed as this campaign's work.
**A clean invariant is a result**: I2's genericity, I1's discrimination, both pairings, I4's
no-restart repair and I7's discriminator held across **66 cycles** of detach, reattach, abort,
backup, restore, kill, redeploy, reboot and hard reset.
---
## 8. What did not run, and why
**Atom coverage: 12 of the brief's ~13 families ran.** Run 1 covered 6; run 2b added the rest.
| Atom (brief B1) | Ran |
|---|---|
| back up Tier 1/2 · restore + verify by discriminator · detach/reattach target · detach/reattach non-target · delete and redeploy · reboot the guest | **YES** (run 1 + 2b) |
| **abort a filesystem in place** (R-117 Q7) · **kill the agent mid-backup** · **hard-reset the VM mid-write** · **reboot the VM** · **concurrent backup+restore** · **concurrent backup+detach** · **fill a drive near-full** | **YES** (run 2b) |
| back up / restore **Tier 3** | **NO** — §3, structurally impossible to isolate |
**Still honest about two limits.** `kill_controller` is not literally "mid-backup" — it lands at a
permuted point, so it exercises recovery rather than the mid-write race; the dedicated
`concurrent backup+detach` atom (`I1-under-load`, 5/5 PASS) covers the mid-operation case properly.
And **I9** — a Tier-1/2 restore with the guest's `app.yaml` moved aside — was **not automated**; it
was proven on this same controller 0.188.0 by the tester-gate run and is cited, not re-claimed.
**Depth reached: 39 consecutive cycles**, past the brief's "the state that quietly drifts on the
thirty-eighth". Cycles 3439 were clean on every invariant. Beyond 39 is untested, not proven clean.
---
## 9. Teardown — OWED, nothing removed yet
The rig is intact **on purpose**, so Phase B can run without repeating Phase A. Every item below is
outstanding and must be removed when the campaign ends:
| Layer | Item | Command |
|---|---|---|
| VM | `311 c10-appliance` on demo-hp | `qm stop 311 && qm destroy 311 --purge` |
| storage | `c10-scratch` dir storage | `pvesm remove c10-scratch` |
| PBS | datastore `felhom-c10`, user `c10@pbs`, token `!box`, both ACLs | `proxmox-backup-manager datastore remove felhom-c10` · `user delete c10@pbs` · remove `/mnt/5_hdd/backup/campaign10` |
| restic | subaccount **`u629488-sub4`** (`felhom-campaign10`) on box 611714 — **unused, since Tier 3 never ran** | `DELETE /v1/storage_boxes/611714/subaccounts/281530` |
| **hub** | **customer `c10-soak`****disposition: DELETE.** Named explicitly per R-131, which is four orphaned scratch customers left by exactly this omission | `POST /configs/c10-soak/delete` with `ack_hosts=1 ack_reset=1 ack_purge=1 confirm_id=c10-soak` |
| secrets | `~/.config/campaign10/` on DooPlex (host + dashboard passwords, API key, app secrets, PBS token, restic password); `/root/c10/` and `/root/c10api.sh` in guest 9201 | `shred -u` both |
`sess-f` ("R-120 golden 0.186.0 proof") is still present and is **not** this run's record — R-131
stands, untouched.
---
## 10. Hygiene note
The customer API key for `c10-soak` was printed into this session's transcript while being read from
the hub page. It belongs to a scratch customer scheduled for deletion (§9) and grants nothing once
that record is gone, but the write-out was avoidable and is recorded here rather than left implicit —
the same class as R-132, where a correctly-made request was undone by how it was reported.