docs: R-182 closed, R-90 closed on measurement, R-86 unblocked, ep0 record corrected
gates / gates (push) Successful in 7s
gates / gates (push) Successful in 7s
R-182 CLOSED (controller v0.194.0 + hub v0.90.0/.1), proven live on demo-hp. The hub's notification_log for the run reads: two per-app failures RECORDED, one digest SENT naming both, and the customer channel SKIPPED with operator_only. Against the measured previous behaviour — two failures, one email naming one app, one leaving no trace anywhere. Scenario D proved itself on an event I had not planned: disk_critical alarmed on two filesystems, the second was collapsed by the cooldown, and that collapse is now visible WITH ITS KEY. Yesterday it would have left nothing at all. A gap the spec did not anticipate is recorded with its fix: the per-app event also fires from the periodic sweep, outside any run, so making it record-only would have created a NEW silence. The sweep emits a digest too, with no run_id, so it stays under the ordinary hourly cooldown. ep0: MEASURED on the box — 7757 MB (8 GB), 4 vCPU, and the 4 GiB swapfile SURVIVED the resize and is active (checked, because a resize is a stop/start). The 40 GB local disk is UNCHANGED, so no disk figure was touched anywhere. Five documents corrected — three of which the task's list did not name, found by searching. Two audit/evidence documents ANNOTATED, body untouched: they record what was true when written and that is their value. R-90 CLOSED. R-86 unblocked and re-ranked, stated honestly: 8 GB is comfortable, not unbounded — the original OOM was a 14.46 GB restore — so the restore-test cadence should still be paced, just not by fear of the endpoint. target-selection.md's "D-d did not name ep0 either way" is deliberately left standing. It is the operator's question, not CC's. STATUS.md 127 -> 83 lines, items rather than sentences.
This commit is contained in:
@@ -3,7 +3,8 @@
|
||||
**Class:** supervised operational run. **No repo version bump** — the only commits are this record
|
||||
and the capacity note. **Nothing was deleted.**
|
||||
|
||||
**Host:** `ep0` / `felhom-hetzner`, `167.233.158.164`, Hetzner CX23, Nuremberg.
|
||||
**Host:** `ep0` / `felhom-hetzner`, `167.233.158.164`, Hetzner **CX33 (4 vCPU / 8 GB RAM)**, Nuremberg.
|
||||
> **Rescaled 2026-08-03** from the CX23 (2 vCPU / 3.8 GB) this runbook was written against. **The 40 GB local disk did NOT change** — this was a CPU/RAM resize — so every disk figure below still stands. The 4 GiB swapfile added on 2026-07-27 survived the resize.
|
||||
**Datastore moved:** `felhom-offsite`, `/srv/pbs-felhom` → **`/mnt/pbs-datastore`** (name unchanged).
|
||||
**Window:** 06:58 → 07:19 UTC. PBS down 07:00 → 07:17 UTC.
|
||||
|
||||
|
||||
@@ -229,7 +229,9 @@ as the hub 400ing an unknown event type. `verify-new` verifies each snapshot as
|
||||
`keep-last 2` that covers essentially the whole datastore and turns a dead check live, for a few
|
||||
minutes of ep0 CPU per weekly backup.
|
||||
|
||||
> Watch item: ep0 is a 3.7 GB CX23 with **no swap**, and inline verification runs within the backup
|
||||
> Watch item (**superseded 2026-08-03**: ep0 is now a **CX33, 8 GB RAM**, and it HAS a 4 GiB swapfile
|
||||
> which survived the resize — so the pressure below is much reduced, though the shape of the concern
|
||||
> stands). As written: ep0 is a 3.7 GB CX23 with **no swap**, and inline verification runs within the backup
|
||||
> window. Today's full forced verify completed fine (~250 MiB/s, 0 errors), but see
|
||||
> `RUNBOOK-ep0-datastore-volume-2026-07-27.md` for the rsync OOM on this same box.
|
||||
|
||||
@@ -307,7 +309,7 @@ Untouched. Rollback remains a two-line `datastore.cfg` revert. Volume: 98 G, 13
|
||||
watching: it is the only thing that reclaims chunks, and nothing has ever exercised it here.
|
||||
4. **Hub PBS-DR gauge granularity** — 0.1 GB steps mean routine incremental backups are invisible to
|
||||
it. Not a fault, but it cannot be used as write-proof evidence for small deltas.
|
||||
5. **ep0 has no swap** (3.7 GB CX23) — see the volume runbook's OOM.
|
||||
5. ~~**ep0 has no swap** (3.7 GB CX23)~~ — **corrected 2026-08-03: ep0 is a CX33 with 8 GB RAM and an active 4 GiB swapfile.** See the volume runbook's OOM for the original incident.
|
||||
|
||||
## 11. Observations
|
||||
|
||||
|
||||
@@ -5,7 +5,8 @@
|
||||
> firewall, and the hub-driven `felhom-peersync` reconcile surface. Re-running it on a fresh VM
|
||||
> re-creates the endpoint from nothing (that is the DR story, step 8).
|
||||
>
|
||||
> **Validated:** 2026-07-03 on the dev/test endpoint `felhom-hetzner` (Hetzner CX23, Debian 13,
|
||||
> **Validated:** 2026-07-03 on the dev/test endpoint `felhom-hetzner` (Hetzner CX23 **at the time — rescaled
|
||||
> to a CX33, 4 vCPU / 8 GB RAM, on 2026-08-03; the 40 GB local disk is unchanged**, Debian 13,
|
||||
> `167.233.158.164` / `2a01:4f8:1c16:7aa1::1`) with hub v0.32.0. The production endpoint is a
|
||||
> later re-run of this runbook on a production VM.
|
||||
>
|
||||
@@ -31,7 +32,7 @@ Parameters used throughout (adjust for a new endpoint):
|
||||
points at nothing (live-run finding). Home-resolver propagation can lag public DNS by
|
||||
minutes — a client-side `wg-quick up` that fails to resolve right after record creation
|
||||
just needs a retry.
|
||||
- [ ] Sanity: `ssh root@167.233.158.164 hostname` → `felhom-hetzner` (the throwaway CX23), not
|
||||
- [ ] Sanity: `ssh root@167.233.158.164 hostname` → `felhom-hetzner` (the throwaway box, **CX33 since 2026-08-03**), not
|
||||
any production box.
|
||||
|
||||
## 1. Base (on the box, as root)
|
||||
|
||||
@@ -98,7 +98,8 @@ enrolled host. No access route from DooPlex, and nothing here needs one.
|
||||
### `ep0` (`felhom-hetzner`, `ep0.felhom.eu`) + the Hetzner Storage Boxes — **not protected by D-d; not scratch either**
|
||||
|
||||
Reads are fine. It is the **offsite of last resort** (PBS-DR datastore, WireGuard hub, operator OOB
|
||||
path) and RAM-constrained (3.8 GB, R-90) so a large restore can OOM it. Do not delete datastores, prune
|
||||
path) and — until 2026-08-03 — RAM-constrained (3.8 GB, R-90); it is now a **CX33 with 8 GB RAM and a
|
||||
4 GiB swapfile**, which is what closed R-90. A very large restore is still worth watching. Do not delete datastores, prune
|
||||
jobs, tunnel config or nftables rules; never a drill target. The Storage Boxes hold the restic copy —
|
||||
customer documents and photos, on a credential that can still delete (R-95).
|
||||
**Access: `ssh root@167.233.158.164` from DooPlex** — *not* `felhom-pve → 10.77.0.1`, the route that
|
||||
|
||||
Reference in New Issue
Block a user