docs: supervisor (03), node_* ruling (08, CONTEXT), per-tier page + tier skip (07), self-bind triggers + PBS-DR lifecycle (05), settings after install (02), park + No TLS Verify runbooks, volunteer prerequisites
gates / gates (push) Successful in 18s

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-15 10:42:53 +02:00
parent d8cd4d4412
commit 4c4e3b3a3f
9 changed files with 109 additions and 1 deletions
@@ -601,3 +601,15 @@ measured; that the restore writes to that path is read, not measured.
catalog's version within 15 minutes (observed). Any "don't touch deployed apps" rule loses that.
- `up -d`-on-restart is what injects `app.yaml` env into a running stack. Reverting to
`docker compose restart` would silently stop doing that.
## App settings after install [DESIGN, recorded 2026-09-15]
**An installed app's settings are read-only on its page** („Ez az alkalmazás már telepítve van. Az alábbi
beállítások csak olvashatók.", `deploy.html`). This is a design, not an oversight: a changed value would
need a guarded re-deploy that nothing performs. **Finding (2026-09-15):** the catalog field flag
`locked_after_deploy` (`stacks/metadata.go`) is parsed and read by **no** controller code — every field
is read-only after install whatever the catalog says. So a page must never tell a household to change
a value after install. **Vaultwarden (R-512)** follows the design instead of breaking it: registration
is closed by default and the household is invited from the admin panel, measured to work without mail
(stranger 400, invite 200, invited 200).
@@ -94,6 +94,21 @@ by verb**:
- **An operator signature is always required** to destroy/overwrite any resource holding the only/primary copy of customer data — live-guest destroy, storage detach/wipe, restore-overwrite, decommission — *regardless of whether it arrives as a job or as a desired-state delta*. A compromised hub cannot forge them because the signing key is **not held by the hub** (it lives with the operator / a separate signing path; the hub only queues opaque signed blobs).
- **Data-bearing-ness is agent-internal evidence, never a caller's claim (slice 8C).** For a customer-driven storage op (`POST /disks/format`, §6) the agent **inspects the actual device** (filesystem signature / partition table / partitions / mount, conservative — ambiguous → data-bearing) to decide the class. A blank device → benign self-serve `mkfs`; a data-bearing device → `ClassStorageWipe` → this gate → `pending_signature`. The **destructive completion of a data-bearing wipe is slice 10** (the operator-signed path); 8C refuses it. This mirrors the provenance rule above: just as the scratch tag is agent-internal (never hub-sourced), data-bearing-ness is agent-observed (never controller-asserted) — a compromised controller cannot relabel a data-bearing drive "blank" to walk the gate.
- **Healing a crashed controller is non-destructive by construction:** it is reconstructable from its image + the guest's persistent volume, so "redeploy" = restart the LXC / `docker compose up -d` **inside the existing guest** — never a guest destroy. (v0.33 precedent: `watchdog.go` restarts stopped stacks, it never destroys the guest.)
- **The controller supervisor is that sentence made real (R-523, agent v0.131.0).** Measured
2026-09-15: after `docker kill`, Docker restarts neither an `unless-stopped` nor an `always`
container, and the golden's `felhom-controller-bootstrap.service` is a oneshot that watches nothing.
So every 30 s the agent checks, for each felhom-pool guest it provisioned
(`/var/lib/felhom-agent/guests/<vmid>/bootstrap` exists) that is running, whether
`felhom-controller` is running; on the **second** consecutive "no" it runs
`systemctl restart felhom-controller-bootstrap.service` in the guest — the swap's own restart, over
the same two sudoers grants. **Guards:** not during a controller swap; not when parked
(`touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` on the host); not on a stopped,
locked or vzdump-busy guest; not on an unknown docker answer; and no thrash — 3 restarts in 15 minutes
stop restarts for 30 minutes. **Events:** the agent has no event channel; its host report carries a
`controller_supervisor` stanza, and the hub mints `controller_restarted_by_agent` (info) and
`controller_crashloop` (error), both operator-only, keyed on timestamps so an agent restart can
neither lose nor invent one. New goldens run the controller with `--restart always`, which covers
only a Docker daemon restart.
Signed payloads carry a **nonce + expiry** (anti-replay: a captured "restore" job cannot be
re-injected later) and a target binding (host + guest id) so a signature can't be retargeted.
@@ -250,4 +250,37 @@ customer instead of a single controller.
- Multi-tenant resource fairness (deferred shared-host case).
- Hub-side desired-state **editing UX** specifics (form/diff wiring) — to be grounded against
`hub/internal/web/configs.go` at implementation.
- Golden-image refresh cadence / fleet versioning (carried from Part 3 §13).
- Golden-image refresh cadence / fleet versioning (carried from Part 3 §13).
## 13. When the self-bind (connect) link is sent [DESIGN, operator decision A 2026-09-15, hub v0.114.0]
The box's console tells the volunteer to open „az e-mailben kapott link". That mail must exist whenever
a customer is **waiting for a box**, without an operator press. `autoMintSelfBindLink` runs on four
events: **customer creation**; **RESET completion**; **an e-mail set or changed on a customer with no
bound host** (config save); **a host delete** that keeps the customer. The last two re-check that no
host is bound, so a customer who has a box is never sent a pairing link. **Appliance registration is
not a trigger** — it knows no customer (`api/appliance.go`). Every send is stored as a hub-internal
`selfbind_link_sent` event with its occasion; the Setup tab shows „Kapcsolódó link elküldve: <date>
(<occasion>)" and keeps the button as the manual resend. Until 2026-09-15 only the first two existed,
and BIGNIGHT's tester-1 never got its mail (R-509).
## 14. The PBS-DR descriptor and its endpoint token — lifecycle [DESIGN, recorded 2026-09-15]
Two records must agree: the **token** on the endpoint (ep0: namespace `<customer>` + token
`felhom@pbs!<customer>`) and the **descriptor** in the host's `desired_json` (`pbs_dr`: namespace,
token id, fingerprint, datastore, tunnel IP, `secret_generation`) with its consume-once secret.
| event | token on ep0 | descriptor on hub |
|---|---|---|
| DR tier ON + WG peer present (form save or WG hook) | **provision** | created, secret stored |
| re-issue, descriptor present | **re-keyed** (delete + recreate) | `secret_generation` bumped |
| **re-issue, descriptor ABSENT, DR flag on** (R-511, v0.114.0) | **adopted**: re-keyed | **rebuilt** from the endpoint's answer; `pbsdr_adopted` audit row |
| host delete | **kept** (tenancy survives) | goes with the host |
| host delete acknowledged through escrow-ack, then a new box | re-keyed automatically (F-14) | rebuilt |
| RESET / customer delete | **deprovisioned** — namespace, every backup group AND token | purged |
**Why adopt exists:** a rebuilt box's WG hook refused („the endpoint already holds a PBS token … use
the explicit Re-issue action") and the re-issue itself then refused with 400 — no button restored the
tier. **Not built:** releasing ONLY the token on host delete. The endpoint's only removal op
(`deprovision`) destroys the backups too, so a token-only release needs a new endpoint operation.
@@ -479,6 +479,20 @@ sentence, covers installed-but-no-resolvable-data-root. **R-356**; reasoning als
---
### 6.4 The customer's page speaks per tier, and a tier without storage is skipped (2026-09-15)
**What the page may claim** (controller v0.243.0 + agent v0.131.0, R-517). The whole-system tile used
to show the agent's single latest record — so a failed 0-byte PBS attempt read „Naprakész" and ticked
„Távoli rendszermentés — külön hardveren" (BIGNIGHT 2026-09-14). Now `GET /backup/status` returns per
tier the newest **success**, the last **attempt** kept apart, and whether the tier's storage exists.
The page shows each tier's newest success; a failed attempt under it as „sikertelen"; a tier whose
storage does not exist as „nincs beállítva". „Naprakész" and the remote tick are computed from
successes only. After an agent restart the success is read back from the tier's storage (size unknown).
**What a run may do** (R-518, cheap half). A tier the agent reports `storage: absent` is dropped before
anything is stopped, logged, and reported once as `backup_tier_skipped`; `unknown` is never skipped.
**Still open:** quiescing per tier, so a slow second tier does not keep every app down.
## 7. The recovery chain (D3) — the reason this document exists
**[DESIGN] 3-2-1 describes copies. It does not describe recovery.**
@@ -185,6 +185,15 @@ signal about a real class of failure, and it is narrower than the phrase suggest
**"How loud" is a separate decision from "does it alarm", and it is made in one place**: the operator
cooldown key at `processOperator`. Everything sharing a key is collapsed for **one hour**.
**Operator ruling 2026-09-15 (decision A) — „the box is down" skips the quiet hour.** `node_stale`,
`node_down` and `node_recovered` no longer share the one-hour operator cooldown; they keep a **5-minute
dedupe** on the same key, so a flapping link cannot mail every sweep (`operatorCooldownFor`, hub
v0.114.0, pinned by `TestOperatorCooldown_NodeLivenessBypassesQuietHour`). **This reverses the design
above for those three types only**, and it is a ruling, not a defect fix. The reason is BIGNIGHT F9
(2026-09-14): the controller was dead for 33 minutes; its `node_stale` mail was suppressed because F8's
`node_stale` had used the hour 39 minutes earlier, and the later `node_recovered` mail was suppressed
the same way. The `host_*` agent-plane siblings are **not** in the ruling and keep the hour.
| Family | Grain | Key carries | Why |
|---|---|---|---|
| app down (`app_start_failed`) | **per APP** | `…:<stack_name>` | no digest exists for it — see below |
@@ -110,6 +110,13 @@ normal route — do not hand-deploy.
4. Verify: `sudo kubectl -n felhom-system logs deploy/hub | grep 'managed floor SERVED'` shows
`from "declared"`, and each box logs `SetFloor: floor "…" → "<VER>"` within about a minute.
### 3.1 Park the controller (stop the agent restarting it) — agent ≥ 0.131.0
The agent restarts a stopped `felhom-controller` within about a minute (R-523). To keep it stopped on
purpose, on the **Proxmox host**: `touch /var/lib/felhom-agent/guests/<vmid>/controller-parked`. The
agent logs `controller is not running and the guest is PARKED — leaving it`. To unpark:
`rm /var/lib/felhom-agent/guests/<vmid>/controller-parked`. A controller swap is never fought either.
## 4. Golden image (fresh Day-0 installs)
The golden is a pre-baked controller-era guest image built in the **drill VM** on 180
@@ -30,6 +30,9 @@
## Az üzemeltető előtte (nem az önkéntes feladata)
1. Létrehozza az ügyfelet a hubon (név, e-mail, domain).
A kapcsolódó linket tartalmazó e-mailt a hub **magától küldi**, amikor az ügyfélnek van e-mail címe és
még nincs gépe (létrehozáskor, e-mail megadásakor, egy korábbi gép törlésekor). Gombot nyomni nem kell;
az ügyfél oldalán látszik, mikor ment ki (R-509, 2026-09-15).
2. **Létrehozza a Cloudflare tunnelt és beírja a tokent** az ügyfél adatlapján — enélkül a
vezérlőpult címe nem nyílik meg (R-494).
3. **Személyesen vagy üzenetben átadja a 5 szóból álló „Tulajdonosi jelmondatot".** Ezt semmilyen
+7
View File
@@ -57,6 +57,13 @@ Tunnels):
Also have ready (optional but recommended): a **Cloudflare API token** with Zone edit rights for the
customer's zone — the hub uses it for geo-restriction management.
**„No TLS Verify" on the public hostname (R-510, 2026-09-15).** The route sends traffic to
`https://traefik`, and traefik answers the name `traefik` with its own default certificate. Tick
**No TLS Verify** under the public hostname's TLS settings (demo-hp's working tunnel has
`originRequest.noTLSVerify: true`). Without it a fresh box answers **502**. *Not yet proven on a
volunteer's route: on 2026-09-15 tester-1 had no box, and three GETs returned 530 (no tunnel
connected).*
### A.2 Create the customer in the hub
Hub UI (`https://hub.felhom.eu`, operator password) → **Customers → New**: