From 4c4e3b3a3f05c2fb474885566a123a4f99e5705c Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 15 Sep 2026 10:42:53 +0200 Subject: [PATCH] docs: supervisor (03), node_* ruling (08, CONTEXT), per-tier page + tier skip (07), self-bind triggers + PBS-DR lifecycle (05), settings after install (02), park + No TLS Verify runbooks, volunteer prerequisites Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 8 +++++ .../architecture/02-controller-module-map.md | 12 +++++++ documentation/architecture/03-host-agent.md | 15 ++++++++ .../architecture/05-hub-architecture.md | 35 ++++++++++++++++++- .../architecture/07-backup-architecture.md | 14 ++++++++ documentation/architecture/08-alarm-ladder.md | 9 +++++ .../runbooks/RUNBOOK-manual-build.md | 7 ++++ .../runbooks/VOLUNTEER-first-hour.md | 3 ++ documentation/runbooks/day0-install.md | 7 ++++ 9 files changed, 109 insertions(+), 1 deletion(-) diff --git a/CONTEXT.md b/CONTEXT.md index 735f7d53..70d8016e 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,14 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## Decision 2026-09-15 — „box is down" mail skips the one-hour quiet rule (operator ruling, decision A) + +`node_stale`, `node_down`, `node_recovered` bypass the 1-hour operator cooldown, with a 5-minute +dedupe (hub v0.114.0). Reverses the documented design for those three types; recorded in +`documentation/architecture/08-alarm-ladder.md` §6.2. Reason: BIGNIGHT F9 — a 33-minute dead box whose +`node_stale` and `node_recovered` mails were both suppressed by an earlier incident's cooldown. Same day, +also the operator's: the self-bind link goes out whenever a customer waits for a box (05 §13). + ## A customer's domain is their own, and the public installer stays interactive (2026-09-14, R-494 / R-495) **Two operator rulings, both 2026-09-14.** diff --git a/documentation/architecture/02-controller-module-map.md b/documentation/architecture/02-controller-module-map.md index 0bf677eb..d093b927 100644 --- a/documentation/architecture/02-controller-module-map.md +++ b/documentation/architecture/02-controller-module-map.md @@ -601,3 +601,15 @@ measured; that the restore writes to that path is read, not measured. catalog's version within 15 minutes (observed). Any "don't touch deployed apps" rule loses that. - `up -d`-on-restart is what injects `app.yaml` env into a running stack. Reverting to `docker compose restart` would silently stop doing that. + +## App settings after install [DESIGN, recorded 2026-09-15] + +**An installed app's settings are read-only on its page** („Ez az alkalmazás már telepítve van. Az alábbi +beállítások csak olvashatók.", `deploy.html`). This is a design, not an oversight: a changed value would +need a guarded re-deploy that nothing performs. **Finding (2026-09-15):** the catalog field flag +`locked_after_deploy` (`stacks/metadata.go`) is parsed and read by **no** controller code — every field +is read-only after install whatever the catalog says. So a page must never tell a household to change +a value after install. **Vaultwarden (R-512)** follows the design instead of breaking it: registration +is closed by default and the household is invited from the admin panel, measured to work without mail +(stranger 400, invite 200, invited 200). + diff --git a/documentation/architecture/03-host-agent.md b/documentation/architecture/03-host-agent.md index 543577e1..87be928a 100644 --- a/documentation/architecture/03-host-agent.md +++ b/documentation/architecture/03-host-agent.md @@ -94,6 +94,21 @@ by verb**: - **An operator signature is always required** to destroy/overwrite any resource holding the only/primary copy of customer data — live-guest destroy, storage detach/wipe, restore-overwrite, decommission — *regardless of whether it arrives as a job or as a desired-state delta*. A compromised hub cannot forge them because the signing key is **not held by the hub** (it lives with the operator / a separate signing path; the hub only queues opaque signed blobs). - **Data-bearing-ness is agent-internal evidence, never a caller's claim (slice 8C).** For a customer-driven storage op (`POST /disks/format`, §6) the agent **inspects the actual device** (filesystem signature / partition table / partitions / mount, conservative — ambiguous → data-bearing) to decide the class. A blank device → benign self-serve `mkfs`; a data-bearing device → `ClassStorageWipe` → this gate → `pending_signature`. The **destructive completion of a data-bearing wipe is slice 10** (the operator-signed path); 8C refuses it. This mirrors the provenance rule above: just as the scratch tag is agent-internal (never hub-sourced), data-bearing-ness is agent-observed (never controller-asserted) — a compromised controller cannot relabel a data-bearing drive "blank" to walk the gate. - **Healing a crashed controller is non-destructive by construction:** it is reconstructable from its image + the guest's persistent volume, so "redeploy" = restart the LXC / `docker compose up -d` **inside the existing guest** — never a guest destroy. (v0.33 precedent: `watchdog.go` restarts stopped stacks, it never destroys the guest.) +- **The controller supervisor is that sentence made real (R-523, agent v0.131.0).** Measured + 2026-09-15: after `docker kill`, Docker restarts neither an `unless-stopped` nor an `always` + container, and the golden's `felhom-controller-bootstrap.service` is a oneshot that watches nothing. + So every 30 s the agent checks, for each felhom-pool guest it provisioned + (`/var/lib/felhom-agent/guests//bootstrap` exists) that is running, whether + `felhom-controller` is running; on the **second** consecutive "no" it runs + `systemctl restart felhom-controller-bootstrap.service` in the guest — the swap's own restart, over + the same two sudoers grants. **Guards:** not during a controller swap; not when parked + (`touch /var/lib/felhom-agent/guests//controller-parked` on the host); not on a stopped, + locked or vzdump-busy guest; not on an unknown docker answer; and no thrash — 3 restarts in 15 minutes + stop restarts for 30 minutes. **Events:** the agent has no event channel; its host report carries a + `controller_supervisor` stanza, and the hub mints `controller_restarted_by_agent` (info) and + `controller_crashloop` (error), both operator-only, keyed on timestamps so an agent restart can + neither lose nor invent one. New goldens run the controller with `--restart always`, which covers + only a Docker daemon restart. Signed payloads carry a **nonce + expiry** (anti-replay: a captured "restore" job cannot be re-injected later) and a target binding (host + guest id) so a signature can't be retargeted. diff --git a/documentation/architecture/05-hub-architecture.md b/documentation/architecture/05-hub-architecture.md index 60639fb3..ff4d50f0 100644 --- a/documentation/architecture/05-hub-architecture.md +++ b/documentation/architecture/05-hub-architecture.md @@ -250,4 +250,37 @@ customer instead of a single controller. - Multi-tenant resource fairness (deferred shared-host case). - Hub-side desired-state **editing UX** specifics (form/diff wiring) — to be grounded against `hub/internal/web/configs.go` at implementation. -- Golden-image refresh cadence / fleet versioning (carried from Part 3 §13). \ No newline at end of file +- Golden-image refresh cadence / fleet versioning (carried from Part 3 §13). + +## 13. When the self-bind (connect) link is sent [DESIGN, operator decision A 2026-09-15, hub v0.114.0] + +The box's console tells the volunteer to open „az e-mailben kapott link". That mail must exist whenever +a customer is **waiting for a box**, without an operator press. `autoMintSelfBindLink` runs on four +events: **customer creation**; **RESET completion**; **an e-mail set or changed on a customer with no +bound host** (config save); **a host delete** that keeps the customer. The last two re-check that no +host is bound, so a customer who has a box is never sent a pairing link. **Appliance registration is +not a trigger** — it knows no customer (`api/appliance.go`). Every send is stored as a hub-internal +`selfbind_link_sent` event with its occasion; the Setup tab shows „Kapcsolódó link elküldve: +()" and keeps the button as the manual resend. Until 2026-09-15 only the first two existed, +and BIGNIGHT's tester-1 never got its mail (R-509). + +## 14. The PBS-DR descriptor and its endpoint token — lifecycle [DESIGN, recorded 2026-09-15] + +Two records must agree: the **token** on the endpoint (ep0: namespace `` + token +`felhom@pbs!`) and the **descriptor** in the host's `desired_json` (`pbs_dr`: namespace, +token id, fingerprint, datastore, tunnel IP, `secret_generation`) with its consume-once secret. + +| event | token on ep0 | descriptor on hub | +|---|---|---| +| DR tier ON + WG peer present (form save or WG hook) | **provision** | created, secret stored | +| re-issue, descriptor present | **re-keyed** (delete + recreate) | `secret_generation` bumped | +| **re-issue, descriptor ABSENT, DR flag on** (R-511, v0.114.0) | **adopted**: re-keyed | **rebuilt** from the endpoint's answer; `pbsdr_adopted` audit row | +| host delete | **kept** (tenancy survives) | goes with the host | +| host delete acknowledged through escrow-ack, then a new box | re-keyed automatically (F-14) | rebuilt | +| RESET / customer delete | **deprovisioned** — namespace, every backup group AND token | purged | + +**Why adopt exists:** a rebuilt box's WG hook refused („the endpoint already holds a PBS token … use +the explicit Re-issue action") and the re-issue itself then refused with 400 — no button restored the +tier. **Not built:** releasing ONLY the token on host delete. The endpoint's only removal op +(`deprovision`) destroys the backups too, so a token-only release needs a new endpoint operation. + diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index 92501d82..b3fda807 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -479,6 +479,20 @@ sentence, covers installed-but-no-resolvable-data-root. **R-356**; reasoning als --- +### 6.4 The customer's page speaks per tier, and a tier without storage is skipped (2026-09-15) + +**What the page may claim** (controller v0.243.0 + agent v0.131.0, R-517). The whole-system tile used +to show the agent's single latest record — so a failed 0-byte PBS attempt read „Naprakész" and ticked +„Távoli rendszermentés — külön hardveren" (BIGNIGHT 2026-09-14). Now `GET /backup/status` returns per +tier the newest **success**, the last **attempt** kept apart, and whether the tier's storage exists. +The page shows each tier's newest success; a failed attempt under it as „sikertelen"; a tier whose +storage does not exist as „nincs beállítva". „Naprakész" and the remote tick are computed from +successes only. After an agent restart the success is read back from the tier's storage (size unknown). + +**What a run may do** (R-518, cheap half). A tier the agent reports `storage: absent` is dropped before +anything is stopped, logged, and reported once as `backup_tier_skipped`; `unknown` is never skipped. +**Still open:** quiescing per tier, so a slow second tier does not keep every app down. + ## 7. The recovery chain (D3) — the reason this document exists **[DESIGN] 3-2-1 describes copies. It does not describe recovery.** diff --git a/documentation/architecture/08-alarm-ladder.md b/documentation/architecture/08-alarm-ladder.md index dd964d9d..72467c02 100644 --- a/documentation/architecture/08-alarm-ladder.md +++ b/documentation/architecture/08-alarm-ladder.md @@ -185,6 +185,15 @@ signal about a real class of failure, and it is narrower than the phrase suggest **"How loud" is a separate decision from "does it alarm", and it is made in one place**: the operator cooldown key at `processOperator`. Everything sharing a key is collapsed for **one hour**. +**Operator ruling 2026-09-15 (decision A) — „the box is down" skips the quiet hour.** `node_stale`, +`node_down` and `node_recovered` no longer share the one-hour operator cooldown; they keep a **5-minute +dedupe** on the same key, so a flapping link cannot mail every sweep (`operatorCooldownFor`, hub +v0.114.0, pinned by `TestOperatorCooldown_NodeLivenessBypassesQuietHour`). **This reverses the design +above for those three types only**, and it is a ruling, not a defect fix. The reason is BIGNIGHT F9 +(2026-09-14): the controller was dead for 33 minutes; its `node_stale` mail was suppressed because F8's +`node_stale` had used the hour 39 minutes earlier, and the later `node_recovered` mail was suppressed +the same way. The `host_*` agent-plane siblings are **not** in the ruling and keep the hour. + | Family | Grain | Key carries | Why | |---|---|---|---| | app down (`app_start_failed`) | **per APP** | `…:` | no digest exists for it — see below | diff --git a/documentation/runbooks/RUNBOOK-manual-build.md b/documentation/runbooks/RUNBOOK-manual-build.md index 51899e12..5fb0a8f1 100644 --- a/documentation/runbooks/RUNBOOK-manual-build.md +++ b/documentation/runbooks/RUNBOOK-manual-build.md @@ -110,6 +110,13 @@ normal route — do not hand-deploy. 4. Verify: `sudo kubectl -n felhom-system logs deploy/hub | grep 'managed floor SERVED'` shows `from "declared"`, and each box logs `SetFloor: floor "…" → ""` within about a minute. +### 3.1 Park the controller (stop the agent restarting it) — agent ≥ 0.131.0 + +The agent restarts a stopped `felhom-controller` within about a minute (R-523). To keep it stopped on +purpose, on the **Proxmox host**: `touch /var/lib/felhom-agent/guests//controller-parked`. The +agent logs `controller is not running and the guest is PARKED — leaving it`. To unpark: +`rm /var/lib/felhom-agent/guests//controller-parked`. A controller swap is never fought either. + ## 4. Golden image (fresh Day-0 installs) The golden is a pre-baked controller-era guest image built in the **drill VM** on 180 diff --git a/documentation/runbooks/VOLUNTEER-first-hour.md b/documentation/runbooks/VOLUNTEER-first-hour.md index 9b924d2a..f301668b 100644 --- a/documentation/runbooks/VOLUNTEER-first-hour.md +++ b/documentation/runbooks/VOLUNTEER-first-hour.md @@ -30,6 +30,9 @@ ## Az üzemeltető előtte (nem az önkéntes feladata) 1. Létrehozza az ügyfelet a hubon (név, e-mail, domain). + A kapcsolódó linket tartalmazó e-mailt a hub **magától küldi**, amikor az ügyfélnek van e-mail címe és + még nincs gépe (létrehozáskor, e-mail megadásakor, egy korábbi gép törlésekor). Gombot nyomni nem kell; + az ügyfél oldalán látszik, mikor ment ki (R-509, 2026-09-15). 2. **Létrehozza a Cloudflare tunnelt és beírja a tokent** az ügyfél adatlapján — enélkül a vezérlőpult címe nem nyílik meg (R-494). 3. **Személyesen vagy üzenetben átadja a 5 szóból álló „Tulajdonosi jelmondatot".** Ezt semmilyen diff --git a/documentation/runbooks/day0-install.md b/documentation/runbooks/day0-install.md index 53edb83d..a71f0f22 100644 --- a/documentation/runbooks/day0-install.md +++ b/documentation/runbooks/day0-install.md @@ -57,6 +57,13 @@ Tunnels): Also have ready (optional but recommended): a **Cloudflare API token** with Zone edit rights for the customer's zone — the hub uses it for geo-restriction management. +**„No TLS Verify" on the public hostname (R-510, 2026-09-15).** The route sends traffic to +`https://traefik`, and traefik answers the name `traefik` with its own default certificate. Tick +**No TLS Verify** under the public hostname's TLS settings (demo-hp's working tunnel has +`originRequest.noTLSVerify: true`). Without it a fresh box answers **502**. *Not yet proven on a +volunteer's route: on 2026-09-15 tester-1 had no box, and three GETs returned 530 (no tunnel +connected).* + ### A.2 Create the customer in the hub Hub UI (`https://hub.felhom.eu`, operator password) → **Customers → New**: