diff --git a/STATUS.md b/STATUS.md index 5b97f478..a521784b 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,9 +2,35 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** -**Updated 2026-10-08 (day): hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller 0.303.0 -(nothing delivered today — tonight is the second kernel night). The open-items list is at 130. Reports: `REPORT.md`, -`REPORT-day-2026-10-08.md`.** +**Updated 2026-10-08 (afternoon): hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller +0.303.0 (nothing delivered today — tonight is the second kernel night). The open-items list is at 130. Reports: +`REPORT-day-2026-10-08.md` (morning), `REPORT-day2-2026-10-08.md` (afternoon).** + +## Afternoon (2026-10-08): your four answers built, and one sheet of decisions + +- **The FAQ is honest now** (you approved the text): it says where backup copies and remote traffic go. Live. +- **Deletion times:** the hub deletes a removed customer's mail and event records 1 year after the deletion (ships + tomorrow). The job that removes a removed customer's copy on DooPlex within 30 days is written and tested, **not + installed** (see D9). Both times are in the privacy-notice draft. +- **You get a mail when a household's code opens, or may open, an old package** (ships tomorrow). +- **Cloudflare:** on the list as a later item. +- **Nine stuck items have a one-page design.** Two were no longer true and are closed; one lost item is back on the list. + +### The decision sheet — answer „all as picked", or name the numbers you change + +| # | Question | Pick | Cost | If you do nothing | +|---|---|---|---|---| +| D1 | May the hub (with your password, no signing key) ask a box to run an off-site backup now, run a named check now, and stop or extend a household's deletion countdown? | Yes — a fixed list of safe, repeatable actions, carried in the box's report reply | ~1 session, controller + hub | A household calling in the first 14 days of a countdown needs a shell on its box; only the household can start an off-site run | +| D2 | When the hub has had no connection from a box for ~6 minutes, may „delete host" go ahead at once after you tick „I checked: the box is off"? | Yes (step 1, showing the live link on the host page, needs no answer) | ~½ session, hub | Deleting a switched-off box (and a reset) waits up to 45 minutes after its last report | +| D3 | May the household's „system health" mail drop the raw technical note and point to the dashboard for the list? | Yes (the dashboard half needs no answer; CC builds it) | ~1 h hub + a hub release | The mail keeps a curly-bracket note with English lines in it | +| D4 | May the box remember a dashboard sign-in across its own restarts (on disk only a fingerprint that cannot be used to sign in)? | Yes | ~½ session, controller | The household is logged out at every settings push and every controller update | +| D5 | Is the measured address block enough for opengist's sign-up, or should the box also close opengist's own switch? | Enough — close the item | Nothing | Same as the pick: the item closes | +| D6 | Should the hub check with Cloudflare that a pasted key reaches only that customer's own domain? | Yes (the „no duplicate or nested domain" guard needs no answer; CC builds it) | ~½ session, hub; a save fails while Cloudflare is down | A key made for the whole account by mistake could change every household's web addresses | +| D7 | Should the hub mail you an error when even one off-site snapshot disappears outside a clean-up window it opened? | Yes | ~½ session, hub | A deletion through any other key stays silent unless it removes over half of a household's history | +| D8 | After a failed off-site restore of one app: keep the app stopped for support now, and „put back exactly as it was" next? | Yes, both, in that order | ~½ session now; ~1 session + a test + disk space for one copy later | The app restarts on a mix of old files and a newer database, and the screen says all is back as it was | +| D9 | May CC install the DooPlex job that removes a deleted customer's copy (dry run first, then daily)? | Yes, after tomorrow's releases are read back | ~30 min; one new local token on DooPlex | A deleted customer's copy stays on DooPlex, and the privacy-notice line „within 30 days" is not true | + +Full designs: `documentation/audits/day-2026-10-08/`. ## Day (2026-10-08): fixes built for tomorrow, the old-code answer made honest, legal drafts diff --git a/documentation/audits/day-2026-10-08/design-R-138.md b/documentation/audits/day-2026-10-08/design-R-138.md new file mode 100644 index 00000000..c68fc324 --- /dev/null +++ b/documentation/audits/day-2026-10-08/design-R-138.md @@ -0,0 +1,52 @@ +# R-138 — a Cloudflare token that can write a shared zone, on a customer's box: a one-page design (2026-10-08) + +Baselines read: felhom.eu `b2dce901`, felhom-controller `a0370b4`. Architecture: `01-topology-and-trust.md` §5 +(trust boundaries: „hub ↔ Cloudflare API") and §7 (networking: „Every customer has their OWN domain — never a name +under `felhom.eu`", operator ruling 2026-09-14). **Status:** design only, nothing built. The token itself is stored +out-of-band; it was not read, printed or used for this page. + +## 1. The problem, and what changed under it +- **The mechanism is still as the row says.** The hub form copies `cf_api_token` verbatim into the customer's config + (`hub/internal/web/configs.go:1682-1684`); the controller writes it to a 0600 `.env` for Traefik + (`controller/internal/infra/infra.go:157-159`; the row's `:123` is stale). An empty token selects HTTP-01 + (`controller/internal/infra/templates/traefik.yml.tmpl:52-61`). The controller also uses it for the geo-WAF + (`controller/internal/web/handlers.go:642`). +- **The row's risk needs a SHARED customer zone, and the design has ruled that out.** `01` §7 (2026-09-14): every + customer has their own domain, never a name under `felhom.eu`. Today's boxes match: demo-felhom, demo-hp and + Tester 1 (`enkicsifelhom.hu`, `operations/nodes.md:44`) each sit on their own zone. The shared-zone plan the row was + written against (`audits/RECON-subdomain-onboarding-2026-07-31.md` §5) is not the plan any more. +- **But nothing ENFORCES the ruling.** The hub accepts any domain, including one equal to or under another customer's + domain, or under `felhom.eu`: `domain TEXT NOT NULL DEFAULT ''` with no uniqueness (`hub/internal/store/store.go:164`), + and the create path refuses only a duplicate customer id (`configs.go:740`). That check was row **R-415** (READY, XS) — + **it was removed from `OPEN-ITEMS.md` in the 2026-10-03 triage (`71b8c8c6`) and never reached `CLOSED-ITEMS.md`**; + it is lost, not closed. +- **A second residue the row does not name:** the four zones sit in ONE Cloudflare account (RECON §2, line 163). A + token minted with account scope (all zones) instead of one zone would let one box rewrite every household's DNS — + the same blast radius as a shared zone, by a typing mistake in the token wizard. Nobody checks the token's reach. + +## 2. Options +| | What | Costs | Risk | +|---|---|---|---| +| **A** | Close R-138 on the ruling; do nothing more. | Nothing. | The ruling is a sentence; an operator mistake (a duplicate or nested domain, an account-wide token) passes silently. | +| **B** | **Enforce the ruling on save (revive R-415):** the hub refuses a domain that equals, contains or sits under another customer's domain, or sits under `felhom.eu`. | Hub only; one store query, one form error, tests. ~¼ session. | None to customers: it refuses an operator input that the ruling already forbids. | +| **C** | B **plus a reach check on the token:** when a `cf_api_token` is saved, the hub asks Cloudflare which zones that token can see (`GET /zones`, the call the hub already makes in `hub/internal/cloudflare/unblock.go:115`) and refuses unless it sees exactly the customer's own zone. | Hub only; one outbound call at save time (an existing dependency, not a new one); a save fails while Cloudflare is down — the form must say so and keep what was typed. ~½ session. | A save blocked by a Cloudflare outage; the operator retries. | + +## 3. The pick — B now (follows the ruling, no decision needed); C on the operator's word +B is the guard the row asked for, re-scoped: not „refuse a token for a shared-zone customer" (no such customer may +exist) but „refuse the layout that would make one". C closes the account-wide-token hole, which is the same danger by +another road; it adds a network call to a form save, so it is the operator's choice. + +## 4. First slice and its red test +- **Hub (B):** in the create and edit paths, before any provisioning: `store.DomainConflicts(customerID, domain)` and a + `felhom.eu` suffix check; the form re-renders with the submitted values and one sentence. +- **Red test (fails on today's code):** create customer `a` with `example.hu`; creating `b` with `example.hu`, + `x.example.hu` or `t1.felhom.eu` is refused and nothing is stored. **Controls:** `b` with `example2.hu` is accepted; + editing `a` with its own `example.hu` is accepted (no self-conflict); `notexample.hu` is accepted (suffix match is on a + label boundary). +- Register: R-415 re-filed (or folded into this row) and R-138 closed on B's commit. + +## 5. One question for the operator +**When you paste a customer's Cloudflare key into the hub, should the hub check with Cloudflare that the key reaches +only that customer's own domain, and refuse it otherwise?** My pick: yes (option C). *If you do nothing:* a key that was +made for the whole Cloudflare account by mistake goes onto the customer's box, and that box could change the web +addresses of every other household. diff --git a/documentation/audits/day-2026-10-08/design-R-30.md b/documentation/audits/day-2026-10-08/design-R-30.md new file mode 100644 index 00000000..738af904 --- /dev/null +++ b/documentation/audits/day-2026-10-08/design-R-30.md @@ -0,0 +1,72 @@ +# R-30 — box presence from the live connection, not the report clock: a one-page design (2026-10-08) + +Baselines read: felhom.eu `b2dce901`, felhom-controller `a0370b4`, felhom-agent `b228b44`. Architecture: +`05-hub-architecture.md` §4 (liveness / dead-man's-switch); `03-host-agent.md` §4 (the host-delete guard's reason). +**Status: design only. Nothing is built.** + +## 1. The problem, with today's numbers + +- **Measured 2026-07-21 (the row):** a powered-off box stayed „healthy" on the hub until `host_stale` fired after the + threshold („no report for 30m"). **Today the wait is longer:** the threshold is 45 minutes (live ConfigMap + `felhom-system/hub-config`, `alerting.stale_threshold: "45m"`, operator ruling A on R-549). The row's „30 min" is stale. +- The agent reports every 900 s (hub `internal/api/handler.go:653-655`, `defaultHostPollSeconds = 900`; agent + `internal/config/config.go:849`). „Online" on the host page is report age under the threshold + (`internal/web/hosts.go:31-44`). **Host delete refuses while „online"** (`hosts.go:935-937`, and since R-599 it says when + the refusal ends). **RESET refuses while any host row exists** (`internal/web/customer_reset.go:106-110`). So a forced + teardown of a box that is already off waits up to 45 minutes after its last report. +- **Measured today (read-only, the ingress access log on DooPlex, 07:0x–07:50 UTC):** the controller's wait channel + (`GET /api/v1/wait`) completes every **241–243 s** per box — three sources, 11–12 holds each, every hold 240.00x s. + So a healthy box starts a new wait at most ~243 s after the previous one started. +- **Not measured:** what the hub sees when a box loses power in the middle of a hold. Read in source: the hub writes a + newline every 25 s into nginx (`internal/api/wait.go:10-23`) and ends the hold itself at 240 s, so the hold most likely + ends normally and then **no new wait arrives**. That gap is the signal. Slice 1 measures it. + +## 2. The three signals + +| Signal | Cadence | What it proves | Blind spot | +|---|---|---|---| +| Agent host report | 900 s | the host and its agent | slow; the alarm needs 45 min hysteresis (R-549) | +| Controller wait channel | ~241 s | the guest's controller is running and reaching the hub | a crashed or updating controller looks like a dead box; per customer, not per host | +| WireGuard handshake age on ep0 | ~2 min (protocol rekey; not measured here) | the host's kernel, independent of both processes | needs a new forced command on ep0 and a hub poller; ep0 is protected | + +## 3. Options + +**A. Keep the report clock.** Nothing to build. A teardown waits ≤45 min; alarms stay right. + +**B. The hub records wait-channel presence; the host page shows it; the delete guard may use it.** In `handleWait` +(once per request — not in `intent.Hub.Wait`, which runs once per 25-s window, `internal/intent/hub.go:77-102`) record +per customer: last wait start, holds open now. Presence = **connected** (a hold is open, or one started < 243 s + 90 s +grace ago), **not connected since T**, or **unknown** (hub restarted < 333 s ago — in memory only, so it falls back to +the report clock; the safe direction). Alerting is not changed. Cost: hub only, ~½ session. Wrong case: a crashed +controller on a live host reads „not connected"; if the guard trusts it, the operator deletes a live host and its agent +is locked out until re-enrolled — no household data is touched. + +**C. B + the WireGuard handshake from ep0 as host-level presence.** The best signal (the host itself, a third channel). +Cost: an ep0 change, a new key and forced command, a hub poller; ep0 work needs an operator task. + +**Pick: B, in two slices.** Slice 1 is display only and needs no decision. Slice 2 (the guard) needs the operator's word. +C only if slice 1's measurement shows the wait channel does not go quiet when a box dies. + +## 4. First slice and its red test + +- **Hub:** `intent.Hub` gains `MarkWaitStart / MarkWaitEnd / Presence(customerID, now)` behind a clock seam; + `handleWait` calls them; the host page shows „Box connection: connected now / last connected