Decision sheet D1-D10: hub CHANGELOG, architecture notes (01, 04, 05, 07, 08, 10), register (VERIFY states, R-912), STATUS, report; decoy suite runs in a git worktree
gates / gates (push) Successful in 4m32s
gates / gates (push) Successful in 4m32s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -213,6 +213,14 @@ app. Controller down → the gated app answers an error, never the app. Measured
|
||||
on Cloudflare) is register row R-494, P3, not blocking. Measured reason this was ruled now: the
|
||||
2026-09-14 first-hour drill used a `*.felhom.eu` customer domain with no tunnel, and the dashboard link
|
||||
in the setup-code mail did not resolve.
|
||||
- **A pasted Cloudflare API token must reach exactly the customer's own zone** (R-138 option C, `09` §3 decision 190;
|
||||
hub main 2026-10-08, unreleased). On save the hub asks Cloudflare `GET /zones` with the token and keeps it only when
|
||||
the token sees exactly ONE zone and that zone IS the customer's domain (a zone above it is refused: a sibling
|
||||
customer added later would be reachable). Cloudflare unreachable → the save fails and the old token stays. The token
|
||||
is never logged. **Limit:** `/zones` lists what the token can READ; a hand-built token with Zone:Read on one zone and
|
||||
DNS:Edit on all zones would pass — the token wizard's single „Specific zone" scope does not build that, and a
|
||||
customer token cannot read its own policies. Tokens saved before this check are checked the first time each config
|
||||
is saved with a new token or domain.
|
||||
- **Tunnel placement: INSIDE the guest** (corrected 2026-10-01, R-754 — the operator's brief of that evening: the build is
|
||||
right, correct the document). `cloudflared` is a container the CONTROLLER renders and keeps up (`internal/infra`,
|
||||
`EnsureBaseStack`, a protected stack), with the tunnel token from `controller.yaml`. *This page used to say it ran on
|
||||
|
||||
@@ -162,6 +162,16 @@ op the agent verifies** (same pipeline, §2.3) — never unauthenticated config.
|
||||
queue — then the agent polls, verifies, executes, and audits. One command + passphrase, from the
|
||||
desk. **Never** a site visit.
|
||||
|
||||
### 6.1 Operator actions in the report reply [DESIGN, R-314/R-279 — `09` §3 decision 185, controller + hub main 2026-10-08, unreleased]
|
||||
|
||||
Routine, non-destructive operator requests need no signature (`03` §4 asks for one only to destroy or overwrite the
|
||||
only copy). The hub stores a row, bumps the box's intent, and the report reply carries `operator_actions:[{id,action,arg}]`
|
||||
until the next report answers `operator_action_results:[{id,outcome,message}]`. The list is CLOSED on both sides:
|
||||
`offsite_backup_now`, `abandon_stop`, `abandon_extend` (1–30 days; never earlier than the current date; refused once
|
||||
the deletion is the hub's), `run_job` (`fill-watch`, `offsite-integrity`, `offsite-proof`, `disk-health-check`).
|
||||
Anything else is refused at the hub's POST and again on the box. No action deletes data, starts a countdown or
|
||||
shortens one (`TestOpActions_ClosedList`, `TestOperatorActions_ClosedList`). Unanswered after 24 h: expired.
|
||||
|
||||
## 7. Hardware readiness (Viktor's "build the foundation now")
|
||||
|
||||
Software `ssh-ed25519` now; a FIDO2 `sk-ssh-ed25519@openssh.com` key later is a **no-op on the
|
||||
|
||||
@@ -95,6 +95,16 @@ Evolves the existing staleness checker (60s **cadence**, a **configured** thresh
|
||||
than waiting for a guest report to go stale.
|
||||
- **Guest-report recency = secondary** app-level signal.
|
||||
|
||||
**Box presence from the wait channel [DESIGN, R-30 — `09` §3 decision 186, hub main 2026-10-08, unreleased].** The
|
||||
hub records, in memory, each controller wait (`GET /api/v1/wait`) per customer: **connected** (a hold is open, or one
|
||||
started < 333 s ago = 243 s cadence + 90 s grace), **not connected since T**, or **unknown** (the hub started < 333 s
|
||||
ago). The host page shows it. Alerting is NOT changed. **„Delete host" goes ahead at once** — instead of waiting for
|
||||
the report clock — only when the host is online by its report, the operator ticked „I checked: the box is off", AND
|
||||
presence has been „not connected" for ≥ 360 s, re-checked at the POST; unknown never permits it. The delete writes an
|
||||
INFO line and one `host_deleted_box_off` event. Wrong case: a box whose controller crashed while its host runs — its
|
||||
agent is locked out until re-enrolled; no household data is touched. `hub/internal/intent/presence.go`,
|
||||
`hub/internal/web/r30_presence_delete_test.go`.
|
||||
|
||||
**Backup-deadline checker:** today it is *event-based* — it scans for `backup_completed`/`backup_failed`
|
||||
events since local midnight and alerts if none. Two changes: (1) **mechanism** — move it to a field
|
||||
check on `host_reports`' last-backup-per-target (cleaner now that backup state arrives in the host
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -387,6 +387,19 @@ Design home: `07` §6.1.1 (`09` decisions 109–110; CC decisions 115–116, *op
|
||||
|
||||
---
|
||||
|
||||
## 6.5 One lost off-site copy on a pinned tier [DESIGN, R-435 — `09` §3 decision 191, hub main 2026-10-08, unreleased]
|
||||
|
||||
On a **pinned** off-site tier (the hub holds a confirmed append-only key, decision 69) the only legitimate way the
|
||||
snapshot count can fall is a clean-up window the hub opened (decision 68). So the hub raises
|
||||
`offsite_snapshots_dropped` (**error**) when the count falls by even ONE more than those windows explain since the
|
||||
previous trustworthy report. Each window explains at most its own hub-set `max_remove` — a box's `count_after` cannot
|
||||
widen it, and a window closed by timeout or still open explains exactly that cap; each window explains one fall only;
|
||||
a window stuck open past its deadline explains nothing; when the windows cannot be read, the half-rule decides (fail
|
||||
closed). Non-pinned (NAS) tiers keep the more-than-half rule. **Limits:** the count comes from the box (R-895's
|
||||
caveat holds); it is a NET count — new snapshots between two reports hide the same number of deletions; the „window
|
||||
spent" memory is in-process, so after a hub restart one window can explain one more fall. Pinned by
|
||||
`hub/internal/monitor/r435_pinned_drop_test.go` and `hub/internal/store/r435_windows_between_test.go`.
|
||||
|
||||
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
|
||||
|
||||
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
|
||||
|
||||
@@ -376,6 +376,13 @@ harmless (`TestFlashKeyRoundTrip`). The same shape one layer in: an **alert bann
|
||||
background health cycle and read minutes later, so `Alert` carries `MessageKey` + `MessageArgs` and
|
||||
`GetAlerts(lang)` renders on the way out.
|
||||
|
||||
**[FACT] 2026-10-08 (R-79, `09` §3 decision 187; controller main, unreleased): every health issue and warning now
|
||||
carries its key** — the last six producers (Docker unreachable, protected container down, storage unavailable / not
|
||||
separate / usage high / almost full) joined the seven resource ones; the wire text is unchanged byte for byte. The
|
||||
Docker error inside its sentence stays English (Docker writes it). **The household's health mail no longer carries the
|
||||
raw details note** (hub main, unreleased): it says what the headline says and points to the dashboard; the operator's
|
||||
mail keeps the note (`hub/internal/notify/r79_health_mail_test.go`).
|
||||
|
||||
**[DESIGN] Word order is Go's explicit argument index, not a second placeholder syntax.** The plan
|
||||
proposed a named-parameter (`{{.Name}}`) form for multi-parameter Go messages. English reorders with
|
||||
`%[2]s`, which `fmt` already understands, so the Hungarian value stays **the format string the code
|
||||
|
||||
Reference in New Issue
Block a user