Decision sheet D1-D10: hub CHANGELOG, architecture notes (01, 04, 05, 07, 08, 10), register (VERIFY states, R-912), STATUS, report; decoy suite runs in a git worktree
gates / gates (push) Successful in 4m32s
gates / gates (push) Successful in 4m32s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -16,6 +16,14 @@
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
|
||||
|
||||
> **2026-10-08 (evening) — the decision sheet D1–D10 built on main, unreleased (`09` §3 185–194).** Controller (`3f84f82`):
|
||||
> operator actions (`internal/report/opactions.go`, closed list), dashboard sessions on disk as sha256 + password
|
||||
> fingerprint (`internal/web/session_store.go`), six health keys, the R-893 hold (`restore_mixed`) on every failure after
|
||||
> a definition/volume moved. Hub: `operator_actions` table + reply field, wait-channel presence + the ticked fast delete
|
||||
> (`intent/presence.go`), health mail without the raw note, Cloudflare `/zones` reach check (zone == domain), R-435
|
||||
> window credits (capped, spent once). Catalog `31e9651` (wger 256M). Ship hub + controller the same day; live proofs
|
||||
> listed in `REPORT-day3-2026-10-08.md`. Register 132 → 136 (+4 another session's).
|
||||
|
||||
> **2026-10-08 (morning) — the first real kernel night read back.** demo-felhom: staged exactly the told 7.0.14-20 at
|
||||
> 04:39:04, healthy 04:40:38, default now 7.0.14-20, apps ~1.5 min away, no alarm. demo-hp: no whole-guest backup that
|
||||
> night (yesterday's 08:49 press + 24 h cadence > window end) → no step; R-899 filed. Reply-To proven by the operator's
|
||||
|
||||
@@ -0,0 +1,88 @@
|
||||
# REPORT — 2026-10-08 (evening): the operator's ten answers (D1–D10) built — code and tests, no delivery
|
||||
|
||||
(`REPORT-day3-…` because other sessions write in this clone today.) Rulings recorded first: `09` §3 decisions 185–194
|
||||
(commit `4510dd7d`), and in each row.
|
||||
|
||||
| Part | Result |
|
||||
|---|---|
|
||||
| **A** — operator actions (D1; R-314, R-279) | **Done on main.** Controller `internal/report/opactions.go` + `scheduler.RunNow`; hub table `operator_actions`, four buttons, reply field `operator_actions`, result field `operator_action_results`. Closed list both sides (`TestOpActions_ClosedList`, `TestOperatorActions_ClosedList`); unknown → `refused`, nothing called; once per id; cross-customer result ignored. Fixed on the way: `ExtendAbandon` could move a deletion earlier or edit the box date in the hub phase; `StopAbandon` reported success while the hub's deletion stayed pending; reply read limit 4 → 64 KiB. Review fix: a `run_job` „done" says it ran, not what it found. |
|
||||
| **B** — delete a switched-off box at once (D2; R-30) | **Done on hub main.** Presence from the wait channel (connected / not connected since T / unknown), shown on the host page; the delete goes ahead at once only with the tick AND ≥ 360 s not connected, re-checked at the POST. 7 mutations red-proved. |
|
||||
| **C** — health mail + dashboard (D3; R-79) | **Done on main, both halves.** Hub: the household's health mail drops the raw note and points to the dashboard (operator mail unchanged). Controller: the last six health producers carry a key; wire text unchanged. |
|
||||
| **D** — signed in across a restart (D4; R-35) | **Done on controller main.** Only sha256(cookie) on disk (0600). Review fixes: rows written under another password are dropped at load; a failed revoking save removes the file. 8 tests. |
|
||||
| **E** — Cloudflare key check (D6; R-138) | **Done on hub main.** Fake Cloudflare in tests only; the token is never logged. Review fixes: the token's one zone must BE the customer's domain (a parent zone was accepted, and checking it against today's customers was order-dependent). Limit filed: R-913 (read scope ≠ write scope). |
|
||||
| **F** — one lost off-site copy is an error (D7; R-435) | **Done on hub main.** Review fixes: each window explains at most its hub-set cap; each window explains one fall only; a window stuck open explains nothing; unreadable windows fall back to the half-rule. `08` §6.5 written (count from the box, net count). |
|
||||
| **G** — failed restore keeps the app stopped (D8; R-893) | **Done on controller main.** Review fixes: the hold now covers every failure after the definition or a volume moved (not only a failed replay); the hold is persisted before the stop. R-893 stays open, NARROWED to the second half („put back exactly as it was") — no new row needed. |
|
||||
| **H** — D5, D10 | **R-717 closed** (D5). **wger `mem_request` 256M pushed** (catalog `31e9651`, CI 1559 success), after a hub read showed no box runs wger (app-telemetry page 12:22 UTC: `wger` 0; controls `paperless`, `bentopdf` present). |
|
||||
| R-554, R-462 | **Not done.** R-554 is still stopped on a customer promise (decision 2 below). R-462 not started (no time left after the review fixes). |
|
||||
|
||||
**Rows: 132 before → 136 after. Opened 1 (R-913). Closed 1 (R-717).** The other +4 are another session's (R-909, R-910,
|
||||
R-911, R-912: dashboard layout and logos). State changes: R-314, R-279, R-30, R-79, R-35, R-138, R-435 → VERIFY (built, ships tomorrow,
|
||||
closes after the live proof); R-893 → NARROWED.
|
||||
|
||||
**Commits.** Controller: `ee8a526`…`d3e17e9` (cherry-picked from the helper branches onto main), `3f84f82` (review
|
||||
fixes). Hub/felhom.eu: `4510dd7d` (rulings), then the integration push (this report's commit). Catalog: `31e9651`.
|
||||
|
||||
## Security review — what it found and what happened
|
||||
|
||||
Automatic commit reviews plus two read-only review helpers. Fixed and red-proved the same evening (10): Cloudflare
|
||||
parent zone; order-dependence of the zone check; R-435 fail-open on unreadable windows; a box's `count_after`
|
||||
widening a window; one window explaining several falls; a stale open window; D4 revocation lost on a failed save
|
||||
(two shapes); D8 earlier failure branches started the app on a mix; the hold written after the stop; `run_job`
|
||||
„done" read as a pass. **Written as limits, not fixed:** D2 — presence is the guest controller's, so a live host
|
||||
with a stopped guest reads „not connected" (the cost accepted with decision 186; noted on R-30); D8 — placed files
|
||||
alone still restart after a successful rollback (option C as ruled; option A is next); D6 — tokens saved before the
|
||||
check were never checked (R-138 closes only after each is checked once); D6 — read scope ≠ write scope (R-913).
|
||||
|
||||
## Waits for tomorrow's releases
|
||||
|
||||
- **hub:** the morning's items (R-243, R-901 one-year deletion, R-304 mail, R-415 guard) + D1 hub half, D2, D3 mail, D6, D7.
|
||||
- **controller:** the morning's R-899/R-304 + the afternoon's R-304 mail, R-298 button + another session's dashboard
|
||||
layout fixes + D1, D3 dashboard, D4, D8. MinAgent unchanged (0.131.0). **Release the hub and the controller the
|
||||
same day** (D1's wire fields; an older hub sends no actions, an older controller ignores them and the rows expire).
|
||||
- **agent:** the morning's R-899 + ring label + R-304 424 (nothing new tonight).
|
||||
- **catalog:** already pushed (wger 256M; reaches a box at its next sync).
|
||||
- **DooPlex:** the `ep0-copy` clean-up job (D9) — dry run, then daily — after the releases are read back.
|
||||
|
||||
## Live proofs tomorrow's session must run (scratch 9202 only, after the releases)
|
||||
|
||||
1. **D1:** from the hub's host page, `run_job fill-watch` → the hub's result row; control from another channel: the
|
||||
controller log pull's „checked N filesystem(s)" line. `offsite_backup_now` → result; control: the off-site snapshot
|
||||
list on 9202's backup page. `abandon_stop` / `abandon_extend` only if 9202 has a countdown.
|
||||
2. **D2:** stop guest 9202; time until the host page says „not connected"; control: the ingress log's last
|
||||
`/api/v1/wait` line from 9202. **Do NOT tick-delete demo-hp** — the fast delete removes the Proxmox host; prove it
|
||||
only through `GET /hosts/<id>/delete-impact` (`off_tick_required`). Evidence off before 9202 starts again.
|
||||
3. **D3:** a `health_critical` on 9202 → the household mail in the catch-all has no `{`, has the dashboard line; the
|
||||
operator mail still has the note. The dashboard banner in English for an English household.
|
||||
4. **D4:** sign in on 9202, restart the controller, the next page needs no login; then change the password → the old
|
||||
cookie is refused after the next restart.
|
||||
5. **D6:** re-save each customer's config once with its current token (operator present): each must pass; a refusal
|
||||
keeps the old token. Then close R-138.
|
||||
6. **D7:** a hand-granted window on 9202 removes N → no alarm; control: the hub's own window row.
|
||||
7. **D8:** design slice 0 on 9202 with a throwaway app (version N snapshot, update to N+1, restore with a truncated
|
||||
dump) → the app held, the page sentence, the operator line. Evidence off before teardown.
|
||||
8. **D9:** install the DooPlex job, dry run first, then daily.
|
||||
|
||||
## Machines
|
||||
|
||||
demo-hp, demo-felhom, Tester 1: not touched. 9202 and the bench: not touched. Hub: read only (three page reads with
|
||||
the operator password: `/apps`, `/hosts`, `/configs`). DooPlex: no change (helper work ran in git worktrees under
|
||||
`/mnt/5_hdd/felhom.eu/worktrees/`; no Docker command). No release, no deploy, no reboot, no prune. Provisioned nothing.
|
||||
|
||||
## Small fixes without a row
|
||||
|
||||
`scripts/test_gate_decoys.py` crashed in a linked git worktree (lock path); controller `TestR650_NoBareDockerExec`
|
||||
raced a vanishing temp file; the three D1 abandon fixes above.
|
||||
|
||||
Instruction files: no edit.
|
||||
|
||||
## Decisions for the operator
|
||||
|
||||
1. **R-913 — how sure must the Cloudflare key check be?** The hub can see what a pasted key may READ, not what it may
|
||||
CHANGE. A key made by hand with „read one zone, change all zones" would pass. **Pick: (a) you always make the key
|
||||
with the one recipe** („Edit zone DNS", Zone Resources: Include → Specific zone); it costs nothing. (b) The hub
|
||||
makes the key itself: safest, but the hub then holds a key that can make keys. *If you do nothing:* the check stays
|
||||
as built, and the limit stays written down.
|
||||
2. **R-554 — what should the household's recovery note say once the old setup wizard is deleted?** Today the note on
|
||||
the box tells the household to run a setup script on port 8081 to restore. **Pick: the note says „Call Felhom; we
|
||||
restore your box from its backup"**, and CC deletes the wizard. *If you do nothing:* the wizard stays, unused, and a
|
||||
box whose first start cannot reach the hub shows it to the household.
|
||||
@@ -71,6 +71,8 @@
|
||||
|
||||
| Symbol | File | Short signature | Use for | Gotchas |
|
||||
|---|---|---|---|---|
|
||||
| `store.WindowCreditsBetween` | hub/internal/store/offsite_keys.go | `(customerID, from, to) ([]WindowCredit, unknown)` | R-435: what the hub's clean-up windows may explain | Each window ≤ its `max_remove`; the CALLER spends each window once (`OffsiteChecker.usedWindows`); unknown → the half-rule, never „explained" |
|
||||
| `cloudflare.TokenZones` / `ZoneEquals` | hub/internal/cloudflare/reach.go | `(ctx, base, token) (names, total, err)` | R-138: what zones a pasted token can see | Every failure wraps `ErrReachUnknown` (refuse the save); the token never appears in an error; `/zones` is READ scope, not write |
|
||||
| `offsitekeys.Registrar` (`Install` / `Confirm` / `Audit` / `OpenWindow` / `CloseWindow` / `MoveAside`) | hub/internal/offsitekeys/offsitekeys.go | `(ctx, Target, password, …)` | EVERY write to a sub-account's `.ssh/authorized_keys` and every repo move-aside | **The only writer of that file, and the only deleter on a sub-account (`DeleteSetAside`: `<repo>.orphaned-*` only, decision 74).** `read()` is read-only (R-827). Uses the provider's port-23 restricted shell (`dd of=` takes stdin, `mv` overwrites, `test` does NOT exist — measured); an unpinned line is a deletion route and is dropped on every install; the window line goes FIRST (first match wins). Never `rm`. |
|
||||
| `offsitekeys.Service` (`RegisterKey`, `ConfirmKey`, `AuditAll`, `OpenWindowFor`, `CloseWindowFor`, `SweepExpiredWindows`) | hub/internal/offsitekeys/service.go | — | Binding the registrar to the store, descriptor and operator events | The box-facing API (`/api/v1/offsite/register-key…`) answers with NO credential — pinned by `TestOffsiteKeyEndpoints_AuthAndNoPasswordInAnyResponse`. |
|
||||
| `(*Store).SaveOneTimeSecret` / `OffsitePassword` / `SealLegacyOffsiteSecrets` | hub/internal/store/offsite_seal.go | — | Storing / reading the sub-account password | **Sealed AES-256-GCM; no key → refused (fail-closed).** Under `go test` every store gets a fixed key (`testing.Testing()`); production needs `OFFSITE_SECRET_KEY`. Never serve the value to a box. |
|
||||
@@ -80,6 +82,8 @@
|
||||
|
||||
| Symbol | File | Short signature | Use for | Gotchas |
|
||||
|---|---|---|---|---|
|
||||
| `store.CreateOperatorAction` / `PendingOperatorActions` / `RecordOperatorActionResult` / `ValidateOperatorAction` | hub/internal/store/opactions.go | rows of `operator_actions` | Operator actions carried in the report reply (decision 185); POST `/hosts/{id}/operator-action` → `handleOperatorAction` | `ValidateOperatorAction` IS the closed list (`TestOperatorActions_ClosedList`); a result is matched on (customer, id) — never on id alone |
|
||||
| `intent.Hub.MarkWaitStart` / `MarkWaitEnd` / `Presence` | hub/internal/intent/presence.go | `(customerID, now)` | Box presence from the wait channel (R-30); the delete-host fast path | In memory: a hub younger than 333 s answers `unknown`, which never permits the fast delete; `hostOffDeleteAfter` ≥ the window (pinned) |
|
||||
| `(*Server).hostDetailData` | hub/internal/web/hosts.go (~L282) | `(host *store.Host, r) map[string]interface{}` | The ONE view-model builder for the shared `host_detail_body` sub-template (standalone `/hosts/{id}` + customer Host tab) | Booleans/counts only for DR/escrow; carries `Deletable` (= status != "ok") which gates the danger-zone card. Never add a secret field. |
|
||||
| `parseHostAddresses` + `(*Server).hostNetwork` / `hostNetworkView` (v0.85.0) | hub/internal/web/hosts.go | `(reportJSON) []hostAddressView` · `(host, reportJSON) hostNetworkView` | The host page's Network card: every routable address the box holds + its WireGuard allocation | Needs agent **>= 0.119.0** (`minAgentForAddresses`); below it the wire has no `addresses` key and the card renders **UNKNOWN, never "no addresses"** — an absent signal is not a negative result. WireGuard is TWO facts: the hub's allocation (`GetWGPeerForHost`, authoritative) AND whether the box confirms holding it — the allocation alone cannot distinguish a live tunnel from a peer that was never applied. The WG row is split out by comparing against the ALLOCATION, never by matching the interface name `wg-felhom`, which is a unit name that can change. |
|
||||
| `(*Store).GetHostRecoveryMeta` + `(*Server).handleHostRevealRecoveryCredential` | hub/internal/store/host_recovery.go · hub/internal/web/hosts.go | `(hostID) (*HostRecoveryMeta, error)` · `POST /hosts/{id}/reveal-recovery-credential` | The break-glass console credential, split into a RENDER half and a RETRIEVE half (v0.84.0) | **Use `GetHostRecoveryMeta` on any page-render path** — its struct and its `SELECT` both omit the `secret` column, so it cannot leak one; `GetHostRecoveryCredential` (which does select it) belongs only to the two retrieval handlers. The reveal is POST so the ServeHTTP-level CSRF check applies and no secret is reachable by URL; it writes ONE `recovery_credential_revealed` event via `SaveEvent` and calls NO dispatcher (the `handleRequestLogTail` shape). `api/handler.go handleAdminGetRecoveryCredential` (global key) is the independent fallback for when the UI is down — never route the UI through it. Secret at rest is plaintext → R-133. |
|
||||
|
||||
@@ -2,9 +2,30 @@
|
||||
|
||||
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.**
|
||||
|
||||
**Updated 2026-10-08 (afternoon): hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller
|
||||
0.303.0 (nothing delivered today — tonight is the second kernel night). The open-items list is at 129. Reports:
|
||||
`REPORT-day-2026-10-08.md` (morning), `REPORT-day2-2026-10-08.md` (afternoon).**
|
||||
**Updated 2026-10-08 (evening): hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller
|
||||
0.303.0 (nothing delivered today — tonight is the second kernel night). The open-items list is at 136. Reports:
|
||||
`REPORT-day-2026-10-08.md` (morning), `REPORT-day2-2026-10-08.md` (afternoon), `REPORT-day3-2026-10-08.md` (evening).**
|
||||
|
||||
## Evening (2026-10-08): your ten answers (D1–D10) built — they ship tomorrow
|
||||
|
||||
- **Buttons in the hub** (D1): run an off-site backup now, run one of four checks now, stop or extend a household's
|
||||
deletion countdown. The list is fixed. No button deletes anything or shortens a countdown.
|
||||
- **A switched-off box can be deleted at once** (D2), after you tick „I checked: the box is off" and the hub has heard
|
||||
nothing from it for 6 minutes. The host page shows when the box was last connected.
|
||||
- **The household's „system health" mail is short** (D3): no technical note; it points to the dashboard. The dashboard
|
||||
shows every health warning in the household's language.
|
||||
- **The household stays signed in** when the box restarts or updates (D4). The disk keeps only a fingerprint.
|
||||
- **opengist's item is closed** (D5): the address block is enough.
|
||||
- **The hub checks a pasted Cloudflare key** (D6): it must reach exactly that customer's own domain.
|
||||
- **One lost off-site copy outside a clean-up window = an error mail to you** (D7).
|
||||
- **A failed restore of one app keeps the app stopped** (D8) when its version or its data volumes already moved, and
|
||||
the page says it needs our help. „Put back exactly as it was" is the next step (on the list).
|
||||
- **wger needs 256 MB to install** (D10). Live in the catalog; no box runs wger.
|
||||
- **D9 (the DooPlex clean-up job) is tomorrow**, after the releases are read back.
|
||||
- Security checks read every change. Ten problems were fixed and tested the same evening. Two stay as written limits
|
||||
(one is the cost you accepted with D2; files alone wait for D8's next step). One needs you (decision 1 below).
|
||||
|
||||
**Needs you:** two decisions at the end of `REPORT-day3-2026-10-08.md`.
|
||||
|
||||
## Afternoon (2026-10-08): your four answers built, and one sheet of decisions
|
||||
|
||||
@@ -16,7 +37,7 @@
|
||||
- **Cloudflare:** on the list as a later item.
|
||||
- **Nine stuck items have a one-page design.** Two were no longer true and are closed; one lost item is back on the list.
|
||||
|
||||
### The decision sheet — answer „all as picked", or name the numbers you change
|
||||
### The decision sheet — ANSWERED 2026-10-08 14:16, all as picked (`09` §3 decisions 185–194)
|
||||
|
||||
| # | Question | Pick | Cost | If you do nothing |
|
||||
|---|---|---|---|---|
|
||||
|
||||
@@ -213,6 +213,14 @@ app. Controller down → the gated app answers an error, never the app. Measured
|
||||
on Cloudflare) is register row R-494, P3, not blocking. Measured reason this was ruled now: the
|
||||
2026-09-14 first-hour drill used a `*.felhom.eu` customer domain with no tunnel, and the dashboard link
|
||||
in the setup-code mail did not resolve.
|
||||
- **A pasted Cloudflare API token must reach exactly the customer's own zone** (R-138 option C, `09` §3 decision 190;
|
||||
hub main 2026-10-08, unreleased). On save the hub asks Cloudflare `GET /zones` with the token and keeps it only when
|
||||
the token sees exactly ONE zone and that zone IS the customer's domain (a zone above it is refused: a sibling
|
||||
customer added later would be reachable). Cloudflare unreachable → the save fails and the old token stays. The token
|
||||
is never logged. **Limit:** `/zones` lists what the token can READ; a hand-built token with Zone:Read on one zone and
|
||||
DNS:Edit on all zones would pass — the token wizard's single „Specific zone" scope does not build that, and a
|
||||
customer token cannot read its own policies. Tokens saved before this check are checked the first time each config
|
||||
is saved with a new token or domain.
|
||||
- **Tunnel placement: INSIDE the guest** (corrected 2026-10-01, R-754 — the operator's brief of that evening: the build is
|
||||
right, correct the document). `cloudflared` is a container the CONTROLLER renders and keeps up (`internal/infra`,
|
||||
`EnsureBaseStack`, a protected stack), with the tunnel token from `controller.yaml`. *This page used to say it ran on
|
||||
|
||||
@@ -162,6 +162,16 @@ op the agent verifies** (same pipeline, §2.3) — never unauthenticated config.
|
||||
queue — then the agent polls, verifies, executes, and audits. One command + passphrase, from the
|
||||
desk. **Never** a site visit.
|
||||
|
||||
### 6.1 Operator actions in the report reply [DESIGN, R-314/R-279 — `09` §3 decision 185, controller + hub main 2026-10-08, unreleased]
|
||||
|
||||
Routine, non-destructive operator requests need no signature (`03` §4 asks for one only to destroy or overwrite the
|
||||
only copy). The hub stores a row, bumps the box's intent, and the report reply carries `operator_actions:[{id,action,arg}]`
|
||||
until the next report answers `operator_action_results:[{id,outcome,message}]`. The list is CLOSED on both sides:
|
||||
`offsite_backup_now`, `abandon_stop`, `abandon_extend` (1–30 days; never earlier than the current date; refused once
|
||||
the deletion is the hub's), `run_job` (`fill-watch`, `offsite-integrity`, `offsite-proof`, `disk-health-check`).
|
||||
Anything else is refused at the hub's POST and again on the box. No action deletes data, starts a countdown or
|
||||
shortens one (`TestOpActions_ClosedList`, `TestOperatorActions_ClosedList`). Unanswered after 24 h: expired.
|
||||
|
||||
## 7. Hardware readiness (Viktor's "build the foundation now")
|
||||
|
||||
Software `ssh-ed25519` now; a FIDO2 `sk-ssh-ed25519@openssh.com` key later is a **no-op on the
|
||||
|
||||
@@ -95,6 +95,16 @@ Evolves the existing staleness checker (60s **cadence**, a **configured** thresh
|
||||
than waiting for a guest report to go stale.
|
||||
- **Guest-report recency = secondary** app-level signal.
|
||||
|
||||
**Box presence from the wait channel [DESIGN, R-30 — `09` §3 decision 186, hub main 2026-10-08, unreleased].** The
|
||||
hub records, in memory, each controller wait (`GET /api/v1/wait`) per customer: **connected** (a hold is open, or one
|
||||
started < 333 s ago = 243 s cadence + 90 s grace), **not connected since T**, or **unknown** (the hub started < 333 s
|
||||
ago). The host page shows it. Alerting is NOT changed. **„Delete host" goes ahead at once** — instead of waiting for
|
||||
the report clock — only when the host is online by its report, the operator ticked „I checked: the box is off", AND
|
||||
presence has been „not connected" for ≥ 360 s, re-checked at the POST; unknown never permits it. The delete writes an
|
||||
INFO line and one `host_deleted_box_off` event. Wrong case: a box whose controller crashed while its host runs — its
|
||||
agent is locked out until re-enrolled; no household data is touched. `hub/internal/intent/presence.go`,
|
||||
`hub/internal/web/r30_presence_delete_test.go`.
|
||||
|
||||
**Backup-deadline checker:** today it is *event-based* — it scans for `backup_completed`/`backup_failed`
|
||||
events since local midnight and alerts if none. Two changes: (1) **mechanism** — move it to a field
|
||||
check on `host_reports`' last-backup-per-target (cleaner now that backup state arrives in the host
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -387,6 +387,19 @@ Design home: `07` §6.1.1 (`09` decisions 109–110; CC decisions 115–116, *op
|
||||
|
||||
---
|
||||
|
||||
## 6.5 One lost off-site copy on a pinned tier [DESIGN, R-435 — `09` §3 decision 191, hub main 2026-10-08, unreleased]
|
||||
|
||||
On a **pinned** off-site tier (the hub holds a confirmed append-only key, decision 69) the only legitimate way the
|
||||
snapshot count can fall is a clean-up window the hub opened (decision 68). So the hub raises
|
||||
`offsite_snapshots_dropped` (**error**) when the count falls by even ONE more than those windows explain since the
|
||||
previous trustworthy report. Each window explains at most its own hub-set `max_remove` — a box's `count_after` cannot
|
||||
widen it, and a window closed by timeout or still open explains exactly that cap; each window explains one fall only;
|
||||
a window stuck open past its deadline explains nothing; when the windows cannot be read, the half-rule decides (fail
|
||||
closed). Non-pinned (NAS) tiers keep the more-than-half rule. **Limits:** the count comes from the box (R-895's
|
||||
caveat holds); it is a NET count — new snapshots between two reports hide the same number of deletions; the „window
|
||||
spent" memory is in-process, so after a hub restart one window can explain one more fall. Pinned by
|
||||
`hub/internal/monitor/r435_pinned_drop_test.go` and `hub/internal/store/r435_windows_between_test.go`.
|
||||
|
||||
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
|
||||
|
||||
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
|
||||
|
||||
@@ -376,6 +376,13 @@ harmless (`TestFlashKeyRoundTrip`). The same shape one layer in: an **alert bann
|
||||
background health cycle and read minutes later, so `Alert` carries `MessageKey` + `MessageArgs` and
|
||||
`GetAlerts(lang)` renders on the way out.
|
||||
|
||||
**[FACT] 2026-10-08 (R-79, `09` §3 decision 187; controller main, unreleased): every health issue and warning now
|
||||
carries its key** — the last six producers (Docker unreachable, protected container down, storage unavailable / not
|
||||
separate / usage high / almost full) joined the seven resource ones; the wire text is unchanged byte for byte. The
|
||||
Docker error inside its sentence stays English (Docker writes it). **The household's health mail no longer carries the
|
||||
raw details note** (hub main, unreleased): it says what the headline says and points to the dashboard; the operator's
|
||||
mail keeps the note (`hub/internal/notify/r79_health_mail_test.go`).
|
||||
|
||||
**[DESIGN] Word order is Go's explicit argument index, not a second placeholder syntax.** The plan
|
||||
proposed a named-parameter (`{{.Name}}`) form for multi-parameter Go messages. English reorders with
|
||||
`%[2]s`, which `fmt` already understands, so the Hungarian value stays **the format string the code
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -1,3 +1,37 @@
|
||||
## Unreleased (2026-10-08, evening) — the operator's decision sheet: operator actions (D1), box presence + delete a switched-off box at once (D2), the household's health mail points to the dashboard (D3), the Cloudflare token reach check (D6), one lost off-site snapshot is an error (D7) — ships with tomorrow's hub release
|
||||
|
||||
Operator rulings 2026-10-08 14:16, `09` §3 decisions 185–194. **Operator action on deploy: none.** The controller of the
|
||||
same day carries D1's other half; an older controller ignores `operator_actions` (rows then expire after 24 h).
|
||||
|
||||
- **D1 — operator actions (R-314, R-279).** Table `operator_actions(id, customer_id, action, arg, requested_at,
|
||||
requested_by, done_at, outcome, message)`; four buttons on the host page (POST `/hosts/{id}/operator-action`,
|
||||
operator auth + CSRF); the row is carried in the report reply (`operator_actions`) until the box answers
|
||||
(`operator_action_results`); a result naming another customer's id is ignored; unanswered after 24 h → `expired`; a
|
||||
customer RESET → `cancelled`. CLOSED list, validated at the POST (400, no row): `offsite_backup_now`, `abandon_stop`,
|
||||
`abandon_extend` (1–30), `run_job` (`fill-watch`, `offsite-integrity`, `offsite-proof`, `disk-health-check`)
|
||||
(`TestOperatorActions_ClosedList`). Who pressed = the channel and address (one password, no user names), logged at
|
||||
INFO at the press and at the result; each result is a stored `operator_action` event (never a customer mail). New
|
||||
wire-gate root for the reply field.
|
||||
- **D2 — box presence from the wait channel (R-30, slices 1 and 2).** `intent.Hub` records wait starts/ends per
|
||||
customer: connected / not connected since T / unknown (hub younger than 333 s). The host page shows „Box connection".
|
||||
„Delete host" goes ahead at once when the host is online by report, the operator ticked „I checked: the box is off",
|
||||
and presence has been „not connected" ≥ 360 s — re-checked at the POST; otherwise 409 as before. INFO audit line + one
|
||||
`host_deleted_box_off` event (warning, hub-minted). `TestPresence_*`, `TestR30_*` (red-proved, 7 mutations).
|
||||
- **D3 — the household's health mail (R-79 option B).** `health_degraded` / `health_critical` / `health_recovered`
|
||||
customer mails drop the raw details note and end with „A részleteket a vezérlőpultodon látod." / „You can see the
|
||||
details on your dashboard." (`mail.customer.line.health_dashboard`). The operator's mail keeps the note; other event
|
||||
types unchanged. 6 mail goldens regenerated. `TestR79_*`.
|
||||
- **D6 — the Cloudflare token reach check (R-138 option C).** On create/edit, a new token (or the same token on a new
|
||||
domain) is checked with Cloudflare `GET /zones`: saved only when it sees exactly one zone and that zone IS the
|
||||
customer's domain. Cloudflare down → refused, the old token stays. Never logged. **Security review fixes (same day):**
|
||||
a parent zone that could hold a sibling customer is refused, order-independently (strict equality). Tests with a fake
|
||||
Cloudflare, `TestR138_*`, `TestZoneEquals`, `TestTokenZones_*`. Limit written in `01` §7 (listing = read scope).
|
||||
- **D7 — one lost off-site snapshot is an error (R-435 option C).** Pinned tiers: any fall the hub's own clean-up
|
||||
windows do not explain raises `offsite_snapshots_dropped` (error, „outside any clean-up window the hub opened").
|
||||
**Security review fixes (same day):** each window explains at most its hub-set `max_remove` (a box's `count_after`
|
||||
cannot widen it); each window explains one fall only; a window stuck open past its deadline explains nothing;
|
||||
unreadable windows fall back to the half-rule. NAS tiers keep the half-rule. `TestR435_*` (store + monitor). `08` §6.5.
|
||||
|
||||
## Unreleased (2026-10-08) — an alarm when a box never backs up off-site because its escrow is pending (R-243; `09` §3 decision 179); a deleted customer's audit rows go after 1 year (R-901; decision 181); the operator's older-recovery-package mail (R-304; decision 183); no duplicate or nested customer domain (R-415, R-138 option B) — ships with tomorrow's hub release
|
||||
|
||||
**Operator action on deploy: none.** Expect ONE `offsite_escrow_pending` mail for **Tester 2** on the first sweep after
|
||||
|
||||
@@ -1,3 +1,10 @@
|
||||
## gates — the decoy suite runs in a linked git worktree (2026-10-08, fixed without a row)
|
||||
|
||||
- `test_gate_decoys.py` took its lock at `ROOT/.git/decoy-suite.lock`; in a `git worktree add` copy `.git` is a FILE, so
|
||||
the suite crashed (`NotADirectoryError`) and `repo_gates.py` went red for every helper working in a worktree. The lock
|
||||
now lives in `git rev-parse --absolute-git-dir` (per worktree — each worktree mutates its own files), falling back to
|
||||
`ROOT/.git`. Measured: the full `repo_gates.py --fast` run green inside `worktrees/hub-integ`.
|
||||
|
||||
## gates — site gate 20: the one picture viewer (2026-10-08)
|
||||
|
||||
- Every `data-gallery` opener must be an `<a href>` to an existing file under `website/assets/`; a page with openers must
|
||||
|
||||
@@ -39,7 +39,15 @@ ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
# already mutated, and B's "restore" then wrote A's decoy back for good: website/en/apps.html was left without its
|
||||
# Radicale logo, one step from a push. An exclusive lock makes a second run WAIT for the first.
|
||||
import fcntl
|
||||
_LOCK = open(os.path.join(ROOT, ".git", "decoy-suite.lock"), "w")
|
||||
# The lock lives in THIS worktree's git dir: in a linked worktree (`git worktree add`) `.git` is a FILE, and the old
|
||||
# `ROOT/.git/decoy-suite.lock` crashed with NotADirectoryError (2026-10-08). Each worktree mutates its own files, so a
|
||||
# per-worktree lock is the right scope.
|
||||
try:
|
||||
_GITDIR = subprocess.run(["git", "rev-parse", "--absolute-git-dir"], cwd=ROOT, capture_output=True, text=True,
|
||||
check=True).stdout.strip()
|
||||
except (OSError, subprocess.CalledProcessError):
|
||||
_GITDIR = os.path.join(ROOT, ".git")
|
||||
_LOCK = open(os.path.join(_GITDIR, "decoy-suite.lock"), "w")
|
||||
fcntl.flock(_LOCK, fcntl.LOCK_EX)
|
||||
|
||||
# ── WHAT THIS FILE COVERS ────────────────────────────────────────────────────────────────────────
|
||||
|
||||
Reference in New Issue
Block a user