Part F: nine one-page designs (R-314/279/177, R-30, R-79, R-35, R-717, R-138, R-435, R-893); R-177 and R-298 closed with evidence; R-415 re-filed (lost in the 2026-10-03 triage); STATUS decision sheet D1-D9; 131 -> 130
gates / gates (push) Successful in 3m57s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-08 10:12:46 +02:00
parent 5beedcce1a
commit 78121aa475
11 changed files with 505 additions and 14 deletions
+29 -3
View File
@@ -2,9 +2,35 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.**
**Updated 2026-10-08 (day): hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller 0.303.0
(nothing delivered today — tonight is the second kernel night). The open-items list is at 130. Reports: `REPORT.md`,
`REPORT-day-2026-10-08.md`.**
**Updated 2026-10-08 (afternoon): hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller
0.303.0 (nothing delivered today — tonight is the second kernel night). The open-items list is at 130. Reports:
`REPORT-day-2026-10-08.md` (morning), `REPORT-day2-2026-10-08.md` (afternoon).**
## Afternoon (2026-10-08): your four answers built, and one sheet of decisions
- **The FAQ is honest now** (you approved the text): it says where backup copies and remote traffic go. Live.
- **Deletion times:** the hub deletes a removed customer's mail and event records 1 year after the deletion (ships
tomorrow). The job that removes a removed customer's copy on DooPlex within 30 days is written and tested, **not
installed** (see D9). Both times are in the privacy-notice draft.
- **You get a mail when a household's code opens, or may open, an old package** (ships tomorrow).
- **Cloudflare:** on the list as a later item.
- **Nine stuck items have a one-page design.** Two were no longer true and are closed; one lost item is back on the list.
### The decision sheet — answer „all as picked", or name the numbers you change
| # | Question | Pick | Cost | If you do nothing |
|---|---|---|---|---|
| D1 | May the hub (with your password, no signing key) ask a box to run an off-site backup now, run a named check now, and stop or extend a household's deletion countdown? | Yes — a fixed list of safe, repeatable actions, carried in the box's report reply | ~1 session, controller + hub | A household calling in the first 14 days of a countdown needs a shell on its box; only the household can start an off-site run |
| D2 | When the hub has had no connection from a box for ~6 minutes, may „delete host" go ahead at once after you tick „I checked: the box is off"? | Yes (step 1, showing the live link on the host page, needs no answer) | ~½ session, hub | Deleting a switched-off box (and a reset) waits up to 45 minutes after its last report |
| D3 | May the household's „system health" mail drop the raw technical note and point to the dashboard for the list? | Yes (the dashboard half needs no answer; CC builds it) | ~1 h hub + a hub release | The mail keeps a curly-bracket note with English lines in it |
| D4 | May the box remember a dashboard sign-in across its own restarts (on disk only a fingerprint that cannot be used to sign in)? | Yes | ~½ session, controller | The household is logged out at every settings push and every controller update |
| D5 | Is the measured address block enough for opengist's sign-up, or should the box also close opengist's own switch? | Enough — close the item | Nothing | Same as the pick: the item closes |
| D6 | Should the hub check with Cloudflare that a pasted key reaches only that customer's own domain? | Yes (the „no duplicate or nested domain" guard needs no answer; CC builds it) | ~½ session, hub; a save fails while Cloudflare is down | A key made for the whole account by mistake could change every household's web addresses |
| D7 | Should the hub mail you an error when even one off-site snapshot disappears outside a clean-up window it opened? | Yes | ~½ session, hub | A deletion through any other key stays silent unless it removes over half of a household's history |
| D8 | After a failed off-site restore of one app: keep the app stopped for support now, and „put back exactly as it was" next? | Yes, both, in that order | ~½ session now; ~1 session + a test + disk space for one copy later | The app restarts on a mix of old files and a newer database, and the screen says all is back as it was |
| D9 | May CC install the DooPlex job that removes a deleted customer's copy (dry run first, then daily)? | Yes, after tomorrow's releases are read back | ~30 min; one new local token on DooPlex | A deleted customer's copy stays on DooPlex, and the privacy-notice line „within 30 days" is not true |
Full designs: `documentation/audits/day-2026-10-08/`.
## Day (2026-10-08): fixes built for tomorrow, the old-code answer made honest, legal drafts
@@ -0,0 +1,52 @@
# R-138 — a Cloudflare token that can write a shared zone, on a customer's box: a one-page design (2026-10-08)
Baselines read: felhom.eu `b2dce901`, felhom-controller `a0370b4`. Architecture: `01-topology-and-trust.md` §5
(trust boundaries: „hub ↔ Cloudflare API") and §7 (networking: „Every customer has their OWN domain — never a name
under `felhom.eu`", operator ruling 2026-09-14). **Status:** design only, nothing built. The token itself is stored
out-of-band; it was not read, printed or used for this page.
## 1. The problem, and what changed under it
- **The mechanism is still as the row says.** The hub form copies `cf_api_token` verbatim into the customer's config
(`hub/internal/web/configs.go:1682-1684`); the controller writes it to a 0600 `.env` for Traefik
(`controller/internal/infra/infra.go:157-159`; the row's `:123` is stale). An empty token selects HTTP-01
(`controller/internal/infra/templates/traefik.yml.tmpl:52-61`). The controller also uses it for the geo-WAF
(`controller/internal/web/handlers.go:642`).
- **The row's risk needs a SHARED customer zone, and the design has ruled that out.** `01` §7 (2026-09-14): every
customer has their own domain, never a name under `felhom.eu`. Today's boxes match: demo-felhom, demo-hp and
Tester 1 (`enkicsifelhom.hu`, `operations/nodes.md:44`) each sit on their own zone. The shared-zone plan the row was
written against (`audits/RECON-subdomain-onboarding-2026-07-31.md` §5) is not the plan any more.
- **But nothing ENFORCES the ruling.** The hub accepts any domain, including one equal to or under another customer's
domain, or under `felhom.eu`: `domain TEXT NOT NULL DEFAULT ''` with no uniqueness (`hub/internal/store/store.go:164`),
and the create path refuses only a duplicate customer id (`configs.go:740`). That check was row **R-415** (READY, XS) —
**it was removed from `OPEN-ITEMS.md` in the 2026-10-03 triage (`71b8c8c6`) and never reached `CLOSED-ITEMS.md`**;
it is lost, not closed.
- **A second residue the row does not name:** the four zones sit in ONE Cloudflare account (RECON §2, line 163). A
token minted with account scope (all zones) instead of one zone would let one box rewrite every household's DNS —
the same blast radius as a shared zone, by a typing mistake in the token wizard. Nobody checks the token's reach.
## 2. Options
| | What | Costs | Risk |
|---|---|---|---|
| **A** | Close R-138 on the ruling; do nothing more. | Nothing. | The ruling is a sentence; an operator mistake (a duplicate or nested domain, an account-wide token) passes silently. |
| **B** | **Enforce the ruling on save (revive R-415):** the hub refuses a domain that equals, contains or sits under another customer's domain, or sits under `felhom.eu`. | Hub only; one store query, one form error, tests. ~¼ session. | None to customers: it refuses an operator input that the ruling already forbids. |
| **C** | B **plus a reach check on the token:** when a `cf_api_token` is saved, the hub asks Cloudflare which zones that token can see (`GET /zones`, the call the hub already makes in `hub/internal/cloudflare/unblock.go:115`) and refuses unless it sees exactly the customer's own zone. | Hub only; one outbound call at save time (an existing dependency, not a new one); a save fails while Cloudflare is down — the form must say so and keep what was typed. ~½ session. | A save blocked by a Cloudflare outage; the operator retries. |
## 3. The pick — B now (follows the ruling, no decision needed); C on the operator's word
B is the guard the row asked for, re-scoped: not „refuse a token for a shared-zone customer" (no such customer may
exist) but „refuse the layout that would make one". C closes the account-wide-token hole, which is the same danger by
another road; it adds a network call to a form save, so it is the operator's choice.
## 4. First slice and its red test
- **Hub (B):** in the create and edit paths, before any provisioning: `store.DomainConflicts(customerID, domain)` and a
`felhom.eu` suffix check; the form re-renders with the submitted values and one sentence.
- **Red test (fails on today's code):** create customer `a` with `example.hu`; creating `b` with `example.hu`,
`x.example.hu` or `t1.felhom.eu` is refused and nothing is stored. **Controls:** `b` with `example2.hu` is accepted;
editing `a` with its own `example.hu` is accepted (no self-conflict); `notexample.hu` is accepted (suffix match is on a
label boundary).
- Register: R-415 re-filed (or folded into this row) and R-138 closed on B's commit.
## 5. One question for the operator
**When you paste a customer's Cloudflare key into the hub, should the hub check with Cloudflare that the key reaches
only that customer's own domain, and refuse it otherwise?** My pick: yes (option C). *If you do nothing:* a key that was
made for the whole Cloudflare account by mistake goes onto the customer's box, and that box could change the web
addresses of every other household.
@@ -0,0 +1,72 @@
# R-30 — box presence from the live connection, not the report clock: a one-page design (2026-10-08)
Baselines read: felhom.eu `b2dce901`, felhom-controller `a0370b4`, felhom-agent `b228b44`. Architecture:
`05-hub-architecture.md` §4 (liveness / dead-man's-switch); `03-host-agent.md` §4 (the host-delete guard's reason).
**Status: design only. Nothing is built.**
## 1. The problem, with today's numbers
- **Measured 2026-07-21 (the row):** a powered-off box stayed „healthy" on the hub until `host_stale` fired after the
threshold („no report for 30m"). **Today the wait is longer:** the threshold is 45 minutes (live ConfigMap
`felhom-system/hub-config`, `alerting.stale_threshold: "45m"`, operator ruling A on R-549). The row's „30 min" is stale.
- The agent reports every 900 s (hub `internal/api/handler.go:653-655`, `defaultHostPollSeconds = 900`; agent
`internal/config/config.go:849`). „Online" on the host page is report age under the threshold
(`internal/web/hosts.go:31-44`). **Host delete refuses while „online"** (`hosts.go:935-937`, and since R-599 it says when
the refusal ends). **RESET refuses while any host row exists** (`internal/web/customer_reset.go:106-110`). So a forced
teardown of a box that is already off waits up to 45 minutes after its last report.
- **Measured today (read-only, the ingress access log on DooPlex, 07:0x–07:50 UTC):** the controller's wait channel
(`GET /api/v1/wait`) completes every **241–243 s** per box — three sources, 11–12 holds each, every hold 240.00x s.
So a healthy box starts a new wait at most ~243 s after the previous one started.
- **Not measured:** what the hub sees when a box loses power in the middle of a hold. Read in source: the hub writes a
newline every 25 s into nginx (`internal/api/wait.go:10-23`) and ends the hold itself at 240 s, so the hold most likely
ends normally and then **no new wait arrives**. That gap is the signal. Slice 1 measures it.
## 2. The three signals
| Signal | Cadence | What it proves | Blind spot |
|---|---|---|---|
| Agent host report | 900 s | the host and its agent | slow; the alarm needs 45 min hysteresis (R-549) |
| Controller wait channel | ~241 s | the guest's controller is running and reaching the hub | a crashed or updating controller looks like a dead box; per customer, not per host |
| WireGuard handshake age on ep0 | ~2 min (protocol rekey; not measured here) | the host's kernel, independent of both processes | needs a new forced command on ep0 and a hub poller; ep0 is protected |
## 3. Options
**A. Keep the report clock.** Nothing to build. A teardown waits ≤45 min; alarms stay right.
**B. The hub records wait-channel presence; the host page shows it; the delete guard may use it.** In `handleWait`
(once per request — not in `intent.Hub.Wait`, which runs once per 25-s window, `internal/intent/hub.go:77-102`) record
per customer: last wait start, holds open now. Presence = **connected** (a hold is open, or one started < 243 s + 90 s
grace ago), **not connected since T**, or **unknown** (hub restarted < 333 s ago — in memory only, so it falls back to
the report clock; the safe direction). Alerting is not changed. Cost: hub only, ~½ session. Wrong case: a crashed
controller on a live host reads „not connected"; if the guard trusts it, the operator deletes a live host and its agent
is locked out until re-enrolled — no household data is touched.
**C. B + the WireGuard handshake from ep0 as host-level presence.** The best signal (the host itself, a third channel).
Cost: an ep0 change, a new key and forced command, a hub poller; ep0 work needs an operator task.
**Pick: B, in two slices.** Slice 1 is display only and needs no decision. Slice 2 (the guard) needs the operator's word.
C only if slice 1's measurement shows the wait channel does not go quiet when a box dies.
## 4. First slice and its red test
- **Hub:** `intent.Hub` gains `MarkWaitStart / MarkWaitEnd / Presence(customerID, now)` behind a clock seam;
`handleWait` calls them; the host page shows „Box connection: connected now / last connected <time>" from the host's
customer. Debug log on each change of presence.
- **Red tests** (fail today — no such state): (1) a wait starts at t0 and ends at t0+240 s, and none follows → at t0+300 s presence
is `connected`, at t0+334 s it is `not connected since t0+240 s`; (2) a new start at t0+242 s keeps it `connected`; (3) a new
`Hub` (a hub restart) answers `unknown`, never `not connected`; (4) a render test per branch — the „not connected"
line appears only when absent (the seam-built-but-never-wired trap).
- **Live measurement** (after a hub release; scratch 9202 only): stop guest 9202, then read the time until the host
page says „not connected". Control from a different channel: the ingress access log's last `/api/v1/wait` line from
that box's address. Evidence copied off before 9202 is started again.
- **Slice 2 (after the operator's answer):** when report-age says „online" but presence says „not connected" for longer
than the grace, the delete refusal offers a tick box „I checked: the box is off"; the delete is logged and raises one
operator event.
## 5. One question for the operator
**When the hub has had no connection from a box for about six minutes, may „delete host" let you delete it at once,
after you tick „I checked: the box is off"?** My pick: yes (slice 2). It costs: if the box was really on (only its
controller had crashed), its host agent is locked out until it is re-enrolled; no household data is touched. *If you do
nothing:* slice 1 still shows „last connected", but deleting a powered-off host, and so a RESET, still waits up to
45 minutes after its last report.
@@ -0,0 +1,77 @@
# R-314, R-279, R-177 — the operator's door into a running controller: a one-page design (2026-10-08)
Baselines read: felhom-controller `a0370b4`, felhom-agent `b228b44`, felhom.eu `b2dce901`. Architecture:
`03-host-agent.md` §4 (signature only for destroying or overwriting the only copy of customer data),
`04-control-plane-authorization.md` §6 („routine work: no signing"), `07-backup-architecture.md` (decision 74, R-241/R-245).
**Status: design only. Nothing is built.**
## 1. What is still true (read in source today, not from the rows)
- **R-177 is no longer true — close it.** The row says fill-watch runs only at 03:30 and at start. Since controller
v0.297.0 (R-363, 2026-10-05) it also runs every 10 minutes: `cmd/controller/main.go:1579`
(`sched.Every("fill-watch-interval", …)`), `main.go:2478` (`fillWatchInterval = 10 * time.Minute`), pinned by
`cmd/controller/r363_fillwatch_interval_test.go`. Every run logs a positive line
(`internal/fillwatch/fillwatch.go:267`, „checked N filesystem(s) …"), which the operator reads through the hub's
existing controller-log pull. „Did the warning clear after the household freed space?" is now a ≤10-minute wait,
not a restart. Its general ask („run a named job now") is folded into option A below at no extra cost.
- **R-279 is still true.** The only way to start an off-site run on demand is the household's dashboard
(`internal/web/server.go:752` → `offbox_handlers.go:264`). A search of hub, controller report code and agent for any
run-now path finds only the nightly call (`main.go:1417`).
- **R-314 is half true now.** The CLI levers are unchanged (`main.go:91-93`, `229-270`) and run as a SECOND process —
the lost-update problem is written down at `main.go:208-219`. What changed: **decision 74** — on the pinned
off-site tier the deletion is the HUB's, 7 days after the box asks for it, and the operator can cancel it at the hub
(`hub/internal/web/server.go:685-704` → `offsitekeys/service.go:474`; a route, no button). The box follows that cancel
(`internal/backup/offbox_abandon.go:240`). **The gap that remains:** the box's own 14-day countdown
(`offbox_abandon.go:31`) runs first, and a hub cancel during it does nothing — `CancelOffsiteAbandon` only touches
pending rows (`hub/internal/store/offsite_keys.go:284-291`), and on day 14 the box asks again
(`offbox_abandon.go:266`). On a non-pinned target the box deletes by itself on day 14 (`offbox_abandon.go:287`) with
no hub window at all. So a household that telephones in the first 14 days still needs a shell.
## 2. The door already exists — it is the report reply, not the controller's web surface
The 2026-10-05 note looked at the controller's HTTP surface (the household's session). The running process already
takes operator requests another way: operator presses a button → the hub stores a pending row and bumps the box's
intent generation → the controller's wait channel wakes (`internal/report/waiter.go:131-138`) → an out-of-cycle report
→ the reply carries the request (`internal/report/pusher.go:43-50`; hub `internal/api/handler.go:600-610`) → the running
controller acts. Proven for log pulls (hub `internal/web/logbundle.go:43-67`, controller `internal/report/selftail.go`).
Latency: seconds; worst case the 15-minute cycle. The hub never connects into the box.
**Trust.** `03` §4 asks for a signature only to destroy or overwrite the only copy. None of these does: an off-site
run adds a snapshot (the pinned tier cannot delete, decision 69); a check only reads; stop/extend keep data. The hub can
already cancel the hub-phase deletion (decision 74), so „stop" gives a compromised hub no new power. **Not on the list:**
anything that deletes or starts a countdown, and clearing a restore hold (R-379 — it lets an app start on a possibly
broken database; it stays on the CLI).
## 3. Options
| | What | Cost | Risk |
|---|---|---|---|
| **A** | **Operator actions in the report reply.** Hub table `operator_actions(id, customer_id, action, arg, requested_at, done_at, outcome, message)`, buttons on the host page, reply field `operator_actions:[{id,action,arg}]` until a result arrives. Controller: a CLOSED switch in the running process; result in the next report `operator_action_results:[{id,outcome,message}]`; hub marks it done and saves a hub-minted operator event. | ~1 session, two repos, additive wire fields (wire-contract gate). | A controller that crashes mid-action gets it again — every listed action is safe to repeat (single-flight, refuse-when-nothing-running). |
| **B** | Signed job through the agent, which runs `docker exec … --abandon-stop` in the guest. | New signed-op class + an agent exec path; a signing ceremony for a non-destructive act (against `04` §6). | Still a second process: the lost-update at `main.go:208-219` stays. |
| **C** | A unix socket in the container served by the running controller; the CLI flags become its clients. | ~½ session, controller only. | Fixes the second-process problem but the telephone path is still a shell. Good later for R-379. |
**Pick: A.** It reuses a proven channel, needs no key, and the household-visible state changes inside the process
that owns it.
## 4. First slice and its red tests
- **Controller** `internal/report/opactions.go` (the `selftail.go` shape) + a closed map in `main.go`:
`offsite_backup_now` (the same four checks as `offboxRunHandler`, then `RunOffboxBackup` in a goroutine);
`abandon_stop` → `StopAbandon` (which also cancels a hub-phase request); `abandon_extend` (arg 1–30 days) →
`ExtendAbandon`; `run_job` (arg ∈ fill-watch, offsite-integrity, offsite-proof, disk-health-check) → a new
`Scheduler.RunNow(name)` that refuses unknown or running jobs. An unknown action answers `refused`, never silence.
- **Red tests** (each fails today — the reply field does not exist): (1) a reply with `abandon_stop` during a countdown
→ `AbandonStatus().Active` is false in the SAME manager AND `settings.json` on disk agrees; (2) the same id delivered
twice → `StopAbandon` runs once; (3) an unknown action → result `refused`, nothing called; (4) hub: a host-page POST
stores a row and bumps intent; the reply lists it until a result arrives, then not; a result naming another
customer's id is ignored.
- **Live proof** (scratch 9202 only, after the releases): `run_job fill-watch` and `offsite_backup_now`. Positive
observables: the hub's result event; control from another channel: the controller log pull (the „checked N" line)
and the off-site snapshot list on the box's backup page.
## 5. One question for the operator
**May the hub, with only your hub password and no signing key, ask a box to run an off-site backup now, run a named
check now, and stop or extend a household's deletion countdown?** My pick: yes (option A). It costs one session in the
controller and the hub. *If you do nothing:* nothing is built; a household that telephones in the first 14 days of a
countdown is served by a shell on its box, and an off-site run can only be started by the household.
@@ -0,0 +1,45 @@
# R-35 — a config apply ends the household's dashboard session: a one-page design (2026-10-08)
**Status:** design only. Read in source today (controller `main` a0370b4ed8ef). Architecture documents:
`architecture/02-controller-module-map.md` (config refresh), `07-backup-architecture.md` §5 (the whole-guest archive is
plaintext on the household's premises), `09-update-architecture.md` §3 decisions 63-64 (the family gate).
## The problem, as measured
2026-07-21: the hub raised `config_version` 10→11 at 16:54:58; the controller self-restarted (StartedAt 16:54:59Z, up
16:55:02); the household was logged out mid-flow (row evidence `controller-log-full.txt`). Still true in source:
- `internal/report/config_refresh.go:33-71` — any `config_version` change: re-pull `controller.yaml`, record, then
`Restart` = `api.GracefulSelfRestart` (`cmd/controller/main.go:1100-1109`). No per-setting choice.
- `internal/web/auth.go:15-18, 250-270` — dashboard sessions live only in `s.sessions map[string]*session` (token,
expiry, CSRF token). Any process exit ends every session. **This is not only a config-apply problem:** every
controller update (the floor, the household's own update press, a crash restart) logs the household out too.
**New since the row's last note (2026-10-05):** the family gate (decisions 63-64, controller v0.287.0) already persists
its sessions to `family.json` (0600, tmp+fsync+rename) in the data dir „so a restart keeps every session"
(`internal/family/family.go:11-12, 45-50, 267-290`). It stores the session ID itself. So „login tokens on disk" is
already the product's state for family sign-ins; the dashboard is the odd one out.
## Options
| | What | Costs | Risk |
|---|---|---|---|
| **A — hot-apply** | Adopt changed `controller.yaml` fields in the running process; restart only for fields that need it. | A per-field ruling over the whole config (hub, offbox, cloudflare, paths…), a reload path per module. ~2-3 sessions. | A field applied half-way; fixes config apply only, not updates. |
| **B — persist sessions, token on disk** | Write the session map to `sessions.json` (0600) like `family.json`. | ~½ session. | A copy of the archive (plaintext by design, `07` §5) holds live 7-day dashboard tokens. |
| **C — persist sessions, fingerprint on disk** | As B, but the file holds `sha256(token)` → {expiry, CSRF token}; lookup hashes the cookie. Cleared by `invalidateAllSessions` (password change) and logout as today; expired rows dropped at load and at the 15-min cleanup. | ~½ session + tests. | A stolen archive gives a fingerprint that cannot be turned back into a cookie. A box restored from an archive keeps that day's sessions valid until they expire (≤7 days) — the same as family sessions today. |
**Pick: C.** It ends the logout for every restart, not just config apply, at the smallest cost, and it removes the one
cost the row named (tokens in the archive). A stays a later refinement if restarts themselves become a problem.
**First slice, with its red test:** `internal/web/session_store.go` (load at `NewServer`, save under `sessionsMu` on
create/delete/invalidate, atomic write as in `family.saveLocked`). Tests: (1) create a session, build a NEW `Server` on
the same data dir, the cookie is still valid and returns the same CSRF token — red today (map is fresh); (2) the file
contains no token bytes (grep the file for the token: must be absent); (3) after `invalidateAllSessions`, a new
`Server` rejects the old cookie; (4) an expired row is not loaded. Live check on scratch 9202: sign in, trigger a
controller restart, the next page load needs no login.
## One question for the operator
**May the box remember a dashboard sign-in across its own restarts, keeping on its disk only a fingerprint that cannot
be used to sign in?** *If you do nothing:* the household is logged out at every settings push and every controller
update, as today.
@@ -0,0 +1,56 @@
# R-435 — the snapshot-drop detector does not see one app's off-site history vanish: a one-page design (2026-10-08)
Baselines read: felhom.eu `b2dce901`, felhom-controller `a0370b4`. Architecture: `07-backup-architecture.md` §11 row 10
(ransomware / malicious deletion; the 2026-10-03..05 `[FACT]` entries on the append-only key and the clean-up window);
`09` §3 decisions 68, 69 (the box prunes only in a hub-opened window; the hub-provisioned key is append-only).
**Status:** design only, nothing built.
## 1. The problem, and what changed under it since the row was filed (2026-09-01)
- **Still true in code.** `hub/internal/monitor/offsite.go:254-257`: an alarm needs a fall of more than half the last
count AND at least 5 (`snapshotDropped`, `:285-295`). One app's tag (~9 of 69 on demo-hp when measured) is invisible.
The limit is written in the comment above the constants (`:242-253`), as the row says.
- **The row's second complaint is gone.** `STATUS.md` no longer says „noticed within a day" (`grep -n -i "within a day\|noticed" STATUS.md`
returned nothing today). `07` row 10 says „only a fall of MORE than half (R-435)".
- **The threat the row describes is now mostly PREVENTED, not only undetected.** Since decisions 68–69 (2026-10-03) the
box's off-site key cannot delete (provider refusal measured, `07` row 10). The box deletes only inside a hub-opened
window (`controller/internal/backup/offbox_window.go:203-225`, `t.Pinned()`), and inside it the fake-snapshot guard
refuses any plan that removes a snapshot younger than `keepDaily` days that is not superseded the same day
(`offbox_window.go:24-36`). A one-app wipe must remove that app's newest snapshots, so the guard refuses it.
- **What is left uncovered:** a deletion by something that is NOT the box's key — anyone with a deleting credential
for the sub-account or the provider account — or a box that lies about its plan. The count detector sees those only
above one half.
## 2. The new fact that makes a precise detector cheap
On a pinned tier the **only legitimate way the count can fall is a clean-up window**, and the hub records every window
with `count_before` / `count_after` (`hub/internal/store/store.go:835-845`, `offsite_keys.go:84-101`). So the hub can
say „this fall is explained" exactly, with no threshold: between two trustworthy reports, the count may fall by at
most what the windows closed in that interval removed. **Any fall beyond that is unexplained — even one snapshot.**
## 3. Options
| | What | Costs | Risk |
|---|---|---|---|
| **A** | Keep as is: documented blind spot. | Nothing. | A single-app deletion from outside the box stays silent. |
| **B** | The row's own suggestion: the controller reports a count per app tag; the hub keeps a per-tag baseline and threshold. | Two repos, a new report field (wire-contract gate), a per-tag threshold nobody has measured; retention also moves per-tag counts, so it needs its own calibration. ~1.5 sessions. | Noise from retention on boundary days — the reason the global threshold is high. |
| **C** | Hub only: on a **pinned** tier (hub holds a confirmed append-only key for the customer), compare `prev − cur` with the snapshots removed by windows closed since the previous trustworthy report. Any unexplained fall ≥ 1 raises `offsite_snapshots_dropped`. Non-pinned tiers (the household's NAS) keep today's half-rule. | One repo. One store query (windows closed in an interval), one branch in `snapshotDropped`, tests. ~½ session. | A window that closes by timeout without a box result has no `count_after`; treat it as „explains anything" (no alarm, log one INFO line) — the safe side for noise, recorded as a limit. |
## 4. The pick — C
C covers the shape B was meant to cover — and any size — without a new threshold, a new field or the second repo. It
uses the only fact that is exact (the windows the hub itself opened). It keeps the half-rule for NAS tiers, where the
box still prunes by itself. B stays a note: it would add per-app naming in the message, which C does not have (C says
„N snapshots went, outside any clean-up window"; the operator reads which ones on the box).
## 5. First slice and its red test
- Hub: `store.RemovedByWindowsBetween(customerID, from, to) (removed int, unknown bool)`; `OffsiteChecker` keeps the time
of the last trustworthy baseline beside `lastCounts`; on a pinned customer, alarm when `prev − cur > removed` and not
`unknown`. The message adds „outside any clean-up window".
- **Red test (seen failing on today's code):** pinned customer, baseline 69, next trustworthy report 60, no window →
exactly one `offsite_snapshots_dropped` event. Today: none (9 < 34.5).
- **Controls:** same fall with a closed window that removed 9 → no event; NAS (not pinned) customer 69 → 60 → no
event (half-rule kept); untrustworthy report → no event and baseline unchanged (existing rule).
- Live proof, when releases resume: on scratch 9202 only — a window granted by hand removes N and no alarm; then the
hub's own row for that window as the control from another channel.
## 6. One question for the operator
**Should the hub raise an error mail when even ONE off-site snapshot disappears outside a clean-up window it opened?**
My pick: yes (option C). *If you do nothing:* deletions by the box stay blocked as today, but a deletion through any
other credential stays silent unless it removes more than half of a household's off-site history.
@@ -0,0 +1,49 @@
# R-717 — opengist's sign-up switch: its container has no tool to write it: a one-page design (2026-10-08)
**Status:** design only. Read today: controller `main` a0370b4ed8ef, catalog `main` 32d134639c0c, opengist upstream
`master` (fetched 2026-10-08). Architecture: `09-update-architecture.md` §3 decisions 46-47 (setup gate, sign-up block),
`01-topology-and-trust.md` §5 (the gate).
## Where it stands
- Wishlist is done (controller v0.301.0 `open_command`, proven on 9202 2026-10-06). The row is narrowed to opengist.
- Opengist (`templates/opengist/docker-compose.yml`, image `ghcr.io/thomiceli/opengist:1.15`) is closed by the address
block alone: `signup_block: "PathRegexp(`(?i)^/+-/+register`)"` (`.felhom.yml:42`). Measured 2026-09-29: every trick
shape — trailing slash, upper/mixed case, percent-encoded, double slash, query string — answers 403
(`audits/signup-lock-2026-09-29/B/B6-tricks-all-regex.txt`). The app publishes no port; traefik is its only door.
- **Upstream, read today:** the switch is row `disable-signup` in opengist's admin-settings table; startup seeds it to
`"0"` from a hard-coded map — no config key and no env feeds it (`internal/db/db.go`; `config.yml` has no
signup/register key). **The settings are read from the database on every request** (`dataInit` middleware →
`loadSettings` → `db.GetSettings()`, `internal/web/server/middlewares.go`). So a row written while the app runs takes
effect at the next request — the row's „needs its sqlite with the app stopped" is not needed.
- The controller already embeds a pure-Go SQLite (`modernc.org/sqlite`, `go.mod:11`) and already runs short helper
containers on app volumes (`docker run --rm -v vol:…`, `internal/stacks/undo.go:205-235`, image `alpine`, which has
no sqlite).
## Options
| | What | Costs | Risk |
|---|---|---|---|
| **A — the box writes the row with its own image** | New `after_setup` form `sqlite: {volume, file, close_sql, open_sql, check_sql}`; the controller runs `docker run --rm -v opengist_data:/v <its own image> <subcommand>` that opens the file with modernc sqlite (busy timeout), runs the statement, reads it back and prints the marker. Feeds the existing close/open/loop machinery unchanged. | ~1 session controller + catalog line + a 9202 proof. No new external image. | A second writer on a live SQLite (normal locking; same kernel). A schema rename upstream → the read-back fails → the existing failure record, never silent. |
| **B — a sidecar with a sqlite tool** | Add a small sqlite image as a second service in the template. | ~½ session. | **A new external dependency** (fenced — operator only), RAM and an image to keep current on every box. |
| **C — keep the address block alone** | Accept the measured block as the lock for opengist; close the row as decided. | Nothing. | The switch stays „open" inside the app; only a path traefik does not see could reach it, and the app has no such path today. |
**Pick: C, with A ready.** The block is measured against every trick shape and the app has no door but traefik, so A
buys defence in depth for a P3-LOW. A is the right build if the operator wants the app's own switch closed too: it adds
no dependency and serves any later app whose switch lives only in its database.
**First slice of A (only on the operator's yes), with its red test:** `after_setup.go` accepts the `sqlite` form; a test
with a temp SQLite file runs close → the read-back prints the marker and the row reads `1`; open → `0`; a missing table
→ an error and no marker (red today: the form is unknown and the spec is refused). Then the opengist template line and a
9202 proof in wishlist's shape (stranger 403/refused after setup, a family member signs up inside the window, closed again
after it).
**Small, for the parent now (comments only, no decision):** `internal/stacks/after_setup.go:33` names opengist among the
apps „closed by `command`" — today only wishlist is; `internal/stacks/signup_block.go:22` shows opengist's block as
``PathPrefix(`/-/register`)`` while the template uses the case-insensitive `PathRegexp` above.
## One question for the operator
**Is the measured address block enough for opengist, or should the box also close opengist's own sign-up switch (about
one working session)?** *If you do nothing:* opengist stays closed by the address block alone, and the row closes as
decided.
@@ -0,0 +1,57 @@
# R-79 — health issue and warning text on the household's pages and mail: a one-page design (2026-10-08)
**Status:** design only. Read in source today (controller `main` a0370b4ed8ef, hub `main` b2dce901b23f). Architecture
document: `architecture/10-localisation.md` (§9 „compared, not shown", §10.1 „on the wire = not translated",
§10.6b residue table).
## The row is half stale — what moved since it was filed
The row (filed 2026-07-26) says „every producer is English" and asks for a spike to pick the seam. **The seam was picked
and built** by R-516 item 10 (controller, after 0.252.0): the wire text stays frozen (`internal/monitor/wire_golden_test.go`)
and each entry carries a parallel `MsgRef` (bundle key + args) that the dashboard banner renders in the household's
language (`internal/monitor/healthcheck.go:28-35, 73-86`; `internal/web/alerts.go:238-279`, `MessageKey`/`MessageArgs`
rendered by `GetAlerts(lang)`). That is the render-boundary seam, and it is the right one: the hub report keeps its bytes.
**Done (key attached, pinned by `internal/monitor/r516_resource_msgs_test.go`):** all 7 resource entries —
system/data disk high + critical, memory, CPU, temperature (`healthcheck.go:357-441`).
**Still verbatim on the household's banner (zero `MsgRef`):**
| producer | `healthcheck.go` | wire text | what a household sees |
|---|---|---|---|
| Docker unreachable | 138 | `Docker: %v` (English) | English in both languages |
| protected container down | 157 | `Protected container not running: %s` (English) | English in both languages |
| storage almost full (issue) | 170 via `checkStoragePaths` 336 | `issueFmtStorageAlmostFull` (Hungarian, frozen) | Hungarian for an English household |
| storage unavailable / not separate / usage high (warnings) | 173 via 322, 328, 338 | `warnFmtStorage*` (Hungarian, frozen) | Hungarian for an English household |
(The disconnected-drive warning is already replaced by its own banner, `alerts.go:254-261`.)
**The mail half, not covered by R-516:** `health_critical` is customer-on by default
(`controller internal/settings/settings.go:635`). The controller sends `HealthDetails{issues, warnings}` as the event's
details (`internal/notify/notifier.go:139-144, 383-396`). The hub's customer mail localises the headline
(`mail.event.health_critical`) but then appends the **raw details JSON** as „- Megjegyzés: {…}"
(`hub internal/notify/templates.go:187-189`, called from `dispatcher.go:331, 880`). So the household's mail carries
`{"previous_status":"ok","current_status":"fail","issues":["Protected container not running: cloudflared"]}`.
## Options
| | What | Costs | Risk |
|---|---|---|---|
| **A** | Dashboard only: give the 6 remaining producers a `MsgRef` (`checkStoragePaths` returns a parallel `[]MsgRef`); 6 hu+en keys whose Hungarian equals today's text. | ~2 h, controller only. | None on the wire (the wire golden stays green by construction). |
| **B** | A + mail: the hub stops appending the raw details JSON to the CUSTOMER mail of `health_degraded` / `health_critical` / `health_recovered` (operator mail unchanged); the household's mail says what the headline says and points to the dashboard, where A shows the list in its language. | A + ~1 h hub (one set + one test). Needs a hub release. | The household loses a list it could not read anyway; the operator keeps it. |
| **C** | A + translated list in the mail: details carry `issue_keys` [{key,args}] and the hub renders them from its own bundle. | Two-repo, ~1 session; the hub must hold copies of controller keys (two producers, one sentence — the R-516 trap). | Drift between the two bundles. |
**Pick: A now, B after the operator's word.** A is a small, decision-free fix in the shape R-516 item 10 already set.
B is the cheapest honest mail; C buys a translated list at the price of a second copy of every sentence.
**First slice (A), with its red test:** extend `r516_resource_msgs_test.go` with a case per remaining producer
(`checkStoragePaths` with a fake unavailable / not-separate / >90 % path; the Docker and protected-container issues via
their seams) asserting (1) the wire text is unchanged byte for byte and (2) the entry carries its key. Red-proof: drop one
`MsgRef` → the case fails and the `internal/web` banner test shows the wire text. The i18n Go-parity gate
(`scripts/i18n_go_parity.py`) checks the Hungarian values equal the frozen literals.
## One question for the operator
**May the household's „system health" e-mail stop showing the raw technical note, and send the household to the
dashboard for the list instead?** *If you do nothing:* A ships alone; the dashboard is right in both languages, and the
mail keeps its note — a line of English inside curly brackets — as it does today.
@@ -0,0 +1,56 @@
# R-893 — after a failed off-site database replay, the rollback mixes new and old: a one-page design (2026-10-08)
Baselines read: felhom-controller `a0370b4`, felhom.eu `b2dce901`. Architecture: `07-backup-architecture.md` §6.3
(„[FACT] 2026-10-06 — restoring over a newer schema", the R-893 known limit) and „[DESIGN] 2026-08-22 — the failure
ladder: replay → rollback → hold"; `09` §3 decision 154 (R-638 option A). **Status:** design only, nothing built. Not
measured yet — the 9202 measurement the row asks for is slice 0 below.
## 1. The problem (read in source today, `controller/internal/backup/offbox_reconstitute.go`)
The off-site restore of one app does, in order:
1. `:758` — the undo copy: a **logical dump of the LIVE database** (`writeSafetyDump`, `:214`). Nothing else is saved.
2. `:804-811` — when the snapshot was written at another app version, the **snapshot's older definition** is written.
3. `:812-842` — the snapshot's **files are copied over** the live ones (placements).
4. `:856` — the **named volumes are replaced** with the snapshot's tars (this includes the database volume).
5. `:873-917` — DB-only start, replay of the snapshot's dump. **If the replay fails**, `rollbackSafetyDump` (`:456`)
loads the step-1 dump (NEWER state) into the database that now sits on the step-4 (OLDER) volume, then the app starts.
What the household gets after a failed replay:
- the database: the older volume with the newer dump laid over it — the loader drops only what the dump knows, so
tables only the older version had stay (`07` §6.3);
- files and other volumes: the **snapshot's** state (steps 3–4 are never undone);
- the definition: the **older** one when step 2 ran — the rollback branch never writes the live one back.
And the sentence shown (`hu.json:1261`, `en.json:1266`, key `err.backup.db_restore_failed_rolled_back`) says
„your data is back as it was before the restore, and the app is still running". **That is true only of the database
rows, and not of the files, the other volumes or the version.**
## 2. Options
| | What | Costs | Risk |
|---|---|---|---|
| **A** | **Pre-restore copy of the app's volumes.** After step 1, while stopped, tar the live named volumes to an undo folder beside the undo dump (the recovery-unit volume-dump helper; one implementation). On a failed replay: put back those volumes, the live definition and the live files' copy → the exact state before the restore. A fit check refuses to start when the copy does not fit (R-685's lesson). | Disk = one copy of the app's volumes for the restore's duration; time ≈ one volume dump. ~1 session + a 9202 drill. | Small: the helpers exist; the new part is the undo of steps 2–4. Files placed by step 3 need the same treatment, or step 3 must write to a staging folder first. |
| **B** | **R-638 option B: a loader that rebuilds** (drop and re-create the database before loading). The newer dump over the older volume then gives exactly the newer database. | Changes the loader under EVERY restore path, Postgres and MariaDB. ~2 sessions. | High blast radius; still leaves files, other volumes and the definition at the snapshot's state. |
| **C** | **Hold instead of a mixed start.** On a failed replay when step 2 ran (version changed) or step 4 replaced any volume: roll the database back as now, write the live definition back, and **hold the app stopped** (the ladder's third step, `holdAppAfterFailedRollback`) with an honest sentence and an operator event. | Small: one branch, one sentence pair (hu/en), tests. ~½ session. | The app is down until the operator acts; no household data is changed beyond what happens today. |
## 3. The pick — C now, then A
C removes the worst outcome (an older app running on mixed data while the screen says „back as it was") at no disk
cost and without touching the loader. A is the real fix — it makes „back as it was" true — and is built after the 9202
measurement shows how big the volume copy is for the standing apps. B stays a note: it fixes only the database and
widens the risk to every restore.
## 4. First slices and their red tests
- **Slice 0 — measure (9202 only, throwaway app):** off-site snapshot at version N, Update to N+1 (migrating), then an
off-site restore with the replay forced to fail (the `rollbackImport`/`reimportDBDumpsFrom` seam is for tests; on the
box, a snapshot dump truncated by hand in the scratch copy). Record: tables, files, app version, the sentence. Evidence
copied off before teardown.
- **Slice 1 (C):** red test beside `r379_rollback_test.go` — `ReconstituteFromOffsite` with `defineFromSnapshot` true
and a failing replay must (a) call the definition write with the LIVE definition, (b) NOT start the app, (c) leave a
hold, (d) return a sentence that does not contain „visszakerültek" / „back as it was". Today (a)–(d) all fail —
`TestR379_ScenarioA` pins the opposite for its fixture, which has no version change, so it stays green.
- **Slice 2 (A):** red test — after a failed replay, the volume-replay seam is called a second time with the undo
folder, and the app starts; a too-small free space refuses before step 1 with nothing touched.
## 5. One question for the operator
**When an off-site restore of one app fails half-way, which do you want: the app stopped and waiting for us (no extra
disk), or the app put back exactly as it was (needs free space for one more copy of that app's data during the restore)?**
My pick: stopped now, „exactly as it was" next. *If you do nothing:* after such a failure the app keeps starting on a
mix of the older backup's files and the newer database, and the screen tells the household everything is back as it was.
+2
View File
@@ -32,6 +32,8 @@ The full text of every row below: `git show b2dce901b2:documentation/backlog/OPE
| Row | What | Closed | Evidence |
|---|---|---|---|
| **R-177** | **There was no operator-triggerable „run the fill check now" path.** (P4) | CLOSED 2026-10-08 — NO LONGER TRUE: since controller v0.297.0 (R-363) the fill check also runs every 10 minutes, so no door is needed for it | `felhom-controller/controller/cmd/controller/main.go:1579` (`sched.Every("fill-watch-interval", …)`), `:2478` (`fillWatchInterval = 10 * time.Minute`), pinned by `cmd/controller/r363_fillwatch_interval_test.go`; each run logs „checked N filesystem(s)" (readable through the hub's controller-log pull). Design `audits/day-2026-10-08/design-R-314-279-177.md`. |
| **R-298** | **The `/storage` page's unregistered list filtered on `role==='user-data'`, so a drive that is also the backup target could never be registered.** (P3) | CLOSED 2026-10-08 — NOT REPRODUCIBLE / NOT TRUE: the agent's role comes from the storage type and backing disk and never from being the backup target, so a backup-target drive on a non-system disk IS user-data | `felhom-agent/internal/storage/role.go:172-186` (unchanged since 2026-07-13); `internal/localapi/disks.go:212-233` and `:1239-1249`; August topology `audits/evidence-rehearsal-2026-08-09/GATE0-demo-hp-before.txt:47-84` → user-data; read-only today: demo-felhom `felhom-backup` on `sdb` (root on `sda`). Side fix on controller main `a40729a`: no eject/format button on a backup-target drive (the agent refuses the eject, 403). |
| **R-900** | **The website's FAQ said the household's data is not with a third party, while copies and traffic go to processors.** (P2) | CLOSED 2026-10-08 — PUBLISHED: the GDPR answer on `gyik.html` and `en/faq.html` (visible text and structured data) now names the box, the encrypted copies at Hetzner in the EU, and Cloudflare's view of remote-access traffic; text approved by the operator in chat (`09` §3 decision 180) | website commit `b2dce901`; live read-back 2026-10-08: the new sentence 2× on each page (visible + JSON-LD), the old „nem harmadik félnél" 0×; rest of the site searched (ASCII fragments, positive control) — no other page makes the claim. Left for R-813: the contact form's consent line „harmadik félnek nem adjuk ki", which the consent draft already replaces. |
## 2026-10-07 — the operator's answers on the night designs
File diff suppressed because one or more lines are too long