diff --git a/CONTEXT.md b/CONTEXT.md index ec556dd8..81de5859 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,23 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## Decisions 2026-09-17 — the chaos-night fixes (operator rulings) + +**Ruling A (R-549): the quiet-box alarm waits three report cycles.** `alerting.stale_threshold` 30 m → +**45 m**; `node_down`/`host_down` at 90 m. Mechanism: configuration in `manifests/hub.yaml`, plus hub +v0.117.0 making the dashboard's customer status read the same value (it was hardcoded). Cost accepted: +a dead box pages 15 minutes later. `08-alarm-ladder.md` §6.2. + +**Ruling „fix" (R-550): the restore record is persisted — a reversal of the in-memory design, for the +restore record only.** Controller v0.246.0, `restore-status.json` in DataDir; `restore_interrupted` +(warning, household). Notification cooldowns stay in memory. `07-backup-architecture.md` §3. + +**R-546 (no new ruling — a fix inside R-543's):** the recovery-code reminder and the escrow page wait for +the agent's own preflight `ok`; the start API refuses before staging. Controller v0.246.0. + +**Ruling 3 of 2026-09-16, built (R-539):** the slow crash-loop counter, agent v0.132.0 / hub v0.117.0, +signed to demo-hp and the N100 by CC under ruling 1 of 2026-09-16. `03-host-agent.md`. + ## Decisions 2026-09-16 (evening) — the recovery code: the household is asked, the page says „szünetel" **No new ruling. One correction to the record, from the operator (2026-09-16):** the note saying a diff --git a/documentation/architecture/03-host-agent.md b/documentation/architecture/03-host-agent.md index 80df4fd2..ce942601 100644 --- a/documentation/architecture/03-host-agent.md +++ b/documentation/architecture/03-host-agent.md @@ -121,6 +121,18 @@ a long window (N restarts in 24 h) raising a `warning`. Not built in the 2026-09 told to measure the budget rather than change it. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f9dprime.txt`. +**[BUILT 2026-09-17 — agent v0.132.0, hub v0.117.0] The slow counter.** Beside the unchanged 3-in-15 +brake: restarts the supervisor performed in the last **24 hours**; at the **fifth**, the heartbeat stanza +sets `slow_crashloop_since` (moving at most once per 24 h) and the hub raises +`controller_slow_crashloop` (**warning, operator-only**). It never stops restarting — the fast brake stays +the only brake. **Persisted** per guest at `/var/lib/felhom-agent/guests//controller-restarts-24h.json` +(tmp + rename, 0600) so an agent restart or reboot does not reset it; the fast record stays in memory, +and the reason it does (a persisted „give up" could outlive the fix) does not apply to a counter that +only warns. **Deliberate kills count** — the supervisor cannot tell an operator's `docker kill` from a +crash (measured 2026-09-15). Delivered to demo-hp and the N100 by one operator-signed `agent_update` each +(ruling 1 of 2026-09-16); both logged `slow_crashloop_max=5 slow_crashloop_window=24h0m0s` at start. +Evidence: `audits/evidence-chaos-fixes-2026-09-17/partC-agent-update.txt`. + Signed payloads carry a **nonce + expiry** (anti-replay: a captured "restore" job cannot be re-injected later) and a target binding (host + guest id) so a signature can't be retargeted. Notification-on-destructive-op is an **audit signal, never the guard** — a compromised hub diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index c07eeba8..20a5ba47 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -122,6 +122,17 @@ distinct action (`audits/CAMPAIGN-8-backup-restore-2026-07-27.md:520`). not by a customer. `00-capability-map.md:75` still carries "A customer (not the operator) performs a restore via UI alone" as **MISSING as evidence** — by absence of the run, not by a product gap. +**[RULING 2026-09-17, operator — a design choice REVERSED] The restore record survives a restart +(R-550, controller v0.246.0).** The async restore status (`internal/backup/opstatus.go`) was in memory +by choice — „same precedent as notification cooldowns". Chaos night round 10 measured the cost: a +restore accepted, the box hard-reset four seconds later, and the status answering the Go zero value, so +the household could never learn whether it finished. **For the restore record only**, it is now written +atomically to `restore-status.json` in the controller's state directory at both ends of an op; at +startup a record still marked running becomes a failed, interrupted result („A visszaállítás megszakadt +(a doboz újraindult) — indítsd el újra."), shown per app on the restore page until that app's next +restore, and raised once as `restore_interrupted`. **Notification cooldowns stay in memory** — the +precedent is kept for what it was written for. + ### Lane 2 — the operator: guest and host recovery Rebuilding an LXC guest, rebuilding a host, and re-establishing a box's identity are **operator diff --git a/documentation/architecture/08-alarm-ladder.md b/documentation/architecture/08-alarm-ladder.md index 32c6d36c..828c78d3 100644 --- a/documentation/architecture/08-alarm-ladder.md +++ b/documentation/architecture/08-alarm-ladder.md @@ -198,6 +198,26 @@ same box — the agent's dead-man's-switch rather than the controller's — and would have kept exactly the F9 silence on the host plane. Everything else keeps the hour (pinned by `TestOperatorCooldown_NodeLivenessBypassesQuietHour`, which also asserts an unrelated type still waits). +**Operator ruling A, 2026-09-17 (R-549) — „the box went quiet" waits THREE report cycles, not two.** +The staleness threshold (`alerting.stale_threshold`, configuration in `manifests/hub.yaml`) moves from +**30 m to 45 m**; `node_down` and `host_down` follow at 2× = **90 m**. The reason is chaos night round 9 +(`audits/DRILL-chaos-night-2026-09-17.md`): report cadence 15 m; a failed push is retried for ~100 s and +then given up (correctly — a report is a snapshot); the measured gap to the next good report was +**29 m 59 s** against a 30-minute threshold. One missed push spent the entire budget, so ordinary jitter +would page the operator about a box that was healthy and had already repaired itself. **The cost, +stated:** a truly dead box now pages 15 minutes later (45 m instead of 30), and „the box is down" 30 +minutes later (90 m instead of 60). **One value, everywhere:** both staleness checkers, the host status and +— since hub v0.117.0 — the customer status on the dashboard read the same threshold (`controllerStatus` +hardcoded 30 m / 1 h until then; pinned by `TestControllerStatus_FollowsConfiguredThreshold`). A running +hub prints it at startup (`node_stale after 45m0s, node_down after 1h30m0s`). + +**Two event types added 2026-09-17, with who receives them:** + +| event | severity | who | minted by | why that audience | +|---|---|---|---|---| +| `controller_slow_crashloop` | warning | **operator only** | hub, from the agent's `slow_crashloop_since` moving (agent v0.132.0, hub v0.117.0) | host ids and vmids; the household's side of it is the dashboard coming back each time | +| `restore_interrupted` | warning | **the household** (and the operator) | controller v0.246.0 at startup, once per interruption | they pressed restore, were told it started, and can run it again — Hungarian `customerMessages` entry | + | Family | Grain | Key carries | Why | |---|---|---|---| | app down (`app_start_failed`) | **per APP** | `…:` | no digest exists for it — see below | diff --git a/documentation/audits/evidence-chaos-fixes-2026-09-17/partA-hub-45m.txt b/documentation/audits/evidence-chaos-fixes-2026-09-17/partA-hub-45m.txt new file mode 100644 index 00000000..5096b40d --- /dev/null +++ b/documentation/audits/evidence-chaos-fixes-2026-09-17/partA-hub-45m.txt @@ -0,0 +1,19 @@ +# Part A - R-549, operator ruling A: the quiet-box alarm waits three report cycles +# Deployed 2026-09-17 ~08:22Z. Manifest commit 06334e164c51; hub image 0.117.0. + +## ArgoCD, verified by revision and image - not by the rollout message + before sync: OutOfSync rev=06334e164c519a9926e868c23e6bda2f58f915cc + after: Synced, health Healthy, rev == manifest HEAD: YES + image: gitea.dooplex.hu/admin/felhom-hub:0.117.0 + NOTE: `rollout status` printed "successfully rolled out" while the new pod was still 0/1 ready. + Readiness was waited for separately (ready=true at 08:22:51Z). + +## The live ConfigMap + stale_threshold: "45m" + +## THE PROOF - the running hub's own startup lines (these fields did not exist before v0.117.0) + 2026/09/17 10:22:34 [INFO] Staleness checker initialized: 2 ok, 0 stale, 2 down (node_stale after 45m0s, node_down after 1h30m0s) + 2026/09/17 10:22:34 [INFO] Host staleness checker initialized: 2 ok, 0 stale, 0 down (stale/down left unseeded → first Check emits; host_stale after 45m0s, host_down after 1h30m0s) + (log times are the pod's local CEST; 10:22 CEST = 08:22Z) + +## Hub serving: https://hub.felhom.eu/healthz -> 200 diff --git a/documentation/audits/evidence-chaos-fixes-2026-09-17/partB4a-restore-interrupted.txt b/documentation/audits/evidence-chaos-fixes-2026-09-17/partB4a-restore-interrupted.txt new file mode 100644 index 00000000..f5afd1a9 --- /dev/null +++ b/documentation/audits/evidence-chaos-fixes-2026-09-17/partB4a-restore-interrupted.txt @@ -0,0 +1,25 @@ +2026-09-17T08:49:35Z === PROOF A: a restore interrupted 2 s in, on 9201, throwaway app homebox === +2026-09-17T08:49:35Z login: session=yes csrf_len=64 +tr: write error: Broken pipe +2026-09-17T08:49:35Z deploy homebox -> 202 {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"} +2026-09-17T08:49:46Z homebox container: running +2026-09-17T08:50:06Z backup run -> 200 +2026-09-17T08:51:51Z backup finished: {"ok":true,"data":{"db_dump":{"count":6,"duration":"1m43.068591229s","last_run":"2026-09-17T08:51:49.156859606Z","success":true},"enabled":true,"running":false}} +2026-09-17T08:51:51Z login: session=yes csrf_len=64 +2026-09-17T08:51:51Z restore-status BEFORE: {"ok":true,"data":{"running":false,"started_at":"0001-01-01T00:00:00Z"}} +2026-09-17T08:51:51Z POST /backup/restore homebox -> HTTP/1.1 302 Found +2026-09-17T08:51:53Z restore-status 2 s in: {"ok":true,"data":{"running":true,"op":"restore","stack":"homebox","started_at":"2026-09-17T08:51:51.445442422Z"}} +2026-09-17T08:51:53Z KILL: docker kill felhom-controller -> felhom-controller +2026-09-17T08:51:53Z record on disk right after the kill: {"running":true,"op":"restore","stack":"homebox","started_at":"2026-09-17T08:51:51.445442422Z"} +2026-09-17T08:52:33Z controller back: state=running after 42 s (NOT restarted by hand) +2026-09-17T08:52:59Z login: session=yes csrf_len=64 +2026-09-17T08:52:59Z restore-status AFTER: {"ok":true,"data":{"running":false,"started_at":"0001-01-01T00:00:00Z","last":{"op":"restore","stack":"homebox","ok":false,"message":"A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra.","finished_at":"2026-09-17T08:52:33.050939642Z","interrupted":true},"last_recent":true}} +2026-09-17T08:53:14Z /backups/restore: bytes=68722 card('Megszakadt vissza')=1 msg('megszakadt')=1 NEGATIVE('zzzz-not-present')=0 + card: homebox: A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra. +2026-09-17T08:53:14Z controller log lines: + 2026/09/17 08:52:33 restore_record.go:88: [WARN] [backup] restore of homebox (restore, started 2026-09-17T08:51:51Z) was INTERRUPTED by a controller stop — recorded as failed; the household is told to run it again + 2026/09/17 08:52:33 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: true) + 2026/09/17 08:52:33 notifier.go:208: [DEBUG] PushEvent: type=restore_interrupted severity=warning url=https://hub.felhom.eu/api/v1/event + 2026/09/17 08:52:33 notifier.go:234: [DEBUG] PushEvent: restore_interrupted pushed OK (HTTP 200) + 2026/09/17 08:52:33 notifier.go:236: [INFO] Event pushed: restore_interrupted (warning) — A(z) homebox visszaállítása megszakadt, mert a doboz újraindult — indítsd el újra. +2026-09-17T08:53:14Z homebox after: running diff --git a/documentation/audits/evidence-chaos-fixes-2026-09-17/partB4a-teardown.txt b/documentation/audits/evidence-chaos-fixes-2026-09-17/partB4a-teardown.txt new file mode 100644 index 00000000..5cc1ff64 --- /dev/null +++ b/documentation/audits/evidence-chaos-fixes-2026-09-17/partB4a-teardown.txt @@ -0,0 +1,17 @@ +2026-09-17T08:54:29Z === TEARDOWN A (and the live half of 'the notice lasts until that app's next restore') === +2026-09-17T08:54:29Z second restore of homebox -> HTTP/1.1 302 Found +2026-09-17T08:54:39Z restore-status: {"ok":true,"data":{"running":false,"op":"restore","stack":"homebox","started_at":"2026-09-17T08:54:29.123541604Z","last":{"op":"restore","stack":"homebox","ok":true,"message":"A(z) homebox: 1 adatkötet visszaállítva — az alkalmazás újraindult.","finished_at":"2026-09-17T08:54:38.543283958Z"},"last_recent":true}} +2026-09-17T08:54:50Z /backups/restore card('Megszakadt vissza')=0 (must be 0 now) +2026-09-17T08:54:50Z record on disk: {"running":false,"op":"restore","stack":"homebox","started_at":"2026-09-17T08:54:29.123541604Z","last":{"op":"restore","stack":"homebox","ok":true,"message":"A(z) homebox: 1 adatkötet visszaállítva — az alkalmazás újraindult.","finished_at":"2026-09-17T08:54:38.543283958Z"}} +2026-09-17T08:54:50Z remove homebox (data + backups) -> 409 {"ok":false,"error":"stack \"homebox\" is still running — stop it first before removing"} +2026-09-17T08:56:51Z homebox containers left: 1 volumes left: 1 +2026-09-17T08:56:51Z standing apps still up: 10 of 10 checked; controller: running +2026-09-17T08:56:51Z guest secrets/scripts left: 1 (teardownA.sh removes itself next) +left: 0 +2026-09-17T08:57:13Z === TEARDOWN A, second pass: stop first, then remove (my first pass skipped the stop and was refused 409) === +2026-09-17T08:57:13Z stop homebox -> 200 {"ok":true,"message":"Stack homebox stop completed"} +2026-09-17T08:57:13Z homebox state: +2026-09-17T08:57:14Z remove homebox (data + backups) -> 200 {"ok":true,"data":{"removed":"homebox","volumes_removed":[],"hdd_paths_removed":[],"hdd_paths_preserved":[],"hdd_note":"Az alkalmazás nem tárolt saját adatot külső meghajtón, így ott nem volt m +2026-09-17T08:57:14Z homebox containers left: 0 volumes left: 0 stack dir app.yaml deployed: 0 +2026-09-17T08:57:14Z standing apps: 10 of 10 up; controller running +guest leftovers: 0 diff --git a/documentation/audits/evidence-chaos-fixes-2026-09-17/partC-agent-update.txt b/documentation/audits/evidence-chaos-fixes-2026-09-17/partC-agent-update.txt new file mode 100644 index 00000000..10485e38 --- /dev/null +++ b/documentation/audits/evidence-chaos-fixes-2026-09-17/partC-agent-update.txt @@ -0,0 +1,20 @@ +# Part C - agent v0.132.0 delivered by operator-signed agent_update jobs (ruling 3 of 2026-09-16) +# released sha256 4afe815749a41b327ebe4a98a4557ad2acffb1fa71740a3a8b8003473b835321 +signed+uploaded 2026-09-17T08:27:36Z: demo-hp-bb76ea nonce=dfc91f1c9bb7eb789fc99469adba4392; demo-felhom-8363b5 nonce=8e183e6634041822d6b161a6de7d7768; expires 09:12:36Z +## 2026-09-17T08:30:07Z hp now runs 0.132.0 - waiting 75 s for the dwell/commit + [hp] Sep 17 10:29:25 demo-hp felhom-agent[1526161]: time=2026-09-17T10:29:25.418+02:00 level=WARN msg="signedjobs: AUTHORIZED signed op — executing" job=dbc1b52384ade807 op=agent_update key_id=felhom-op-1 nonce=dfc91f1c9bb7eb789fc99469adba4392 + [hp] Sep 17 10:29:25 demo-hp felhom-agent[1526161]: time=2026-09-17T10:29:25.886+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=dbc1b52384ade807 op=agent_update + [hp] Sep 17 10:29:29 demo-hp felhom-agent[1195783]: time=2026-09-17T10:29:29.858+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_window=24h0m0s guests_dir=/var/lib/felhom-agent/guests + [hp] Sep 17 10:29:29 demo-hp felhom-agent[1195783]: time=2026-09-17T10:29:29.859+02:00 level=WARN msg="selfupdate: new version running — dwelling before commit" version=0.132.0 prev=0.131.0 dwell=1m0s + [hp] Sep 17 10:30:29 demo-hp felhom-agent[1195783]: time=2026-09-17T10:30:29.884+02:00 level=WARN msg="selfupdate: update committed" version=0.132.0 wrapper="" + [hp] active: active + [hp] felhom-agent 0.132.0 +## 2026-09-17T08:35:11Z felhom-pve now runs 0.132.0 - waiting 75 s for the dwell/commit + [felhom-pve] Sep 17 10:34:48 demo-felhom felhom-agent[2416121]: time=2026-09-17T10:34:48.018+02:00 level=WARN msg="signedjobs: AUTHORIZED signed op — executing" job=640bd623316627c5 op=agent_update key_id=felhom-op-1 nonce=8e183e6634041822d6b161a6de7d7768 + [felhom-pve] Sep 17 10:34:48 demo-felhom felhom-agent[2416121]: time=2026-09-17T10:34:48.396+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=640bd623316627c5 op=agent_update + [felhom-pve] Sep 17 10:34:52 demo-felhom felhom-agent[3550256]: time=2026-09-17T10:34:52.247+02:00 level=WARN msg="selfupdate: new version running — dwelling before commit" version=0.132.0 prev=0.131.0 dwell=1m0s + [felhom-pve] Sep 17 10:34:52 demo-felhom felhom-agent[3550256]: time=2026-09-17T10:34:52.247+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_window=24h0m0s guests_dir=/var/lib/felhom-agent/guests + [felhom-pve] Sep 17 10:35:52 demo-felhom felhom-agent[3550256]: time=2026-09-17T10:35:52.283+02:00 level=WARN msg="selfupdate: update committed" version=0.132.0 wrapper="" + [felhom-pve] active: active + [felhom-pve] felhom-agent 0.132.0 +## 2026-09-17T08:36:26Z watcher end: hp=1 n100=1 diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 7340d364..9483bff1 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -318,3 +318,6 @@ Compressed here to title, shipping version, evidence, and the sentences that sta | **R-487** | **A removed app whose backups were kept was listed on neither backup page (P2).** Closed in controller **v0.242.0** (`d698ce3`): the local lists are keyed on the DRIVES the way R-237 keyed the off-site list on the store — `ListRemovedAppUnits` walks `backups/primary/` on the system path and every connected registered drive; the Mentések page lists the unit after the deployed rows („Eltávolítva — visszaállítható", one action), the Visszaállítás picker lists it in its own group, `GET /api/backup/snapshots` answers for it, and the restore opens the unit where it sits (`primaryUnitDirFor` — a unit kept on a data drive was unreachable before, the fallback named the system path). Proven live on the scratch guest 9202 with an opengist throwaway: removed with data, backups kept → row + picker + API answered; „Visszaállítás a mentésből" reinstalled it running. Red-proofs: lister inert, picker 404, wrong unit dir, row not built, row not rendered, picker not rendered — all fail. `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN-LIVE** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` | | **R-490** | **The monitoring page's memory-distribution card never rendered — `/api/system/info` was 404 (P3).** Closed in controller **v0.242.0** (`d698ce3`): an exact-path mount ahead of the web layer's `/api/system/` prefix, and `systemInfo` reads the default storage path like every other reader of the empty global. Live on 9202: 200 with the drive figures. Red-proofs: mount removed, fallback removed — both fail. **The global's deletion stays deferred → R-492.** `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN-LIVE** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` | | **R-491** | **Removing an app left its update hold in the store, so a reinstall started held (P2).** Closed in controller **v0.242.0** (`d698ce3`): `removeStack` clears an UPDATE hold (`Settings.ClearUpdateHold`, never an R-379 restore hold), logged. Proven live on 9202: a held opengist removed → the store no longer carries the hold, the app redeployed without refusal. Red-proof: the removal without the clear fails the wiring test. `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN-LIVE** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` | +| **R-546** | **The first-hour guide and the reminder bar sent the household to create their recovery code before the box could (P2).** Closed in controller **v0.246.0** (`0fe315b`): the bar consults the agent's OWN preflight `ok` (all five blocking items, not a copy of `pbs_storage_id`), cached 60 s, probed only while paused, held back while not ready; `/backup/escrow` shows a waiting card that polls and reloads; `POST /api/escrow/start` refuses 409 before staging (the direct path that produced the raw `-storage` stderr); unknown readiness keeps the bar. The guide moves the step after the first apps: „amikor a sárga sáv megjelenik”. Red-proofs: bar held back, waiting card, start refusal; controls ready and unknown. **Proven by tests through ServeHTTP, NOT live** — no Tier-0 box is paused and agent-connected (**R-551**); chaos night measured live the ~17-minute red window and its self-heal. | **CLOSED 2026-09-17 — PROVEN (tests); live walk owed by R-551** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` | +| **R-549** | **The staleness alarm's budget was two report cycles, so one failed push spent all of it (P2).** Closed by operator ruling A (2026-09-17): `alerting.stale_threshold` 30 m → **45 m**, `node_down`/`host_down` at 90 m (`manifests/hub.yaml`, commit `06334e1`). The dashboard's customer status hardcoded 30 m / 1 h and would have disagreed with the alarms, so hub **v0.117.0** (`37ae31f`) makes `controllerStatus` read the same value — red-proof `TestControllerStatus_FollowsConfiguredThreshold` (report 40m old: status warn, want ok). **Proven live:** the running hub printed `node_stale after 45m0s, node_down after 1h30m0s` and `host_stale after 45m0s, host_down after 1h30m0s` at 08:22Z. **Reasoning kept:** the threshold is configuration and every reader — both checkers, host status, customer status — reads the one value; a dead box now pages 15 minutes later, a cost the ruling accepts. `audits/evidence-chaos-fixes-2026-09-17/partA-hub-45m.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` | +| **R-550** | **The restore record was in-memory only: after the machine stopped, nothing told the household their restore did not finish (P2).** Closed in controller **v0.246.0** (`0fe315b`) by operator ruling „fix” — a reversal of the in-memory design for the restore record only (cooldowns stay in memory): `restore-status.json` in DataDir, atomic at both ends of an op; a record still running at startup becomes a failed, interrupted result per app, shown on `/backups/restore` until that app's next restore, raised once as `restore_interrupted` (hub v0.117.0, household). Red-proofs: record across restart (`StartedAt:0001-01-01`), `main()` wiring (AST), startup helper, page card. **Proven live on demo-hp 9201:** a throwaway homebox restore killed 2 s in; after the supervisor's restart the status read `ok:false … megszakadt … interrupted:true`, the card showed, the event reached the hub (HTTP 200, stored under demo-hp); a second restore cleared the card. **Known gap filed: R-552** (a removed app keeps its notice). `audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 5e320dda..818c9312 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -728,11 +728,10 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-543** | **[P1-HIGH] Off-site ON by default is not off-site WORKING: on a fresh box tier 3 sits at „Kulcsletétre vár" until the household does the escrow ceremony, and nothing asks them to — while the tier-1 row now tells them their files are protected by that very copy.** MEASURED 2026-09-16 on the fresh box (controller 0.244.0, hub 0.116.0, off-site provisioned automatically by the new default): the app-backup page reads „3. mentés — Kulcsletétre vár · A távoli mentés a titkosítási kulcs letétbe helyezéséig szünetel", the remote page reads „Helyreállítási kód szükséges", and `POST /backup/offbox/run` returns 302 while producing no snapshot (the controller log shows only `offsite-credential-retry`, no restic activity). **Why it matters more than before today:** hub v0.116.0 makes off-site the default *because* a one-drive box otherwise keeps the household's files in no tier at all (R-537/R-538), and controller v0.244.0 now prints „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi" under the tier-1 row. On day one both are true-in-intent and false-in-fact: the copy is paused. **Fix shape (one of):** prompt the escrow ceremony as part of first-run when off-site is enabled and un-escrowed; and/or make the tier-1 sentence state the tier's actual state („…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel"). The ceremony itself works and is customer-facing („Helyreállítási kód létrehozása"); what is missing is that anyone is told to do it. **CLOSED 2026-09-16 — controller v0.245.0, both halves proven live.** The pause is untouched: it is the zero-knowledge escrow design, and this row was never about the mechanism. (a) **The household is asked:** while the off-site tier is configured and its escrow is not complete, every authenticated page carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking `/backup/escrow`. It is the R-241 bar, second instance — same session-cookie dismissal, back at the next visit, gone for good when escrowed; **no second banner system**. It hangs off `executeTemplate`, the single render choke point, so it cannot reach only the pages someone remembered. (b) **The tier-1 sentence renders by state:** `driveFilesNoteFor` takes `tier3State`'s own vocabulary — `active` → „védi", `escrow_pending` → „védené … a helyreállítási kód létrehozásáig szünetel" + the route, no off-site and no second drive → „nincs másolat" + both ways out. **Measured live on 0.245.0:** on a paused box (9202, off-site configured through the product's own endpoint, `escrow_state=pending`) the bar renders on /dashboard, /launcher, /backups/apps and /settings; a manual `POST /backup/offbox/run` is refused by the fork-4 gate („A távoli mentés a kulcs letétbe helyezésére vár.") with `last_run=None, snapshot_count=None`; and a throwaway class-A app's row reads „…védené … szünetel" with „védi"=0. On an escrowed box (9201) the bar is absent on all three pages and the row reads „védi". Both red-proofed (the bar test fails on BOTH pages with the one hook line removed; the sentence test quotes the exact v0.244.0 promise when the state is ignored). The first-hour guide now asks for the code right after the dashboard password and before the first app. Evidence: `audits/evidence-recovery-code-2026-09-16/`. | **CLOSED 2026-09-16 — shipped in controller v0.245.0 and proven live** | | **R-544** | **[P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody.** MEASURED 2026-09-16 during the teardown of the fresh box: an unacknowledged delete was correctly refused 409 („has key escrow (acknowledgement missing)") and the refusal text promises the acknowledgement „moves it to retained custody"; the acknowledged delete then logged `host deleted: tester-1-33b6a9 (escrow deleted: true)`. The hub's own customer page states the truth — „host deletion only demotes custody, never destroys it… recovery-key custody is demoted to retained custody, not destroyed", with the customer delete named as „the one true purge point". **Nothing is broken; the log is.** An operator reading that line during an incident would believe a household's last key had just been destroyed, and the R-304 retention exists precisely so it is not. **Fix shape:** log what happened — `escrow custody demoted to retained (host delete)` — and keep the boolean's name out of operator-facing text. | **READY — rank P3-LOW; owner: CC (hub)** | | **R-545** | **[P3-LOW] There is no product action that un-configures an off-site target — only one that re-starts an ORPHANED repository.** FOUND 2026-09-16 while exercising R-543's paused state on the scratch guest: `POST /backup/offbox/config` configures a target and can disable it (`enabled` unchecked), but nothing removes it. `POST /backup/offbox/reset` refuses unless `OffboxOrphaned()` is true („Az offsite tároló nincs elárvult állapotban."), and it means „start a new remote backup, set the old history aside" — not „forget this destination". So a household that sets up the wrong NAS, or a box being handed to someone else, keeps the host, user, path, ssh key and minted repo password on disk with no route to clear them; a disabled target still holds its secrets in `data/offbox/`. **Why P3 and not higher:** a disabled target runs nothing and the secrets are 0600 on the box's own disk, so nothing leaks and no copy is lost. **Fix shape:** a „Távoli cél törlése" action beside the config form that clears the target and shreds `data/offbox/`, REFUSING while the hub holds a sealed package for this box (the R-241 rule — dropping the key would orphan the history that package protects). Teardown for this session's proof had to clear it out-of-band for exactly this reason, which is the measurement. | **READY — rank P3-LOW; owner: CC** | -| **R-546** | **[P2-MEDIUM] The first-hour guide sends the household to create their recovery code at a moment when the box cannot yet do it — and the new reminder bar urges them there on every page.** MEASURED 2026-09-16/17 on a fresh box (`tester-1-022354`, guest 9201, controller 0.245.0, agent 0.131.0, installed from the published ISO 1.28.0). `VOLUNTEER-first-hour.md` §6 — added hours earlier in controller v0.245.0 — places „A helyreállítási kód" immediately after the dashboard password and **before the first app**, because until it is done the off-site copy does not run. At exactly that point the ceremony FAILS: `POST /api/escrow/start` → 200, then `GET /api/escrow/status` → `detail: "exit 2: … selftest=escrow-create requires -storage (or escrow.pbs_storage…"`, and `POST /api/escrow/claim` → **409** „A folyamat jelenlegi állapotában a kód nem kérhető le." **Cause, measured on both sides:** the hub had auto-provisioned the DR descriptor at 20:19 (no press — see R-534/R-511's acknowledged-delete path) and its Backup & DR panel itself read „descriptor provisioned … **waiting** · ceremony possible once the descriptor is applied on the box"; the box had no PBS storage (`pvesm status` = local + local-lvm only) and `/etc/felhom-agent/agent.json` had **no `escrow` section at all**. **It is a TIMING gap and it self-heals:** a watcher left the box alone and polled — `pbs_storage` and `escrow.pbs_storage_id` both became `felhom-pbs` at **20:35:16Z, ~17 minutes after the bind**; the retried ceremony then passed every preflight item and the claim returned 200 (83-character code, entropy 129.2 bits), and `escrow_state` flipped to `escrowed`. **Why it still matters:** for those ~17 minutes the R-543 reminder bar (also v0.245.0) is on *every* page telling the household to do the one thing that refuses, and nothing on the page says „wait a few minutes" — the volunteer meets a stderr fragment about a `-storage` flag. **Fix shape (one of):** have the escrow page/bar consult `preflight` and say „a doboz még készül — pár perc múlva próbáld újra" while `pbs_storage_id` is unset; or move the guide's step to after the first app; or make the bar appear only once preflight is green. **No product code was changed tonight** (validation run). Evidence: `audits/evidence-chaos-night-2026-09-17/phase0-escrow-failure.txt` and `phase0-escrow-retry.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller copy + guide timing)** | | **R-547** | **[P3-LOW] A disk that fills and empties between sweeps is never mentioned to anyone: `disk_critical` is defined at ≥95 % used, but the fill-watch runs once a day.** MEASURED 2026-09-17 (chaos night) on a fresh box (`tester-1-022354`, controller 0.245.0): the customer guest’s root filesystem was held at **96 % for ten minutes** (29 G used, 1.5 G free) and **no alarm of any kind fired** — checked twice, once by the round’s own runner and once independently after the fill was released. Cause, established from the ladder BEFORE the round rather than after: `fillwatch` runs **daily at 03:30** plus once ~90 s after a controller start, so a ten-minute window contains no check unless a restart lands inside it. The timing here was almost comic — the controller restarted at 21:28 after the previous round’s power cut, so its one opportunistic check ran about **twenty seconds before** the disk filled. Meanwhile all twelve apps kept serving and the background household loop logged 12 operations with 0 failures, so nothing else would have hinted at it either. **This is the ladder working as designed, not a missed alarm** — the fill-watch is a daily sweep, not a monitor. It is filed because the honest answer to „would the household be told their disk is full?” is **no, unless the controller happens to restart while it is full**, and that answer is not written down anywhere. **Fix shape (one of):** sample the fill more often than daily (a cheap `statfs` on the 5-minute health pass would do it); or say plainly in `08-alarm-ladder.md` that a transient full disk is out of scope. Evidence: `audits/evidence-chaos-night-2026-09-17/round-3.txt`. | **READY — rank P3-LOW; owner: CC** | | **R-548** | **[P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever.** MEASURED 2026-09-17 (chaos night) on `tester-1-022354`: a whole-guest backup wrote a **~29 GB source** (`mp0` = `local-lvm:vm-9201-disk-1`, 70 G provisioned, 40.58 % used, `backup=1`) into `pve-root`, which on a 32 GB system disk is **14 GB total with ~4.9 GB free**. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — **~16 MB/s, i.e. under four minutes to a full `/`** on the nested PVE. **The product’s behaviour is correct and legible throughout:** it failed the tier and said which one — `whole_guest_backup_failed` (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`), and the **off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all**. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can **never** succeed, and it keeps retrying on a backoff for ever, burning I/O and risking `/` each time. **Honest caveat:** the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. **Fix shape:** compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: `audits/evidence-chaos-night-2026-09-17/round-6.txt`. | **READY — rank P3-LOW; owner: CC** | -| **R-549** | **[P2-MEDIUM] The staleness alarm's budget is exactly two report cycles, so ONE failed push spends all of it.** MEASURED 2026-09-17 (chaos night, round 9) on `tester-1-022354` (controller 0.245.0): hub reachability was removed for ten minutes from the VM's side. The controller built its 23:08:42Z report, retried the push **three times over 1 m 40.8 s**, and gave up at 23:10:23Z (`hub push failed after 3 attempts`) - **31 seconds before the link returned**. Nothing is queued, and that is correct: a report is a snapshot, so the next one carries the same truth, and the controller says so itself ("backing off (the 15-min cycle still reconciles)"). **The arithmetic is the finding.** Last good report 22:53:43Z; next scheduled 23:23:42Z; measured cadence **15m0s**; `node_stale` trips at **30 minutes**. The gap is **29 m 59 s** - one second inside the threshold. So a single missed push spends the whole staleness budget, and any ordinary jitter pages the operator about a box that is healthy, serving every app, and has already repaired itself unaided. **The product behaved correctly throughout:** no false alarm fired, no app stopped, and both the hub link and the host-agent link recovered by themselves the moment the block lifted. What is filed is the margin, not a misbehaviour. **Fix shape:** either set the staleness threshold to a clear multiple of the cadence (three cycles, not two), or let a push that has failed all three attempts retry once off-cycle instead of waiting for the next scheduled report. **Honest caveat:** the cut was injected by the drill and also severed the controller from its host agent, which a real ISP outage would not do - but the report arithmetic above depends only on the hub being unreachable. Evidence: `audits/evidence-chaos-night-2026-09-17/round-9.txt`. | **READY - rank P2-MEDIUM; owner: CC** | -| **R-550** | **[P2-MEDIUM] The restore record is in-memory only: after the machine stops, the status surface is blank and nothing tells the household their restore did not finish.** MEASURED 2026-09-17 (chaos night, round 10) on `tester-1-022354` (controller 0.245.0): an app restore was accepted at 23:26:08Z (`302`, flash "Visszaallitas elindult") and the guest's host was hard-reset **four seconds later**, mid-write. Afterwards `/api/backup/restore-status` - the endpoint the restore page itself polls - answers `{"ok":true,"data":{"running":false,"started_at":"0001-01-01T00:00:00Z"}}`: the Go **zero value**, with **no `last` field at all**, while the page's own script renders " sikertelen." from `st.last.message` and " folyamatban" from `st.op`. So the UI has a last-operation branch with nothing to populate it after a restart. On disk, in the real data directory, no restore, lock or state file exists and **no file at all was modified in the reset window**. **The box recovered perfectly** - 26/26 containers back in 150 s, boot reconciliation naming the app it recovered, every front door serving, one true `controller_started` alarm and no false one. **CORRECTED 2026-09-17T00:24Z:** this row first claimed no restore surface existed at all, citing four endpoints that 404'd - all four were paths I GUESSED, and all four were wrong. The real routes (`/backups/restore`, `/stacks//backup`, `/api/backup/restore-status`) came from the controller's own rendered links. The corrected claim is narrower and stands on the endpoint's own answer. **Honest limit:** only four seconds elapsed, so `started_at` may be zero because the restore never truly began rather than because the reboot erased it; the pre-reset log is unrecoverable (the stream holds zero lines before the reboot). Either way nothing tells the customer. **Fix shape:** persist the last restore outcome the way the backup tiers already persist theirs, and have boot reconciliation mark an in-flight restore as abandoned so the page can say so. Evidence: `audits/evidence-chaos-night-2026-09-17/round-10.txt`. | **READY - rank P2-MEDIUM; owner: CC** | +| **R-551** | **[P3-LOW] No Tier-0 box can put the escrow ceremony in the state R-546 fixes — paused AND connected to its agent — so the readiness branches are proven only by tests.** FOUND 2026-09-17 while live-validating controller v0.246.0 (R-546). The branches (reminder bar held back while the agent's preflight is not `ok`; the waiting card on `/backup/escrow`; `POST /api/escrow/start` refused 409 before staging) need a box whose off-site tier is configured, whose escrow is NOT done, and whose controller reaches the agent. Measured on demo-hp: **9201** reaches the agent but is escrowed (the bar is off by design there; making it paused would be hand-set state on the standing demo box, and a real ceremony would supersede its live escrow — the hub keeps ONE `host_escrow` row per host, `host_id PRIMARY KEY`); **9202** is paused-capable but has **no local-API token at all** (its `bootstrap.json` holds only `schema`, `customer.id`, `disposition`), so its readiness is always UNKNOWN and the bar always shows. The fresh-bind window where this state occurs naturally (~17 min, chaos night Phase 0) needs a fresh install. **What IS proven:** five tests driving the real pages and handler through `ServeHTTP` with a fake agent, each red-proofed; and chaos night measured live that the agent's preflight is red for ~17 minutes after a bind and turns green by itself (`evidence-chaos-night-2026-09-17/phase0-escrow-*`). **Fix shape:** give the scratch guest a local-API token through the agent's own provisioning path, or walk R-546 on the next fresh install. | **READY - rank P3-LOW; owner: CC** | +| **R-552** | **[P3-LOW] An interrupted-restore notice for an app that is then REMOVED stays on the restore page for ever.** FOUND 2026-09-17 by CC reviewing its own controller v0.246.0 (R-550) during live validation. The per-app notice (`Manager.opInterrupted`, persisted in `restore-status.json`) is cleared in exactly one place — `BeginRestoreOp` for that app (`internal/backup/opstatus.go`) — and `removeStack` (`internal/api/router.go`) never touches the restore record. So a household that answers „A visszaállítás megszakadt … indítsd el újra" by REMOVING the app instead of restoring it keeps a „Megszakadt visszaállítás" card about an app that no longer exists. Measured shape, not hypothetical: on 9201 the notice cleared only when homebox was restored again (08:54:50Z, card count 0) — the teardown deliberately took that path before removing it. **Fix shape:** `removeStack` clears the app's notice (a `ClearInterruptedRestore(stack)` beside the existing update-hold clear, R-491's precedent), with a wiring test. Not fixed in v0.246.0: found after the release was built; one release per repo per session. | **READY - rank P3-LOW; owner: CC** | | **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** | | **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** | | **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** | diff --git a/documentation/runbooks/VOLUNTEER-first-hour.md b/documentation/runbooks/VOLUNTEER-first-hour.md index 3a395335..b58c472d 100644 --- a/documentation/runbooks/VOLUNTEER-first-hour.md +++ b/documentation/runbooks/VOLUNTEER-first-hour.md @@ -102,12 +102,27 @@ nem kell újra összekötni.)* **legalább 12 karakteres** jelszót. Ez lesz a vezérlőpult jelszava. 4. Ha nem jött meg a kód: **„Új kód kérése"** — mindig ugyanarra az e-mail címre érkezik. -## 6. A helyreállítási kód (~2 perc) — ezt ne hagyd ki +## 6. Az első két alkalmazás telepítése (~2 perc) + +1. A vezérlőpulton: **Alkalmazások**. Keresd meg például a **BookStack**-et (családi wiki) és a + **PrivateBin**-t (titkosított jegyzet), és nyomd meg a **Telepítés** gombot. +2. A telepítő oldalon csak az **aldomain** a kérdés — hagyd az alapértelmezettet (`wiki`, `paste`). + Az „Automatikusan generált értékek" részt nem kell felírnod. +3. **Telepítés indítása.** A BookStack kb. 1 perc, a PrivateBin kb. 20 másodperc. + +## 7. A helyreállítási kód (~2 perc) — ezt ne hagyd ki A doboz a fájljaidról **titkosított** másolatot küld a Felhom távoli tárhelyére. A titkosítás kulcsát **csak te** kapod meg: ez a **helyreállítási kód**. Amíg nem hozod létre, **a távoli -mentés nem indul el** — a vezérlőpult minden oldalán látszó sáv ezt írja: +mentés nem indul el**. + +**Amikor a sárga sáv megjelenik (a beállítás után néhány perccel),** akkor jött el az ideje. A doboz az +első perceket a távoli tárhely előkészítésével tölti, és addig a kódot még nem tudja létrehozni — ezért +a sáv csak akkor jelenik meg, amikor már lehet. A sáv ezt írja: „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." +Ha a sáv még nincs ott, telepítsd nyugodtan az első alkalmazásokat (előző lépés), és nézz vissza. +Ha mégis túl korán nyitod meg az oldalt, azt írja: „A doboz még készül — …pár perc múlva…", és +magától frissül. 1. **Biztonsági mentés → Távoli mentés**, vagy egyszerűen kattints a sávon a **„Helyreállítási kód létrehozása"** hivatkozásra. @@ -121,14 +136,6 @@ mentés nem indul el** — a vezérlőpult minden oldalán látszó sáv ezt ír Ha ezzel megvagy, a sáv eltűnik, és a távoli mentés magától elindul. -## 7. Az első két alkalmazás telepítése (~2 perc) - -1. A vezérlőpulton: **Alkalmazások**. Keresd meg például a **BookStack**-et (családi wiki) és a - **PrivateBin**-t (titkosított jegyzet), és nyomd meg a **Telepítés** gombot. -2. A telepítő oldalon csak az **aldomain** a kérdés — hagyd az alapértelmezettet (`wiki`, `paste`). - Az „Automatikusan generált értékek" részt nem kell felírnod. -3. **Telepítés indítása.** A BookStack kb. 1 perc, a PrivateBin kb. 20 másodperc. - ## 8. Első belépés az alkalmazásokba - **BookStack:** az alkalmazás oldalán az „Első lépések" rész a címet `wiki.DOMAIN` alakban írja —