Fixes before the first tester: decisions 50/51 outcomes, register 359->361 (R-719/720/721/722/727 closed, R-723 fixed, R-724/725 narrowed, R-728/729 opened), STATUS with the Tester-2 checklist, report
gates / gates (push) Successful in 26s

Secret scan before commit: 6162 files, 0 hits, control 1.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-30 11:01:08 +02:00
parent 7e50098088
commit 11591f3a93
13 changed files with 265 additions and 28 deletions
+5 -1
View File
@@ -20,7 +20,11 @@
> restore test takes only the box's own).** `09` §3. 50: when the customer has off-site, every newly installed app is
> in the off-site copy; over quota the page names the largest and the household chooses; history is never deleted to
> make room without that choice. 51: the three drill archives in `tester-1`'s ep0 namespace go, nothing else. Also: the
> first real tester is a NEW record `Tester-2`; the volunteer guide is rewritten as measured (R-722).
> first real tester is a NEW record `Tester-2`; the volunteer guide is rewritten as measured (R-722). **Outcome (same day):** controller v0.283.0/v0.283.1 (apps off-site by default + one press + size card; a Stop holds
> at the quiesce resume, the volume dump, the update leg and the crash recovery — v0.283.0 missed the dump's production
> adapter, caught live), agent v0.138.0 (the restore test skips an archive written with another key; signed delivery to
> both demo boxes), hub v0.126.0 („Új linket kérek" on an expired/used bind link; no day-one false mails), ep0's three
> drill archives removed. Floor 0.283.1. Report: `REPORT-fixes-first-tester-2026-09-30.md`.
> **2026-09-29 evening — operator rulings 48 (wanderer stays, A) and 49 ("close sign-up now", A).** Controller
> **v0.282.0**: `after_setup:` (the app's own sign-up switch, env merged + one `compose up -d`, when the gate opens or on
+79
View File
@@ -0,0 +1,79 @@
# REPORT — 2026-09-30: the fixes before the first real tester
Architecture read first: `07` §6 (the 2026-09-16 ruling), `03` (the restore test's selection), `05` (self-bind),
`04` §3.1 (signed delivery), `09` §3 decisions 45–49; rows R-719 … R-727, R-95, R-240, R-494, R-600, R-688.
Releases: **controller v0.283.0 + v0.283.1**, **agent v0.138.0**, **hub v0.126.0**. Evidence:
`documentation/audits/evidence-fixes-first-tester-2026-09-30/` (part0, partA, partC, partD, partE, release).
## Tester-2 — read-only checklist (nothing on Tester-2, Cloudflare, ep0 or the Storage Box was changed)
| # | item | state | where to fix |
|---|---|---|---|
| 1 | Tunnel token pasted | **done** — tunnel `3ce0eccd…` (Peti's original, reused) | — |
| 2 | DNS of `sajatfelhom.hu` | **done** — ONE record `*.sajatfelhom.hu` → that same tunnel; no leftover pointing elsewhere; today 530 (no connector, right with no box) | — |
| 3 | The tunnel's route `*.sajatfelhom.hu` → `https://traefik`, No TLS Verify | **UNKNOWN** — the record's Cloudflare key reads DNS only (`Authentication error` on the tunnel; stopped there) | Cloudflare → Zero Trust → Networks → Tunnels → this tunnel → Published application routes |
| 4 | Off-site | **done** — shared 100 GB, new sub-account 322460 (username `sub2` reused; Peti's was emptied and deleted 2026-09-25) | — |
| 5 | DR tier (ep0) | **done** — ON; nothing of Tester-2 or Peti on ep0 yet (made at the first connection) | — |
| 6 | The connect e-mail | **sent twice at 07:36 UTC** (the customer was created twice, R-728) — **only one of the two links works**; valid until 2026-10-07 07:36 UTC | Tell your friend: if one link says „expired", use the other; or press „Send self-bind link" once just before the install |
| 7 | E-mail language | **English** — mail, bind page and the box start in English | Hub → Tester-2 → Edit, if Hungarian is wanted |
| 8 | Owner passphrase | yours to hand over | in person / by phone |
| 9 | Customer id `Tester-2` has a capital letter | never walked before; no known break | note only |
## The Parts
| Part | state | note |
|---|---|---|
| 0 Tester-2 read-only | **done** | checklist above; items 3 and 6 need you |
| A1 measure the over-quota path | **done** | it deletes nothing extra: refuses new pushes, runs only the ruled retention; now pinned |
| A2 apps off-site by default + one press for older apps | **done** | controller v0.283.0; live on 9202 |
| A3 size warning | **done** | page card; unit + parity proven (a household NAS target has no quota to test live) |
| B guide + slips | **done / narrowed** | six stale lines rewritten (the sixth found on the way: auto-reboot); R-724/R-725 partly, the rest narrowed |
| C1 ep0 cleanup | **done** | three archives removed; every other namespace byte-identical |
| C2 agent v0.138.0 | **done** | signed delivery to both demo boxes (340 s); due-check normal on both |
| D fresh link for a returning customer | **changed** | the brief's trigger cannot be built (a registering box is unclaimed); built „Új linket kérek" on the expired/used page |
| E Stop holds | **done, after a live failure** | v0.283.0 was wrong in production (adapter); v0.283.1 fixed and proven live |
| F day-one mails | **done** | unit + red-proof; live proof at Tester-2's first hour |
## Claims in the brief that turned out wrong, named
- **"The over-quota path prunes history"** — **wrong.** Over the quota the box refuses new pushes and runs the SAME
retention as every night; with no new snapshots nothing extra ages out. No P1.
- **"Peti's Cloudflare records for `sajatfelhom.hu` may still be there"** — **there is exactly one record, and it
points at the tunnel Tester-2 now carries** (Peti's tunnel, reused). Nothing stale to remove.
- **"A PBS archive carries its box's key fingerprint or host id"** — **half right:** the key fingerprint yes (PVE
content `encrypted`), a host id no (the comment is only „felhom local-api", the owner is the customer's token).
- **"The hub sends no link at registration"** — **right, and it cannot:** the registration carries nothing of a
customer. The fix was changed to a button on the old link's page.
- **"Stop is lost at the backup's resume only"** — **wrong:** also at the nightly volume dump, the update leg and the
startup crash recovery; all four fixed.
- **Mine, from 2026-09-30:** "the ✗ names the wrong tier" — misread; the local-tier heading was the NEXT section.
## Red-proofs (each seen failing on its assertion, then restored)
RP31 new app not ON · RP32 earlier OFF overridden · RP33 hook not wired · RP34 over-quota forget differs · RP35 exact
quota "does not fit" · RP36 quiesce restarts a stopped app · RP37 dump restarts it · RP38 update leg presses it ·
RP39 (agent) an earlier box's archive picked · RP40 new box "recovered" · RP41 first-hour skip mailed · RP42 no fresh
link · RP43 the production adapter does not answer · RP44 crash recovery restarts it. Outputs:
`evidence-fixes-first-tester-2026-09-30/` and the scratchpad `rp/` copies.
## Slips of mine, said
- The first hub/evidence commit went out after my secret-scan script crashed (it looked for last night's shredded
files). Scanned right after: 5 125 files, 7 secrets, **0 hits**, control 1. Nothing leaked.
- v0.283.0 shipped the Stop fix un-wired in production; the live test caught it; v0.283.1 is the second controller
release this session (the one-release rule bent, reason in its CHANGELOG).
## Rows
Closed **R-719, R-720, R-721, R-722, R-727**. Fixed pending live proof **R-723**. Narrowed **R-724, R-725**. Opened
**R-728** (customer created twice), **R-729** (no way to remove an off-site target). **Register 359 → 361 rows.**
R-726 (a returning customer's orphaned repository) stays open — a NEW record does not meet it.
## Teardown
- **9202:** controller left on 0.283.1 (scratch; the floor does not reach it); the throwaway NAS target switched off
through the form and then removed from `settings.json` with the controller stopped (R-729: no product path);
`glance` removed with its data; paperless-ngx running; the off-site switches back OFF.
- **Demo boxes:** controller 0.283.1 by the floor, agent 0.138.0 by signed jobs — nothing else.
- **ep0:** only the three archives of decision 51. **Hub:** floor 0.283.1 (MinAgent 0.131.0 declared); hub 0.126.0.
- **Tester-2 / Cloudflare:** read only. DooPlex: pushes, builds, the hub deploy, signing.
+19 -15
View File
@@ -1,23 +1,27 @@
# STATUS — what works, what's broken, what's next
**Ready for a first real tester: YES, on a new customer record — a fresh box went from the download to restored apps with no help at all; fix the volunteer guide first, and decide whether their apps go off-site by themselves (today they do not).**
**Ready for the first real tester (Tester-2): yes, once you check two things — the tunnel's route in Cloudflare, and which of the two connect mails your friend uses.**
**Updated 2026-09-30 morning. Both demo boxes run controller 0.282.0 and host agent 0.137.0. Hub 0.125.0. New installs get golden 0.282.0 with agent 0.137.0 (baked and vouched last night).**
**Updated 2026-09-30. Both demo boxes run controller 0.283.1 and host agent 0.138.0. Hub 0.126.0. New installs get golden 0.282.0 with agent 0.137.0.**
**Decisions today** (yours, recorded): every new app goes off-site by itself (50); the three drill backups on ep0 go, and the restore test uses only the box's own backups (51).
**Tester-2 — what I checked, read only.**
- Done: tunnel key pasted; the domain has exactly one DNS record, and it points at that same tunnel; off-site on (100 GB, a fresh storage account); DR tier on; e-mail language English.
- **Check yourself:** the tunnel's route. My Cloudflare key could read DNS but not tunnels. It must say `*.sajatfelhom.hu` → `https://traefik`, with „No TLS Verify" ticked.
- **Two connect mails went out** (the customer was created twice). Only one link works. Tell your friend: if one says "expired", use the other. The live one is valid until 7 October.
**What I did, and it worked.**
- **A fresh box, walked like a volunteer, needed no help.** Download, install on the Hungarian keyboard, the e-mailed link, the dashboard through the internet, the recovery code, three apps, a family member, backup, restore, delete, a power cut, typos, a phone. Last walk needed one intervention; this one needed none.
- **The new safety parts held on a fresh box.** Strangers saw only the gate page. The gate opened by itself after the household's setup. Sign-up stayed closed. The random first password worked, and the old default password was refused.
- **Apps go off-site by themselves now.** Older apps get one button. If the apps do not fit in 100 GB, the page names the biggest. The box never deletes old copies to make room — I checked; it never did.
- **A Stop holds during a backup.** The first version was wrong in real use; my live test caught it, and the second version is proven live.
- **ep0 is clean:** the three old drill backups are gone; the other customers' backups are unchanged.
- **A returning customer gets a new link** from a button on the expired link's page.
- **The volunteer guide is current** (six stale lines fixed).
- **Two false operator mails on a new box's first day are gone.**
**What broke.**
- **No off-site copy on night one.** Every app starts with its off-site copy switched off, and nothing tells the household. On top of that, tester-1's old off-site store belonged to an earlier box, so the box skipped it all night. The household's data stayed safe on the box.
- **The night's restore test tried an old box's backup** and failed on its key. The page shows a bare ✗ and names the wrong place.
- **A Stop pressed during a backup was undone by that backup.**
- **The volunteer guide is out of date in five places.** For example, it still gives BookStack's old default password, which is now refused.
**Rows.** 9 opened, 1 closed. The list went from 350 to 359 rows.
**Rows.** 5 closed, 2 opened. The list went from 359 to 361 rows.
**What needs you.**
1. **Decide: should apps go off-site by themselves?** Option A: yes, every app is switched on when the customer has off-site (my pick — that is what the guide already promises). Option B: no, the guide tells the household to switch each app on. If you do nothing, a new household's apps have no off-site copy.
2. **Approve the guide fix** (five lines, I write them). If you do nothing, a volunteer tries BookStack's old password and fails.
3. **Three whole-machine backups of deleted drill boxes are still in tester-1's space on ep0** (two from 16 September, one from last night). I only looked; I did not touch them. The old ones make the restore test fail. If you do nothing, every new tester-1 box repeats that failure. Say "remove them" and I will do it with before/after checks.
4. **D4, the image copies, the Peti leftovers:** unchanged. If you do nothing, nothing changes.
1. **The Tester-2 tunnel route** (above). If you do nothing and it is missing, your friend's dashboard answers an error page.
2. **A golden before your friend installs?** The rule says bake one before any fresh install; this task did not include it. Option A (my pick): I bake 0.283.1 before the install — the box then starts on today's code. Option B: skip it — the box starts on 0.282.0 and updates itself to 0.283.1 within seconds (proven on both demo boxes today). If you do nothing, B happens.
3. **D4, the image copies:** unchanged. If you do nothing, nothing changes.
File diff suppressed because one or more lines are too long
@@ -553,12 +553,19 @@ R-636's louder repeated alarm.
which a per-app switch starting OFF had undone. If the apps will not fit the customer's quota, the page says so,
names the largest, and the household chooses which stay off-site. Apps already installed are not switched by a
release (the 2026-09-29 Part 0 rule); the backups page offers one press to switch them all on.
**The box never deletes off-site history to make room without the household's choice.** *(Outcome: filled in by
the 2026-09-30 fixes session.)*
**The box never deletes off-site history to make room without the household's choice.** **Outcome (2026-09-30,
controller v0.283.0):** measured first — over the quota the box already refused NEW pushes and ran only the ruled
retention, and an app whose files would cross the quota went up settings + database only; no history was ever
deleted to make room (now pinned). A fresh install switches the app ON (`DefaultOffboxOnForNewApp`; an earlier
choice is kept); older apps get one press on both backup pages; the size card names the three largest. Live on
9202: the press, and a fresh app joining by itself.
51. **ep0: the three whole-guest archives of deleted drill boxes in `tester-1`'s namespace are removed, and only
those** — *operator ruling 2026-09-30 (R-727).* With before/after controls on every other namespace. The restore
test is fixed to take only the current box's own archives, so an archive of an earlier box in the same namespace
can never be tested again. *(Outcome: filled in by the 2026-09-30 fixes session.)*
can never be tested again. **Outcome (2026-09-30):** the three archives (2026-09-16T17:27:32Z, 2026-09-16T21:59:54Z,
2026-09-29T19:37:07Z), mapped to their drill boxes by key fingerprint and host record, were forgotten on ep0 one by
one; every other namespace byte-identical before and after; chunks go at ep0's own weekly GC. Agent v0.138.0: an
archive carries its key FINGERPRINT (not a host id), and the restore test skips one written with another key.
Same day, operator: the first real tester gets a NEW customer record (`Tester-2`, domain `sajatfelhom.hu`), not
`tester-1`; and CC rewrites the volunteer guide as measured (R-722).
@@ -0,0 +1,3 @@
gitea.dooplex.hu/admin/felhom-controller:0.283.0
gitea.dooplex.hu/admin/felhom-controller:0.283.0 Up 25 seconds (healthy)
gitea.dooplex.hu/admin/felhom-controller:0.283.1 Up 25 seconds (healthy)
@@ -0,0 +1,31 @@
## 2026-09-30T08:40:33Z 9202: the household's own NAS target, through the page's form (192.0.2.1 = RFC 5737 documentation address, never reachable; throwaway key)
HTTP/2 302
location: /backups/remote?flash_error=flash.offbox.target_saved_escrow_agent_down
## 2026-09-30T08:41:15Z before: paperless-ngx=OFF privatebin=OFF
/backups/apps shows the offer: 1
press „Igen, mindegyikre“ → 302 https://192.168.0.114/backups/remote?flash=flash.offbox.enabled_all
after: paperless-ngx=ON privatebin=ON
offer still shown: 0
## 08:41:27Z fresh install of glance on 9202, fields ['SUBDOMAIN']
deploy → {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
state after 18s: running
off-site switches now: glance=ON paperless-ngx=ON privatebin=ON
2026/09/30 08:41:30 [INFO] [backup] glance: off-site copy switched ON at install (decision 50)
## 2026-09-30T08:55:05Z 9202 teardown
paperless-ngx Start → {"ok":true,"data":{"state":"starting"},"message":"Stack paperless-ngx start requ
glance Stop → {"ok":true,"message":"Stack glance stop completed"}
glance remove (data + backups) → {"ok":true,"data":{"removed":"glance","volumes_removed":["glance_glance_config"],"hdd_paths_removed":[],"hdd_paths_preserved":[],"hdd_note":"Az alkalmazás nem
off-site OFF paperless-ngx → 302
off-site OFF privatebin → 302
target switched off via the form → 302 https://192.168.0.114/backups/remote?flash=flash.offbox.target_saved
data=
Stopping 'felhom-controller-bootstrap.service', but its triggering units are still active:
felhom-controller-bootstrap.path
Traceback (most recent call last):
File "<stdin>", line 2, in <module>
FileNotFoundError: [Errno 2] No such file or directory: '/settings.json'
gitea.dooplex.hu/admin/felhom-controller:0.283.1 Up 27 seconds (healthy)
/backups/remote now: 0 × „Még nincs beállítva távoli mentési cél“
file=/var/lib/docker/volumes/felhom-controller-data/_data/data/settings.json
gitea.dooplex.hu/admin/felhom-controller:0.283.1 Up 30 seconds (healthy)
/backups/remote: „Még nincs beállítva távoli mentési cél“ × 1; paperless: "state":"running"
@@ -0,0 +1,55 @@
== demo-hp
felhom-agent 0.138.0
Sep 30 10:40:28 demo-hp felhom-agent[2355282]: time=2026-09-30T10:40:28.593+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=11d75b6842043927 op=agent_update
Sep 30 10:40:30 demo-hp felhom-agent[2355282]: time=2026-09-30T10:40:30.718+02:00 level=INFO msg="backup: restore-test scheduler shutting down" reason="context canceled"
Sep 30 10:40:31 demo-hp felhom-agent[2135127]: time=2026-09-30T10:40:31.715+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
Sep 30 10:40:32 demo-hp felhom-agent[2135127]: time=2026-09-30T10:40:32.658+02:00 level=INFO msg="janitor: starting (restore-test scratch retry + stale-lock sweep)" interval=10m0s
Sep 30 10:41:32 demo-hp felhom-agent[2135127]: time=2026-09-30T10:41:32.681+02:00 level=WARN msg="selfupdate: update committed" version=0.138.0 wrapper=""
== felhom-pve
felhom-agent 0.138.0
Sep 30 10:37:12 demo-felhom felhom-agent[2880126]: time=2026-09-30T10:37:12.223+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=ccafbfb170d1b714 op=agent_update
Sep 30 10:37:14 demo-felhom felhom-agent[2081460]: time=2026-09-30T10:37:14.960+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
Sep 30 10:37:16 demo-felhom felhom-agent[2081460]: time=2026-09-30T10:37:16.219+02:00 level=INFO msg="janitor: starting (restore-test scratch retry + stale-lock sweep)" interval=10m0s
Sep 30 10:38:16 demo-felhom felhom-agent[2081460]: time=2026-09-30T10:38:16.243+02:00 level=WARN msg="selfupdate: update committed" version=0.138.0 wrapper=""
== demo-hp 08:58:03 — selftest restore-test-due (READ-ONLY: the scheduler's own verdict)
* demo-hp status=online fp=07:B5:72:5D:5E:C1…
[ ok ] node status up 953h13m36s, load [0.60 0.67 0.58], mem 5.2GiB/29.3GiB, root 23.0GiB/38.6GiB
[ ok ] list lxc 1 guest(s)
- 9201 "demo-hp" status=running
[ ok ] pool read pool "felhom", 1 member(s)
- 9201 type=lxc
[ ok ] storage 4 store(s)
- local-lvm type=lvmthin content=images,rootdir used=32.4GiB/53.9GiB
- local type=dir content=vztmpl,backup,import,iso used=23.0GiB/38.6GiB
- felhom-pbs type=pbs content=backup used=0.0GiB/0.0GiB
- nvme-scratch type=dir content=images,rootdir used=73.5GiB/937.8GiB
=== selftest OK ===
== felhom-pve 08:58:04 — selftest restore-test-due (READ-ONLY: the scheduler's own verdict)
* demo-felhom status=online fp=60:8F:4C:50:C3:8E…
[ ok ] node status up 1225h32m30s, load [0.29 0.30 0.26], mem 2.4GiB/15.4GiB, root 27.4GiB/93.9GiB
[ ok ] list lxc 1 guest(s)
- 9201 "demo-felhom" status=running
[ ok ] pool read pool "felhom", 1 member(s)
- 9201 type=lxc
[ ok ] storage 4 store(s)
- felhom-backup type=dir content=backup used=11.9GiB/915.8GiB
- local type=dir content=import,backup,iso,vztmpl used=27.4GiB/93.9GiB
- local-lvm type=lvmthin content=rootdir,images used=11.9GiB/348.8GiB
- felhom-pbs type=pbs content=backup used=0.0GiB/0.0GiB
=== selftest OK ===
== demo-hp 08:58:13 — selftest=restore-test-due (READ-ONLY)
time=2026-09-30T10:58:13.922+02:00 level=INFO msg="backup: restore-test skips an entry that is not a backup of a guest" target=local volid=local:backup/felhom-golden-0.236.0.tar.zst size_bytes=654115664 reason="not a backup of a guest (the storage reports no v
time=2026-09-30T10:58:13.924+02:00 level=INFO msg="backup: restore-test skips an entry that is not a backup of a guest" target=local volid=local:backup/vzdump-lxc-9100-2026_08_21-17_59_15.tar.zst size_bytes=656970239 reason="guest 9100 does not exist on this n
tier=felhom-pbs due=false archive="felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z" landed=2026-09-24T20:06:25Z proven="felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z"
reason: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven
tier=local due=false archive="local:backup/vzdump-lxc-9201-2026_09_29-04_42_11.tar.zst" landed=2026-09-29T02:42:11Z proven="local:backup/vzdump-lxc-9201-2026_09_29-04_42_11.tar.zst"
reason: newest settled archive (landed 2026-09-29T02:42:11Z) is already proven
cost tier=felhom-pbs one_lookup=322ms
cost tier=local one_lookup=45ms
== felhom-pve 08:58:14 — selftest=restore-test-due (READ-ONLY)
tier=felhom-backup due=true archive="felhom-backup:backup/vzdump-lxc-9201-2026_09_29-07_36_43.tar.zst" landed=2026-09-29T05:36:43Z proven="felhom-backup:backup/vzdump-lxc-9201-2026_09_28-07_36_39.tar.zst"
reason: newest settled archive (landed 2026-09-29T05:36:43Z) has not been proven (last proven archive was a different one)
tier=felhom-pbs due=false archive="felhom-pbs:backup/ct/9201/2026-09-29T04:16:43Z" landed=2026-09-29T04:16:43Z proven="felhom-pbs:backup/ct/9201/2026-09-29T04:16:43Z"
reason: newest settled archive (landed 2026-09-29T04:16:43Z) is already proven
cost tier=felhom-backup one_lookup=20ms
cost tier=felhom-pbs one_lookup=306ms
@@ -0,0 +1,5 @@
## 2026-09-30T08:38:33Z live, hub 0.126.0 — a made-up link (never real; same page as an expired one by design)
GET /bind/<made-up> → 200; page: Felhom — Doboz összekötése | Felhom | doboz | összekötése | Ez a hivatkozás érvénytelen vagy lejárt. | A hivatkozás 7 napig érvényes. Ha lejárt, kérj újat az ügyfélszolgálattól, vagy az összekötést az üzemeltető is elvégezheti. | Ha ez a te linked volt, és a dobozod még nincs összekötve, új linket küldünk arra az e-mail címre, amelyet a Felhomnál megadtál. | Új linket kérek | Felhom.eu
the button's form: <form method="POST" action="/bind/<token>/resend">
POST …/resend → 200; page: Felhom — Doboz összekötése | Felhom | doboz | összekötése | Kész. | Ha ez egy valódi hivatkozás volt, és a dobozod még nincs összekötve, néhány percen belül új e-mailt kapsz a regisztrált címedre. Ha nem jön, szólj az ügyfélszolgálatnak. | Felhom.eu
hub log (self-bind lines, last minute): 0
@@ -0,0 +1,29 @@
## 2026-09-30T08:42:08Z 9202: „Mentés most“, and the household presses Stop on paperless-ngx while the dump holds it down
paperless-ngx before: "state":"running"
„Mentés most“ → {"ok":true,"message":"Mentés elindítva"}
08:42:13.593 saw: 2026/09/30 08:42:13 backup.go:955: [INFO] [backup] Stopping paperless-ngx for safe volume
08:42:13.600 Stop pressed → {"ok":true,"message":"Stack paperless-ngx stop completed"}
backup run: "success":true
paperless-ngx after: "state":"starting"
2026/09/30 08:42:13 backup.go:955: [INFO] [backup] Stopping paperless-ngx for safe volume dump
2026/09/30 08:42:13 manager.go:1312: [INFO] [stacks] Stopping stack: paperless-ngx
2026/09/30 08:42:13 router.go:601: [INFO] [api] stop requested for stack: paperless-ngx
2026/09/30 08:42:13 desiredstate.go:81: [INFO] [stacks] desired state for paperless-ngx recorded as "stopped" (was "running")
2026/09/30 08:42:13 manager.go:1312: [INFO] [stacks] Stopping stack: paperless-ngx
2026/09/30 08:42:21 backup.go:971: [INFO] [backup] Restarting paperless-ngx after volume dump
2026/09/30 08:42:21 manager.go:1223: [INFO] [stacks] Starting stack: paperless-ngx
## 2026-09-30T08:54:15Z RETEST on v0.283.1 — „Mentés most“, Stop on paperless-ngx while the dump holds it
paperless-ngx before: "state":"running"
„Mentés most“ → {"ok":true,"message":"Mentés elindítva"}
08:54:19.823 saw: 2026/09/30 08:54:19 backup.go:955: [INFO] [backup] Stopping paperless-ngx for safe volume
08:54:19.828 Stop pressed → {"ok":true,"message":"Stack paperless-ngx stop completed"}
backup run: "success":true
paperless-ngx 20 s after the run: "state":"stopped"
2026/09/30 08:54:14 desiredstate.go:81: [INFO] [stacks] desired state for paperless-ngx recorded as "running" (was "stopped")
2026/09/30 08:54:14 manager.go:1223: [INFO] [stacks] Starting stack: paperless-ngx
2026/09/30 08:54:19 backup.go:955: [INFO] [backup] Stopping paperless-ngx for safe volume dump
2026/09/30 08:54:19 manager.go:1312: [INFO] [stacks] Stopping stack: paperless-ngx
2026/09/30 08:54:19 router.go:601: [INFO] [api] stop requested for stack: paperless-ngx
2026/09/30 08:54:19 desiredstate.go:81: [INFO] [stacks] desired state for paperless-ngx recorded as "stopped" (was "running")
2026/09/30 08:54:19 manager.go:1312: [INFO] [stacks] Stopping stack: paperless-ngx
2026/09/30 08:54:28 backup.go:967: [INFO] [backup] paperless-ngx NOT restarted after the volume dump: the household stopped it meanwhile
@@ -2,3 +2,6 @@
HTTP/1.1 303 See Other
Location: /configuration?flash=floor_set
2026/09/30 10:33:19 [INFO] Global controller-version floor set to "0.283.0" (declared MinAgent "0.131.0")
## 2026-09-30T08:53:29Z floor → 0.283.1 (MinAgent 0.131.0 declared)
Location: /configuration?flash=floor_set
2026/09/30 10:53:29 [INFO] Global controller-version floor set to "0.283.1" (declared MinAgent "0.131.0")
@@ -0,0 +1,15 @@
## 2026-09-30T08:37:05Z hub 0.126.0 via ArgoCD (HEAD 7e5009808856179dea9cafcde472cb462943ae6f)
revision seen: 7e5009808856179dea9cafcde472cb462943ae6f
deployment "hub" successfully rolled out
image: gitea.dooplex.hu/admin/felhom-hub:0.125.0
argocd: sync=OutOfSync health=Healthy rev=7e5009808856179dea9cafcde472cb462943ae6f
2026/09/30 10:34:43 [INFO] enqueued signed-op job ccafbfb170d1b714 for host demo-felhom-8363b5 (775 bytes)
2026/09/30 10:35:14 [INFO] PBS-DR box refreshed: 18.7% full (18.3 GB of 97.9 GB)
2026/09/30 10:37:11 [INFO] host-report from demo-felhom-8363b5 (1 guests, 4 storage targets, 2 backups, 2 restore-tests, 2 pbs-snapshots, 14458 bytes)
2026/09/30 10:37:11 [INFO] DR-recipe host-half stored for customer demo-felhom (host demo-felhom-8363b5, v1)
2026/09/30 10:37:12 [INFO] host demo-felhom-8363b5 cleared signed-op job ccafbfb170d1b714 (executed or rejected)
08:38:00 image=gitea.dooplex.hu/admin/felhom-hub:0.126.0 argocd=Synced/Healthy
running pod image: gitea.dooplex.hu/admin/felhom-hub:0.126.0
2026/09/30 10:37:19 [INFO] felhom-hub 0.126.0 starting
2026/09/30 10:37:19 [INFO] Default controller-version floor: 0.120.0
2026/09/30 10:37:20 [INFO] Gitea artifact browser enabled (Day-0 version dropdowns) via http://gitea.gitea-system.svc.cluster.local:3000
+10 -8
View File
@@ -830,15 +830,17 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-716** | **[P3-LOW] Apps installed before controller 0.281.0 keep their open sign-up — decision 47 closes it only on apps whose gate the box opened.** READ 2026-09-29 on the demo boxes after catalog `6faf432` synced: demo-hp's adventurelog and opengist, demo-felhom's opengist carry `signup_block:` in their synced template and no gate record, so no block (`audits/gate-rollout-2026-09-29/0/P0-3-demo-boxes-after-push.txt`). This is Part 0's rule working as designed (a catalog change never touches an installed app). **Needs an operator word** before anything changes on an installed app: a one-time "close sign-up now" press on the app page for an installed app, or leave them. Only the demo boxes have such installs today. **Operator ruled A (decision 49); built in controller v0.282.0 and pressed** on demo-hp's adventurelog and opengist and demo-felhom's opengist: before, sign-up served; after, refused; adventurelog's own switch on (`audits/signup-lock-2026-09-29/`C). | **CLOSED — 2026-09-29** |
| **R-717** | **[P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone.** MEASURED 2026-09-29: opengist `disable-signup` is an admin-panel setting (no env, no CLI); wishlist `system_config.enableSignup` (Prisma). Their blocks are case-insensitive and refused every trick shape (`audits/signup-lock-2026-09-29/B/`). **Fix direction:** an `after_setup` command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). | **OPEN — P3; owner: CC** |
| **R-718** | **[P3-LOW] "Close sign-up now" restarts an app with its own switch, and the card does not say so.** MEASURED 2026-09-29 on demo-hp: pressing it on adventurelog recreated its backend (~30 s, one 500 on its login page). The window's card says the app restarts; the close card does not. **Fix direction:** the close card and the gate-open moment say "the app restarts once" where `after_setup.env` exists. **ALSO MEASURED 2026-09-29 (new-household drill):** the gate-open press on a fresh vikunja restarted it for its own switch — the front door answered 404 for ~2 s and nothing said so. | **OPEN — P3; owner: CC** |
| **R-719** | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. | **READY — rank P2-MEDIUM; owner: operator (which shape) / CC** |
| **R-720** | **[P2-MEDIUM] A new household's apps are not in the off-site copy: every app starts with „3. mentés Kikapcsolva", and nothing the household is told says to switch it on.** MEASURED 2026-09-29 on a fresh box (customer with off-site ON, the default since 2026-09-16): `/backups/remote` read „Aktív — nincs kijelölt alkalmazás"; `/backups/apps` read „3. mentés Kikapcsolva — Ez az alkalmazás nincs kijelölve távoli mentésre" for all three apps. From source, an app is off-site only after `POST /backup/offbox/toggle` (`settings.SetAppOffbox`); no deploy or claim path sets it. The volunteer guide §7 says that after the recovery code „a távoli mentés magától elindul" — it starts, and copies nothing. So a household following the guide has **no off-site copy of its apps on night one**, on a one-drive box where the whole-guest tiers do not carry the data drive (`07` §6). The drill pressed „Bekapcsolás" for BookStack only, as a household reading the page might, to measure both paths on the night. R-240 (the empty run's wording) is the same gap seen from the other end. **Needs an operator decision** (it changes what the product promises): apps default to off-site ON when the customer has off-site, or the guide adds the step. | **READY — rank P2-MEDIUM; owner: operator (decide) / CC** |
| **R-721** | **[P2-MEDIUM] The household presses Stop during a whole-guest backup, and the backup starts the app again.** MEASURED 2026-09-29 19:37 UTC on a fresh box: the first off-site whole-guest backup quiesced three apps at 19:37:01; the household pressed „Leállítás" on actualbudget at 19:37:16 and the controller recorded `desired state for actualbudget recorded as "stopped"`; the same second the backup's early resume ran `unquiescing … restarting 3 stack(s)` and `Starting stack: actualbudget`. The app page then read „Fut" while `app.yaml` kept `desired_state: stopped` (so the next reboot would stop it). The household had to press Stop again, and the removal was refused „Az alkalmazáson mentés vagy visszaállítás fut" for ~4 min (honest). Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase1/step10-stop-undone-by-quiesce.log`. **Fix direction:** unquiesce restarts only stacks whose desired state is still running; a test pins it (Stop during a quiesce → still stopped after the resume). | **READY — rank P2-MEDIUM; owner: CC** |
| **R-722** | **[P2-MEDIUM] The volunteer guide is stale in five places a volunteer reads literally.** MEASURED 2026-09-29 walking `runbooks/VOLUNTEER-first-hour.md` on golden 0.282.0: (1) §8 says BookStack's login is `admin@admin.com / password` — since the random first password (controller v0.280.0) that login is REFUSED; the app page says the right thing („a Beállítások oldalon látható első jelszó"). (2) §7 says the recovery-code bar comes „néhány perccel" after setup — measured ~15 min after enrolment (the off-site tier arrives with the agent's next 15-minute host report). (3) Nothing says to switch each app's off-site copy on (R-720). (4) §2 says a 2 GB USB stick and Rufus; the download page says at least 4 GB and Balena Etcher. (5) The operator part says no button is needed for the link (R-719), and §5 names the mail „Elindult a Felhom szervered" while a customer with an earlier box gets „Új beállító kód — újratelepült a szervered … A korábbi jelszavad már nem érvényes". Also, from the screens: the installer pre-selects `/dev/sda` (the guide says it never chooses), and its Summary screen rests on „Previous", not „Install". **Fix:** a guide edit; the operator approves the text. | **READY — rank P2-MEDIUM; owner: CC (text) / operator (approve)** |
| **R-723** | **[P3-LOW] A fresh box sends the operator two mails on day one that describe nothing wrong.** MEASURED 2026-09-29: `Operator email sent for tester-1/node_recovered` 2 s after the new box's first controller report (the customer's previous box had been silent 12 days — the new box is not a recovery); and `backup_tier_skipped (warning) — Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed)` → operator mail at 19:27 UTC, 7 min after enrolment, because the first whole-guest run fired before the off-site tier's descriptor arrived (applied ~19:35 UTC, backed up fine at 19:37). Operator-only, so no household is alarmed, but an operator learns to ignore both. **Fix direction:** `node_recovered` not for a host enrolled < N min ago; the skipped tier inside the first hour after enrolment is `info`, not a mailed warning. | **READY — rank P3-LOW; owner: CC** |
| **R-724** | **[P3-LOW] The status pages disagree with each other in small ways a household notices.** MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw `2026-09-29T19:22:30Z`; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). | **READY — rank P3-LOW; owner: CC** |
| **R-725** | **[P3-LOW] Small copy slips on the first-hour path.** MEASURED 2026-09-29: the self-bind PAGE says the passphrase was received „a beállításkor" while the mail and console say „a Felhom üzemeltetőjétől" (R-497 unified the mail and console, not the page); the recovery-code wizard addresses the household formally („Írja fel", „adja meg") while every other screen says „te"; the console's linked banner ends „a doboz össze van kötve. V" (a stray glyph); a gated app answers a phone app's API call with English JSON „this app is waiting for its first setup" (the browser gets the Hungarian gate page). | **READY — rank P3-LOW; owner: CC** |
| **R-719** | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** |
| **R-720** | **[P2-MEDIUM] A new household's apps are not in the off-site copy: every app starts with „3. mentés Kikapcsolva", and nothing the household is told says to switch it on.** MEASURED 2026-09-29 on a fresh box (customer with off-site ON, the default since 2026-09-16): `/backups/remote` read „Aktív — nincs kijelölt alkalmazás"; `/backups/apps` read „3. mentés Kikapcsolva — Ez az alkalmazás nincs kijelölve távoli mentésre" for all three apps. From source, an app is off-site only after `POST /backup/offbox/toggle` (`settings.SetAppOffbox`); no deploy or claim path sets it. The volunteer guide §7 says that after the recovery code „a távoli mentés magától elindul" — it starts, and copies nothing. So a household following the guide has **no off-site copy of its apps on night one**, on a one-drive box where the whole-guest tiers do not carry the data drive (`07` §6). The drill pressed „Bekapcsolás" for BookStack only, as a household reading the page might, to measure both paths on the night. R-240 (the empty run's wording) is the same gap seen from the other end. **Needs an operator decision** (it changes what the product promises): apps default to off-site ON when the customer has off-site, or the guide adds the step. **BUILT 2026-09-30 (controller v0.283.0, decision 50):** a fresh install on a box with off-site switches the app ON; older apps get one press on both backup pages; over the quota the page names the largest apps. Measured first: over the quota nothing but the ruled retention runs — no history is deleted (now pinned). Live on 9202: the press put both older apps ON; a fresh `glance` joined by itself. The size card is unit + parity proven (a household NAS target has no quota). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partA/`. | **CLOSED 2026-09-30 — controller v0.283.0** |
| **R-721** | **[P2-MEDIUM] The household presses Stop during a whole-guest backup, and the backup starts the app again.** MEASURED 2026-09-29 19:37 UTC on a fresh box: the first off-site whole-guest backup quiesced three apps at 19:37:01; the household pressed „Leállítás" on actualbudget at 19:37:16 and the controller recorded `desired state for actualbudget recorded as "stopped"`; the same second the backup's early resume ran `unquiescing … restarting 3 stack(s)` and `Starting stack: actualbudget`. The app page then read „Fut" while `app.yaml` kept `desired_state: stopped` (so the next reboot would stop it). The household had to press Stop again, and the removal was refused „Az alkalmazáson mentés vagy visszaállítás fut" for ~4 min (honest). Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase1/step10-stop-undone-by-quiesce.log`. **Fix direction:** unquiesce restarts only stacks whose desired state is still running; a test pins it (Stop during a quiesce → still stopped after the resume). **FIXED 2026-09-30 (controller v0.283.0 + v0.283.1):** the quiesce resume, the nightly volume dump, the update leg and the startup crash recovery skip an app the household stopped. **v0.283.0 was wrong in production** — its adapter did not answer the question and the dump restarted the app 8 s after the Stop (measured live on 9202); v0.283.1 wires it and pins the PRODUCTION types. Live on 0.283.1: `paperless-ngx NOT restarted after the volume dump: the household stopped it meanwhile`. Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partE/`. | **CLOSED 2026-09-30 — controller v0.283.1, proven live** |
| **R-722** | **[P2-MEDIUM] The volunteer guide is stale in five places a volunteer reads literally.** MEASURED 2026-09-29 walking `runbooks/VOLUNTEER-first-hour.md` on golden 0.282.0: (1) §8 says BookStack's login is `admin@admin.com / password` — since the random first password (controller v0.280.0) that login is REFUSED; the app page says the right thing („a Beállítások oldalon látható első jelszó"). (2) §7 says the recovery-code bar comes „néhány perccel" after setup — measured ~15 min after enrolment (the off-site tier arrives with the agent's next 15-minute host report). (3) Nothing says to switch each app's off-site copy on (R-720). (4) §2 says a 2 GB USB stick and Rufus; the download page says at least 4 GB and Balena Etcher. (5) The operator part says no button is needed for the link (R-719), and §5 names the mail „Elindult a Felhom szervered" while a customer with an earlier box gets „Új beállító kód — újratelepült a szervered … A korábbi jelszavad már nem érvényes". Also, from the screens: the installer pre-selects `/dev/sda` (the guide says it never chooses), and its Summary screen rests on „Previous", not „Install". **Fix:** a guide edit; the operator approves the text. **REWRITTEN 2026-09-30** (operator: CC rewrites as measured): both guides — BookStack's generated password on the app page, the ~15 min recovery-code delay, apps off-site by default + the one-press offer, a 4 GB stick, the returning customer's mail and the expired-link button, the pre-selected disk, focus on „Previous", and the auto-reboot (a sixth stale line: with the defaults the „reboot now?" prompt never appears). Operator part: set the e-mail language, check the domain's DNS for an earlier box's records, the tunnel route with No TLS Verify. Every changed line dated. | **CLOSED 2026-09-30** |
| **R-723** | **[P3-LOW] A fresh box sends the operator two mails on day one that describe nothing wrong.** MEASURED 2026-09-29: `Operator email sent for tester-1/node_recovered` 2 s after the new box's first controller report (the customer's previous box had been silent 12 days — the new box is not a recovery); and `backup_tier_skipped (warning) — Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed)` → operator mail at 19:27 UTC, 7 min after enrolment, because the first whole-guest run fired before the off-site tier's descriptor arrived (applied ~19:35 UTC, backed up fine at 19:37). Operator-only, so no household is alarmed, but an operator learns to ignore both. **Fix direction:** `node_recovered` not for a host enrolled < N min ago; the skipped tier inside the first hour after enrolment is `info`, not a mailed warning. **FIXED 2026-09-30 (hub v0.126.0):** no `node_recovered` when the customer's host was enrolled after the outage began; `backup_tier_skipped` in a box's first hour recorded, not mailed. Real recoveries and old boxes' skips still mail (controls). Unit + red-proof (RP40, RP41); the live proof is Tester-2's first hour. | **FIXED — hub v0.126.0; live proof at the first real install; owner: CC** |
| **R-724** | **[P3-LOW] The status pages disagree with each other in small ways a household notices.** MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw `2026-09-29T19:22:30Z`; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). **FIXED 2026-09-30 (controller v0.283.0):** Beállítások shows the real schedule and a local time; the dashboard's last backup is local time (R-500); „az előző N órája készült" replaces „0 órája". **NARROWED — remaining:** „Helyi cím (LAN)" / „Átjáró" read through the file-sharing container, so a box without it says „nem állapítható meg" — not a text fix (the read needs another path). | **NARROWED — the LAN/gateway read only; owner: CC** |
| **R-725** | **[P3-LOW] Small copy slips on the first-hour path.** MEASURED 2026-09-29: the self-bind PAGE says the passphrase was received „a beállításkor" while the mail and console say „a Felhom üzemeltetőjétől" (R-497 unified the mail and console, not the page); the recovery-code wizard addresses the household formally („Írja fel", „adja meg") while every other screen says „te"; the console's linked banner ends „a doboz össze van kötve. V" (a stray glyph); a gated app answers a phone app's API call with English JSON „this app is waiting for its first setup" (the browser gets the Hungarian gate page). **FIXED 2026-09-30:** the recovery wizard speaks „te" (controller v0.283.0; formal ceiling 18 → 17); the bind page says „a Felhom üzemeltetőjétől kaptál" (hub v0.126.0). **NARROWED — remaining:** the console's stray „V" (the installer/agent's banner, not these repos' text); the gate's English JSON to a phone app (the app shows its own error; left, deliberately); and the expired bind page still says „kérj újat az ügyfélszolgálattól" ABOVE the new „Új linket kérek" button (hub copy, next hub release). | **NARROWED — three small copy items; owner: CC** |
| **R-726** | **[P2-MEDIUM] A new box for a customer who had one before makes NO off-site copy on night one: the old repository is found orphaned, and the fix is a button nobody pointed the household to.** MEASURED 2026-09-30 00:15 UTC on the new-household drill box (`tester-1`, whose previous box was deleted 2026-09-17): `[offbox] offsite repo ORPHANED — remote holds backups written under a previous, no-longer-available key; runs will skip until reset` → `offbox_repo_orphaned` (warning) to the household's timeline and an operator mail; `offsite-integrity` then checked nothing. The household's page is honest („A távoli tároló másik kulccsal készült mentéseket tartalmaz … Új távoli mentés indítása…", old history set aside, never deleted), but the evening before, the recovery-code ceremony and the off-site page raised nothing, and the guide does not mention it. The hub re-issued the off-site credentials on re-enroll by itself; it could have known the repository would orphan. Customer data was never at risk (the old repository is untouched); the household simply has no off-site copy until someone presses the button. **Fix direction:** offer the reset at the recovery-code ceremony when the repository already holds another key's snapshots, or the hub's re-enroll re-issue sets the old history aside the same way (it is the same move-aside), and the guide says so. | **READY — rank P2-MEDIUM; owner: CC (design first) / operator (which)** |
| **R-727** | **[P2-MEDIUM] The whole-guest restore test picks a PREVIOUS box's archive, fails on its key, and the household sees a bare ✗ labelled with the wrong tier.** MEASURED 2026-09-30 01:52 UTC on the drill box: the agent's restore test chose `felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z` („newest settled archive … has not been proven") — written by a drill box of 2026-09-16, still in the customer's ep0 namespace next to tonight's own `2026-09-29T19:37:07Z` — and failed `wrong key - unable to verify signature since manifest's key 6b:ca:5f:3f… does not match provided key de:51:7a:18…`. One operator mail (`restore_test_failed`); the hub then re-logged the stored failure at every 15-minute report for 5 hours. The household's backups page read „✗ Visszaállítás ellenőrizve 2026-09-30 03:52 — Helyi tároló (local)" — the local tier had not failed; the pbs tier had, and the page says neither which nor why. Root: host delete leaves the old box's archives in the namespace (R-526's shape). **Fix direction:** the restore test skips (and reports as foreign) archives whose key fingerprint is not this host's; the page names the tier that failed. | **READY — rank P2-MEDIUM; owner: CC** |
| **R-727** | **[P2-MEDIUM] The whole-guest restore test picks a PREVIOUS box's archive, fails on its key, and the household sees a bare ✗ labelled with the wrong tier.** MEASURED 2026-09-30 01:52 UTC on the drill box: the agent's restore test chose `felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z` („newest settled archive … has not been proven") — written by a drill box of 2026-09-16, still in the customer's ep0 namespace next to tonight's own `2026-09-29T19:37:07Z` — and failed `wrong key - unable to verify signature since manifest's key 6b:ca:5f:3f… does not match provided key de:51:7a:18…`. One operator mail (`restore_test_failed`); the hub then re-logged the stored failure at every 15-minute report for 5 hours. The household's backups page read „✗ Visszaállítás ellenőrizve 2026-09-30 03:52 — Helyi tároló (local)" — the local tier had not failed; the pbs tier had, and the page says neither which nor why. Root: host delete leaves the old box's archives in the namespace (R-526's shape). **Fix direction:** the restore test skips (and reports as foreign) archives whose key fingerprint is not this host's; the page names the tier that failed. **FIXED 2026-09-30 (agent v0.138.0 + decision 51):** measured — a PBS archive carries its key FINGERPRINT (PVE content `encrypted`), not a host id; the storage carries its own (`encryption-key`). The restore test skips an archive written with another key (logged by name). The ✗ card names the tier (controller v0.283.0) — **and the 2026-09-30 claim "the page blames the local tier" was my misreading**: the „Helyi tároló (local)" after the ✗ was the next section's heading. ep0: the three drill archives in `tester-1` removed, other namespaces byte-identical. Delivered by signed jobs to both demo boxes (340 s); their due-check reads normally on 0.138.0. Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partC/`. | **CLOSED 2026-09-30 — agent v0.138.0** |
| **R-728** | **[P3-LOW] A customer created with one press was created TWICE, and the first of its two connect mails holds a dead link.** MEASURED 2026-09-30 on `Tester-2`: the hub logged `Customer config created: Tester-2` twice in the same second and two self-bind mints (hashes `c40df008…`, `6a1cbef4…`); a mint replaces the previous link (single-active), so one of the two identical mails the tester received answers „expired". Cause not established (a double form submit, or the handler run twice). **Fix direction:** make the create idempotent within a few seconds (or disable the button on submit), and pin it. The workaround for the tester is in STATUS. | **READY — rank P3-LOW; owner: CC (hub)** |
| **R-729** | **[P3-LOW] An off-site target, once saved on the page, cannot be removed through the product.** MEASURED 2026-09-30 on 9202: `/backup/offbox/config` refuses an empty address and no route clears the target; the session removed its throwaway target from `settings.json` by hand, with the controller stopped (harness teardown on a scratch guest). A household that tries its own NAS and gives up keeps a disabled target forever. **Fix direction:** a „Távoli mentési cél törlése" press that clears the target (never the repository). | **READY — rank P3-LOW; owner: CC (controller)** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.