Fixes before the first tester: decisions 50/51 outcomes, register 359->361 (R-719/720/721/722/727 closed, R-723 fixed, R-724/725 narrowed, R-728/729 opened), STATUS with the Tester-2 checklist, report
gates / gates (push) Successful in 26s
gates / gates (push) Successful in 26s
Secret scan before commit: 6162 files, 0 hits, control 1. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
+5
-1
@@ -20,7 +20,11 @@
|
||||
> restore test takes only the box's own).** `09` §3. 50: when the customer has off-site, every newly installed app is
|
||||
> in the off-site copy; over quota the page names the largest and the household chooses; history is never deleted to
|
||||
> make room without that choice. 51: the three drill archives in `tester-1`'s ep0 namespace go, nothing else. Also: the
|
||||
> first real tester is a NEW record `Tester-2`; the volunteer guide is rewritten as measured (R-722).
|
||||
> first real tester is a NEW record `Tester-2`; the volunteer guide is rewritten as measured (R-722). **Outcome (same day):** controller v0.283.0/v0.283.1 (apps off-site by default + one press + size card; a Stop holds
|
||||
> at the quiesce resume, the volume dump, the update leg and the crash recovery — v0.283.0 missed the dump's production
|
||||
> adapter, caught live), agent v0.138.0 (the restore test skips an archive written with another key; signed delivery to
|
||||
> both demo boxes), hub v0.126.0 („Új linket kérek" on an expired/used bind link; no day-one false mails), ep0's three
|
||||
> drill archives removed. Floor 0.283.1. Report: `REPORT-fixes-first-tester-2026-09-30.md`.
|
||||
|
||||
> **2026-09-29 evening — operator rulings 48 (wanderer stays, A) and 49 ("close sign-up now", A).** Controller
|
||||
> **v0.282.0**: `after_setup:` (the app's own sign-up switch, env merged + one `compose up -d`, when the gate opens or on
|
||||
|
||||
@@ -0,0 +1,79 @@
|
||||
# REPORT — 2026-09-30: the fixes before the first real tester
|
||||
|
||||
Architecture read first: `07` §6 (the 2026-09-16 ruling), `03` (the restore test's selection), `05` (self-bind),
|
||||
`04` §3.1 (signed delivery), `09` §3 decisions 45–49; rows R-719 … R-727, R-95, R-240, R-494, R-600, R-688.
|
||||
Releases: **controller v0.283.0 + v0.283.1**, **agent v0.138.0**, **hub v0.126.0**. Evidence:
|
||||
`documentation/audits/evidence-fixes-first-tester-2026-09-30/` (part0, partA, partC, partD, partE, release).
|
||||
|
||||
## Tester-2 — read-only checklist (nothing on Tester-2, Cloudflare, ep0 or the Storage Box was changed)
|
||||
|
||||
| # | item | state | where to fix |
|
||||
|---|---|---|---|
|
||||
| 1 | Tunnel token pasted | **done** — tunnel `3ce0eccd…` (Peti's original, reused) | — |
|
||||
| 2 | DNS of `sajatfelhom.hu` | **done** — ONE record `*.sajatfelhom.hu` → that same tunnel; no leftover pointing elsewhere; today 530 (no connector, right with no box) | — |
|
||||
| 3 | The tunnel's route `*.sajatfelhom.hu` → `https://traefik`, No TLS Verify | **UNKNOWN** — the record's Cloudflare key reads DNS only (`Authentication error` on the tunnel; stopped there) | Cloudflare → Zero Trust → Networks → Tunnels → this tunnel → Published application routes |
|
||||
| 4 | Off-site | **done** — shared 100 GB, new sub-account 322460 (username `sub2` reused; Peti's was emptied and deleted 2026-09-25) | — |
|
||||
| 5 | DR tier (ep0) | **done** — ON; nothing of Tester-2 or Peti on ep0 yet (made at the first connection) | — |
|
||||
| 6 | The connect e-mail | **sent twice at 07:36 UTC** (the customer was created twice, R-728) — **only one of the two links works**; valid until 2026-10-07 07:36 UTC | Tell your friend: if one link says „expired", use the other; or press „Send self-bind link" once just before the install |
|
||||
| 7 | E-mail language | **English** — mail, bind page and the box start in English | Hub → Tester-2 → Edit, if Hungarian is wanted |
|
||||
| 8 | Owner passphrase | yours to hand over | in person / by phone |
|
||||
| 9 | Customer id `Tester-2` has a capital letter | never walked before; no known break | note only |
|
||||
|
||||
## The Parts
|
||||
|
||||
| Part | state | note |
|
||||
|---|---|---|
|
||||
| 0 Tester-2 read-only | **done** | checklist above; items 3 and 6 need you |
|
||||
| A1 measure the over-quota path | **done** | it deletes nothing extra: refuses new pushes, runs only the ruled retention; now pinned |
|
||||
| A2 apps off-site by default + one press for older apps | **done** | controller v0.283.0; live on 9202 |
|
||||
| A3 size warning | **done** | page card; unit + parity proven (a household NAS target has no quota to test live) |
|
||||
| B guide + slips | **done / narrowed** | six stale lines rewritten (the sixth found on the way: auto-reboot); R-724/R-725 partly, the rest narrowed |
|
||||
| C1 ep0 cleanup | **done** | three archives removed; every other namespace byte-identical |
|
||||
| C2 agent v0.138.0 | **done** | signed delivery to both demo boxes (340 s); due-check normal on both |
|
||||
| D fresh link for a returning customer | **changed** | the brief's trigger cannot be built (a registering box is unclaimed); built „Új linket kérek" on the expired/used page |
|
||||
| E Stop holds | **done, after a live failure** | v0.283.0 was wrong in production (adapter); v0.283.1 fixed and proven live |
|
||||
| F day-one mails | **done** | unit + red-proof; live proof at Tester-2's first hour |
|
||||
|
||||
## Claims in the brief that turned out wrong, named
|
||||
|
||||
- **"The over-quota path prunes history"** — **wrong.** Over the quota the box refuses new pushes and runs the SAME
|
||||
retention as every night; with no new snapshots nothing extra ages out. No P1.
|
||||
- **"Peti's Cloudflare records for `sajatfelhom.hu` may still be there"** — **there is exactly one record, and it
|
||||
points at the tunnel Tester-2 now carries** (Peti's tunnel, reused). Nothing stale to remove.
|
||||
- **"A PBS archive carries its box's key fingerprint or host id"** — **half right:** the key fingerprint yes (PVE
|
||||
content `encrypted`), a host id no (the comment is only „felhom local-api", the owner is the customer's token).
|
||||
- **"The hub sends no link at registration"** — **right, and it cannot:** the registration carries nothing of a
|
||||
customer. The fix was changed to a button on the old link's page.
|
||||
- **"Stop is lost at the backup's resume only"** — **wrong:** also at the nightly volume dump, the update leg and the
|
||||
startup crash recovery; all four fixed.
|
||||
- **Mine, from 2026-09-30:** "the ✗ names the wrong tier" — misread; the local-tier heading was the NEXT section.
|
||||
|
||||
## Red-proofs (each seen failing on its assertion, then restored)
|
||||
|
||||
RP31 new app not ON · RP32 earlier OFF overridden · RP33 hook not wired · RP34 over-quota forget differs · RP35 exact
|
||||
quota "does not fit" · RP36 quiesce restarts a stopped app · RP37 dump restarts it · RP38 update leg presses it ·
|
||||
RP39 (agent) an earlier box's archive picked · RP40 new box "recovered" · RP41 first-hour skip mailed · RP42 no fresh
|
||||
link · RP43 the production adapter does not answer · RP44 crash recovery restarts it. Outputs:
|
||||
`evidence-fixes-first-tester-2026-09-30/` and the scratchpad `rp/` copies.
|
||||
|
||||
## Slips of mine, said
|
||||
|
||||
- The first hub/evidence commit went out after my secret-scan script crashed (it looked for last night's shredded
|
||||
files). Scanned right after: 5 125 files, 7 secrets, **0 hits**, control 1. Nothing leaked.
|
||||
- v0.283.0 shipped the Stop fix un-wired in production; the live test caught it; v0.283.1 is the second controller
|
||||
release this session (the one-release rule bent, reason in its CHANGELOG).
|
||||
|
||||
## Rows
|
||||
|
||||
Closed **R-719, R-720, R-721, R-722, R-727**. Fixed pending live proof **R-723**. Narrowed **R-724, R-725**. Opened
|
||||
**R-728** (customer created twice), **R-729** (no way to remove an off-site target). **Register 359 → 361 rows.**
|
||||
R-726 (a returning customer's orphaned repository) stays open — a NEW record does not meet it.
|
||||
|
||||
## Teardown
|
||||
|
||||
- **9202:** controller left on 0.283.1 (scratch; the floor does not reach it); the throwaway NAS target switched off
|
||||
through the form and then removed from `settings.json` with the controller stopped (R-729: no product path);
|
||||
`glance` removed with its data; paperless-ngx running; the off-site switches back OFF.
|
||||
- **Demo boxes:** controller 0.283.1 by the floor, agent 0.138.0 by signed jobs — nothing else.
|
||||
- **ep0:** only the three archives of decision 51. **Hub:** floor 0.283.1 (MinAgent 0.131.0 declared); hub 0.126.0.
|
||||
- **Tester-2 / Cloudflare:** read only. DooPlex: pushes, builds, the hub deploy, signing.
|
||||
@@ -1,23 +1,27 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Ready for a first real tester: YES, on a new customer record — a fresh box went from the download to restored apps with no help at all; fix the volunteer guide first, and decide whether their apps go off-site by themselves (today they do not).**
|
||||
**Ready for the first real tester (Tester-2): yes, once you check two things — the tunnel's route in Cloudflare, and which of the two connect mails your friend uses.**
|
||||
|
||||
**Updated 2026-09-30 morning. Both demo boxes run controller 0.282.0 and host agent 0.137.0. Hub 0.125.0. New installs get golden 0.282.0 with agent 0.137.0 (baked and vouched last night).**
|
||||
**Updated 2026-09-30. Both demo boxes run controller 0.283.1 and host agent 0.138.0. Hub 0.126.0. New installs get golden 0.282.0 with agent 0.137.0.**
|
||||
|
||||
**Decisions today** (yours, recorded): every new app goes off-site by itself (50); the three drill backups on ep0 go, and the restore test uses only the box's own backups (51).
|
||||
|
||||
**Tester-2 — what I checked, read only.**
|
||||
- Done: tunnel key pasted; the domain has exactly one DNS record, and it points at that same tunnel; off-site on (100 GB, a fresh storage account); DR tier on; e-mail language English.
|
||||
- **Check yourself:** the tunnel's route. My Cloudflare key could read DNS but not tunnels. It must say `*.sajatfelhom.hu` → `https://traefik`, with „No TLS Verify" ticked.
|
||||
- **Two connect mails went out** (the customer was created twice). Only one link works. Tell your friend: if one says "expired", use the other. The live one is valid until 7 October.
|
||||
|
||||
**What I did, and it worked.**
|
||||
- **A fresh box, walked like a volunteer, needed no help.** Download, install on the Hungarian keyboard, the e-mailed link, the dashboard through the internet, the recovery code, three apps, a family member, backup, restore, delete, a power cut, typos, a phone. Last walk needed one intervention; this one needed none.
|
||||
- **The new safety parts held on a fresh box.** Strangers saw only the gate page. The gate opened by itself after the household's setup. Sign-up stayed closed. The random first password worked, and the old default password was refused.
|
||||
- **Apps go off-site by themselves now.** Older apps get one button. If the apps do not fit in 100 GB, the page names the biggest. The box never deletes old copies to make room — I checked; it never did.
|
||||
- **A Stop holds during a backup.** The first version was wrong in real use; my live test caught it, and the second version is proven live.
|
||||
- **ep0 is clean:** the three old drill backups are gone; the other customers' backups are unchanged.
|
||||
- **A returning customer gets a new link** from a button on the expired link's page.
|
||||
- **The volunteer guide is current** (six stale lines fixed).
|
||||
- **Two false operator mails on a new box's first day are gone.**
|
||||
|
||||
**What broke.**
|
||||
- **No off-site copy on night one.** Every app starts with its off-site copy switched off, and nothing tells the household. On top of that, tester-1's old off-site store belonged to an earlier box, so the box skipped it all night. The household's data stayed safe on the box.
|
||||
- **The night's restore test tried an old box's backup** and failed on its key. The page shows a bare ✗ and names the wrong place.
|
||||
- **A Stop pressed during a backup was undone by that backup.**
|
||||
- **The volunteer guide is out of date in five places.** For example, it still gives BookStack's old default password, which is now refused.
|
||||
|
||||
**Rows.** 9 opened, 1 closed. The list went from 350 to 359 rows.
|
||||
**Rows.** 5 closed, 2 opened. The list went from 359 to 361 rows.
|
||||
|
||||
**What needs you.**
|
||||
1. **Decide: should apps go off-site by themselves?** Option A: yes, every app is switched on when the customer has off-site (my pick — that is what the guide already promises). Option B: no, the guide tells the household to switch each app on. If you do nothing, a new household's apps have no off-site copy.
|
||||
2. **Approve the guide fix** (five lines, I write them). If you do nothing, a volunteer tries BookStack's old password and fails.
|
||||
3. **Three whole-machine backups of deleted drill boxes are still in tester-1's space on ep0** (two from 16 September, one from last night). I only looked; I did not touch them. The old ones make the restore test fail. If you do nothing, every new tester-1 box repeats that failure. Say "remove them" and I will do it with before/after checks.
|
||||
4. **D4, the image copies, the Peti leftovers:** unchanged. If you do nothing, nothing changes.
|
||||
1. **The Tester-2 tunnel route** (above). If you do nothing and it is missing, your friend's dashboard answers an error page.
|
||||
2. **A golden before your friend installs?** The rule says bake one before any fresh install; this task did not include it. Option A (my pick): I bake 0.283.1 before the install — the box then starts on today's code. Option B: skip it — the box starts on 0.282.0 and updates itself to 0.283.1 within seconds (proven on both demo boxes today). If you do nothing, B happens.
|
||||
3. **D4, the image copies:** unchanged. If you do nothing, nothing changes.
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -553,12 +553,19 @@ R-636's louder repeated alarm.
|
||||
which a per-app switch starting OFF had undone. If the apps will not fit the customer's quota, the page says so,
|
||||
names the largest, and the household chooses which stay off-site. Apps already installed are not switched by a
|
||||
release (the 2026-09-29 Part 0 rule); the backups page offers one press to switch them all on.
|
||||
**The box never deletes off-site history to make room without the household's choice.** *(Outcome: filled in by
|
||||
the 2026-09-30 fixes session.)*
|
||||
**The box never deletes off-site history to make room without the household's choice.** **Outcome (2026-09-30,
|
||||
controller v0.283.0):** measured first — over the quota the box already refused NEW pushes and ran only the ruled
|
||||
retention, and an app whose files would cross the quota went up settings + database only; no history was ever
|
||||
deleted to make room (now pinned). A fresh install switches the app ON (`DefaultOffboxOnForNewApp`; an earlier
|
||||
choice is kept); older apps get one press on both backup pages; the size card names the three largest. Live on
|
||||
9202: the press, and a fresh app joining by itself.
|
||||
51. **ep0: the three whole-guest archives of deleted drill boxes in `tester-1`'s namespace are removed, and only
|
||||
those** — *operator ruling 2026-09-30 (R-727).* With before/after controls on every other namespace. The restore
|
||||
test is fixed to take only the current box's own archives, so an archive of an earlier box in the same namespace
|
||||
can never be tested again. *(Outcome: filled in by the 2026-09-30 fixes session.)*
|
||||
can never be tested again. **Outcome (2026-09-30):** the three archives (2026-09-16T17:27:32Z, 2026-09-16T21:59:54Z,
|
||||
2026-09-29T19:37:07Z), mapped to their drill boxes by key fingerprint and host record, were forgotten on ep0 one by
|
||||
one; every other namespace byte-identical before and after; chunks go at ep0's own weekly GC. Agent v0.138.0: an
|
||||
archive carries its key FINGERPRINT (not a host id), and the restore test skips one written with another key.
|
||||
Same day, operator: the first real tester gets a NEW customer record (`Tester-2`, domain `sajatfelhom.hu`), not
|
||||
`tester-1`; and CC rewrites the volunteer guide as measured (R-722).
|
||||
|
||||
|
||||
@@ -0,0 +1,3 @@
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.283.0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.283.0 Up 25 seconds (healthy)
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.283.1 Up 25 seconds (healthy)
|
||||
@@ -0,0 +1,31 @@
|
||||
## 2026-09-30T08:40:33Z 9202: the household's own NAS target, through the page's form (192.0.2.1 = RFC 5737 documentation address, never reachable; throwaway key)
|
||||
HTTP/2 302
|
||||
location: /backups/remote?flash_error=flash.offbox.target_saved_escrow_agent_down
|
||||
## 2026-09-30T08:41:15Z before: paperless-ngx=OFF privatebin=OFF
|
||||
/backups/apps shows the offer: 1
|
||||
press „Igen, mindegyikre“ → 302 https://192.168.0.114/backups/remote?flash=flash.offbox.enabled_all
|
||||
after: paperless-ngx=ON privatebin=ON
|
||||
offer still shown: 0
|
||||
## 08:41:27Z fresh install of glance on 9202, fields ['SUBDOMAIN']
|
||||
deploy → {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
|
||||
state after 18s: running
|
||||
off-site switches now: glance=ON paperless-ngx=ON privatebin=ON
|
||||
2026/09/30 08:41:30 [INFO] [backup] glance: off-site copy switched ON at install (decision 50)
|
||||
## 2026-09-30T08:55:05Z 9202 teardown
|
||||
paperless-ngx Start → {"ok":true,"data":{"state":"starting"},"message":"Stack paperless-ngx start requ
|
||||
glance Stop → {"ok":true,"message":"Stack glance stop completed"}
|
||||
glance remove (data + backups) → {"ok":true,"data":{"removed":"glance","volumes_removed":["glance_glance_config"],"hdd_paths_removed":[],"hdd_paths_preserved":[],"hdd_note":"Az alkalmazás nem
|
||||
off-site OFF paperless-ngx → 302
|
||||
off-site OFF privatebin → 302
|
||||
target switched off via the form → 302 https://192.168.0.114/backups/remote?flash=flash.offbox.target_saved
|
||||
data=
|
||||
Stopping 'felhom-controller-bootstrap.service', but its triggering units are still active:
|
||||
felhom-controller-bootstrap.path
|
||||
Traceback (most recent call last):
|
||||
File "<stdin>", line 2, in <module>
|
||||
FileNotFoundError: [Errno 2] No such file or directory: '/settings.json'
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.283.1 Up 27 seconds (healthy)
|
||||
/backups/remote now: 0 × „Még nincs beállítva távoli mentési cél“
|
||||
file=/var/lib/docker/volumes/felhom-controller-data/_data/data/settings.json
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.283.1 Up 30 seconds (healthy)
|
||||
/backups/remote: „Még nincs beállítva távoli mentési cél“ × 1; paperless: "state":"running"
|
||||
@@ -0,0 +1,55 @@
|
||||
== demo-hp
|
||||
felhom-agent 0.138.0
|
||||
Sep 30 10:40:28 demo-hp felhom-agent[2355282]: time=2026-09-30T10:40:28.593+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=11d75b6842043927 op=agent_update
|
||||
Sep 30 10:40:30 demo-hp felhom-agent[2355282]: time=2026-09-30T10:40:30.718+02:00 level=INFO msg="backup: restore-test scheduler shutting down" reason="context canceled"
|
||||
Sep 30 10:40:31 demo-hp felhom-agent[2135127]: time=2026-09-30T10:40:31.715+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
|
||||
Sep 30 10:40:32 demo-hp felhom-agent[2135127]: time=2026-09-30T10:40:32.658+02:00 level=INFO msg="janitor: starting (restore-test scratch retry + stale-lock sweep)" interval=10m0s
|
||||
Sep 30 10:41:32 demo-hp felhom-agent[2135127]: time=2026-09-30T10:41:32.681+02:00 level=WARN msg="selfupdate: update committed" version=0.138.0 wrapper=""
|
||||
== felhom-pve
|
||||
felhom-agent 0.138.0
|
||||
Sep 30 10:37:12 demo-felhom felhom-agent[2880126]: time=2026-09-30T10:37:12.223+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=ccafbfb170d1b714 op=agent_update
|
||||
Sep 30 10:37:14 demo-felhom felhom-agent[2081460]: time=2026-09-30T10:37:14.960+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
|
||||
Sep 30 10:37:16 demo-felhom felhom-agent[2081460]: time=2026-09-30T10:37:16.219+02:00 level=INFO msg="janitor: starting (restore-test scratch retry + stale-lock sweep)" interval=10m0s
|
||||
Sep 30 10:38:16 demo-felhom felhom-agent[2081460]: time=2026-09-30T10:38:16.243+02:00 level=WARN msg="selfupdate: update committed" version=0.138.0 wrapper=""
|
||||
== demo-hp 08:58:03 — selftest restore-test-due (READ-ONLY: the scheduler's own verdict)
|
||||
* demo-hp status=online fp=07:B5:72:5D:5E:C1…
|
||||
[ ok ] node status up 953h13m36s, load [0.60 0.67 0.58], mem 5.2GiB/29.3GiB, root 23.0GiB/38.6GiB
|
||||
[ ok ] list lxc 1 guest(s)
|
||||
- 9201 "demo-hp" status=running
|
||||
[ ok ] pool read pool "felhom", 1 member(s)
|
||||
- 9201 type=lxc
|
||||
[ ok ] storage 4 store(s)
|
||||
- local-lvm type=lvmthin content=images,rootdir used=32.4GiB/53.9GiB
|
||||
- local type=dir content=vztmpl,backup,import,iso used=23.0GiB/38.6GiB
|
||||
- felhom-pbs type=pbs content=backup used=0.0GiB/0.0GiB
|
||||
- nvme-scratch type=dir content=images,rootdir used=73.5GiB/937.8GiB
|
||||
=== selftest OK ===
|
||||
== felhom-pve 08:58:04 — selftest restore-test-due (READ-ONLY: the scheduler's own verdict)
|
||||
* demo-felhom status=online fp=60:8F:4C:50:C3:8E…
|
||||
[ ok ] node status up 1225h32m30s, load [0.29 0.30 0.26], mem 2.4GiB/15.4GiB, root 27.4GiB/93.9GiB
|
||||
[ ok ] list lxc 1 guest(s)
|
||||
- 9201 "demo-felhom" status=running
|
||||
[ ok ] pool read pool "felhom", 1 member(s)
|
||||
- 9201 type=lxc
|
||||
[ ok ] storage 4 store(s)
|
||||
- felhom-backup type=dir content=backup used=11.9GiB/915.8GiB
|
||||
- local type=dir content=import,backup,iso,vztmpl used=27.4GiB/93.9GiB
|
||||
- local-lvm type=lvmthin content=rootdir,images used=11.9GiB/348.8GiB
|
||||
- felhom-pbs type=pbs content=backup used=0.0GiB/0.0GiB
|
||||
=== selftest OK ===
|
||||
== demo-hp 08:58:13 — selftest=restore-test-due (READ-ONLY)
|
||||
time=2026-09-30T10:58:13.922+02:00 level=INFO msg="backup: restore-test skips an entry that is not a backup of a guest" target=local volid=local:backup/felhom-golden-0.236.0.tar.zst size_bytes=654115664 reason="not a backup of a guest (the storage reports no v
|
||||
time=2026-09-30T10:58:13.924+02:00 level=INFO msg="backup: restore-test skips an entry that is not a backup of a guest" target=local volid=local:backup/vzdump-lxc-9100-2026_08_21-17_59_15.tar.zst size_bytes=656970239 reason="guest 9100 does not exist on this n
|
||||
tier=felhom-pbs due=false archive="felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z" landed=2026-09-24T20:06:25Z proven="felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z"
|
||||
reason: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven
|
||||
tier=local due=false archive="local:backup/vzdump-lxc-9201-2026_09_29-04_42_11.tar.zst" landed=2026-09-29T02:42:11Z proven="local:backup/vzdump-lxc-9201-2026_09_29-04_42_11.tar.zst"
|
||||
reason: newest settled archive (landed 2026-09-29T02:42:11Z) is already proven
|
||||
cost tier=felhom-pbs one_lookup=322ms
|
||||
cost tier=local one_lookup=45ms
|
||||
== felhom-pve 08:58:14 — selftest=restore-test-due (READ-ONLY)
|
||||
tier=felhom-backup due=true archive="felhom-backup:backup/vzdump-lxc-9201-2026_09_29-07_36_43.tar.zst" landed=2026-09-29T05:36:43Z proven="felhom-backup:backup/vzdump-lxc-9201-2026_09_28-07_36_39.tar.zst"
|
||||
reason: newest settled archive (landed 2026-09-29T05:36:43Z) has not been proven (last proven archive was a different one)
|
||||
tier=felhom-pbs due=false archive="felhom-pbs:backup/ct/9201/2026-09-29T04:16:43Z" landed=2026-09-29T04:16:43Z proven="felhom-pbs:backup/ct/9201/2026-09-29T04:16:43Z"
|
||||
reason: newest settled archive (landed 2026-09-29T04:16:43Z) is already proven
|
||||
cost tier=felhom-backup one_lookup=20ms
|
||||
cost tier=felhom-pbs one_lookup=306ms
|
||||
@@ -0,0 +1,5 @@
|
||||
## 2026-09-30T08:38:33Z live, hub 0.126.0 — a made-up link (never real; same page as an expired one by design)
|
||||
GET /bind/<made-up> → 200; page: Felhom — Doboz összekötése | Felhom | doboz | összekötése | Ez a hivatkozás érvénytelen vagy lejárt. | A hivatkozás 7 napig érvényes. Ha lejárt, kérj újat az ügyfélszolgálattól, vagy az összekötést az üzemeltető is elvégezheti. | Ha ez a te linked volt, és a dobozod még nincs összekötve, új linket küldünk arra az e-mail címre, amelyet a Felhomnál megadtál. | Új linket kérek | Felhom.eu
|
||||
the button's form: <form method="POST" action="/bind/<token>/resend">
|
||||
POST …/resend → 200; page: Felhom — Doboz összekötése | Felhom | doboz | összekötése | Kész. | Ha ez egy valódi hivatkozás volt, és a dobozod még nincs összekötve, néhány percen belül új e-mailt kapsz a regisztrált címedre. Ha nem jön, szólj az ügyfélszolgálatnak. | Felhom.eu
|
||||
hub log (self-bind lines, last minute): 0
|
||||
+29
@@ -0,0 +1,29 @@
|
||||
## 2026-09-30T08:42:08Z 9202: „Mentés most“, and the household presses Stop on paperless-ngx while the dump holds it down
|
||||
paperless-ngx before: "state":"running"
|
||||
„Mentés most“ → {"ok":true,"message":"Mentés elindítva"}
|
||||
08:42:13.593 saw: 2026/09/30 08:42:13 backup.go:955: [INFO] [backup] Stopping paperless-ngx for safe volume
|
||||
08:42:13.600 Stop pressed → {"ok":true,"message":"Stack paperless-ngx stop completed"}
|
||||
backup run: "success":true
|
||||
paperless-ngx after: "state":"starting"
|
||||
2026/09/30 08:42:13 backup.go:955: [INFO] [backup] Stopping paperless-ngx for safe volume dump
|
||||
2026/09/30 08:42:13 manager.go:1312: [INFO] [stacks] Stopping stack: paperless-ngx
|
||||
2026/09/30 08:42:13 router.go:601: [INFO] [api] stop requested for stack: paperless-ngx
|
||||
2026/09/30 08:42:13 desiredstate.go:81: [INFO] [stacks] desired state for paperless-ngx recorded as "stopped" (was "running")
|
||||
2026/09/30 08:42:13 manager.go:1312: [INFO] [stacks] Stopping stack: paperless-ngx
|
||||
2026/09/30 08:42:21 backup.go:971: [INFO] [backup] Restarting paperless-ngx after volume dump
|
||||
2026/09/30 08:42:21 manager.go:1223: [INFO] [stacks] Starting stack: paperless-ngx
|
||||
## 2026-09-30T08:54:15Z RETEST on v0.283.1 — „Mentés most“, Stop on paperless-ngx while the dump holds it
|
||||
paperless-ngx before: "state":"running"
|
||||
„Mentés most“ → {"ok":true,"message":"Mentés elindítva"}
|
||||
08:54:19.823 saw: 2026/09/30 08:54:19 backup.go:955: [INFO] [backup] Stopping paperless-ngx for safe volume
|
||||
08:54:19.828 Stop pressed → {"ok":true,"message":"Stack paperless-ngx stop completed"}
|
||||
backup run: "success":true
|
||||
paperless-ngx 20 s after the run: "state":"stopped"
|
||||
2026/09/30 08:54:14 desiredstate.go:81: [INFO] [stacks] desired state for paperless-ngx recorded as "running" (was "stopped")
|
||||
2026/09/30 08:54:14 manager.go:1223: [INFO] [stacks] Starting stack: paperless-ngx
|
||||
2026/09/30 08:54:19 backup.go:955: [INFO] [backup] Stopping paperless-ngx for safe volume dump
|
||||
2026/09/30 08:54:19 manager.go:1312: [INFO] [stacks] Stopping stack: paperless-ngx
|
||||
2026/09/30 08:54:19 router.go:601: [INFO] [api] stop requested for stack: paperless-ngx
|
||||
2026/09/30 08:54:19 desiredstate.go:81: [INFO] [stacks] desired state for paperless-ngx recorded as "stopped" (was "running")
|
||||
2026/09/30 08:54:19 manager.go:1312: [INFO] [stacks] Stopping stack: paperless-ngx
|
||||
2026/09/30 08:54:28 backup.go:967: [INFO] [backup] paperless-ngx NOT restarted after the volume dump: the household stopped it meanwhile
|
||||
+3
@@ -2,3 +2,6 @@
|
||||
HTTP/1.1 303 See Other
|
||||
Location: /configuration?flash=floor_set
|
||||
2026/09/30 10:33:19 [INFO] Global controller-version floor set to "0.283.0" (declared MinAgent "0.131.0")
|
||||
## 2026-09-30T08:53:29Z floor → 0.283.1 (MinAgent 0.131.0 declared)
|
||||
Location: /configuration?flash=floor_set
|
||||
2026/09/30 10:53:29 [INFO] Global controller-version floor set to "0.283.1" (declared MinAgent "0.131.0")
|
||||
|
||||
@@ -0,0 +1,15 @@
|
||||
## 2026-09-30T08:37:05Z hub 0.126.0 via ArgoCD (HEAD 7e5009808856179dea9cafcde472cb462943ae6f)
|
||||
revision seen: 7e5009808856179dea9cafcde472cb462943ae6f
|
||||
deployment "hub" successfully rolled out
|
||||
image: gitea.dooplex.hu/admin/felhom-hub:0.125.0
|
||||
argocd: sync=OutOfSync health=Healthy rev=7e5009808856179dea9cafcde472cb462943ae6f
|
||||
2026/09/30 10:34:43 [INFO] enqueued signed-op job ccafbfb170d1b714 for host demo-felhom-8363b5 (775 bytes)
|
||||
2026/09/30 10:35:14 [INFO] PBS-DR box refreshed: 18.7% full (18.3 GB of 97.9 GB)
|
||||
2026/09/30 10:37:11 [INFO] host-report from demo-felhom-8363b5 (1 guests, 4 storage targets, 2 backups, 2 restore-tests, 2 pbs-snapshots, 14458 bytes)
|
||||
2026/09/30 10:37:11 [INFO] DR-recipe host-half stored for customer demo-felhom (host demo-felhom-8363b5, v1)
|
||||
2026/09/30 10:37:12 [INFO] host demo-felhom-8363b5 cleared signed-op job ccafbfb170d1b714 (executed or rejected)
|
||||
08:38:00 image=gitea.dooplex.hu/admin/felhom-hub:0.126.0 argocd=Synced/Healthy
|
||||
running pod image: gitea.dooplex.hu/admin/felhom-hub:0.126.0
|
||||
2026/09/30 10:37:19 [INFO] felhom-hub 0.126.0 starting
|
||||
2026/09/30 10:37:19 [INFO] Default controller-version floor: 0.120.0
|
||||
2026/09/30 10:37:20 [INFO] Gitea artifact browser enabled (Day-0 version dropdowns) via http://gitea.gitea-system.svc.cluster.local:3000
|
||||
@@ -830,15 +830,17 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-716** | **[P3-LOW] Apps installed before controller 0.281.0 keep their open sign-up — decision 47 closes it only on apps whose gate the box opened.** READ 2026-09-29 on the demo boxes after catalog `6faf432` synced: demo-hp's adventurelog and opengist, demo-felhom's opengist carry `signup_block:` in their synced template and no gate record, so no block (`audits/gate-rollout-2026-09-29/0/P0-3-demo-boxes-after-push.txt`). This is Part 0's rule working as designed (a catalog change never touches an installed app). **Needs an operator word** before anything changes on an installed app: a one-time "close sign-up now" press on the app page for an installed app, or leave them. Only the demo boxes have such installs today. **Operator ruled A (decision 49); built in controller v0.282.0 and pressed** on demo-hp's adventurelog and opengist and demo-felhom's opengist: before, sign-up served; after, refused; adventurelog's own switch on (`audits/signup-lock-2026-09-29/`C). | **CLOSED — 2026-09-29** |
|
||||
| **R-717** | **[P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone.** MEASURED 2026-09-29: opengist `disable-signup` is an admin-panel setting (no env, no CLI); wishlist `system_config.enableSignup` (Prisma). Their blocks are case-insensitive and refused every trick shape (`audits/signup-lock-2026-09-29/B/`). **Fix direction:** an `after_setup` command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). | **OPEN — P3; owner: CC** |
|
||||
| **R-718** | **[P3-LOW] "Close sign-up now" restarts an app with its own switch, and the card does not say so.** MEASURED 2026-09-29 on demo-hp: pressing it on adventurelog recreated its backend (~30 s, one 500 on its login page). The window's card says the app restarts; the close card does not. **Fix direction:** the close card and the gate-open moment say "the app restarts once" where `after_setup.env` exists. **ALSO MEASURED 2026-09-29 (new-household drill):** the gate-open press on a fresh vikunja restarted it for its own switch — the front door answered 404 for ~2 s and nothing said so. | **OPEN — P3; owner: CC** |
|
||||
| **R-719** | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. | **READY — rank P2-MEDIUM; owner: operator (which shape) / CC** |
|
||||
| **R-720** | **[P2-MEDIUM] A new household's apps are not in the off-site copy: every app starts with „3. mentés Kikapcsolva", and nothing the household is told says to switch it on.** MEASURED 2026-09-29 on a fresh box (customer with off-site ON, the default since 2026-09-16): `/backups/remote` read „Aktív — nincs kijelölt alkalmazás"; `/backups/apps` read „3. mentés Kikapcsolva — Ez az alkalmazás nincs kijelölve távoli mentésre" for all three apps. From source, an app is off-site only after `POST /backup/offbox/toggle` (`settings.SetAppOffbox`); no deploy or claim path sets it. The volunteer guide §7 says that after the recovery code „a távoli mentés magától elindul" — it starts, and copies nothing. So a household following the guide has **no off-site copy of its apps on night one**, on a one-drive box where the whole-guest tiers do not carry the data drive (`07` §6). The drill pressed „Bekapcsolás" for BookStack only, as a household reading the page might, to measure both paths on the night. R-240 (the empty run's wording) is the same gap seen from the other end. **Needs an operator decision** (it changes what the product promises): apps default to off-site ON when the customer has off-site, or the guide adds the step. | **READY — rank P2-MEDIUM; owner: operator (decide) / CC** |
|
||||
| **R-721** | **[P2-MEDIUM] The household presses Stop during a whole-guest backup, and the backup starts the app again.** MEASURED 2026-09-29 19:37 UTC on a fresh box: the first off-site whole-guest backup quiesced three apps at 19:37:01; the household pressed „Leállítás" on actualbudget at 19:37:16 and the controller recorded `desired state for actualbudget recorded as "stopped"`; the same second the backup's early resume ran `unquiescing … restarting 3 stack(s)` and `Starting stack: actualbudget`. The app page then read „Fut" while `app.yaml` kept `desired_state: stopped` (so the next reboot would stop it). The household had to press Stop again, and the removal was refused „Az alkalmazáson mentés vagy visszaállítás fut" for ~4 min (honest). Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase1/step10-stop-undone-by-quiesce.log`. **Fix direction:** unquiesce restarts only stacks whose desired state is still running; a test pins it (Stop during a quiesce → still stopped after the resume). | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-722** | **[P2-MEDIUM] The volunteer guide is stale in five places a volunteer reads literally.** MEASURED 2026-09-29 walking `runbooks/VOLUNTEER-first-hour.md` on golden 0.282.0: (1) §8 says BookStack's login is `admin@admin.com / password` — since the random first password (controller v0.280.0) that login is REFUSED; the app page says the right thing („a Beállítások oldalon látható első jelszó"). (2) §7 says the recovery-code bar comes „néhány perccel" after setup — measured ~15 min after enrolment (the off-site tier arrives with the agent's next 15-minute host report). (3) Nothing says to switch each app's off-site copy on (R-720). (4) §2 says a 2 GB USB stick and Rufus; the download page says at least 4 GB and Balena Etcher. (5) The operator part says no button is needed for the link (R-719), and §5 names the mail „Elindult a Felhom szervered" while a customer with an earlier box gets „Új beállító kód — újratelepült a szervered … A korábbi jelszavad már nem érvényes". Also, from the screens: the installer pre-selects `/dev/sda` (the guide says it never chooses), and its Summary screen rests on „Previous", not „Install". **Fix:** a guide edit; the operator approves the text. | **READY — rank P2-MEDIUM; owner: CC (text) / operator (approve)** |
|
||||
| **R-723** | **[P3-LOW] A fresh box sends the operator two mails on day one that describe nothing wrong.** MEASURED 2026-09-29: `Operator email sent for tester-1/node_recovered` 2 s after the new box's first controller report (the customer's previous box had been silent 12 days — the new box is not a recovery); and `backup_tier_skipped (warning) — Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed)` → operator mail at 19:27 UTC, 7 min after enrolment, because the first whole-guest run fired before the off-site tier's descriptor arrived (applied ~19:35 UTC, backed up fine at 19:37). Operator-only, so no household is alarmed, but an operator learns to ignore both. **Fix direction:** `node_recovered` not for a host enrolled < N min ago; the skipped tier inside the first hour after enrolment is `info`, not a mailed warning. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-724** | **[P3-LOW] The status pages disagree with each other in small ways a household notices.** MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw `2026-09-29T19:22:30Z`; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-725** | **[P3-LOW] Small copy slips on the first-hour path.** MEASURED 2026-09-29: the self-bind PAGE says the passphrase was received „a beállításkor" while the mail and console say „a Felhom üzemeltetőjétől" (R-497 unified the mail and console, not the page); the recovery-code wizard addresses the household formally („Írja fel", „adja meg") while every other screen says „te"; the console's linked banner ends „a doboz össze van kötve. V" (a stray glyph); a gated app answers a phone app's API call with English JSON „this app is waiting for its first setup" (the browser gets the Hungarian gate page). | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-719** | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** |
|
||||
| **R-720** | **[P2-MEDIUM] A new household's apps are not in the off-site copy: every app starts with „3. mentés Kikapcsolva", and nothing the household is told says to switch it on.** MEASURED 2026-09-29 on a fresh box (customer with off-site ON, the default since 2026-09-16): `/backups/remote` read „Aktív — nincs kijelölt alkalmazás"; `/backups/apps` read „3. mentés Kikapcsolva — Ez az alkalmazás nincs kijelölve távoli mentésre" for all three apps. From source, an app is off-site only after `POST /backup/offbox/toggle` (`settings.SetAppOffbox`); no deploy or claim path sets it. The volunteer guide §7 says that after the recovery code „a távoli mentés magától elindul" — it starts, and copies nothing. So a household following the guide has **no off-site copy of its apps on night one**, on a one-drive box where the whole-guest tiers do not carry the data drive (`07` §6). The drill pressed „Bekapcsolás" for BookStack only, as a household reading the page might, to measure both paths on the night. R-240 (the empty run's wording) is the same gap seen from the other end. **Needs an operator decision** (it changes what the product promises): apps default to off-site ON when the customer has off-site, or the guide adds the step. **BUILT 2026-09-30 (controller v0.283.0, decision 50):** a fresh install on a box with off-site switches the app ON; older apps get one press on both backup pages; over the quota the page names the largest apps. Measured first: over the quota nothing but the ruled retention runs — no history is deleted (now pinned). Live on 9202: the press put both older apps ON; a fresh `glance` joined by itself. The size card is unit + parity proven (a household NAS target has no quota). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partA/`. | **CLOSED 2026-09-30 — controller v0.283.0** |
|
||||
| **R-721** | **[P2-MEDIUM] The household presses Stop during a whole-guest backup, and the backup starts the app again.** MEASURED 2026-09-29 19:37 UTC on a fresh box: the first off-site whole-guest backup quiesced three apps at 19:37:01; the household pressed „Leállítás" on actualbudget at 19:37:16 and the controller recorded `desired state for actualbudget recorded as "stopped"`; the same second the backup's early resume ran `unquiescing … restarting 3 stack(s)` and `Starting stack: actualbudget`. The app page then read „Fut" while `app.yaml` kept `desired_state: stopped` (so the next reboot would stop it). The household had to press Stop again, and the removal was refused „Az alkalmazáson mentés vagy visszaállítás fut" for ~4 min (honest). Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase1/step10-stop-undone-by-quiesce.log`. **Fix direction:** unquiesce restarts only stacks whose desired state is still running; a test pins it (Stop during a quiesce → still stopped after the resume). **FIXED 2026-09-30 (controller v0.283.0 + v0.283.1):** the quiesce resume, the nightly volume dump, the update leg and the startup crash recovery skip an app the household stopped. **v0.283.0 was wrong in production** — its adapter did not answer the question and the dump restarted the app 8 s after the Stop (measured live on 9202); v0.283.1 wires it and pins the PRODUCTION types. Live on 0.283.1: `paperless-ngx NOT restarted after the volume dump: the household stopped it meanwhile`. Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partE/`. | **CLOSED 2026-09-30 — controller v0.283.1, proven live** |
|
||||
| **R-722** | **[P2-MEDIUM] The volunteer guide is stale in five places a volunteer reads literally.** MEASURED 2026-09-29 walking `runbooks/VOLUNTEER-first-hour.md` on golden 0.282.0: (1) §8 says BookStack's login is `admin@admin.com / password` — since the random first password (controller v0.280.0) that login is REFUSED; the app page says the right thing („a Beállítások oldalon látható első jelszó"). (2) §7 says the recovery-code bar comes „néhány perccel" after setup — measured ~15 min after enrolment (the off-site tier arrives with the agent's next 15-minute host report). (3) Nothing says to switch each app's off-site copy on (R-720). (4) §2 says a 2 GB USB stick and Rufus; the download page says at least 4 GB and Balena Etcher. (5) The operator part says no button is needed for the link (R-719), and §5 names the mail „Elindult a Felhom szervered" while a customer with an earlier box gets „Új beállító kód — újratelepült a szervered … A korábbi jelszavad már nem érvényes". Also, from the screens: the installer pre-selects `/dev/sda` (the guide says it never chooses), and its Summary screen rests on „Previous", not „Install". **Fix:** a guide edit; the operator approves the text. **REWRITTEN 2026-09-30** (operator: CC rewrites as measured): both guides — BookStack's generated password on the app page, the ~15 min recovery-code delay, apps off-site by default + the one-press offer, a 4 GB stick, the returning customer's mail and the expired-link button, the pre-selected disk, focus on „Previous", and the auto-reboot (a sixth stale line: with the defaults the „reboot now?" prompt never appears). Operator part: set the e-mail language, check the domain's DNS for an earlier box's records, the tunnel route with No TLS Verify. Every changed line dated. | **CLOSED 2026-09-30** |
|
||||
| **R-723** | **[P3-LOW] A fresh box sends the operator two mails on day one that describe nothing wrong.** MEASURED 2026-09-29: `Operator email sent for tester-1/node_recovered` 2 s after the new box's first controller report (the customer's previous box had been silent 12 days — the new box is not a recovery); and `backup_tier_skipped (warning) — Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed)` → operator mail at 19:27 UTC, 7 min after enrolment, because the first whole-guest run fired before the off-site tier's descriptor arrived (applied ~19:35 UTC, backed up fine at 19:37). Operator-only, so no household is alarmed, but an operator learns to ignore both. **Fix direction:** `node_recovered` not for a host enrolled < N min ago; the skipped tier inside the first hour after enrolment is `info`, not a mailed warning. **FIXED 2026-09-30 (hub v0.126.0):** no `node_recovered` when the customer's host was enrolled after the outage began; `backup_tier_skipped` in a box's first hour recorded, not mailed. Real recoveries and old boxes' skips still mail (controls). Unit + red-proof (RP40, RP41); the live proof is Tester-2's first hour. | **FIXED — hub v0.126.0; live proof at the first real install; owner: CC** |
|
||||
| **R-724** | **[P3-LOW] The status pages disagree with each other in small ways a household notices.** MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw `2026-09-29T19:22:30Z`; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). **FIXED 2026-09-30 (controller v0.283.0):** Beállítások shows the real schedule and a local time; the dashboard's last backup is local time (R-500); „az előző N órája készült" replaces „0 órája". **NARROWED — remaining:** „Helyi cím (LAN)" / „Átjáró" read through the file-sharing container, so a box without it says „nem állapítható meg" — not a text fix (the read needs another path). | **NARROWED — the LAN/gateway read only; owner: CC** |
|
||||
| **R-725** | **[P3-LOW] Small copy slips on the first-hour path.** MEASURED 2026-09-29: the self-bind PAGE says the passphrase was received „a beállításkor" while the mail and console say „a Felhom üzemeltetőjétől" (R-497 unified the mail and console, not the page); the recovery-code wizard addresses the household formally („Írja fel", „adja meg") while every other screen says „te"; the console's linked banner ends „a doboz össze van kötve. V" (a stray glyph); a gated app answers a phone app's API call with English JSON „this app is waiting for its first setup" (the browser gets the Hungarian gate page). **FIXED 2026-09-30:** the recovery wizard speaks „te" (controller v0.283.0; formal ceiling 18 → 17); the bind page says „a Felhom üzemeltetőjétől kaptál" (hub v0.126.0). **NARROWED — remaining:** the console's stray „V" (the installer/agent's banner, not these repos' text); the gate's English JSON to a phone app (the app shows its own error; left, deliberately); and the expired bind page still says „kérj újat az ügyfélszolgálattól" ABOVE the new „Új linket kérek" button (hub copy, next hub release). | **NARROWED — three small copy items; owner: CC** |
|
||||
| **R-726** | **[P2-MEDIUM] A new box for a customer who had one before makes NO off-site copy on night one: the old repository is found orphaned, and the fix is a button nobody pointed the household to.** MEASURED 2026-09-30 00:15 UTC on the new-household drill box (`tester-1`, whose previous box was deleted 2026-09-17): `[offbox] offsite repo ORPHANED — remote holds backups written under a previous, no-longer-available key; runs will skip until reset` → `offbox_repo_orphaned` (warning) to the household's timeline and an operator mail; `offsite-integrity` then checked nothing. The household's page is honest („A távoli tároló másik kulccsal készült mentéseket tartalmaz … Új távoli mentés indítása…", old history set aside, never deleted), but the evening before, the recovery-code ceremony and the off-site page raised nothing, and the guide does not mention it. The hub re-issued the off-site credentials on re-enroll by itself; it could have known the repository would orphan. Customer data was never at risk (the old repository is untouched); the household simply has no off-site copy until someone presses the button. **Fix direction:** offer the reset at the recovery-code ceremony when the repository already holds another key's snapshots, or the hub's re-enroll re-issue sets the old history aside the same way (it is the same move-aside), and the guide says so. | **READY — rank P2-MEDIUM; owner: CC (design first) / operator (which)** |
|
||||
| **R-727** | **[P2-MEDIUM] The whole-guest restore test picks a PREVIOUS box's archive, fails on its key, and the household sees a bare ✗ labelled with the wrong tier.** MEASURED 2026-09-30 01:52 UTC on the drill box: the agent's restore test chose `felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z` („newest settled archive … has not been proven") — written by a drill box of 2026-09-16, still in the customer's ep0 namespace next to tonight's own `2026-09-29T19:37:07Z` — and failed `wrong key - unable to verify signature since manifest's key 6b:ca:5f:3f… does not match provided key de:51:7a:18…`. One operator mail (`restore_test_failed`); the hub then re-logged the stored failure at every 15-minute report for 5 hours. The household's backups page read „✗ Visszaállítás ellenőrizve 2026-09-30 03:52 — Helyi tároló (local)" — the local tier had not failed; the pbs tier had, and the page says neither which nor why. Root: host delete leaves the old box's archives in the namespace (R-526's shape). **Fix direction:** the restore test skips (and reports as foreign) archives whose key fingerprint is not this host's; the page names the tier that failed. | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-727** | **[P2-MEDIUM] The whole-guest restore test picks a PREVIOUS box's archive, fails on its key, and the household sees a bare ✗ labelled with the wrong tier.** MEASURED 2026-09-30 01:52 UTC on the drill box: the agent's restore test chose `felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z` („newest settled archive … has not been proven") — written by a drill box of 2026-09-16, still in the customer's ep0 namespace next to tonight's own `2026-09-29T19:37:07Z` — and failed `wrong key - unable to verify signature since manifest's key 6b:ca:5f:3f… does not match provided key de:51:7a:18…`. One operator mail (`restore_test_failed`); the hub then re-logged the stored failure at every 15-minute report for 5 hours. The household's backups page read „✗ Visszaállítás ellenőrizve 2026-09-30 03:52 — Helyi tároló (local)" — the local tier had not failed; the pbs tier had, and the page says neither which nor why. Root: host delete leaves the old box's archives in the namespace (R-526's shape). **Fix direction:** the restore test skips (and reports as foreign) archives whose key fingerprint is not this host's; the page names the tier that failed. **FIXED 2026-09-30 (agent v0.138.0 + decision 51):** measured — a PBS archive carries its key FINGERPRINT (PVE content `encrypted`), not a host id; the storage carries its own (`encryption-key`). The restore test skips an archive written with another key (logged by name). The ✗ card names the tier (controller v0.283.0) — **and the 2026-09-30 claim "the page blames the local tier" was my misreading**: the „Helyi tároló (local)" after the ✗ was the next section's heading. ep0: the three drill archives in `tester-1` removed, other namespaces byte-identical. Delivered by signed jobs to both demo boxes (340 s); their due-check reads normally on 0.138.0. Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partC/`. | **CLOSED 2026-09-30 — agent v0.138.0** |
|
||||
| **R-728** | **[P3-LOW] A customer created with one press was created TWICE, and the first of its two connect mails holds a dead link.** MEASURED 2026-09-30 on `Tester-2`: the hub logged `Customer config created: Tester-2` twice in the same second and two self-bind mints (hashes `c40df008…`, `6a1cbef4…`); a mint replaces the previous link (single-active), so one of the two identical mails the tester received answers „expired". Cause not established (a double form submit, or the handler run twice). **Fix direction:** make the create idempotent within a few seconds (or disable the button on submit), and pin it. The workaround for the tester is in STATUS. | **READY — rank P3-LOW; owner: CC (hub)** |
|
||||
| **R-729** | **[P3-LOW] An off-site target, once saved on the page, cannot be removed through the product.** MEASURED 2026-09-30 on 9202: `/backup/offbox/config` refuses an empty address and no route clears the target; the session removed its throwaway target from `settings.json` by hand, with the controller stopped (harness teardown on a scratch guest). A household that tries its own NAS and gives up keeps a disabled target forever. **Fix direction:** a „Távoli mentési cél törlése" press that clears the target (never the repository). | **READY — rank P3-LOW; owner: CC (controller)** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user