STATUS: the 2026-10-10 decision sheet (S1-S38, T1-T19) and the register-shrink report; R-903 live read-back
gates / gates (push) Successful in 5m55s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-10 15:57:06 +02:00
parent b8db1383bf
commit a263061183
3 changed files with 169 additions and 2 deletions
+83 -2
View File
@@ -2,8 +2,89 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.**
**Updated 2026-10-10: hub 0.145.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller
0.305.0. The open-items list is at 138. Reports: `REPORT-new-apps-2026-10-10.md`, `REPORT-release-2026-10-10.md`, `REPORT-break-the-circle-2026-10-09.md`, `REPORT-dooplex-survival-2026-10-09.md`, `REPORT-day4-2026-10-09.md`.**
**Updated 2026-10-10 (afternoon): hub 0.145.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller
0.307.0. The open-items list is at 125 (was 138). Reports: `REPORT-register-shrink-2026-10-10.md`, `REPORT-new-apps-2026-10-10.md`, `REPORT-release-2026-10-10.md`, `REPORT-break-the-circle-2026-10-09.md`, `REPORT-dooplex-survival-2026-10-09.md`, `REPORT-day4-2026-10-09.md`.**
## Saturday 2026-10-10 (late afternoon): the list is shorter, and one sheet for you
- **The list went from 138 to 125.** 13 items closed, each with a check on a real box, the live website or the off-site server. No new item was opened.
- **Controller 0.307.0 runs on demo-hp, demo-felhom and Tester 1.** It removes an old fallback that no box needs any more, and it fixes a bug I found today: after a controller restart, the backup page listed each database several times. I checked the fix on demo-hp: 6 databases, 6 rows.
- **The website:** the phone menu now works with JavaScript off (it did not open before), and the eight dashboard pictures show today's screens. A new website check makes sure every app has its logo file.
- **The app catalog:** its slow container check now refuses to run on DooPlex, so it cannot act on the production machine again.
- **Two checks wait for tomorrow:** tonight's backup on Tester 1 must still run after the daytime press I made (it proves a fix from Thursday), and on Sunday evening the hub must mail you that Tester 2 has never made an off-site copy (it is due then, not before).
- **Good news found on the way:** DooPlex's saved login for the image registry works again, so builds push normally.
- **Not done:** no hub release and no agent release (none was needed). The R-925 work and its Secrets were not touched.
### The decision sheet (2026-10-10 afternoon) — answer „all as picked", or name the numbers you change
Each line: the question, **my pick**, what the other option costs, what happens if you do nothing. The full reasoning is in each row of `OPEN-ITEMS.md`.
| # | Row | Question | Pick | The other option | If you do nothing |
|---|---|---|---|---|---|
| S1 | R-922 | When a household clears its mail address, do we also stop mails to the address you registered for them? | **No** — the registered address is the contract contact; the privacy notice says so | B (clear both): you lose your only mail contact; kernel notices go nowhere | Behaviour is the pick, but the privacy notice does not say it |
| S2 | R-831 | The Hetzner storage token that leaked on 2026-10-03 (Secret not changed since 2026-07-09) — rotate it? | **Rotate**, after R-925's rotation is finished (~5 min: new token, patch the Secret, restart the hub Deployment, delete the old) | A (accept): a token that can create/reset/delete sub-accounts stays valid | The token stays valid |
| S3 | R-870 | Tester 1's two leaked tokens (rulings 10-04/10-05: not now) — close as accepted? | **Close as accepted**; rotate when Tester 1 is retired | B (rotate now): ~15 min in Cloudflare + the hub | Same as the pick, without the record |
| S4 | R-242 | Build a check that the newest baked golden was also vouched? | **Yes**, a hub alarm, before the first external install (~1 hub session + an attended deploy) | A (no): a forgotten vouch is found only when a fresh install lands on an old controller | Same as A |
| S5 | R-388 | Redesign the household's notification settings to „only what you can act on" before the first sale? | **Yes, design now**, build in slices (several sessions) | B (after the first customers): the first customer gets the per-detector toggle page | The page grows by one toggle per new detector |
| S6 | R-244 | May CC remove a deleted customer's id from shared app-log issue rows (and sweep once)? | **Yes**, before the first real customer is deleted (1 hub session + an attended deploy) | B (accept): one id per torn-down customer stays in aggregate rows, no secret | Same as B |
| S7 | R-264 | The six owed hub readers are already decided (2026-08-12) — may CC own the row and build them when the queue is idle? | **Yes, CC owns it** (~1 session per group, attended hub deploys) | B (drop the emitters): a two-repo change that breaks the host-report golden | Six facts stay sent and unread |
| S8 | R-698 | A restore of an app version whose maker deleted the image cannot start — accept it? | **Accept**: write it as a known limit in `07` §6.6 and close | B (mirror every installed image to DooPlex): storage + bandwidth, a new part on the recovery path | Behaviour is the pick, unwritten |
| S9 | R-255 | Build one test that renders every dashboard page with planted secrets and fails if any shows? | **Yes** (CC, one session; ~23 page fixtures) | B (no): a secret under a neutral page key can still ship unseen | A fourth leak of the R-254 shape would not be caught |
| S10 | R-266 | Carry „disk reading failed" to the hub so it is not read as an empty disk? | **Yes**, in the next hub + controller release (two-repo change) | B (no): while the disk cannot be read, the hub sees 0 % and stays quiet | A missed alarm during a fault that has louder symptoms |
| S11 | R-333 | A healthy NVMe under load can pass 60 °C and show „Hiba" — how do we band NVMe heat? | **No heat band for NVMe**; rely on the drive's own critical warning (smallest change) | A (separate NVMe bands, warn 70): one small controller change + test | A busy NVMe can raise a false drive alarm |
| S12 | R-577 | Do guest share pages get the language globe (the visitor's own cookie)? | **Yes** (CC, ~1 hour, next controller release) | B (no): a foreign visitor sees the household's language | No harm today |
| S13 | R-327 | Which status does the code-naming arc get in the capability map? | **„built"** — fixed and shipped, no customer has used it yet; CC rewrites the title | B („walked"): an over-claim | The map keeps showing a fixed defect as open |
| S14 | R-230 | May dated history entries in the memory index keep version numbers? | **Yes**; close (a) and drop the pilot (c) | B (bulk-remove them, start the pilot): one session, less detail | The gate keeps warning; no harm |
| S15 | R-719 | The expired self-bind link now offers „Új linket kérek" — accept this shape and close? | **Accept and close** | B (another shape): a new hub change | The row stays open; behaviour as built |
| S16 | R-352 | The data-placement spec: hot data in the guest, bulk on a drive is by design — close? | **Close as by-design** (the visibility fix shipped; R-368 fixed the comment) | B (keep the 5-point spec open): a P4 row nobody works | Nothing changes for a household |
| S17 | R-526 | Build a „release only the PBS token" operation on ep0 for a host delete? | **No, close** — the R-511 adopt path already reuses the kept token | B (yes): a new operation on a protected box, one session | Kept tokens stay on ep0; no customer impact |
| S18 | R-213 | When do we design „see what a restore would change"? | **After the first customers** | B (now): several sessions | Same as the pick |
| S19 | R-288 | Give CC one session to restructure the capability map (211 KB; one status line per capability + a dated pointer)? | **Yes** (documentation only) | B (accept): the map keeps growing (~75 KB in two months) and nobody can re-verify it | Same as B |
| S20 | R-231 | Put DooPlex's own backup scripts (`/opt/backup/scripts`, 16 files, changed on the box only) under version control? | **Yes** — CC copies them into a repo + an install script, you present for the install | B (accept): a DooPlex rebuild has to rewrite them from memory | Same as B |
| S21 | R-887 | CI jobs are still lost (4 today, one of them this session's — re-run once, it passed) — change DooPlex to stop it? | **Yes**: raise the act-runner fetch timeout and rate-limit the public Gitea pages the crawler walks (two DooPlex changes, yours) | B (no): sessions re-run lost jobs once; red runs with no mail | A real failure can hide among the lost ones |
| S22 | R-886 | Alertmanager cannot write its state (root-owned volume, the pod runs as nobody) — may CC add `fsGroup: 65534` on DooPlex and prove a silence survives a restart? | **Yes** (one manifest line, CC with your word) | B (no): silences and send history die at every pod restart | Same as B |
| S23 | R-211 | Prometheus has no config reloader, so an alert-rule change waits for a manual reload — may CC add the reloader sidecar on DooPlex? | **Yes** (one manifest change, CC with your word) | B (no): every rule change needs a `POST /-/reload` by hand | A rule change can sit unread with no error |
| S24 | R-770 | Add Invidious to the catalog? | **No — close** | B (go): a rolling companion image, PostgreSQL 14 near end of life, YouTube can block the household's IP for 24 h | The idea waits; nothing breaks |
| S25 | R-771 | Add moonlight-web (game streaming) to the catalog? | **No — close** | B (go): a new LAN-only publishing model (UDP/WebRTC) | The idea waits; nothing breaks |
| S26 | R-782 | Glance (the link dashboard) has no login: anyone with the address sees it. Give it one? | **Yes, behind the family gate** like Grimmory (catalog change, CC) | B (no): the household's link list is public by address | Same as B |
| S27 | R-554 | Deleting the obsolete setup wizard makes `recovery-info.txt` (on every box) point at a route that no longer exists. What should that file say instead? | **Point at the dashboard's Restore page and the recovery code** (CC writes it, you approve the Hungarian) | B (keep the wizard): obsolete code stays reachable | The obsolete wizard stays reachable |
| S28 | R-562 | Hungarian numbers and dates on the dashboard: decimal comma and `2026. 10. 10. 15:04` everywhere? | **Yes, Hungarian style in Hungarian, ISO-like in English** (CC, one controller session) | B (leave as is): pages disagree with themselves | Same as B |
| S29 | R-250 | A new customer's off-site setup can fail on its first try (the DNS/IPv6 settle is ~100 s, the scan waits 60 s) and says only „error". May CC fix it? | **Yes**: the scan prefers the IPv4 address and waits longer, and the error says „safe to press again" (CC, hub, attended deploy) | B (leave): the first thing a new customer's setup does can fail looking like an outage | Same as B |
| S30 | R-246 | The hub's `stale_at` column changes behaviour, but nothing sets it and nothing shows it. Retire it or give it a visible setter? | **Retire it** (CC, hub, small) | B (a visible setter with evidence): one hub session | A trap only a database read can spring |
| S31 | R-913 | The Cloudflare token check sees what a token can READ, not what it can WRITE. Which guard? | **(a) you mint every customer token from one recipe; the hub checks nothing more** (free) | (b) the hub mints the token itself with your account token — a new privileged secret on the hub | The check stays as built; the limit stays written in `01` §7 |
| S32 | R-521 | When a household's drive is unplugged, does the household get a mail too (today only you do)? | **Yes, one mail per lost drive in the household's language** (CC, hub + controller) | B (no): the household learns it only on the dashboard | Same as B; the hub's cooldown fix (F6/F7) is CC's either way |
| S33 | R-402 | What should the hub's host page say about the off-site integrity check (verdict + depth are sent, nothing reads them)? | **One line: „Last full check: OK/FAILED, depth N, date"** on the host page (CC, hub) | B (drop the two fields from the wire): a two-repo change | Two facts stay sent and unread |
| S34 | R-290 | 20 of 28 green capability-map rows cite no evidence file. Fold this into the map restructure (R-288)? | **Yes, fold** (one row closes) | B (fix row by row now): longer | Same as today |
| S35 | R-209a | DooPlex's SSD2 move has never been proven across a reboot, and you ruled „no reboot". Close as accepted until a planned reboot? | **Close as accepted**; the next planned reboot checks it (the mechanism is proven) | B (keep watching): a row nobody can act on | Same as B |
| S36 | R-769 | Add Pinchflat (upstream paused, no image tag)? | **No — close** | B (go via a fork): an unmaintained base | The idea waits |
| S37 | R-884 | ArgoCD shows Prometheus OutOfSync on DooPlex (cause unknown). May CC read the diff and propose git-or-live? | **Yes, read only first**, then one line for you | B (leave): drift stays unexplained | Same as B |
| S38 | R-616 | The catalog-clone credential leak is fixed and live (no box's clone holds a credential, read 2026-10-10). Was the Gitea admin token it exposed rotated in the R-925 work? | **If yes: close** | If no: rotate it (Gitea → Settings → Applications) | The row stays open |
### Your own tasks (no decision — something only you can do)
| # | Row | What is true now | What you do | Does it block the first paying customer? |
|---|---|---|---|---|
| T1 | R-789 | Tandoor is offered with no written yes from its authors | Ask the Tandoor authors for written permission for a paid service | **Blocks the first paying customer** — without a yes, CC hides Tandoor first (one catalog line) |
| T2 | R-802 | The non-OSI licence table has had no lawyer's review | Send `audits/licences-2026-10-02/TABLE.md` to a lawyer | **Blocks the first paying customer** (your ruling, decision 66) |
| T3 | R-784 | SparkyFitness's licence forbids commercial use without the author's written permission; it is offered today | Ask the author for written permission | **Blocks the first paying customer** — without a yes, CC hides it first (one catalog line) |
| T4 | R-813 | The legal pages are a closed-test version, unreviewed | When the company exists: the lawyer's review (with R-802) and the full ÁSZF + imprint | **Blocks the first paying customer** (with R-802) |
| T5 | R-923 | The break-glass sheet is not known to be printed | Print it, write the keys on it, store it away from home; say „done" | Before the first paying customer: losing DooPlex would lose the keys to its own off-site copy |
| T6 | R-924 | The two signing keys (sheet lines S6a, S6b) are not known to be printed | Print and staple them; say „done" | Not a hard block |
| T7 | R-504 | `iso.felhom.eu/` answers 404 | Add a Cloudflare redirect to `felhom.eu/letoltes` (5 min), or say „drop" | Not a block |
| T8 | R-779 | Phone sign-in over mobile data never proven | 2 minutes on your phone while CC reads the logs | Not a block |
| T9 | R-862 | Tester 2 lacks its one-time bootstrap step | When Tester 2 is on: the tunnel + three commands (`runbooks/config-bundle.md`) | Not a block |
| T10 | R-883 | 9 DooPlex Deployments + 1 Pod run a moving tag (one is Felhom's own contact-mailer) | Pin them in homelab-manifests | Not a block |
| T11 | R-814 | The old Storage Box #611421 (~€4/month) is still paid; ep0 does not mount it | Delete it in the Hetzner console | Not a block; costs money |
| T12 | R-917 | The first Facebook post goes out Monday 12 Oct 19:00 | After it goes out, pin it; CC reads it back | Not a block |
| T13 | R-902 | No source for the contact-mailer | Look for the February folder on your Windows machine; if absent, say so and CC rebuilds it | Not a block |
| T14 | R-433 | „Can the MAIN Storage Box account read single files from a snapshot?" — Hetzner said „should be possible"; nobody has tried | Read one file from a snapshot with the main account (or give CC a main-account credential for that one read) | Not a block |
| T15 | R-132 | The hub operator password was printed into a session transcript (2026-07-31, again 2026-09-18) | Rotate `HUB_PW` (after R-925's work) | Not a block |
| T16 | R-908 | An old Resend key sits in `homelab-manifests` history | Check in the Resend console that it is revoked | Not a block |
| T17 | R-232 | DooPlex's own backups: items (c)–(h) of the 2026-08-06 survey are still open | Your DooPlex list; say which you want CC to take | Not a block |
| T18 | R-882 | Longhorn's instance-manager can go stale and block volume growth | Before growing any volume: check the instance-manager's age (or have CC check it) | Not a block |
| T19 | R-919 | The rebuilt Facebook cover is not yet checked in the phone app | Upload the rebuilt cover, look at it in the Facebook app, say „whole" | Not a block |
Not on the sheet, on purpose: R-832, R-904 and R-920 are deferred by your earlier rulings; R-243 is a dated check (above); R-925 belongs to another session.
## Saturday 2026-10-10 (afternoon): two new apps in the catalogue, and one that was stopped