decisions 44-45, FIRST-ADMIN link, register (R-701/R-702 closed, R-707..R-709), STATUS, report, evidence
gates / gates (push) Successful in 26s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-28 19:02:06 +02:00
parent aedaab8944
commit 56d6de9808
42 changed files with 653 additions and 27 deletions
+8
View File
@@ -15,6 +15,14 @@
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-09-28 evening — operator rulings (decisions 44, 45):** 44 — demo-hp's restore test uses `nvme-scratch` (the
> agent's storage role was granted there with the operator's word; proven). 45 — no app is published with a login a
> stranger knows: generated first password via `after_install:` where the app's own CLI can set it, else the page
> says the default. Controller v0.279.0 (after_install, known-login page rule, Part D empty-backup alarm, night chain,
> R-706). claper + bookstack fixed; 37 apps remain (R-707, several sessions, operator's choice). Report:
> `REPORT-logins-nvme-2026-09-28.md`.
> **2026-09-28 — decided by CC unattended (operator may reverse): `09` §3 decision 43 — claper → PostgreSQL 17, calcom → 18** (by
> each upstream; decision 42's rule; calcom's memory 768M → 1536M first, R-703 closed). Also that day: golden 0.276.0 baked + vouched with agent 0.137.0; controller
> v0.277.0 (kept data loads from the off-site copy, R-691 (2)); floor 0.277.0; R-702 (claper default admin, P1) and
+46
View File
@@ -0,0 +1,46 @@
# REPORT — 2026-09-28 evening: no app goes live with a login a stranger knows; demo-hp's restore test on the NVMe; an empty backup is an alarm; ep0; small leftovers
Architecture read first: `09` §3 (decisions 11–43, and the new 44–45), `01` §5, `03` (restore storage), `07` §6,
`06` + R-600. Controller **v0.279.0** (one release). Scope ruled mid-session by the operator: the brief as written for
all 40 apps, across several sessions; this session starts it and hands over.
## The Parts
| Part | Step | State | Note |
|---|---|---|---|
| A | audit of 53 apps | done | `app-catalog-felhom.eu/FIRST-ADMIN.md`; linked from `01` §5 and `09` decision 45. Measured: claper, bookstack, calibre-web, calcom (9202); bookstack, calibre-web, romm (demo-hp). The rest is marked "read" |
| A | measure every class-3/4 app on 9202 | **partly** | 4 of 39 measured today; the rest is R-707 |
| B1 | route (a)/(b) per app | **2 done** | claper, bookstack: fresh-install proof (default fails, generated works, wrong fails); claper after a restore too. calibre-web: route found, blocked by our generator (no special character) |
| B2 | route (c) sentence | done (mechanism) | shown for every template with `default_creds` and no working `after_install` — today bookstack(installed)/calibre-web/mealie/romm/wger/zipline pages; class-4 apps have no default to name |
| B3 | demo boxes read-only | done | demo-hp: bookstack + calibre-web defaults still log in (page now warns); romm's note is stale (401). Nothing changed |
| B4 | box-side mechanism | done | `after_install:` in v0.279.0, tests + red-proofs; failures recorded and shown, app stays running |
| C | demo-hp restore test on NVMe | done, **changed** | "set restore_storage, nothing else" did not work: 403 without a grant; operator approved the grant; one test passed in 8m46s; `local-lvm` unchanged. Next scheduled cycle: see below |
| D | empty-backup alarm | done | measured before: a WARN line only (demo-hp, yesterday). Built: digest once/app/tier/day + page sentence; tests + red-proofs; live negative on 9202 (no false alarm). Live positive not reproducible (its known cause, R-704, is fixed) |
| E | ep0 peer | done — **nothing to remove** | the peer was already gone (the hub's sync) |
| F1 | R-706 | done | v0.279.0, red-proofed; not seen live |
| F2 | R-705 controller half | done | proven on 9202 |
| G | night watch | not done | optional; not run (see STATUS) |
## Claims in the brief that turned out wrong
- **"Six apps ship a hard-coded default"** — wrong: 5 (bookstack, calibre-web, claper, mealie, wger); romm's and
zipline's notes are stale (romm's does not log in). And **34 more** have an open first-run screen.
- **"The box publishes every app on the internet at install"** — right (the tunnel's `*.domain` route).
- **"No architecture document covers default logins"** — right (only `10-localisation.md` mentions `default_creds`
for translation); the audit's home is now `FIRST-ADMIN.md`, linked from `01` and `09`.
- **"`nvme-scratch` can hold a restored guest"** — the disk and content types could; the agent could not use it
without a storage grant (403). Granted with the operator's word.
- **"The test-install peer is still on ep0"** — wrong: gone.
- **"Nothing flags a running app with an empty backup today"** — right for yesterday's controller (a WARN line only).
## Rows
Opened: R-707 (37 apps), R-708 (grafana `admin` fallback), R-709 (password fields in the page HTML). Closed: R-701,
R-702. Narrowed: R-705 (agent half left). Watching: R-706. R-600 annotated. Register 342 → 345 rows.
## Teardown
Machines: 9202 — claper, bookstack, calibre-web installed and removed through the product; back on the live catalog;
drill catalog = live. demo-hp 9201 — nothing changed (read-only logins). Host: demo-hp agent config
(`restore_storage`, saved `agent.json.pre-d44`) and one ACL grant — both kept (decision 44). Hub: floor 0.279.0. ep0:
read only.
+17 -21
View File
@@ -1,30 +1,26 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-28 afternoon. Both demo boxes run controller 0.278.0 and host agent 0.137.0. Hub 0.125.0. New installs get golden 0.276.0 with agent 0.137.0.**
**Updated 2026-09-28 evening. Both demo boxes run controller 0.279.0 and host agent 0.137.0. Hub 0.125.0. New installs get golden 0.276.0 with agent 0.137.0.**
**Decisions I took on my own** (you may reverse each):
1. **claper goes to PostgreSQL 17, calcom to 18.** Each follows what its own makers run: claper's makers use 15, so 17 (no disk-layout change); calcom's makers use 18.
2. **calcom gets 1536 MB of memory instead of 768 MB.** At 768 MB it was killed at every start, so it could never run. At 1536 MB its own use peaked at about half.
**Decisions today** (yours, recorded): the HP box's restore test uses the big NVMe disk. No app goes live with a login a stranger knows.
**What I did, and it worked.**
- **The weekly golden is built and vouched** (0.276.0, with host agent 0.137.0). A test install in the throwaway machine came up on it with the right versions. The test customer is removed from the hub.
- **"Use my kept data" can now load the database from the off-site copy.** Proven on the HP box: the page named "the off-site copy, 2026-09-28 15:40", the app came back with its account and its files.
- **The first real restore from the off-site copy worked.** Account back, a later change gone (as it must be), same version.
- **Two more apps can move to a new database version: claper and calcom.** Each passed the throwaway test machine and the scratch box.
- **One more fix (controller 0.278.0): a "stopped" mark from an app's earlier install no longer sticks to a new install.** On the HP box such a mark from 13 September made the backup skip a freshly installed app. That app had no real backup at all. This is my second controller release today; the rules say one. It blocked the off-site proof, and nothing in the product could clear the mark.
- **The HP box's night:** the database step ran, paperless moved to PostgreSQL 18 by itself (all rows equal), then the full-system backup ran. Nothing had to wait for anything.
- **claper and bookstack now get their own random first password.** The box sets it right after install and shows it on the app page. The old shared password no longer works. Tested on fresh installs; for claper also after a restore.
- **Every app with a known default login now says so** on the install page and the app page: "This app starts with a known, shared password: … Change it right after the install."
- **The HP box's full-system restore test works now, on the NVMe disk.** It needed one Proxmox permission, which you approved. The test passed in 9 minutes; the other disk was not touched.
- **A running app whose backup holds no data now raises an alarm** for us and a sentence on its backup page. Yesterday's nextcloud case was silent.
- **A debug button runs the night's backup chain now**, in the night's order. Tested on the scratch box.
- **Removing an app "with its backups" now also deletes its off-site check copy.**
- **ep0:** the test install's network link was already gone. Nothing to remove.
**What broke, or is not done.**
- **claper creates an admin account with the public password "claper" on every install.** Anyone who knows that could log in. No box runs claper now. I did not change it; you choose the fix (below).
- **The HP box's full-system restore test still cannot run.** Freeing space gave 26.6 GB; it needs 31 GB. It refuses safely, and the hub sees each refusal.
- **The demo-felhom box's off-site step of last night is not readable** (today's upgrades erased the logs). Not a fault seen; just not read.
- **Small gaps filed:** removing an app with its backups leaves its 1 GB off-site check copy; there is no button to run the whole night now.
**What is not done.**
- **37 apps still start with a login a stranger can take.** 3 have a known default (calibre-web, mealie, wger). 34 let the first visitor create the admin. You chose to fix them over several sessions. The list and the route for each are written down.
- **On the HP box, the installed bookstack and calibre-web still accept their default passwords.** I changed nothing there, as the brief said. Their pages now warn.
- **romm's page shows a default login that does not exist.** Filed with the 37.
- **The night watch did not run.** It was optional.
**Rows.** 5 opened, 2 closed. The list went from 337 to 342 rows.
**Rows.** 3 opened, 2 closed. The list went from 342 to 345 rows.
**What needs you.**
1. **The claper admin password:** (A) remove claper from the catalog until it is fixed (I recommend this), or (B) have the box change that password after install. If you do nothing, a new claper install keeps the public password.
2. **The HP box's restore test:** (A) let the test use the big NVMe disk (I recommend this; it has 880 GB free), or (B) accept that this small box cannot test it. If you do nothing, it refuses every 6 hours.
3. **D4** (what a restore brings back when the backup is older): A stays in force. Nobody is blocked.
4. **The image-copy question** (a maker deletes an old version): if you do nothing, nothing changes.
5. **From before:** Peti's box in the project text; Peti's Cloudflare leftovers; the old Storage Box. If you do nothing, they stay.
1. **The installed bookstack and calibre-web on the HP box:** (A) I change their admin passwords on the box and show them to you (I recommend this; they are on the internet), or (B) leave them. If you do nothing, anyone who knows those defaults can log in to the two demo apps.
2. **D4, the image copies, the Peti leftovers:** unchanged. If you do nothing, nothing changes.
@@ -119,6 +119,11 @@ credentials.
| guest ↔ Proxmox host | **(none direct)** | the guest holds no Proxmox creds; all via the agent | — |
| hub ↔ Cloudflare API | geo-restriction WAF (enforcement) | the **hub** holds the CF API token; reconciles geo desired-state → WAF | the customer's zone/WAF |
**Every app is on the internet from its first minute** (`*.domain` through the tunnel), so an app's FIRST admin login is
a trust boundary too: a default password, or a "first visitor creates the admin" screen, is open to a stranger until the
household acts. Rule and per-app status: `09` §3 decision 45 and `app-catalog-felhom.eu/FIRST-ADMIN.md` (the audit
of all 53 apps, 2026-09-28).
---
## 6. Enrollment & identity
+1 -1
View File
@@ -317,7 +317,7 @@ per Part 1: **snapshot** (LVM-thin, transient, whole-guest rollback — not a ba
> vzdump log's "Total bytes written" or the PBS snapshot size. **Never the archive file:** 9201's file
> was 6.9 GB and its restore wrote 22.6 GB, so "file × 1.2 + 5 GiB" would have let that test run.
> - **Off the tested guest's pool** when another storage is eligible (active, `rootdir`, and the agent
> holds `Datastore.AllocateSpace` there) and fits. demo-hp has none: `nvme-scratch` carries no grant.
> holds `Datastore.AllocateSpace` there) and fits. demo-hp had none until 2026-09-28: `nvme-scratch` carried no grant. **Since then (decision 44):** demo-hp's `restore_storage` is `nvme-scratch`, with the agent's `FelhomAgentStore` role granted there (user + token) — one restore test passed in 8m46s, `local-lvm` untouched (`audits/logins-nvme-2026-09-28/C/`). Saved config: `/etc/felhom-agent/agent.json.pre-d44`.
> - **Unknown refuses.** A refusal is the test's result — `pass=false`, `skipped`, "skipped: not enough
> space on …" — so the hub raises `restore_test_failed`; it is never a pass and never dropped.
> - **Leftovers on a timer.** A failed scratch teardown and the stale-lock sweep (R-673) run every 10
@@ -487,6 +487,16 @@ R-636's louder repeated alarm.
runs untagged `postgres` at `/var/lib/postgresql` (Prisma 6.16.1 in v6.2.0). First its memory limit had to be fixed
(768M → 1536M, R-703); then both venues proved it (bench: 122 tables equal, `anon` 61.5 %; box: converted through the
guarded Update in 68 s), catalog `037f956`. Reversible: the catalog pins the major.
44. **demo-hp's scheduled restore test restores onto `nvme-scratch`** — *operator ruling 2026-09-28 (R-701 option (b)).*
`local-lvm` (53.9 GiB, the guest's own pool) cannot hold a 31 GiB restore; the NVMe can (~880 GiB free). Needed, and
added with the operator's word: the agent's storage role on `/storage/nvme-scratch` (user + token) — without it
Proxmox refused the restore with 403 `Datastore.AllocateSpace`. Proven: one restore test passed there in 8m46s,
`local-lvm` untouched. `03-host-agent.md`, `audits/logins-nvme-2026-09-28/C/`.
45. **No app is published with a login a stranger knows** — *operator ruling 2026-09-28 (R-702 widened).* Where the box
can set the first admin password, it generates one at install and shows it on the app page (`after_install:`,
controller v0.279.0). Where it cannot, the app stays in the catalog; the install dialog and the app page say what
the default login is and to change it at once. Operator: *"I don't think we should exclude apps if we can't change
the first PW."* The per-app audit and status: `app-catalog-felhom.eu/FIRST-ADMIN.md`.
---
@@ -0,0 +1,10 @@
BEFORE: default password works:
test password: generated in the guest, length 24, not printed
FELHOM_AFTER_INSTALL_OK
AFTER: default password works:
AFTER: the new password works:
control: a wrong password:
after the change: the default password claper works: ANS=false
control: a wrong password works: ANS=false
@@ -0,0 +1,2 @@
before: gitea.dooplex.hu/admin/felhom-controller:0.277.0
gitea.dooplex.hu/admin/felhom-controller:0.279.0 Up 25 seconds (healthy)
@@ -0,0 +1,12 @@
app page shows a first password for ADMIN_PASSWORD: True length 30
app page: default-login card shown: False | known-login warning shown: False
NEGATIVE the known default admin@claper.co / claper: ANS=false
POSITIVE the generated first password from the app page: ANS=true
CONTROL a wrong password: ANS=false
after_install:
at: "2026-09-28T16:43:07Z"
ok: true
control: /apps/claper loaded: 44661 bytes, has the tagline: True | has the first-steps line: True
default-login card (Alapértelmezett belépés) shown: False
@@ -0,0 +1,17 @@
before the restore: the page shows a first password: True
18:45:38 [R] restoring claper from snapshot 'helyi' (of 1 offered)
18:45:38 [R] POST /backup/restore -> HTTP/2 302 ['location: /backups/restore?flash=flash.restore.started']
18:45:38 + 0.0s restore (True, None, None)
18:46:06 + 28.4s restore (False, None, None)
18:46:06 [R] after restore: state=running hold=None phase=None
restore: {'ok': True, 'http': 'HTTP/2 302', 'seconds': 28.4, 'state_after': 'running'}
after the restore: the page shows a password value: True | same as before: True
after the restore: the page says the backup's login is the one that works: False
NEGATIVE the known default: ANS=false
the first password from BEFORE the restore: ANS=true
CONTROL wrong: ANS=false
after_install:
at: "2026-09-28T16:43:07Z"
ok: true
@@ -0,0 +1,8 @@
bookstack admin@admin.com / password : 302 -> / | control wrong: 302 -> /login
calibre-web admin / admin123 : 302 -> / | control wrong: 200 ->
romm admin / admin (POST /api/login) : 401 | control wrong: 401
bookstack | page bytes 43697 | warning: Ez az alkalmazás egy ismert, közös jelszóval indul: admin@admin.com / password. Telepítés után azonnal változtasd meg.
calibre-web | page bytes 42081 | warning: Ez az alkalmazás egy ismert, közös jelszóval indul: admin / admin123. Telepítés után azonnal változtasd meg.
romm | page bytes 41671 | warning: Ez az alkalmazás egy ismert, közös jelszóval indul: admin / admin. Telepítés után azonnal változtasd meg.
docmost | page bytes 43690 | warning: None
@@ -0,0 +1,4 @@
## 2026-09-28T16:46:28Z floor 0.278.0 -> 0.279.0 (min_agent 0.131.0 from the v0.279.0 header)
HTTP/1.1 303 See Other
Location: /configuration?flash=floor_set
name="min_controller_version" value="0.279.0"
@@ -0,0 +1,9 @@
BEFORE: default admin@admin.com / password -> 302->/ | control wrong -> 302->/login
--email[=EMAIL] The email address for the new admin user
--name[=NAME] The name of the new admin user
--password[=PASSWORD] The password to assign to the new admin user
--generate-password Generate a random password for the new admin user
--initial Indicate if this should set/update the details of the initial admin user
command output: The default admin user has been updated with the provided details!
AFTER: default -> 302->/login | new -> 302->/ | control wrong -> 302->/login
@@ -0,0 +1,10 @@
app page shows a first password: True
app page loaded: True | default-login card shown: False | warning shown: False
NEGATIVE the known default admin@admin.com / password: 302->/login
POSITIVE the generated first password: 302->/
CONTROL wrong: 302->/login
after_install:
at: "2026-09-28T16:53:15Z"
ok: true
@@ -0,0 +1,24 @@
BEFORE: default admin / admin123 -> 302->/ | control wrong -> 200->
uid=1000(abc) gid=1000(abc) groups=1000(abc),100(users)
-rw-r--r-- 1 1000 1000 200704 Sep 28 18:55 /config/app.db
command output: [2026-09-28 18:56:04,711] INFO {cps:92} ProxyFix configured to trust 1 proxy(ies) for X-Forwarded-* headers [2026-09-28 18:56:04,789] INFO {cps.ub:1066} Migrating system magic shelves... [2026-09-28 18:56:04,935] INFO {cps:140} SESSION_COOKIE_SECURE set to False (Standard/LDAP login) Password doe
AFTER: default -> 302->/ | new -> 200-> | control wrong -> 200->
Password doesn't comply with password validation rules
/app/calibre-web-automated/cps/helper.py:796:def valid_password(check_password):
/app/calibre-web-automated/cps/helper.py-797- if config.config_password_policy:
/app/calibre-web-automated/cps/helper.py-798- verify = ""
/app/calibre-web-automated/cps/helper.py-799- if config.config_password_min_length > 0:
/app/calibre-web-automated/cps/helper.py-800- verify += r"^(?=.{" + str(config.config_password_min_length) + ",}$)"
/app/calibre-web-automated/cps/helper.py-801- if config.config_password_number:
/app/calibre-web-automated/cps/helper.py-802- verify += r"(?=.*?\d)"
/app/calibre-web-automated/cps/helper.py-803- if config.config_password_lower:
/app/calibre-web-automated/cps/helper.py-804- verify += r"(?=.*?[\p{Ll}])"
/app/calibre-web-automated/cps/helper.py-805- if config.config_password_upper:
/app/calibre-web-automated/cps/helper.py-806- verify += r"(?=.*?[\p{Lu}])"
/app/calibre-web-automated/cps/helper.py-807- if config.config_password_character:
/app/calibre-web-automated/cps/helper.py-808- verify += r"(?=.*?[\p{Letter}])"
/app/calibre-web-automated/cps/helper.py-809- if config.config_password_special:
/app/calibre-web-automated/cps/helper.py-810- verify += r"(?=.*?[^\p{Letter}\s0-9])"
@@ -0,0 +1,18 @@
18:56:57 [X] stop -> 200 {'ok': True, 'message': 'Stack calibre-web stop completed'}
18:57:02 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/calibre-web tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghaj
18:57:02 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
18:57:29 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'calibre-web', 'volumes_removed': ['calibre-web_calibre_web_config'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'back
18:57:37 [X] after remove: deployed=False leftovers='/opt/docker/stacks/calibre-web'
felhom-controller filebrowser paperless-postgres paperless-redis paperless-webserver privatebin traefik
0
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
sync_interval: 15m
token: <redacted>
username: ""
hub:
0
drill = live ccd17da
@@ -0,0 +1,20 @@
# Read-only: one login attempt per app with its KNOWN catalog default, through the app's own login form,
# from inside the guest to traefik (https://127.0.0.1, Host header). A wrong-password control per app.
set -u
DOM=enkisfelhom.hu
J=/tmp/b3jar.$$
try_bookstack() { # $1 password
rm -f $J; T=$(curl -sk -c $J -b $J -H "Host: wiki.$DOM" https://127.0.0.1/login | grep -o 'name="_token" value="[^"]*"' | head -1 | sed 's/.*value="//;s/"$//')
curl -sk -o /dev/null -w '%{http_code} -> %{redirect_url}' -c $J -b $J -H "Host: wiki.$DOM" -X POST --data-urlencode "_token=$T" --data-urlencode "email=admin@admin.com" --data-urlencode "password=$1" https://127.0.0.1/login | sed "s|https\?://[^/]*||"
}
try_calibre() {
rm -f $J; T=$(curl -sk -c $J -b $J -H "Host: books.$DOM" https://127.0.0.1/login | grep -o 'name="csrf_token" value="[^"]*"' | head -1 | sed 's/.*value="//;s/"$//')
curl -sk -o /dev/null -w '%{http_code} -> %{redirect_url}' -c $J -b $J -H "Host: books.$DOM" -X POST --data-urlencode "csrf_token=$T" --data-urlencode "username=admin" --data-urlencode "password=$1" --data-urlencode "remember_me=on" https://127.0.0.1/login | sed "s|https\?://[^/]*||"
}
try_romm() {
curl -sk -o /dev/null -w '%{http_code}' -u "admin:$1" -H "Host: jatek.$DOM" -X POST https://127.0.0.1/api/login
}
echo "bookstack admin@admin.com / password : $(try_bookstack password) | control wrong: $(try_bookstack wrongxyz123)"
echo "calibre-web admin / admin123 : $(try_calibre admin123) | control wrong: $(try_calibre wrongxyz123)"
echo "romm admin / admin (POST /api/login) : $(try_romm admin) | control wrong: $(try_romm wrongxyz123)"
rm -f $J
@@ -0,0 +1,12 @@
set -u
DOM=enkisfelhom.hu; J=/tmp/bsjar.$$
login() { rm -f $J; T=$(curl -sk -c $J -b $J -H "Host: c-bs.$DOM" https://127.0.0.1/login | grep -o 'name="_token" value="[^"]*"' | head -1 | sed 's/.*value="//;s/"$//')
curl -sk -o /dev/null -w '%{http_code}->%{redirect_url}' -c $J -b $J -H "Host: c-bs.$DOM" -X POST --data-urlencode "_token=$T" --data-urlencode "email=admin@admin.com" --data-urlencode "password=$1" https://127.0.0.1/login | sed "s|https\?://[^/]*||"; }
for i in $(seq 1 30); do [ "$(curl -sk -o /dev/null -w '%{http_code}' -H "Host: c-bs.$DOM" https://127.0.0.1/login)" = 200 ] && break; sleep 5; done
echo "BEFORE: default admin@admin.com / password -> $(login password) | control wrong -> $(login wrongxyz123)"
docker exec bookstack sh -c 'php /app/www/artisan bookstack:create-admin --help 2>&1 | grep -E "initial|email|password|name" | head -8'
NEWPW=$(head -c 300 /dev/urandom | tr -dc A-Za-z0-9 | head -c 24)
OUT=$(docker exec bookstack php /app/www/artisan bookstack:create-admin --email=admin@admin.com --name=Admin --password="$NEWPW" --initial 2>&1)
echo "command output: $(echo "$OUT" | tr '\n' ' ' | cut -c1-200)"
echo "AFTER: default -> $(login password) | new -> $(login "$NEWPW") | control wrong -> $(login wrongxyz123)"
unset NEWPW; rm -f $J
@@ -0,0 +1,12 @@
set -u
DOM=enkisfelhom.hu; J=/tmp/cwjar.$$
login() { rm -f $J; T=$(curl -sk -c $J -b $J -H "Host: c-cw.$DOM" https://127.0.0.1/login | grep -o 'name="csrf_token" value="[^"]*"' | head -1 | sed 's/.*value="//;s/"$//')
curl -sk -o /dev/null -w '%{http_code}->%{redirect_url}' -c $J -b $J -H "Host: c-cw.$DOM" -X POST --data-urlencode "csrf_token=$T" --data-urlencode "username=admin" --data-urlencode "password=$1" https://127.0.0.1/login | sed "s|https\?://[^/]*||"; }
for i in $(seq 1 40); do [ "$(curl -sk -o /dev/null -w '%{http_code}' -H "Host: c-cw.$DOM" https://127.0.0.1/login)" = 200 ] && break; sleep 5; done
echo "BEFORE: default admin / admin123 -> $(login admin123) | control wrong -> $(login wrongxyz123)"
docker exec calibre-web sh -c 'id abc 2>&1 | head -1; ls -ln /config/app.db'
NEWPW=$(head -c 300 /dev/urandom | tr -dc A-Za-z0-9 | head -c 24)
OUT=$(docker exec -u abc calibre-web sh -c "cd /app/calibre-web-automated && python3 cps.py -p /config/app.db -s admin:$NEWPW" 2>&1; echo "rc=$?")
echo "command output: $(echo "$OUT" | tr '\n' ' ' | sed "s/$NEWPW/<redacted>/g" | cut -c1-300)"
echo "AFTER: default -> $(login admin123) | new -> $(login "$NEWPW") | control wrong -> $(login wrongxyz123)"
unset NEWPW; rm -f $J
@@ -0,0 +1,5 @@
set -u
AUTH='IO.puts("ANS=#{Claper.Accounts.get_user_by_email_and_password("admin@claper.co", "PW") != nil}")'
auth() { docker exec -e ELIXIR_ERL_OPTIONS=+fnu claper /app/bin/claper rpc "${AUTH//PW/$1}" 2>&1 | grep -o 'ANS=[a-z]*'; }
echo "after the change: the default password claper works: $(auth claper)"
echo "control: a wrong password works: $(auth wrongxyz)"
@@ -0,0 +1,16 @@
set -u
# Elixir code files inside the container, so no shell quoting reaches the code.
docker exec claper sh -c 'cat > /tmp/felhom_auth.exs' <<'EXS'
IO.puts("ANS=#{Claper.Accounts.get_user_by_email_and_password("admin@claper.co", System.get_env("P") || "") != nil}")
EXS
auth() { docker exec -e ELIXIR_ERL_OPTIONS=+fnu claper sh -c "/app/bin/claper rpc \"$(docker exec claper cat /tmp/felhom_auth.exs | sed "s/System.get_env(\"P\") || \"\"/\"$1\"/")\"" 2>&1 | grep -o 'ANS=[a-z]*'; }
echo "BEFORE: default password works: $(auth claper)"
NEWPW=$(head -c 300 /dev/urandom | tr -dc A-Za-z0-9 | head -c 24)
echo "test password: generated in the guest, length ${#NEWPW}, not printed"
CODE='u = Claper.Accounts.get_user_by_email("admin@claper.co"); r = if u, do: Claper.Accounts.update_user_password(u, "claper", %{password: "PW", password_confirmation: "PW"}), else: :no_default_admin; case r do {:ok, _} -> IO.puts("FELHOM_AFTER_INSTALL_OK"); other -> IO.puts("FELHOM_AFTER_INSTALL_FAILED #{inspect(other, limit: 3)}") end'
docker exec -e ELIXIR_ERL_OPTIONS=+fnu claper /app/bin/claper rpc "${CODE//PW/$NEWPW}" 2>&1 | grep -o 'FELHOM_AFTER_INSTALL_[A-Z]*.*' | cut -c1-200
echo "AFTER: default password works: $(auth claper)"
echo "AFTER: the new password works: $(auth "$NEWPW")"
echo "control: a wrong password: $(auth wrongxyz)"
docker exec claper rm -f /tmp/felhom_auth.exs
unset NEWPW
@@ -0,0 +1,20 @@
## 2026-09-28T14:25:22Z Part C — demo-hp restore_storage local-lvm -> nvme-scratch (decision 44)
-rw------- 1 felhom-agent felhom-agent 2451 Sep 24 21:57 /etc/felhom-agent/agent
restore_storage: local-lvm -> nvme-scratch
19c19
< "restore_storage": "local-lvm",
---
> "restore_storage": "nvme-scratch",
600 felhom-agent
active
Sep 28 16:25:22 demo-hp felhom-agent[1370926]: time=2026-09-28T16:25:22.505+02:00 level=INFO msg="backup: restore-test scheduler shutting down" reason="context canceled"
Sep 28 16:25:22 demo-hp felhom-agent[1370926]: time=2026-09-28T16:25:22.505+02:00 level=INFO msg="storage: watchdog shutting down" reason="context canceled"
Sep 28 16:25:22 demo-hp systemd[1]: Stopping felhom-agent.service - Felhom host agent (Proxmox host tier; hub control loop + PBS verify + storage watchdog)...
Sep 28 16:25:22 demo-hp systemd[1]: Stopped felhom-agent.service - Felhom host agent (Proxmox host tier; hub control loop + PBS verify + storage watchdog).
Sep 28 16:25:22 demo-hp systemd[1]: Started felhom-agent.service - Felhom host agent (Proxmox host tier; hub control loop + PBS verify + storage watchdog).
Sep 28 16:25:22 demo-hp felhom-agent[2355282]: time=2026-09-28T16:25:22.547+02:00 level=INFO msg="felhom-agent daemon starting" version=0.137.0 host_id=demo-hp-bb76ea hub_url=https://hub.felhom.eu interval_s=900
Sep 28 16:25:23 demo-hp felhom-agent[2355282]: time=2026-09-28T16:25:23.491+02:00 level=INFO msg="storage: watchdog starting" interval=5s debounce=15s
Sep 28 16:25:23 demo-hp felhom-agent[2355282]: time=2026-09-28T16:25:23.491+02:00 level=INFO msg="pbs: verify loop starting" cadence=6h0m0s
## pool before:
local-lvm lvmthin active 56487936 33017198 23470737 58.45%
nvme-scratch dir active 983379700 50330164 883022924 5.12%
@@ -0,0 +1,27 @@
## 2026-09-28T14:25:36Z one real restore test (selftest=restore-test), as the agent user
=== felhom-agent 0.137.0 selftest=restore-test ===
--- recover: reaping any leaked scratch from a prior crashed test ---
recover: examined=0 scratch_destroyed=0 scratch_clean=0
restoring felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z into scratch band [990000,990009] on nvme-scratch …
time=2026-09-28T16:25:37.513+02:00 level=INFO msg="restore-test: space preflight passed" storage=nvme-scratch required_bytes=33252620982 avail_bytes=904215474176
time=2026-09-28T16:25:37.634+02:00 level=INFO msg="restore-test: full-fidelity restore params derived from the archive config" scratch=990000 params=4
--- restore-test record ---
{
"source_archive": "felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z",
"source_tier": "pbs",
"scratch_vmid": 990000,
"pass": false,
"verified": "",
"error": "reconcile: restore-test restore: proxmox: POST /nodes/demo-hp/lxc -\u003e HTTP 403: permission denied at /storage/nvme-scratch (missing privilege Datastore.AllocateSpace)",
"tested_at": "2026-09-28T14:25:37Z",
"duration_seconds": 1.239386731
}
space preflight passed: storage=nvme-scratch required=33252620982 avail=904215474176
[FAIL] restore-test (scratch 990000): reconcile: restore-test restore: proxmox: POST /nodes/demo-hp/lxc -> HTTP 403: permission denied at /storage/nvme-scratch (missing privilege Datastore.AllocateSpace)
rc=1
## pool samples (time storage used% avail):
16:25:36 local-lvm 23470737 58.45%
16:25:36 nvme-scratch 883022924 5.12%
VMID Status Lock Name
9201 running demo-hp
9202 running demo-hp-scratch
@@ -0,0 +1,18 @@
## 2026-09-28T15:58:32Z operator approved: grant FelhomAgentStore on /storage/nvme-scratch (user + token), demo-hp only
BEFORE:
| /storage/felhom-pbs | FelhomAgentStore | user | felhom-agent@pve | 1 |
| /storage/felhom-pbs | FelhomAgentStore | token | felhom-agent@pve!agent | 1 |
| /storage/local | FelhomAgentStore | user | felhom-agent@pve | 1 |
| /storage/local | FelhomAgentStore | token | felhom-agent@pve!agent | 1 |
| /storage/local-lvm | FelhomAgentStore | user | felhom-agent@pve | 1 |
| /storage/local-lvm | FelhomAgentStore | token | felhom-agent@pve!agent | 1 |
COMMAND (target first): pveum acl modify /storage/nvme-scratch --roles FelhomAgentStore (user felhom-agent@pve, token felhom-agent@pve!agent)
AFTER:
| /storage/felhom-pbs | FelhomAgentStore | user | felhom-agent@pve | 1 |
| /storage/felhom-pbs | FelhomAgentStore | token | felhom-agent@pve!agent | 1 |
| /storage/local | FelhomAgentStore | user | felhom-agent@pve | 1 |
| /storage/local | FelhomAgentStore | token | felhom-agent@pve!agent | 1 |
| /storage/local-lvm | FelhomAgentStore | user | felhom-agent@pve | 1 |
| /storage/local-lvm | FelhomAgentStore | token | felhom-agent@pve!agent | 1 |
| /storage/nvme-scratch | FelhomAgentStore | user | felhom-agent@pve | 1 |
| /storage/nvme-scratch | FelhomAgentStore | token | felhom-agent@pve!agent | 1 |
@@ -0,0 +1,99 @@
## 2026-09-28T15:58:42Z one real restore test on nvme-scratch (selftest=restore-test), as the agent user
=== felhom-agent 0.137.0 selftest=restore-test ===
--- recover: reaping any leaked scratch from a prior crashed test ---
recover: examined=0 scratch_destroyed=0 scratch_clean=0
restoring felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z into scratch band [990000,990009] on nvme-scratch …
time=2026-09-28T17:58:43.148+02:00 level=INFO msg="restore-test: space preflight passed" storage=nvme-scratch required_bytes=33252620981 avail_bytes=904215433216
time=2026-09-28T17:58:43.254+02:00 level=INFO msg="restore-test: full-fidelity restore params derived from the archive config" scratch=990000 params=4
time=2026-09-28T18:07:22.310+02:00 level=INFO msg="audit: gate decision" class=guest_destroy host=demo-hp-bb76ea guest=990000 source=one_shot_job disposition=benign allowed=true reason=benign key_id="" nonce="" durable_id=""
time=2026-09-28T18:07:22.310+02:00 level=INFO msg="gate decision" class=guest_destroy guest=990000 source=one_shot_job disposition=benign allowed=true reason=benign
time=2026-09-28T18:07:28.488+02:00 level=INFO msg="restore-test: scratch guest torn down" vmid=990000
--- restore-test record ---
{
"source_archive": "felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z",
"source_tier": "pbs",
"scratch_vmid": 990000,
"pass": true,
"verified": "boot+running",
"tested_at": "2026-09-28T16:07:28Z",
"duration_seconds": 526.400585242,
"mount_parity": "ok",
"mount_inventory": [
"mp0=/var/lib/felhom (70G)",
"mp8=/mnt/felhom-drives (throwaway for the archived bind)",
"mp9=/etc/felhom-bootstrap (throwaway for the archived bind)"
]
}
space preflight passed: storage=nvme-scratch required=33252620981 avail=904215433216
=== selftest=restore-test OK (scratch 990000 restored+booted+verified+torn-down in 8m46s) ===
rc=0
## pool samples (time storage used_KiB used%):
17:58:42 local-lvm 33022847 58.46%
17:58:42 nvme-scratch 50330204 5.12%
17:58:58 local-lvm 33022847 58.46%
17:58:58 nvme-scratch 51384200 5.23%
17:59:14 local-lvm 33022847 58.46%
17:59:14 nvme-scratch 52273728 5.32%
17:59:30 local-lvm 33022847 58.46%
17:59:30 nvme-scratch 53087720 5.40%
17:59:46 local-lvm 33022847 58.46%
17:59:46 nvme-scratch 53732108 5.46%
18:00:02 local-lvm 33022847 58.46%
18:00:02 nvme-scratch 54567904 5.55%
18:00:18 local-lvm 33022847 58.46%
18:00:18 nvme-scratch 54665524 5.56%
18:00:34 local-lvm 33022847 58.46%
18:00:34 nvme-scratch 55543996 5.65%
18:00:51 local-lvm 33022847 58.46%
18:00:51 nvme-scratch 56598724 5.76%
18:01:07 local-lvm 33022847 58.46%
18:01:07 nvme-scratch 57386612 5.84%
18:01:23 local-lvm 33022847 58.46%
18:01:23 nvme-scratch 58709180 5.97%
18:01:39 local-lvm 33022847 58.46%
18:01:39 nvme-scratch 58709180 5.97%
18:01:55 local-lvm 33022847 58.46%
18:01:55 nvme-scratch 59700476 6.07%
18:02:11 local-lvm 33022847 58.46%
18:02:11 nvme-scratch 60306788 6.13%
18:02:27 local-lvm 33022847 58.46%
18:02:27 nvme-scratch 61193424 6.22%
18:02:44 local-lvm 33022847 58.46%
18:02:44 nvme-scratch 61865512 6.29%
18:03:00 local-lvm 33022847 58.46%
18:03:00 nvme-scratch 63145012 6.42%
18:03:16 local-lvm 33022847 58.46%
18:03:16 nvme-scratch 63145012 6.42%
18:03:32 local-lvm 33022847 58.46%
18:03:32 nvme-scratch 64424332 6.55%
18:03:48 local-lvm 33022847 58.46%
18:03:48 nvme-scratch 64457976 6.55%
18:04:04 local-lvm 33022847 58.46%
18:04:04 nvme-scratch 65295280 6.64%
18:04:20 local-lvm 33022847 58.46%
18:04:20 nvme-scratch 66332308 6.75%
18:04:36 local-lvm 33022847 58.46%
18:04:36 nvme-scratch 67011780 6.81%
18:04:53 local-lvm 33022847 58.46%
18:04:53 nvme-scratch 67144228 6.83%
18:05:09 local-lvm 33022847 58.46%
18:05:09 nvme-scratch 67835968 6.90%
18:05:25 local-lvm 33022847 58.46%
18:05:25 nvme-scratch 68712568 6.99%
18:05:41 local-lvm 33022847 58.46%
18:05:41 nvme-scratch 69680724 7.09%
18:05:57 local-lvm 33022847 58.46%
18:05:57 nvme-scratch 69768120 7.09%
18:06:13 local-lvm 33022847 58.46%
18:06:13 nvme-scratch 70670236 7.19%
18:06:29 local-lvm 33022847 58.46%
18:06:29 nvme-scratch 71616032 7.28%
18:06:45 local-lvm 33022847 58.46%
18:06:45 nvme-scratch 72544896 7.38%
18:07:01 local-lvm 33022847 58.46%
18:07:01 nvme-scratch 72623192 7.39%
18:07:18 local-lvm 33022847 58.46%
18:07:18 nvme-scratch 74121424 7.54%
VMID Status Lock Name
9201 running demo-hp
9202 running demo-hp-scratch
@@ -0,0 +1 @@
2026/09/28 09:40:18 [INFO] wg registered: host=drill-g0276-b861fc pubkey=Ly0yjK/BwcPQqDdwIhcd2W96UwdPhQsIQ8fY5GcevVQ= ip=10.77.0.5/32 changed=true gen=1 sync=ok
@@ -0,0 +1,22 @@
## 2026-09-28T16:08:13Z ep0 read-only
felhom-hetzner
wg0 KNaFXHiY9l08UhPO4gxUBFQz2WF3RYWyUGL8k+hzDVo= 10.77.0.2/32
wg0 yNVpWt9Ekid9p459/VLHbmCDHIi8xTUYHbgazUp5S00= 10.77.0.250/32
wg0 kdhOHMyADRpaTaFmh8ypwsWngvBpIyMyon80DJK+00M= 10.77.0.4/32
wg0 snkgWxlcN7zXxjy1rG/jefe6uysGq/27E736CVPkbBc= 10.77.0.3/32
wg0 KNaFXHiY9l08UhPO4gxUBFQz2WF3RYWyUGL8k+hzDVo= 1790611636
wg0 yNVpWt9Ekid9p459/VLHbmCDHIi8xTUYHbgazUp5S00= 1790611612
wg0 kdhOHMyADRpaTaFmh8ypwsWngvBpIyMyon80DJK+00M= 1786601749
wg0 snkgWxlcN7zXxjy1rG/jefe6uysGq/27E736CVPkbBc= 1790611684
## saved WireGuard config: count of the test key / of drill-g0276
0
0
## control: a key that IS configured is found:
1
## PBS datastores + namespaces:
felhom-offsite /mnt/pbs-datastore
/mnt/pbs-datastore/ns/demo-felhom
/mnt/pbs-datastore/ns/demo-hp
/mnt/pbs-datastore/ns/tester-1
## grep drill-g0276 across /etc /root /srv (names only):
(end)
@@ -0,0 +1,25 @@
<string>:9: SyntaxWarning: invalid escape sequence '\]'
<string>:9: SyntaxWarning: invalid escape sequence '\]'
<string>:10: SyntaxWarning: invalid escape sequence '\]'
POST /api/debug/backup/night-chain -> ('202', {'data': {'legs': ['db-dump', 'tier2', 'update-leg']}, 'message': 'started', 'ok': True})
second press while it runs -> ('409', {'error': 'refused: a backup or restore is running', 'ok': False})
2026/09/28 16:44:22 night_chain.go:108: [INFO] [night-chain] manual run started from 172.18.0.6:43528: [db-dump tier2 update-leg]
2026/09/28 16:44:22 night_chain.go:84: [INFO] [night-chain] manual run of tonight's chain: dump → second drive → off-site → update leg
2026/09/28 16:44:22 night_chain.go:77: [INFO] [night-chain] db-dump: started
2026/09/28 16:44:22 night_chain.go:104: [WARN] [night-chain] manual run REFUSED: a backup or restore is running
2026/09/28 16:45:05 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for claper → /mnt/sys_drive/felhom-data/backups/primary/claper (images=2, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
2026/09/28 16:45:05 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for paperless-ngx → /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx (images=3, secrets-referenced=3, data_keys=
2026/09/28 16:45:05 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for privatebin → /mnt/sys_drive/felhom-data/backups/primary/privatebin (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld
2026/09/28 16:45:05 night_chain.go:82: [INFO] [night-chain] db-dump: done in 43s
2026/09/28 16:45:05 night_chain.go:77: [INFO] [night-chain] tier2: started
2026/09/28 16:45:05 tier2.go:425: [INFO] [backup] Tier 2 copied claper → /mnt/felhom-drives/scratch_hdd/backups/secondary/claper (48.7 MB, 0 leg(s), 0s)
2026/09/28 16:45:06 tier2.go:425: [INFO] [backup] Tier 2 copied paperless-ngx → /mnt/sys_drive/felhom-data/backups/secondary/paperless-ngx (76.1 MB, 1 leg(s), 0s) [SSD: state-only]
2026/09/28 16:45:06 tier2.go:425: [INFO] [backup] Tier 2 copied privatebin → /mnt/felhom-drives/scratch_hdd/backups/secondary/privatebin (20.3 KB, 0 leg(s), 0s)
2026/09/28 16:45:06 night_chain.go:82: [INFO] [night-chain] tier2: done in 0s
2026/09/28 16:45:06 night_chain.go:90: [INFO] [night-chain] offsite: no off-site target on this box — skipped, as at night
2026/09/28 16:45:06 night_chain.go:77: [INFO] [night-chain] update-leg: started
2026/09/28 16:45:06 unattended.go:268: [INFO] [update-leg] started (manual-chain): window 02:30, no step starts at or after 22:00
2026/09/28 16:45:06 unattended.go:246: [INFO] [update-leg] update leg (manual-chain): done=0 undone=0 held=0 failed=0 skipped=0 in 0s [skipped: ]
2026/09/28 16:45:06 night_chain.go:82: [INFO] [night-chain] update-leg: done in 0s
2026/09/28 16:45:06 night_chain.go:93: [INFO] [night-chain] finished in 44s
@@ -0,0 +1,13 @@
# logins-nvme — 2026-09-28 evening: default logins (decision 45), demo-hp restore test on the NVMe (decision 44), Part D, ep0, R-705/R-706
Architecture read: `09` §3 decisions 44–45 (new), `01` §5 (the new trust line), `03` (restore storage), `07` §6
(Part D), `06` + R-600 (ep0). Controller v0.279.0. Method: endpoint level (the calls the pages make) and each app's
own CLI/login form; generated passwords are read from the app page and never printed.
| folder | what |
|---|---|
| `B/` | claper + bookstack: probe of the app's own CLI route, then the fresh-install proof through `after_install` (default fails, generated works, wrong fails); claper after a restore; calibre-web's route refused an alphanumeric password (its policy); demo-hp's INSTALLED apps: default logins tried read-only (bookstack + calibre-web still log in; romm's note is stale) and the page warnings; floor 0.279.0 |
| `C/` | demo-hp: `restore_storage` → nvme-scratch; the 403 without a grant; the grant (operator's word); one restore test passed in 8m46s with pool samples |
| `E/` | ep0 read-only: the test install's peer is gone (hub record vs `wg show` vs saved configs, with a positive control) |
| `F/` | the night chain on 9202 (R-705): order, the refusal, the day deadline; Part D raised no false alarm |
| `redproofs/` | RP7–RP15 for v0.279.0, each failing on an assertion |
@@ -0,0 +1,10 @@
# RP10 — the manual leg keeps the night deadline
263c263
< if manual {
---
> if false && manual { // RED-PROOF RP10
unattended_test.go:490: the manual leg's deadline 2026-09-25 07:30:00 +0000 UTC, want 2026-09-25 21:15:00 +0000 UTC
--- FAIL: TestR705_TheManualLegRunsByDay (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.008s
FAIL
@@ -0,0 +1,10 @@
# RP11 — the data check does nothing (pre-0.279.0: a log line only)
87c87
< hollow := m.watchesForData(app) && !unitCarriesData(unitDir)
---
> hollow := false && m.watchesForData(app) && !unitCarriesData(unitDir) // RED-PROOF RP11
r_partd_hollow_watch_test.go:63: operator told [], want exactly [cloud/offsite] (held+stopped and volume-less apps are not flagged)
--- FAIL: TestPartD_ARunningAppWithAnEmptyCopyIsAnAlarm (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.008s
FAIL
@@ -0,0 +1,10 @@
# RP12 — the page never gets the sentences
1134c1134
< data["HollowCopyLines"] = s.hollowCopyLines(s.langFor(r)) // Part D: running apps whose copy holds no data
---
> // RED-PROOF RP12
r_partd_hollow_page_test.go:27: the flagged app's sentence is not on the page (code 200)
--- FAIL: TestPartD_TheBackupPageSaysARunningAppHasNoData (0.06s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.068s
FAIL
@@ -0,0 +1,10 @@
# RP13 — no success-marker check (the exit code alone decides)
139c139
< if err == nil && strings.Contains(out, ai.Success) {
---
> if err == nil && strings.Contains(out, "") { // RED-PROOF RP13: the exit code alone decides
after_install_test.go:76: an exit-0 command without the success marker reported success
--- FAIL: TestAfterInstall_ReplacesTheDefaultLoginOnceAndRecordsIt (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.011s
FAIL
@@ -0,0 +1,10 @@
# RP14 — an undeclared/empty value is filled in
67c67
< if !ok[k] || env[k] == "" {
---
> if false { // RED-PROOF RP14
after_install_test.go:72: expanded with env map[]
--- FAIL: TestAfterInstall_ReplacesTheDefaultLoginOnceAndRecordsIt (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.011s
FAIL
@@ -0,0 +1,12 @@
# RP15 — the default card ignores after_install
23c23
< if meta.AfterInstall == nil {
---
> if true { // RED-PROOF RP15
known_login_test.go:34: after_install declared: before install: in effect = true, want false
known_login_test.go:34: after_install succeeded: in effect = true, want false
known_login_test.go:34: after_install not run yet: in effect = true, want false
--- FAIL: TestKnownLogin_InEffectOnlyUntilReplaced (0.01s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.016s
FAIL
@@ -0,0 +1,10 @@
# RP7 — removeStack does not delete the verification copy (0.278.0)
1035c1035
< resp.BackupPathsRemoved = append(resp.BackupPathsRemoved, r.removeVerificationCopy(name)...)
---
> // RED-PROOF RP7
r706_verify_copy_test.go:74: removeStack does not call removeVerificationCopy (R-706)
--- FAIL: TestR706_RemovalWithBackupsDeletesTheVerificationCopy (0.01s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.015s
FAIL
@@ -0,0 +1,12 @@
# RP8 — two legs swapped (off-site before Tier 2)
86d85
< step("tier2", func() error { c.tier2(); return nil })
88a88,90
> }
> step("tier2", func() error { c.tier2(); return nil })
> if false {
r705_night_chain_test.go:47: order [db-dump offsite tier2 update-leg], want [db-dump tier2 offsite update-leg]
--- FAIL: TestR705_NightChainRunsTheLegsInOrderAndRefusesWhenBusy (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.008s
FAIL
@@ -0,0 +1,11 @@
# RP9 — no busy() refusal
57c57
< if busy, why := c.busy(); busy {
---
> if busy, why := c.busy(); false && busy { // RED-PROOF RP9
r705_night_chain_test.go:54: busy start: ok=true why=""
--- FAIL: TestR705_NightChainRunsTheLegsInOrderAndRefusesWhenBusy (0.00s)
FAIL
/mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/web/r705_night_chain_test.go:25 +0x36
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.012s
FAIL
@@ -0,0 +1,19 @@
"""claperproof.py — decision 45 proof on a FRESH claper: the known default fails, the generated one (read
from the app page the household sees) works, through claper's own authentication. The value is never printed."""
import re, sys, html as H
import walk as w
w.login()
page = w.page("/stacks/claper/deploy")
m = re.search(r'name="ADMIN_PASSWORD"[^>]*value="([^"]*)"', page) or re.search(r'value="([^"]*)"[^>]*name="ADMIN_PASSWORD"', page)
gen = H.unescape(m.group(1)) if m else ""
print("app page shows a first password for ADMIN_PASSWORD:", bool(gen), "length", len(gen))
info = w.page("/apps/claper")
print("app page: default-login card shown:", "admin@claper.co / claper" in info, "| known-login warning shown:", "ismert, közös jelszóval" in info)
auth = 'IO.puts("ANS=#{Claper.Accounts.get_user_by_email_and_password("admin@claper.co", "PW") != nil}")'
script = "set -u\n" + "\n".join(
f"echo \"{label}: $(docker exec -e ELIXIR_ERL_OPTIONS=+fnu claper /app/bin/claper rpc '{auth.replace('PW', pw)}' 2>&1 | grep -o 'ANS=[a-z]*')\""
for label, pw in (("NEGATIVE the known default admin@claper.co / claper", "claper"),
("POSITIVE the generated first password from the app page", gen),
("CONTROL a wrong password", "wrongxyz123")))
print(w.guest(script))
print(w.guest("grep -A3 '^after_install:' /opt/docker/stacks/claper/app.yaml"))
@@ -0,0 +1,22 @@
"""After a RESTORE of claper from its own unit: the default must still fail, the first password (read from
the page BEFORE the restore) must still work, and the page must not show a value that does not work."""
import re, html as H, os
import walk as w
w.login()
def pw_on_page():
m = re.search(r'name="ADMIN_PASSWORD"[^>]*value="([^"]*)"', w.page("/stacks/claper/deploy"))
return H.unescape(m.group(1)) if m else ""
before = pw_on_page()
print("before the restore: the page shows a first password:", bool(before))
res = w.restore("claper")
print("restore:", {k: res.get(k) for k in ("ok", "http", "seconds", "state_after")})
after = pw_on_page()
print("after the restore: the page shows a password value:", bool(after), "| same as before:", after == before and bool(after))
d = w.page("/stacks/claper/deploy")
print("after the restore: the page says the backup's login is the one that works:", "restored" in d.lower() or "mentésből" in d)
auth = 'IO.puts("ANS=#{Claper.Accounts.get_user_by_email_and_password("admin@claper.co", "PW") != nil}")'
script = "set -u\n" + "\n".join(
f"echo \"{l}: $(docker exec -e ELIXIR_ERL_OPTIONS=+fnu claper /app/bin/claper rpc '{auth.replace('PW', p)}' 2>&1 | grep -o 'ANS=[a-z]*')\""
for l, p in (("NEGATIVE the known default", "claper"), ("the first password from BEFORE the restore", before), ("CONTROL wrong", "wrongxyz123")))
print(w.guest(script))
print(w.guest("grep -E -A2 '^(after_install|restored_logins):' /opt/docker/stacks/claper/app.yaml"))
@@ -0,0 +1,18 @@
"""loginproof.py <app> <sub> <email> — decision 45 proof on a FRESH install of a BookStack-shaped app (a Laravel
login form): the known default fails, the generated first password read from the app page works, a wrong one fails.
The value is never printed."""
import re, sys, html as H
import walk as w
app, sub, email, default = sys.argv[1:5]
w.login()
m = re.search(r'name="ADMIN_PASSWORD"[^>]*value="([^"]*)"', w.page(f"/stacks/{app}/deploy"))
gen = H.unescape(m.group(1)) if m else ""
print("app page shows a first password:", bool(gen))
info = w.page(f"/apps/{app}")
print("app page loaded:", len(info) > 20000, "| default-login card shown:", "Alapértelmezett belépés" in info, "| warning shown:", "known-login-line" in info)
fn = f'''J=/tmp/lp.$$; login() {{ rm -f $J; T=$(curl -sk -c $J -b $J -H "Host: {sub}.{w.DOMAIN}" https://127.0.0.1/login | grep -o 'name="_token" value="[^"]*"' | head -1 | sed 's/.*value="//;s/"$//'); curl -sk -o /dev/null -w '%{{http_code}}->%{{redirect_url}}' -c $J -b $J -H "Host: {sub}.{w.DOMAIN}" -X POST --data-urlencode "_token=$T" --data-urlencode "email={email}" --data-urlencode "password=$1" https://127.0.0.1/login | sed "s|https\\?://[^/]*||"; }}'''
print(w.guest(fn + f'''
echo "NEGATIVE the known default {email} / {default}: $(login '{default}')"
echo "POSITIVE the generated first password: $(login '{gen}')"
echo "CONTROL wrong: $(login wrongxyz123)"; rm -f $J'''))
print(w.guest(f"grep -A3 '^after_install:' /opt/docker/stacks/{app}/app.yaml"))
+8 -5
View File
@@ -757,7 +757,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-597** | **[P2-MEDIUM] The setup code is three Hungarian words, inside an otherwise fully English e-mail, sent to a household the hub knows is English.** FOUND 2026-09-20 by the slice-6 drill. The mail is English end to end (slice 3 working); the code it carries was **`képző-szkítia-ásatás`** — 20 characters, 3 words, **5 of them outside ASCII** (ő, í, á×2, é). An English speaker must copy three words they cannot read, spell or say aloud, and type them into a box on a keyboard that has no ő. They can paste — until the day they read the code to someone over the telephone, which is precisely what a three-word code is FOR. **The same generator feeds the recovery code (10 words) and the owner passphrase (5 words)**, so the fault is one wordlist wide, not one mail wide: this walk saw the passphrase too and it is Hungarian. **Fix shape:** an English wordlist chosen per `customer.language`, with the same word count and the same entropy, and a test that pins BOTH lists' entropy and that no word in either needs a character outside the reader's keyboard. **Not a rename of the existing words** — a second list. **CLOSED 2026-09-21, hub v0.119.0.** **One third of the row was wrong: the recovery code was never Hungarian.** `felhom-agent` mints it (`internal/escrow`) from the **EFF large wordlist** and always has — ten English words, ≈129 bits. The hub does not own that secret and no row was opened for it: a second definition here is the drift `backupTargetAbsentText` already demonstrates across two repos. The two the hub DOES mint now follow the household: setup code 3 hu words (44.6 bits) → **4 en words (51.7)**, owner passphrase 5 hu (74.3) → **6 en (77.5)**, list and count chosen together by `RandomPassphraseFor(lang, use)` so a caller cannot pair an English list with a Hungarian count. **The floor is computed from the embedded lists at test time, not compared with a constant** — red-proofed at 3 English words (38.77 vs 44.56). Hungarian is byte-unchanged, and the list length is pinned so a swap cannot move it quietly. **The task's proposed "read it over the phone" filter was MEASURED and NOT adopted** — it removes 5270 of 7772 words (68%, 12.92 → 11.29 bits/word) and would make this list stricter than the one the product already uses for the code a household writes on paper during a disaster; what it reached for is kept as an assertion (`TestEnglishListIsTranscribable`: 3-9 lower-case ASCII letters, no digit, no separator). Decision recorded in source, **operator may reverse**. Also: **no claim mail ever stated a word count** — the only count wording was the bind page's passphrase hint, whose English half is now count-free. | **CLOSED 2026-09-21 — hub v0.119.0** |
| **R-598** | **[P2-MEDIUM] The Backup page's two protection warnings — the ones that say whether the household's files are safe — are Hungarian on an English dashboard.** FOUND 2026-09-20 by the slice-6 drill on a fresh box, and confirmed on the demo box. Of 73 lines on `/backups` exactly four are Hungarian: **„Csak egy másolat készül (nincs második meghajtó) — a 3-2-1 mentéshez csatlakoztasson egy második meghajtót vagy offsite tárolót"**, **„A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok ellen véd, lemezhiba ellen nem"**, and the two backup-target names **„Helyi tároló (local)"** and **„Biztonsági szerver – külön hardver (PBS)"**. They come from `internal/web/backup_handlers.go` (12 Hungarian literals) and `internal/web/backup_target_offer.go` (10) — again composed sentences handed to the page, the R-573/R-596 shape. **It matters more than its line count:** those two warnings are the only place the product tells a household that one copy on one disk is not protection, and the volunteer guide's §9 sends every tester to exactly this page to read exactly these two sentences. **Fix shape:** keys + args for both files, with the retrieval-promise gate run over the English (these sentences are about what a backup does and does not protect). **CLOSED 2026-09-21, controller v0.259.0.** The row's count of `backup_handlers.go` was 12; **nine are code and three are Hungarian inside COMMENTS**. The offer file's ten is right. `degradedMessageFor` now returns a **KEY** — the decision stays language-free and in one place, the words are chosen by the caller that knows the reader — and `buildTierViews` / `backupTargetLabel` / `loadGuestBackup` take the language the way `buildDataPathCards` already did. **The English is asserted to carry the same NEGATION the Hungarian does** ("protects against corrupted files, **but not** against a disk failure"); an English sentence that promised disk-failure protection would be worse than leaving it Hungarian. **Proven LIVE on guest 9201** for the two tier names ("Local storage (felhom-backup)", "Backup server – separate hardware (PBS)"); **the two warnings themselves were NOT walked live** — that box is healthy and a healthy box renders nothing by design, and producing the state would mean un-assigning a live backup target. They are covered by render tests through the real handler in both states. **An apostrophe cost a render:** the first English absent-drive sentence never matched because `html/template` escapes `'` to `&#39;` — caught by the test, not by review. | **CLOSED 2026-09-21 — controller v0.259.0; the two warnings proven by render test, not live** |
| **R-599** | **[P3-LOW] A drill's teardown is blocked for 30 minutes by design, and nothing says so.** FOUND 2026-09-20 tearing the slice-6 drill down. VM destroyed at 17:12Z; the hub then refused **both** `POST /configs/<id>/delete` (409, *host … is ONLINE*) and the host delete (`deletable:false`) — correctly, because an online host would receive permanent 401s. But "online" is not a liveness probe: it is a **report-staleness window**, and the window is **45 minutes** — `manifests/hub.yaml` sets `alerting.stale_threshold: "45m"`, which `hostStatus()` reads (`ok` under it, `stale` over, `down` at 2x). A machine that no longer exists therefore reads ONLINE for three quarters of an hour. **Measured the boring way, and worth recording:** this row first said 30 minutes, because `monitor/host_staleness.go`'s literal default says 30m — the DEPLOYED value is in the manifest, and the 409s kept coming after the half hour was up. Reading a default and calling it the live value is the same mistake in a smaller coat. **The consequence is not theoretical:** a session that destroys its VM and then tears down the hub side walks away believing the delete failed, or leaves the customer behind — and the 2026-09-14 drill's teardown had the same shape without recording this. **Fix shape (smallest first):** the 409 body says *how long* it will refuse ("the last report was N minutes ago; deletion opens at HH:MM"), and `runbooks/target-selection.md`'s drill section names the wait. A force flag is NOT proposed — the refusal is right, only silent about its own clock. | **READY — rank P3-LOW; owner: CC (hub)** |
| **R-600** | **[P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0.** FOUND 2026-09-20 by the slice-6 drill's teardown, **measured on ep0 rather than inferred from the hub**. The customer delete cascade finished at 19:41:54 with `customer DELETE cascade COMPLETE for drill-en-0920 (journal #18) — full teardown`, and every hub-side row was gone (0 configs, 0 hosts, 17 residue rows purged, PBS tenancy deprovisioned, escrow demoted). **Three minutes later `wg show wg0 allowed-ips` on ep0 still listed `10.77.0.5/32`** — the drill box's peer — because `wgsync` pushes on its own cycle. **Watched to the end rather than assumed: the peer was gone by 17:47:46Z — it outlived the *full teardown* line by about 6 minutes.** (My first estimate said ~35, read off the gap between two log lines; the sync runs oftener and only LOGS when something changes. That is the second time in this session that a period inferred from two log lines was wrong — the other was the delete's own staleness window. **A period read off two log lines is not a measurement.**) **The 2026-09-14 drill's findings say "the teardown removes it through the host delete"; measured, the host delete removes the hub's RECORD and the peer goes on the next push.** The mechanism is not broken — it is asynchronous, and the log line claims a completeness it does not yet have. **Why it is P2 rather than P3:** a session that tears down, reads *full teardown*, and leaves is the normal case; the peer outlives it by minutes, and the third teardown layer is the one the workspace rules single out as the one that gets forgotten. Six minutes is short — but the session that reads *full teardown* and leaves has no way to know it is six and not six hours. **Fix shape (smallest first):** the cascade triggers a wgsync push before it logs COMPLETE, or the log line says what is still pending and when ("wg peer removal queued; next push in N min"). A session should not have to read ep0 to know whether a teardown finished. **2026-09-25 (Peti's retirement):** nothing to remove for `peti-felhom-86d37d` — its host was deleted 2026-07-15 and ep0's live `wg show` carries no peer beyond the demo boxes', drill-r50 and the operator OOB (`audits/retire-peti-2026-09-25/A1-ep0-before.txt`); whether that July delete removed a peer, or none existed, is not recorded. | **READY - rank P2-MEDIUM; owner: CC (hub)** |
| **R-600** | **[P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0.** FOUND 2026-09-20 by the slice-6 drill's teardown, **measured on ep0 rather than inferred from the hub**. The customer delete cascade finished at 19:41:54 with `customer DELETE cascade COMPLETE for drill-en-0920 (journal #18) — full teardown`, and every hub-side row was gone (0 configs, 0 hosts, 17 residue rows purged, PBS tenancy deprovisioned, escrow demoted). **Three minutes later `wg show wg0 allowed-ips` on ep0 still listed `10.77.0.5/32`** — the drill box's peer — because `wgsync` pushes on its own cycle. **Watched to the end rather than assumed: the peer was gone by 17:47:46Z — it outlived the *full teardown* line by about 6 minutes.** (My first estimate said ~35, read off the gap between two log lines; the sync runs oftener and only LOGS when something changes. That is the second time in this session that a period inferred from two log lines was wrong — the other was the delete's own staleness window. **A period read off two log lines is not a measurement.**) **The 2026-09-14 drill's findings say "the teardown removes it through the host delete"; measured, the host delete removes the hub's RECORD and the peer goes on the next push.** The mechanism is not broken — it is asynchronous, and the log line claims a completeness it does not yet have. **Why it is P2 rather than P3:** a session that tears down, reads *full teardown*, and leaves is the normal case; the peer outlives it by minutes, and the third teardown layer is the one the workspace rules single out as the one that gets forgotten. Six minutes is short — but the session that reads *full teardown* and leaves has no way to know it is six and not six hours. **Fix shape (smallest first):** the cascade triggers a wgsync push before it logs COMPLETE, or the log line says what is still pending and when ("wg peer removal queued; next push in N min"). A session should not have to read ep0 to know whether a teardown finished. **2026-09-25 (Peti's retirement):** nothing to remove for `peti-felhom-86d37d` — its host was deleted 2026-07-15 and ep0's live `wg show` carries no peer beyond the demo boxes', drill-r50 and the operator OOB (`audits/retire-peti-2026-09-25/A1-ep0-before.txt`); whether that July delete removed a peer, or none existed, is not recorded. **-- 2026-09-28: the Day-0 test install's peer (`drill-g0276`, key `Ly0yjK…`, 10.77.0.5) was checked on ep0 the day after its customer DELETE: gone from `wg show`, absent from `/etc/wireguard/*.conf` (control: a live key found), no namespace or file names the customer — the hub's sync removed it; nothing was removed by hand.** `audits/logins-nvme-2026-09-28/E/`. | **READY - rank P2-MEDIUM; owner: CC (hub)** |
| **R-601** | **[P2-MEDIUM] ~~demo-hp is unreachable~~ — WRONG, WITHDRAWN THE SAME DAY. The box was never down; MY ROUTES WERE.** Filed 2026-09-21 morning after `ssh demo-hp`, the hub-vaulted break-glass over the tailnet, `demo-hp-lan`, a ping and `ip neigh` on felhom-pve all failed, and `tailscale status` said *`demo-hp … offline, last seen 30d ago`*. **The operator looked at the hub and said it was ONLINE. It was**: it had reported 13 minutes earlier, and it has been up **4 weeks 2 days**. **Two stale facts, each enough on its own:** (1) `~/.ssh/config` sends `demo-hp` to the tailnet address `100.76.96.79`, and **tailscale is not installed on that box at all** (checked on it: no `tailscaled`, no `tailscale` binary) — so that entry is a dead peer from an earlier build and can never answer; (2) `demo-hp-lan` and `nodes.md` both say `192.168.0.87`, and the box is **statically** on **`192.168.0.104/24`**, bridge-port `nic0` (nodes.md says `enp2s0f0`). **The hub knew the right address the whole time** — every host report carries `addresses: [{iface: vmbr0, cidr: 192.168.0.104/24}, …]`. **What I actually did wrong, and it is the part worth keeping:** I ran `ip neigh` on felhom-pve, and `192.168.0.104 … STALE` was *in that output*, four lines above the `192.168.0.87 … FAILED` I quoted. I searched the output for the address I expected instead of reading it for the address that was there. **The standing rule says a "no access" claim must list what was tried; it does not say the list makes the claim true.** Six failed routes to a stale address are six failures of one assumption, not six pieces of evidence. **FIXED:** both `~/.ssh/config` entries repointed to `.104` (each carrying a comment saying why, including that there is no tailscale on this box), both verified live; `nodes.md` corrected. | **CLOSED 2026-09-21 — withdrawn, the claim was false; the routes are fixed** |
| **R-604** | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** |
| **R-605** | **[P3-LOW] A catalog gate that REFUSED TO RUN and a gate that ran and could not decide print the same word, so a reader cannot tell which happened.** FOUND 2026-09-21 while answering why the chaos night's update round could not run. On 2026-09-17 `check-image-resolvable` and `check-volume-persistence` both returned INCONCLUSIVE and the drawn `update` action was replaced with `use` (`audits/DRILL-chaos-night-2026-09-17.md:692-695`). **Neither script is defective — they behaved exactly as designed**, and both headers say why: a detector that cannot prove itself must refuse to report rather than guess (`check-image-resolvable.py` cites the 2026-07-21 incident where a Docker Hub throttle read as 24 of 65 pins falsely dead). **What is missing is the DISTINCTION.** `check-volume-persistence.py`'s `self_test` refuses to evaluate ANY app when it cannot build its canary image — a harness-level refusal — while `classify()` returns a per-app UNDETERMINED for an app that wrote nothing; `check-image-resolvable.py` likewise separates a harness-level canary failure (exit 2 at `check()` L180-183) from a per-pin throttle (L121-128). **`catalog_gates.py`'s VERDICT map collapses all of them into one `INCONCLUSIVE` label**, so the operator-facing summary cannot say whether the gate ran at all. **The cost is real and already paid:** no raw stdout of the 2026-09-17 run survives in either evidence directory, so the exact triggering path is INFERRED from the code plus the documented throttle precedent, not observed — a distinct summary line would have recorded it for free. **Fix shape:** have each gate's exit distinguish "the harness refused" from "the result is undetermined" (a third exit code, or a marker line the runner matches), and have `catalog_gates.py` print the two differently. **Ships with a decoy each way (R-421): a run whose canary fails must NOT read as a per-app undetermined, and vice versa.** Small. | **READY — rank P3-LOW; owner: CC (catalog)** |
@@ -812,12 +812,15 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-693** | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** |
| **R-698** | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. | **OPEN — P3; owner: operator (a decision), CC measures** |
| **R-700** | **[P2] A drive move unpinned the app — its next start took the catalog's newest version, past the ladder.** Found 2026-09-27 reading the code for R-697 (not seen on a box): `doFlipRedeploy` (the per-app and whole-drive move) persisted through the restore's fresh `app.yaml` write, which drops `pinned_images`, `desired_state`, `installed_images`, the update records and the kept conversion copies. Unpinned, the catalog syncer copies the catalog's compose verbatim (`sync.renderSource`'s table) and the next `up` runs the newest version — for a PostgreSQL app past its conversion step, i.e. a new engine on an old datadir. Pin adoption repairs it only at a controller restart. **-- 2026-09-27 (controller v0.276.0): FIXED** — `persistDriveFlip` changes `HDD_PATH` and nothing else; red-proofed (`audits/records-carried-2026-09-27/redproofs/RP3`, `RP4`). **STILL OPEN: the live proof** — no Tier-0 guest has two drives (9202 has one); prove a move on a box with a second drive, reading `pinned_images` before and after and the running image after the next sync. | **WATCHING — P2; owner: CC (live proof)** |
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. **-- 2026-09-28 option (a) MEASURED — NOT ENOUGH:** `pct fstrim` 9201 + 9202 from the host (rc 0): `local-lvm` 62.32 % → **50.60 %**, free 20.3 → **26.6 GiB** — still under the 31 GiB the preflight needs. The pool is 53.9 GiB and guest 9201 itself holds ~26 GiB, so no trim can reach 31 GiB. The agent's own verdict after the trim: `tier=felhom-pbs due=true` (archive 2026-09-24T20:06:25Z, not proven). `nvme-scratch` has ~820 GiB free. **Left to the operator: (b) or (c).** `audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt` **-- 2026-09-28 14:13 CEST:** the first scheduled cycle after the trim was refused again ("needs 31.0 GiB free, has 23.9 GiB"). **Answered: each refusal DOES reach the hub** — `[WARN] host demo-hp-bb76ea restore-test FAILED …` at every host-report (every 15 min). | **OPEN — P3; owner: operator ((b) or (c)); (a) measured, not enough** |
| **R-702** | **[P1-HIGH] Every claper install creates an admin `admin@claper.co` with the public password `claper`, and the app is published on the household's domain.** Measured 2026-09-28 on scratch guest 9202 (catalog `claper` template, `ghcr.io/claperco/claper:2.5` = 2.5.1): the image's own start command runs `Claper.Release.seeds`, which logs `Created default admin user: Email: admin@claper.co`; asked through claper's own CLI (`bin/claper rpc`), `get_user_by_email_and_password("admin@claper.co", "claper")` answered **true**, an unknown e-mail answered false (control). The template routes `<sub>.<domain>` through traefik and the tunnel, so any claper a household installs can be logged into by anyone who knows claper's README. Not measured: whether any box runs claper today (R-632 lists it as never deployed), whether upstream reads an env var for the seed admin. **Needs:** a decision on the fix shape — (a) the catalog passes a generated admin password (if upstream supports it), (b) the controller changes the seeded admin's password after the first start, (c) pull claper from the catalog until (a)/(b). Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-claper-default-admin.txt`. | **OPEN — P1; owner: operator (fix shape), CC implements** |
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. **-- 2026-09-28 option (a) MEASURED — NOT ENOUGH:** `pct fstrim` 9201 + 9202 from the host (rc 0): `local-lvm` 62.32 % → **50.60 %**, free 20.3 → **26.6 GiB** — still under the 31 GiB the preflight needs. The pool is 53.9 GiB and guest 9201 itself holds ~26 GiB, so no trim can reach 31 GiB. The agent's own verdict after the trim: `tier=felhom-pbs due=true` (archive 2026-09-24T20:06:25Z, not proven). `nvme-scratch` has ~820 GiB free. **Left to the operator: (b) or (c).** `audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt` **-- 2026-09-28 14:13 CEST:** the first scheduled cycle after the trim was refused again ("needs 31.0 GiB free, has 23.9 GiB"). **Answered: each refusal DOES reach the hub** — `[WARN] host demo-hp-bb76ea restore-test FAILED …` at every host-report (every 15 min). **-- 2026-09-28 evening: CLOSED by option (b) (operator, `09` §3 decision 44).** demo-hp `restore_storage` → `nvme-scratch`; Proxmox first refused it (403 `Datastore.AllocateSpace` — the agent had no grant there, the 03 doc said so); with the operator's word the agent's `FelhomAgentStore` role was granted on `/storage/nvme-scratch` (user + token), and one restore test passed: restored + booted + verified + torn down in 8m46s, NVMe 50.3 → 74.1 GiB used at peak → back, `local-lvm` 58.46 % throughout. `audits/logins-nvme-2026-09-28/C/`. | **CLOSED — option (b), 2026-09-28** |
| **R-702** | **[P1-HIGH] Every claper install creates an admin `admin@claper.co` with the public password `claper`, and the app is published on the household's domain.** Measured 2026-09-28 on scratch guest 9202 (catalog `claper` template, `ghcr.io/claperco/claper:2.5` = 2.5.1): the image's own start command runs `Claper.Release.seeds`, which logs `Created default admin user: Email: admin@claper.co`; asked through claper's own CLI (`bin/claper rpc`), `get_user_by_email_and_password("admin@claper.co", "claper")` answered **true**, an unknown e-mail answered false (control). The template routes `<sub>.<domain>` through traefik and the tunnel, so any claper a household installs can be logged into by anyone who knows claper's README. Not measured: whether any box runs claper today (R-632 lists it as never deployed), whether upstream reads an env var for the seed admin. **Needs:** a decision on the fix shape — (a) the catalog passes a generated admin password (if upstream supports it), (b) the controller changes the seeded admin's password after the first start, (c) pull claper from the catalog until (a)/(b). Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-claper-default-admin.txt`. **-- 2026-09-28 evening: claper FIXED (catalog `9dc8a05`, controller v0.279.0 `after_install`)** — on a fresh install the seeded password is replaced by a generated one shown on the app page; measured on 9202: the default no longer authenticates, the generated one does, a wrong one does not, a restore keeps it. **The widened question (every app, operator ruling = decision 45) continues in R-707.** | **CLOSED — claper fixed 2026-09-28; the class → R-707** |
| **R-703** | **[P2] calcom v6.2.0 cannot start at its catalog memory limit — a fresh install crash-loops and the box stops it.** Measured 2026-09-28 on 9202: install from the live catalog → `crash_loop — 6 in 10m0s; STOPPING it (decision 28)`; one Start later, the container's own cgroup counted `oom_kill 1` per start at `memory.max` 805306368 (768 MiB) while `anon` reached ~700 MB during `turbo run start` (`signal: 'SIGKILL'`); Docker reported `OOMKilled=false` (R-528's shape). So calcom in the live catalog cannot run on any box. Not measured: the limit it needs. **Needs:** a measured limit (a bench watch at 1.5–2 GiB), then the catalog change. Until then calcom's PostgreSQL move is `inconclusive — the FROM version does not run`. Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-calcom-crash.txt`, `C0-calcom-memory.txt`. **-- 2026-09-28 later: FIXED in the catalog (`9555e73`, alone in its commit): memory 768M → 1536M.** Measured on 9202 through the drill catalog: at 2048M a 12-minute watch read `anon` steady ~780 MiB, 0 kills; at 1536M, sampled every 2 s from the container's birth, `anon` peaked at **817 MiB (53 %)** during start, `memory.peak` 1075 MiB, 0 kills, healthy. 1024M would sit at the 80 % `memory_tight` line. The seed route works at the new limit (`Calcom` fixture, catalog `b35fc7f`). `…/box/R703-0*.txt` | **CLOSED — catalog `9555e73`, 2026-09-28** |
| **R-704** | **[P3-LOW] The box's crash-loop stop (decision 28) outlives the app: after remove and reinstall, the new install is still held.** Measured 2026-09-28 on 9202: calcom crash-looped at 08:22 and 08:28 (`unhealthy_stop`, `crash_loop`, trip 2, recorded 08:28:59Z); it was then REMOVED through the product twice and installed fresh twice (09:14:51Z the last). At 09:45 the new, healthy install's Update was refused `409 held` with the crash-loop sentence („…újra és újra összeomlott…"), and `GET /api/stacks/calcom` carried the old `hold_reason` while `state=running`. Start lifted it (`the unhealthy stop is LIFTED by Start`). So a household that removes a crash-looping app and installs it again (the obvious fix) finds its updates refused for a crash of a previous install. Not measured: whether the nightly update leg also skips it; whether other holds (restore hold) behave the same. **Fix direction:** the remove clears the app's box-set holds, as `DeleteAppBackupPrefs` clears its backup preferences (R-474). Evidence: `audits/pg-calcom-claper-2026-09-28/box/calcom/hold.txt`, `…/box/calcom/move.txt`. **-- 2026-09-28 later: SECOND and worse instance, then FIXED in controller v0.278.0.** demo-hp's fresh nextcloud (installed 10:13) carried an UPDATE hold from a nextcloud of 2026-09-13 (set before v0.242.0 made removals clear update holds; nothing ever swept it). At the manual off-site run (15:17) the backup leg logged `Skipping volume dump for nextcloud — the app is HELD stopped`, captured no unit, and pushed a snapshot that `carried NO database dump and NO volume tar` — a freshly installed app silently NOT backed up. **Fix:** a removal also clears the crash-loop stop (`settings.ClearUpdateHold`), and a new install (plain or "use my kept data") drops a leftover update/crash-loop hold of an app that is not installed (`Router.dropLeftoverHold`); restore holds (R-379) untouched. Red-proofed RP4–RP6 (`audits/kept-offsite-2026-09-28/redproofs/`). Floor 0.278.0. **STILL OPEN: live proof of the install-time drop** (a box with a leftover hold on an uninstalled app). | **WATCHING — P2; owner: CC (install-time drop, live)** |
| **R-705** | **[P3-LOW] There is no way to run the night's chain now — only its pieces.** Asked by the operator 2026-09-28 (to finish a proof in the day). What exists (read from source, controller v0.278.0): the backup page's off-site run-now (`POST /backup/offbox/run`) runs the dump leg first (the R-44 pre-phase: DB dumps, volume dumps with brief app stops, unit capture) and then the push — used live on demo-hp 2026-09-28 15:17, 3m57s; the debug API has `backup/dbdump`, `backup/crossdrive` (Tier 2), `backup/integrity`, `backup/offsite-proof`. **Missing:** the automatic update leg (`RunUpdateLeg`, chained only to the scheduled off-site job) and the whole-guest backup (the agent's, on its own 24 h / 7 d cadence) have no manual trigger; the only way to run the chain in order is to move the backup window (`POST /backups/window`), which takes W..W+2h at least. **Needs:** a debug action "run tonight's chain now" (dump → Tier 2 → off-site → update leg, in order, one at a time), and an agent-side "whole-guest backup now" for demo boxes. Not built. | **OPEN — P3; owner: CC** |
| **R-706** | **[P3-LOW] Removing an app "with its backups" leaves its off-site verification copy on the drive.** Measured 2026-09-28 on demo-hp: after a full off-site restore of nextcloud (which leaves the downloaded copy in `backups/offsite-restore/nextcloud`, ~1 GB, by design, for the household to inspect), `POST /api/stacks/nextcloud/remove` with `remove_backups: true` removed the unit and listed `backup_paths_removed` WITHOUT the verification copy; it stayed until the restore page's own delete (`POST /backup/offbox/verify-copy/delete`, 302 `scratch_deleted`). A household that removes an app to free space keeps 1 GB it cannot see on the app list. **Fix direction:** the removal with backups also deletes the app's verification copy (the same `DeleteOffsiteRestoreCopy`). Evidence: `audits/kept-offsite-2026-09-28/E/E9-teardown.txt`. | **OPEN — P3; owner: CC** |
| **R-705** | **[P3-LOW] There is no way to run the night's chain now — only its pieces.** Asked by the operator 2026-09-28 (to finish a proof in the day). What exists (read from source, controller v0.278.0): the backup page's off-site run-now (`POST /backup/offbox/run`) runs the dump leg first (the R-44 pre-phase: DB dumps, volume dumps with brief app stops, unit capture) and then the push — used live on demo-hp 2026-09-28 15:17, 3m57s; the debug API has `backup/dbdump`, `backup/crossdrive` (Tier 2), `backup/integrity`, `backup/offsite-proof`. **Missing:** the automatic update leg (`RunUpdateLeg`, chained only to the scheduled off-site job) and the whole-guest backup (the agent's, on its own 24 h / 7 d cadence) have no manual trigger; the only way to run the chain in order is to move the backup window (`POST /backups/window`), which takes W..W+2h at least. **Needs:** a debug action "run tonight's chain now" (dump → Tier 2 → off-site → update leg, in order, one at a time), and an agent-side "whole-guest backup now" for demo boxes. Not built. **-- 2026-09-28 evening: the controller half BUILT (v0.279.0):** debug `POST /api/debug/backup/night-chain` runs dump → Tier 2 → off-site → update leg in order, one at a time, refused while anything else runs; the leg gets its normal length from its own start. Proven on 9202: 44 s, a second press 409, the leg deadline 22:00. **Still open: the whole-guest backup (agent side) has no manual trigger.** | **OPEN — P3, agent half only; owner: CC** |
| **R-706** | **[P3-LOW] Removing an app "with its backups" leaves its off-site verification copy on the drive.** Measured 2026-09-28 on demo-hp: after a full off-site restore of nextcloud (which leaves the downloaded copy in `backups/offsite-restore/nextcloud`, ~1 GB, by design, for the household to inspect), `POST /api/stacks/nextcloud/remove` with `remove_backups: true` removed the unit and listed `backup_paths_removed` WITHOUT the verification copy; it stayed until the restore page's own delete (`POST /backup/offbox/verify-copy/delete`, 302 `scratch_deleted`). A household that removes an app to free space keeps 1 GB it cannot see on the app list. **Fix direction:** the removal with backups also deletes the app's verification copy (the same `DeleteOffsiteRestoreCopy`). Evidence: `audits/kept-offsite-2026-09-28/E/E9-teardown.txt`. **-- 2026-09-28 evening: FIXED in controller v0.279.0** — a removal with its backups also deletes the verification copy and lists it among the removed paths (`TestR706_…`, red-proofed RP7). Not yet seen live (needs a full off-site restore then a removal). | **WATCHING — P3; owner: CC (live)** |
| **R-707** | **[P2] 37 apps still start with a login a stranger can take (`09` §3 decision 45).** Audit of all 53 apps: `app-catalog-felhom.eu/FIRST-ADMIN.md` (class, fix route, status, measured or read). Open: **3 hard-coded defaults** — calibre-web (`admin / admin123`, measured working on demo-hp and 9202; its own `cps.py -s` route needs a generated password WITH a special character — our generator is letters+digits, a controller change), mealie (`changeme@example.com / MyPassword`), wger (`admin / adminadmin`); **34 open first-run screens** (the first visitor creates the admin: actualbudget, adventurelog, audiobookshelf, calcom, docmost, emby, ghost, gitea, gramps-web, home-assistant, homebox, immich, jellyfin, komga, n8n, navidrome, opengist, outline, papra, plant-it, radarr, rallly, recipe-importer, romm, seerr, sonarr, sparkyfitness, tandoor, termix, uptime-kuma, vikunja, wanderer, wishlist, zipline). **Stale notes:** romm's `default_creds` `admin / admin` answers 401 on demo-hp (like a wrong password) — the page now warns with a login that does not exist; zipline's looks stale too. **Measured on demo-hp 2026-09-28 (read-only):** bookstack's default still logs in on the INSTALLED app (the fix is for new installs; the page now warns). Each fix: route (a) env or (b) the app's own CLI/API via `after_install:`, proven on 9202 with the default failing and the generated password working; route (c) a page sentence. Several sessions (operator, 2026-09-28). | **OPEN — P2; owner: CC** |
| **R-708** | **[P3-LOW] grafana falls back to password `admin` when its admin field is empty.** `templates/grafana/docker-compose.yml:18` `GF_SECURITY_ADMIN_PASSWORD=${…:-admin}` (read 2026-09-28, the audit). Today the field is generated and required, so it is never empty on a normal install — but an edit, an import or a restore that drops the value would publish grafana with `admin / admin`. **Fix direction:** no default in the compose (`${GF_SECURITY_ADMIN_PASSWORD:?}` refuses to start instead). | **OPEN — P3; owner: CC** |
| **R-709** | **[P3-LOW] The deploy page writes the generated admin passwords of installed apps into its HTML.** `internal/web/templates/deploy.html` renders a `type: password` field's decrypted value into a disabled `<input value=…>` (read 2026-09-28; used by the proofs of R-702/R-707 to read the first password as the household sees it). `type: secret` fields got a fetch-on-demand reveal in R-254; `type: password` fields did not. The page needs a login, so this is exposure to a logged-in session's HTML (browser cache, a shared screen, a saved page), not to strangers. **Fix direction:** the R-254 reveal for password fields too. | **OPEN — P3; owner: CC** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.