09-update-architecture.md: the update path finally has a document, and it is a living one
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
R-438's document half. It records how an update works AS MEASURED, quotes the RestartStack comment that proves the restart half was CHOSEN (a design decision is not a defect), carries the three operator rulings of 2026-09-02, strikes the word 'rollback' (once a migration has run the old image will not start), states the target shape, and lists the seven slices with a status each. R-438 and R-440 amended and BOTH STAY OPEN: the mechanism is documented, not changed. Nothing closed, so CLOSED-ITEMS.md is untouched. Eight new register rows, 194 -> 202: R-446 (Naprakesz can be false for the 23 floating pins), R-447..R-451 (one per remaining slice, with a rank and an owner), R-452 (no gate enforces catalog_since - the runner fetches at --depth 1), and R-453 (the vaulted dashboard password is stale on BOTH demo boxes, which is what stopped the badge render from being validated live). Live evidence for slices 1 and 2 in documentation/tests/. The record is PROVEN LIVE through the boot reconciler on demo-hp - one entry per compose service, digests matching ground truth read independently. The badge RENDER is not, and the five attempts are listed rather than summarised.
This commit is contained in:
@@ -1,214 +1,63 @@
|
||||
# REPORT — SPIKE: what an app update actually does, and which other paths do it too (2026-09-01)
|
||||
# REPORT — update arc slices 1 & 2: documents, register, roadmap, capability map (2026-09-02)
|
||||
|
||||
**Task class: Spike.** No production code was written in any repo. Output is a findings doc, register
|
||||
rows, two architecture updates and one operator decision.
|
||||
*Overwritten each session. Nothing durable lives only here.*
|
||||
|
||||
**Findings doc:** `documentation/audits/SPIKE-app-update-2026-09-01.md`
|
||||
**Evidence:** `documentation/audits/evidence-spike-app-update-2026-09-01/` (19 files)
|
||||
## What this session changed in THIS repo
|
||||
|
||||
---
|
||||
|
||||
## 1. The Phase 1 answer, stated unambiguously
|
||||
|
||||
**YES — `docker compose up -d` upgrades an app whose compose file has already moved, and the Restart
|
||||
button does it.**
|
||||
|
||||
Produced by **variants 1a and 1b, and independently by 1c-ii**:
|
||||
|
||||
- **1a** (target image ABSENT locally): the restart took **18.3 s**, ended with the container on the
|
||||
new tag, and **left the new image in the local store where it had not been** — so `up -d` performed
|
||||
a network pull, in a code path that contains no pull step.
|
||||
- **1b** (target image PRESENT): **0.5 s**, container recreated onto the new tag, image id unchanged.
|
||||
- **1d, the negative control** (file NOT edited): container id, image id, digest and `StartedAt` all
|
||||
identical — `up -d` did not even recreate. **The method can show "no change" when nothing changed.**
|
||||
- **1c-ii** (unattended): the boot reconciler brought a failed app back on the **new** version with
|
||||
nobody pressing anything.
|
||||
|
||||
**And the answer to 1c is NO, which narrows the exposure:** a hard guest reset upgraded nothing.
|
||||
Docker's `restart: unless-stopped` restored the existing containers on the old image and the
|
||||
reconciler logged its own verdict — `Boot reconciliation: no boot-orphaned apps (nothing to start)`.
|
||||
|
||||
**Per the task's own rule, no fix is proposed.** The behaviour may have been chosen — `RestartStack`
|
||||
says so in a comment — so it goes to the operator as a decision.
|
||||
|
||||
---
|
||||
|
||||
## 2. Confirmed baselines actually used
|
||||
|
||||
| Repo | at §1 of the task | measured at start | end | moved? |
|
||||
|---|---|---|---|---|
|
||||
| felhom-controller | `960d29b0612c` | `960d29b0612c` | `960d29b0612c` | **no — read only** |
|
||||
| felhom.eu | `1d59353df437` | `1d59353df437` | this session's documents | as planned |
|
||||
| app-catalog-felhom.eu | `29edad9c5bf4` | `29edad9c5bf4` | `5d8f25f`, **tree identical to `29edad9c5bf4`** | reverted |
|
||||
| felhom-agent | `4586f0f7f6d1` | `4586f0f7f6d1` | `4586f0f7f6d1` | not touched |
|
||||
|
||||
**None had moved since the task was written.** Live controller on demo-hp: 0.232.0.
|
||||
|
||||
---
|
||||
|
||||
## 3. Per-phase results, each with its control
|
||||
|
||||
| Phase | Result | Control |
|
||||
|---|---|---|
|
||||
| **0** | R-438, R-439, R-440 filed, each with rank and owner | ids confirmed free by grepping all three register files (R-437 was highest) |
|
||||
| **1** | **THE GATE: YES.** 1a/1b upgrade; 1c does not; 1c-ii upgrades unattended | **1d** — unchanged file, no change at all |
|
||||
| **2** | Sync rewrote a DEPLOYED app's file at 17:45:17Z; container unchanged; nothing told the customer | positive (`BentoPDF`, 4 hits) + negative (`zzz-never-present`, 0) on both customer pages |
|
||||
| **3a** | Pull failure: HTTP 500, **app untouched and still running** | re-run as a Restart: also 500, app still ran |
|
||||
| **3b** | **HTTP 200 "update completed" over a crash-looping app**; alarm fired 5m16s later | the alarm is a positive observable (`Event pushed: app_start_failed`), plus the detector heartbeat `1 currently down` |
|
||||
| **4** | demo-hp and demo-felhom: **zero drift by tag**. But **2 of 6 floating pins have already MOVED upstream** | both fully-pinned images (`romm:5.0.0`, `pdo:2.0.5`) were SAME |
|
||||
| **5** | Existing safety dump is **database-only**; DB dumps are 48 KB–395 KB; the **file** half is the cost, and it does not fit for a large app | demo-hp is too young to price it — stated, and Campaign 10's measured figures cited instead |
|
||||
| **6** | **App data CANNOT be rolled back.** Old image refuses to start on migrated data | returning to 32.0.9 restored **both** seeded markers byte-identical — the data is not destroyed, only the downgrade is blocked |
|
||||
|
||||
---
|
||||
|
||||
## 4. The exact symbols that bring an app back — found by reading
|
||||
|
||||
| path | `file:symbol` |
|
||||
| file | change |
|
||||
|---|---|
|
||||
| boot reconciler | `felhom-controller/controller/internal/bootrecon/bootrecon.go:269` — `Reconciler.Run` → `StartStack` |
|
||||
| its scheduler | `controller/cmd/controller/main.go:2127` — `runBootReconcile`, called at `:450` |
|
||||
| app-stop guard | `controller/internal/backup/appstop_marker.go:283` — `AppStopGuard.Recover` → `StartStack` |
|
||||
| its hold-aware wrapper | `controller/cmd/controller/main.go:2008` — `gatedAppStopStarter.StartStack` |
|
||||
| drive-return gate | `controller/internal/web/intermediary.go:222` — `Server.restartStacks` → `StartStack` |
|
||||
| guest-boot change | `controller/internal/web/intermediary.go:458` — `Server.processGuestBootChange` |
|
||||
| quiesce restart | `controller/internal/quiesce/quiesce.go:733` — `Loop.restartAll` |
|
||||
| off-site reconstitution | `controller/internal/backup/offbox_reconstitute.go:692` — `Manager.ReconstituteFromOffsite` |
|
||||
| **`documentation/architecture/09-update-architecture.md`** | **CREATED — the deliverable.** Its absence was R-438. A LIVING document: every slice of this arc updates it in the same session. |
|
||||
| `documentation/tests/VALIDATION-update-slice12-2026-09-02.md` | CREATED — the live evidence, copied off demo-hp at the end of the phase that produced it. |
|
||||
| `documentation/backlog/OPEN-ITEMS.md` | R-438 and R-440 amended (both stay OPEN); **8 new rows: R-446..R-453**. |
|
||||
| `documentation/backlog/ROADMAP.md` | the update arc added as one item, naming the capability-map rows it flips. |
|
||||
| `documentation/architecture/00-capability-map.md` | one new row: what version a box runs, and whether it is behind. |
|
||||
| `STATUS.md` | one new operator item (9) and a new lead paragraph. |
|
||||
|
||||
Full table of all 13 non-API call sites: findings doc §8.
|
||||
## The architecture document — what it settles
|
||||
|
||||
---
|
||||
1. **How an update works today, as measured** — `UpdateStack` is `pull` then `up -d --remove-orphans`;
|
||||
the syncer overwrites a deployed app's compose on a 15-minute cycle with no deployed check;
|
||||
**thirteen** non-API call sites end in `compose up -d`.
|
||||
2. **What was chosen, and by whom** — `RestartStack`'s own comment, quoted. **A design decision is not
|
||||
a defect.** What was never decided is what the syncer does underneath a deployed app.
|
||||
3. **The three operator rulings of 2026-09-02** — verified backup as a precondition; the support window
|
||||
runs on how far behind the CATALOG a box is; automatic within a major, never across one.
|
||||
4. **The vocabulary ruling** — "rollback" is struck. The available shapes are ABORT and RESTORE.
|
||||
5. **The target shape** — the live compose file becomes derived from a pin in `app.yaml`.
|
||||
6. **The seven slices**, each with a status line. 1 and 2 are shipped.
|
||||
7. **Known limitations**, including the floating-tag one.
|
||||
|
||||
## 5. Register rows opened, updated or re-ranked
|
||||
## Register
|
||||
|
||||
| row | what | owner |
|
||||
**194 rows before, 202 after.** Nothing closed, and that is stated rather than implied: R-438 and
|
||||
R-440 are **amended and stay OPEN** — the mechanism is now documented, not changed — so nothing moved
|
||||
to `CLOSED-ITEMS.md` and that file is untouched.
|
||||
|
||||
| row | what | state |
|
||||
|---|---|---|
|
||||
| **R-438** | opened P1-HIGH, then **updated with the live evidence** (sync overwrite measured; consequence measured; the in-source design intent found and recorded, which narrows it) | **VIKTOR rules, CC measures** |
|
||||
| **R-439** | opened P3-LOW, then **updated — the severity survives but its stated reason was imprecise** (`isOperationalState` counts `restarting`/`degraded` as operational) | CC |
|
||||
| **R-440** | opened P2-MEDIUM, then **updated: two floating pins have ALREADY moved, with a passing control** | CC |
|
||||
| **R-441** | NEW — the restore path and the sync disagree about the image, and the sync wins within 15 minutes | CC measures, VIKTOR rules |
|
||||
| **R-442** | NEW — **P1-HIGH**: `remove_hdd_data:true` is inert with no `paths.hdd_path`; 128 MB left, API said neither removed nor preserved | CC |
|
||||
| **R-443** | NEW — the Update button reports success over an app it has broken | CC proposes, VIKTOR rules |
|
||||
| **R-444** | NEW — nothing runs `pct fstrim`; demo-hp's thin pool held ~23.8 GB of freed blocks | CC |
|
||||
| **R-445** | NEW — hub app telemetry outlives the app and sets a fleet-wide recommendation | VIKTOR rules, CC implements |
|
||||
| R-446 | „Naprakész" can be FALSE for the 23 floating pins | OPEN, P2-MEDIUM, CC |
|
||||
| R-447 | slice 3 — make the live compose DERIVED | **BLOCKED** on an operator ruling, P1-HIGH |
|
||||
| R-448 | slice 4 — a guarded update (subsumes R-443) | READY, P2-MEDIUM |
|
||||
| R-449 | slice 5 — an upgrade test that runs again | READY, P2-MEDIUM |
|
||||
| R-450 | slice 6 — version sequence; an engine change gets its own edge | READY, P2-MEDIUM |
|
||||
| R-451 | slice 7 — a fleet sweep (needs a hub change: no image field is reported) | READY, P3-LOW |
|
||||
| R-452 | no gate enforces `catalog_since` (`--depth 1` has no parent to diff) | READY, P3-LOW |
|
||||
| R-453 | **the vaulted dashboard password is stale on BOTH demo boxes** | **WAITING-ON-OPERATOR**, P2-MEDIUM |
|
||||
|
||||
**Nothing was closed and nothing was re-ranked.** The ranking of R-438 relative to existing rows is
|
||||
Viktor's, and I have not moved anything.
|
||||
## Live validation
|
||||
|
||||
---
|
||||
Full evidence: `documentation/tests/VALIDATION-update-slice12-2026-09-02.md`.
|
||||
|
||||
## 6. Claims in the task that turned out to be wrong, named
|
||||
**PROVEN LIVE on demo-hp at controller 0.233.0**, through a real production caller (the boot
|
||||
reconciler — no hand-set state): `bentopdf` recorded 1 service, `bookstack` recorded **2**, keyed by
|
||||
compose service name, **all three digests matching ground truth read independently beforehand**.
|
||||
Encrypted secrets byte-identical across the write. Both apps up and healthy; nothing provisioned.
|
||||
|
||||
1. **"a power cut … is an unattended three-major-version upgrade"** — not as stated. A plain power cut
|
||||
upgraded nothing (measured). The unattended upgrade needs *"and the app did not come back"*.
|
||||
**The exposure is smaller than the operator page claims.**
|
||||
2. **"Five other code paths end in `compose up -d`"** — **thirteen** non-API call sites, nine files.
|
||||
3. **"Phase 4 — demo-hp, demo-felhom and Peti's box"** — Peti's box is DOWN, not enrolled, and
|
||||
`runbooks/target-selection.md:161` says *"No access route from DooPlex"*. Two boxes measured live;
|
||||
Peti's row is UNKNOWN, with what is knowable taken read-only from the hub and the catalog history.
|
||||
4. **R-439's severity reason** — right conclusion, imprecise reason (see §5).
|
||||
5. **R-440's "23 pins"** — **exactly right** (79 image lines / 53 apps / 66 distinct; 23 with no patch
|
||||
component). One arguable 24th named rather than rounded away.
|
||||
6. **The catalog history figures** — spot-checked and **all correct**: 153 commits, 53 apps, 0 files
|
||||
with upgrade metadata, and all four multi-major bumps confirmed by commit hash.
|
||||
7. **My own method, corrected in-flight:** the first customer-page search used `grep -o "2.8.6"`, whose
|
||||
unescaped `.` produced two false hits; `grep -F` gives zero. The controls caught it.
|
||||
**NOT live-validated: the rendered badge.** The vaulted dashboard password is stale on both demo
|
||||
controllers (R-453) and there is no operator route to a customer's password. Five attempts are listed
|
||||
in §4 of the validation file. The render is covered by tests that render the PRODUCTION templates.
|
||||
|
||||
---
|
||||
## Sibling repos
|
||||
|
||||
## 7. Evidence handling
|
||||
|
||||
**Evidence was written directly into `documentation/audits/evidence-spike-app-update-2026-09-01/` on
|
||||
DooPlex as each phase produced it — before every revert, including the intermediate ones.** The
|
||||
catalog revert (18:10:29Z), the compose reverts, the guest reset and the Nextcloud teardown all
|
||||
happened after their evidence was already off the machine. **Nothing was lost and nothing had to be
|
||||
reproduced.**
|
||||
|
||||
---
|
||||
|
||||
## 8. Observations
|
||||
|
||||
1. **The restore path writes the recovery unit's OLD image pin into the live stack dir, and the
|
||||
catalog syncer overwrites it again within 15 minutes.** The overwrite half is measured on demo-hp;
|
||||
that the restore writes to that same path is read at `cmd/controller/main.go:2570`, not measured —
|
||||
both halves are graded as such in the row. **FILED: R-441**
|
||||
|
||||
2. **`remove_hdd_data: true` removed nothing.** 128 MB of app data stayed on the drive while the API
|
||||
returned 200 with `hdd_paths_removed:null, hdd_paths_preserved:null`. Root cause established with
|
||||
controls: `Paths.HDDPath` has no default and demo-hp's `controller.yaml` does not set it, so
|
||||
`ParseComposeHDDMounts` returns nil on its first line. **FILED: R-442**
|
||||
|
||||
3. **An update can report success over an app it has just broken.** HTTP 200 and "updated
|
||||
successfully" while the container was already crash-looping; the truth reached the customer 5m16s
|
||||
later through the dead-app alarm rather than through the update itself. **FILED: R-443**
|
||||
|
||||
4. **The PVE thin pool was holding ~23.8 GB of blocks the guest had already freed**, and `fstrim` from
|
||||
inside the unprivileged container is refused; `pct fstrim` from the host reclaimed it. Nothing runs
|
||||
it on the fleet. **FILED: R-444**
|
||||
|
||||
5. **The hub kept app telemetry for an app that no longer exists anywhere**, and it now sets a
|
||||
fleet-wide suggested memory limit for Nextcloud derived from a 15-minute crash-looping throwaway.
|
||||
Retained rather than cleared, because the reset is irreversible and on the operator's own surface.
|
||||
**FILED: R-445**
|
||||
|
||||
6. **The syncer's debug hash line cannot show what it claims to show.** `logFileHashes`
|
||||
(`internal/sync/sync.go:386`) reads the destination *after* the write, so it printed
|
||||
`src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed)` — the same hash twice, with the word
|
||||
"changed". **NOT-A-FINDING:** it is DEBUG-only and the `Updated <app>/<file>` INFO line immediately
|
||||
above it carries the fact correctly, so nothing is lost and no behaviour is wrong. Recorded so the
|
||||
next person reading a sync log does not try to learn from those two hashes, which cannot differ.
|
||||
|
||||
---
|
||||
|
||||
## 9. Teardown — all three layers
|
||||
|
||||
**Layer 1 — the machine.** `bentopdf` restored to its catalog tag `v2.8.6`, digest
|
||||
`sha256:eaeea1e447205a79…`, **byte-identical to the run's baseline**. The throwaway `nextcloud` stack
|
||||
removed via the product's own endpoint; containers, all three volumes and `app.yaml` gone. The 128 MB
|
||||
the product failed to remove (observation 2) deleted by hand along with its backup dirs;
|
||||
`find /mnt -iname "*nextcloud*"` returns nothing. Five images this run pulled removed by **targeted
|
||||
`docker rmi`** — **no `prune` of any kind was run anywhere**. No guest was created; guest 9201 was
|
||||
hard-reset once by design and returned with all nine apps.
|
||||
|
||||
**Layer 2 — the host.** `local-lvm` 68.97% → 70.91% during the run → **26.78%** after `pct fstrim
|
||||
9201`. The run's ~1.05 GiB was returned and 23.8 GB more that predated it. Guest filesystems back to
|
||||
pre-run values (`/` 957 M, `/mnt/sys_drive` 12 G, `hdd_1` 5.5 G).
|
||||
|
||||
**Layer 3 — the hub. This run provisioned NOTHING.** No customer record and no appliance record was
|
||||
created; the customers list is unchanged at five rows, identical to the list read at the start. The
|
||||
existing `demo-hp` customer was used. What the run *did* create is hub **events** (`app_start_failed`,
|
||||
plus deploy/remove for the throwaway app) — **retained deliberately**, because the event log is an
|
||||
append-only record and deleting from it to tidy a test damages the surface this project relies on for
|
||||
history. One residue is **retained rather than cleared** with the reason and the exact one-line command
|
||||
recorded in the findings doc §13 (observation 5).
|
||||
|
||||
**Nothing on `demo-felhom`, `ep0`, DooPlex or Peti's box was modified. Peti's box was never contacted.**
|
||||
|
||||
---
|
||||
|
||||
## 10. Final verification
|
||||
|
||||
```
|
||||
felhom-controller: git status --porcelain → empty
|
||||
HEAD = origin/main = 960d29b0612c
|
||||
go build ./... → OK
|
||||
go vet ./... → OK
|
||||
go test ./... → rc=0, 28 packages, 0 FAIL
|
||||
```
|
||||
|
||||
**The controller tree was left untouched, and that is proven rather than asserted.**
|
||||
|
||||
---
|
||||
|
||||
## 11. My own mistakes
|
||||
|
||||
1. **My first customer-page search used a regex where I needed a literal.** `grep -o "2.8.6"` treats
|
||||
`.` as a wildcard and reported two hits on a page that contains none. I caught it only because the
|
||||
task requires a control on every such search, and re-ran with `grep -F`. **The rule earned its
|
||||
place in the same session it was applied.**
|
||||
2. **My first poll for "nextcloud is healthy" matched the wrong container.** The break condition
|
||||
matched `healthy` anywhere in the line and fired on `nextcloud-redis`. Corrected to an exact-name
|
||||
filter. No result depended on it.
|
||||
3. **I renumbered `STATUS.md` badly on the first attempt**, leaving the list running 1–6, 8, 9 with no
|
||||
item 7, and left the section's own lead line saying one thing was waiting when there were two.
|
||||
Both fixed. **This is the ranking-paragraph-goes-stale trap (R-405) in miniature**, and it appeared
|
||||
within minutes of my writing about it.
|
||||
- `felhom-controller` **v0.233.0** — `8025304acc0a`, deployed to demo-hp and verified healthy.
|
||||
- `app-catalog-felhom.eu` — `69761cf91bfc` (backfill) + `8220f8d` (REPORT).
|
||||
|
||||
@@ -1,6 +1,11 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-01 (third pass) — the backup work is FINISHED for beta, and I have written down
|
||||
**Updated 2026-09-02 — the box now writes down which version of each app it is running, and shows one
|
||||
small label saying whether it is up to date. No version numbers, and nothing about updating changed.
|
||||
The writing-down half is proven on the real machine; for the label I need one password from you
|
||||
(item 9). Nothing is broken while it waits.**
|
||||
|
||||
**Earlier 2026-09-01 (third pass) — the backup work is FINISHED for beta, and I have written down
|
||||
where it stops. One alarm that was telling you something untrue is fixed and live (hub 0.111.1).
|
||||
ONE THING is waiting on you: two short e-mails to Hetzner, drafted and ready to paste.**
|
||||
|
||||
@@ -18,7 +23,7 @@ not an evening's work.**
|
||||
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
||||
nothing.*
|
||||
|
||||
1. **Two things are waiting on you — item 4 (send two e-mails) and item 7 (one design decision, new today).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on
|
||||
1. **Three things are waiting on you — item 4 (send two e-mails), item 7 (one design decision) and item 9 (one password, new today, two minutes).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on
|
||||
the real machines:
|
||||
- the background job that could delete a live restore's lock now waits its turn — and the check
|
||||
that finds the next one like it is a test, not a comment, so it cannot come back quietly;
|
||||
@@ -111,6 +116,36 @@ nothing.*
|
||||
Full measurement, with the controls and the quoted output:
|
||||
`felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md`.
|
||||
|
||||
9. **I need the dashboard password for `demo-hp`, or your go-ahead to reset it. Two minutes, and
|
||||
nothing is broken while it waits.**
|
||||
|
||||
Today the box learned to write down which version of each app it is really running, and to show
|
||||
the customer one small label: **„Naprakész"** or **„Frissítés elérhető — 45 napja"**. No version
|
||||
numbers — a household cannot act on `26.05.2`.
|
||||
|
||||
**The writing-down half is proven on the real machine.** I made two apps fail to come back, the box
|
||||
repaired them by itself, and it wrote down exactly what it installed — one line per container, with
|
||||
the fingerprint that cannot lie. I checked those fingerprints against the machine independently and
|
||||
they match. Both apps are up and healthy, and no data was touched.
|
||||
|
||||
**The label half I could not look at.** To open a customer page I have to log in as the customer,
|
||||
and the password we keep on file no longer works — on `demo-hp` **or** on `demo-felhom`. I tried
|
||||
both, and the box's own log says "wrong password", not "wrong address". There is no operator route
|
||||
to a customer's password: it is only ever e-mailed to them.
|
||||
|
||||
**Two ways forward. Pick one:**
|
||||
|
||||
- **Send me the current `demo-hp` dashboard password.** I use it for two page loads and nothing
|
||||
else. Simplest, and it changes nothing on the box.
|
||||
- **Let me reset it back to the one on file.** Same thing we did on 9 August. It also repairs the
|
||||
stored password, so the next session does not lose this half hour again.
|
||||
|
||||
**My pick: let me reset it** — otherwise the same wall is there next week, on both machines.
|
||||
|
||||
**If you do nothing:** nothing breaks and no customer is affected. The label is covered by tests
|
||||
that render the real pages, so this is a confirmation, not a discovery. The box goes on recording
|
||||
versions either way.
|
||||
|
||||
8. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
||||
as it misled one by an hour.
|
||||
|
||||
@@ -104,6 +104,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
|
||||
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
|
||||
| Protected infra stacks can't be stopped/removed from UI | controller | **PROVEN-LIVE** | `CAMPAIGN-nomercy` + `RERUN-p1p3` T-SEC-PROTECTED (refuse stop/remove, stay Up) | (Cited `CAMPAIGN-2` T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal) |
|
||||
| **What VERSION a box is running, and whether it is behind the catalog** | controller v0.233.0 + catalog `69761cf` | **PROVEN-LIVE for the RECORD; IMPLEMENTED for the BADGE** — and the split is the point, not a hedge. | **`tests/VALIDATION-update-slice12-2026-09-02.md`** — on demo-hp 0.233.0, through a REAL production caller (`bootrecon → StartStack → compose up -d → recordInstalledImages`, no hand-set state): `bentopdf` recorded **1** service and `bookstack` recorded **2**, keyed by compose SERVICE name, and **all three digests match the ground truth read independently from the containers before anything was touched**. `catalog_since` reached the box on the normal 15-minute sync. bookstack's two encrypted secrets are byte-identical across the write. Unit side: `installed_test.go` + `updatebadge_test.go`, incl. a wiring test through a real `RestartStack`, an AST walk of all four call sites, and three companion red-proofs. | **THE UNEXERCISED LEG, NAMED: the rendered badge has never been seen on a live page.** The vaulted dashboard password is stale on BOTH demo controllers (`Hibás jelszó`, confirmed against the controller's own log, and the same on demo-felhom), and there is no operator-side route to a customer's dashboard password (R-119). Every INPUT the badge reads is verified live; the render is covered only by tests that render the PRODUCTION templates. **Absent means UNKNOWN, never current** — a legacy `app.yaml` renders NOTHING, red-proved. **No version number is shown to the customer** and **no registry is queried**, so „Naprakész" CAN BE FALSE for the 23 floating pins (**R-446**). Nothing about updating changed: R-438, R-440, R-441, R-443 all stand. Reasoning: `architecture/09-update-architecture.md`; remaining slices R-447..R-452 |
|
||||
| Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad `backup:` block degrades to legacy, loudly) | controller v0.132, catalog | **PROVEN-LIVE** | `CAMPAIGN-2` T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs | |
|
||||
| Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) | agent v0.94.0→**v0.95.0**, controller v0.169.0→v0.171.0→**v0.215.0**, hub v0.73.1 | **PROVEN-LIVE (healthy path + delivery + the severity wire).** **IMPLEMENTED, NOT proven-live: the Hiba-from-counters path** (v0.215.0) — it has never fired on real hardware, only against the committed fixture's values in unit tests (**R-332**) | **2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix):** the card on guest 9201 now shows BOTH real disks with **real verdicts + human model labels** — **„AirDisk 512GB SSD" → Rendben (34°C)** (the system SSD, via LVM/dm resolution) and **„TOSHIBA MQ04ABF100" → Rendben (30°C)** (the USB, via union-path SMART). `/disks` carries `smart.health=PASSED` + `model_name` for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (`SPIKE-smart-coverage-2026-07-25.md` had proven both disks answer `smartctl -a -j` PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. **Notification pipeline PROVEN-LIVE 2026-07-24** — a `disk_health_degraded` POST (the exact `notify.PushEvent` wire call) was **400-rejected by hub v0.73.0** and **200-accepted + „Operator email sent" by hub v0.73.1** | No new smartctl load; feature-detect by payload presence → **MinAgent floor unchanged**; no sudoers/`-d sat` change. **No global banner** (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin `local` on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + `model_name` capture. **2026-08-14 — a genuinely failing disk HAS now been seen, and it broke three assumptions** (`audits/DIAG-smart-passed-trap-2026-08-14.md` + two committed fixtures: raw `smartctl -a -j` and 406 `smartd` lines from ST3000VX010 S/N Z6A07P2G). **(1)** `smart_status.passed` is STRUCTURALLY incapable of failing on unreadable sectors — attrs 187/197/198 all carry `thresh: 0` and a normalized value floors at 1 — so the drive read PASSED at 352 pending sectors and 1001 uncorrectable reads. **(2)** The alert it did produce carried severity `"warn"`, which the hub coerces to `info` and never emails: **the counterfactual is ZERO emails about this drive** (R-328, fixed controller v0.215.0, and the `warning`-vs-`warn` pair proven side by side in `notification_log` on 2026-08-14 — `sent` vs no row at all). **(3)** The old check spoke once and forgot on restart, so between 8 and 352 sectors it emitted nothing. v0.215.0 adds the sustained/count/heat Hiba rules, persisted state and an hourly cadence. **The verdict half of that arm remains unit+red-proof covered only** — no live drive has reached Hiba from counters (R-332). **SMART history/trending (hub-side) PARKED** (ROADMAP R-73) |
|
||||
| App crashes → customer notified (one event per transition, no flapping spam) | controller v0.120, hub v0.48 | **IMPLEMENTED** | controller v0.120.0 (dead-app alerting, `app_start_failed`, one-event-per-transition red-proofs); `CAMPAIGN-3` F11 surfaced the gap | End-to-end crash→customer-email delivery never live-confirmed (6B deferred / 6C inconclusive: clean stop ≠ crash); anti-spam unit-proven |
|
||||
|
||||
@@ -0,0 +1,230 @@
|
||||
# 09 — How an app update works, and what it is becoming
|
||||
|
||||
> **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.**
|
||||
> Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen
|
||||
> deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who
|
||||
> did not know it was one.
|
||||
|
||||
**This file carries the REASONING. The register (`backlog/OPEN-ITEMS.md`) carries the work. The
|
||||
source is the truth.** Nothing here is invented: every mechanism claim is cited either to
|
||||
`audits/SPIKE-app-update-2026-09-01.md`, which measured it live, or to live source at `file:symbol`.
|
||||
|
||||
---
|
||||
|
||||
## 1. How an update works today, as measured
|
||||
|
||||
### 1.1 The button
|
||||
|
||||
`Manager.UpdateStack` (`felhom-controller/controller/internal/stacks/manager.go:1199`) is two compose
|
||||
commands and nothing else:
|
||||
|
||||
```
|
||||
compose pull → compose up -d --remove-orphans
|
||||
```
|
||||
|
||||
**No safety copy. No rollback. No hold. No verification.** Confirmed by reading and across six live
|
||||
updates (spike §10 item 6). A pull FAILURE is handled correctly — `UpdateStack` returns after the
|
||||
failed pull and never reaches `up -d`, so the running app survives untouched (measured twice, spike
|
||||
§4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443.
|
||||
|
||||
### 1.2 The catalog syncer moves the file underneath a deployed app
|
||||
|
||||
`Syncer.copyTemplates` (`felhom-controller/controller/internal/sync/sync.go:319`) copies
|
||||
`docker-compose.yml` and `.felhom.yml` into **every** stack folder on a 15-minute cycle
|
||||
(`internal/config/config.go:351`, default `15m`). **It has no deployed check of any kind.** Its only
|
||||
guard is a sha256 content compare in `copyIfChanged` (`sync.go:403`) and its only exclusion is
|
||||
`app.yaml`. The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the
|
||||
sync does not restart anything.**
|
||||
|
||||
That is why a deployed app's compose file and its running containers can disagree **indefinitely**.
|
||||
Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the
|
||||
deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3).
|
||||
|
||||
### 1.3 Thirteen other paths end in `compose up -d`
|
||||
|
||||
Excluding the three API actions, **13 call sites across 9 files** call `StartStack` or `RestartStack`,
|
||||
and every one ends in `compose up -d` against the live compose file (spike §8 — the task that
|
||||
commissioned the spike said five; the count is thirteen). They include the **boot reconciler**
|
||||
(`bootrecon.go:269`), the **app-stop guard's** recovery (`appstop_marker.go:283`), the **drive-return
|
||||
gate** (`intermediary.go:222`), the quiesce restart-after-backup, off-site reconstitution, and every
|
||||
restore path.
|
||||
|
||||
**So an upgrade can happen with nobody pressing anything** — measured, spike §2 variant 1c-ii, where
|
||||
a boot reconciliation started an app on a newer image at 17:55:44Z.
|
||||
|
||||
### 1.4 One fear is measured SMALLER than it was stated
|
||||
|
||||
A plain power cut does **not** upgrade anything. Docker's own `restart: unless-stopped` puts the
|
||||
existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs
|
||||
`up -d` — it says so in its own words: `no boot-orphaned apps (nothing to start)` (spike §2, variant
|
||||
1c, a positive observable and not an absent log line).
|
||||
|
||||
**The unattended upgrade needs the narrower precondition: *"and the app did not come back."*** Saying
|
||||
so is more useful than leaving the scarier version standing.
|
||||
|
||||
---
|
||||
|
||||
## 2. What was chosen, and by whom
|
||||
|
||||
`Manager.RestartStack` (`internal/stacks/manager.go:1161`) carries this comment, and it predates the
|
||||
whole arc:
|
||||
|
||||
> *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template
|
||||
> changes (new images, healthchecks) are picked up. Plain `docker compose restart` only sends
|
||||
> SIGTERM+start to existing containers without re-reading the compose file or env."*
|
||||
|
||||
**So the restart behaviour was chosen, deliberately, and written down. A design decision is not a
|
||||
defect.** What was never decided — and is recorded nowhere — is what happens once the catalog syncer
|
||||
moves the file underneath a *deployed* app, and whether the choice was meant to extend to the thirteen
|
||||
unattended call sites. **That gap is R-438, and it stays open**: this document records the mechanism;
|
||||
it does not change it.
|
||||
|
||||
---
|
||||
|
||||
## 3. The three operator decisions (2026-09-02)
|
||||
|
||||
These are rulings, not proposals. Anything specced against a different assumption is wrong.
|
||||
|
||||
1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the
|
||||
update path. The guest-snapshot alternative is to be **spiked before anything is designed around
|
||||
it**. Context: the existing safety machinery (`Manager.writeSafetyDump`,
|
||||
`internal/backup/offbox_reconstitute.go:207`) is **database-only**, which is the headline of spike
|
||||
§6 — the file half was never priced, and demo-hp is too young a box to price it.
|
||||
|
||||
2. **The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is.**
|
||||
A customer on the newest version is supported however old that version is. This is why
|
||||
`catalog_since` exists and why no version string is shown.
|
||||
|
||||
3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human,
|
||||
because §4 says it cannot be undone.
|
||||
|
||||
---
|
||||
|
||||
## 4. The vocabulary ruling — "rollback" is struck
|
||||
|
||||
**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually
|
||||
run, putting the old image tag back produces a container that refuses to start —
|
||||
|
||||
> *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and
|
||||
> downgrading is not supported"*
|
||||
|
||||
with a positive control proving the data is intact, only unreachable by the old version (§7 6d).
|
||||
|
||||
**So "rollback" must not appear in any spec for this arc.** The two shapes actually available are:
|
||||
|
||||
| shape | when it applies | what it does |
|
||||
|---|---|---|
|
||||
| **ABORT** | before anything migrated | stop, put the old image back, the app runs again |
|
||||
| **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy |
|
||||
|
||||
There is no third. And per §5 below, a restore's image-level undo currently has a ≤15-minute
|
||||
half-life because the syncer overwrites it (**R-441**).
|
||||
|
||||
---
|
||||
|
||||
## 5. The target shape
|
||||
|
||||
**The live `docker-compose.yml` becomes DERIVED from a pin recorded in `app.yaml`** — the one file the
|
||||
syncer never touches (`sync.go:319`'s exclusion). The catalog then proposes; `app.yaml` decides; the
|
||||
rendered compose file is an output rather than an input, and the thirteen unattended `up -d` paths
|
||||
stop being able to change a version by accident.
|
||||
|
||||
**Nothing in slices 1 or 2 implements this.** They make the current state *visible*, which is the
|
||||
prerequisite for judging how urgent it is.
|
||||
|
||||
---
|
||||
|
||||
## 6. The seven slices
|
||||
|
||||
| # | slice | status |
|
||||
|---|---|---|
|
||||
| **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** |
|
||||
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** |
|
||||
| **3** | **The compose file becomes DERIVED** — stop the syncer overwriting a deployed app's file; the pin in `app.yaml` wins. Needs the operator's ruling on R-438 first. | OPEN — R-447 |
|
||||
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 |
|
||||
| **5** | **An upgrade test** — prove a real one-major upgrade end to end, including the abort path. | OPEN — R-449 |
|
||||
| **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 |
|
||||
| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 |
|
||||
|
||||
**The rule slice 6 inherits, recorded now while it is cheap:** an engine change gets its own edge,
|
||||
never bundled with an app version bump. `bookstack` moved the application *and* MariaDB 11.6 → 12.3 in
|
||||
one commit (`0b73e5e`); that is two migrations behind one edge, and an unreadable failure when it
|
||||
breaks.
|
||||
|
||||
---
|
||||
|
||||
## 7. What slices 1 and 2 actually built
|
||||
|
||||
### 7.1 The record (slice 1)
|
||||
|
||||
`Manager.recordInstalledImages` (`felhom-controller/controller/internal/stacks/installed.go`) runs
|
||||
after a successful compose up from `StartStack`, `RestartStack`, `UpdateStack` and `runComposeDeploy`,
|
||||
and writes `app.yaml`:
|
||||
|
||||
```yaml
|
||||
installed_images:
|
||||
web:
|
||||
ref: lscr.io/linuxserver/bookstack:26.05.2
|
||||
digest: sha256:… # "" if the image was never pulled from a registry
|
||||
at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service
|
||||
```
|
||||
|
||||
Three rules, each with its reason:
|
||||
|
||||
- **It reads the CONTAINER, never `docker-compose.yml`.** §1.2 is why: that file is the value that has
|
||||
already moved. A record built from it would answer "what will happen next time something runs
|
||||
`up -d`", which is a different question.
|
||||
- **A failed write NEVER refuses the action** — deliberately the opposite of `SetDesiredState`.
|
||||
Intent refused, observation logged. Refusing to start a customer's app because we could not write
|
||||
down which version it is trades a real outage for a bookkeeping gap.
|
||||
- **It is NOT called from `StartStackServices`** — the R-47 DB-only restore window would overwrite a
|
||||
complete record with a partial one.
|
||||
|
||||
**Nothing reads it to take a decision.** Slice 2 reads it to render a label.
|
||||
|
||||
### 7.2 The label (slice 2)
|
||||
|
||||
`web.updateBadge` (`internal/web/updatebadge.go`) compares the recorded reference per service against
|
||||
what the current template pins, and renders through the existing `meta_badge` partial — no new markup,
|
||||
no new CSS.
|
||||
|
||||
**Absent means UNKNOWN and never means current.** Every `app.yaml` written before v0.233.0 has no
|
||||
record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were
|
||||
current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by
|
||||
a test with a companion red-proof.
|
||||
|
||||
**No version number reaches the customer** (operator ruling: a household cannot act on `26.05.2`).
|
||||
Version strings stay in the logs, the API and the hub.
|
||||
|
||||
---
|
||||
|
||||
## 8. Known limitations, stated plainly
|
||||
|
||||
1. **„Naprakész" can be FALSE for the 23 floating pins.** The comparison is reference-to-reference and
|
||||
queries no registry — a customer's box must not depend on reaching eight upstream registries to
|
||||
render a page. For `postgres:16-alpine`, `mariadb:11.6` and 21 others the reference can be
|
||||
identical while the image behind it has moved. **Measured, not theorised:** spike §5 found
|
||||
`mariadb:11.4` and `mariadb:12.3` had both already moved upstream, with two fully-pinned controls
|
||||
holding. Digest-level comparison needs a registry query and is deferred — **R-446**.
|
||||
2. **Nothing enforces `catalog_since`.** A commit that moves an `image:` line and forgets the date
|
||||
under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent
|
||||
commit to diff against, so the gate needs a deeper fetch — **R-452**.
|
||||
3. **The record only appears after the next lifecycle action.** An app that is running and untouched
|
||||
keeps a legacy `app.yaml` and therefore no badge, until someone restarts, updates or redeploys it.
|
||||
That is correct — the alternative is inventing a record from the file §1.2 says has already moved —
|
||||
but it means the fleet view fills in gradually rather than at upgrade.
|
||||
4. **The hub does not record image tags at all.** Its report's container payload carries name, state,
|
||||
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
|
||||
change; it is not derivable from what is already reported.
|
||||
|
||||
---
|
||||
|
||||
## 9. Where the rest lives
|
||||
|
||||
| what | where |
|
||||
|---|---|
|
||||
| the measurements this document rests on | `audits/SPIKE-app-update-2026-09-01.md` |
|
||||
| the work | `backlog/OPEN-ITEMS.md` — R-438..R-445, R-446..R-452 |
|
||||
| the syncer, described accurately but without the consequence | `architecture/02-controller-module-map.md` |
|
||||
| what the lifecycle actions are proven to do | `architecture/00-capability-map.md` |
|
||||
| the implementation | `felhom-controller/controller/README.md` §"What is installed, and is it current?" |
|
||||
@@ -675,14 +675,22 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** |
|
||||
| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. **STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner.** `docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic` reads: *"Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH."* So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses `rclone:`, measured with a control). **AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all.** That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. **Corroboration, unlooked for:** that page's table of port-23 commands matches, item for item, the `help` output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. | **OPEN — ask the vendor before building anything** |
|
||||
| **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** |
|
||||
| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. **MEASURED LIVE 2026-09-01 — CONFIRMED, and the mechanism is now attributed to an exact symbol.** A real catalog pin change (`bentopdf` v2.8.6 -> v2.8.5, commit `214d448`) travelled the real 15-minute cycle: at **17:45:17Z** `[INFO] [sync] Updated bentopdf/docker-compose.yml` rewrote the DEPLOYED app's file (mtime 17:45:17.646) while the container went on running v2.8.6 (started 17:36:35Z, unchanged). **Nothing told the customer** — no event, no notification, no email, and the customer's own pages carry NO version string at all (searched with ASCII fragments and BOTH controls; a first pass using an unescaped `.` over-counted and was corrected with `grep -F`). **THE CONSEQUENCE IS ALSO MEASURED:** with the file moved, `POST /api/stacks/bentopdf/restart` upgraded the container — **18.3 s and a network PULL** when the target image was absent, 0.5 s when present — and a **boot reconciliation upgraded it with NOBODY PRESSING ANYTHING** (`bootrecon.go:259` -> `StartStack` -> `compose up -d`). **THE DESIGN INTENT IS ALREADY IN THE SOURCE and it narrows this row:** `Manager.RestartStack` comments that `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*. So the RESTART half was chosen and written down; what is recorded nowhere is what the syncer then does to a deployed app, and whether the choice was meant to extend to the 13 UNATTENDED call sites. **AND ONE FEAR IS MEASURED SMALLER THAN FEARED:** a plain power cut does NOT upgrade — Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. The unattended upgrade needs the narrower precondition *"and the app did not come back"*. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** |
|
||||
| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. **MEASURED LIVE 2026-09-01 — CONFIRMED, and the mechanism is now attributed to an exact symbol.** A real catalog pin change (`bentopdf` v2.8.6 -> v2.8.5, commit `214d448`) travelled the real 15-minute cycle: at **17:45:17Z** `[INFO] [sync] Updated bentopdf/docker-compose.yml` rewrote the DEPLOYED app's file (mtime 17:45:17.646) while the container went on running v2.8.6 (started 17:36:35Z, unchanged). **Nothing told the customer** — no event, no notification, no email, and the customer's own pages carry NO version string at all (searched with ASCII fragments and BOTH controls; a first pass using an unescaped `.` over-counted and was corrected with `grep -F`). **THE CONSEQUENCE IS ALSO MEASURED:** with the file moved, `POST /api/stacks/bentopdf/restart` upgraded the container — **18.3 s and a network PULL** when the target image was absent, 0.5 s when present — and a **boot reconciliation upgraded it with NOBODY PRESSING ANYTHING** (`bootrecon.go:259` -> `StartStack` -> `compose up -d`). **THE DESIGN INTENT IS ALREADY IN THE SOURCE and it narrows this row:** `Manager.RestartStack` comments that `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*. So the RESTART half was chosen and written down; what is recorded nowhere is what the syncer then does to a deployed app, and whether the choice was meant to extend to the 13 UNATTENDED call sites. **AND ONE FEAR IS MEASURED SMALLER THAN FEARED:** a plain power cut does NOT upgrade — Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. The unattended upgrade needs the narrower precondition *"and the app did not come back"*. `audits/SPIKE-app-update-2026-09-01.md` **THE DOCUMENT HALF IS NOW DISCHARGED, 2026-09-02: `documentation/architecture/09-update-architecture.md` exists and is a LIVING document, updated by every slice of this arc.** It records the mechanism as measured (§1), quotes the `RestartStack` comment that proves the restart half was CHOSEN (§2), carries the three operator rulings of 2026-09-02 (§3), strikes the word "rollback" (§4), states the target shape (§5) and lists the seven slices with a status each (§6). **THE ROW STAYS OPEN AND THE REASON IS THE POINT: the mechanism is now DOCUMENTED, not CHANGED.** Whether the syncer should go on overwriting a deployed app's compose file is still the operator's ruling, and acting on it is slice 3 (R-447). | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** |
|
||||
| **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). **RE-READ AND CONFIRMED 2026-09-01; the severity argument SURVIVES but its stated reason was imprecise and is corrected here.** The task's reason was *"the UI only renders Frissites when the app is operational, and a held app is stopped"*. Half right: `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded` as operational too**, and this was OBSERVED live — the green `Frissites` button rendered over a crash-looping app during the spike's Phase 3b. **So the button is hidden specifically because a held app is `StateStopped`, not because broken apps hide it.** LOW stands; the reason must be stated precisely or the next reader will widen it. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-441** | **[P2-MEDIUM] The restore path and the catalog sync disagree about which image the app should run, and the SYNC WINS within 15 minutes.** `stackAdapter.RecreateStackDefinitionFromUnit` (`felhom-controller/controller/cmd/controller/main.go:2570`) writes the recovery unit's CAPTURED `docker-compose.yml` — carrying the OLD image pin — straight into the live stack dir, and `restore_unit.go:317` states the intent: *"Resolved from the UNIT's compose, because that file is about to BECOME the live one."* But `Syncer.copyIfChanged` overwrites any stack file whose content differs from the catalog, on the next 15-minute tick, with no deployed check (R-438). **So a restore's image-level rollback has a <=15-minute half-life, and the next `compose up -d` from any of the 13 unattended call sites re-applies the catalog pin.** **GRADED HONESTLY — the two halves have different evidence:** the overwrite is **MEASURED** (a locally-modified compose on demo-hp was overwritten by the sync at 18:10:29Z, `[INFO] [sync] Updated bentopdf/docker-compose.yml`); that the restore writes to that same path is **READ, not measured**. Settling it needs one live restore with a stale pin, which is a phase, not a check. **Why it matters more than it reads:** R-361's undo copy plus this is the only route back that exists, and Phase 6 proved putting the old TAG back is not a rollback at all (R-443's sibling finding) — so the data restore is the whole remedy, and it is fighting the syncer. Owner: **CC to measure, Viktor to rule on which wins.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC measures, VIKTOR rules** |
|
||||
| **R-442** | **[P1-HIGH] `remove_hdd_data: true` is INERT on a box whose `controller.yaml` has no `paths.hdd_path` — the customer's data stays on the drive and the API reports NEITHER removed NOR preserved.** MEASURED on demo-hp 2026-09-01: removing an app with `{"remove_hdd_data":true,"remove_backups":true}` returned **HTTP 200** with `"hdd_paths_removed":null,"hdd_paths_preserved":null` and left **128 MB** at `/mnt/felhom-drives/hdd_1/appdata/nextcloud`. **ROOT CAUSE, with controls:** `Paths.HDDPath` (`internal/config/config.go:117`) has **NO default** — only an env override at `:403` — and demo-hp's `controller.yaml` `paths:` block holds only `data_dir`, `stacks_dir`, `system_data_path`; the container has **no `FELHOM_PATHS_*` variable at all** (measured, count 0). So `cfg.Paths.HDDPath == ""` and `ParseComposeHDDMounts` (`internal/stacks/delete.go:600-603`) returns `nil` on its FIRST line — logging `found 0 HDD mounts` — for a compose that plainly contains `- ${HDD_PATH}/appdata/nextcloud:/var/www/html/data`. **The second half of the same removal ALSO no-op'd:** `[WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps`. **Why P1:** a customer who removes an app and asks for the data to be deleted is told it worked, and it was not. This is a privacy answer, not a tidiness one. **NOT ESTABLISHED: whether the fleet shares this config shape** — demo-felhom and any customer box must be checked before sizing it. **The fix needs a test that FAILS when `hdd_path` is empty**, or the guard comes back. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: CC** |
|
||||
| **R-443** | **[P2-MEDIUM] The Update button reports SUCCESS over an app it has just broken, and the truth arrives 5m16s later by a different road.** MEASURED on demo-hp 2026-09-01: `POST /api/stacks/bentopdf/update` against an image that pulls cleanly and then fails to run returned **HTTP 200 `{"ok":true,"message":"Stack bentopdf update completed"}`** and logged `Stack bentopdf updated successfully (took 3.5s)`, while the container went to `status=restarting RestartCount=9`. **The controller's own post-start line told the truth (`manager.go:1403 ... alpine:3.20 restarting`) — but it runs AFTER the API has already answered.** This is this repo's own `up -d` exits 0 on a crash-loop invariant surfacing at the customer's most consequential button. **What the customer's page then said:** badge **`Ujraindites...`**, `Restarting (0) 15 seconds ago`, and the full green button row — because `isOperationalState` counts `StateRestarting` as operational (see R-439). *"Restarting"* reads as transient, not as failure, and nothing says the update caused it. **THE HONEST OTHER HALF, and it must travel with this row: the customer IS told.** `app_start_failed` fired at 18:05:59Z with severity `warning` (inside the hub's exact vocabulary, so it really delivers) — 5m16s after the update, from `crashLoopAfter = 5 * time.Minute`, a threshold whose own comment argues it well. **So this is NOT the silent-dead-app class; it is a TRUTHFULNESS-AT-THE-MOMENT-OF-ACTION problem.** Also recorded: on a pull FAILURE the product behaves correctly — HTTP 500, and `compose up -d` resolves images before touching a container, so the running app survives (measured twice). Owner: **CC to propose, VIKTOR to rule on whether Update should wait and verify.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules** |
|
||||
| **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** |
|
||||
| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-447** | **[P1-HIGH] UPDATE ARC SLICE 3 — make the live compose file DERIVED, so the syncer stops changing a deployed app's version.** The target shape (`architecture/09-update-architecture.md` §5): the pin lives in `app.yaml`, the one file `Syncer.copyTemplates` never touches, and the live `docker-compose.yml` becomes an OUTPUT rather than an input. The catalog then proposes and `app.yaml` decides, and the thirteen unattended `compose up -d` call sites (spike §8) stop being able to change a version by accident. **BLOCKED ON AN OPERATOR RULING, and that is the whole reason this is a separate slice:** R-438 established that the restart half of this behaviour was CHOSEN and written down in `Manager.RestartStack`'s own comment. Changing it is not a bug fix; it is reversing a decision, and the decision-maker is Viktor. **Slices 1 and 2 shipped first ON PURPOSE** — the fleet's real state has to be visible before anyone can judge how urgent this is. Do NOT add a deployed check to any of the thirteen paths ahead of the ruling. `architecture/09-update-architecture.md` §5 | **BLOCKED — rank P1-HIGH; owner: VIKTOR rules, CC implements** |
|
||||
| **R-448** | **[P2-MEDIUM] UPDATE ARC SLICE 4 — a guarded update: a verified backup as a precondition, an abort path, and the truth at the moment of action.** Three parts, each already evidenced. (a) **The precondition is a VERIFIED RECENT BACKUP, not a new copy** (operator ruling 2026-09-02); the guest-snapshot alternative must be SPIKED before anything is designed around it. Today's safety machinery is DATABASE-ONLY (`Manager.writeSafetyDump`, `internal/backup/offbox_reconstitute.go:207`) and the file half was never priced (spike §6). (b) **The abort path, never a "rollback"** — spike §7 proved the word is wrong: once a migration has run, the old image refuses to start on the migrated data. The two available shapes are ABORT (before anything migrated) and RESTORE FROM A COPY (after). (c) **Truth at the moment of action** — this subsumes **R-443**: the Update button returns HTTP 200 over an app it has just broken and the alarm arrives 5m16s later. `architecture/09-update-architecture.md` §3, §4 | **READY — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules on (c)** |
|
||||
| **R-449** | **[P2-MEDIUM] UPDATE ARC SLICE 5 — an upgrade test that proves a real one-major upgrade end to end, INCLUDING the abort path.** Spike §7 performed the pieces by hand on Nextcloud (31.0.14 → 32.0.9 ran the migration; 31.0.14 → 34.0.1 was refused by the app and reported as SUCCESS by the product; the downgrade attempt was refused with the data intact). **None of it is a test that runs again.** A one-off measurement that nothing repeats decays into a claim — this project's most-repeated defect class. The test must assert the CONSEQUENCE (does the app serve after the upgrade? does the abort put it back?), not the mechanism. Venue: a Tier-0 box; `runbooks/target-selection.md` names which. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-452** | **[P3-LOW] Nothing enforces `catalog_since`, so the one number the update badge shows can silently under-report.** `app-catalog-felhom.eu` `CLAUDE.md` now states the rule — any commit that changes an `image:` line must set that app's `catalog_since` to the same day — and all 53 apps were backfilled from git history on 2026-09-02 (`69761cf`). **A rule with no instrument is a wish; that is this project's most-repeated finding and this row exists so it is not repeated silently.** A stale `catalog_since` makes „Frissítés elérhető — N napja" under-report N, which is the single number the badge exists to give. **WHY IT WAS NOT BUILT IN THE SAME SESSION, stated rather than implied:** the gate would have to diff an `image:` line against the PARENT commit, and `catalog_gates.py` runs under a runner that fetches at `--depth 1` — there is no parent to diff against. The gate therefore needs a deeper fetch, which is a change to the CI shape and not to a script. **This is the R-421 class in advance: an enumerated gap becomes a row in the same session it is enumerated.** `architecture/09-update-architecture.md` §8.2 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-453** | **[P2-MEDIUM] The vaulted customer dashboard password is STALE ON BOTH DEMO BOXES, and there is no operator-side route to the real one — so no session can drive a customer page.** MEASURED 2026-09-02 while trying to live-validate the update badge: `PASSWORD` from DooPlex `~/.config/credentials` returns **HTTP 200 with the login page and the body string `Hibás jelszó`** against demo-hp guest 9201 (`https://192.168.0.138:443`, `Host: felhom.enkisfelhom.hu`) **and** against demo-felhom guest 9201 (`https://192.168.0.149:443`, `Host: felhom.demo-felhom.eu`). **The discriminator is the controller's own log, not the status code** — `auth.go:176: [WARN] [web] Failed login` proves wrong PASSWORD rather than wrong Host header, which is the trap this class always presents (a rejected login renders no flash and looks exactly like a routing problem). Every other key in the credentials file was checked and none is a dashboard password (`HUB_PW`, `TS_KEY`, `HETZNER_API`, `ISO_S3_*`, and the `R_*` keys are escrow recovery codes). **THIS IS THE SECOND TIME:** it drifted on demo-hp on 2026-08-09 and was put back on operator instruction by writing a fresh bcrypt hash into the guest's `data/settings.json`; demo-hp was then reinstalled and re-claimed on 2026-08-21, and demo-felhom has now drifted as well. **Why it is not merely inconvenient: it silently converts "endpoint-level validation" — this project's STANDARD method, because there is no browser on DooPlex — into "unit tests only" for anything that renders a customer page.** The cost is paid per session and rediscovered each time. **The fix is a decision, not a command:** the claim code is bcrypt-hashed hub-side and only e-mailed (R-119), so either the operator records the current demo passwords out-of-band, or a re-set becomes a standing authorisation for the two Tier-0 demo boxes. **NOT TAKEN UNILATERALLY:** re-setting a dashboard password is a decision about a customer account, and it was done under operator instruction last time. Raised in `STATUS.md` item 9. Evidence: `tests/VALIDATION-update-slice12-2026-09-02.md` §4, which lists all five attempts. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR decides, CC executes** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
@@ -52,6 +52,7 @@
|
||||
|
||||
| ID | Item | Size | Status | Notes |
|
||||
|----|------|------|--------|-------|
|
||||
| **UPDATE-ARC** | **The app-update arc — seven slices, from "nobody knows what any box runs" to "an update is a decision the box can take safely."** Opened after `audits/SPIKE-app-update-2026-09-01.md` measured what an update actually does. | L | **slices 1 & 2 SHIPPED 2026-09-02 (controller v0.233.0 + catalog `69761cf`); slices 3–7 open** | **The reasoning has a home and it is the point of the exercise: `architecture/09-update-architecture.md`** — a LIVING document, updated by every slice in the same session, created because its ABSENCE was a finding (R-438: the mechanism was chosen deliberately and written down nowhere). **Capability-map rows this flips:** the App-lifecycle row (`00-capability-map.md`, "start/stop/restart/update/logs/remove/redeploy") and its 2026-09-01 sibling ("what restart and update do to a deployed app whose compose file the catalog already moved") — both recorded that the ACTIONS work and said nothing about VERSIONS. Slices 1 and 2 add the version half. **The findings live in the register, per the ONE REGISTER ruling:** R-438, R-440, R-441, R-443 (existing) and R-446..R-452 (this session) — one row per remaining slice, each with a rank and an owner, so the arc is visible in the register and not only in a task file. **Slice 3 (R-447) is BLOCKED ON AN OPERATOR RULING and that is deliberate:** changing what the syncer does to a deployed app reverses a decision, not a bug, and slices 1 and 2 shipped first so the fleet's real state is visible before anyone judges how urgent it is. |
|
||||
| R-6 | **Spike: LAN service discovery from the guest** — SSDP multicast (UDP 1900, DLNA), WSD (Windows discovery), mDNS; host-network vs macvlan; is the customer LXC LAN-bridged in appliance deployments? | M | **spiked (2026-07-18)** | **VERDICT: appliance guest IS LAN-bridged (own DHCP lease on the household /24); multicast discovery works ONLY in the guest netns — guest-direct or Docker `--network host` (SSDP/mDNS/WSD all PASS both ways); the default docker bridge is categorically DEAF to LAN multicast (WSD/mDNS RX FAIL, unicast-publish PASS). Real samba+wsdd on host-net → Windows 11 ProbeMatch + FELHOM-SPIKE renders in Explorer + 445 + authenticated SMB round-trip all PASS; real SSDP `MediaServer:1` advert reaches both LAN clients. → R-7 SMB stack MUST be host-network LAN-bound; R-8 Jellyfin-DLNA plausible if host-network. Caveat: `vmbr0 multicast_snooping=1` worked only because the household router is a live querier — customer LANs w/ snooping+no-querier, and Peti's BYO bridge, are UNTESTED gaps.** **S4b (human leg, the sharpest finding): wsdd makes the box VISIBLE but the Explorer double-click FAILS `0x80070035` — WSD gives no name resolution; the flat `\\FELHOM-SPIKE` resolved by no path. Adding `nmbd` (NetBIOS) fixed it live (flat name resolves + mounts). → R-7 needs smbd+wsdd+nmbd (+avahi/.local for modern clients), not wsdd alone.** Doc: `audits/SPIKE-lan-discovery-2026-07-18.md`. |
|
||||
| R-7b | **Share backup EXECUTION** — put share data into the live tier-2 + offsite runs (the design fork reported by R-7 slice 1) | M | **SHIPPED (controller v0.145.0, 2026-07-18)** | **Viktor's ruling: Model B′ — a SIBLING shares source.** New, additive job/leg code reusing the proven primitives (tier-2 mirror seam, restic wrappers, soft-quota/enlargement gate, status recorders) while leaving **every per-app engine path byte-identical** — NOT a synthetic recovery unit (breaks on multi-drive shares, wraps 1 KB of JSON in dump machinery) and NOT engine-loop surgery. The B′ invariant is enforced by test in both tiers, red-proofed. Tier 2 → `RunSharesTier2` (legs grouped by SOURCE drive → `backups/secondary/_shares/<driveKey>/<share>`, payload at `_payload/`, layout marker LAST). Tier 3 → `runOffboxSharesLeg`: ONE extra `restic backup --tag felhom-offbox --tag _shares` placed after the app loop and BEFORE retention, so `forget --group-by host,tags` covers the new group with no flag change; a quota-blocked push degrades to the **manifest only, never to nothing**. Restore → „Megosztások" on `/backups/restore`: scratch, then a missing-only merge whose every destination is PREFIX-ASSERTED against live storage roots, definitions merged existing-wins, then `ReconcileSamba`, then the credential. The **payload** (`_shares-manifest.json` + a best-effort secret-bearing `passdb.tar`) is what makes DR return files + configuration + password rather than loose bytes. **Fold-in: samba joins the liveness set** — `EffectiveProtected` adds the CONTAINER `felhom-samba` exactly while sharing is on. **FULLY PROVEN-LIVE on demo (2026-07-18), all four legs.** (1) tier-2: real `/api/backup/tier2` trigger → `_shares` tree + marker + payload on the cross-drive target, mirrored file md5-identical, payload 0600 preserved. (2) offsite: Viktor's manual run 12:18:16Z → snapshot **`e0b9d723`** (tags `felhom-offbox,_shares`) with the payload dir + both share folders; a second run via the „Távoli mentés" button → **`4e2b15ec`**, containing `_shares-manifest.json` (418 B) AND `passdb.tar` (855 040 B), both 0600, share files with uid 1000 preserved. (3) restore round-trip: probe file + the `dokumentumok` DEFINITION deleted via the real endpoints, then „Megosztások" restore + place → `1 file(s), 1 definition(s) re-added, 1 kept, 0 refused, credential=true`; probe back md5-identical, the two pre-existing files NOT overwritten (missing-only proven on live data), definition back with its ORIGINAL flags and created_at, `smb.conf` re-rendered, `filmek` untouched. (4) liveness: samba stopped → `health_critical` pushed and hub-accepted (200) → self-healed. Remaining human leg: SMB positive auth with the real household password (never persisted by design). **Correction:** an earlier revision of this row and of the ship REPORT wrongly claimed the demo box had no offsite target — the verification read a guessed settings key (`offbox_target`) instead of the real one (`offbox`); root cause dissected in REPORT §7b. Findings: the reserved-name assumption was FALSE (`nbNameRe` accepted „_shares" as a share name — now refused); the alert/e-mail pipeline needed NO change and adds no new event type. Docs: `controller/sharing.md`; ship report `felhom-controller/REPORT.md`. |
|
||||
| R-8 | DLNA (**gate input now exists — R-6 spiked 2026-07-18: SSDP reaches LAN clients from host-net**): validate Jellyfin's built-in DLNA server first; only add minidlna to the catalog if Jellyfin-DLNA fails | S | idea (unblocked) | Don't add catalog weight before proving the cheap path. **R-6 confirmed the cheap path is physically viable — Jellyfin DLNA must run host-network (same multicast constraint as R-7)** |
|
||||
|
||||
@@ -0,0 +1,170 @@
|
||||
# VALIDATION — update arc slices 1 & 2, live on demo-hp (2026-09-02)
|
||||
|
||||
**Controller `0.233.0`, guest 9201 on `demo-hp` (Tier 0 — disposable).** Evidence copied off the box
|
||||
at the end of the phase that produced it, not at the end of the session.
|
||||
|
||||
**Method: endpoint/production-path level.** `claude-in-chrome` is not available on DooPlex.
|
||||
**Which path was used, stated exactly: the BOOT RECONCILER**, `bootrecon.Run → StackProvider.StartStack
|
||||
→ compose up -d → recordInstalledImages` — a REAL production caller, the same one
|
||||
`SPIKE-app-update-2026-09-01` §2 variant 1c-ii used, and no hand-set state anywhere.
|
||||
**Why not the customer's Restart button: see §4 — the vaulted dashboard password no longer opens
|
||||
either demo controller.**
|
||||
|
||||
---
|
||||
|
||||
## 1. Deployed version
|
||||
|
||||
```
|
||||
$ pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.233.0 Up 19 seconds (healthy)
|
||||
```
|
||||
|
||||
Previous: `0.232.0`. Image digest `sha256:df5940ccf5a548ceee9065a2cc55a467f8941c636138441e0d77e320dc2023ab`.
|
||||
|
||||
## 2. The record — PROVEN LIVE, on a single-service AND a multi-service app
|
||||
|
||||
### 2.0 Before — nine deployed apps, ZERO with a record
|
||||
|
||||
```
|
||||
bentopdf deployed=1 installed_images=0 paperless-ngx deployed=1 installed_images=0
|
||||
bookstack deployed=1 installed_images=0 privatebin deployed=1 installed_images=0
|
||||
calibre-web deployed=1 installed_images=0 romm deployed=1 installed_images=0
|
||||
docmost deployed=1 installed_images=0
|
||||
kimai deployed=1 installed_images=0
|
||||
opengist deployed=1 installed_images=0
|
||||
```
|
||||
|
||||
**That is the legacy case, and it is the state the badge must render NOTHING for.**
|
||||
|
||||
### 2.0b Ground truth, read from the containers BEFORE anything was touched
|
||||
|
||||
```
|
||||
bentopdf ref=ghcr.io/alam00000/bentopdf:v2.8.6 rd=…@sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
|
||||
bookstack ref=lscr.io/linuxserver/bookstack:26.05.2 rd=…@sha256:3db259db582808ab498d49ae96b0a63f935d9cf3635c9d5bd8b8815c6ff1f8a1
|
||||
bookstack-db ref=mariadb:12.3 rd=…@sha256:a02fe89cb597d4375812b2eac90cf9d0775d4686daa7f7cc750ebbcad7525bbc
|
||||
```
|
||||
|
||||
**This is the discriminator.** Every digest below is compared against these, taken independently.
|
||||
|
||||
### 2.1 Single service — `bentopdf`
|
||||
|
||||
Staged exactly as the spike stages a boot orphan: `docker rm -f bentopdf` (chosen because it has **no
|
||||
database, no volume and no data of any kind**), then the controller restarted so the reconciler runs.
|
||||
`desired_state: running` was left untouched. **Nothing else was staged; the reconciler selected the app
|
||||
on its own.**
|
||||
|
||||
```
|
||||
18:27:32 main.go:2162: [INFO] [bootrecon] boot window: fleet settled after 30s — sweeping
|
||||
18:27:32 bootrecon.go:259: [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]
|
||||
18:27:32 manager.go:1066: [INFO] [stacks] Starting stack: bentopdf
|
||||
18:27:32 manager.go:1081: [INFO] [stacks] Stack bentopdf started successfully (took 0.3s)
|
||||
18:27:32 installed.go:408: [INFO] [stacks] installed-images bentopdf: recorded 1 service(s) (bentopdf=ghcr.io/alam00000/bentopdf:v2.8.6 (sha256:eaeea1e44720…))
|
||||
```
|
||||
|
||||
`/opt/docker/stacks/bentopdf/app.yaml` after:
|
||||
|
||||
```yaml
|
||||
installed_images:
|
||||
bentopdf:
|
||||
ref: ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
digest: sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
|
||||
at: "2026-09-02T18:27:32Z"
|
||||
```
|
||||
|
||||
**Digest MATCHES §2.0b exactly.**
|
||||
|
||||
### 2.2 Multi-service — `bookstack`. **The case that matters.**
|
||||
|
||||
**One entry for the whole stack is the wrong outcome, and only a multi-container app can tell.**
|
||||
|
||||
Same staging (`docker rm -f bookstack bookstack-db`; both volumes are NAMED —
|
||||
`bookstack_bookstack_config`, `bookstack_bookstack_db_data` — and were verified present, so no data
|
||||
was at risk), then the controller restarted.
|
||||
|
||||
```
|
||||
18:28:48 bootrecon.go:259: [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [bookstack]
|
||||
18:28:48 manager.go:1066: [INFO] [stacks] Starting stack: bookstack
|
||||
18:28:54 installed.go:408: [INFO] [stacks] installed-images bookstack: recorded 2 service(s) (bookstack=lscr.io/linuxserver/bookstack:26.05.2 (sha256:3db259db5828…), bookstack-db=mariadb:12.3 (sha256:a02fe89cb597…))
|
||||
18:28:54 bootrecon.go:303: [INFO] [bootrecon] Boot reconciliation complete: 1 app(s) recovered in 1 attempt(s)
|
||||
```
|
||||
|
||||
```yaml
|
||||
installed_images:
|
||||
bookstack:
|
||||
ref: lscr.io/linuxserver/bookstack:26.05.2
|
||||
digest: sha256:3db259db582808ab498d49ae96b0a63f935d9cf3635c9d5bd8b8815c6ff1f8a1
|
||||
at: "2026-09-02T18:28:54Z"
|
||||
bookstack-db:
|
||||
ref: mariadb:12.3
|
||||
digest: sha256:a02fe89cb597d4375812b2eac90cf9d0775d4686daa7f7cc750ebbcad7525bbc
|
||||
at: "2026-09-02T18:28:54Z"
|
||||
```
|
||||
|
||||
**TWO entries, keyed by COMPOSE SERVICE NAME, both digests matching §2.0b exactly.**
|
||||
|
||||
**A NOTE ON WHY IT TOOK TWO PASSES, because it is a real fact about the reconciler and not a mishap.**
|
||||
The first pass removed only the `bookstack` app container and left `bookstack-db` running; the
|
||||
reconciler did **not** select bookstack — it found `[bentopdf]` alone. A stack with one live member is
|
||||
not a boot orphan to it. Removing the DB container as well made the whole stack orphaned and it was
|
||||
selected on the next pass. Both apps are running and healthy at the end (§5).
|
||||
|
||||
### 2.3 The encrypted secrets survived the write — checked, not assumed
|
||||
|
||||
`bookstack`'s `app.yaml` holds two encrypted values. Before and after the record was written, both
|
||||
`ENC:` strings are **byte-identical**, and `deployed`, `deployed_at`, `locked_fields` and
|
||||
`desired_state` are unchanged. This is the copy-and-overlay property (R-100) holding under a new field.
|
||||
|
||||
## 3. `catalog_since` reached the box by the NORMAL 15-minute sync
|
||||
|
||||
No force, no hand-edit:
|
||||
|
||||
```
|
||||
$ grep -n catalog_since /opt/docker/stacks/bookstack/.felhom.yml
|
||||
13:catalog_since: "2026-07-18"
|
||||
```
|
||||
|
||||
## 4. The badge render — **NOT LIVE-VALIDATED. What was tried, in full.**
|
||||
|
||||
**A "no access" claim must list its attempts.** These are the attempts:
|
||||
|
||||
| # | attempt | result |
|
||||
|---|---|---|
|
||||
| 1 | `POST /login` to demo-hp guest `https://192.168.0.138:443`, `Host: felhom.enkisfelhom.hu`, `-k`, password from DooPlex `~/.config/credentials` `PASSWORD` (extracted with `sed`, never `cut` — the values are quoted) | **HTTP 200 with the login page and the body string `Hibás jelszó`** |
|
||||
| 2 | the controller's own log, as the discriminator between "wrong host header" and "wrong password" | `auth.go:176: [WARN] [web] Failed login from 172.18.0.3` — **wrong password, not a routing problem** |
|
||||
| 3 | the same password against the OTHER demo box, demo-felhom guest `https://192.168.0.149:443`, `Host: felhom.demo-felhom.eu` | **HTTP 200 + `Hibás jelszó` as well** |
|
||||
| 4 | every other key in `~/.config/credentials` — `HUB_PW`, `R_DEMO-HP`, `R_DEMO-FELHOM`, `R_C11_REWALK`, `R_PART4`, `TS_KEY`, `HETZNER_API`, `ISO_S3_*` | none is a dashboard password; the `R_*` keys are escrow recovery codes |
|
||||
| 5 | reading the source for an unauthenticated route to an app page | only `/claim`, `/claim/request-new-code`, `/api/health` and `/static/` are exempt (`internal/web/auth.go`) |
|
||||
|
||||
**So the vaulted `PASSWORD` is stale on BOTH demo controllers.** It was already recorded as having
|
||||
drifted once (memory `demo-hp-guest-controller-access`, 2026-08-09, put back on operator instruction);
|
||||
demo-hp was reinstalled and re-claimed on 2026-08-21, and demo-felhom has now drifted too.
|
||||
|
||||
**There is no operator-side route to a customer's dashboard password** — the claim code is
|
||||
bcrypt-hashed hub-side and only emailed (R-119). Re-setting it means writing a new bcrypt hash into the
|
||||
guest's `data/settings.json`, which is a decision about a customer account and was done under operator
|
||||
instruction last time. **It is therefore a HUMAN step, by this project's own rule**, and it is not
|
||||
taken here.
|
||||
|
||||
**What IS established about the badge, and it stops short of the render:** every INPUT the badge reads
|
||||
is verified live and consistent on this box — `installed_images` present with refs matching the
|
||||
template's pins exactly (§2.2 vs the compose file's `lscr.io/linuxserver/bookstack:26.05.2` and
|
||||
`mariadb:12.3`), and `catalog_since: "2026-07-18"` present (§3). So bookstack's inputs are the
|
||||
`Naprakész` case and the seven undisturbed apps are the no-record case. **That is an inference from
|
||||
verified inputs, not an observation of the rendered page, and it is not counted as evidence.**
|
||||
|
||||
The render itself is covered by unit tests that render the PRODUCTION templates
|
||||
(`TestGroupD_BadgeRendersOnBothSurfaces`, both surfaces, all states, with a companion red-proof), which
|
||||
is the strongest statement available without the password.
|
||||
|
||||
## 5. End state — nothing left broken, nothing provisioned
|
||||
|
||||
```
|
||||
$ pct exec 9201 -- docker ps --format '{{.Names}}|{{.Image}}|{{.Status}}' | grep -E 'bookstack|bentopdf'
|
||||
bookstack|lscr.io/linuxserver/bookstack:26.05.2|Up 10 seconds (health: starting)
|
||||
bookstack-db|mariadb:12.3|Up 16 seconds (healthy)
|
||||
bentopdf|ghcr.io/alam00000/bentopdf:v2.8.6|Up About a minute (healthy)
|
||||
```
|
||||
|
||||
All three containers are back on the SAME images they ran before, both named volumes untouched, and
|
||||
both apps carry a complete record. **This run provisioned nothing** — no guest, no storage, no hub
|
||||
record — so there is no teardown to report on any of the three layers.
|
||||
Reference in New Issue
Block a user