night 2026-09-13/14: first "be a customer" rotation (adventurelog) — 7 defects found, 13 rows closed
gates / gates (push) Successful in 18s
gates / gates (push) Successful in 18s
New runbooks/nightly-rotation.md; observations_gate.py reads every section (R-471); target-selection.md names real paths (R-461); R-93 carries the fact that drill-r50 is gone. Register: R-473/R-474/R-466/R-471/R-453/R-461 and v0.240.0's R-477/R-478/R-480/R-482/R-484/R-485/R-486 closed; R-481, R-483, R-487, R-488, R-489 opened. 09 §6.1, 07 §6, CONTEXT, STATUS note. Evidence: audits/nightly-2026-09-13-adventurelog/, audits/v0240-2026-09-13/.
This commit is contained in:
@@ -264,3 +264,16 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-470** | **Four controller CHANGELOG headers (v0.233.0–v0.236.0) carried no `MinAgent:` line, while the vouch and now the declared floor read it from the header.** Closed 2026-09-13 (`felhom-controller` `f946b0d`): the four headers backfilled with `**MinAgent: 0.129.0** (unchanged)` — v0.232.0's value, proven unchanged (no commit under `internal/agentapi` since 2026-09-01; highest `featureMinAgent` 0.129.0) — and `controller/scripts/minagent_header_gate.py` (fast, blocking) refuses a newest header without the line; a prose or code-span mention does not count (decoy; red-proof F). | **CLOSED 2026-09-13 — GATED** | full text: `git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-472** | **The golden cadence ruling and the hub's floor rule contradicted each other: a floor above the vouched golden delivered nothing.** Operator ruling 2026-09-13, hub **v0.112.0** (`f181efd`): a floor saved with the release's declared MinAgent is served above the golden under the same agent comparison; an undeclared one is still held and both forms refuse it (`floor_needs_min_agent`). Proven live: controller 0.239.0 reached demo-hp in 14 s and demo-felhom in 15 s from the save, hub `managed floor SERVED … from declared`. Evidence: `audits/rulings-r472-r475-2026-09-13/` 02, 03. **Reasoning kept:** *the manifest leads the floor inside the golden; above it, the release's own declared MinAgent does* (publish-train rule 1); a declaration binds to its exact floor. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-475** | **The update precondition was Tier-2-only, so an app with no second-drive copy could not be updated.** Operator ruling 2026-09-13, controller **v0.239.0** (`b93c154`): the first fresh copy in the order Tier 2, Tier 1, Tier 3 (bounded, unreachable = absent + WARN); `backup_max_age` applies to the chosen tier; nothing anywhere → back up first; refused only when no copy and no backup can be taken; the hold names the tier; a Tier-2 failure in the pre-backup is a WARN. Proven live on demo-hp: nothing anywhere → backed up first (04), Tier 1 alone (05), held naming „saját meghajtó” (07), restored from „helyi” (08). Red-proof M (age only on Tier 2) fails. **Reasoning kept:** *first FRESH copy, not first copy* — a stale mirror must not force a backup while the own unit is minutes old. Follow-ups: R-477..R-480. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-473** | **The `glance` catalog template crash-looped on every fresh install — the image ships no default config.** Closed 2026-09-13 (catalog `50ad286`): an `entrypoint` wrapper seeds a small Hungarian start page on first boot only when `/app/config/glance.yml` is absent, then execs the image's own command (read from the image, not guessed); the seed validated with the image's `config:validate`. Proven live on demo-hp the same evening: a fresh throwaway install healthy in 21 s, restart count 0, front door 200, seed present in the volume (`audits/nightly-2026-09-13-adventurelog/08-R473-glance-fresh-install.txt`). No version moved. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-471** | **`observations_gate.py` read only the FIRST observations section, so an appended second section with an unmarked item passed — and the R-419 decoy had been reading LIVE HOLE at HEAD.** Closed 2026-09-13 (felhom.eu `scripts/observations_gate.py`): `observation_sections` collects every observations heading and `observation_items` pools their items; the heading shown is all of them joined. Cause established: first-heading-wins (the parser `break`s on the first match). Red-proof: the old parser exits 0 on the decoy shape, the new one 1 (`audits/v0240-2026-09-13/rp-R471.txt`); all 12 felhom.eu decoys behave; the felhom.eu and controller REPORTs still pass through the shared script. The "run by hand" half stands: the decoy suite is still not in any runner (R-426's exemptions) — that is a separate row if wanted. | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-486** | **Removing an app with its backups KEPT forgot its Tier-2 record, so the second-drive restore was refused over an intact mirror.** Closed in controller **v0.240.0** (`bdcbd50`): the record goes only with `remove_backups`. Proven live on demo-hp: remove keeping backups → `cross_drive` record kept → „Teljes visszaállítás" restored 2 volumes and the database, data identical. Red-proof: an unconditional `SetCrossDriveConfig(nil)` fails the wiring test. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-484** | **PostGIS was not a database.** Closed in v0.240.0: `dbTypeForImage` matches `postgis`, `pgvector`, `timescaledb` as Postgres. Proven live: adventurelog's unit now carries `db-dumps/adventurelog-postgres.sql` (8 MB) and the unit restore reports „2 adatkötet és az adatbázis visszaállítva". Red-proof: dropping `postgis` fails the table test. Immich's own image already matched. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-485** | **The backup card read two dead paths and said `has_backups:false` over 484 MB.** Closed in v0.240.0: it sizes the recovery unit and the app's Tier-2 mirror(s). Proven live: `244M` + `244M`, `has_backups:true`. Red-proof: the old paths fail `TestR485`. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-480** | **The card kept a held update's „leállítva marad" sentence after a successful restore and after removal.** Closed in v0.240.0: the stack remembers its last update ended held; `fillHoldReason` hides that outcome once the hold is gone or the app is not deployed; a pull failure keeps its sentence. Red-proof: disabling the block fails three assertions. Live: the card after a restore reads clean. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-477** | **The update's Tier-3 lookup paid the full 15 s bound and blamed another app.** Closed in v0.240.0: `OffsiteSnapshotTimes` — one `snapshots --json`, no per-app `stats`; the inventory page and the update share `offsiteNewestPerTag` (registered read-only for R-408). Red-proof: routing the lookup through the inventory fails `TestR477`. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-478** | **A unit left by a removed install counted as the reinstall's fresh copy.** Closed in v0.240.0: `usableRestorePoint` refuses a copy older than the app's `deployed_at`; R-474's fix removes such units anyway. Red-proof: dropping the check fails `TestR478`. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-474** | **„Delete backups" deleted only `db-dumps`; the unit and the Tier-2 mirror survived; prefs stayed.** Closed in v0.240.0 for the backups half: the whole unit, every mirror (`Tier2MirrorDirsForApp` / `RemoveTier2Mirrors`, exact-path guarded) and the prefs go; `backup_paths_removed` lists them. Proven live: 244M + 243.6 MB removed, nothing left on either drive, prefs `None`. **The `volumes_removed: null` half is re-filed as R-489.** `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-466** | **Removal with „Mentési adatok törlése" left the unit's `compose/` + `manifest.json`.** Subsumed by R-474's fix in v0.240.0 (the whole unit goes). Decided by the same change: the button means the WHOLE unit. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-482** | **`adventurelog` ran Django with DEBUG=True on the public origin.** Closed in catalog `ed2c018`: `DEBUG=False` on the backend service. Proven live after the sync: the same CSRF failure renders the 300-byte production page, no debug text. `wger`, `tandoor`, `paperless-ngx` are recorded as unmeasured. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-453** | **`~/.config/credentials` values are single-quoted, and a half-applied strip produced a confidently wrong "password is stale" verdict — twice.** Closed 2026-09-13: the instrument the row asked for exists — `felhom.eu/scripts/read_credential.py KEY <0600-file>` (`66156c6`, 2026-08-31) — and the whole 2026-09-13 night used it for every controller and hub password with zero quoting incidents; the memory `credentials-file-values-are-quoted` now names it. The instrumentation lesson (a discriminator that rules out one alternative does not rule in the rest) stays in the memory. | **CLOSED 2026-09-13 — INSTRUMENTED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-461** | **`runbooks/target-selection.md` named a venue that does not exist and fenced a fixture that is gone.** Closed 2026-09-13, both halves checked against both boxes first: (a) no `/mnt/nvme-1tb` on demo-hp or demo-felhom; demo-hp's NVMe is `nvme0n1` at `/mnt/hdd_1` (demo-felhom's `/mnt/hdd_1` is `sdb`) — the runbook now names `/mnt/hdd_1` and says it is the same disk as the data drive; (b) `qm list` is empty on BOTH boxes — `drill-r50` (VM 300) exists nowhere; the fence text stays with the measured absence written beside it, and R-93 carries the fact. | **CLOSED 2026-09-13 — DOCUMENTED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
@@ -319,7 +319,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing
|
||||
| **R-124** | **The recipe spells PBS's root namespace `"root"`, but the PBS API spells it `""`** and no namespace is literally named `root` — an operator pasting the field into `pct restore --ns root` gets a failure | READY (XS) | — | Pre-existing wire convention (`ToHub` has normalised empty→`"root"` since slice 6), deliberately NOT changed under R-106 so the field's meaning did not shift mid-fix. Documented at `hub.PBSRootNamespace`. Affects only a box with no `namespace` line — **no real customer today**, all three are per-customer. Fix = emit `""` + rely on `namespace_state`, or emit a `--ns`-ready form | CC |
|
||||
| **R-89** | Retention as a per-customer **commercial** policy on the hub | READY (increment 2) | — | Policy object + reconciler → ep0 prune job; keep box tokens write-only | CC |
|
||||
| **R-92** | Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable | READY (XS) | — | Widen precision when retention becomes customer-visible | CC |
|
||||
| **R-93** | `drill-r50` is both a blocked customer and the only drift fixture | READY (XS) | — | Retire it for a synthetic fixture, or unblock + silence per-customer | CC |
|
||||
| **R-93** | `drill-r50` is both a blocked customer and the only drift fixture **FACT 2026-09-13 (R-461): the fixture is GONE — `qm list` is empty on both demo boxes, so neither option is available and the row's premise no longer holds; the operator decides whether that closes it or reopens it as "build a drift fixture".** | READY (XS) | — | Retire it for a synthetic fixture, or unblock + silence per-customer | CC |
|
||||
| **R-129** | **Every doc says demo-hp has "no baked SSH key"** and needs the G1 break-glass password — but `ssh -o BatchMode=yes demo-hp` authenticated **by key**, first try, 2026-07-31 | READY (XS) | — | Stale in the expensive direction: a session that believes it sends itself to the hub vault for a credential it does not need. Verify who owns the key and when it landed, then correct `CLAUDE.md`, `runbooks/target-selection.md:41-42`, `runbooks/workspace-CLAUDE.md` and `felhom-agent/CLAUDE.md` together — or remove the key if it was not deliberate | CC |
|
||||
| **R-130** | **A "hard min" that only warns.** A fresh box's `local-lvm` was ~75 GiB against `HARD_MIN_LVM_GIB=120` (`scripts/felhom-host-install.sh`); the installer logged `[WARN] local-lvm free ~75 GiB < hard min 120 GiB` and went on to a **fully successful** install | READY (S) | — | Either the minimum is not hard (rename it and state the real floor) or it is wrong (and 120 GiB is not what a working appliance needs). Leaving it is the R-29 shape: a check that reads as coverage while providing none. Evidence: same audit §8 | CC |
|
||||
| **R-131** | **`sess-f` is a fourth orphaned scratch customer** on the hub ("R-120 golden 0.186.0 proof", DOWN), left by the 2026-07-30 session | READY (XS) | — | After `drill-r50`, `sess-c`, `sess-d` — the accumulation `runbooks/target-selection.md:86-87` and `PROMPT-TEMPLATE.md` §13 both warn about, now on its fourth instance. Delete it (see the recorded command in `audits/tester-gate-golden-0.188.0-2026-07-31.md` §7.1); the recurrence itself argues for a periodic scratch-customer sweep rather than another reminder | CC |
|
||||
@@ -682,31 +682,27 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-452** | **[P3-LOW] Nothing enforces `catalog_since`, so the one number the update badge shows can silently under-report.** `app-catalog-felhom.eu` `CLAUDE.md` now states the rule — any commit that changes an `image:` line must set that app's `catalog_since` to the same day — and all 53 apps were backfilled from git history on 2026-09-02 (`69761cf`). **A rule with no instrument is a wish; that is this project's most-repeated finding and this row exists so it is not repeated silently.** A stale `catalog_since` makes „Frissítés elérhető — N napja" under-report N, which is the single number the badge exists to give. **WHY IT WAS NOT BUILT IN THE SAME SESSION, stated rather than implied:** the gate would have to diff an `image:` line against the PARENT commit, and `catalog_gates.py` runs under a runner that fetches at `--depth 1` — there is no parent to diff against. The gate therefore needs a deeper fetch, which is a change to the CI shape and not to a script. **This is the R-421 class in advance: an enumerated gap becomes a row in the same session it is enumerated.** `architecture/09-update-architecture.md` §8.2 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-453** | **[P3-LOW] `~/.config/credentials` quotes its values with SINGLE quotes, and a half-applied strip cost a session an hour and produced a WRONG diagnosis that was published before the operator corrected it.** MEASURED 2026-09-02: extracting `PASSWORD` with `sed 's/^"//;s/"$//'` — double quotes only — sent the literal `'` characters as part of the password, and `POST /login` returned **HTTP 200 + `Hibás jelszó`** on BOTH demo boxes. **THE DIAGNOSIS THAT FOLLOWED WAS WRONG AND WAS WRITTEN INTO A REGISTER ROW, A STATUS ITEM AND A MEMORY BEFORE IT WAS CHECKED:** "the vaulted password is stale on both boxes". The operator answered in one line — *the value is in single quotes* — and one retry with `sed "s/^['\\"]//;s/['\\"]$//"` returned **302 + `felhom_session`**. **THE INSTRUMENTATION LESSON, which is the actual finding and outlives the typo:** the controller's own log line `auth.go:176 [WARN] Failed login` was quoted as the discriminator, and it IS a true and useful one — **it separates "wrong password" from "wrong Host header", and that is ALL it separates.** It cannot distinguish a wrong password from wrong password HANDLING, and it was read as if it could. A discriminator that rules out one alternative is not a discriminator that rules in the remaining one. **This is the SECOND time this exact file's quoting has produced a confidently wrong verdict** — the memory `credentials-file-values-are-quoted` was minted for the first (`cut -d=` keeps the quotes → a wrong "the password is stale" diagnosis), and this session applied that memory HALF, stripping one quote character and not the other. **The fix is an instrument, not a resolution to be careful:** one shared helper that extracts a value from that file correctly, used everywhere, so the next session cannot get it half right. Evidence: `tests/VALIDATION-update-slice12-2026-09-02.md` §4.0, which states the correction rather than quietly removing the claim. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-456** | **[P3-LOW] A partly-dead stack is not a boot orphan, and that is written down nowhere.** MEASURED 2026-09-02 on demo-hp while validating v0.233.0: `docker rm -f bookstack` (leaving `bookstack-db` running) then a controller restart produced `Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]` — **bookstack was NOT selected**, although the app container was gone and `desired_state: running` was recorded. Removing `bookstack-db` as well made the whole stack orphaned and the very next pass repaired it in 6.3 s. **So `bootrecon.isBootOrphan` requires the stack as a WHOLE to be down; one live member is enough to make it invisible to the reconciler.** **NOT called a defect, and the reason is part of the row:** `StateDegraded` IS in `IsDownState`, and the crash-loop/dead-app alarm path (`classifyRunStates`) does count a degraded stack as down — so the customer IS told; it is the automatic REPAIR that does not fire, and there may be a good reason (repairing half a stack while its DB is live is not obviously safe). **What is certain is that nobody has written the rule down**, so the next session re-derives it the same way this one did — by watching a reconciliation not happen, which is an absent observable and the weakest possible evidence. Either state the rule in `02-controller-module-map.md` with a test pinning it, or change it. Owner: **CC.** `tests/VALIDATION-update-slice12-2026-09-02.md` §2.2 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-461** | **[P3-LOW] `runbooks/target-selection.md` names a venue that does not exist and fences a fixture that is gone.** MEASURED 2026-09-06 on demo-hp while siting the R-449 guest. (a) The runbook says to put VM disks on a dir storage at **`/mnt/nvme-1tb`, at its root**. **There is no `/mnt/nvme-1tb`** — the 1 TB NVMe is mounted at **`/mnt/hdd_1`**, which is the enrolled user-data drive and the `felhom-backup` target, i.e. the same disk under a different path. The instruction's REASON is still exactly right (`local-lvm` is an over-subscribed thin pool backing the live guest 9201, and this run kept off it — `local-lvm` read **30.50 % before and after**), so the fence held; only its address is stale. (b) The runbook forbids destroying **`drill-r50` (VM 300)**, "the only drift fixture (R-93)". **`qm list` returns nothing on demo-hp** — there are no VMs at all. **The fence currently protects nothing, and R-93's premise that a drift fixture exists is false.** **Why it is a row and not a quiet edit:** a runbook that names a path nobody can find is one a session works around, and working around a safety instruction is how the instruction stops being followed. Both halves need checking against the box before the text is changed — (b) in particular may mean R-93 should be closed or reopened as "the drift fixture is gone", which is a different fact from "do not destroy it". Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §7 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 | **READY — rank P2-MEDIUM; owner: VIKTOR rules on scope, CC implements** |
|
||||
| **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-465** | **[P3-LOW] `cfg.Paths.HDDPath` — the global that R-442 proved is set on NO box (demo-hp AND demo-felhom: 0 `hdd_path`, 0 `FELHOM_PATHS_*`) — still has SIX readers, each reading an always-empty value:** `report/builder.go:69`, `monitor/healthcheck.go:35`, `api/router.go` (system-info), `web/server.go:740`, `cmd/controller/main.go` (auto-discovery seed + metrics HDD path). Removal was silently inert for months on the very same read. Whether any of these is inert the same way — a report field that is always empty, a health check that never fires, a metric never collected — is a one-hour audit: for each reader, name the POSITIVE observable that must appear when it works and check it on the box. R-442 §5 said "do not delete it here"; this row is the audit it deferred. Owner: **CC.** `felhom-controller/REPORT.md` (v0.236.0, Observations 1) | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-466** | **[P3-LOW] Removing an app with „Mentési adatok törlése" ticked leaves its recovery unit's `compose/` + `manifest.json` on the drive.** The router passes only `backup.AppDBDumpPath(nsRoot, name)` to `RemoveStack`, so `backups/primary/<app>/db-dumps/` goes and the unit root keeps `compose/` (the app's `app.yaml` with the portable secret class at 0600) and `manifest.json`. MEASURED 2026-09-13 on demo-hp after the R-442 Scenario A removal — `audits/R442-2026-09-13/teardown-and-log.txt`, residue check: `backups/primary/nextcloud` still listed with the app gone; removed by hand at teardown. A customer who asked for the backups to go is left with the app's definition and a manifest. **Decide:** the button means the WHOLE unit (pass the unit root, under the same `backups/`-prefix guard) or stays db-dumps-only (then the modal must say so). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
| **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** |
|
||||
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). | **READY — unblocked by R-448; rank P3-LOW; owner: CC (removal is a deliberate act)** |
|
||||
|
||||
| **R-471** | **[P3-LOW] `observations_gate.py` reads the FIRST `Observations` section of `REPORT.md` and nothing after it, so an appended second section with an unmarked item passes — and the decoy that proves it has been reading "LIVE HOLE" at HEAD, unnoticed, because the decoy suite is run by hand.** MEASURED 2026-09-13 on a clean worktree of `4b2e560`: `python3 scripts/test_gate_decoys.py` → `FAIL: observations/R-419: decoy PASSED - LIVE HOLE (rc=0)`; the gate's own output shows it scanned `## 11. Observations — noticed, documented, NOT acted on` (3 items, all marked) and never reached the appended `## Observations` block the decoy planted. Whether the cause is "first heading wins" or a heading-shape filter is NOT established — only that the planted unmarked item was not seen. **Consequence:** a report with two observation sections gets the second one unchecked. **Two fixes, both owed:** scan every section whose heading contains `Observations`, and put `test_gate_decoys.py` where something runs it (it is the instrument for R-421 and nothing in `repo_gates.py` invokes it). | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
| **R-473** | **[P2-MEDIUM] The `glance` catalog template crash-loops on EVERY fresh install.** MEASURED 2026-09-13 on demo-hp (deployed as a throwaway for the slice-4 live test): `restarts=13 status=restarting`, the log repeating `parsing config: reading /app/config/glance.yml: open /app/config/glance.yml: no such file or directory`, the `glance_config` volume empty. The image does not create a default config and the template seeds none. The install page reports the deploy as started and the card then shows a restarting app. Not a version problem: the template's current tag v0.8.5 has the same config requirement (the catalog walk that pinned it never did a fresh install). **Fix shape:** seed a minimal `glance.yml` the way `gokapi`'s catalog entry seeds `config.json` (memory `gokapi-headless-setup`), then prove a fresh install lands healthy. Evidence `audits/slice4-2026-09-13/live/02-glance-abandoned.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-474** | **[P3-LOW] "Remove app" with *also delete backups* deletes only the `db-dumps` directory, and reports `volumes_removed: null` over a volume it DID remove.** MEASURED 2026-09-13 removing the glance throwaway on controller v0.237.0 (`remove_hdd_data:true, remove_backups:true`): response `{"removed":"glance","volumes_removed":null,"hdd_paths_removed":[],…}`; afterwards `docker volume ls` shows no glance volume, while `backups/primary/glance/{compose,manifest.json}` and the stack dir's `applied-compose.yml` remain. The router passes ONLY `backup.AppDBDumpPath(nsRoot, name)` (router.go, beside a comment that disk-tier backup "moved to the host agent" — stale since Tier 2 returned to the controller). So the recovery unit and any Tier-2 copy survive a removal that promised to delete backups, and the answer names no volume. Same class as R-442: the removal's answer does not describe what happened. (Removal also refuses a crash-looping app as "still running — stop it first", which is correct and was observed.) **REPRODUCED a second time the same day** removing the uptime-kuma throwaway: `volumes_removed: null`, and `backups/primary/uptime-kuma/{compose,manifest.json,volume-dumps}` plus `backups/secondary/uptime-kuma/recovery-unit` survived `remove_backups:true` (residue cleared by hand; `audits/slice4-2026-09-13/` `live/15-teardown.txt`). **REPRODUCED a third time 2026-09-13 (afternoon, v0.239.0)** removing the gokapi and actualbudget throwaways with `remove_backups:true`: both `backups/primary/<app>/` units survived (actualbudget's with its 74 KB volume tar), `volumes_removed: null` again, and both `app_backup` prefs remained (`10-teardown.txt`). It also fed R-478. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-476** | **[P3-LOW] The Mentések page names a Tier-2 copy's date from the unit MANIFEST, which moves only when the app's DEFINITION changes — so it can undersell a fresh copy by a day or more.** MEASURED on demo-hp 2026-09-13: bookstack's mirror held `bookstack-mariadb.sql` written 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z; the Tier-2 run succeeded at 01:30Z. A capture rewrites the manifest only when checksums, dump NAMES or controller version change, and nightly dumps keep their names. R-403 chose the package date so a PRESERVED package is never shown as fresh — correct — but for a normal run it is the older, flattering-in-reverse date. The update (slice 4) deliberately ages the copy by the last successful copy instead (`Tier2RestorePoint.ProvenCopyTime`), so the two can name different dates. Fix shape: record the unit's DATA time (newest dump mtime) in the manifest, or name `LastSuccess` when the leg was not preserved. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-477** | **[P3-LOW] The update's Tier-3 lookup pays its full 15 s bound on demo-hp, and the bound shows up as a WARN about a DIFFERENT app.** MEASURED 2026-09-13 (controller v0.239.0, Scenario K): the `checking` phase ran 15:25:45 → 15:26:00, and the only line in the window was `[WARN] [offbox] inventory: size of kimai's newest snapshot unknown: offbox stats b4350226: signal: killed`. Cause: `UpdateRestorePoints` calls `OffsiteInventoryList`, which runs one `snapshots` AND one `stats` per app; the update needs only the snapshot times, and the 15 s context killed a `stats`. **Consequence:** every update with no fresh Tier 2/Tier 1 copy waits up to 15 s before backing up, and the operator log blames kimai for an update of actualbudget. **Fix shape:** a snapshots-only lookup for the update path (no sizes), so the bound is a guard and not the normal duration. `audits/rulings-r472-r475-2026-09-13/04-K-no-unit-anywhere.txt` | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-478** | **[P3-LOW] A recovery unit LEFT BEHIND by a removed install counted as the fresh Tier-1 restore point of a reinstall.** MEASURED 2026-09-13 (v0.239.0, Scenario H): gokapi had been removed earlier; its unit (manifest 06:59:17Z) survived (R-474). A new gokapi was deployed at 15:30:46Z and updated at 15:31:10Z: `precondition met — Tier 1 (own recovery unit) copy from 2026-09-13T06:59:17Z (8h32m0s old)`. That unit belonged to the OLD install — a different `app.yaml` (subdomain, generated password). The capture sweep rewrote it only at 15:34:22Z. **Consequence:** in that window a failed update would name, and a restore would apply, the previous install's definition. **Fix shape:** R-474 deleting the unit on removal closes most of it; independently, a unit whose manifest predates the stack's current deploy should not count. `05-H-gokapi-tier1.txt`, `06-F-setup-unit-rewritten.txt` | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-479** | **[P2-MEDIUM] For an app whose data is a bind mount, the Tier-1 unit holds settings only — so the route back a Tier-1 hold names restores the definition and NOT the data.** MEASURED 2026-09-13 restoring gokapi from „helyi” after a held update: `a beállítások visszaálltak … FIGYELEM: ez a mentés csak a beállításokat tartalmazta, adatot nem`. The restore message is honest; the HOLD sentence („Visszaállítható … saját meghajtó”) does not say it. The ruling (R-475) accepts any tier, in the order 2, 1, 3 — so for such an app a fresh Tier-1 unit is chosen ahead of an off-site snapshot that WOULD carry the data. **Consequence:** an update whose migration rewrote bind-mounted data has no data route back through the copy it named. **Decision-shaped:** either Tier 1 counts only for apps whose unit carries their data (DB dump / volume tar), or the hold sentence says "settings only" for that case. `07-F-hold-names-own-drive.txt`, `08-restore-from-helyi.txt` | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-480** | **[P2-MEDIUM] After a successful restore, the app card still shows the failed update's sentence — which says the running app „leállítva marad”.** MEASURED 2026-09-13 on demo-hp (v0.239.0): the restore cleared the hold (`restore hold CLEARED for gokapi`), gokapi ran healthy, `hold_reason` was empty — but `update_phase=failed` and `update_error` stayed, and `GET /stacks` rendered the sentence inside gokapi's card (control: absent from actualbudget's card). It survived the app's removal too (`not_deployed … phase=failed err=…`). Present since v0.238.0's page; this morning's restore walk did not check the card text. **Fix shape:** a successful restore (and a removal) clears the stack's last update outcome. `09-card-after-restore.txt`, `10-teardown.txt` | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-481** | **[P2-MEDIUM] There is no scratch guest on demo-hp, so the nightly rotation cannot restore a throwaway "into a scratch guest", and the nine standing apps cannot be tonight's throwaway at all.** MEASURED 2026-09-13 (`pct list` / `qm list` on demo-hp: only 9201). The rotation brief needs a second controller guest for two steps — the cross-guest restore, and a throwaway deploy of an app that is already standing on 9201 (the stack name collides, and the standing apps may not be touched). Every earlier cross-guest walk built a whole appliance from the published ISO (VM 323/325, hours each) and enrolled it as a new customer; a `pct clone` of 9201 would carry demo-hp's identity, tunnel and hub enrolment. **Decision-shaped:** which route makes the scratch guest — a persistent second LXC on demo-hp born from the golden template and enrolled as its own customer (`nightly-scratch`), or an ISO-built appliance per night. Until then the rotation restores in place (remove → restore from the unit on the same guest) and skips the standing nine (`runbooks/nightly-rotation.md`). | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-483** | **[P3-LOW] `adventurelog` v0.12.1: a photo upload through the app's own API proxy fails 500 from every non-browser client — whether a household can upload photos in a browser is UNKNOWN and cannot be measured here.** MEASURED 2026-09-13 on demo-hp (throwaway `travelnight`): `POST /api/images` (multipart: `image`, `location`, `is_primary`) through the public origin returns `{"error":"Internal Server Error"}` from the **frontend** (the backend log shows no request at all); the frontend container logs `RequestContentLengthMismatchError: Request body length does not match content-length header` from its undici forwarder. Same result with an ASCII filename, an accented one, urllib, curl `-F`, and curl with `Transfer-Encoding: chunked`. Everything else on the API — sign-up, the frontend's login form, collections, locations, visits, notes, edit, delete — works. **What is NOT established:** whether the app's own browser page uploads succeed (a browser's FormData body may satisfy the proxy). Per `CLAUDE.md`, that is a manual click-through: open `https://<sub>.<domain>`, add a photo to a place. **Not ours to fix in the template** (a frontend proxy bug); a newer catalog pin is a version promotion, not tonight's. Evidence: `audits/nightly-2026-09-13-adventurelog/03e-photo-500.txt`. | **WAITING-ON-OPERATOR — needs a browser click-through; rank P3-LOW; owner: VIKTOR checks, CC re-files** |
|
||||
| **R-487** | **[P2-MEDIUM] A removed app whose backups were kept is listed on NEITHER backup page, so the restore that brings it back has no button — the customer's remove-by-mistake route exists only as an endpoint.** MEASURED 2026-09-13 on demo-hp (nightly rotation, `adventurelog` removed with backups kept, unit + mirror on disk): `GET /backups/apps` and `GET /backups/restore` contain the string `adventurelog` zero times; `POST /backup/restore stack_name=adventurelog snapshot_id=helyi` then restored it in 22 s with the data byte-identical. Cause: `buildAppBackupRows` walks `status.AppDataInfo` = `DiscoverAppData` over DEPLOYED stacks only. The off-site list had exactly this defect and was fixed by keying it on the store (R-237, v0.204.0); the local and Tier-2 lists were not. **Fix shape:** list every app with a recovery unit on a registered drive (`ListRestorePoints` over the primary dirs), marking removed ones „eltávolítva — visszaállítható"; the unit restore already reinstalls (R-253). Not a design reversal — the same rule R-237 set. Evidence: `audits/nightly-2026-09-13-adventurelog/05b-restore-tier1.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user