From 07059427835da7938afe1c025378489afa7dd826 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 2 Sep 2026 20:47:53 +0200 Subject: [PATCH] the badge IS proven live, and the 'stale password' finding was mine, not the box's I reported that the vaulted dashboard password no longer worked on either demo box, and quoted the controller's own 'Failed login' as the discriminator. The password was fine. ~/.config/credentials quotes its values with SINGLE quotes and my sed stripped only double quotes, so the quote characters went out as part of the password. The operator corrected it in one line; one retry returned 302. The instrumentation lesson is the finding and R-453 now carries it: 'Failed login' separates wrong-password from wrong-Host-header, and that is ALL it separates. It cannot tell a wrong password from wrong password HANDLING, and I read it as if it could. This is the second time this file's quoting has produced a confident wrong verdict, so the fix is one shared extraction helper, not a resolution to be careful. With the session recovered, the badge is validated on live pages: Naprakesz twice on /stacks and on /apps/bookstack; NO badge at all on /apps/docmost (a deployed app with no record - absent is UNKNOWN, not current); and 'Frissites elerheto - 52 napja' on both surfaces, the age being real arithmetic on bentopdf's catalog_since. The behind state was staged by editing one compose tag, with no restart and no up -d, and reverted byte-identically (sha256 equal, diff empty, container never touched). Capability-map row upgraded to PROVEN-LIVE with the one unexercised badge state named. STATUS item 9 now needs nothing from the operator. --- REPORT.md | 11 +- STATUS.md | 49 +++----- .../architecture/00-capability-map.md | 2 +- documentation/backlog/OPEN-ITEMS.md | 2 +- .../VALIDATION-update-slice12-2026-09-02.md | 115 +++++++++++++----- 5 files changed, 116 insertions(+), 63 deletions(-) diff --git a/REPORT.md b/REPORT.md index 4aeef77d..aa1481b6 100644 --- a/REPORT.md +++ b/REPORT.md @@ -42,7 +42,7 @@ to `CLOSED-ITEMS.md` and that file is untouched. | R-450 | slice 6 — version sequence; an engine change gets its own edge | READY, P2-MEDIUM | | R-451 | slice 7 — a fleet sweep (needs a hub change: no image field is reported) | READY, P3-LOW | | R-452 | no gate enforces `catalog_since` (`--depth 1` has no parent to diff) | READY, P3-LOW | -| R-453 | **the vaulted dashboard password is stale on BOTH demo boxes** | **WAITING-ON-OPERATOR**, P2-MEDIUM | +| R-453 | **CORRECTED** — the password was fine; `~/.config/credentials` uses SINGLE quotes and the strip was half-applied. The finding is the instrumentation lesson, not the typo | READY, P3-LOW | | R-454 | five `gofmt`-unclean `internal/web` test files, and no gate notices | READY, P3-LOW | | R-455 | **DooPlex has no Docker Hub login — the ceiling now blocks BUILDS, not just gates** | **WAITING-ON-OPERATOR**, P2-MEDIUM | | R-456 | a partly-dead stack is not a boot orphan, and that rule is written down nowhere | READY, P3-LOW | @@ -56,9 +56,12 @@ reconciler — no hand-set state): `bentopdf` recorded 1 service, `bookstack` re compose service name, **all three digests matching ground truth read independently beforehand**. Encrypted secrets byte-identical across the write. Both apps up and healthy; nothing provisioned. -**NOT live-validated: the rendered badge.** The vaulted dashboard password is stale on both demo -controllers (R-453) and there is no operator route to a customer's password. Five attempts are listed -in §4 of the validation file. The render is covered by tests that render the PRODUCTION templates. +**The badge is ALSO proven live**, on both surfaces and in three of its four states — including the +load-bearing one: a deployed app with **no record renders no badge at all**. „Frissítés elérhető — +**52 napja**" was produced by editing bentopdf's compose tag only (no restart, no `up -d`) and reverted +byte-identically. **One correction of my own is stated in §4.0 of the validation file rather than +buried: I first reported the vaulted password as stale on both boxes. It was not — I stripped only +double quotes from a single-quoted value.** The operator caught it in one line. ## Sibling repos diff --git a/STATUS.md b/STATUS.md index 511b54bd..a97003d0 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,9 +1,9 @@ # STATUS — what works, what's broken, what's next **Updated 2026-09-02 — the box now writes down which version of each app it is running, and shows one -small label saying whether it is up to date. No version numbers, and nothing about updating changed. -The writing-down half is proven on the real machine; for the label I need one password from you -(item 9). Nothing is broken while it waits.** +small label saying whether it is up to date: „Naprakész" or „Frissítés elérhető — 52 napja". No version +numbers, and nothing about updating changed. Both halves are proven on the real machine, labels read off +the real pages. NOTHING NEW IS WAITING ON YOU.** **Earlier 2026-09-01 (third pass) — the backup work is FINISHED for beta, and I have written down where it stops. One alarm that was telling you something untrue is fixed and live (hub 0.111.1). @@ -23,7 +23,7 @@ not an evening's work.** *This section is allowed to be longer than one screen, and each item says what happens if you do nothing.* -1. **Three things are waiting on you — item 4 (send two e-mails), item 7 (one design decision) and item 9 (one password, new today, two minutes).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on +1. **Two things are waiting on you — item 4 (send two e-mails) and item 7 (one design decision).** Item 9 is new today and needs nothing from you. Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on the real machines: - the background job that could delete a live restore's lock now waits its turn — and the check that finds the next one like it is a test, not a comment, so it cannot come back quietly; @@ -116,35 +116,26 @@ nothing.* Full measurement, with the controls and the quoted output: `felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md`. -9. **I need the dashboard password for `demo-hp`, or your go-ahead to reset it. Two minutes, and - nothing is broken while it waits.** +9. **Nothing here needs you. I got in on my own — after getting it wrong first, which is worth one + line.** - Today the box learned to write down which version of each app it is really running, and to show - the customer one small label: **„Naprakész"** or **„Frissítés elérhető — 45 napja"**. No version - numbers — a household cannot act on `26.05.2`. + The box now writes down which version of each app it is really running, and shows the customer one + small label: **„Naprakész"** or **„Frissítés elérhető — 52 napja"**. No version numbers — a + household cannot act on `26.05.2`. - **The writing-down half is proven on the real machine.** I made two apps fail to come back, the box - repaired them by itself, and it wrote down exactly what it installed — one line per container, with - the fingerprint that cannot lie. I checked those fingerprints against the machine independently and - they match. Both apps are up and healthy, and no data was touched. + **Both halves are proven on the real machine.** I made two apps fail to come back, the box repaired + them itself and wrote down exactly what it installed — one line per container, with the fingerprint + that cannot lie. I checked those fingerprints against the machine independently and they match. Then + I opened the real customer pages and read the labels off them. - **The label half I could not look at.** To open a customer page I have to log in as the customer, - and the password we keep on file no longer works — on `demo-hp` **or** on `demo-felhom`. I tried - both, and the box's own log says "wrong password", not "wrong address". There is no operator route - to a customer's password: it is only ever e-mailed to them. + **Where I went wrong, because you should not have to spot it twice.** I told you the saved password + no longer worked on either machine. **It worked fine.** I read it out of the file wrongly — the + value is wrapped in quote marks and I only removed one kind. You caught it in one line. The machine + told me "wrong password", which was true, and I took it to mean the password was wrong when it meant + *what I sent* was wrong. I have filed that as a small job for myself: one shared way of reading that + file, so the next session cannot get it half right. - **Two ways forward. Pick one:** - - - **Send me the current `demo-hp` dashboard password.** I use it for two page loads and nothing - else. Simplest, and it changes nothing on the box. - - **Let me reset it back to the one on file.** Same thing we did on 9 August. It also repairs the - stored password, so the next session does not lose this half hour again. - - **My pick: let me reset it** — otherwise the same wall is there next week, on both machines. - - **If you do nothing:** nothing breaks and no customer is affected. The label is covered by tests - that render the real pages, so this is a confirmation, not a discovery. The box goes on recording - versions either way. + **If you do nothing:** nothing. This one is closed. 8. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session, diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 979b14dc..6e8cc8a0 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -104,7 +104,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. | | **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. | | Protected infra stacks can't be stopped/removed from UI | controller | **PROVEN-LIVE** | `CAMPAIGN-nomercy` + `RERUN-p1p3` T-SEC-PROTECTED (refuse stop/remove, stay Up) | (Cited `CAMPAIGN-2` T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal) | -| **What VERSION a box is running, and whether it is behind the catalog** | controller v0.233.0 + catalog `69761cf` | **PROVEN-LIVE for the RECORD; IMPLEMENTED for the BADGE** — and the split is the point, not a hedge. | **`tests/VALIDATION-update-slice12-2026-09-02.md`** — on demo-hp 0.233.0, through a REAL production caller (`bootrecon → StartStack → compose up -d → recordInstalledImages`, no hand-set state): `bentopdf` recorded **1** service and `bookstack` recorded **2**, keyed by compose SERVICE name, and **all three digests match the ground truth read independently from the containers before anything was touched**. `catalog_since` reached the box on the normal 15-minute sync. bookstack's two encrypted secrets are byte-identical across the write. Unit side: `installed_test.go` + `updatebadge_test.go`, incl. a wiring test through a real `RestartStack`, an AST walk of all four call sites, and three companion red-proofs. | **THE UNEXERCISED LEG, NAMED: the rendered badge has never been seen on a live page.** The vaulted dashboard password is stale on BOTH demo controllers (`Hibás jelszó`, confirmed against the controller's own log, and the same on demo-felhom), and there is no operator-side route to a customer's dashboard password (R-119). Every INPUT the badge reads is verified live; the render is covered only by tests that render the PRODUCTION templates. **Absent means UNKNOWN, never current** — a legacy `app.yaml` renders NOTHING, red-proved. **No version number is shown to the customer** and **no registry is queried**, so „Naprakész" CAN BE FALSE for the 23 floating pins (**R-446**). Nothing about updating changed: R-438, R-440, R-441, R-443 all stand. Reasoning: `architecture/09-update-architecture.md`; remaining slices R-447..R-452 | +| **What VERSION a box is running, and whether it is behind the catalog** | controller v0.233.0 + catalog `69761cf` | **PROVEN-LIVE — both the record and the rendered badge.** | **`tests/VALIDATION-update-slice12-2026-09-02.md`** — on demo-hp 0.233.0, through a REAL production caller (`bootrecon → StartStack → compose up -d → recordInstalledImages`, no hand-set state): `bentopdf` recorded **1** service and `bookstack` recorded **2**, keyed by compose SERVICE name, and **all three digests match the ground truth read independently from the containers before anything was touched**. `catalog_since` reached the box on the normal 15-minute sync. bookstack's two encrypted secrets are byte-identical across the write. Badge evidence: same file §4. Unit side: `installed_test.go` + `updatebadge_test.go`, incl. a wiring test through a real `RestartStack`, an AST walk of all four call sites, and three companion red-proofs. | **THE UNEXERCISED LEG, NAMED: one badge STATE of four.** „Frissítés elérhető" WITHOUT an age needs an app whose `catalog_since` is absent, malformed or future-dated, and all 53 now carry a valid one — unit-tested only (`TestGroupF`). The other three are live: „Naprakész" ×2 on `/stacks` and on `/apps/bookstack`; **NO badge at all on `/apps/docmost`**, a deployed app with no record — absent is UNKNOWN and is not rendered as current; and „Frissítés elérhető — **52 napja**" on both surfaces, the age being real arithmetic on bentopdf's `catalog_since` 2026-07-12. **The behind state was staged by editing bentopdf's compose tag ONLY** — no container restarted, no `up -d` — then reverted byte-identically (`sha256` equal, `diff` empty); so the RENDER is measured and the syncer's own half stays measured separately in the spike. Searched with `grep -oF` ASCII fragments plus positive AND negative controls. **The Frissítés/Újraindítás/Leállítás buttons are unchanged on the live page in the behind state.** **Absent means UNKNOWN, never current** — a legacy `app.yaml` renders NOTHING, red-proved. **No version number is shown to the customer** and **no registry is queried**, so „Naprakész" CAN BE FALSE for the 23 floating pins (**R-446**). Nothing about updating changed: R-438, R-440, R-441, R-443 all stand. Reasoning: `architecture/09-update-architecture.md`; remaining slices R-447..R-452 | | Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad `backup:` block degrades to legacy, loudly) | controller v0.132, catalog | **PROVEN-LIVE** | `CAMPAIGN-2` T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs | | | Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) | agent v0.94.0→**v0.95.0**, controller v0.169.0→v0.171.0→**v0.215.0**, hub v0.73.1 | **PROVEN-LIVE (healthy path + delivery + the severity wire).** **IMPLEMENTED, NOT proven-live: the Hiba-from-counters path** (v0.215.0) — it has never fired on real hardware, only against the committed fixture's values in unit tests (**R-332**) | **2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix):** the card on guest 9201 now shows BOTH real disks with **real verdicts + human model labels** — **„AirDisk 512GB SSD" → Rendben (34°C)** (the system SSD, via LVM/dm resolution) and **„TOSHIBA MQ04ABF100" → Rendben (30°C)** (the USB, via union-path SMART). `/disks` carries `smart.health=PASSED` + `model_name` for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (`SPIKE-smart-coverage-2026-07-25.md` had proven both disks answer `smartctl -a -j` PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. **Notification pipeline PROVEN-LIVE 2026-07-24** — a `disk_health_degraded` POST (the exact `notify.PushEvent` wire call) was **400-rejected by hub v0.73.0** and **200-accepted + „Operator email sent" by hub v0.73.1** | No new smartctl load; feature-detect by payload presence → **MinAgent floor unchanged**; no sudoers/`-d sat` change. **No global banner** (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin `local` on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + `model_name` capture. **2026-08-14 — a genuinely failing disk HAS now been seen, and it broke three assumptions** (`audits/DIAG-smart-passed-trap-2026-08-14.md` + two committed fixtures: raw `smartctl -a -j` and 406 `smartd` lines from ST3000VX010 S/N Z6A07P2G). **(1)** `smart_status.passed` is STRUCTURALLY incapable of failing on unreadable sectors — attrs 187/197/198 all carry `thresh: 0` and a normalized value floors at 1 — so the drive read PASSED at 352 pending sectors and 1001 uncorrectable reads. **(2)** The alert it did produce carried severity `"warn"`, which the hub coerces to `info` and never emails: **the counterfactual is ZERO emails about this drive** (R-328, fixed controller v0.215.0, and the `warning`-vs-`warn` pair proven side by side in `notification_log` on 2026-08-14 — `sent` vs no row at all). **(3)** The old check spoke once and forgot on restart, so between 8 and 352 sectors it emitted nothing. v0.215.0 adds the sustained/count/heat Hiba rules, persisted state and an hourly cadence. **The verdict half of that arm remains unit+red-proof covered only** — no live drive has reached Hiba from counters (R-332). **SMART history/trending (hub-side) PARKED** (ROADMAP R-73) | | App crashes → customer notified (one event per transition, no flapping spam) | controller v0.120, hub v0.48 | **IMPLEMENTED** | controller v0.120.0 (dead-app alerting, `app_start_failed`, one-event-per-transition red-proofs); `CAMPAIGN-3` F11 surfaced the gap | End-to-end crash→customer-email delivery never live-confirmed (6B deferred / 6C inconclusive: clean stop ≠ crash); anti-spam unit-proven | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index a84e25e0..5d403601 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -690,7 +690,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** | | **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 | **READY — rank P3-LOW; owner: CC** | | **R-452** | **[P3-LOW] Nothing enforces `catalog_since`, so the one number the update badge shows can silently under-report.** `app-catalog-felhom.eu` `CLAUDE.md` now states the rule — any commit that changes an `image:` line must set that app's `catalog_since` to the same day — and all 53 apps were backfilled from git history on 2026-09-02 (`69761cf`). **A rule with no instrument is a wish; that is this project's most-repeated finding and this row exists so it is not repeated silently.** A stale `catalog_since` makes „Frissítés elérhető — N napja" under-report N, which is the single number the badge exists to give. **WHY IT WAS NOT BUILT IN THE SAME SESSION, stated rather than implied:** the gate would have to diff an `image:` line against the PARENT commit, and `catalog_gates.py` runs under a runner that fetches at `--depth 1` — there is no parent to diff against. The gate therefore needs a deeper fetch, which is a change to the CI shape and not to a script. **This is the R-421 class in advance: an enumerated gap becomes a row in the same session it is enumerated.** `architecture/09-update-architecture.md` §8.2 | **READY — rank P3-LOW; owner: CC** | -| **R-453** | **[P2-MEDIUM] The vaulted customer dashboard password is STALE ON BOTH DEMO BOXES, and there is no operator-side route to the real one — so no session can drive a customer page.** MEASURED 2026-09-02 while trying to live-validate the update badge: `PASSWORD` from DooPlex `~/.config/credentials` returns **HTTP 200 with the login page and the body string `Hibás jelszó`** against demo-hp guest 9201 (`https://192.168.0.138:443`, `Host: felhom.enkisfelhom.hu`) **and** against demo-felhom guest 9201 (`https://192.168.0.149:443`, `Host: felhom.demo-felhom.eu`). **The discriminator is the controller's own log, not the status code** — `auth.go:176: [WARN] [web] Failed login` proves wrong PASSWORD rather than wrong Host header, which is the trap this class always presents (a rejected login renders no flash and looks exactly like a routing problem). Every other key in the credentials file was checked and none is a dashboard password (`HUB_PW`, `TS_KEY`, `HETZNER_API`, `ISO_S3_*`, and the `R_*` keys are escrow recovery codes). **THIS IS THE SECOND TIME:** it drifted on demo-hp on 2026-08-09 and was put back on operator instruction by writing a fresh bcrypt hash into the guest's `data/settings.json`; demo-hp was then reinstalled and re-claimed on 2026-08-21, and demo-felhom has now drifted as well. **Why it is not merely inconvenient: it silently converts "endpoint-level validation" — this project's STANDARD method, because there is no browser on DooPlex — into "unit tests only" for anything that renders a customer page.** The cost is paid per session and rediscovered each time. **The fix is a decision, not a command:** the claim code is bcrypt-hashed hub-side and only e-mailed (R-119), so either the operator records the current demo passwords out-of-band, or a re-set becomes a standing authorisation for the two Tier-0 demo boxes. **NOT TAKEN UNILATERALLY:** re-setting a dashboard password is a decision about a customer account, and it was done under operator instruction last time. Raised in `STATUS.md` item 9. Evidence: `tests/VALIDATION-update-slice12-2026-09-02.md` §4, which lists all five attempts. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR decides, CC executes** | +| **R-453** | **[P3-LOW] `~/.config/credentials` quotes its values with SINGLE quotes, and a half-applied strip cost a session an hour and produced a WRONG diagnosis that was published before the operator corrected it.** MEASURED 2026-09-02: extracting `PASSWORD` with `sed 's/^"//;s/"$//'` — double quotes only — sent the literal `'` characters as part of the password, and `POST /login` returned **HTTP 200 + `Hibás jelszó`** on BOTH demo boxes. **THE DIAGNOSIS THAT FOLLOWED WAS WRONG AND WAS WRITTEN INTO A REGISTER ROW, A STATUS ITEM AND A MEMORY BEFORE IT WAS CHECKED:** "the vaulted password is stale on both boxes". The operator answered in one line — *the value is in single quotes* — and one retry with `sed "s/^['\\"]//;s/['\\"]$//"` returned **302 + `felhom_session`**. **THE INSTRUMENTATION LESSON, which is the actual finding and outlives the typo:** the controller's own log line `auth.go:176 [WARN] Failed login` was quoted as the discriminator, and it IS a true and useful one — **it separates "wrong password" from "wrong Host header", and that is ALL it separates.** It cannot distinguish a wrong password from wrong password HANDLING, and it was read as if it could. A discriminator that rules out one alternative is not a discriminator that rules in the remaining one. **This is the SECOND time this exact file's quoting has produced a confidently wrong verdict** — the memory `credentials-file-values-are-quoted` was minted for the first (`cut -d=` keeps the quotes → a wrong "the password is stale" diagnosis), and this session applied that memory HALF, stripping one quote character and not the other. **The fix is an instrument, not a resolution to be careful:** one shared helper that extracts a value from that file correctly, used everywhere, so the next session cannot get it half right. Evidence: `tests/VALIDATION-update-slice12-2026-09-02.md` §4.0, which states the correction rather than quietly removing the claim. | **READY — rank P3-LOW; owner: CC** | | **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** | | **R-455** | **[P2-MEDIUM] DooPlex has no Docker Hub login, and the unauthenticated ceiling now blocks BUILDS, not just gates.** MEASURED 2026-09-02: `build.sh 0.233.0 --push` failed at `[internal] load metadata for docker.io/library/debian:bookworm-slim: 429 Too Many Requests`, with `golang:1.24-bookworm` cancelled behind it. Neither base image was in the local store, so there was no fallback. The same window also left the catalog's `image-resolvable` gate INCONCLUSIVE (6 of 65 lookups throttled). **This is R-41's known note — *"the full sweep is still OWED — DooPlex is not logged in to Docker Hub"* — arriving with a bigger bill: it is no longer only a gate that cannot reach a verdict, it is a release that cannot be built.** **WORKAROUND USED AND RECORDED RATHER THAN BURIED:** both base images were pulled from Google's official Docker Hub mirror (`mirror.gcr.io/library/...`) and retagged locally, after which `build.sh` ran unmodified; the two digests are written into `felhom-controller/REPORT.md` §5 so the identity check against Hub is one command when the window clears. **GRADED: that the mirror is byte-identical to Hub is KNOWN, not MEASURED — the throttle was still in force at the end of the session.** **The fix is a credential, and it is a decision:** a Docker Hub account for DooPlex (free tier lifts the ceiling ~6x for authenticated pulls), or an explicit standing ruling that the mirror is the sanctioned source and `build.sh` should name it. **A build that depends on an anonymous third-party quota is not a build you can run when you need to.** Owner: **VIKTOR decides, CC executes.** | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR decides, CC executes** | | **R-456** | **[P3-LOW] A partly-dead stack is not a boot orphan, and that is written down nowhere.** MEASURED 2026-09-02 on demo-hp while validating v0.233.0: `docker rm -f bookstack` (leaving `bookstack-db` running) then a controller restart produced `Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]` — **bookstack was NOT selected**, although the app container was gone and `desired_state: running` was recorded. Removing `bookstack-db` as well made the whole stack orphaned and the very next pass repaired it in 6.3 s. **So `bootrecon.isBootOrphan` requires the stack as a WHOLE to be down; one live member is enough to make it invisible to the reconciler.** **NOT called a defect, and the reason is part of the row:** `StateDegraded` IS in `IsDownState`, and the crash-loop/dead-app alarm path (`classifyRunStates`) does count a degraded stack as down — so the customer IS told; it is the automatic REPAIR that does not fire, and there may be a good reason (repairing half a stack while its DB is live is not obviously safe). **What is certain is that nobody has written the rule down**, so the next session re-derives it the same way this one did — by watching a reconciliation not happen, which is an absent observable and the weakest possible evidence. Either state the rule in `02-controller-module-map.md` with a test pinning it, or change it. Owner: **CC.** `tests/VALIDATION-update-slice12-2026-09-02.md` §2.2 | **READY — rank P3-LOW; owner: CC** | diff --git a/documentation/tests/VALIDATION-update-slice12-2026-09-02.md b/documentation/tests/VALIDATION-update-slice12-2026-09-02.md index 06781e2f..61d7b43a 100644 --- a/documentation/tests/VALIDATION-update-slice12-2026-09-02.md +++ b/documentation/tests/VALIDATION-update-slice12-2026-09-02.md @@ -7,8 +7,8 @@ at the end of the phase that produced it, not at the end of the session. **Which path was used, stated exactly: the BOOT RECONCILER**, `bootrecon.Run → StackProvider.StartStack → compose up -d → recordInstalledImages` — a REAL production caller, the same one `SPIKE-app-update-2026-09-01` §2 variant 1c-ii used, and no hand-set state anywhere. -**Why not the customer's Restart button: see §4 — the vaulted dashboard password no longer opens -either demo controller.** +**Why not the customer's Restart button: §4.0 — the dashboard login was lost to a mistake of mine +for part of the run, and the record half was measured before it was recovered.** --- @@ -123,38 +123,97 @@ $ grep -n catalog_since /opt/docker/stacks/bookstack/.felhom.yml 13:catalog_since: "2026-07-18" ``` -## 4. The badge render — **NOT LIVE-VALIDATED. What was tried, in full.** +## 4. The badge render — **PROVEN LIVE on both surfaces, three of the four states** -**A "no access" claim must list its attempts.** These are the attempts: +### 4.0 A false diagnosis of my own, corrected here rather than buried -| # | attempt | result | -|---|---|---| -| 1 | `POST /login` to demo-hp guest `https://192.168.0.138:443`, `Host: felhom.enkisfelhom.hu`, `-k`, password from DooPlex `~/.config/credentials` `PASSWORD` (extracted with `sed`, never `cut` — the values are quoted) | **HTTP 200 with the login page and the body string `Hibás jelszó`** | -| 2 | the controller's own log, as the discriminator between "wrong host header" and "wrong password" | `auth.go:176: [WARN] [web] Failed login from 172.18.0.3` — **wrong password, not a routing problem** | -| 3 | the same password against the OTHER demo box, demo-felhom guest `https://192.168.0.149:443`, `Host: felhom.demo-felhom.eu` | **HTTP 200 + `Hibás jelszó` as well** | -| 4 | every other key in `~/.config/credentials` — `HUB_PW`, `R_DEMO-HP`, `R_DEMO-FELHOM`, `R_C11_REWALK`, `R_PART4`, `TS_KEY`, `HETZNER_API`, `ISO_S3_*` | none is a dashboard password; the `R_*` keys are escrow recovery codes | -| 5 | reading the source for an unauthenticated route to an app page | only `/claim`, `/claim/request-new-code`, `/api/health` and `/static/` are exempt (`internal/web/auth.go`) | +The first pass reported "the vaulted dashboard password is stale on both demo boxes", with the +controller's own `[WARN] Failed login` quoted as the discriminator. **The password was fine. The +extraction was wrong.** `~/.config/credentials` quotes its values with **single** quotes and the `sed` +used stripped only double quotes, so the literal `'` characters were sent as part of the password. +The operator said so, one retry with `sed "s/^['\"]//;s/['\"]$//"` returned **HTTP 302 + +`felhom_session`**, and everything below followed. -**So the vaulted `PASSWORD` is stale on BOTH demo controllers.** It was already recorded as having -drifted once (memory `demo-hp-guest-controller-access`, 2026-08-09, put back on operator instruction); -demo-hp was reinstalled and re-claimed on 2026-08-21, and demo-felhom has now drifted too. +**This is exactly the trap the `credentials-file-values-are-quoted` memory records** — it was applied +half-way. The controller's log was a true observation and a misleading one: `Failed login` proves the +BYTES did not match the hash; it says nothing about whose fault that is. **A discriminator that +separates "wrong password" from "wrong host header" does not separate "wrong password" from "wrong +password handling", and I read it as if it did.** Register row R-453 was opened on the wrong premise +and is corrected there. -**There is no operator-side route to a customer's dashboard password** — the claim code is -bcrypt-hashed hub-side and only emailed (R-119). Re-setting it means writing a new bcrypt hash into the -guest's `data/settings.json`, which is a decision about a customer account and was done under operator -instruction last time. **It is therefore a HUMAN step, by this project's own rule**, and it is not -taken here. +### 4.1 „Naprakész" — quoted from the live pages -**What IS established about the badge, and it stops short of the render:** every INPUT the badge reads -is verified live and consistent on this box — `installed_images` present with refs matching the -template's pins exactly (§2.2 vs the compose file's `lscr.io/linuxserver/bookstack:26.05.2` and -`mariadb:12.3`), and `catalog_since: "2026-07-18"` present (§3). So bookstack's inputs are the -`Naprakész` case and the seven undisturbed apps are the no-record case. **That is an inference from -verified inputs, not an observation of the rendered page, and it is not counted as evidence.** +`GET /apps/bookstack` and `GET /stacks`, session-authenticated, `Host: felhom.enkisfelhom.hu`: -The render itself is covered by unit tests that render the PRODUCTION templates -(`TestGroupD_BadgeRendersOnBothSurfaces`, both surfaces, all states, with a companion red-proof), which -is the strongest statement available without the password. +```html +Naprakész +``` + +Byte-identical to the string table in the task. **On `/stacks` it appears exactly TWICE** — the two +apps that have a record — while the **seven other deployed apps render NOTHING**. That is the +no-record case, observed live and not inferred. + +### 4.2 „Frissítés elérhető — N napja" — the age comes from `catalog_since`, live + +To reach the behind state the badge needs the template to pin something the container is not running. +Staged on **`bentopdf`** — the app with no database, no volume and no data of any kind — by editing +**only** its `docker-compose.yml` tag `v2.8.6 → v2.8.5` and touching nothing else. **The container was +never restarted and no `up -d` ran**; the file is the badge's input, and this is the same shape the +syncer produces on its own. `ScanStacks` (2-minute cadence) picked it up: + +```html +Frissítés elérhető — 52 napja +``` + +**52 napja is arithmetic on a real catalog value, not a placeholder:** bentopdf's `catalog_since` is +`2026-07-12` and the run is 2026-09-02 — 52 days. It rendered on the app page **and** the app list. + +**REVERTED IMMEDIATELY, and verified byte-identical:** `sha256` before and after both +`39679e28cdd6ee8f46e359290b4d631cf234a1cfdae374cd84c8fe7588870ed1`, `diff` empty, container still +`ghcr.io/alam00000/bentopdf:v2.8.6 Up 17 minutes (healthy)` throughout. The badge then returned to +„Naprakész" with `napja` count **0**. + +**GRADED HONESTLY: the compose file was placed by hand, not by the syncer.** What is proven is the +render — compose file → `ScanStacks` → `TemplateImages` → `updateBadge` → page — with the file +carrying exactly what a sync would have written. The syncer's own half is separately measured in +`SPIKE-app-update-2026-09-01` §3. + +### 4.3 ASCII-fragment counts, with a positive and a negative control + +Accented text is never grepped directly here. `grep -oF`, so no `.` is a wildcard — the correction the +spike had to make on itself. + +| fragment | `/stacks` (current) | `/apps/bookstack` | `/apps/docmost` (legacy) | `/stacks` (behind) | `/apps/bentopdf` (behind) | +|---|---|---|---|---|---| +| `Naprak` | **2** | **1** | **0** | 1 | 0 | +| `napja` | 0 | 0 | 0 | **1** | **1** | +| `52 napja` | 0 | 0 | 0 | **1** | **1** | +| `legfrissebb` (hover) | 2 | 1 | 0 | — | — | +| `BookStack` (positive control) | 1 | 4 | 0 | — | — | +| `Docmost` (positive control) | 1 | 0 | 5 | — | — | +| `zzz-never-present` (**negative control**) | **0** | **0** | **0** | **0** | **0** | +| `Nem-karbantartott-XYZ` (**negative control**) | **0** | **0** | **0** | — | — | + +`/apps/docmost` is the load-bearing column: a deployed app with no record renders **no badge at all**, +live. **Absent is UNKNOWN, and it is not rendered as current.** + +### 4.4 Nothing about updating changed — checked on the live page + +In the behind state, `/stacks` still carries exactly one of each for bentopdf: + +``` +stackAction(event, 'bentopdf', 'update') ×1 +stackAction(event, 'bentopdf', 'restart') ×1 +stackAction(event, 'bentopdf', 'stop') ×1 +``` + +The badge is wired to nothing. + +### 4.5 The one state NOT reachable live + +**„Frissítés elérhető" with no age** — it needs an app whose `catalog_since` is absent, malformed or +future-dated, and all 53 catalog apps now carry a valid one. Covered by `TestGroupF`, which walks +absent, blank, `tegnap`, `18/07/2026`, `2026-13-45` and a future date. ## 5. End state — nothing left broken, nothing provisioned