Correct the placement mis-framing, and file what we wrote down and never filed (R-368..R-375)
gates / gates (push) Successful in 16s
gates / gates (push) Successful in 16s
Documentation and survey only. No code, no machine contacted. THE CORRECTION. The 40 catalogue templates without a configurable path are not missing a choice: 01-topology-and-trust.md:150-152 classes each volume hot (DB/config/cache -> fast storage, ENFORCED) or bulk (media/files), and the 40 are all-hot apps. The deploy page has been saying so to the customer all along (deploy.html:624-625). SPEC-app-data-placement and R-352 are corrected in place with the framing MARKED, not deleted; every measurement stands. R-356 was re-checked and survives, strengthened - an absent HDD_PATH is the normal state, so reading it as "not installed" misreads a correct configuration. The disk claim, precisely: since R-165 there is ONE guest data volume with two binds, not two volumes (build-golden.sh:29-40, 99). A physical-disk failure losing data and first-tier copy together is REAL and is what the other tiers exist for. A full data volume stopping the OS is NOT real and was the overstated one. THE SWEEP. 113 survey-class documents examined, 14 statements of "not filed", 2 already filed. Its positive control convicted the sweep itself twice before it convicted the corpus - markdown bold broke the strongest pattern, and the reporter re-searched a truncated line - both false zeros of the exact class being hunted, and together worth 2 of the 14. THE HEADLINE. The gap the 2026-08-21 drill rediscovered WAS filed - as R-107, ROADMAP.md:122, M/READY, 2026-07-28 - and is absent from OPEN-ITEMS.md, which calls itself the single source of truth. OPEN-ITEMS and that rule both landed 2026-07-27; R-107 went to ROADMAP alone the day after. 72 ids live only in ROADMAP, 29 not done, some of them findings. Filed as R-369 (HIGH). Five more still-open gaps filed with their ages: R-371 (17d), R-372 (38d, the oldest), R-373 (20d), R-374 (14d), R-375 (4d). R-368 corrects Part 4: the storage default IS applied at deploy time via deploy.html:612 - the earlier "the deploy route never reads it" came from grepping Go and never the templates. R-370 records the process failure and is closed by the template change. PROMPT-TEMPLATE gains the two rules it lacked: name the architecture document for the area and say what it says (with a file->area map and the test "is this something we chose?"), and an enumerated gap becomes a register row in the same session - a ROADMAP row alone does not count. Ceiling R-367 -> R-375.
This commit is contained in:
@@ -0,0 +1,383 @@
|
||||
# REPORT — the database nobody backed up, and the restore that returned most apps nothing (2026-08-22)
|
||||
|
||||
**Controller v0.217.0 → v0.218.0.** Two fixes, both found by watching a machine on the night of
|
||||
2026-08-21, both confirmed the same way. **Part 1 first, because it is the only place in the product
|
||||
where one customer action causes permanent total loss.**
|
||||
|
||||
The drill that found them is preserved at
|
||||
`documentation/audits/REPORT-DRILL-backup-truth-2026-08-21.md`; its evidence is in
|
||||
`documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`.
|
||||
|
||||
---
|
||||
|
||||
## 0. Baselines, re-established — not trusted from the sheet
|
||||
|
||||
| item | expected | confirmed |
|
||||
|---|---|---|
|
||||
| controller | 0.217.0 → 0.218.0 | ✔ repo head `v0.217.0`; live on `demo-hp` `…:0.217.0` |
|
||||
| agent | 0.130.0 | ✔ repo head and `felhom-agent --version` on the box |
|
||||
| golden vouched | 0.217.0 | ✔ `<option value="0.217.0" … selected>` |
|
||||
| floor | 0.217.0 | ✔ „Effective floor: v0.217.0 — source: DB (hub_settings)" |
|
||||
| min agent | 0.129.0 | ✔ |
|
||||
| register ceiling | R-366 | ✔ (now **R-367**) |
|
||||
| `ssh hp` → 192.168.0.104 | works | ✔ |
|
||||
| clean trees, all four repos | `HEAD == origin/main`, 0 dirty | ✔ |
|
||||
|
||||
---
|
||||
|
||||
## 1. PART 1 — where the wrong name comes from
|
||||
|
||||
**It comes from a guess, made in a place where an answer was available.**
|
||||
|
||||
`deriveStackName` (`controller/internal/appbackup/dbdump.go:770-798`) is handed a container name and
|
||||
asked which app owns it. For `paperless-postgres` it strips the role suffix to `paperless`; finds
|
||||
`paperless` is **not** a deployed stack; finds the container name is not one either; finds no known
|
||||
stack is a prefix of it (`paperless-ngx` is not a prefix of `paperless-postgres`) — **and then returns
|
||||
the unresolved candidate anyway**, on its last line, silently.
|
||||
|
||||
**Every place that value is used, and what each did with it:**
|
||||
|
||||
| consumer | file:line | effect |
|
||||
|---|---|---|
|
||||
| dump directory | `backup/backup.go:482,506` | wrote to `…/primary/**paperless**/db-dumps/` on the SYSTEM drive |
|
||||
| dump filename | `appbackup/dbdump.go:200-212` | `paperless-postgres.sql`, not `paperless-ngx-postgres.sql` |
|
||||
| unit assembler | `backup/recovery_unit.go:131` | reads `…/primary/**paperless-ngx**/db-dumps` → finds nothing → `db_dumps: null` |
|
||||
| off-site collector | `backup/offbox_capture.go:32` | resolves from the unit path → the orphan is invisible to it |
|
||||
| restore's DB leg | `backup/restore_db.go:69` | `db.StackName != stackName` → never replays |
|
||||
| **safety-dump filter** | `backup/offbox_reconstitute.go:134` | `mine` empty → `hasDB=false` → **no undo copy, and the refusal is never reached** |
|
||||
| `.fab` export | `appexport/export.go:600` | same filter — the bundle carries no database either |
|
||||
| the "has a DB" flag | `web/handlers.go:1245` | the app displays as having no database |
|
||||
|
||||
**Corroboration from the code itself:** `ListDumpFiles` (`dbdump.go:562`) parses a dump filename under
|
||||
the comment *"Parse stack name and DB type from filename: `paperless-ngx-postgres.sql`"* — the reader
|
||||
was written for a name the writer never produced.
|
||||
|
||||
### Controller or catalogue? — **THE CONTROLLER. No halt.**
|
||||
|
||||
Every container the controller starts already carries `com.docker.compose.project`, read live:
|
||||
|
||||
```
|
||||
paperless-postgres paperless-ngx kimai-db kimai romm-db romm
|
||||
```
|
||||
|
||||
**And that label is the stack name BY CONSTRUCTION, not by luck:** `composeExecCustomEnv`
|
||||
(`internal/stacks/manager.go:1218-1230`) runs compose with `cmd.Dir` set to
|
||||
`/opt/docker/stacks/<stack>` and **never passes `-p`**, so compose derives the project from that
|
||||
directory. Renaming the container in the catalogue would have fixed this one app and left the guessing
|
||||
for the next one. **No catalogue change was made.**
|
||||
|
||||
---
|
||||
|
||||
## 2. PART 1.2 — the sweep, proven before its answer was trusted
|
||||
|
||||
| step | result |
|
||||
|---|---|
|
||||
| scratch copy, unplanted | `MISMATCHES: 1` (rc 1) — the known one |
|
||||
| **plant a second mismatch** (`kimai-db` → `timetrack-db`) | **`MISMATCHES: 2`**, convicted by name: `stack=kimai container=timetrack-db -> derived=timetrack` |
|
||||
| plant removed | back to `1` |
|
||||
| **every mismatch removed** | **`MISMATCHES: 0`, rc 0** — "1" is not a stuck value |
|
||||
| scratch discarded | `kimai/docker-compose.yml` md5-identical to the real catalogue |
|
||||
|
||||
**The real count: 1 affected app of 53**, from 15 DB containers checked — `paperless-ngx`.
|
||||
|
||||
---
|
||||
|
||||
## 3. PART 1 — the dumps already written under the wrong name
|
||||
|
||||
**They stay, untouched, and nothing will ever collect them.**
|
||||
`/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql`, 312 381 B,
|
||||
last written 07:38 UTC by the final pre-fix cycle — **verified still present after the fix.**
|
||||
|
||||
Nothing deletes them by design: the F5 stale-primary prune (`backup/backup.go:1248`) skips any
|
||||
directory that is not a deployed app, under the guard *"an undeployed app's last backup is still its
|
||||
restore point"*.
|
||||
|
||||
**They can be adopted, by hand:** move the file to
|
||||
`…/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql` and it becomes a readable restore point.
|
||||
**Deliberately not automatic** — it would sit beside volume tars from a different time (an incoherent
|
||||
pair, the shape the R-43/R-44 stamp exists to surface), and a controller that silently relocates a
|
||||
customer's data on upgrade is a migration, not a fix. **Filed as R-367.**
|
||||
|
||||
---
|
||||
|
||||
## 4. PART 2 — the mechanism, and the comment corrected
|
||||
|
||||
`ReconstituteFromOffsite` skipped every placement flagged as the unit (`offbox_reconstitute.go:344`),
|
||||
and the volume archives live **inside** the unit. A search of the whole off-site path found no
|
||||
volume-restore call; the only two callers of `restoreDockerVolumes` were `restore.go:67` and
|
||||
`restore_unit.go:255`, both local.
|
||||
|
||||
**The comment as it stood:**
|
||||
|
||||
> *"The live recovery unit is still never overwritten — it is the LOCAL restore path's source and
|
||||
> clobbering it would trade one recovery route for another. The snapshot's dump is replayed from the
|
||||
> scratch unit instead, so nothing is lost by skipping it."*
|
||||
|
||||
**Which half still holds: the FIRST.** The live unit is the local restore path's own source and must
|
||||
never be clobbered — that is why the skip stays, and scenario D now fingerprints the whole live unit
|
||||
across the operation to keep it true.
|
||||
|
||||
**The second half was false and is corrected, not left.** True of the *database* dump, false of the
|
||||
*volume* archives, which live in the same unit and were replayed by nothing at all. Skipping the
|
||||
placement is correct; treating the skip as harmless was not — and that sentence is exactly why a
|
||||
reader would not look.
|
||||
|
||||
**The fix:** `restoreDockerVolumesFrom(stackName, dumpDir) (int, error)` — the local path's own replay
|
||||
with an explicit directory, the same shape `reimportDBDumpsFrom` already had beside `reimportDBDumps`.
|
||||
**One implementation, two callers.** Volumes replay **before** the database (a logical dump must still
|
||||
win over a volume-tar copy of the same database) and **inside the stopped window** (Docker will not
|
||||
replace a volume a container holds).
|
||||
|
||||
---
|
||||
|
||||
## 5. THE CUSTOMER'S MESSAGE, BEFORE AND AFTER, VERBATIM
|
||||
|
||||
**R-354 — calibre-web, identical fixture, identical steps, both on `demo-hp`:**
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **before** (0.217.0, 09:41) | `ok=true` — „A(z) calibre-web: **0 fájl visszaállítva** (mentés: 2026-08-22 09:38) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and `ls: /vol/R354-2026-08-22: No such file or directory` |
|
||||
| **after** (0.218.0, 09:56) | `ok=true` — „A(z) calibre-web: **0 fájl és 1 adatkötet visszaállítva** (mentés: 2026-08-22 09:52) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and **5/5 files byte-identical** |
|
||||
|
||||
**R-355 — paperless-ngx:**
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **before** (2026-08-21) | „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**" over a live 72-table PostgreSQL, with **no undo copy taken at all** |
|
||||
| **after** (0.218.0, 09:57) | „A(z) paperless-ngx: **0 fájl és 3 adatkötet és az adatbázis visszaállítva** (mentés: 2026-08-22 09:52) — az alkalmazás újraindult." |
|
||||
|
||||
**A third sentence now exists for the case that had no honest wording** — an app that HAS a database
|
||||
whose snapshot carried no dump: „… FIGYELEM: ennek az alkalmazásnak **VAN adatbázisa**, de a mentés nem
|
||||
tartalmazott adatbázis-mentést, ezért az adatbázis **NEM állt vissza**. A visszaállítás előtti állapot
|
||||
mentése megvan: `<undo>`". That case previously printed the same confident „nincs adatbázisa".
|
||||
|
||||
---
|
||||
|
||||
## 6. THE HARDWARE WALK
|
||||
|
||||
**Method: endpoint-level — the exact endpoints the dashboard's JS calls (`/api/backup/run`,
|
||||
`/backup/offbox/run`, `/backup/offbox/restore`, `/backup/offbox/reconstitute`), with no shell inside
|
||||
the machine for any step of the walk.** Planting the fixture and reading the result back used a shell
|
||||
and is setup/verification, not the walk. No browser exists on DooPlex.
|
||||
|
||||
**The fixture, hashed before anything ran:**
|
||||
|
||||
```
|
||||
7708bf6582f990ee5c98e3fb6638214d4916b517ea80638baf20b0662b56506e SENTINEL.txt
|
||||
fb67d42bdfe09815922c2f4f4a2086025bd1a074757d17b5e5164d29b1a5e8d5 binary-512k.bin
|
||||
847fbae0868fc9a2c985d7411fa545c4853e8561f7b8cc8dd485437b1f1d76c4 nested/őszibarack.md
|
||||
f15b6d0f88e5499836f585403e3b59b4f2006a0e6a96ce4ec598dd7abe38bff5 plain.txt
|
||||
857e8594005d11f9802f61018603b7746e7af37ac02631aaef2db57f9700fb58 árvíztűrő-tükörfúrógép.txt
|
||||
|
||||
accented names as RAW BYTES (UTF-8 NFC):
|
||||
árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
|
||||
őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64
|
||||
```
|
||||
|
||||
**The comparator was proved able to convict first:** one byte flipped at offset 300 000 of
|
||||
`binary-512k.bin` (`3b` → `00`) → `binary-512k.bin: FAILED`, rc 1, the other four `OK`; the unmodified
|
||||
set rc 0. Mutant discarded.
|
||||
|
||||
### Results
|
||||
|
||||
| scenario | result |
|
||||
|---|---|
|
||||
| **1-A** the dump is inside the unit, and inside the off-site snapshot | **PASS.** 0.217.0 09:37: `db_dumps = None`, dump refreshed into the phantom dir. 0.218.0 09:45: `db_dumps = ['paperless-ngx-postgres.sql']` (312 669 B) in the app's own unit, and `restic ls` shows it inside the snapshot **for the first time** |
|
||||
| **1-B** an undo copy taken and verified before anything stops; the message names the database | **PASS.** Before: `find /mnt -name "pre-restore-*paperless*"` → **0**. After: `pre-restore-20260822T075658Z-paperless-ngx-postgres.sql`, 312 957 B, in the app's own unit |
|
||||
| **1-C** the undo cannot be taken → the whole restore refuses, nothing changed | **PASS**, on the app that could never reach this guard before. `ok=false`; the marker kept its mutation (`51d37b0f…`) and `StartedAt` was unchanged — **the app was never stopped** |
|
||||
| **1-D** every other app byte-identical | **PASS.** `romm-mariadb.sql`, `kimai-mariadb.sql` unchanged in name and location; the only phantom directory is the one pre-existing orphan; no new one created |
|
||||
| **2-A/B** the volume comes back and the message says so | **PASS.** calibre-web 5/5 byte-identical incl. both accented names; paperless-ngx 3 volumes + the database |
|
||||
| **2-C** a snapshot with no volume archives is unchanged | **PASS** — the unit test asserts the byte-identical sentence; live, romm/kimai unaffected |
|
||||
| **2-D** the live recovery unit is never written | **PASS** — fingerprinted across the whole operation (see red-proof 6) |
|
||||
| **2-E** a failed replay is reported as a failure naming the volume | **PASS** (unit test; the live path returns the same error) |
|
||||
|
||||
All 15 containers healthy after the walk.
|
||||
|
||||
---
|
||||
|
||||
## 7. RED-PROOFS — seven, each mutation asserted applied and reverted
|
||||
|
||||
| # | mutation | outcome |
|
||||
|---|---|---|
|
||||
| 1 | `resolveStackName` ignores the compose label | **FAIL as required.** `= "paperless", want "paperless-ngx"`, and the consequence printed both paths — written to `…/primary/paperless/db-dumps` vs read `…/primary/paperless-ngx/db-dumps`. **The empty database record, returned** |
|
||||
| 2 | outcome message back to the counter-only predicate | **FAIL as required**, reproducing the 21 August sentence verbatim: `"A(z) paperless-ngx: 0 fájl visszaállítva — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."` |
|
||||
| 3 | the fail-closed refusal removed from `writeSafetyDump` | **FAIL as required** — *a restore was seen proceeding with no undo copy* |
|
||||
| 4 | the volume leg removed | **FAIL as required.** `VolumesReplayed = 0` and the replay never called — last night's silent loss, returned |
|
||||
| 5 | the volume count dropped from the message | **FAIL as required**, reproducing „A(z) calibre-web: 5 fájl visszaállítva …" verbatim |
|
||||
| 6 | the live-unit guard removed | **PASSED FIRST — a defect in MY TEST, not the code.** The fingerprint had been narrowed to the volume directory and was blind to a placement writing into the unit **root**. Widened to the whole unit (excluding only the documented undo copies) it convicts: `PLACED:17603ba1…` appears. **The R-181 class, reproduced inside its own regression test — and the reason the red-proof is mandatory** |
|
||||
| 7 | the volume-replay error swallowed | **FAIL as required** — a partial replay reporting success |
|
||||
|
||||
**Green gate:** `go build ./... && go vet ./... && go test ./...` — **28 packages ok, 0 FAIL lines.**
|
||||
Controller gates: all 11 OK. felhom.eu gates: all OK except the golden-currency gate, which was red
|
||||
until the bake (§8) — correctly, and never bypassed.
|
||||
|
||||
*(Instrument note: `go test ./... | grep -vE '^ok'; echo rc=$?` reports the **grep's** exit code, which
|
||||
is 1 when every test passed and nothing was left to print. The verdict above is from an explicit
|
||||
`FAIL`-line count, not from that.)*
|
||||
|
||||
---
|
||||
|
||||
## 8. WHAT THIS DOES **NOT** FIX — the blocker for the apps that need it most
|
||||
|
||||
**The off-site restore still refuses outright for the 40 of 53 apps that declare no data drive
|
||||
(R-356)**, telling the customer that a running app „nincs telepítve" and to reinstall it "to the same
|
||||
place" — which those apps give them no way to choose. Those are exactly the apps whose entire dataset
|
||||
is a named volume, **so R-354's fix cannot reach them until R-356 is closed.**
|
||||
|
||||
That is why the live confirmation used `calibre-web` and `paperless-ngx`: they declare a drive and can
|
||||
actually reach the restore. R-356 is out of scope by the task's own list, and it is now the first thing
|
||||
worth doing.
|
||||
|
||||
---
|
||||
|
||||
## 9. REGISTER
|
||||
|
||||
- **R-354 — CLOSED**, shipped + proven-live (v0.218.0).
|
||||
- **R-355 — CLOSED**, shipped + proven-live (v0.218.0).
|
||||
- **R-367 — OPENED (LOW)**: dumps already written under the wrong name are stranded; adoptable by hand,
|
||||
deliberately not automatic, never on an upgrade path.
|
||||
- **Ceiling moved: R-366 → R-367.**
|
||||
|
||||
---
|
||||
|
||||
## 10. DELIBERATELY OUT OF SCOPE — so it does not read as forgotten
|
||||
|
||||
R-353 (the empty restore reported as success — real, and next), R-352 (where the 40 apps' data lives,
|
||||
awaiting a ruling), R-356 (§8), deploy-and-restore as one act, R-360 (the delete guard on verification
|
||||
copies), R-364 (the accented-search instrument defect), and the remaining drill rows R-357–R-359,
|
||||
R-361–R-363, R-365, R-366.
|
||||
|
||||
---
|
||||
|
||||
## 11. COMMITS, VERSIONS, CI
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| `felhom-controller` | `5ce3a44` the fix + tests; `2da259a` README + CONTEXT |
|
||||
| image | `gitea.dooplex.hu/admin/felhom-controller:0.218.0` (150M) |
|
||||
| deployed | `demo-hp` guest 9201 — `…:0.218.0 Up (healthy)` |
|
||||
| `demo-felhom` | **untouched this session**, still 0.217.0 |
|
||||
|
||||
---
|
||||
|
||||
## 12. BAKE
|
||||
|
||||
`RUNBOOK-manual-build.md` §4.0/§4.1, in the drill VM. **The golden-currency gate went red the moment
|
||||
v0.218.0 was released and stayed red until the bake — that is correct, and it was never bypassed.**
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| `GOLDEN_VERSION` | **0.218.0** |
|
||||
| `GOLDEN_SHA256` | **8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b** |
|
||||
| package | `…/generic/felhom-golden/0.218.0/golden.tar.zst`, 657 026 013 B |
|
||||
| template | `debian-13-standard_13.6-1_amd64.tar.zst` — listed fresh, not assumed |
|
||||
|
||||
**Pass markers, each checked, with the negative controls:**
|
||||
|
||||
```
|
||||
docker OK (overlay2 : 1 -> " docker OK (overlay2; data-root /var/lib/docker)"
|
||||
including mount point : 2 -> rootfs ('/') and mp0 ('/var/lib/felhom') [there is no mp1]
|
||||
upload OK (HTTP 201) : 1 -> pre-delete HTTP 404 (404/204 expected)
|
||||
excluding : 0 <- negative control
|
||||
FATAL : 0 <- negative control
|
||||
```
|
||||
|
||||
**Token hygiene.** Copied file→file and read by a runner script inside the VM, so it never crossed a
|
||||
shell or a unit property: `systemctl show golden-bake … | grep -c -F "$(cat …)"` → **0**. **The leak
|
||||
grep on the committed log was itself proven before its zero was believed** — token appended to a
|
||||
throwaway copy, grepped (**1**), copy shredded, then the committed log's **0** accepted.
|
||||
|
||||
**Teardown:** `pct destroy 9100 --purge`; token, runner, build script and in-VM log shredded **after**
|
||||
the log was copied out; VM powered off; qemu exited; **`qemu-img snapshot -a virgin` restored.**
|
||||
|
||||
Evidence: `documentation/tests/golden-0.218.0-2026-08-22/`.
|
||||
|
||||
---
|
||||
|
||||
## 13. STOP — THE APPROVAL
|
||||
|
||||
**Nothing below has been changed. Vouching and the floor are yours.**
|
||||
|
||||
### The three Day-0 values, each with BOTH checks
|
||||
|
||||
| field | set to | downloadable | selectable |
|
||||
|---|---|---|---|
|
||||
| **`golden_version`** | **0.218.0** — sha `8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b` | ✔ HTTP 200, 657 026 013 B, and the **downloaded bytes** hash to exactly the bake's reported sha | ✔ offered in the hub dropdown with `data-sha=8e427869d13eafb7…` |
|
||||
| **`agent_version`** | **0.130.0** — sha `a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3` | ✔ HTTP 200, 14 141 158 B, downloaded sha matches the hub's `data-sha` exactly | ✔ already the SELECTED option — **no change needed** |
|
||||
| **`min_agent`** | **0.129.0** | — | ✔ already set to 0.129.0 — **no change needed** |
|
||||
|
||||
**So the vouch is a ONE-field change this time, and that is the safe direction, not a shortcut.** The
|
||||
three-field rule exists so a golden is never shipped onto an agent older than it needs. The controller
|
||||
being baked declares **`MinAgent: 0.129.0` (unchanged)**, the vouched agent is already **0.130.0**,
|
||||
and `0.130.0 ≥ 0.129.0` — so `agent_version` and `min_agent` are already correct and only
|
||||
`golden_version` moves. Setting `min_agent` above the vouched agent is the R-216 shape; it is not
|
||||
happening here.
|
||||
|
||||
**If the hub refuses the save, the guard is working.** The R-120 gate on that POST
|
||||
(`hub/internal/web/configs.go:1165`) refuses a golden older than the newest controller the fleet
|
||||
reports. 0.218.0 is the newest, so it should pass — and if it does not, fix the cause, never the guard.
|
||||
|
||||
**Vouching is reversible:** re-select 0.217.0 and Save. The old package is never deleted by a bake.
|
||||
|
||||
### The one line on the floor
|
||||
|
||||
**Raising the floor to 0.218.0 is what puts this on both machines — and yes, you will want to.**
|
||||
`demo-hp` already runs 0.218.0 (deployed directly for the live proof). **`demo-felhom` is still on
|
||||
0.217.0 and will not move until the floor is raised**, so today it still has both defects. I did not
|
||||
raise it; that is your call, as is vouching.
|
||||
|
||||
---
|
||||
|
||||
## 14. WHAT WAS DROPPED, AND OBSERVATIONS
|
||||
|
||||
**Dropped: nothing.** Both parts completed in the required order, with the live walk and all seven
|
||||
red-proofs. The session did not run short.
|
||||
|
||||
**Observations, noticed and not acted on:**
|
||||
|
||||
- **`demo-felhom` was not touched** and still runs 0.217.0 — so the fleet is deliberately non-uniform
|
||||
until the floor moves. Its two preserved fixtures were not involved in any step of this session.
|
||||
- **The `.fab` export has the same R-355 blind spot** and is fixed by the same change — `export.go:600`
|
||||
filters on the same `db.StackName`. Not separately verified live; the unit tests cover the resolver
|
||||
that feeds it.
|
||||
- **`paperless-ngx`'s orphan directory is now a permanent fixture on `demo-hp`** until R-367 is
|
||||
decided. It is the only reproduction of the pre-fix state left anywhere, which is an argument for
|
||||
leaving it alone for now.
|
||||
- **The safety dump still overwrites the unit's own DB dump before renaming it** (R-361, out of scope)
|
||||
— visible in this session's own evidence, where `romm`'s `db-dumps/` holds both a `pre-restore-*` and
|
||||
the regular dump.
|
||||
|
||||
---
|
||||
|
||||
## 15. THE PICTURE — `where-felhom-stands` brought up to date (same day, after the fixes)
|
||||
|
||||
Regenerated from `where-felhom-stands.yaml` (the HTML is a build product; the YAML is the source).
|
||||
**Eight claims re-checked, three moved — all downward.**
|
||||
|
||||
| claim | was | now | why |
|
||||
|---|---|---|---|
|
||||
| `backup.offsite` | walked | **partial** | "18 snapshots, daily, unbroken" was true on 2026-08-09 and false by 2026-08-21 — the next snapshot after it was put there by hand, twelve days later, and nothing reported the gap |
|
||||
| `fail.wiped-reinstalled.data` | walked | **partial** | a real reinstall orphans BOTH off-premises tiers: restic silently for 12 days (R-193), and the pre-reinstall PBS archives cannot be opened by the rebuilt box at all (R-366) |
|
||||
| `backup.fill-warning` | walked | **partial** | the warning fires correctly, but once a day — a filesystem filling at 03:31 is unannounced for ~24 h, and it was watched staying silent while a volume sat at 99% |
|
||||
|
||||
**Held, with their notes rewritten:** `backup.tier1` and `recover.byte-identical` now carry the
|
||||
R-355/R-354 story *and* their fixes; `backup.restore-proof` stays grey for a sharper reason (its
|
||||
archives are orphaned, not merely untested); `backup.sikeres` gained two fresh instances;
|
||||
`fail.customer-self-restore` records that R-356 blocks 40 of 53 apps regardless of who is driving.
|
||||
|
||||
**`unproven.py --summary` moved, and the checklist asks that it be said:**
|
||||
**NOT WALKED went from 32 of 55 to 35 of 55.** That is the three downgrades above. Nothing was raised
|
||||
— the dataset's own rule forbids raising a status here, and nothing needed it.
|
||||
|
||||
### A defect found in the page's own toolchain
|
||||
|
||||
**The rendered header cited the wrong commits and would have gone on doing so.** `render_stands.py`
|
||||
parsed `verified_on` from the YAML but had the three commit shas **hardcoded**, so the page printed
|
||||
the 2026-08-09 commits while the dataset said 2026-08-22. That is precisely the stale-build-product
|
||||
failure the renderer's own docstring says it exists to prevent. Parsed now, literals gone from both
|
||||
the script and the output, checked. The count beside it also read *"15 status(es) moved in that pass"*
|
||||
when 15 was every recorded move ever accumulated; it now separates the two numbers.
|
||||
|
||||
**`check_stands.py` was proven able to convict before its OK was accepted:** a claim marked `missing`
|
||||
flipped to `walked` in a scratch copy fired rule 5 by name (`use.dlna: status 'walked' but NO evidence
|
||||
document cited`), and the real file still passed. Gates all OK; CI **id=380 / run_number=248**, green.
|
||||
Reference in New Issue
Block a user