Correct the placement mis-framing, and file what we wrote down and never filed (R-368..R-375)
gates / gates (push) Successful in 16s
gates / gates (push) Successful in 16s
Documentation and survey only. No code, no machine contacted. THE CORRECTION. The 40 catalogue templates without a configurable path are not missing a choice: 01-topology-and-trust.md:150-152 classes each volume hot (DB/config/cache -> fast storage, ENFORCED) or bulk (media/files), and the 40 are all-hot apps. The deploy page has been saying so to the customer all along (deploy.html:624-625). SPEC-app-data-placement and R-352 are corrected in place with the framing MARKED, not deleted; every measurement stands. R-356 was re-checked and survives, strengthened - an absent HDD_PATH is the normal state, so reading it as "not installed" misreads a correct configuration. The disk claim, precisely: since R-165 there is ONE guest data volume with two binds, not two volumes (build-golden.sh:29-40, 99). A physical-disk failure losing data and first-tier copy together is REAL and is what the other tiers exist for. A full data volume stopping the OS is NOT real and was the overstated one. THE SWEEP. 113 survey-class documents examined, 14 statements of "not filed", 2 already filed. Its positive control convicted the sweep itself twice before it convicted the corpus - markdown bold broke the strongest pattern, and the reporter re-searched a truncated line - both false zeros of the exact class being hunted, and together worth 2 of the 14. THE HEADLINE. The gap the 2026-08-21 drill rediscovered WAS filed - as R-107, ROADMAP.md:122, M/READY, 2026-07-28 - and is absent from OPEN-ITEMS.md, which calls itself the single source of truth. OPEN-ITEMS and that rule both landed 2026-07-27; R-107 went to ROADMAP alone the day after. 72 ids live only in ROADMAP, 29 not done, some of them findings. Filed as R-369 (HIGH). Five more still-open gaps filed with their ages: R-371 (17d), R-372 (38d, the oldest), R-373 (20d), R-374 (14d), R-375 (4d). R-368 corrects Part 4: the storage default IS applied at deploy time via deploy.html:612 - the earlier "the deploy route never reads it" came from grepping Go and never the templates. R-370 records the process failure and is closed by the template change. PROMPT-TEMPLATE gains the two rules it lacked: name the architecture document for the area and say what it says (with a file->area map and the test "is this something we chose?"), and an enumerated gap becomes a register row in the same session - a ROADMAP row alone does not count. Ceiling R-367 -> R-375.
This commit is contained in:
+29
@@ -14,6 +14,35 @@
|
||||
> language, one screen, no identifiers in the prose. Same subjects, different readers; merging them
|
||||
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
## Read the architecture folder BEFORE calling something a defect (2026-08-22, R-370 / R-369)
|
||||
|
||||
**Between 19 and 22 August the reviewing side called a documented architectural decision a defect, in
|
||||
four places.** The decision is in `architecture/01-topology-and-trust.md:150-152`: `.felhom.yml`
|
||||
classes each volume **hot** (DB/config/cache → fast storage, **enforced**) or **bulk** (media/files →
|
||||
may be slow). The 40 catalogue templates with no `HDD_PATH` are all-hot apps — their data is app
|
||||
internals and belongs inside the guest, on its own volume so it cannot fill the OS
|
||||
(`felhom-agent/configs/build-golden.sh:29-40`). The 13 with a path are the bulk apps. **There is no
|
||||
choice being denied**, and the deploy page has been telling the customer exactly this all along:
|
||||
*„A kiválasztott meghajtón az alkalmazás fájljai (média, dokumentumok) tárolódnak. Az adatbázis a
|
||||
gyors belső SSD-n fut"* (`deploy.html:624-625`).
|
||||
|
||||
**Why it happened, mechanically: the register and live source were read; `documentation/architecture/`
|
||||
was not.** Nothing in `PROMPT-TEMPLATE.md` §4 required naming the architecture document for the area
|
||||
being touched — item 4 named only the controller module map, and the S-1 rule governs *updating* a
|
||||
design doc at the end, not *reading* one at the start. **Fixed in the same session:** §4 now carries
|
||||
the file→area map, the ordering *architecture = reasoning, register = work, source = truth*, and the
|
||||
test **"is what I am about to call a defect something we chose?"**
|
||||
|
||||
**Corrected in place, not rewritten** — `SPEC-app-data-placement-2026-08-21.md` and R-352 carry
|
||||
`[CORRECTED 2026-08-22]` blocks that say what they got wrong; a document that quietly changes its mind
|
||||
teaches nobody.
|
||||
|
||||
**And the sibling failure, R-369: there are two registers.** `OPEN-ITEMS.md` calls itself the single
|
||||
source of truth; `ROADMAP.md` also holds open work. 72 `R-` ids live only in ROADMAP, 29 of them not
|
||||
marked done. **R-107 — the off-site restore never unpacking its own volume tars — was filed there on
|
||||
2026-07-28, one day after OPEN-ITEMS was created, and never entered the register. It was rediscovered
|
||||
from scratch 25 days later by the 2026-08-21 drill and shipped as R-354.** The template now says
|
||||
plainly that a ROADMAP row alone does not satisfy the filing rule.
|
||||
|
||||
## Box REACHABILITY is a separate signal from box FILL — and the cadence difference is deliberate (2026-08-18, R-339, hub v0.106.0)
|
||||
|
||||
|
||||
@@ -1,383 +1,307 @@
|
||||
# REPORT — the database nobody backed up, and the restore that returned most apps nothing (2026-08-22)
|
||||
# REPORT — correcting what we mis-called a defect, and finding what we wrote down and never filed (2026-08-22)
|
||||
|
||||
**Controller v0.217.0 → v0.218.0.** Two fixes, both found by watching a machine on the night of
|
||||
2026-08-21, both confirmed the same way. **Part 1 first, because it is the only place in the product
|
||||
where one customer action causes permanent total loss.**
|
||||
|
||||
The drill that found them is preserved at
|
||||
`documentation/audits/REPORT-DRILL-backup-truth-2026-08-21.md`; its evidence is in
|
||||
`documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`.
|
||||
**Documentation and survey only. No code, no version bump, no bake, no deploy, no machine contacted.**
|
||||
The previous report is preserved at `documentation/audits/REPORT-v0.218.0-r354-r355-2026-08-22.md`.
|
||||
|
||||
---
|
||||
|
||||
## 0. Baselines, re-established — not trusted from the sheet
|
||||
## 1. §2's four claims — ALL FOUR HELD. No halt.
|
||||
|
||||
| item | expected | confirmed |
|
||||
| # | claim | verdict |
|
||||
|---|---|---|
|
||||
| controller | 0.217.0 → 0.218.0 | ✔ repo head `v0.217.0`; live on `demo-hp` `…:0.217.0` |
|
||||
| agent | 0.130.0 | ✔ repo head and `felhom-agent --version` on the box |
|
||||
| golden vouched | 0.217.0 | ✔ `<option value="0.217.0" … selected>` |
|
||||
| floor | 0.217.0 | ✔ „Effective floor: v0.217.0 — source: DB (hub_settings)" |
|
||||
| min agent | 0.129.0 | ✔ |
|
||||
| register ceiling | R-366 | ✔ (now **R-367**) |
|
||||
| `ssh hp` → 192.168.0.104 | works | ✔ |
|
||||
| clean trees, all four repos | `HEAD == origin/main`, 0 dirty | ✔ |
|
||||
| 1 | `07-backup-architecture.md:297` — 53 templates, 52 named volumes, exactly 13 with a `backup:` block, and those 13 are exactly the ones that bind a configurable path | **HOLDS**, at `:296-299`. The heading above it reads *"Coverage per app class — and an unresolved count"*, which the sweep then followed |
|
||||
| 2 | `_recovery-inventory-2026-07-28.md` Tier-3 row — *"…unpacked by no offsite action… no single action does it and no UI routes it. This is my own enumeration; it is not currently filed as a finding."* | **HOLDS**, verbatim, at `:373` |
|
||||
| 3 | `:392-400` — `rootfs`, `mp0 /var/lib/docker`, `mp1 /mnt/sys_drive` as separate volumes | **HOLDS**, at `:392-396` — **but it describes a layout no current box has.** See below |
|
||||
| 4 | `00-capability-map.md:94` — one data volume; a local backup bounded by free space | **HOLDS**, at `:94` |
|
||||
|
||||
**Claims 3 and 4 describe different layouts, six days apart, and 4 is the live one.** The inventory
|
||||
(2026-07-28) records the pre-R-165 split. `felhom-agent/configs/build-golden.sh:29-40` — live source —
|
||||
bakes **ONE** data volume at `/var/lib/felhom`, with `/var/lib/docker` and `/mnt/sys_drive` as two
|
||||
**binds of that same volume**, and states at `:99` that *"There is deliberately no mp1"*.
|
||||
|
||||
**So the prompt's own §1.2 instruction is out of date:** it asked me to say that named volumes and
|
||||
first-tier backups are *"on two different guest volumes on one physical disk"*. They are on **one**
|
||||
guest volume, by two binds. The disk claim is stated precisely in §4 below.
|
||||
|
||||
**A correction to my own §2 working.** In verifying claim 2 I reported that a literal grep matched
|
||||
`:373`. It did not — it returned nothing, and what printed was the second grep in the same cell.
|
||||
`**not** currently filed` does not match `not currently filed`, because of the markdown bold. The
|
||||
claim still holds (I read the line), but the check I quoted was wrong — and that false zero turned out
|
||||
to matter, see §3.
|
||||
|
||||
---
|
||||
|
||||
## 1. PART 1 — where the wrong name comes from
|
||||
## 2. The sweep — three counts, and the control that earned them
|
||||
|
||||
**It comes from a guess, made in a place where an answer was available.**
|
||||
**Method** (`scripts` not committed; it is a throwaway, and the terms below are the durable part):
|
||||
enumerate every **survey-class** document — inventory / audit / spike / review / diagnosis / recon /
|
||||
campaign / report / validation / rehearsal / walk / finding / spec / triage — under
|
||||
`documentation/{audits,architecture,pilot,backlog}`, excluding runbooks, templates, READMEs,
|
||||
CHANGELOGs and ROADMAP, because those instruct rather than conclude. For each, ask whether any
|
||||
`OPEN-ITEMS.md` row cites it by filename, then search every line for the **shapes** an unfiled gap
|
||||
takes.
|
||||
|
||||
`deriveStackName` (`controller/internal/appbackup/dbdump.go:770-798`) is handed a container name and
|
||||
asked which app owns it. For `paperless-postgres` it strips the role suffix to `paperless`; finds
|
||||
`paperless` is **not** a deployed stack; finds the container name is not one either; finds no known
|
||||
stack is a prefix of it (`paperless-ngx` is not a prefix of `paperless-postgres`) — **and then returns
|
||||
the unresolved candidate anyway**, on its last line, silently.
|
||||
| count | value |
|
||||
|---|---|
|
||||
| **documents examined** | **113** (survey-class; 114 after this session added one) |
|
||||
| **gaps found** (documents saying, in their own words, that something was not filed) | **14** statements |
|
||||
| **of those, already filed** | **2** — and the other 12 break down in §3 |
|
||||
|
||||
**Every place that value is used, and what each did with it:**
|
||||
**The search terms, so the next person can widen them:**
|
||||
|
||||
| consumer | file:line | effect |
|
||||
```
|
||||
not (currently )?filed never filed no finding not a finding
|
||||
unresolved unfiled disagreement no single action
|
||||
nothing (does|routes|reads|consults) never been should be
|
||||
ought to is not (currently )?(tested|proven|observed|covered|wired)
|
||||
has (never|not) been (seen|observed|walked|proven|done)
|
||||
no (test|row|register) (pins|covers|cites) TODO FIXME ⚠ not covered gap
|
||||
```
|
||||
|
||||
Every word-gap in the first four patterns is `[*_`~ ]*`, i.e. **markdown emphasis is tolerated** —
|
||||
see §3.
|
||||
|
||||
### The positive control — and it convicted the sweep twice before it convicted the corpus
|
||||
|
||||
**Plant → find → remove → fail to find**, on a scratch copy of the whole `documentation/` tree:
|
||||
|
||||
| step | strongest-shape hits |
|
||||
|---|---|
|
||||
| unplanted scratch copy | **14** |
|
||||
| a synthetic gap planted, **bolded and on one line** | **15** — convicted by name and quoted back |
|
||||
| scratch discarded, real corpus re-run | **14** |
|
||||
|
||||
**The control found two defects in my own sweep before it found anything in the corpus, and both were
|
||||
false zeros of exactly the class this session is about:**
|
||||
|
||||
1. **Markdown emphasis broke the strongest pattern.** The single most important line in the corpus is
|
||||
written `it is **not** currently filed as a finding`, and `not (currently )?filed` does not match
|
||||
it. Fixed by tolerating `[*_`~ ]*` between words.
|
||||
2. **The reporter re-searched a truncated copy of the line.** Hits were stored as `line[:200]`; that
|
||||
line is a long table row and the phrase sits past column 200, so the detector caught it and the
|
||||
**report** dropped it. Fixed by matching the full line and truncating only for display. This alone
|
||||
moved the count **12 → 14**.
|
||||
|
||||
A sweep whose first two runs would have under-reported by 2 is exactly why the control is mandatory.
|
||||
|
||||
---
|
||||
|
||||
## 3. What the 14 were, judged
|
||||
|
||||
| verdict | n | which |
|
||||
|---|---|---|
|
||||
| dump directory | `backup/backup.go:482,506` | wrote to `…/primary/**paperless**/db-dumps/` on the SYSTEM drive |
|
||||
| dump filename | `appbackup/dbdump.go:200-212` | `paperless-postgres.sql`, not `paperless-ngx-postgres.sql` |
|
||||
| unit assembler | `backup/recovery_unit.go:131` | reads `…/primary/**paperless-ngx**/db-dumps` → finds nothing → `db_dumps: null` |
|
||||
| off-site collector | `backup/offbox_capture.go:32` | resolves from the unit path → the orphan is invisible to it |
|
||||
| restore's DB leg | `backup/restore_db.go:69` | `db.StackName != stackName` → never replays |
|
||||
| **safety-dump filter** | `backup/offbox_reconstitute.go:134` | `mine` empty → `hasDB=false` → **no undo copy, and the refusal is never reached** |
|
||||
| `.fab` export | `appexport/export.go:600` | same filter — the bundle carries no database either |
|
||||
| the "has a DB" flag | `web/handlers.go:1245` | the app displays as having no database |
|
||||
| **never a gap** — a reasoned "no finding" | 4 | `_design-review.md:9` (nothing deferred), `CAMPAIGN-11:501` (a measurement was honest), and the two **explicit `### Not filed` sections** in `CAMPAIGN-10-two-storage-soak` and `SPIKE-recovery-unit-space` — each item disposed with a reason. **This is good practice, not a failure, and it is what the rest should look like.** |
|
||||
| **already filed** | 2 | `REPORT-DRILL-backup-truth:321` (its Part 5 became R-360 and R-364, filed the same session) and `REPORT-DRILL:569` (the hub's controller-version blindness — the register already covers it, 2 hits; not re-filed) |
|
||||
| **still open → NEWLY FILED** | 5 | R-371, R-372, R-373, R-374, R-375 |
|
||||
| **the headline** | 3 | `_recovery-inventory:373`, `:1043` and `07-backup-architecture.md:337` — all three the same gap, and it turns out it **was** given a number |
|
||||
|
||||
**Corroboration from the code itself:** `ListDumpFiles` (`dbdump.go:562`) parses a dump filename under
|
||||
the comment *"Parse stack name and DB type from filename: `paperless-ngx-postgres.sql`"* — the reader
|
||||
was written for a name the writer never produced.
|
||||
### The headline: it was filed. In the other register.
|
||||
|
||||
### Controller or catalogue? — **THE CONTROLLER. No halt.**
|
||||
The gap the 2026-08-21 drill rediscovered **was already numbered R-107**, with a full write-up:
|
||||
|
||||
Every container the controller starts already carries `com.docker.compose.project`, read live:
|
||||
> `ROADMAP.md:122` — *"**No offsite action unpacks the named-volume tars Tier-3 captures on every
|
||||
> run.** `ReconstituteFromOffsite` skips the unit outright … `PlaceOffsiteRestore` places it only when
|
||||
> the live unit is ABSENT"* — **M, READY, 2026-07-28**
|
||||
|
||||
and cross-referenced twice in `07-backup-architecture.md` (`:337`, `:902`). **It is absent from
|
||||
`OPEN-ITEMS.md`**, which opens with the words *"the single source of truth for open work"*.
|
||||
|
||||
**The dating makes it a rule violation, not a gap in the rules.** `OPEN-ITEMS.md` was created and the
|
||||
template gained its *"single source of truth"* bullet on **2026-07-27** (`655b69f`). R-107 went into
|
||||
ROADMAP alone on **2026-07-28** (`070b0ce`) — **the day after**.
|
||||
|
||||
**Measured scope: 72 `R-` ids are in ROADMAP and not in OPEN-ITEMS; 29 are not marked shipped, closed
|
||||
or killed.** Most of the 29 are feature *ideas*, which arguably belong only in ROADMAP. A minority are
|
||||
**findings**: R-30, R-31, R-32 (all P2-HIGH), R-35, R-40, R-76, R-79, R-25, R-49, R-10, R-107.
|
||||
|
||||
**The risk is not double-minting** — the two files share ids and OPEN-ITEMS' ceiling (368) is above
|
||||
ROADMAP's (331). **The risk is rediscovery**: a session greps one file, finds nothing, and redoes the
|
||||
work. That is precisely what happened, and it cost an evening and a night. **Filed as R-369, HIGH.**
|
||||
|
||||
---
|
||||
|
||||
## 4. Part 1 — the record corrected
|
||||
|
||||
### The specification
|
||||
|
||||
`SPEC-app-data-placement-2026-08-21.md` now opens with a marked **⚠ CORRECTED 2026-08-22** block and
|
||||
carries four inline `[CORRECTED 2026-08-22]` marks. **Nothing was deleted; every measurement stands.**
|
||||
|
||||
What it got wrong: it treated the absence of a storage field on 40 templates as a choice being denied.
|
||||
The architecture states the rule and states it as **enforced**:
|
||||
|
||||
> *"**App data placement is per-volume, not per-app:** `.felhom.yml` classifies each volume **hot**
|
||||
> (DB/config/cache → fast storage, **enforced**) vs **bulk** (media/files → may be slow)."*
|
||||
> — `01-topology-and-trust.md:150-152`
|
||||
|
||||
The 40 are all-hot apps; the 13 are the bulk apps. **And the deploy page has been telling the customer
|
||||
exactly this the whole time** — `deploy.html:624-625`: *„A kiválasztott meghajtón az alkalmazás
|
||||
**fájljai** (média, dokumentumok) tárolódnak. Az **adatbázis a gyors belső SSD-n** fut."*
|
||||
|
||||
### The disk claim, stated precisely
|
||||
|
||||
- **REAL:** a **physical-disk failure** loses the app's data and its first-tier copy together. That is
|
||||
what the off-site and whole-machine tiers exist for, and it is **equally true of a drive-resident
|
||||
app**, whose unit sits beside its data on the drive deliberately so a restore needs the drive and
|
||||
nothing else.
|
||||
- **OVERSTATED:** a **full data volume stopping the operating system.** The OS rootfs is a separate
|
||||
volume, and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`).
|
||||
Watched working 2026-08-21 with `/var/lib/felhom` at 99%: the reserve refused one app per run, told
|
||||
the hub, and all 15 containers stayed healthy.
|
||||
- **WITHDRAWN:** the comparison to Tier 2's same-disk refusal (`tier2.go:329`). Tier 2 refuses a
|
||||
**second** copy on the same disk; Tier 1's unit is *meant* to sit beside the data.
|
||||
|
||||
### Rows re-framed
|
||||
|
||||
- **R-352** — measurements (1)–(4) all stand; the conclusions drawn from (1) and (3) are marked
|
||||
re-framed. (2) is now filed on its own as **R-368**; (4) needs no ruling.
|
||||
- **R-356** — **re-checked and it SURVIVES UNCHANGED, strengthened.** Because the 40-class correctly
|
||||
has no `HDD_PATH`, a restore that reads that as *"the app is not installed"* is misreading a correct
|
||||
configuration. One sentence in it that leaned on the old framing is corrected in place.
|
||||
|
||||
### My own errors, named — R-370
|
||||
|
||||
**Four instances, not three.** The prompt said three; the evidenced count is four, all authored
|
||||
2026-08-21: SPEC §1, SPEC §2.3, SPEC §5, and R-352 point (3). Recorded as a **process** failure with
|
||||
its mechanism — the register and live source were read, `documentation/architecture/` was not — and
|
||||
the missing step is now in the template. Also written into `CONTEXT.md`.
|
||||
|
||||
---
|
||||
|
||||
## 5. What still needs the operator's ruling — **from this specification, almost nothing**
|
||||
|
||||
- **The placement question is ANSWERED and the answer is no.** Points 1, 2 and 3 of the spec's §5 are
|
||||
withdrawn as a live question; hot data belongs where it is.
|
||||
- **Point 4 is the only thing left**, and it is smaller than it looked — see R-368 below.
|
||||
- **Point 5 needs no ruling.** The Drives count is honest; it is a wording question the design system
|
||||
already owns.
|
||||
|
||||
**One new thing does want your ruling, and it is R-369:** whether the two registers become one, or
|
||||
whether a gate enforces that a not-done `ROADMAP` row has an `OPEN-ITEMS` counterpart. Triage the 29
|
||||
first — most are ideas, a minority are findings.
|
||||
|
||||
---
|
||||
|
||||
## 6. Part 4 — the three answers, and the finding is smaller and different than believed
|
||||
|
||||
**1. Is the default store consulted at deploy time, for the 13 apps that take a path? — YES.**
|
||||
`internal/web/templates/deploy.html:612`:
|
||||
|
||||
```
|
||||
paperless-postgres paperless-ngx kimai-db kimai romm-db romm
|
||||
{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}
|
||||
```
|
||||
|
||||
**And that label is the stack name BY CONSTRUCTION, not by luck:** `composeExecCustomEnv`
|
||||
(`internal/stacks/manager.go:1218-1230`) runs compose with `cmd.Dir` set to
|
||||
`/opt/docker/stacks/<stack>` and **never passes `-p`**, so compose derives the project from that
|
||||
directory. Renaming the container in the catalogue would have fixed this one app and left the guessing
|
||||
for the next one. **No catalogue change was made.**
|
||||
For a new deploy the default drive **is pre-selected**. `DeployStoragePath` embeds
|
||||
`settings.StoragePath` (`web/handlers.go:89-99`), so `.IsDefault` resolves.
|
||||
|
||||
**So `// new apps use this by default` (`settings.go:453`) is IMPRECISE ABOUT THE MECHANISM, NOT
|
||||
FALSE.** The earlier claim — *"the deploy route never reads it"* — is **wrong**, and it is wrong for
|
||||
the same reason as everything else this week: the grep behind it
|
||||
(`grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' deploy.go manager.go`) **searched Go files
|
||||
and never the templates.**
|
||||
|
||||
**The residual, and it is the whole finding:** the default lives in the **template**, not the server.
|
||||
`POST /api/stacks/<n>/deploy` takes `values` verbatim; omit `HDD_PATH` and `withPathVars`
|
||||
(`stacks/deploy.go:584`) gets `""` and no default applies. **That is why the invariant has no test —
|
||||
there is nothing server-side to test.** Filed **R-368, LOW**.
|
||||
|
||||
**2. What the label promises the customer, quoted:** `storage.html:469` —
|
||||
**„Legyen alapértelmezett új telepítéseknél"** ("Be the default for new installations"), with the
|
||||
badge **„Alapértelmezett"** at `:34`. **The promise is kept**: it is the default for new installs,
|
||||
which is exactly what the pre-selection does. It does not promise that every app's data goes there.
|
||||
|
||||
**3. What the Drives count counts:** `countAppsUsingPath` (`web/handlers.go:2162-2175`) counts
|
||||
deployed apps where `appCfg.Env["HDD_PATH"] == storagePath`. **A named-volume app has no `HDD_PATH`,
|
||||
so it can never be counted — by construction, not by accident.** The number is honest; it means *"apps
|
||||
that place bulk data on this drive"*, not *"apps using storage"*. Worth one sentence of wording, not a
|
||||
ruling.
|
||||
|
||||
---
|
||||
|
||||
## 2. PART 1.2 — the sweep, proven before its answer was trusted
|
||||
## 7. Part 3 — the template's two new rules
|
||||
|
||||
| step | result |
|
||||
|---|---|
|
||||
| scratch copy, unplanted | `MISMATCHES: 1` (rc 1) — the known one |
|
||||
| **plant a second mismatch** (`kimai-db` → `timetrack-db`) | **`MISMATCHES: 2`**, convicted by name: `stack=kimai container=timetrack-db -> derived=timetrack` |
|
||||
| plant removed | back to `1` |
|
||||
| **every mismatch removed** | **`MISMATCHES: 0`, rc 0** — "1" is not a stuck value |
|
||||
| scratch discarded | `kimai/docker-compose.yml` md5-identical to the real catalogue |
|
||||
**It had neither.** `enumerat` → 0 hits, `register row` → 0 hits; `architecture` appeared 9 times but
|
||||
§4 item 4 named only `02-controller-module-map.md`, and the S-1 rule at the end governs *updating* a
|
||||
design doc, not *reading* one first. **No halt.** There is no separate authoring companion.
|
||||
|
||||
**The real count: 1 affected app of 53**, from 15 DB containers checked — `paperless-ngx`.
|
||||
**Rule 1, in §4 where files are read** — abridged; the full text carries a file→area table for all
|
||||
eight architecture documents:
|
||||
|
||||
> **THE ARCHITECTURE DOCUMENT FOR THE AREA THIS TASK TOUCHES — NAME IT AND SAY WHAT IT SAYS.** Not
|
||||
> "read the architecture folder": name the file, and state in one line what it rules about this area.
|
||||
> **A prompt that cannot name one says so explicitly, and that absence is itself recorded** — an
|
||||
> undocumented architectural decision is how a deliberate design gets "fixed" by someone who did not
|
||||
> know it was one.
|
||||
>
|
||||
> **Three sources, in this order, before any claim: the architecture folder holds the REASONING, the
|
||||
> register holds the WORK, source holds the TRUTH.**
|
||||
>
|
||||
> **And the test that catches it: _is what I am about to call a defect something we chose?_** If it
|
||||
> was chosen and the choice is wrong, that is **a proposal to change a decision** — it goes to the
|
||||
> operator as a decision, not filed as a bug. **Cost of learning this (R-370):** between 19 and 22
|
||||
> August a documented placement decision was called a defect in four places.
|
||||
|
||||
**Rule 2, on the `OPEN-ITEMS.md` bullet, where a survey session lands:**
|
||||
|
||||
> **AN ENUMERATED GAP BECOMES A ROW, IN THE SAME SESSION. PROSE IS NOT A RECORD.**
|
||||
>
|
||||
> This binds **surveys, inventories, spikes, reviews and diagnoses**, not only implementation
|
||||
> sessions… If a document says a thing is missing, unhandled, unreachable or *"not currently filed"*,
|
||||
> it does not leave the session as prose. It leaves as a row here, with a rank and an owner. Writing
|
||||
> *"not filed"* is not a disposition; it is a note that the work was seen and dropped.
|
||||
>
|
||||
> **A row in `ROADMAP.md` alone does not satisfy this** … invisible to every standing rule that says
|
||||
> *"grep the register before minting"* (**R-369**).
|
||||
>
|
||||
> **The cost, recorded so the rule can be narrowed later rather than becoming permanent by accident:**
|
||||
> R-107 … was enumerated on **2026-07-28** … and never entered here. **It was rediscovered from
|
||||
> scratch 25 days later by an overnight drill**, and shipped as R-354.
|
||||
|
||||
Two rules, one place each, ~40 lines added to a 521-line document.
|
||||
|
||||
---
|
||||
|
||||
## 3. PART 1 — the dumps already written under the wrong name
|
||||
## 8. Rows opened — ceiling R-367 → **R-375**
|
||||
|
||||
**They stay, untouched, and nothing will ever collect them.**
|
||||
`/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql`, 312 381 B,
|
||||
last written 07:38 UTC by the final pre-fix cycle — **verified still present after the fix.**
|
||||
| id | rank | what | first written down | age |
|
||||
|---|---|---|---|---|
|
||||
| **R-368** | LOW | The storage default **does** apply at deploy time (`deploy.html:612`); the earlier "never reads it" was wrong. Residual: the default lives in the template, not the server, so the API has none and nothing server-side can be tested | 2026-08-21 (as the wrong claim) | 1 d |
|
||||
| **R-369** | **HIGH** | **Two registers.** 72 ids in ROADMAP only, 29 not done, some of them findings. R-107 sat there 25 days and was rediscovered by a drill | 2026-07-28 | **25 d** |
|
||||
| **R-370** | CLOSED | Process: a documented decision called a defect four times; architecture folder never read. Rule now in the template | 2026-08-19 | 3 d |
|
||||
| **R-371** | LOW | The off-site tier is the only tier that announces nothing on success | 2026-08-05 | 17 d |
|
||||
| **R-372** | LOW | A Tier-2 copy that has **never** been produced (source missing) is not surfaced distinctly | **2026-07-15** | **38 d — the oldest** |
|
||||
| **R-373** | LOW | `SysDataGrowGB` is the intended sizing lever, it works, and nothing sets it | 2026-08-02 | 20 d |
|
||||
| **R-374** | LOW | Three C1 refusal cases judged borderline, left unfiled **and never named**, so nobody can re-judge them | 2026-08-08 | 14 d |
|
||||
| **R-375** | LOW | A PBS datastore audit signal noted and explicitly not filed | 2026-08-18 | 4 d |
|
||||
|
||||
Nothing deletes them by design: the F5 stale-primary prune (`backup/backup.go:1248`) skips any
|
||||
directory that is not a deployed app, under the guard *"an undeployed app's last backup is still its
|
||||
restore point"*.
|
||||
|
||||
**They can be adopted, by hand:** move the file to
|
||||
`…/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql` and it becomes a readable restore point.
|
||||
**Deliberately not automatic** — it would sit beside volume tars from a different time (an incoherent
|
||||
pair, the shape the R-43/R-44 stamp exists to surface), and a controller that silently relocates a
|
||||
customer's data on upgrade is a migration, not a fix. **Filed as R-367.**
|
||||
**Oldest gap recovered: R-372, written down 2026-07-15, 38 days.**
|
||||
**Most consequential: R-369**, because it explains the other seven.
|
||||
|
||||
---
|
||||
|
||||
## 4. PART 2 — the mechanism, and the comment corrected
|
||||
## 9. What was dropped, and observations
|
||||
|
||||
`ReconstituteFromOffsite` skipped every placement flagged as the unit (`offbox_reconstitute.go:344`),
|
||||
and the volume archives live **inside** the unit. A search of the whole off-site path found no
|
||||
volume-restore call; the only two callers of `restoreDockerVolumes` were `restore.go:67` and
|
||||
`restore_unit.go:255`, both local.
|
||||
**Dropped: nothing.** Parts 1, 2, 3 and 4 all completed. Part 4 was droppable-first and was done.
|
||||
|
||||
**The comment as it stood:**
|
||||
**Observations — noticed, not acted on:**
|
||||
|
||||
> *"The live recovery unit is still never overwritten — it is the LOCAL restore path's source and
|
||||
> clobbering it would trade one recovery route for another. The snapshot's dump is replayed from the
|
||||
> scratch unit instead, so nothing is lost by skipping it."*
|
||||
|
||||
**Which half still holds: the FIRST.** The live unit is the local restore path's own source and must
|
||||
never be clobbered — that is why the skip stays, and scenario D now fingerprints the whole live unit
|
||||
across the operation to keep it true.
|
||||
|
||||
**The second half was false and is corrected, not left.** True of the *database* dump, false of the
|
||||
*volume* archives, which live in the same unit and were replayed by nothing at all. Skipping the
|
||||
placement is correct; treating the skip as harmless was not — and that sentence is exactly why a
|
||||
reader would not look.
|
||||
|
||||
**The fix:** `restoreDockerVolumesFrom(stackName, dumpDir) (int, error)` — the local path's own replay
|
||||
with an explicit directory, the same shape `reimportDBDumpsFrom` already had beside `reimportDBDumps`.
|
||||
**One implementation, two callers.** Volumes replay **before** the database (a logical dump must still
|
||||
win over a volume-tar copy of the same database) and **inside the stopped window** (Docker will not
|
||||
replace a volume a container holds).
|
||||
- **The two `### Not filed` sections are the model.** `CAMPAIGN-10-two-storage-soak:383` and
|
||||
`SPIKE-recovery-unit-space:228` list what they decided not to file **and why, item by item**. That
|
||||
is a disposition, not a shrug. If the new rule needs an exemplar, those are it.
|
||||
- **`07-backup-architecture.md:337` already names this gap and points at R-107**, so the architecture
|
||||
doc was right and current while the register was empty. The document was not the weak link.
|
||||
- **The `_recovery-inventory` cites `07-backup-architecture.md:81` for a phrase that is no longer
|
||||
there** (`grep 'missing-only merge'` → only the inventory's own copy). A stale line-number citation;
|
||||
not filed, because the doc it points into has since been rewritten wholesale and the claim it
|
||||
supported is now handled by R-354.
|
||||
- **The sweep's "cited by no row" count (89) is not a defect count.** Many audits are cited by the
|
||||
capability map, by `where-felhom-stands.yaml` or by other reports rather than by a register row. It
|
||||
is a *candidate* filter, and it earned its keep only in combination with the shape search.
|
||||
- **`REPORT-hub-blindness.md` exists at the repo root and is cited by no register row**, but the
|
||||
register does cover its subject (2 hits). Left alone.
|
||||
|
||||
---
|
||||
|
||||
## 5. THE CUSTOMER'S MESSAGE, BEFORE AND AFTER, VERBATIM
|
||||
## 10. CI
|
||||
|
||||
**R-354 — calibre-web, identical fixture, identical steps, both on `demo-hp`:**
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **before** (0.217.0, 09:41) | `ok=true` — „A(z) calibre-web: **0 fájl visszaállítva** (mentés: 2026-08-22 09:38) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and `ls: /vol/R354-2026-08-22: No such file or directory` |
|
||||
| **after** (0.218.0, 09:56) | `ok=true` — „A(z) calibre-web: **0 fájl és 1 adatkötet visszaállítva** (mentés: 2026-08-22 09:52) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and **5/5 files byte-identical** |
|
||||
|
||||
**R-355 — paperless-ngx:**
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **before** (2026-08-21) | „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**" over a live 72-table PostgreSQL, with **no undo copy taken at all** |
|
||||
| **after** (0.218.0, 09:57) | „A(z) paperless-ngx: **0 fájl és 3 adatkötet és az adatbázis visszaállítva** (mentés: 2026-08-22 09:52) — az alkalmazás újraindult." |
|
||||
|
||||
**A third sentence now exists for the case that had no honest wording** — an app that HAS a database
|
||||
whose snapshot carried no dump: „… FIGYELEM: ennek az alkalmazásnak **VAN adatbázisa**, de a mentés nem
|
||||
tartalmazott adatbázis-mentést, ezért az adatbázis **NEM állt vissza**. A visszaállítás előtti állapot
|
||||
mentése megvan: `<undo>`". That case previously printed the same confident „nincs adatbázisa".
|
||||
|
||||
---
|
||||
|
||||
## 6. THE HARDWARE WALK
|
||||
|
||||
**Method: endpoint-level — the exact endpoints the dashboard's JS calls (`/api/backup/run`,
|
||||
`/backup/offbox/run`, `/backup/offbox/restore`, `/backup/offbox/reconstitute`), with no shell inside
|
||||
the machine for any step of the walk.** Planting the fixture and reading the result back used a shell
|
||||
and is setup/verification, not the walk. No browser exists on DooPlex.
|
||||
|
||||
**The fixture, hashed before anything ran:**
|
||||
|
||||
```
|
||||
7708bf6582f990ee5c98e3fb6638214d4916b517ea80638baf20b0662b56506e SENTINEL.txt
|
||||
fb67d42bdfe09815922c2f4f4a2086025bd1a074757d17b5e5164d29b1a5e8d5 binary-512k.bin
|
||||
847fbae0868fc9a2c985d7411fa545c4853e8561f7b8cc8dd485437b1f1d76c4 nested/őszibarack.md
|
||||
f15b6d0f88e5499836f585403e3b59b4f2006a0e6a96ce4ec598dd7abe38bff5 plain.txt
|
||||
857e8594005d11f9802f61018603b7746e7af37ac02631aaef2db57f9700fb58 árvíztűrő-tükörfúrógép.txt
|
||||
|
||||
accented names as RAW BYTES (UTF-8 NFC):
|
||||
árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
|
||||
őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64
|
||||
```
|
||||
|
||||
**The comparator was proved able to convict first:** one byte flipped at offset 300 000 of
|
||||
`binary-512k.bin` (`3b` → `00`) → `binary-512k.bin: FAILED`, rc 1, the other four `OK`; the unmodified
|
||||
set rc 0. Mutant discarded.
|
||||
|
||||
### Results
|
||||
|
||||
| scenario | result |
|
||||
|---|---|
|
||||
| **1-A** the dump is inside the unit, and inside the off-site snapshot | **PASS.** 0.217.0 09:37: `db_dumps = None`, dump refreshed into the phantom dir. 0.218.0 09:45: `db_dumps = ['paperless-ngx-postgres.sql']` (312 669 B) in the app's own unit, and `restic ls` shows it inside the snapshot **for the first time** |
|
||||
| **1-B** an undo copy taken and verified before anything stops; the message names the database | **PASS.** Before: `find /mnt -name "pre-restore-*paperless*"` → **0**. After: `pre-restore-20260822T075658Z-paperless-ngx-postgres.sql`, 312 957 B, in the app's own unit |
|
||||
| **1-C** the undo cannot be taken → the whole restore refuses, nothing changed | **PASS**, on the app that could never reach this guard before. `ok=false`; the marker kept its mutation (`51d37b0f…`) and `StartedAt` was unchanged — **the app was never stopped** |
|
||||
| **1-D** every other app byte-identical | **PASS.** `romm-mariadb.sql`, `kimai-mariadb.sql` unchanged in name and location; the only phantom directory is the one pre-existing orphan; no new one created |
|
||||
| **2-A/B** the volume comes back and the message says so | **PASS.** calibre-web 5/5 byte-identical incl. both accented names; paperless-ngx 3 volumes + the database |
|
||||
| **2-C** a snapshot with no volume archives is unchanged | **PASS** — the unit test asserts the byte-identical sentence; live, romm/kimai unaffected |
|
||||
| **2-D** the live recovery unit is never written | **PASS** — fingerprinted across the whole operation (see red-proof 6) |
|
||||
| **2-E** a failed replay is reported as a failure naming the volume | **PASS** (unit test; the live path returns the same error) |
|
||||
|
||||
All 15 containers healthy after the walk.
|
||||
|
||||
---
|
||||
|
||||
## 7. RED-PROOFS — seven, each mutation asserted applied and reverted
|
||||
|
||||
| # | mutation | outcome |
|
||||
|---|---|---|
|
||||
| 1 | `resolveStackName` ignores the compose label | **FAIL as required.** `= "paperless", want "paperless-ngx"`, and the consequence printed both paths — written to `…/primary/paperless/db-dumps` vs read `…/primary/paperless-ngx/db-dumps`. **The empty database record, returned** |
|
||||
| 2 | outcome message back to the counter-only predicate | **FAIL as required**, reproducing the 21 August sentence verbatim: `"A(z) paperless-ngx: 0 fájl visszaállítva — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."` |
|
||||
| 3 | the fail-closed refusal removed from `writeSafetyDump` | **FAIL as required** — *a restore was seen proceeding with no undo copy* |
|
||||
| 4 | the volume leg removed | **FAIL as required.** `VolumesReplayed = 0` and the replay never called — last night's silent loss, returned |
|
||||
| 5 | the volume count dropped from the message | **FAIL as required**, reproducing „A(z) calibre-web: 5 fájl visszaállítva …" verbatim |
|
||||
| 6 | the live-unit guard removed | **PASSED FIRST — a defect in MY TEST, not the code.** The fingerprint had been narrowed to the volume directory and was blind to a placement writing into the unit **root**. Widened to the whole unit (excluding only the documented undo copies) it convicts: `PLACED:17603ba1…` appears. **The R-181 class, reproduced inside its own regression test — and the reason the red-proof is mandatory** |
|
||||
| 7 | the volume-replay error swallowed | **FAIL as required** — a partial replay reporting success |
|
||||
|
||||
**Green gate:** `go build ./... && go vet ./... && go test ./...` — **28 packages ok, 0 FAIL lines.**
|
||||
Controller gates: all 11 OK. felhom.eu gates: all OK except the golden-currency gate, which was red
|
||||
until the bake (§8) — correctly, and never bypassed.
|
||||
|
||||
*(Instrument note: `go test ./... | grep -vE '^ok'; echo rc=$?` reports the **grep's** exit code, which
|
||||
is 1 when every test passed and nothing was left to print. The verdict above is from an explicit
|
||||
`FAIL`-line count, not from that.)*
|
||||
|
||||
---
|
||||
|
||||
## 8. WHAT THIS DOES **NOT** FIX — the blocker for the apps that need it most
|
||||
|
||||
**The off-site restore still refuses outright for the 40 of 53 apps that declare no data drive
|
||||
(R-356)**, telling the customer that a running app „nincs telepítve" and to reinstall it "to the same
|
||||
place" — which those apps give them no way to choose. Those are exactly the apps whose entire dataset
|
||||
is a named volume, **so R-354's fix cannot reach them until R-356 is closed.**
|
||||
|
||||
That is why the live confirmation used `calibre-web` and `paperless-ngx`: they declare a drive and can
|
||||
actually reach the restore. R-356 is out of scope by the task's own list, and it is now the first thing
|
||||
worth doing.
|
||||
|
||||
---
|
||||
|
||||
## 9. REGISTER
|
||||
|
||||
- **R-354 — CLOSED**, shipped + proven-live (v0.218.0).
|
||||
- **R-355 — CLOSED**, shipped + proven-live (v0.218.0).
|
||||
- **R-367 — OPENED (LOW)**: dumps already written under the wrong name are stranded; adoptable by hand,
|
||||
deliberately not automatic, never on an upgrade path.
|
||||
- **Ceiling moved: R-366 → R-367.**
|
||||
|
||||
---
|
||||
|
||||
## 10. DELIBERATELY OUT OF SCOPE — so it does not read as forgotten
|
||||
|
||||
R-353 (the empty restore reported as success — real, and next), R-352 (where the 40 apps' data lives,
|
||||
awaiting a ruling), R-356 (§8), deploy-and-restore as one act, R-360 (the delete guard on verification
|
||||
copies), R-364 (the accented-search instrument defect), and the remaining drill rows R-357–R-359,
|
||||
R-361–R-363, R-365, R-366.
|
||||
|
||||
---
|
||||
|
||||
## 11. COMMITS, VERSIONS, CI
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| `felhom-controller` | `5ce3a44` the fix + tests; `2da259a` README + CONTEXT |
|
||||
| image | `gitea.dooplex.hu/admin/felhom-controller:0.218.0` (150M) |
|
||||
| deployed | `demo-hp` guest 9201 — `…:0.218.0 Up (healthy)` |
|
||||
| `demo-felhom` | **untouched this session**, still 0.217.0 |
|
||||
|
||||
---
|
||||
|
||||
## 12. BAKE
|
||||
|
||||
`RUNBOOK-manual-build.md` §4.0/§4.1, in the drill VM. **The golden-currency gate went red the moment
|
||||
v0.218.0 was released and stayed red until the bake — that is correct, and it was never bypassed.**
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| `GOLDEN_VERSION` | **0.218.0** |
|
||||
| `GOLDEN_SHA256` | **8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b** |
|
||||
| package | `…/generic/felhom-golden/0.218.0/golden.tar.zst`, 657 026 013 B |
|
||||
| template | `debian-13-standard_13.6-1_amd64.tar.zst` — listed fresh, not assumed |
|
||||
|
||||
**Pass markers, each checked, with the negative controls:**
|
||||
|
||||
```
|
||||
docker OK (overlay2 : 1 -> " docker OK (overlay2; data-root /var/lib/docker)"
|
||||
including mount point : 2 -> rootfs ('/') and mp0 ('/var/lib/felhom') [there is no mp1]
|
||||
upload OK (HTTP 201) : 1 -> pre-delete HTTP 404 (404/204 expected)
|
||||
excluding : 0 <- negative control
|
||||
FATAL : 0 <- negative control
|
||||
```
|
||||
|
||||
**Token hygiene.** Copied file→file and read by a runner script inside the VM, so it never crossed a
|
||||
shell or a unit property: `systemctl show golden-bake … | grep -c -F "$(cat …)"` → **0**. **The leak
|
||||
grep on the committed log was itself proven before its zero was believed** — token appended to a
|
||||
throwaway copy, grepped (**1**), copy shredded, then the committed log's **0** accepted.
|
||||
|
||||
**Teardown:** `pct destroy 9100 --purge`; token, runner, build script and in-VM log shredded **after**
|
||||
the log was copied out; VM powered off; qemu exited; **`qemu-img snapshot -a virgin` restored.**
|
||||
|
||||
Evidence: `documentation/tests/golden-0.218.0-2026-08-22/`.
|
||||
|
||||
---
|
||||
|
||||
## 13. STOP — THE APPROVAL
|
||||
|
||||
**Nothing below has been changed. Vouching and the floor are yours.**
|
||||
|
||||
### The three Day-0 values, each with BOTH checks
|
||||
|
||||
| field | set to | downloadable | selectable |
|
||||
|---|---|---|---|
|
||||
| **`golden_version`** | **0.218.0** — sha `8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b` | ✔ HTTP 200, 657 026 013 B, and the **downloaded bytes** hash to exactly the bake's reported sha | ✔ offered in the hub dropdown with `data-sha=8e427869d13eafb7…` |
|
||||
| **`agent_version`** | **0.130.0** — sha `a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3` | ✔ HTTP 200, 14 141 158 B, downloaded sha matches the hub's `data-sha` exactly | ✔ already the SELECTED option — **no change needed** |
|
||||
| **`min_agent`** | **0.129.0** | — | ✔ already set to 0.129.0 — **no change needed** |
|
||||
|
||||
**So the vouch is a ONE-field change this time, and that is the safe direction, not a shortcut.** The
|
||||
three-field rule exists so a golden is never shipped onto an agent older than it needs. The controller
|
||||
being baked declares **`MinAgent: 0.129.0` (unchanged)**, the vouched agent is already **0.130.0**,
|
||||
and `0.130.0 ≥ 0.129.0` — so `agent_version` and `min_agent` are already correct and only
|
||||
`golden_version` moves. Setting `min_agent` above the vouched agent is the R-216 shape; it is not
|
||||
happening here.
|
||||
|
||||
**If the hub refuses the save, the guard is working.** The R-120 gate on that POST
|
||||
(`hub/internal/web/configs.go:1165`) refuses a golden older than the newest controller the fleet
|
||||
reports. 0.218.0 is the newest, so it should pass — and if it does not, fix the cause, never the guard.
|
||||
|
||||
**Vouching is reversible:** re-select 0.217.0 and Save. The old package is never deleted by a bake.
|
||||
|
||||
### The one line on the floor
|
||||
|
||||
**Raising the floor to 0.218.0 is what puts this on both machines — and yes, you will want to.**
|
||||
`demo-hp` already runs 0.218.0 (deployed directly for the live proof). **`demo-felhom` is still on
|
||||
0.217.0 and will not move until the floor is raised**, so today it still has both defects. I did not
|
||||
raise it; that is your call, as is vouching.
|
||||
|
||||
---
|
||||
|
||||
## 14. WHAT WAS DROPPED, AND OBSERVATIONS
|
||||
|
||||
**Dropped: nothing.** Both parts completed in the required order, with the live walk and all seven
|
||||
red-proofs. The session did not run short.
|
||||
|
||||
**Observations, noticed and not acted on:**
|
||||
|
||||
- **`demo-felhom` was not touched** and still runs 0.217.0 — so the fleet is deliberately non-uniform
|
||||
until the floor moves. Its two preserved fixtures were not involved in any step of this session.
|
||||
- **The `.fab` export has the same R-355 blind spot** and is fixed by the same change — `export.go:600`
|
||||
filters on the same `db.StackName`. Not separately verified live; the unit tests cover the resolver
|
||||
that feeds it.
|
||||
- **`paperless-ngx`'s orphan directory is now a permanent fixture on `demo-hp`** until R-367 is
|
||||
decided. It is the only reproduction of the pre-fix state left anywhere, which is an argument for
|
||||
leaving it alone for now.
|
||||
- **The safety dump still overwrites the unit's own DB dump before renaming it** (R-361, out of scope)
|
||||
— visible in this session's own evidence, where `romm`'s `db-dumps/` holds both a `pre-restore-*` and
|
||||
the regular dump.
|
||||
|
||||
---
|
||||
|
||||
## 15. THE PICTURE — `where-felhom-stands` brought up to date (same day, after the fixes)
|
||||
|
||||
Regenerated from `where-felhom-stands.yaml` (the HTML is a build product; the YAML is the source).
|
||||
**Eight claims re-checked, three moved — all downward.**
|
||||
|
||||
| claim | was | now | why |
|
||||
|---|---|---|---|
|
||||
| `backup.offsite` | walked | **partial** | "18 snapshots, daily, unbroken" was true on 2026-08-09 and false by 2026-08-21 — the next snapshot after it was put there by hand, twelve days later, and nothing reported the gap |
|
||||
| `fail.wiped-reinstalled.data` | walked | **partial** | a real reinstall orphans BOTH off-premises tiers: restic silently for 12 days (R-193), and the pre-reinstall PBS archives cannot be opened by the rebuilt box at all (R-366) |
|
||||
| `backup.fill-warning` | walked | **partial** | the warning fires correctly, but once a day — a filesystem filling at 03:31 is unannounced for ~24 h, and it was watched staying silent while a volume sat at 99% |
|
||||
|
||||
**Held, with their notes rewritten:** `backup.tier1` and `recover.byte-identical` now carry the
|
||||
R-355/R-354 story *and* their fixes; `backup.restore-proof` stays grey for a sharper reason (its
|
||||
archives are orphaned, not merely untested); `backup.sikeres` gained two fresh instances;
|
||||
`fail.customer-self-restore` records that R-356 blocks 40 of 53 apps regardless of who is driving.
|
||||
|
||||
**`unproven.py --summary` moved, and the checklist asks that it be said:**
|
||||
**NOT WALKED went from 32 of 55 to 35 of 55.** That is the three downgrades above. Nothing was raised
|
||||
— the dataset's own rule forbids raising a status here, and nothing needed it.
|
||||
|
||||
### A defect found in the page's own toolchain
|
||||
|
||||
**The rendered header cited the wrong commits and would have gone on doing so.** `render_stands.py`
|
||||
parsed `verified_on` from the YAML but had the three commit shas **hardcoded**, so the page printed
|
||||
the 2026-08-09 commits while the dataset said 2026-08-22. That is precisely the stale-build-product
|
||||
failure the renderer's own docstring says it exists to prevent. Parsed now, literals gone from both
|
||||
the script and the output, checked. The count beside it also read *"15 status(es) moved in that pass"*
|
||||
when 15 was every recorded move ever accumulated; it now separates the two numbers.
|
||||
|
||||
**`check_stands.py` was proven able to convict before its OK was accepted:** a claim marked `missing`
|
||||
flipped to `walked` in a scratch copy fired rule 5 by name (`use.dlna: status 'walked' but NO evidence
|
||||
document cited`), and the real file still passed. Gates all OK; CI **id=380 / run_number=248**, green.
|
||||
Confirmed by ID in §11 of the commit trail below.
|
||||
|
||||
@@ -130,8 +130,35 @@ never park it on a branch.
|
||||
conventions, access). Versioned copy: `felhom.eu/documentation/runbooks/workspace-CLAUDE.md`.
|
||||
2. `<repo>/CLAUDE.md` — repo build/deploy, code-quality rules, trunk-based + live-validation rules.
|
||||
3. `<repo>/CONTEXT.md` — current project state / decisions / roadmap.
|
||||
4. `felhom.eu/documentation/architecture/02-controller-module-map.md` — KEEP/PORT/DELETE/MODIFY
|
||||
per-package classification. **Read before touching `backup/`, `storage/`, `system/`, `config/`.**
|
||||
4. **THE ARCHITECTURE DOCUMENT FOR THE AREA THIS TASK TOUCHES — NAME IT AND SAY WHAT IT SAYS.**
|
||||
Not "read the architecture folder": name the file, and state in one line what it rules about this
|
||||
area. **A prompt that cannot name one says so explicitly, and that absence is itself recorded** —
|
||||
an undocumented architectural decision is how a deliberate design gets "fixed" by someone who did
|
||||
not know it was one.
|
||||
|
||||
| area | file |
|
||||
|---|---|
|
||||
| what the platform does today, per scenario | `00-capability-map.md` |
|
||||
| topology, trust, **app-data placement (hot vs bulk)**, backup scoping | `01-topology-and-trust.md` |
|
||||
| controller packages: KEEP/PORT/DELETE/MODIFY | `02-controller-module-map.md` |
|
||||
| the host agent | `03-host-agent.md` |
|
||||
| control-plane authorization | `04-control-plane-authorization.md` |
|
||||
| the hub | `05-hub-architecture.md` |
|
||||
| off-site connectivity | `06-offsite-connectivity.md` |
|
||||
| tiers, capture sets, restore paths, recovery model | `07-backup-architecture.md` |
|
||||
|
||||
**Three sources, in this order, before any claim: the architecture folder holds the REASONING, the
|
||||
register holds the WORK, source holds the TRUTH.** Skipping the first is how a decision gets
|
||||
reported as a bug.
|
||||
|
||||
**And the test that catches it: _is what I am about to call a defect something we chose?_** If it
|
||||
was chosen and the choice is wrong, that is **a proposal to change a decision** — it goes to the
|
||||
operator as a decision, not filed as a bug. **Cost of learning this (R-370):** between 19 and 22
|
||||
August a documented placement decision was called a defect in four places, because the register and
|
||||
live source were read and `documentation/architecture/` was not.
|
||||
|
||||
`02-controller-module-map.md` remains **required reading before touching `backup/`, `storage/`,
|
||||
`system/`, `config/`.**
|
||||
5. `felhom.eu/documentation/controller/<feature>.md` — code-verified feature docs (authoritative; match
|
||||
code, not summaries, if they drift).
|
||||
6. `controller/README.md` (or `felhom-agent/README.md`) — module map, feature reference, REST API.
|
||||
@@ -347,6 +374,25 @@ or **what is open changed**), update **all four** in the SAME session:
|
||||
that changes what is open without touching it re-creates exactly the thread-loss the register was
|
||||
built to solve: `REPORT.md` is overwritten every session, so nothing durable may live only there.
|
||||
|
||||
> **AN ENUMERATED GAP BECOMES A ROW, IN THE SAME SESSION. PROSE IS NOT A RECORD.**
|
||||
>
|
||||
> This binds **surveys, inventories, spikes, reviews and diagnoses**, not only implementation
|
||||
> sessions — those are the documents that enumerate gaps, and they are the ones that have lost them.
|
||||
> If a document says a thing is missing, unhandled, unreachable or *"not currently filed"*, it does
|
||||
> not leave the session as prose. It leaves as a row here, with a rank and an owner. Writing
|
||||
> *"not filed"* is not a disposition; it is a note that the work was seen and dropped.
|
||||
>
|
||||
> **A row in `ROADMAP.md` alone does not satisfy this.** Both files hold open work and only this one
|
||||
> calls itself the source of truth, so a finding recorded solely there is invisible to every
|
||||
> standing rule that says *"grep the register before minting"* (**R-369**).
|
||||
>
|
||||
> **The cost, recorded so the rule can be narrowed later rather than becoming permanent by
|
||||
> accident:** R-107 — *"no offsite action unpacks the named-volume tars Tier-3 captures"* — was
|
||||
> enumerated on **2026-07-28**, given a number, written into `ROADMAP.md` and
|
||||
> `07-backup-architecture.md`, and never entered here. **It was rediscovered from scratch 25 days
|
||||
> later by an overnight drill that planted files and watched them not come back**, and shipped as
|
||||
> R-354. The work was right the first time; only the filing was missing.
|
||||
|
||||
**Report which `OPEN-ITEMS.md` rows the task opened, closed or re-ranked** (§15). Every row carries
|
||||
an owner — a row nobody owns is how items got lost in the first place.
|
||||
|
||||
|
||||
@@ -0,0 +1,383 @@
|
||||
# REPORT — the database nobody backed up, and the restore that returned most apps nothing (2026-08-22)
|
||||
|
||||
**Controller v0.217.0 → v0.218.0.** Two fixes, both found by watching a machine on the night of
|
||||
2026-08-21, both confirmed the same way. **Part 1 first, because it is the only place in the product
|
||||
where one customer action causes permanent total loss.**
|
||||
|
||||
The drill that found them is preserved at
|
||||
`documentation/audits/REPORT-DRILL-backup-truth-2026-08-21.md`; its evidence is in
|
||||
`documentation/audits/DRILL-backup-truth-2026-08-21/evidence/`.
|
||||
|
||||
---
|
||||
|
||||
## 0. Baselines, re-established — not trusted from the sheet
|
||||
|
||||
| item | expected | confirmed |
|
||||
|---|---|---|
|
||||
| controller | 0.217.0 → 0.218.0 | ✔ repo head `v0.217.0`; live on `demo-hp` `…:0.217.0` |
|
||||
| agent | 0.130.0 | ✔ repo head and `felhom-agent --version` on the box |
|
||||
| golden vouched | 0.217.0 | ✔ `<option value="0.217.0" … selected>` |
|
||||
| floor | 0.217.0 | ✔ „Effective floor: v0.217.0 — source: DB (hub_settings)" |
|
||||
| min agent | 0.129.0 | ✔ |
|
||||
| register ceiling | R-366 | ✔ (now **R-367**) |
|
||||
| `ssh hp` → 192.168.0.104 | works | ✔ |
|
||||
| clean trees, all four repos | `HEAD == origin/main`, 0 dirty | ✔ |
|
||||
|
||||
---
|
||||
|
||||
## 1. PART 1 — where the wrong name comes from
|
||||
|
||||
**It comes from a guess, made in a place where an answer was available.**
|
||||
|
||||
`deriveStackName` (`controller/internal/appbackup/dbdump.go:770-798`) is handed a container name and
|
||||
asked which app owns it. For `paperless-postgres` it strips the role suffix to `paperless`; finds
|
||||
`paperless` is **not** a deployed stack; finds the container name is not one either; finds no known
|
||||
stack is a prefix of it (`paperless-ngx` is not a prefix of `paperless-postgres`) — **and then returns
|
||||
the unresolved candidate anyway**, on its last line, silently.
|
||||
|
||||
**Every place that value is used, and what each did with it:**
|
||||
|
||||
| consumer | file:line | effect |
|
||||
|---|---|---|
|
||||
| dump directory | `backup/backup.go:482,506` | wrote to `…/primary/**paperless**/db-dumps/` on the SYSTEM drive |
|
||||
| dump filename | `appbackup/dbdump.go:200-212` | `paperless-postgres.sql`, not `paperless-ngx-postgres.sql` |
|
||||
| unit assembler | `backup/recovery_unit.go:131` | reads `…/primary/**paperless-ngx**/db-dumps` → finds nothing → `db_dumps: null` |
|
||||
| off-site collector | `backup/offbox_capture.go:32` | resolves from the unit path → the orphan is invisible to it |
|
||||
| restore's DB leg | `backup/restore_db.go:69` | `db.StackName != stackName` → never replays |
|
||||
| **safety-dump filter** | `backup/offbox_reconstitute.go:134` | `mine` empty → `hasDB=false` → **no undo copy, and the refusal is never reached** |
|
||||
| `.fab` export | `appexport/export.go:600` | same filter — the bundle carries no database either |
|
||||
| the "has a DB" flag | `web/handlers.go:1245` | the app displays as having no database |
|
||||
|
||||
**Corroboration from the code itself:** `ListDumpFiles` (`dbdump.go:562`) parses a dump filename under
|
||||
the comment *"Parse stack name and DB type from filename: `paperless-ngx-postgres.sql`"* — the reader
|
||||
was written for a name the writer never produced.
|
||||
|
||||
### Controller or catalogue? — **THE CONTROLLER. No halt.**
|
||||
|
||||
Every container the controller starts already carries `com.docker.compose.project`, read live:
|
||||
|
||||
```
|
||||
paperless-postgres paperless-ngx kimai-db kimai romm-db romm
|
||||
```
|
||||
|
||||
**And that label is the stack name BY CONSTRUCTION, not by luck:** `composeExecCustomEnv`
|
||||
(`internal/stacks/manager.go:1218-1230`) runs compose with `cmd.Dir` set to
|
||||
`/opt/docker/stacks/<stack>` and **never passes `-p`**, so compose derives the project from that
|
||||
directory. Renaming the container in the catalogue would have fixed this one app and left the guessing
|
||||
for the next one. **No catalogue change was made.**
|
||||
|
||||
---
|
||||
|
||||
## 2. PART 1.2 — the sweep, proven before its answer was trusted
|
||||
|
||||
| step | result |
|
||||
|---|---|
|
||||
| scratch copy, unplanted | `MISMATCHES: 1` (rc 1) — the known one |
|
||||
| **plant a second mismatch** (`kimai-db` → `timetrack-db`) | **`MISMATCHES: 2`**, convicted by name: `stack=kimai container=timetrack-db -> derived=timetrack` |
|
||||
| plant removed | back to `1` |
|
||||
| **every mismatch removed** | **`MISMATCHES: 0`, rc 0** — "1" is not a stuck value |
|
||||
| scratch discarded | `kimai/docker-compose.yml` md5-identical to the real catalogue |
|
||||
|
||||
**The real count: 1 affected app of 53**, from 15 DB containers checked — `paperless-ngx`.
|
||||
|
||||
---
|
||||
|
||||
## 3. PART 1 — the dumps already written under the wrong name
|
||||
|
||||
**They stay, untouched, and nothing will ever collect them.**
|
||||
`/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql`, 312 381 B,
|
||||
last written 07:38 UTC by the final pre-fix cycle — **verified still present after the fix.**
|
||||
|
||||
Nothing deletes them by design: the F5 stale-primary prune (`backup/backup.go:1248`) skips any
|
||||
directory that is not a deployed app, under the guard *"an undeployed app's last backup is still its
|
||||
restore point"*.
|
||||
|
||||
**They can be adopted, by hand:** move the file to
|
||||
`…/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql` and it becomes a readable restore point.
|
||||
**Deliberately not automatic** — it would sit beside volume tars from a different time (an incoherent
|
||||
pair, the shape the R-43/R-44 stamp exists to surface), and a controller that silently relocates a
|
||||
customer's data on upgrade is a migration, not a fix. **Filed as R-367.**
|
||||
|
||||
---
|
||||
|
||||
## 4. PART 2 — the mechanism, and the comment corrected
|
||||
|
||||
`ReconstituteFromOffsite` skipped every placement flagged as the unit (`offbox_reconstitute.go:344`),
|
||||
and the volume archives live **inside** the unit. A search of the whole off-site path found no
|
||||
volume-restore call; the only two callers of `restoreDockerVolumes` were `restore.go:67` and
|
||||
`restore_unit.go:255`, both local.
|
||||
|
||||
**The comment as it stood:**
|
||||
|
||||
> *"The live recovery unit is still never overwritten — it is the LOCAL restore path's source and
|
||||
> clobbering it would trade one recovery route for another. The snapshot's dump is replayed from the
|
||||
> scratch unit instead, so nothing is lost by skipping it."*
|
||||
|
||||
**Which half still holds: the FIRST.** The live unit is the local restore path's own source and must
|
||||
never be clobbered — that is why the skip stays, and scenario D now fingerprints the whole live unit
|
||||
across the operation to keep it true.
|
||||
|
||||
**The second half was false and is corrected, not left.** True of the *database* dump, false of the
|
||||
*volume* archives, which live in the same unit and were replayed by nothing at all. Skipping the
|
||||
placement is correct; treating the skip as harmless was not — and that sentence is exactly why a
|
||||
reader would not look.
|
||||
|
||||
**The fix:** `restoreDockerVolumesFrom(stackName, dumpDir) (int, error)` — the local path's own replay
|
||||
with an explicit directory, the same shape `reimportDBDumpsFrom` already had beside `reimportDBDumps`.
|
||||
**One implementation, two callers.** Volumes replay **before** the database (a logical dump must still
|
||||
win over a volume-tar copy of the same database) and **inside the stopped window** (Docker will not
|
||||
replace a volume a container holds).
|
||||
|
||||
---
|
||||
|
||||
## 5. THE CUSTOMER'S MESSAGE, BEFORE AND AFTER, VERBATIM
|
||||
|
||||
**R-354 — calibre-web, identical fixture, identical steps, both on `demo-hp`:**
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **before** (0.217.0, 09:41) | `ok=true` — „A(z) calibre-web: **0 fájl visszaállítva** (mentés: 2026-08-22 09:38) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and `ls: /vol/R354-2026-08-22: No such file or directory` |
|
||||
| **after** (0.218.0, 09:56) | `ok=true` — „A(z) calibre-web: **0 fájl és 1 adatkötet visszaállítva** (mentés: 2026-08-22 09:52) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa." …and **5/5 files byte-identical** |
|
||||
|
||||
**R-355 — paperless-ngx:**
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **before** (2026-08-21) | „A(z) paperless-ngx: 0 fájl visszaállítva … **Ennek az alkalmazásnak nincs adatbázisa.**" over a live 72-table PostgreSQL, with **no undo copy taken at all** |
|
||||
| **after** (0.218.0, 09:57) | „A(z) paperless-ngx: **0 fájl és 3 adatkötet és az adatbázis visszaállítva** (mentés: 2026-08-22 09:52) — az alkalmazás újraindult." |
|
||||
|
||||
**A third sentence now exists for the case that had no honest wording** — an app that HAS a database
|
||||
whose snapshot carried no dump: „… FIGYELEM: ennek az alkalmazásnak **VAN adatbázisa**, de a mentés nem
|
||||
tartalmazott adatbázis-mentést, ezért az adatbázis **NEM állt vissza**. A visszaállítás előtti állapot
|
||||
mentése megvan: `<undo>`". That case previously printed the same confident „nincs adatbázisa".
|
||||
|
||||
---
|
||||
|
||||
## 6. THE HARDWARE WALK
|
||||
|
||||
**Method: endpoint-level — the exact endpoints the dashboard's JS calls (`/api/backup/run`,
|
||||
`/backup/offbox/run`, `/backup/offbox/restore`, `/backup/offbox/reconstitute`), with no shell inside
|
||||
the machine for any step of the walk.** Planting the fixture and reading the result back used a shell
|
||||
and is setup/verification, not the walk. No browser exists on DooPlex.
|
||||
|
||||
**The fixture, hashed before anything ran:**
|
||||
|
||||
```
|
||||
7708bf6582f990ee5c98e3fb6638214d4916b517ea80638baf20b0662b56506e SENTINEL.txt
|
||||
fb67d42bdfe09815922c2f4f4a2086025bd1a074757d17b5e5164d29b1a5e8d5 binary-512k.bin
|
||||
847fbae0868fc9a2c985d7411fa545c4853e8561f7b8cc8dd485437b1f1d76c4 nested/őszibarack.md
|
||||
f15b6d0f88e5499836f585403e3b59b4f2006a0e6a96ce4ec598dd7abe38bff5 plain.txt
|
||||
857e8594005d11f9802f61018603b7746e7af37ac02631aaef2db57f9700fb58 árvíztűrő-tükörfúrógép.txt
|
||||
|
||||
accented names as RAW BYTES (UTF-8 NFC):
|
||||
árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
|
||||
őszibarack.md = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64
|
||||
```
|
||||
|
||||
**The comparator was proved able to convict first:** one byte flipped at offset 300 000 of
|
||||
`binary-512k.bin` (`3b` → `00`) → `binary-512k.bin: FAILED`, rc 1, the other four `OK`; the unmodified
|
||||
set rc 0. Mutant discarded.
|
||||
|
||||
### Results
|
||||
|
||||
| scenario | result |
|
||||
|---|---|
|
||||
| **1-A** the dump is inside the unit, and inside the off-site snapshot | **PASS.** 0.217.0 09:37: `db_dumps = None`, dump refreshed into the phantom dir. 0.218.0 09:45: `db_dumps = ['paperless-ngx-postgres.sql']` (312 669 B) in the app's own unit, and `restic ls` shows it inside the snapshot **for the first time** |
|
||||
| **1-B** an undo copy taken and verified before anything stops; the message names the database | **PASS.** Before: `find /mnt -name "pre-restore-*paperless*"` → **0**. After: `pre-restore-20260822T075658Z-paperless-ngx-postgres.sql`, 312 957 B, in the app's own unit |
|
||||
| **1-C** the undo cannot be taken → the whole restore refuses, nothing changed | **PASS**, on the app that could never reach this guard before. `ok=false`; the marker kept its mutation (`51d37b0f…`) and `StartedAt` was unchanged — **the app was never stopped** |
|
||||
| **1-D** every other app byte-identical | **PASS.** `romm-mariadb.sql`, `kimai-mariadb.sql` unchanged in name and location; the only phantom directory is the one pre-existing orphan; no new one created |
|
||||
| **2-A/B** the volume comes back and the message says so | **PASS.** calibre-web 5/5 byte-identical incl. both accented names; paperless-ngx 3 volumes + the database |
|
||||
| **2-C** a snapshot with no volume archives is unchanged | **PASS** — the unit test asserts the byte-identical sentence; live, romm/kimai unaffected |
|
||||
| **2-D** the live recovery unit is never written | **PASS** — fingerprinted across the whole operation (see red-proof 6) |
|
||||
| **2-E** a failed replay is reported as a failure naming the volume | **PASS** (unit test; the live path returns the same error) |
|
||||
|
||||
All 15 containers healthy after the walk.
|
||||
|
||||
---
|
||||
|
||||
## 7. RED-PROOFS — seven, each mutation asserted applied and reverted
|
||||
|
||||
| # | mutation | outcome |
|
||||
|---|---|---|
|
||||
| 1 | `resolveStackName` ignores the compose label | **FAIL as required.** `= "paperless", want "paperless-ngx"`, and the consequence printed both paths — written to `…/primary/paperless/db-dumps` vs read `…/primary/paperless-ngx/db-dumps`. **The empty database record, returned** |
|
||||
| 2 | outcome message back to the counter-only predicate | **FAIL as required**, reproducing the 21 August sentence verbatim: `"A(z) paperless-ngx: 0 fájl visszaállítva — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."` |
|
||||
| 3 | the fail-closed refusal removed from `writeSafetyDump` | **FAIL as required** — *a restore was seen proceeding with no undo copy* |
|
||||
| 4 | the volume leg removed | **FAIL as required.** `VolumesReplayed = 0` and the replay never called — last night's silent loss, returned |
|
||||
| 5 | the volume count dropped from the message | **FAIL as required**, reproducing „A(z) calibre-web: 5 fájl visszaállítva …" verbatim |
|
||||
| 6 | the live-unit guard removed | **PASSED FIRST — a defect in MY TEST, not the code.** The fingerprint had been narrowed to the volume directory and was blind to a placement writing into the unit **root**. Widened to the whole unit (excluding only the documented undo copies) it convicts: `PLACED:17603ba1…` appears. **The R-181 class, reproduced inside its own regression test — and the reason the red-proof is mandatory** |
|
||||
| 7 | the volume-replay error swallowed | **FAIL as required** — a partial replay reporting success |
|
||||
|
||||
**Green gate:** `go build ./... && go vet ./... && go test ./...` — **28 packages ok, 0 FAIL lines.**
|
||||
Controller gates: all 11 OK. felhom.eu gates: all OK except the golden-currency gate, which was red
|
||||
until the bake (§8) — correctly, and never bypassed.
|
||||
|
||||
*(Instrument note: `go test ./... | grep -vE '^ok'; echo rc=$?` reports the **grep's** exit code, which
|
||||
is 1 when every test passed and nothing was left to print. The verdict above is from an explicit
|
||||
`FAIL`-line count, not from that.)*
|
||||
|
||||
---
|
||||
|
||||
## 8. WHAT THIS DOES **NOT** FIX — the blocker for the apps that need it most
|
||||
|
||||
**The off-site restore still refuses outright for the 40 of 53 apps that declare no data drive
|
||||
(R-356)**, telling the customer that a running app „nincs telepítve" and to reinstall it "to the same
|
||||
place" — which those apps give them no way to choose. Those are exactly the apps whose entire dataset
|
||||
is a named volume, **so R-354's fix cannot reach them until R-356 is closed.**
|
||||
|
||||
That is why the live confirmation used `calibre-web` and `paperless-ngx`: they declare a drive and can
|
||||
actually reach the restore. R-356 is out of scope by the task's own list, and it is now the first thing
|
||||
worth doing.
|
||||
|
||||
---
|
||||
|
||||
## 9. REGISTER
|
||||
|
||||
- **R-354 — CLOSED**, shipped + proven-live (v0.218.0).
|
||||
- **R-355 — CLOSED**, shipped + proven-live (v0.218.0).
|
||||
- **R-367 — OPENED (LOW)**: dumps already written under the wrong name are stranded; adoptable by hand,
|
||||
deliberately not automatic, never on an upgrade path.
|
||||
- **Ceiling moved: R-366 → R-367.**
|
||||
|
||||
---
|
||||
|
||||
## 10. DELIBERATELY OUT OF SCOPE — so it does not read as forgotten
|
||||
|
||||
R-353 (the empty restore reported as success — real, and next), R-352 (where the 40 apps' data lives,
|
||||
awaiting a ruling), R-356 (§8), deploy-and-restore as one act, R-360 (the delete guard on verification
|
||||
copies), R-364 (the accented-search instrument defect), and the remaining drill rows R-357–R-359,
|
||||
R-361–R-363, R-365, R-366.
|
||||
|
||||
---
|
||||
|
||||
## 11. COMMITS, VERSIONS, CI
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| `felhom-controller` | `5ce3a44` the fix + tests; `2da259a` README + CONTEXT |
|
||||
| image | `gitea.dooplex.hu/admin/felhom-controller:0.218.0` (150M) |
|
||||
| deployed | `demo-hp` guest 9201 — `…:0.218.0 Up (healthy)` |
|
||||
| `demo-felhom` | **untouched this session**, still 0.217.0 |
|
||||
|
||||
---
|
||||
|
||||
## 12. BAKE
|
||||
|
||||
`RUNBOOK-manual-build.md` §4.0/§4.1, in the drill VM. **The golden-currency gate went red the moment
|
||||
v0.218.0 was released and stayed red until the bake — that is correct, and it was never bypassed.**
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| `GOLDEN_VERSION` | **0.218.0** |
|
||||
| `GOLDEN_SHA256` | **8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b** |
|
||||
| package | `…/generic/felhom-golden/0.218.0/golden.tar.zst`, 657 026 013 B |
|
||||
| template | `debian-13-standard_13.6-1_amd64.tar.zst` — listed fresh, not assumed |
|
||||
|
||||
**Pass markers, each checked, with the negative controls:**
|
||||
|
||||
```
|
||||
docker OK (overlay2 : 1 -> " docker OK (overlay2; data-root /var/lib/docker)"
|
||||
including mount point : 2 -> rootfs ('/') and mp0 ('/var/lib/felhom') [there is no mp1]
|
||||
upload OK (HTTP 201) : 1 -> pre-delete HTTP 404 (404/204 expected)
|
||||
excluding : 0 <- negative control
|
||||
FATAL : 0 <- negative control
|
||||
```
|
||||
|
||||
**Token hygiene.** Copied file→file and read by a runner script inside the VM, so it never crossed a
|
||||
shell or a unit property: `systemctl show golden-bake … | grep -c -F "$(cat …)"` → **0**. **The leak
|
||||
grep on the committed log was itself proven before its zero was believed** — token appended to a
|
||||
throwaway copy, grepped (**1**), copy shredded, then the committed log's **0** accepted.
|
||||
|
||||
**Teardown:** `pct destroy 9100 --purge`; token, runner, build script and in-VM log shredded **after**
|
||||
the log was copied out; VM powered off; qemu exited; **`qemu-img snapshot -a virgin` restored.**
|
||||
|
||||
Evidence: `documentation/tests/golden-0.218.0-2026-08-22/`.
|
||||
|
||||
---
|
||||
|
||||
## 13. STOP — THE APPROVAL
|
||||
|
||||
**Nothing below has been changed. Vouching and the floor are yours.**
|
||||
|
||||
### The three Day-0 values, each with BOTH checks
|
||||
|
||||
| field | set to | downloadable | selectable |
|
||||
|---|---|---|---|
|
||||
| **`golden_version`** | **0.218.0** — sha `8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b` | ✔ HTTP 200, 657 026 013 B, and the **downloaded bytes** hash to exactly the bake's reported sha | ✔ offered in the hub dropdown with `data-sha=8e427869d13eafb7…` |
|
||||
| **`agent_version`** | **0.130.0** — sha `a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3` | ✔ HTTP 200, 14 141 158 B, downloaded sha matches the hub's `data-sha` exactly | ✔ already the SELECTED option — **no change needed** |
|
||||
| **`min_agent`** | **0.129.0** | — | ✔ already set to 0.129.0 — **no change needed** |
|
||||
|
||||
**So the vouch is a ONE-field change this time, and that is the safe direction, not a shortcut.** The
|
||||
three-field rule exists so a golden is never shipped onto an agent older than it needs. The controller
|
||||
being baked declares **`MinAgent: 0.129.0` (unchanged)**, the vouched agent is already **0.130.0**,
|
||||
and `0.130.0 ≥ 0.129.0` — so `agent_version` and `min_agent` are already correct and only
|
||||
`golden_version` moves. Setting `min_agent` above the vouched agent is the R-216 shape; it is not
|
||||
happening here.
|
||||
|
||||
**If the hub refuses the save, the guard is working.** The R-120 gate on that POST
|
||||
(`hub/internal/web/configs.go:1165`) refuses a golden older than the newest controller the fleet
|
||||
reports. 0.218.0 is the newest, so it should pass — and if it does not, fix the cause, never the guard.
|
||||
|
||||
**Vouching is reversible:** re-select 0.217.0 and Save. The old package is never deleted by a bake.
|
||||
|
||||
### The one line on the floor
|
||||
|
||||
**Raising the floor to 0.218.0 is what puts this on both machines — and yes, you will want to.**
|
||||
`demo-hp` already runs 0.218.0 (deployed directly for the live proof). **`demo-felhom` is still on
|
||||
0.217.0 and will not move until the floor is raised**, so today it still has both defects. I did not
|
||||
raise it; that is your call, as is vouching.
|
||||
|
||||
---
|
||||
|
||||
## 14. WHAT WAS DROPPED, AND OBSERVATIONS
|
||||
|
||||
**Dropped: nothing.** Both parts completed in the required order, with the live walk and all seven
|
||||
red-proofs. The session did not run short.
|
||||
|
||||
**Observations, noticed and not acted on:**
|
||||
|
||||
- **`demo-felhom` was not touched** and still runs 0.217.0 — so the fleet is deliberately non-uniform
|
||||
until the floor moves. Its two preserved fixtures were not involved in any step of this session.
|
||||
- **The `.fab` export has the same R-355 blind spot** and is fixed by the same change — `export.go:600`
|
||||
filters on the same `db.StackName`. Not separately verified live; the unit tests cover the resolver
|
||||
that feeds it.
|
||||
- **`paperless-ngx`'s orphan directory is now a permanent fixture on `demo-hp`** until R-367 is
|
||||
decided. It is the only reproduction of the pre-fix state left anywhere, which is an argument for
|
||||
leaving it alone for now.
|
||||
- **The safety dump still overwrites the unit's own DB dump before renaming it** (R-361, out of scope)
|
||||
— visible in this session's own evidence, where `romm`'s `db-dumps/` holds both a `pre-restore-*` and
|
||||
the regular dump.
|
||||
|
||||
---
|
||||
|
||||
## 15. THE PICTURE — `where-felhom-stands` brought up to date (same day, after the fixes)
|
||||
|
||||
Regenerated from `where-felhom-stands.yaml` (the HTML is a build product; the YAML is the source).
|
||||
**Eight claims re-checked, three moved — all downward.**
|
||||
|
||||
| claim | was | now | why |
|
||||
|---|---|---|---|
|
||||
| `backup.offsite` | walked | **partial** | "18 snapshots, daily, unbroken" was true on 2026-08-09 and false by 2026-08-21 — the next snapshot after it was put there by hand, twelve days later, and nothing reported the gap |
|
||||
| `fail.wiped-reinstalled.data` | walked | **partial** | a real reinstall orphans BOTH off-premises tiers: restic silently for 12 days (R-193), and the pre-reinstall PBS archives cannot be opened by the rebuilt box at all (R-366) |
|
||||
| `backup.fill-warning` | walked | **partial** | the warning fires correctly, but once a day — a filesystem filling at 03:31 is unannounced for ~24 h, and it was watched staying silent while a volume sat at 99% |
|
||||
|
||||
**Held, with their notes rewritten:** `backup.tier1` and `recover.byte-identical` now carry the
|
||||
R-355/R-354 story *and* their fixes; `backup.restore-proof` stays grey for a sharper reason (its
|
||||
archives are orphaned, not merely untested); `backup.sikeres` gained two fresh instances;
|
||||
`fail.customer-self-restore` records that R-356 blocks 40 of 53 apps regardless of who is driving.
|
||||
|
||||
**`unproven.py --summary` moved, and the checklist asks that it be said:**
|
||||
**NOT WALKED went from 32 of 55 to 35 of 55.** That is the three downgrades above. Nothing was raised
|
||||
— the dataset's own rule forbids raising a status here, and nothing needed it.
|
||||
|
||||
### A defect found in the page's own toolchain
|
||||
|
||||
**The rendered header cited the wrong commits and would have gone on doing so.** `render_stands.py`
|
||||
parsed `verified_on` from the YAML but had the three commit shas **hardcoded**, so the page printed
|
||||
the 2026-08-09 commits while the dataset said 2026-08-22. That is precisely the stale-build-product
|
||||
failure the renderer's own docstring says it exists to prevent. Parsed now, literals gone from both
|
||||
the script and the output, checked. The count beside it also read *"15 status(es) moved in that pass"*
|
||||
when 15 was every recorded move ever accumulated; it now separates the two numbers.
|
||||
|
||||
**`check_stands.py` was proven able to convict before its OK was accepted:** a claim marked `missing`
|
||||
flipped to `walked` in a scratch copy fired rule 5 by name (`use.dlna: status 'walked' but NO evidence
|
||||
document cited`), and the real file still passed. Gates all OK; CI **id=380 / run_number=248**, green.
|
||||
@@ -651,11 +651,11 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-349** | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC |
|
||||
| **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:<HUB_PW>@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes |
|
||||
| **R-351** | **The restore never read back where the backup said the data lived, and a second press started a second restore.** Two findings, one session, both shipped. **(a) The blindness.** Every recovery-unit `manifest.json` has carried `drive` and `namespace_root` since schema 1 (`controller/internal/backup/recovery_unit.go:48-49`), written at capture from the app's own placement. `grep -rE '\.Drive\b\|\.NamespaceRoot\b' --include=*.go` found **no non-test reader anywhere** — the reconstitution opened the manifest (`offbox_reconstitute.go:235`) purely for the coherence stamp and resolved its destination from the LIVE app instead. **A restore into a destination different from the recorded one therefore succeeded silently, under a green message.** **(b) The second press.** All seven restore handlers gated on `backupMgr.IsRunning()` — the CONCURRENCY flag, acquired *inside* the goroutine (`offbox_reconstitute.go:180`) **after** the handler returned. Established with a test before any change: both the reconstitute and place handlers answered „…elindult" and **overwrote the first restore's op/stack**. The wizard had read the correct flag since v0.154.0 and said so in a comment; the handlers were never moved over. **(c)** The banner gated its terminal result on a page-local `sawRunning`, so a restore that finished before the page opened — the 8.666 s OpenGist restore — was shown to nobody. | **CLOSED 2026-08-21** — controller | — | **Shipped:** `backup/offbox_placement.go` (`CheckPlacement`, `PlacementMismatchMessage`, `RecordedUnitForStack`); mismatch **named and refused** before the safety dump, with `ack_placement` as a **separate** field from `confirm=1`; the not-installed refusal names the recorded drive; deploy page **prefills the address and folder from the app's own backup**; `Server.restoreOpBlocked()` reads BOTH flags; `RestoreOpStatus.LastRecent` + `RestoreResultWindow` moved to `internal/backup` as ONE expression for two surfaces. Red-proofs: B with **both** guards removed **was seen starting a restore with no drive attached** (no error, full 3.00 s run into `/tmp/mutant-destination`); C, E and A each returned their wrong outcome; D forced on broke 8 ordinary reconstitute tests, proving reachability both ways. | CC |
|
||||
| **R-352** | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes |
|
||||
| **R-352** | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes |
|
||||
| **R-353** | **A restore reported success having returned configuration and no data — and no screen could have told the customer.** `demo-hp`, 2026-08-21, OpenGist. The off-site reconstitution refused at 16:37:14 (not installed); the person reinstalled and ran the local unit restore, which reported `Restore-from-unit completed: opengist in 8.666896042s`. **The unit it restored from contains `manifest.json` + `compose/{app.yaml,.felhom.yml,docker-compose.yml}` and NOTHING else — `volume_dumps: None`, `db_dumps: None`** — and the off-site snapshot was **182.3 KB**. So the restore returned the app's configuration; there was no data leg in the unit to return, and the outcome said only that it had completed. **A warning beside a success is read as a success, and an unknown must never be drawn as healthy.** **Compounding, and recorded as UNKNOWN rather than fine:** whether the 40-class reaches the off-site tier at all has **not been observed** — `runVolumeDumps` (`backup/backup.go:607+`) covers them on paper, but every unit on the box reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly dump run had happened on a one-hour-old box. | **OPEN — NEXT SESSION'S FIRST ITEM** | — | **Two things, in order. (1)** A restore whose unit carries no `db_dumps` and no `volume_dumps` must **say so in its outcome** — „a mentés csak a beállításokat tartalmazta, adatot nem" — instead of reporting a bare completion. The verdict must consult what was actually placed, not merely that the operation ended. **(2)** Then *prove* the off-site coverage of a named-volume app by running a dump cycle and reading the resulting manifest, rather than inferring it from the gate order. Do not close (1) on the strength of (2) being likely. **(2) IS NOW SATISFIED — drill 2026-08-21.** A dump cycle was run and the manifests read: `privatebin volume_dumps=[privatebin_privatebin_data.tar]`, `opengist volume_dumps=[opengist_opengist_data.tar]`, `kimai volume_dumps=[kimai_kimai_db_data.tar, kimai_kimai_var.tar]` — the 40-class DOES reach the off-site tier, and PrivateBin's planted 1 MB came back byte-identical from its off-site snapshot into the checking folder. **(1) stands and is now strictly larger than when written:** R-354 shows the bare completion is also reported over a unit that DID carry a data leg, because the off-site restore never replays volume dumps at all. | CC |
|
||||
| **R-354** | **The off-site full restore has NO named-volume leg — the tar is in the unit, in the snapshot and in the checking folder, and is never replayed.** Proven live on `demo-hp` 2026-08-21 22:23 with planted files. `calibre-web`'s `calibre_web_config` tar (1 422 848 B) was present at every stage and the restore returned 5 declared user files and **not the volume**, under „5 fájl visszaállítva … az alkalmazás újraindult". Cause: `ReconstituteFromOffsite` skips every placement flagged `isUnit` (`controller/internal/backup/offbox_reconstitute.go:341-346`) and the volume tars live INSIDE the unit; a grep for a volume-restore call across the whole off-site path returns nothing. The **local** restore does have one (`restore.go:99 restoreDockerVolumes`) — proven the same night by returning PrivateBin's planted 1 MB byte-identical from the same tar. **For the 13 drive-declaring apps the lost leg is the app's own configuration; for the 40 no-drive apps it is the entire dataset.** | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22** (controller **v0.218.0**) | — | `restoreDockerVolumesFrom` is the LOCAL path's own replay with an explicit directory — ONE implementation, two callers, because a second copy of that loop is what produced the divergence. It reads the SCRATCH unit; the live unit is still never written. Volumes replay BEFORE the database (a logical dump must still win over a volume-tar copy of the same database) and inside the stopped window (Docker will not replace a volume a container holds). `VolumesReplayed` is on the result and in the sentence. **The comment beside the skip was half false and is corrected, not left:** it justified the skip by saying the dump is replayed from the scratch "so nothing is lost" — true of the database, false of the volumes. The half that still holds (the live unit is the local path's source) is named. **PROVEN LIVE on `demo-hp`, negative control first.** On 0.217.0 with the fixture planted, hashed and then deleted from the live volume: „A(z) calibre-web: **0 fájl visszaállítva** … — az alkalmazás újraindult.", `ok=true`, and `ls` reported the directory absent. On 0.218.0, same fixture, same steps: „A(z) calibre-web: **0 fájl és 1 adatkötet visszaállítva** …" and **5/5 files byte-identical**, both Hungarian accented filenames included. **Four red-proofs, each mutation asserted applied:** the volume leg removed returned the silent loss; the count dropped from the message returned the true-but-incomplete sentence verbatim; the unit guard removed was SEEN writing into the live unit; the error swallowed let a partial replay report success. **The unit-guard proof initially PASSED against the mutation** — the fingerprint had been narrowed to the volume directory and was blind to a placement writing into the unit root (the R-181 class, in the test rather than the code). Widened, and it convicts. **NOT reached by this fix, and it is the blocker:** the 40 apps that declare no data drive still cannot run this restore at all — **R-356** refuses first. | CC |
|
||||
| **R-355** | **`paperless-ngx`'s PostgreSQL is dumped into a directory for a stack that does not exist, so its unit has never contained a database dump — and the destructive restore therefore takes no safety dump and tells the customer the app has no database.** `deriveStackName("paperless-postgres", known)` (`controller/internal/appbackup/dbdump.go:770-798`) strips the `postgres` suffix to `paperless`, finds it is NOT a known stack, finds no known stack is a prefix of the container name, and then **returns the unresolved candidate anyway** — no warning, no refusal. Observed live 2026-08-21: the dump (284 617 B, 72 tables, valid) landed in `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/` on the SYSTEM drive while the app's unit sits on `/mnt/felhom-drives/hdd_1` recording `"db_dumps": null`. The orphan directory is outside the app's off-site capture set, so the only copy of that dump is on the machine it protects. `writeSafetyDump` filters on the same wrong name, so `hasDB` is false: **no undo is taken and the fail-closed refusal cannot fire** — verified, `find /mnt -name "pre-restore-*"` empty before AND after a destructive restore. Outcome said „0 fájl visszaállítva … Ennek az alkalmazásnak nincs adatbázisa." **A catalogue-wide sweep of every DB-bearing template shows this is the ONLY affected app (1 of 53).** | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22** (controller **v0.218.0**) | — | **Neither candidate: a third that removes the guessing.** Every container the controller starts carries `com.docker.compose.project`, and that label IS the stack name BY CONSTRUCTION — compose is run with `cmd.Dir` set to `/opt/docker/stacks/<stack>` and never `-p` (`stacks/manager.go:1218`). `resolveStackName` prefers it whenever it names a deployed stack; `deriveStackName` stays as the fallback for containers not started by compose; an attribution that resolves to NO known stack is now **loud** instead of silently returned. **The fix is in the CONTROLLER, not the catalogue** — renaming the container would have fixed this one app and left the guessing for the next. **The sweep was proven before its answer was trusted:** a second mismatch planted in a scratch copy (`kimai-db`→`timetrack-db`) was convicted by name, removal returned it to 1, and a catalogue with every mismatch removed exits **0** — so "1" is not a stuck value. **1 affected app of 53**, 15 DB containers checked. **PROVEN LIVE on `demo-hp`, negative control first.** 0.217.0 at 09:37: unit `db_dumps = None`, dump refreshed into the phantom `…/primary/paperless/db-dumps/paperless-postgres.sql`. 0.218.0 at 09:45: `db_dumps = ['paperless-ngx-postgres.sql']` **inside the app's own unit**, and present in the off-site snapshot for the first time. Scenario B: the destructive restore said „**0 fájl és 3 adatkötet és az adatbázis visszaállítva**" and wrote `pre-restore-20260822T075658Z-paperless-ngx-postgres.sql` (312 957 B) where **zero** undo copies had existed. Scenario C: with the undo made impossible the restore REFUSED, the marker kept its mutation and the container's `StartedAt` was unchanged — **the app was never stopped**. Scenario D: `romm-mariadb.sql` and `kimai-mariadb.sql` unchanged in name and location. **Three red-proofs, each asserted applied:** the name fix reverted printed both divergent paths; the message predicate reverted returned the false „nincs adatbázisa" sentence verbatim; the refusal removed was seen letting a restore proceed with no undo. **The dumps already written under the wrong name are NOT deleted** — see **R-367**. | CC |
|
||||
| **R-356** | **The off-site restore refuses for all 40 no-drive apps, says the app "is not installed" when it is running, and then gives an instruction those apps make impossible.** `ReconstituteFromOffsite` refuses when `GetStackHDDPath(stack)` is empty (`offbox_reconstitute.go:208-227`); for a 40-class app that is ALWAYS empty, because they are offered no storage field at deploy time (R-352's own measurement). Observed 2026-08-21 22:21 on `privatebin` while it was `deployed=true, state=running, healthy`: „a(z) privatebin nincs telepítve, ezért nincs hová visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra az alkalmazást ugyanerre a helyre…". **The predicate is "has an HDD path"; the sentence says "is not installed"; for this class they are different things**, and the remedy offered cannot be carried out. This is also what the 2026-08-21 afternoon OpenGist journey hit before falling back to the local restore (R-353). | **OPEN — HIGH** | — | Separate the two questions. A 40-class app has a destination — the system data path — and the restore already knows it. | CC |
|
||||
| **R-356** | **The off-site restore refuses for all 40 no-drive apps, says the app "is not installed" when it is running, and then gives an instruction those apps make impossible.** `ReconstituteFromOffsite` refuses when `GetStackHDDPath(stack)` is empty (`offbox_reconstitute.go:208-227`); for a 40-class app that is ALWAYS empty, because they are offered no storage field at deploy time (R-352's own measurement). Observed 2026-08-21 22:21 on `privatebin` while it was `deployed=true, state=running, healthy`: „a(z) privatebin nincs telepítve, ezért nincs hová visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra az alkalmazást ugyanerre a helyre…". **The predicate is "has an HDD path"; the sentence says "is not installed"; for this class they are different things**, and the remedy offered cannot be carried out. This is also what the 2026-08-21 afternoon OpenGist journey hit before falling back to the local restore (R-353). | **OPEN — HIGH** | — | Separate the two questions. A 40-class app has a destination — the system data path — and the restore already knows it. **RE-CHECKED 2026-08-22 against the corrected placement framing: this row SURVIVES UNCHANGED and is strengthened by it.** The 40-class has no `HDD_PATH` *because the architecture puts its hot data inside the guest by design* (`01-topology-and-trust.md:150-152`) — so an absent HDD path is the NORMAL state for most of the catalogue, and a restore that reads it as "the app is not installed" is misreading a correct configuration, not reporting a misconfiguration. The one sentence in this row that leaned on the old framing — "because they are offered no storage field at deploy time (R-352's own measurement)" — should read: *because these apps have no bulk volume to place, so the field correctly does not exist.* | CC |
|
||||
| **R-357** | **The DESTRUCTIVE restore has no free-space gate; the three that exist are all on non-destructive paths.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the gates sit at `offbox_restore.go:231` (scratch restore), `:297` (prepare-full) and `:423` (place-to-live). Proven 2026-08-21 23:11 with 300 KB free and 1 MB to write: it stopped `paperless-ngx`, failed halfway (`rsync … No space left on device (28)`), left the data directory holding **2 of 5** planted entries, and restarted the app. The message is honest but is raw rsync output. | **OPEN — MEDIUM** | — | Same gate, same wording as `:297`, before the stop. | CC |
|
||||
| **R-358** | **A FAILED scratch restore leaves a partial copy that the product then offers as a full restore source — and the destructive restore runs from it and reports success.** `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only whether the directory exists and is non-empty; its comment defers completeness to `PlaceOffsiteRestore`, which stats top-level placements, not files. Proven 2026-08-21 22:54-22:56 against a deliberately corrupted store: the restore failed honestly (`ciphertext verification failed`, 54 files, 15 of 16 originals), the wizard then offered „Teljes visszaállítás indítása", and pressing it reported `ok=true`. **The failure is detected and then forgotten.** | **OPEN — MEDIUM** | — | Record the failure against the scratch and refuse to place from it until it is re-prepared. | CC |
|
||||
| **R-359** | **The off-site restic store is never verified by anything, ever.** The complete set of restic verbs in the controller is `restore, snapshots, backup, unlock, stats, init, forget, prune, cat` — **no `check`**. The agent's `RestoreTest` is PBS-tier only. Established 2026-08-21 by deliberately corrupting one pack: `restic check` catches it immediately („ciphertext verification failed", „Fatal: repository contains errors"), and the product only meets the damage when a customer is already trying to recover. | **OPEN — MEDIUM** | — | A periodic `restic check` (structure) with an occasional `--read-data`, reported like any other backup verdict. Note PBS already has verify jobs; this is the tier that does not. | CC |
|
||||
@@ -669,6 +669,14 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-367** | **The database dumps already written under the wrong name are stranded, and nothing will ever collect them.** R-355's fix sends `paperless-ngx`'s dump to the right place from now on; it does not move the ones already written. On `demo-hp` that is `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql` (312 381 B, 2026-08-22 07:38, the last pre-fix cycle). **Nothing deletes them and that is by design, not by luck:** the F5 stale-primary prune (`backup.go:1248`) skips any directory whose name is not a deployed app, under the guard *"an undeployed app's last backup is still its restore point"* — verified still present after the fix. They are equally invisible to the off-site push, which resolves paths from the app's own unit. **They CAN be adopted, by hand:** move the file to `…/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql` and it becomes a readable restore point for that app. **It is deliberately not automatic.** The adopted dump would sit beside volume tars taken at a different time, i.e. an INCOHERENT pair — the exact shape R-43/R-44's coherence stamp exists to make visible — and a controller that silently relocates a customer's data on upgrade is a migration, not a fix. **Filed rather than done**, because whether a stale orphan is worth adopting at all is a judgement about one machine's history, not a rule. | **OPEN — LOW** | follows R-355 | Decide per box: adopt (and say the pair is skewed), or delete deliberately. Neither on an upgrade path. | Viktor rules, CC executes |
|
||||
|
||||
| **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** Both box checkers (`OffsiteBoxChecker` over the Hetzner API, `PBSDRBoxChecker` over ep0's `usage` op) held their last snapshot and returned quietly on a failed fetch. That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. During the 2026-08-18 ep0 incident the hub said nothing for the entire outage; the only mails came from the boxes' own backup failures, and **only because the WEEKLY offsite run happened to fall inside the window**. Two days earlier, nothing would have fired at all | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | — | **PROVEN-LIVE still owed.** No real or constructed outage has exercised the emit path end to end, and one cannot be manufactured without making ep0 or the Hetzner API unreachable — ep0 is Tier 2 protected, so that is forbidden. The honest route is a constructed outage against a scratch hub instance with the tenantsync client pointed at a blackholed address. **Do not close this row on the unit tests** | CC |
|
||||
| **R-368** | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** R-352 and `SPEC-app-data-placement-2026-08-21.md` §2.2 stated *"the deploy route never reads it"*, from `grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' internal/stacks/deploy.go internal/stacks/manager.go` → nothing. **That grep searched Go files only and never the templates.** `internal/web/templates/deploy.html:612` reads `.IsDefault` directly off each `DeployStoragePath` (which embeds `settings.StoragePath`, `web/handlers.go:89-99`) and **pre-selects the default drive for a new deploy**: `{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}`. So `// new apps use this by default` (`settings.go:453`) is **IMPRECISE ABOUT THE MECHANISM, NOT FALSE** — nobody calls `GetDefaultStoragePath()` on that route, but the value is honoured. The customer-facing label promises exactly this and no more: **„Legyen alapértelmezett új telepítéseknél"** (`storage.html:469`). **THE RESIDUAL, and it is the whole finding:** the default lives in the TEMPLATE, not in the server. `POST /api/stacks/<n>/deploy` accepts `values` verbatim; omit `HDD_PATH` and `withPathVars` (`stacks/deploy.go:584`) receives `""` and no default is applied. **That is why the invariant has no test — there is nothing server-side to test.** | **OPEN — LOW** | corrects R-352(2); supersedes SPEC §2.2 | Either move the default into the server so the API and the form agree and a test can pin it, or reword the comment to say the template owns it. Do not "fix" the behaviour: it is correct on the path customers use. | CC |
|
||||
| **R-369** | **There are TWO registers, only one calls itself the source of truth, and work filed in the other is invisible to every standing rule that says "grep the register".** `OPEN-ITEMS.md` opens *"the single source of truth for open work"*; `ROADMAP.md` opens *"the prioritized decision log of planned/open work"*. Both hold open work. **Measured 2026-08-22: 72 `R-` ids appear in ROADMAP and not in OPEN-ITEMS; 29 of those are not marked shipped/closed/killed.** Most are feature ideas that arguably belong only in ROADMAP — but some are FINDINGS: R-30/R-31/R-32 (all marked P2-HIGH), R-35, R-40, R-76, R-79, R-25, R-49, R-10 and **R-107**. **The cost is measured, not hypothetical:** R-107 — *"No offsite action unpacks the named-volume tars Tier-3 captures on every run"* — was filed **2026-07-28, READY, severity M**, and re-stated in `07-backup-architecture.md:337,902`. It is absent from OPEN-ITEMS. On 2026-08-21 an overnight drill rediscovered it from scratch by planting files and watching them not come back, and it shipped as R-354 on 2026-08-22. **25 days.** **The dating is exact and makes it a rule violation rather than a gap in the rules:** OPEN-ITEMS.md was created and the template gained its *"single source of truth"* bullet on **2026-07-27** (`655b69f`); R-107 went into ROADMAP alone on **2026-07-28** (`070b0ce`) — the day after. **The risk is NOT double-minting** (ids are shared; OPEN-ITEMS' ceiling 368 exceeds ROADMAP's 331) — it is that a session greps one file, finds nothing, and redoes the work. | **OPEN — HIGH** | R-107 (ROADMAP-only), R-354 | Decide the division and make it mechanical: either one register, or a gate that fails when a ROADMAP row that is not shipped/killed has no OPEN-ITEMS counterpart. Triage the 29 first — most are ideas, a minority are findings. | **Viktor rules**, CC executes |
|
||||
| **R-370** | **PROCESS: between 2026-08-19 and 2026-08-22 the reviewing side called a documented architectural decision a defect, in four places, because it read the register and live source and never `documentation/architecture/`.** The decision: `.felhom.yml` classes each volume **hot** (DB/config/cache → fast storage, **enforced**) or **bulk** (media/files → may be slow) — `01-topology-and-trust.md:150-152`. The 40 templates with no `HDD_PATH` are all-hot apps; the 13 with one are the bulk apps. **The four instances**, all authored 2026-08-21: `SPEC-app-data-placement-2026-08-21.md` §1 (absence of a storage field framed as the finding), §2.3 (*"the defect … is that these apps are on the system drive in the first place"*), §5 (posed a question the architecture already answers), and **R-352 point (3)** (compared it to Tier 2's same-disk refusal). The prompt that commissioned this correction said *three*; the evidenced count is **four**. **Not a personal failure — a missing step:** nothing in `PROMPT-TEMPLATE.md` §4 required naming the architecture document for the area being touched. §4 item 4 named only `02-controller-module-map.md`, a package classification, and the S-1 rule at the end governs UPDATING a design doc, not READING one first. Fixed in the same session — see the template's new §4 rule. | **CLOSED — corrected 2026-08-22** | R-352 (re-framed), R-369 | The record is corrected in place with the framing marked rather than deleted, per the standing rule that a document which quietly changes its mind teaches nobody. | CC |
|
||||
| **R-371** | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
|
||||
| **R-372** | **A Tier-2 copy that has NEVER been produced because its source path is missing is not surfaced prominently to the operator.** Written down **2026-07-15**, in `audits/CAMPAIGN-6E-2026-07-15.md:128` (F-6E-1), whose disposition ends *"Optional product idea: surface 'tier-2 has never produced a copy (source missing)' more prominently in the operator UI — not filed."* The finding it sits on was correctly judged demo-data churn rather than a product defect (the code warns loudly and does not silently succeed), but the surfacing idea was never carried anywhere. **Age when filed: 38 days — the oldest gap this sweep recovered.** | **OPEN — LOW** | — | Decide whether "never produced a copy" deserves its own operator surface, distinct from "last copy failed". | CC |
|
||||
| **R-373** | **`SysDataGrowGB` is the intended lever for the system-data volume, it works, and nothing sets it.** Written down 2026-08-02 in `audits/SPIKE-recovery-unit-space-2026-08-02.md:230-232`, under an explicit *"### Not filed"* heading: the 20 G / 50 G mismatch was ruled a tier-sizing decision rather than a defect, *"`SysDataGrowGB` is the intended lever and it works; nothing sets it."* A lever with no caller is the same shape as R-368's comment — a setting that names a behaviour nothing invokes. **Age when filed: 20 days.** | **OPEN — LOW** | R-368 (same shape) | Either wire it to something an operator can reach, or remove it and record the sizing decision where a reader will meet it. | CC |
|
||||
| **R-374** | **Three C1 refusal cases were judged borderline, left unfiled, and never named — so nobody can re-open the judgement.** `audits/CAMPAIGN-12-class-sweep-2026-08-08.md:123`: *"'Names a route' is a judgement, not a predicate — two readers could disagree on the borderline cases, and three of the 19 were called borderline and left unfiled."* The disclosure is honest and is exactly the right thing to write; **what is missing is WHICH three.** An unnamed borderline case cannot be re-judged by a second reader, which is the only remedy a judgement call has. **Age when filed: 14 days.** | **OPEN — LOW** | — | Name the three in that document, or file them as one row listing them. No code. | CC |
|
||||
| **R-375** | **A PBS datastore signal was noted and explicitly not filed.** `audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:171`: *"Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here."* Recorded so the note has a number and stops depending on someone re-reading that report. **Age when filed: 4 days.** | **OPEN — LOW** | — | Confirm the cause on ep0 the next time it is touched; it is a read-only check. | CC |
|
||||
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
|
||||
@@ -1,5 +1,38 @@
|
||||
# SPEC — where an app's data is placed, and who decides
|
||||
|
||||
> ## ⚠ CORRECTED 2026-08-22 — THE MEASUREMENTS STAND, THE FRAMING WAS WRONG
|
||||
>
|
||||
> **This document called a documented architectural decision a defect.** It is left in place rather
|
||||
> than rewritten, because a document that quietly changes its mind teaches nobody. Every measurement
|
||||
> below is good and was re-checked on 2026-08-22; the corrections are marked inline as
|
||||
> **[CORRECTED 2026-08-22]**.
|
||||
>
|
||||
> **What it got wrong.** It treated the absence of a storage field on 40 of 53 templates as a choice
|
||||
> being denied. There is no choice to deny. The architecture states the placement rule and states it
|
||||
> as *enforced*:
|
||||
>
|
||||
> > *"**App data placement is per-volume, not per-app:** `.felhom.yml` classifies each volume **hot**
|
||||
> > (DB/config/cache → fast storage, **enforced**) vs **bulk** (media/files → may be slow). A photo
|
||||
> > app's DB stays on SSD while its blobs go to the USB."*
|
||||
> > — `documentation/architecture/01-topology-and-trust.md:150-152`
|
||||
>
|
||||
> The 40 templates are **all-hot apps**: their data *is* app internals — database, config, cache —
|
||||
> and hot data belongs on fast storage inside the guest, by design. The 13 with a configurable path
|
||||
> are the media and library apps, where **bulk** content legitimately belongs on an attached drive.
|
||||
> The split is not 40 apps missing a feature; it is the hot/bulk contract working.
|
||||
>
|
||||
> **And the guest volume is deliberately its own volume**, so that filling it cannot stop the
|
||||
> operating system: the golden ships a small OS rootfs plus a single data volume at `/var/lib/felhom`
|
||||
> (`felhom-agent/configs/build-golden.sh:29-40`), and a capture that would exhaust it is refused per
|
||||
> app rather than allowed to stop the container runtime
|
||||
> (`documentation/architecture/00-capability-map.md:94`, PROVEN-LIVE 2026-08-03).
|
||||
>
|
||||
> **One real defect survives the correction and is unchanged:** `settings.go:453` says
|
||||
> `// new apps use this by default` and no app has ever used it. That is §2.2 below and it is now
|
||||
> filed on its own, scoped by measurement rather than by assumption — see the register.
|
||||
>
|
||||
> **What still needs a ruling: far less than §5 implies.** See the corrected §5.
|
||||
|
||||
**Filed 2026-08-21. Status: SPECIFICATION ONLY — nothing here is implemented, and nothing here may be
|
||||
implemented without the operator's ruling. It deserves its own session.**
|
||||
|
||||
@@ -20,6 +53,15 @@ The first hypothesis — that a customer had typed a bad path, or had failed to
|
||||
|
||||
> **The configured default data store is not consulted on the deploy route at all.**
|
||||
|
||||
> **[CORRECTED 2026-08-22]** The observation is right; the conclusion drawn from it was not. There is
|
||||
> nothing to choose **because these apps have no bulk volume to place** — the hot/bulk contract
|
||||
> (`01-topology-and-trust.md:150-152`) puts their data on fast storage inside the guest by design, and
|
||||
> OpenGist is a pure hot app. The sentence in the quote above is still true as a statement about the
|
||||
> deploy route, and it is still a defect **for the 13 templates that DO take a path** — that is the
|
||||
> real remainder, established in Part 4 of the 2026-08-22 session and filed separately. For the other
|
||||
> 40 it is not a defect at all: a default drive cannot be "consulted" for an app that has no path to
|
||||
> put it in.
|
||||
|
||||
## 2. The four measured facts
|
||||
|
||||
### 2.1 The affected class is 40 of 53 catalogue templates
|
||||
@@ -77,6 +119,29 @@ Live:
|
||||
So for these 40 apps **the data and its nearest copy sit on the same physical device**, reached by a
|
||||
customer doing nothing wrong.
|
||||
|
||||
> **[CORRECTED 2026-08-22] — and be precise about which risk is real.**
|
||||
>
|
||||
> **The device claim is true. The framing "by accident" was not.** Since R-165 (golden `build-golden.sh`
|
||||
> v3.0.0, proven live 2026-08-03) the guest carries a small OS rootfs plus **ONE** data volume at
|
||||
> `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume**, and
|
||||
> the script says so in terms — *"There is deliberately no mp1"* (`build-golden.sh:99`). The
|
||||
> `mp0`/`mp1` split this document's era assumed was retired six days before this document was written
|
||||
> and is on no current box.
|
||||
>
|
||||
> **Which risk is real, stated exactly:**
|
||||
> - **A failure of the physical disk loses both** the app's data and its first-tier copy. **REAL, and
|
||||
> it is what the off-site and whole-machine tiers exist for.** It is not specific to the 40-class:
|
||||
> a drive-resident app's unit lives beside its data on the drive too, deliberately, so a restore
|
||||
> needs the drive and nothing else.
|
||||
> - **A full data volume taking out the operating system. NOT REAL, and it was the overstated one.**
|
||||
> The OS rootfs is a separate volume, and the capture floor refuses a capture that would exhaust the
|
||||
> data volume rather than letting it stop the container runtime (`00-capability-map.md:94`). Watched
|
||||
> working on 2026-08-21: with `/var/lib/felhom` at 99% the reserve refused one app per run and told
|
||||
> the hub, while all 15 containers stayed healthy.
|
||||
>
|
||||
> So the sentence *"The defect is not Tier 1's rule; it is that these apps are on the system drive in
|
||||
> the first place"* below is **withdrawn**. Being on the guest's data volume is where hot data belongs.
|
||||
|
||||
The project already treats that posture as unacceptable — for the *other* tier. Tier 2 refuses it
|
||||
outright at `internal/backup/tier2.go:329`, recording
|
||||
`a kiválasztott cél ugyanazon a fizikai lemezen van`. Tier 1 has no such notion, and for a
|
||||
@@ -129,7 +194,28 @@ selected drive for the 13. Visibility only. **No placement changed. Nothing was
|
||||
|
||||
## 5. The open question this document exists to hand over
|
||||
|
||||
> **[CORRECTED 2026-08-22] — THIS QUESTION IS ALREADY ANSWERED, AND THE ANSWER IS NO.**
|
||||
>
|
||||
> The architecture answers it: hot data (DB/config/cache) belongs on fast storage inside the guest and
|
||||
> that placement is **enforced**, not preferred (`01-topology-and-trust.md:150-152`). A named-volume
|
||||
> app is all-hot. Moving it to the attached drive would be moving hot data onto storage the same
|
||||
> document classes as possibly-slow, and would put it outside the guest vzdump that currently protects
|
||||
> it (`01-topology-and-trust.md:154-156`).
|
||||
>
|
||||
> **So points 1, 2 and 3 below are withdrawn as a live question.** They remain a good record of what a
|
||||
> change would have cost, which is why they are not deleted. Point 3 was already arguing against the
|
||||
> premise of the question it sat under.
|
||||
>
|
||||
> **WHAT ACTUALLY STILL NEEDS YOUR RULING FROM THIS DOCUMENT: point 4 only, and it is smaller than it
|
||||
> looks** — `IsDefault`'s comment must become true or go away, for the 13 apps it could apply to. That
|
||||
> is filed as its own row with its scope measured rather than assumed. Point 5 (the Drives page count)
|
||||
> needs **no ruling**: the count is honest, it is only easy to misread, and that is a wording question
|
||||
> the design system already owns.
|
||||
>
|
||||
> **Everything else in this section: nothing to decide.**
|
||||
|
||||
**Should a named-volume app's data live on the default data drive rather than the system drive?**
|
||||
*(superseded — see the correction directly above)*
|
||||
|
||||
Points the next session must settle, each of which is a reason this was not decided tonight:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user