Files
felhom.eu/REPORT.md
T
admin fddfe00ce2
gates / gates (push) Successful in 15s
REPORT: record CI run 382/250 by id
2026-08-22 11:18:46 +02:00

18 KiB
Raw Blame History

REPORT — correcting what we mis-called a defect, and finding what we wrote down and never filed (2026-08-22)

Documentation and survey only. No code, no version bump, no bake, no deploy, no machine contacted. The previous report is preserved at documentation/audits/REPORT-v0.218.0-r354-r355-2026-08-22.md.


1. §2's four claims — ALL FOUR HELD. No halt.

# claim verdict
1 07-backup-architecture.md:297 — 53 templates, 52 named volumes, exactly 13 with a backup: block, and those 13 are exactly the ones that bind a configurable path HOLDS, at :296-299. The heading above it reads "Coverage per app class — and an unresolved count", which the sweep then followed
2 _recovery-inventory-2026-07-28.md Tier-3 row — "…unpacked by no offsite action… no single action does it and no UI routes it. This is my own enumeration; it is not currently filed as a finding." HOLDS, verbatim, at :373
3 :392-400 — rootfs, mp0 /var/lib/docker, mp1 /mnt/sys_drive as separate volumes HOLDS, at :392-396 — but it describes a layout no current box has. See below
4 00-capability-map.md:94 — one data volume; a local backup bounded by free space HOLDS, at :94

Claims 3 and 4 describe different layouts, six days apart, and 4 is the live one. The inventory (2026-07-28) records the pre-R-165 split. felhom-agent/configs/build-golden.sh:29-40 — live source — bakes ONE data volume at /var/lib/felhom, with /var/lib/docker and /mnt/sys_drive as two binds of that same volume, and states at :99 that "There is deliberately no mp1".

So the prompt's own §1.2 instruction is out of date: it asked me to say that named volumes and first-tier backups are "on two different guest volumes on one physical disk". They are on one guest volume, by two binds. The disk claim is stated precisely in §4 below.

A correction to my own §2 working. In verifying claim 2 I reported that a literal grep matched :373. It did not — it returned nothing, and what printed was the second grep in the same cell. **not** currently filed does not match not currently filed, because of the markdown bold. The claim still holds (I read the line), but the check I quoted was wrong — and that false zero turned out to matter, see §3.


2. The sweep — three counts, and the control that earned them

Method (scripts not committed; it is a throwaway, and the terms below are the durable part): enumerate every survey-class document — inventory / audit / spike / review / diagnosis / recon / campaign / report / validation / rehearsal / walk / finding / spec / triage — under documentation/{audits,architecture,pilot,backlog}, excluding runbooks, templates, READMEs, CHANGELOGs and ROADMAP, because those instruct rather than conclude. For each, ask whether any OPEN-ITEMS.md row cites it by filename, then search every line for the shapes an unfiled gap takes.

count value
documents examined 113 (survey-class; 114 after this session added one)
gaps found (documents saying, in their own words, that something was not filed) 14 statements
of those, already filed 2 — and the other 12 break down in §3

The search terms, so the next person can widen them:

not (currently )?filed      never filed        no finding          not a finding
unresolved                  unfiled            disagreement        no single action
nothing (does|routes|reads|consults)           never been          should be
ought to                    is not (currently )?(tested|proven|observed|covered|wired)
has (never|not) been (seen|observed|walked|proven|done)
no (test|row|register) (pins|covers|cites)     TODO       FIXME       ⚠       not covered       gap

Every word-gap in the first four patterns is [*_~ ]*`, i.e. markdown emphasis is tolerated — see §3.

The positive control — and it convicted the sweep twice before it convicted the corpus

Plant → find → remove → fail to find, on a scratch copy of the whole documentation/ tree:

step strongest-shape hits
unplanted scratch copy 14
a synthetic gap planted, bolded and on one line 15 — convicted by name and quoted back
scratch discarded, real corpus re-run 14

The control found two defects in my own sweep before it found anything in the corpus, and both were false zeros of exactly the class this session is about:

  1. Markdown emphasis broke the strongest pattern. The single most important line in the corpus is written it is **not** currently filed as a finding, and not (currently )?filed does not match it. Fixed by tolerating [*_~ ]*` between words.
  2. The reporter re-searched a truncated copy of the line. Hits were stored as line[:200]; that line is a long table row and the phrase sits past column 200, so the detector caught it and the report dropped it. Fixed by matching the full line and truncating only for display. This alone moved the count 12 → 14.

A sweep whose first two runs would have under-reported by 2 is exactly why the control is mandatory.


3. What the 14 were, judged

verdict n which
never a gap — a reasoned "no finding" 4 _design-review.md:9 (nothing deferred), CAMPAIGN-11:501 (a measurement was honest), and the two explicit ### Not filed sections in CAMPAIGN-10-two-storage-soak and SPIKE-recovery-unit-space — each item disposed with a reason. This is good practice, not a failure, and it is what the rest should look like.
already filed 2 REPORT-DRILL-backup-truth:321 (its Part 5 became R-360 and R-364, filed the same session) and REPORT-DRILL:569 (the hub's controller-version blindness — the register already covers it, 2 hits; not re-filed)
still open → NEWLY FILED 5 R-371, R-372, R-373, R-374, R-375
the headline 3 _recovery-inventory:373, :1043 and 07-backup-architecture.md:337 — all three the same gap, and it turns out it was given a number

The headline: it was filed. In the other register.

The gap the 2026-08-21 drill rediscovered was already numbered R-107, with a full write-up:

ROADMAP.md:122 — "No offsite action unpacks the named-volume tars Tier-3 captures on every run. ReconstituteFromOffsite skips the unit outright … PlaceOffsiteRestore places it only when the live unit is ABSENT" — M, READY, 2026-07-28

and cross-referenced twice in 07-backup-architecture.md (:337, :902). It is absent from OPEN-ITEMS.md, which opens with the words "the single source of truth for open work".

The dating makes it a rule violation, not a gap in the rules. OPEN-ITEMS.md was created and the template gained its "single source of truth" bullet on 2026-07-27 (655b69f). R-107 went into ROADMAP alone on 2026-07-28 (070b0ce) — the day after.

Measured scope: 72 R- ids are in ROADMAP and not in OPEN-ITEMS; 29 are not marked shipped, closed or killed. Most of the 29 are feature ideas, which arguably belong only in ROADMAP. A minority are findings: R-30, R-31, R-32 (all P2-HIGH), R-35, R-40, R-76, R-79, R-25, R-49, R-10, R-107.

The risk is not double-minting — the two files share ids and OPEN-ITEMS' ceiling (368) is above ROADMAP's (331). The risk is rediscovery: a session greps one file, finds nothing, and redoes the work. That is precisely what happened, and it cost an evening and a night. Filed as R-369, HIGH.


4. Part 1 — the record corrected

The specification

SPEC-app-data-placement-2026-08-21.md now opens with a marked ⚠ CORRECTED 2026-08-22 block and carries four inline [CORRECTED 2026-08-22] marks. Nothing was deleted; every measurement stands.

What it got wrong: it treated the absence of a storage field on 40 templates as a choice being denied. The architecture states the rule and states it as enforced:

"App data placement is per-volume, not per-app: .felhom.yml classifies each volume hot (DB/config/cache → fast storage, enforced) vs bulk (media/files → may be slow)." — 01-topology-and-trust.md:150-152

The 40 are all-hot apps; the 13 are the bulk apps. And the deploy page has been telling the customer exactly this the whole time — deploy.html:624-625: „A kiválasztott meghajtón az alkalmazás fájljai (média, dokumentumok) tárolódnak. Az adatbázis a gyors belső SSD-n fut."

The disk claim, stated precisely

  • REAL: a physical-disk failure loses the app's data and its first-tier copy together. That is what the off-site and whole-machine tiers exist for, and it is equally true of a drive-resident app, whose unit sits beside its data on the drive deliberately so a restore needs the drive and nothing else.
  • OVERSTATED: a full data volume stopping the operating system. The OS rootfs is a separate volume, and the capture floor refuses per app before exhaustion (00-capability-map.md:94). Watched working 2026-08-21 with /var/lib/felhom at 99%: the reserve refused one app per run, told the hub, and all 15 containers stayed healthy.
  • WITHDRAWN: the comparison to Tier 2's same-disk refusal (tier2.go:329). Tier 2 refuses a second copy on the same disk; Tier 1's unit is meant to sit beside the data.

Rows re-framed

  • R-352 — measurements (1)–(4) all stand; the conclusions drawn from (1) and (3) are marked re-framed. (2) is now filed on its own as R-368; (4) needs no ruling.
  • R-356 — re-checked and it SURVIVES UNCHANGED, strengthened. Because the 40-class correctly has no HDD_PATH, a restore that reads that as "the app is not installed" is misreading a correct configuration. One sentence in it that leaned on the old framing is corrected in place.

My own errors, named — R-370

Four instances, not three. The prompt said three; the evidenced count is four, all authored 2026-08-21: SPEC §1, SPEC §2.3, SPEC §5, and R-352 point (3). Recorded as a process failure with its mechanism — the register and live source were read, documentation/architecture/ was not — and the missing step is now in the template. Also written into CONTEXT.md.


5. What still needs the operator's ruling — from this specification, almost nothing

  • The placement question is ANSWERED and the answer is no. Points 1, 2 and 3 of the spec's §5 are withdrawn as a live question; hot data belongs where it is.
  • Point 4 is the only thing left, and it is smaller than it looked — see R-368 below.
  • Point 5 needs no ruling. The Drives count is honest; it is a wording question the design system already owns.

One new thing does want your ruling, and it is R-369: whether the two registers become one, or whether a gate enforces that a not-done ROADMAP row has an OPEN-ITEMS counterpart. Triage the 29 first — most are ideas, a minority are findings.


6. Part 4 — the three answers, and the finding is smaller and different than believed

1. Is the default store consulted at deploy time, for the 13 apps that take a path? — YES. internal/web/templates/deploy.html:612:

{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}

For a new deploy the default drive is pre-selected. DeployStoragePath embeds settings.StoragePath (web/handlers.go:89-99), so .IsDefault resolves.

So // new apps use this by default (settings.go:453) is IMPRECISE ABOUT THE MECHANISM, NOT FALSE. The earlier claim — "the deploy route never reads it" — is wrong, and it is wrong for the same reason as everything else this week: the grep behind it (grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' deploy.go manager.go) searched Go files and never the templates.

The residual, and it is the whole finding: the default lives in the template, not the server. POST /api/stacks/<n>/deploy takes values verbatim; omit HDD_PATH and withPathVars (stacks/deploy.go:584) gets "" and no default applies. That is why the invariant has no test — there is nothing server-side to test. Filed R-368, LOW.

2. What the label promises the customer, quoted: storage.html:469 — „Legyen alapértelmezett új telepítéseknél" ("Be the default for new installations"), with the badge „Alapértelmezett" at :34. The promise is kept: it is the default for new installs, which is exactly what the pre-selection does. It does not promise that every app's data goes there.

3. What the Drives count counts: countAppsUsingPath (web/handlers.go:2162-2175) counts deployed apps where appCfg.Env["HDD_PATH"] == storagePath. A named-volume app has no HDD_PATH, so it can never be counted — by construction, not by accident. The number is honest; it means "apps that place bulk data on this drive", not "apps using storage". Worth one sentence of wording, not a ruling.


7. Part 3 — the template's two new rules

It had neither. enumerat → 0 hits, register row → 0 hits; architecture appeared 9 times but §4 item 4 named only 02-controller-module-map.md, and the S-1 rule at the end governs updating a design doc, not reading one first. No halt. There is no separate authoring companion.

Rule 1, in §4 where files are read — abridged; the full text carries a file→area table for all eight architecture documents:

THE ARCHITECTURE DOCUMENT FOR THE AREA THIS TASK TOUCHES — NAME IT AND SAY WHAT IT SAYS. Not "read the architecture folder": name the file, and state in one line what it rules about this area. A prompt that cannot name one says so explicitly, and that absence is itself recorded — an undocumented architectural decision is how a deliberate design gets "fixed" by someone who did not know it was one.

Three sources, in this order, before any claim: the architecture folder holds the REASONING, the register holds the WORK, source holds the TRUTH.

And the test that catches it: is what I am about to call a defect something we chose? If it was chosen and the choice is wrong, that is a proposal to change a decision — it goes to the operator as a decision, not filed as a bug. Cost of learning this (R-370): between 19 and 22 August a documented placement decision was called a defect in four places.

Rule 2, on the OPEN-ITEMS.md bullet, where a survey session lands:

AN ENUMERATED GAP BECOMES A ROW, IN THE SAME SESSION. PROSE IS NOT A RECORD.

This binds surveys, inventories, spikes, reviews and diagnoses, not only implementation sessions… If a document says a thing is missing, unhandled, unreachable or "not currently filed", it does not leave the session as prose. It leaves as a row here, with a rank and an owner. Writing "not filed" is not a disposition; it is a note that the work was seen and dropped.

A row in ROADMAP.md alone does not satisfy this … invisible to every standing rule that says "grep the register before minting" (R-369).

The cost, recorded so the rule can be narrowed later rather than becoming permanent by accident: R-107 … was enumerated on 2026-07-28 … and never entered here. It was rediscovered from scratch 25 days later by an overnight drill, and shipped as R-354.

Two rules, one place each, ~40 lines added to a 521-line document.


8. Rows opened — ceiling R-367 → R-375

id rank what first written down age
R-368 LOW The storage default does apply at deploy time (deploy.html:612); the earlier "never reads it" was wrong. Residual: the default lives in the template, not the server, so the API has none and nothing server-side can be tested 2026-08-21 (as the wrong claim) 1 d
R-369 HIGH Two registers. 72 ids in ROADMAP only, 29 not done, some of them findings. R-107 sat there 25 days and was rediscovered by a drill 2026-07-28 25 d
R-370 CLOSED Process: a documented decision called a defect four times; architecture folder never read. Rule now in the template 2026-08-19 3 d
R-371 LOW The off-site tier is the only tier that announces nothing on success 2026-08-05 17 d
R-372 LOW A Tier-2 copy that has never been produced (source missing) is not surfaced distinctly 2026-07-15 38 d — the oldest
R-373 LOW SysDataGrowGB is the intended sizing lever, it works, and nothing sets it 2026-08-02 20 d
R-374 LOW Three C1 refusal cases judged borderline, left unfiled and never named, so nobody can re-judge them 2026-08-08 14 d
R-375 LOW A PBS datastore audit signal noted and explicitly not filed 2026-08-18 4 d

Oldest gap recovered: R-372, written down 2026-07-15, 38 days. Most consequential: R-369, because it explains the other seven.


9. What was dropped, and observations

Dropped: nothing. Parts 1, 2, 3 and 4 all completed. Part 4 was droppable-first and was done.

Observations — noticed, not acted on:

  • The two ### Not filed sections are the model. CAMPAIGN-10-two-storage-soak:383 and SPIKE-recovery-unit-space:228 list what they decided not to file and why, item by item. That is a disposition, not a shrug. If the new rule needs an exemplar, those are it.
  • 07-backup-architecture.md:337 already names this gap and points at R-107, so the architecture doc was right and current while the register was empty. The document was not the weak link.
  • The _recovery-inventory cites 07-backup-architecture.md:81 for a phrase that is no longer there (grep 'missing-only merge' → only the inventory's own copy). A stale line-number citation; not filed, because the doc it points into has since been rewritten wholesale and the claim it supported is now handled by R-354.
  • The sweep's "cited by no row" count (89) is not a defect count. Many audits are cited by the capability map, by where-felhom-stands.yaml or by other reports rather than by a register row. It is a candidate filter, and it earned its keep only in combination with the shape search.
  • REPORT-hub-blindness.md exists at the repo root and is cited by no register row, but the register does cover its subject (2 hits). Left alone.

10. CI, confirmed by ID

felhom.eu id=382 / run_number=250, head 091a4b74 — success. Commit 091a4b7, pushed to main, all 10 gates OK, tree clean. No other repo was touched.