R-356 docs: correct R-107 in the architecture, record the design, refresh STATUS, compress the register
gates / gates (push) Failing after 17s
gates / gates (push) Failing after 17s
07-backup-architecture.md: three places said no offsite action unpacks the named-volume tars. R-107 closed in controller v0.218.0; all three corrected with a dated [FACT], the old sentence kept in the past tense. R-102 is NOT closed and the correction says so explicitly. New [DESIGN] paragraph in 6.3: the restore destination is resolved by the same rule as the capture destination, and the wrong-disk refusal applies to apps that have a drive to get wrong. Carries the 13/40 measurement. STATUS.md was internally contradictory - nothing waiting, and one decision waiting, for something the same page recorded as shipped. 218 -> 102 lines; the deciding section now says what happens if nothing is done. R-356 compressed into CLOSED-ITEMS.md; OPEN-ITEMS 327109 -> 325236 bytes. Drill record and 16 evidence files for the live walk on demo-hp.
This commit is contained in:
@@ -525,7 +525,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-349** | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC |
|
||||
| **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:<HUB_PW>@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes |
|
||||
| **R-353** | **A restore reported success having returned configuration and no data — and no screen could have told the customer.** `demo-hp`, 2026-08-21, OpenGist. The off-site reconstitution refused at 16:37:14 (not installed); the person reinstalled and ran the local unit restore, which reported `Restore-from-unit completed: opengist in 8.666896042s`. **The unit it restored from contains `manifest.json` + `compose/{app.yaml,.felhom.yml,docker-compose.yml}` and NOTHING else — `volume_dumps: None`, `db_dumps: None`** — and the off-site snapshot was **182.3 KB**. So the restore returned the app's configuration; there was no data leg in the unit to return, and the outcome said only that it had completed. **A warning beside a success is read as a success, and an unknown must never be drawn as healthy.** **Compounding, and recorded as UNKNOWN rather than fine:** whether the 40-class reaches the off-site tier at all has **not been observed** — `runVolumeDumps` (`backup/backup.go:607+`) covers them on paper, but every unit on the box reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly dump run had happened on a one-hour-old box. | **OPEN — NEXT SESSION'S FIRST ITEM** | — | **Two things, in order. (1)** A restore whose unit carries no `db_dumps` and no `volume_dumps` must **say so in its outcome** — „a mentés csak a beállításokat tartalmazta, adatot nem" — instead of reporting a bare completion. The verdict must consult what was actually placed, not merely that the operation ended. **(2)** Then *prove* the off-site coverage of a named-volume app by running a dump cycle and reading the resulting manifest, rather than inferring it from the gate order. Do not close (1) on the strength of (2) being likely. **(2) IS NOW SATISFIED — drill 2026-08-21.** A dump cycle was run and the manifests read: `privatebin volume_dumps=[privatebin_privatebin_data.tar]`, `opengist volume_dumps=[opengist_opengist_data.tar]`, `kimai volume_dumps=[kimai_kimai_db_data.tar, kimai_kimai_var.tar]` — the 40-class DOES reach the off-site tier, and PrivateBin's planted 1 MB came back byte-identical from its off-site snapshot into the checking folder. **(1) stands and is now strictly larger than when written:** R-354 shows the bare completion is also reported over a unit that DID carry a data leg, because the off-site restore never replays volume dumps at all. | CC |
|
||||
| **R-356** | **The off-site restore refuses for all 40 no-drive apps, says the app "is not installed" when it is running, and then gives an instruction those apps make impossible.** `ReconstituteFromOffsite` refuses when `GetStackHDDPath(stack)` is empty (`offbox_reconstitute.go:208-227`); for a 40-class app that is ALWAYS empty, because they are offered no storage field at deploy time (R-352's own measurement). Observed 2026-08-21 22:21 on `privatebin` while it was `deployed=true, state=running, healthy`: „a(z) privatebin nincs telepítve, ezért nincs hová visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra az alkalmazást ugyanerre a helyre…". **The predicate is "has an HDD path"; the sentence says "is not installed"; for this class they are different things**, and the remedy offered cannot be carried out. This is also what the 2026-08-21 afternoon OpenGist journey hit before falling back to the local restore (R-353). | **OPEN — HIGH** | — | Separate the two questions. A 40-class app has a destination — the system data path — and the restore already knows it. **RE-CHECKED 2026-08-22 against the corrected placement framing: this row SURVIVES UNCHANGED and is strengthened by it.** The 40-class has no `HDD_PATH` *because the architecture puts its hot data inside the guest by design* (`01-topology-and-trust.md:150-152`) — so an absent HDD path is the NORMAL state for most of the catalogue, and a restore that reads it as "the app is not installed" is misreading a correct configuration, not reporting a misconfiguration. The one sentence in this row that leaned on the old framing — "because they are offered no storage field at deploy time (R-352's own measurement)" — should read: *because these apps have no bulk volume to place, so the field correctly does not exist.* | CC |
|
||||
| **R-357** | **The DESTRUCTIVE restore has no free-space gate; the three that exist are all on non-destructive paths.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the gates sit at `offbox_restore.go:231` (scratch restore), `:297` (prepare-full) and `:423` (place-to-live). Proven 2026-08-21 23:11 with 300 KB free and 1 MB to write: it stopped `paperless-ngx`, failed halfway (`rsync … No space left on device (28)`), left the data directory holding **2 of 5** planted entries, and restarted the app. The message is honest but is raw rsync output. | **OPEN — MEDIUM** | — | Same gate, same wording as `:297`, before the stop. | CC |
|
||||
| **R-358** | **A FAILED scratch restore leaves a partial copy that the product then offers as a full restore source — and the destructive restore runs from it and reports success.** `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only whether the directory exists and is non-empty; its comment defers completeness to `PlaceOffsiteRestore`, which stats top-level placements, not files. Proven 2026-08-21 22:54-22:56 against a deliberately corrupted store: the restore failed honestly (`ciphertext verification failed`, 54 files, 15 of 16 originals), the wizard then offered „Teljes visszaállítás indítása", and pressing it reported `ok=true`. **The failure is detected and then forgotten.** | **OPEN — MEDIUM** | — | Record the failure against the scratch and refuse to place from it until it is re-prepared. | CC |
|
||||
| **R-359** | **The off-site restic store is never verified by anything, ever.** The complete set of restic verbs in the controller is `restore, snapshots, backup, unlock, stats, init, forget, prune, cat` — **no `check`**. The agent's `RestoreTest` is PBS-tier only. Established 2026-08-21 by deliberately corrupting one pack: `restic check` catches it immediately („ciphertext verification failed", „Fatal: repository contains errors"), and the product only meets the damage when a customer is already trying to recover. | **OPEN — MEDIUM** | — | A periodic `restic check` (structure) with an occasional `--read-data`, reported like any other backup verdict. Note PBS already has verify jobs; this is the tier that does not. | CC |
|
||||
|
||||
Reference in New Issue
Block a user