docs: P6 EXECUTED — containerd store moved to SSD2; reboot deferred (R-209, R-209a)
gates / gates (push) Successful in 7s

Operator ruled "proceed" after reading the pre-analysis; the storageReserved
condition CC made the move conditional on was applied with it. The pre-analysis
is retained verbatim as §11-bis so the ruling can be audited against what it
was given.

MOVED WITH ZERO LOSS, verified BEFORE the original was touched on four
independent observables: 550,891 = 550,891 filesystem entries; 448 = 448
trusted.overlay xattrs (load-bearing — overlayfs stacking rides them, which is
why rsync -X is not optional); 37,243 = 37,243 hardlinks; byte-identical
meta.db sha256. Verified again after: identical image/tag/volume ID SETS,
cache 2.782 GB / 38 records, ceiling still 30 GiB, pg 4 DBs / 31 tables /
175,135,767 B, redis DBSIZE 2437. End-to-end proof was a real build on the
relocated store, rc=0 — not an inventory count.

k3s was never at risk, and that was established BEFORE stopping anything: it
runs a separate containerd, so Gitea, the registry, the hub, PBS, Longhorn and
~160 pods stayed up. Only the two jarr-* dev containers were affected.

TRAP FOUND WHILE PROVING THE GUARD, and it is the reusable part:
RequiresMountsFor on a path with NO mount unit is a SILENT NO-OP — containerd
started normally against an absent-but-unmounted path. A typo'd guard buys
nothing and says nothing: the built-but-never-wired shape again. The guard was
therefore verified positively at the unit level (Requires= AND
After=mnt-ssd_2.mount on both units), and refusal proven with a genuinely
absent DEVICE via a temporary synthetic .mount unit — /mnt/ssd_2 hosts 12 live
Longhorn replicas and must never be unmounted, and editing fstab on a
production host risks emergency mode at boot. Result:
"Job containerd.service/start failed with result 'dependency'", is-active
inactive. It refused rather than creating an empty store on the root disk.

storageReserved on SSD2 0 -> 80 GB; still Schedulable=True at 76.34%.

NOT PROVEN, AND NOT GLOSSED (R-209a): the operator ruled DO NOT REBOOT, so the
move has not survived a reboot — which is this task's own standard for
validated. The MECHANISM is proven; the CONSEQUENCE (that a real boot mounts
/mnt/ssd_2 before containerd starts, in this host's actual ordering) is not,
and mount-ordering reasoning is exactly the class this project has been burned
by. So: the pre-move tree stays MOVED ASIDE, not deleted (34.3 GB, which is why
/ sits at 54% and not lower), and felhom-store-postboot-check.service — oneshot,
enabled, dry-run PASS at install — validates it automatically at the next boot,
asserting positively that /var/lib/containerd does NOT exist. Read
/var/log/felhom-store-postboot-check.log after the next boot; on PASS the
rollback copy can go and ~34 GB returns.

Gates: repo_gates.py --fast — all OK, rc=0 (run separately from this commit).
This commit is contained in:
2026-08-05 10:27:40 +02:00
parent ee9d9bf203
commit c917251eeb
3 changed files with 198 additions and 12 deletions
+15 -8
View File
@@ -56,6 +56,13 @@ over; and the file was restored **byte for byte identical**. *(R-201)*
quietly around it: the warning meant to catch exactly this **could never fire** (fixed and proved), quietly around it: the warning meant to catch exactly this **could never fire** (fixed and proved),
and the weekly cleanup is still forbidden from touching the thing that grows (next session). and the weekly cleanup is still forbidden from touching the thing that grows (next session).
*(R-205 … R-211)* *(R-205 … R-211)*
- **2026-08-05** — **The build data now lives on the second SSD, moved with nothing lost.** All 345
images, every saved volume and both development databases came through identical — checked before
the original was touched and again afterwards, and confirmed by running a real build on the moved
copy. **The server's own services never went down:** Gitea, the registry, the hub and the backup
system run on a separate system and stayed up throughout. The second SSD now also **reserves 80 GB**
for this, so the storage layer can no longer quietly claim the space and repeat what happened to the
first disk. The old copy is kept as the way back until the machine next restarts. *(R-209)*
- **2026-08-05** — **Three of the four recovery crutches removed.** The reset code works first time; - **2026-08-05** — **Three of the four recovery crutches removed.** The reset code works first time;
a credential re-issue no longer blocks off-site backups on a healthy machine; and the default a credential re-issue no longer blocks off-site backups on a healthy machine; and the default
restore no longer quietly returns the wrong thing. *(R-204 items 13, R-196)* restore no longer quietly returns the wrong thing. *(R-204 items 13, R-196)*
@@ -79,14 +86,14 @@ over; and the file was restored **byte for byte identical**. *(R-201)*
## Waiting on you ## Waiting on you
- **Whether to move the build data onto the second SSD at all — and CC's advice is now "probably - **One thing to read after the machine next restarts — and nothing to do until then.** You told me
not".** You ruled "cap it, then move it". The cap is in and it did the job on its own: the disk sits not to restart DooPlex, so I did not, and the move to the second SSD has therefore never been
at 54% with 199 GB free, and moving is no longer a rescue. The second SSD looks roomy (91% free) and through a restart. It works right now and nothing was lost, but a restart is the one test that
today the move is safe by a wide margin — but it is the **only** disk of the four with **no space matters for this kind of change, and it has not happened. **I made it check itself:** whenever the
reserved for itself**, and its volumes are allowed to claim 6.9× more than they currently use. If machine next starts, for any reason, it writes a plain PASS or FAIL line to
they ever grow into that, the move would push it below the same floor that just took the first disk `/var/log/felhom-store-postboot-check.log`. **If it says PASS, the old copy can be deleted and 34 GB
out of service. **If you want the move, do it together with reserving ~80 GB on that disk; otherwise comes back.** Until then I have deliberately kept that old copy, which is the only quick way back —
leaving it where it is costs nothing now.** *(R-209)* it is why the disk sits at 54% rather than lower. *(R-209a)*
- **One list to rule on: 193 old images that exist only on this machine.** 131 controller versions and - **One list to rule on: 193 old images that exist only on this machine.** 131 controller versions and
62 hub versions are not in the registry, so they cannot be re-downloaded — all of them old 62 hub versions are not in the registry, so they cannot be re-downloaded — all of them old
(controller up to 0.135.0, hub up to 0.57.0; everything newer is safely in the registry). **Nothing (controller up to 0.135.0, hub up to 0.57.0; everything newer is safely in the registry). **Nothing
@@ -488,10 +488,188 @@ happened to restart the pod. `Synced/Healthy` was true and beside the point. Fil
--- ---
## 11. P6 — the SSD2 move: pre-analysis only. **NOT EXECUTED.** ## 11. P6 — the SSD2 move: **EXECUTED 2026-08-05 on operator authorisation**
**Stopped here as specified. Nothing was moved, no daemon was reconfigured for SSD2, no mount unit was > **Executed after the operator ruled "proceed", having read the pre-analysis below. The
written.** The operator must be present. > `storageReserved` condition CC named was applied with it. The operator subsequently ruled
> explicitly: DO NOT REBOOT DooPlex — so the reboot validation is DEFERRED, and what that leaves
> unproven is stated in §11.9 rather than glossed.**
### 11.0 Execution record
**Blast radius, established before anything was stopped:** k3s runs a **separate** containerd
(`/run/k3s/containerd/containerd.sock`), so **Gitea, the registry, the hub, PBS, Longhorn and all
~160 pods stayed up throughout**. The only workloads affected were the two `jarr-*` dev containers.
`ctr -n moby containers list` confirmed exactly those two on the system containerd.
| # | step | result |
|---|---|---|
| 1 | stop `docker.service`, `docker.socket`, `containerd.service` | all `inactive` |
| 2 | `mkdir /mnt/ssd_2/containerd` (mode 711, root:root) | ok |
| 3 | `rsync -aHAX --numeric-ids` | **550,891 files, 33,276,370,041 B, 4 m 58 s** |
| 4 | `mv /var/lib/containerd → /var/lib/containerd.pre-move-2026-08-05` | **moved aside, never deleted** |
| 5 | `root = '/mnt/ssd_2/containerd'` in `/etc/containerd/config.toml` (backup kept) | one-line diff |
| 6 | `RequiresMountsFor=/mnt/ssd_2` drop-ins on **both** `containerd.service` and `docker.service` | resolved into both units |
| 7 | start + verify | see below |
| 8 | `storageReserved = 80 GB` on SSD2 | applied |
**`-X` is load-bearing** — overlayfs snapshots carry `trusted.overlay.*` xattrs, and losing them
would silently corrupt layer stacking.
### 11.1 Copy verified BEFORE the original was touched
Four independent observables, all matching:
| check | source | destination |
|---|---:|---:|
| filesystem entries | 550,891 | **550,891** |
| `trusted.overlay` xattrs | 448 | **448** |
| files with link count >1 (hardlinks) | 37,243 | **37,243** |
| `io.containerd.metadata.v1.bolt/meta.db` sha256 | `c60ecfb4ee134255…` | **`c60ecfb4ee134255…`** |
(Byte totals differ by +14 MB — ext4 block allocation, not content.)
### 11.2 Nothing was lost
| | pre-move | post-move |
|---|---:|---:|
| unique images | 345 | **345 — identical ID set** |
| tag lines | 350 | **350 — identical set** |
| docker volumes | 8 | **8 — identical set** |
| build cache | 2.782 GB / 38 records | **2.782 GB / 38 records** |
| GC ceiling | 30 GiB | **30 GiB, still in force** |
| postgres databases | 4 | **4** |
| postgres `jarr_dev` bytes | 175,135,767 | **175,135,767** |
| postgres public tables | 31 | **31** |
| redis `DBSIZE` | 2437 | **2437** |
**End-to-end proof, not just inventory:** a real `./build.sh` run against the relocated store
completed `rc=0`, produced an image, and grew the cache 2.782 → 4.012 GB. The scratch tag was removed
and `:latest` restored to the anchor digest afterwards.
### 11.3 The store is demonstrably on SSD2
```
containerd config dump → root = '/mnt/ssd_2/containerd'
du /mnt/ssd_2/containerd → 35,500,183,552 B
/var/lib/containerd → does not exist
```
### 11.4 The guard — and a trap found while proving it
**First attempt did NOT refuse, and that is a finding worth more than the test.** With
`RequiresMountsFor=/mnt/absent-guard-test-2026-08-05` — a path with **no mount unit at all**
containerd **started normally**. `RequiresMountsFor` on a path systemd has no mount unit for is a
**silent no-op**. So a typo in the guard path buys nothing and says nothing: the same
*built-but-never-wired* shape this project has shipped repeatedly.
**So the guard was verified positively, at the unit level:**
```
containerd.service Requires=mnt-ssd_2.mount After=mnt-ssd_2.mount
docker.service Requires=mnt-ssd_2.mount
```
**Then the refusal itself was proven with a genuinely absent DEVICE.** `/mnt/ssd_2` hosts **12 live
Longhorn replicas** and must never be unmounted, and editing `fstab` on a production host risks
emergency mode at boot — so a **temporary synthetic `.mount` unit** pointing at a nonexistent UUID
was used instead, with a temporary `RequiresMountsFor` on the real `containerd.service`:
```
A dependency job for containerd.service failed.
systemd: Dependency failed for containerd.service - containerd container runtime.
systemd: containerd.service: Job containerd.service/start failed with result 'dependency'.
containerd is-active: inactive
```
**containerd REFUSED to start rather than creating an empty store on the root disk** — exactly the
required behaviour, and exactly the signature the ep0 runbook describes. All test artifacts were
removed; only `10-ssd2-mount.conf` remains on each unit and the guard resolves to `/mnt/ssd_2` alone.
### 11.5 Longhorn — the condition CC made the move conditional on
`storageReserved` on SSD2 raised **0 → 80,000,000,000 B**, covering the containerd tree so Longhorn's
scheduler can no longer oversubscribe into the space the store needs.
| disk | available | reserved | `Schedulable` |
|---|---:|---:|---|
| **Crucial-SSD-240G (SSD2)** | 179,516,211,200 (**76.34%**) | **80,000,000,000** | **True** |
| default-disk (SSD1) | 236,978,176,000 (49.64%) | 143,217,954,816 | **True** |
| HDD4 | 1,793,798,963,200 (60.82%) | 536,870,912,000 | True |
| HDD2 | 2,470,864,486,400 (24.91%) | 536,870,912,000 | **False** — pre-existing, untouched |
SSD2 remains schedulable with the reserve applied. **Rollback: patch `storageReserved` back to 0.**
### 11.6 A standing post-boot assertion was installed
`felhom-store-postboot-check.service` (oneshot, `After=docker.service`, enabled) runs
`/usr/local/sbin/felhom-store-postboot-check.sh` at **every** boot and logs to
`/var/log/felhom-store-postboot-check.log`. It asserts, with **positive** observables:
- `/mnt/ssd_2` is mounted;
- containerd's effective `root` is `/mnt/ssd_2/containerd`;
- **`/var/lib/containerd` does NOT exist** — i.e. no empty store was recreated on the root disk, which
is the precise trap the guard exists to prevent;
- docker sees ≥100 images;
- both dev containers are running.
Dry-run at install time: **RESULT: PASS** on all six lines. It was proven working *before* being
relied upon.
### 11.7 Disk position after the move
| | value |
|---|---:|
| `/` | 224 GB used, **199 GB free, 54%** |
| `/mnt/ssd_2` | 52 GB used, **157 GB free, 25%** |
| store on SSD2 | 35,500,183,552 B |
| **rollback copy retained on `/`** | **34,299,498,496 B** |
**`/` did not drop below 54% because the rollback copy is deliberately still there.** Reclaiming it
returns ~34 GB — see §11.9.
### 11.8 Rollback, complete and one sequence
1. `sudo systemctl stop docker docker.socket containerd`
2. `sudo mv /var/lib/containerd.pre-move-2026-08-05 /var/lib/containerd`
3. `sudo cp /etc/containerd/config.toml.pre-move-2026-08-05 /etc/containerd/config.toml`
4. `sudo rm -rf /etc/systemd/system/{containerd,docker}.service.d/10-ssd2-mount.conf`
5. `sudo systemctl disable --now felhom-store-postboot-check.service && sudo rm /etc/systemd/system/felhom-store-postboot-check.service /usr/local/sbin/felhom-store-postboot-check.sh`
6. `sudo kubectl patch nodes.longhorn.io -n longhorn-system dooplex --type merge -p '{"spec":{"disks":{"Crucial-SSD-240G":{"storageReserved":0}}}}'`
7. `sudo systemctl daemon-reload && sudo systemctl start containerd docker`
8. `sudo rm -rf /mnt/ssd_2/containerd`
### 11.9 **NOT PROVEN: the reboot. Stated, not glossed.**
**The operator ruled explicitly that DooPlex must not be rebooted, so the reboot validation did not
happen.** Uptime is unbroken (7 weeks 6 days, since 2026-06-10) — verified, not assumed.
**By this task's own standard — *"a move that has not survived a reboot has not been validated"* — the
move is NOT fully validated,** and it would be dishonest to close it as if it were. Precisely:
- **PROVEN (mechanism):** the guard is wired into both units (`Requires=` + `After=mnt-ssd_2.mount`),
and containerd **refuses to start** when a required mount's device is genuinely absent.
- **NOT PROVEN (consequence):** that a real boot sequence mounts `/mnt/ssd_2` before containerd starts
*on this host, in the real ordering*. Mount-ordering reasoning is exactly the class this project has
been burned by, and the mechanism proof does not substitute for it. `CLAUDE.md` prefers the
consequence over the mechanism; here only the mechanism is available.
**Two consequences follow, and both are deliberate:**
1. **The rollback copy `/var/lib/containerd.pre-move-2026-08-05` (34.3 GB) STAYS** until a reboot has
validated the move. Deleting it now would trade a cheap 34 GB for the only cheap way back.
2. **The next reboot — whenever it happens, planned or not — validates this automatically.** The
post-boot service writes `RESULT: PASS` or `RESULT: FAIL` to
`/var/log/felhom-store-postboot-check.log`. **Read that file after the next boot**; on `PASS`, the
rollback copy can be removed and ~34 GB returned to `/`.
Tracked as **R-209a**.
---
## 11-bis. P6 — the pre-analysis the ruling was made on
*(Retained as written, before execution, so the ruling can be audited against what it was given.)*
**First, the framing changed underneath this probe.** P3's prune took `/` to 53% used and un-stuck **First, the framing changed underneath this probe.** P3's prune took `/` to 53% used and un-stuck
SSD1's Longhorn disk. **The move is no longer a rescue; it is a deliberate choice.** Per standing rule 4, SSD1's Longhorn disk. **The move is no longer a rescue; it is a deliberate choice.** Per standing rule 4,
+2 -1
View File
@@ -185,7 +185,8 @@ drill needed — the fourth is R-193, and until it goes the recovery is still op
| **R-206** | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC | | **R-206** | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC |
| **R-207** | **`DRY_RUN=1` on `node-housekeeping.sh` is NOT non-mutating — it destroys the metric history it is supposed to let you inspect** | **READY (S) — NEW 2026-08-05** | — | `write_metrics()` (`node-housekeeping.sh.j2:119-150`) has **no `DRY_RUN` guard at all**`DRY_RUN` appears in it only inside a log line — and it is called from an **unconditional `EXIT` trap** (`:150`). A dry run therefore atomically renames over the live node_exporter textfile, overwriting `node_housekeeping_last_success_timestamp_seconds` with *now* and `reclaimed_bytes` with ~0 — **resetting the staleness clock `HousekeepingStale` watches and erasing the 8-week reclaim history**. Confirmed by reading the source in both the 2026-08-05 audit and this spike; **neither run executed it**, so the history survives. Fix: guard `write_metrics` on `DRY_RUN`, or have the trap skip it. Pairs naturally with R-206 (same file, same role) | CC | | **R-207** | **`DRY_RUN=1` on `node-housekeeping.sh` is NOT non-mutating — it destroys the metric history it is supposed to let you inspect** | **READY (S) — NEW 2026-08-05** | — | `write_metrics()` (`node-housekeeping.sh.j2:119-150`) has **no `DRY_RUN` guard at all**`DRY_RUN` appears in it only inside a log line — and it is called from an **unconditional `EXIT` trap** (`:150`). A dry run therefore atomically renames over the live node_exporter textfile, overwriting `node_housekeeping_last_success_timestamp_seconds` with *now* and `reclaimed_bytes` with ~0 — **resetting the staleness clock `HousekeepingStale` watches and erasing the 8-week reclaim history**. Confirmed by reading the source in both the 2026-08-05 audit and this spike; **neither run executed it**, so the history survives. Fix: guard `write_metrics` on `DRY_RUN`, or have the trap skip it. Pairs naturally with R-206 (same file, same role) | CC |
| **R-208** | **Every Felhom Go build re-downloads its modules because `ARG VERSION` sits ABOVE the module-download layer — ~440 MB of dead cache per build, 90.5 GB of the 157 GB** | **READY (S) — NEW 2026-08-05, ROOT CAUSE PROVEN** | — | **Measured, not inferred.** All **208** retained `go mod download` records carried **`Usage count: 1`** — not one was ever reused in ~a month of builds. The mechanism was isolated by four controlled builds: an unchanged tree rebuilt with the **same** `--build-arg VERSION``RUN go mod download` **CACHED**; the **same** tree with a **new** `VERSION`**executed**. `COPY go.mod ./` stays CACHED either way, which is the tell: a `COPY`'s key is content-based, while a `RUN`'s key includes the stage **environment**, and `ARG VERSION`/`ARG GIT_COMMIT` are declared *before* the download in `felhom-controller/controller/Dockerfile`. Since every real build passes a fresh version, the layer is invalidated **every single time**. **`felhom.eu/hub/Dockerfile` has the identical defect** (`ARG VERSION`/`ARG BUILD_TIME` above `COPY go.mod go.sum*``RUN go mod download`) — and because both Dockerfiles produce byte-identical `buildx du` description strings, the 208 records are a COMBINED count and must not be attributed to one project. **Fix shape (one line each, not applied here):** move the `ARG VERSION`/`ARG GIT_COMMIT`/`ARG BUILD_TIME` declarations down to just above the final `go build`. **Worth more than the cap and the move combined** — the cap bounds the symptom, this removes the source. `build.sh`'s `rm -rf` + `cp -a` and its host-side `go mod tidy` were **ruled out by fingerprinting**: the tree is byte-identical across runs and `tidy` is a no-op | CC | | **R-208** | **Every Felhom Go build re-downloads its modules because `ARG VERSION` sits ABOVE the module-download layer — ~440 MB of dead cache per build, 90.5 GB of the 157 GB** | **READY (S) — NEW 2026-08-05, ROOT CAUSE PROVEN** | — | **Measured, not inferred.** All **208** retained `go mod download` records carried **`Usage count: 1`** — not one was ever reused in ~a month of builds. The mechanism was isolated by four controlled builds: an unchanged tree rebuilt with the **same** `--build-arg VERSION``RUN go mod download` **CACHED**; the **same** tree with a **new** `VERSION`**executed**. `COPY go.mod ./` stays CACHED either way, which is the tell: a `COPY`'s key is content-based, while a `RUN`'s key includes the stage **environment**, and `ARG VERSION`/`ARG GIT_COMMIT` are declared *before* the download in `felhom-controller/controller/Dockerfile`. Since every real build passes a fresh version, the layer is invalidated **every single time**. **`felhom.eu/hub/Dockerfile` has the identical defect** (`ARG VERSION`/`ARG BUILD_TIME` above `COPY go.mod go.sum*``RUN go mod download`) — and because both Dockerfiles produce byte-identical `buildx du` description strings, the 208 records are a COMBINED count and must not be attributed to one project. **Fix shape (one line each, not applied here):** move the `ARG VERSION`/`ARG GIT_COMMIT`/`ARG BUILD_TIME` declarations down to just above the final `go build`. **Worth more than the cap and the move combined** — the cap bounds the symptom, this removes the source. `build.sh`'s `rm -rf` + `cp -a` and its host-side `go mod tidy` were **ruled out by fingerprinting**: the tree is byte-identical across runs and `tidy` is a no-op | CC |
| **R-209** | **Should the containerd store move to SSD2 at all? — SSD2 is the ONLY Longhorn disk with `storageReserved = 0`, and the move is safe today but breaches the 25% floor if its volumes ever inflate** | **WAITING-ON-OPERATOR** — NEW 2026-08-05 | an operator ruling | **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator | | **R-209** | ~~**Should the containerd store move to SSD2 at all?**~~ | **EXECUTED 2026-08-05 on operator ruling — reboot validation DEFERRED → R-209a** | — | **Operator ruled "proceed" having read the pre-analysis; CC's `storageReserved` condition was applied with it.** Moved with **zero loss, verified on four independent observables BEFORE the original was touched** (550,891 entries = 550,891; **448 = 448 `trusted.overlay` xattrs**; 37,243 = 37,243 hardlinks; byte-identical `meta.db` sha256) and again after (identical image/tag/volume ID **sets**, cache 2.782 GB/38 records, pg 4 DBs / 31 tables / 175,135,767 B, redis 2437). **`-X` is load-bearing** — overlayfs stacking rides `trusted.overlay.*`. End-to-end proof was a **real build** on the relocated store, `rc=0`. **k3s was never at risk and this was established before stopping anything:** it runs a separate containerd, so Gitea, the registry, the hub, PBS, Longhorn and ~160 pods stayed up; only the two `jarr-*` dev containers were affected. `storageReserved` on SSD2 **0 → 80 GB**, still `Schedulable=True` at 76.34%. **A TRAP was found while proving the guard, and it is the reusable part: `RequiresMountsFor` on a path with NO mount unit is a SILENT NO-OP** — containerd started normally against an absent-but-unmounted path, so a typo'd guard buys nothing and says nothing (the *built-but-never-wired* shape again). The guard was therefore verified positively at the unit level (`Requires=` **and** `After=mnt-ssd_2.mount` on both units), and refusal was then proven with a genuinely absent **device** — via a temporary synthetic `.mount` unit, because `/mnt/ssd_2` hosts 12 live Longhorn replicas and must never be unmounted, and editing `fstab` on a production host risks emergency mode at boot: `Job containerd.service/start failed with result 'dependency'`, `is-active: inactive`. **Rollback is one documented sequence** (audit §11.8); the pre-move tree is **moved aside, not deleted**. Evidence: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §11 | |
| **R-209a** | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** | operator + CC | **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator |
| **R-210** | **Which of the 345 local images may be deleted — 131 controller tags and 62 hub tags exist ONLY on this box and are not recoverable by `docker pull`** | **WAITING-ON-OPERATOR** — NEW 2026-08-05 | an operator ruling | **Nothing was deleted; this is a list, not an action.** The registry was queried directly: `felhom-controller` has **76** tags in Gitea vs **207** locally, `felhom-hub` **45** vs **107**. The **131 + 62 local-only tags are all OLD** — controller `0.39.0``0.135.0` plus `v0.35.0``v0.39.0`, hub `0.9.0``0.57.0` plus `v0.7.2``v0.13.0` — while everything from controller `0.136.0` and hub `0.58.0` upward IS in the registry and therefore re-pullable. **Size the prize honestly before spending a decision on it:** per-tag sizes sum to 139.29 GB, but that double-counts shared layers — `docker system df` puts the **real** dedup'd image footprint at **31.02 GB with 27.02 GB reclaimable**, i.e. an order of magnitude less than the build cache P3 already returned. `docker image prune -a` would remove 343 of 345 (only `redis:7-alpine` and `postgres:16-alpine` are held by running containers). **CC's view: not worth doing for the space** — it buys ~27 GB against 199 GB now free, and its only real benefit is dropping unrecoverable clutter | operator | | **R-210** | **Which of the 345 local images may be deleted — 131 controller tags and 62 hub tags exist ONLY on this box and are not recoverable by `docker pull`** | **WAITING-ON-OPERATOR** — NEW 2026-08-05 | an operator ruling | **Nothing was deleted; this is a list, not an action.** The registry was queried directly: `felhom-controller` has **76** tags in Gitea vs **207** locally, `felhom-hub` **45** vs **107**. The **131 + 62 local-only tags are all OLD** — controller `0.39.0``0.135.0` plus `v0.35.0``v0.39.0`, hub `0.9.0``0.57.0` plus `v0.7.2``v0.13.0` — while everything from controller `0.136.0` and hub `0.58.0` upward IS in the registry and therefore re-pullable. **Size the prize honestly before spending a decision on it:** per-tag sizes sum to 139.29 GB, but that double-counts shared layers — `docker system df` puts the **real** dedup'd image footprint at **31.02 GB with 27.02 GB reclaimable**, i.e. an order of magnitude less than the build cache P3 already returned. `docker image prune -a` would remove 343 of 345 (only `redis:7-alpine` and `postgres:16-alpine` are held by running containers). **CC's view: not worth doing for the space** — it buys ~27 GB against 199 GB now free, and its only real benefit is dropping unrecoverable clutter | operator |
| **R-211** | **Prometheus has no config-reloader — a rules change reaches the pod and is never read** | **READY (S) — NEW 2026-08-05** | — | Found while verifying R-205 rather than by looking for it. The `mon-system/prometheus` Deployment runs **one** container (`prom/prometheus:v3.12.0`) with **no `configmap-reload`/`prometheus-config-reloader` sidecar**. After the ArgoCD sync the updated `node-housekeeping-alerts.yml` was present **inside the pod** (`grep -c "and on(instance)"` → 3 on the mounted symlink) while the Prometheus **rules API still served the old expression** — for **4+ minutes**, with no error anywhere. It only took effect after an explicit `POST /-/reload`. **The consequence is general, not specific to R-205: every rule edit in this repo since the stack was built has silently not applied until something happened to restart the pod** — so "committed and synced" has never meant "in force", and ArgoCD reporting `Synced/Healthy` is true and beside the point. `--web.enable-lifecycle` IS already set, so the fix is small: add a reloader sidecar watching the ConfigMap, or a `checksum/config` pod annotation so a rules change rolls the pod. **Same class as the four *built-but-never-wired* seams** — the control exists, nothing walks it | CC | | **R-211** | **Prometheus has no config-reloader — a rules change reaches the pod and is never read** | **READY (S) — NEW 2026-08-05** | — | Found while verifying R-205 rather than by looking for it. The `mon-system/prometheus` Deployment runs **one** container (`prom/prometheus:v3.12.0`) with **no `configmap-reload`/`prometheus-config-reloader` sidecar**. After the ArgoCD sync the updated `node-housekeeping-alerts.yml` was present **inside the pod** (`grep -c "and on(instance)"` → 3 on the mounted symlink) while the Prometheus **rules API still served the old expression** — for **4+ minutes**, with no error anywhere. It only took effect after an explicit `POST /-/reload`. **The consequence is general, not specific to R-205: every rule edit in this repo since the stack was built has silently not applied until something happened to restart the pod** — so "committed and synced" has never meant "in force", and ArgoCD reporting `Synced/Healthy` is true and beside the point. `--web.enable-lifecycle` IS already set, so the fix is small: add a reloader sidecar watching the ConfigMap, or a `checksum/config` pod annotation so a rules change rolls the pod. **Same class as the four *built-but-never-wired* seams** — the control exists, nothing walks it | CC |