Files
felhom.eu/STATUS.md
T
admin ee9d9bf203
gates / gates (push) Successful in 7s
docs: SPIKE — DooPlex build-cache containment; R-205..R-211 (2026-08-05)
Spike output only; no production Go code. The one shipped change rides in
homelab-manifests 6808a4b (R-205, the monitoring rule).

VERDICT: mechanism confirmed, with one correction and one refutation.

- CONFIRMED: builder.gc IS honoured under the containerd worker and DOES evict.
  Proven by naming a 440 MB `go mod download` record present at build N and
  absent by N+2 — not by absence of an error.
- CORRECTED: honoured ONLY in the `policy` array form. The flat form is
  SILENTLY ignored — daemon starts, logs nothing, keeps its defaults.
  `dockerd --validate` returned "configuration OK" for a bogus key AND for a
  config that then crashed the daemon. The oracle is `docker buildx inspect`.
- REFUTED: Docker's `data-root` would NOT move the cache — it moves 0.62 GB.
  The 181.4 GB belongs to the system containerd (`root` in
  /etc/containerd/config.toml).

P3 (operator-approved) executed: prune claimed 156.9 GB, the filesystem
returned 150.35 GB (the 6.5 GB gap is layers shared with images), SYNCHRONOUSLY
— / went 86% -> 53% used, and Longhorn's default disk went
Schedulable=False (DiskPressure) -> Schedulable=True (18.85% -> 50.32%).

P7 root-caused the largest item and it is NOT the cap: all 208 `go mod download`
records had Usage count 1. Isolated by controlled builds — same VERSION build-arg
-> CACHED, new VERSION -> executed, byte-identical tree. `ARG VERSION`/`ARG
GIT_COMMIT` sit ABOVE the module-download step, and a RUN's cache key includes
the stage environment. Both Dockerfiles have it. One line each to fix -> R-208.

P6 NOT EXECUTED — stops at the operator, as specified. Pre-analysis: the move is
safe as measured (+38.8 pp above the 25% floor) but SSD2 is the only Longhorn
disk with storageReserved=0 and is overcommitted 6.9x; at full inflation the move
lands 12 pp BELOW the floor. The prune removed the move's urgency, so CC
recommends against it unless ~80 GB is reserved on SSD2 -> R-209.

Register: R-205 (CLOSED, shipped), R-206 (Ansible: cap + prune + narrowed Docker
ban), R-207 (DRY_RUN guard), R-208 (ARG ordering), R-209/R-210 (operator),
R-211 (Prometheus has no config-reloader — rules changes have never applied
until something restarted the pod; found while verifying R-205).

Gates: repo_gates.py --fast — all OK, rc=0 (run separately from this commit).
2026-08-05 10:02:46 +02:00

8.3 KiB
Raw Blame History

STATUS — what works, what's broken, what's next

Updated 2026-08-05.

A view, not a source. documentation/backlog/OPEN-ITEMS.md is the authority on open work; this page restates part of it in plain words, and nothing may exist only here. Not CONTEXT.md, which is technical state written for Claude Code — keep the two separate. Maintenance: update at the end of every session in which something shipped, broke, or was decided. One screen; cut items rather than extend it.

What works right now

A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who sets their own password. They install apps from a catalogue of fifty-three, share files over the home network, and open apps from a launcher or a shared link. Backups run on their own to three places — the machine's drive, a second drive, and an encrypted off-site copy.

And the whole backup promise is now proved. On 4 August we destroyed a machine on purpose and deleted a marked file from its disk. Using the recovery code you saved: the backup key came back identical, character for character; the existing off-site store opened rather than starting over; and the file was restored byte for byte identical. (R-201)

What's broken

  • A customer still cannot do that recovery alone — but only one step is left. Getting from "the key is recoverable" to "the file is back" took four steps that appeared in no instructions. Three are fixed today (R-204): the local reset-code tool works on the first try instead of needing the controller restarted; re-issuing the storage credential no longer falsely marks the recovery key "stale" (which used to stop every off-site backup and invite the one act that would have destroyed the recovered key); and the everyday restore now says in plain Hungarian that it returned the app's settings and database and not your documents, and names the button that does. The step that remains is the first one: a rebuilt machine cannot get a storage credential by itself, because the one-time password was used up by its predecessor — so you still have to press Re-issue. That is a design decision waiting on you, below. (R-193)
  • Rebuilding a machine still throws away its off-site backup HISTORY. The machine invents the key that encrypts its own off-site backups, and a rebuilt machine invents a brand-new one. Both demo machines did this on 34 August — 51 backups (~1.2 GB) between them. The old key now survives the recovery ceremony, and a changed key now raises an alarm the same day, but the rebuild itself still starts a fresh history. (R-193)
  • One screen still tells the customer something we cannot yet promise. The „elárvult tároló" card says the old backups may later be restorable with the matching recovery code. That is true for machines that re-seal from now on and false for anything already orphaned — and the machine cannot tell which case it is in. We deliberately did not patch the sentence: a conditional promise that can still be wrong is worse there than a vague one. (R-202)
  • The off-site copy can be erased by the machine that made it. The credential that writes it can also delete it. A daily snapshot is armed as a stopgap. (R-95, R-87)

What shipped recently

  • 2026-08-05The DooPlex server's disk is out of danger: 86% full → 54%, and the storage layer is unstuck. The cause was leftover working data from building our own software — 157 GB of it, growing about 5 GB a day, which nothing was allowed to delete. 148 GB came back in 86 seconds. The storage layer had already stopped accepting new copies of any volume onto that disk; that is fixed the same day. A 30 GB ceiling is now in place and was proved to work by deliberately overfilling it and watching it evict — not by assuming the setting took. Two things that failed quietly around it: the warning meant to catch exactly this could never fire (fixed and proved), and the weekly cleanup is still forbidden from touching the thing that grows (next session). (R-205 … R-211)
  • 2026-08-05Three of the four recovery crutches removed. The reset code works first time; a credential re-issue no longer blocks off-site backups on a healthy machine; and the default restore no longer quietly returns the wrong thing. (R-204 items 13, R-196)
  • 2026-08-04 (night)The drill PASSED, end to end, on real hardware. (R-201)
  • 2026-08-04 — The folder-left-out-of-the-backup problem fixed both halves: the app and its backup look in the same directory, and a backup that misses a folder marked essential reports incomplete instead of success. (R-203)
  • 2026-08-04 — The hub now keeps the off-site backup key when a machine re-seals, instead of only the whole-machine one, and a machine can fetch its own sealed package back. (R-198, R-199)
  • 2026-08-04 — The daily false alarm about David is gone; a vanished permission now repairs itself and says it had to; the weekly off-site backup stopped reporting failure after a successful upload. (R-195, R-190, R-191)

What we're working on

  • Next: the retention proof. The hub keeping the old sealed key when a machine re-seals is the one remaining link that has never run outside a test. Proving it needs a second deliberate wipe on the spare demo machine, and it is its own procedure. (R-198)
  • Then: the orphaned-backup deletion you asked for (below), and the off-site copy the machine can still erase. (R-193, R-95)

Waiting on you

  • Whether to move the build data onto the second SSD at all — and CC's advice is now "probably not". You ruled "cap it, then move it". The cap is in and it did the job on its own: the disk sits at 54% with 199 GB free, and moving is no longer a rescue. The second SSD looks roomy (91% free) and today the move is safe by a wide margin — but it is the only disk of the four with no space reserved for itself, and its volumes are allowed to claim 6.9× more than they currently use. If they ever grow into that, the move would push it below the same floor that just took the first disk out of service. If you want the move, do it together with reserving ~80 GB on that disk; otherwise leaving it where it is costs nothing now. (R-209)
  • One list to rule on: 193 old images that exist only on this machine. 131 controller versions and 62 hub versions are not in the registry, so they cannot be re-downloaded — all of them old (controller up to 0.135.0, hub up to 0.57.0; everything newer is safely in the registry). Nothing was deleted. Worth knowing before you spend time on it: they only account for about 27 GB against 199 GB now free, so this is about clutter, not space. (R-210)
  • The one-shot credential decision — this is now the last thing between a customer and an unaided recovery. A rebuilt machine has no storage credential of its own, so an operator must press Re-issue. Everything after that point is now self-service. Deciding how a rebuilt machine should get a credential is the remaining design question. (R-193, R-204 item 4)
  • The recovery screen you described has been priced, and it can be built. A freshly installed machine that finds a sealed package waiting should say so, offer a box for the recovery code, and show what would come back before doing anything. One thing to weigh, deliberately not decided: that screen is reachable by anyone with the household's dashboard password, and the preview reveals backup dates and app names. (R-193)
  • The orphaned backups on the storage box — you said delete, and it is still owed. About 1.2 GB across the two demo machines, in set-aside stores nobody can open and nothing prunes. It wants its own session rather than riding along with other work. (R-193)
  • (decided 4 Aug) You chose not to keep a copy of the backup key on the Proxmox host, which makes the customer's own recovery code the only route back from a rebuild. (R-193)
  • A job, not a decision: the hub password needs changing. A diagnostic command printed it into a session log; nothing suggests anyone else saw it. (R-132)
  • One small question, not urgent. The automatic version check cannot see which version you have told machines to install, only which ones exist. (R-184)