aa62449694
gates / gates (push) Successful in 8s
Two boxes, two DIFFERENT supply paths, so the session proved the disk shape and the delivery route rather than one of them twice. demo-hp (layout proof, --golden <local volid>): mp0 at /var/lib/felhom, backup=1, 70G, no mp1; /var/lib/docker and /mnt/sys_drive both real mounts of its subdirectories via fstab; one df figure and one device id (64519) on all three paths; reboots 3/3 with the binds surviving each. demo-felhom (pipeline proof, --force-gitea-golden): fetch_verify succeeding against the vouched manifest for BOTH artifacts -- 'verified sha256 54e2a4c431daf580... matches the hub manifest' for the golden, a7763d31... for the agent. 250G single volume, grep -c '^mp1:' = 0, reboots 3/3. Journey proven on both, endpoint-level: claim -> deploy -> back up -> restore, with a planted marker returning byte-identical on each box. Ceiling measured gone: 65 GiB and 233 GiB available to a recovery unit, against 19 and 45. R-165 -> IMPLEMENTED, not PROVEN-LIVE, on the operator's ruling. B2, which that row records as the bulkhead's replacement, fired live for the first time and does refuse per app, delete nothing and alert -- but it is checked only in captureAllRecoveryUnits while runVolumeDumps writes the bulk unguarded, and its 'the previous unit is untouched' claim was measured false (182,272 B dump replaced by 2,147,666,432 B under a manifest still dated 06:34:26). -> R-181. New: R-179 (uninstall leaves NAS network-storage units), R-180 (--archive-storage not cross-checked against the ACL grant; 403 at step 8/8 after root@pam is rotated), R-181. Third instance of R-115 recorded (agent 0.120.0 unpublished). No code written, no version bumps -- this was a runbook.
151 lines
10 KiB
Markdown
151 lines
10 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
|
|
|
**Updated 2026-08-03.**
|
|
|
|
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority on open work; this
|
|
> page restates part of it in plain words, and **nothing may exist only here**. **Not `CONTEXT.md`**,
|
|
> which is technical state written for Claude Code — keep the two separate. **Maintenance:** update
|
|
> at the end of every session in which something shipped, broke, or was decided. One screen; cut
|
|
> items rather than extend it.
|
|
|
|
## What works right now
|
|
|
|
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer,
|
|
who sets their own password. They install apps from a catalogue of fifty-three, share files over the
|
|
home network, and open apps from a launcher or a shared link. Backups run on their own to three
|
|
places — the machine's drive, a second drive, and an encrypted off-site copy — and a customer can
|
|
restore files and app data from the drive alone. Proven end to end on real hardware.
|
|
|
|
**Apps come back after a power cut.** The machine tells an app the customer switched off from one
|
|
that simply did not come back, and waits for the system to finish starting before deciding instead of
|
|
glancing once, five seconds in. Hard-reset the demo box six times in a row: everything came back every
|
|
time, and an app switched off deliberately stayed off every time.
|
|
|
|
## What's broken
|
|
|
|
**The off-site copy can be erased by the machine that made it** — the credential that writes it can
|
|
also delete it. A daily snapshot is armed as a stopgap, and we have never restored from that copy.
|
|
*(R-95, R-87)*
|
|
|
|
**Three apps out of fifty-three kept their data where backups never looked.** They reported healthy;
|
|
the data would vanish on the next update. Two are fixed, the third is now clear to fix because it is
|
|
installed nowhere. *(R-156)*
|
|
|
|
**The reserve does not guard the step that actually fills the disk.** The rule that refuses a backup
|
|
when space runs low is checked at the wrong moment: the big write happens first, unchecked, and only
|
|
the small write after it is refused. So the thing meant to stop a runaway backup is the thing it
|
|
runs past. Worse, when it does refuse it says *"your last good copy is untouched"* — and we measured
|
|
that copy being overwritten by the earlier step anyway. Found by deliberately filling a rebuilt demo
|
|
machine. Nothing was deleted and you were emailed, both correctly. *(R-181)*
|
|
|
|
**When that happens, only one page says so** — no email, no alert. The page that answers "is this app
|
|
backed up?" is the one that stays silent. *(R-158)*
|
|
|
|
**The checks now have two nets, and the second one emails you.** Every repository has one command
|
|
that runs all of its checks; it runs by itself before every push and refuses a push that fails. That
|
|
one lives on the workstation and can be skipped. So the build server now runs the same checks again,
|
|
on a machine that does not care who pushed or what they typed — and **when they fail it sends you an
|
|
email**, because a red mark on a page nobody watches is not a warning. Proven with a real broken
|
|
change, not assumed. The one thing it still cannot do is *stop* the change: every change here goes
|
|
straight to the main copy with no review step, so there is no point in the road for it to stand at.
|
|
It notices, quickly, and tells you. *(R-29, R-161, R-168, R-169)*
|
|
|
|
**A filling disk now warns the customer before anything breaks, and a failed backup now reaches you.**
|
|
Until today the first sign that a disk was filling up was a backup that did not happen — nothing said
|
|
anything beforehand. Two things changed. The customer is now warned while there is still room to act,
|
|
naming the drive and how much space is left, in plain Hungarian that says what to do about it. And
|
|
when one app's backup fails for any reason, **you** are told which app and why, with the disk figures
|
|
attached — the page that answers "is this app backed up?" was, until now, the one page that never
|
|
said. The customer is deliberately *not* told about that second one: they can free up space, but they
|
|
can do nothing about a backup that failed, so telling them would only alarm them.
|
|
|
|
Both were proven on the demo machine by actually filling a disk. One detail is worth knowing because
|
|
it is why there are two rules and not one: the serious warning fired when free space dropped below a
|
|
fixed amount while the disk was only 91% full — a percentage on its own would have missed it.
|
|
|
|
**These went in *before* the partition change deliberately.** The partition being removed is also a
|
|
barrier against a runaway backup filling the space the machine needs to run; putting the warnings in
|
|
first means that when it comes down, the thing watching is already working and already tested.
|
|
*(R-167, R-158)*
|
|
|
|
**The backup partition is gone — and both demo machines now run on the new shape.** Each was wiped
|
|
and rebuilt from the new base image on 3 August and taken through the whole journey a customer takes:
|
|
set the machine up, install an app, back it up, restore it. One storage area instead of two, and the
|
|
space a backup can use went from 19 GB to 65 GB on the small machine and from 45 GB to 233 GB on the
|
|
big one. The wall does not move; it stops existing. Each machine was rebooted three times over and
|
|
came back correctly every time. **The two machines were rebuilt deliberately differently** — the
|
|
first from a copy of the image held locally, the second by the ordinary route a real customer takes,
|
|
fetching the image and checking it against the fingerprint we publish, so both the disk shape and the
|
|
delivery route are now proven rather than one twice.
|
|
|
|
**Their previous demo apps and data are gone.** That was the point of a wipe and you approved it; the
|
|
machines now carry a couple of small test apps instead.
|
|
|
|
**What replaced the wall.** It was quietly doing a second job — keeping a runaway backup from eating
|
|
the space the machine needs to keep running. That job is now explicit: if a backup would push the disk
|
|
below a safe reserve, **that one app's backup is refused, its last good copy is left exactly as it
|
|
was, and you are told.** Nothing is ever deleted to make room; every app has only one local copy, so
|
|
"delete the oldest" would always mean destroying some other app's only copy.
|
|
|
|
**How the shape was chosen — worth one line, because it was not the obvious one.** Three ways of doing
|
|
it were built and rebooted rather than argued about. All three worked. They differed in what they
|
|
quietly broke: one put your backups inside Docker's own storage, where the normal way of fixing a sick
|
|
Docker would wipe them; another exposed all of Docker's internals to the part of the system that
|
|
manages your drives. The third does neither, and costs one extra line of configuration.
|
|
|
|
**The new base image is now switched on**, so any machine installed from here on gets the new shape.
|
|
The tester's box is untouched and keeps working exactly as before; it gets the new shape whenever it
|
|
is reinstalled. *(R-165, R-178)*
|
|
|
|
## What we're working on
|
|
|
|
- **Now:** the last app whose data was never saved; today's decisions written down.
|
|
- **Next:** fixing the reserve so it guards the step that fills the disk, and stopping it from
|
|
claiming your last good copy is untouched when it is not *(R-181)*. The partition merge itself is
|
|
done and proven on both machines.
|
|
- **After:** rebuilding how the machine records whether an app is meant to be running.
|
|
|
|
## Waiting on you
|
|
|
|
- **How a new version reaches a machine — and it has now been forgotten a third time.** Pushing the
|
|
installer publishes it: half a minute later every new machine downloads it, with no staging and no
|
|
way back but another push. And publishing is a step we remember rather than one the release
|
|
performs. On 3 August the new agent — the half of the partition merge that runs on the machine —
|
|
turned out to have been built and installed on both demo machines but **never published**, so a
|
|
rebuild would have quietly put the *old* one back and proved a version nobody ships. Caught before
|
|
the wipe and fixed in ten minutes, but only because someone happened to look. Nothing is installing
|
|
today, so this is the cheapest moment to settle it. *(R-110, R-115)*
|
|
- **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a
|
|
session log; nothing suggests anyone else saw it. *(R-132)*
|
|
- **Nothing on the partition merge itself** — the shape is chosen, built, and now proven on both demo
|
|
machines. The tester's box needs no conversion: it will simply be reinstalled. What is left is the
|
|
reserve defect above, which is ours to fix, not yours to decide. *(R-165, R-181)*
|
|
|
|
## Changed since last update
|
|
|
|
- **2026-08-03** — Both demo machines wiped and rebuilt from the new base image, and taken through
|
|
set-up → install an app → back it up → restore it. The backup space ceiling is gone and measured.
|
|
The hard stop that replaced the old partition was fired for the first time on real hardware — it
|
|
refused, it deleted nothing, it emailed you, **and it turned out to be watching the wrong step**;
|
|
that is now the top thing to fix.
|
|
|
|
- **2026-08-02** — The false "host offline" warning is fixed. The hub's database was supposed to be in
|
|
a mode where reading a page cannot block a machine's status update; a one-word difference meant that
|
|
setting had **never taken effect**, for the hub's whole life. Fixed and verified live. **Also found:
|
|
the hub's own database is in no automatic backup** — it holds every machine's emergency password.
|
|
Filed, not yet fixed.
|
|
|
|
- **2026-08-02** — Boot recovery finished; six hard resets, everything back every time. Two instances
|
|
of the same hole — starting an app whose external drive was missing — were found by reading the code
|
|
and fixed the same day.
|
|
|
|
- **2026-08-02** — Thirteen mechanical checks had built up and nothing ran most of them; two were
|
|
failing quietly. Fixed, and the arrangement that replaced them is described above.
|
|
- **2026-08-02** — Decided: the 20 GB backup partition goes away and shares space with app data. That
|
|
changes the disk layout, so it happens before any machine is installed outside the house. **Measured
|
|
since:** no machine outside the house is registered yet, so this is as cheap now as it will ever be;
|
|
and the two demo machines can simply be reinstalled rather than converted.
|
|
- **2026-08-02** — Decided: only this machine and the tester's box are protected; every other box,
|
|
demo boxes included, may be broken or reinstalled freely. Two of the three apps that never saved
|
|
their data are fixed; this page created.
|