41dbecb264
gates / gates (push) Successful in 8s
R-167/R-158 shipped and proven live (controller v0.191.x, hub v0.89.0): two new capability-map rows PROVEN-LIVE with live citations, and 07-backup-architecture.md §7.5's closing claim "nothing warns when an app crosses the line" is now false and rewritten (S-1: an architectural contract changed in the same session). §7.5 also gains the caveat that its size bound is ONE BOX'S, not the fleet's. Part 3 SPIKE (audits/SPIKE-r165-mp1-merge-2026-08-02.md): M1-M5 measured, NO layout touched. Three findings the merge session must not re-derive: "the layout" is not one thing (200G/50G vs 50G/20G vs 16G/8G); mp1 is a BULKHEAD and not only a ceiling, so after the merge an overflow reaches /var/lib/docker; the golden fails closed on the split in four places. D-a's condition (1) is currently SATISFIED — no external box is in the hub's register, and both demo boxes are Tier 0 and reinstallable. Recommendation given, choice NOT made — it ends at the operator's ruling. CONTEXT.md S-11 (D-c's routing, and why R-158's own backup_failed proposal was overruled) and S-12 (the monitoring landed BEFORE the merge). STATUS.md gains the plain-language section and the merge decision, with two older entries trimmed so the page did not grow. New rows R-174 (closed same session), R-175, R-176, R-177; each ID grepped free before minting.
121 lines
8.2 KiB
Markdown
121 lines
8.2 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
|
|
|
**Updated 2026-08-02.**
|
|
|
|
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority on open work; this
|
|
> page restates part of it in plain words, and **nothing may exist only here**. **Not `CONTEXT.md`**,
|
|
> which is technical state written for Claude Code — keep the two separate. **Maintenance:** update
|
|
> at the end of every session in which something shipped, broke, or was decided. One screen; cut
|
|
> items rather than extend it.
|
|
|
|
## What works right now
|
|
|
|
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer,
|
|
who sets their own password. They install apps from a catalogue of fifty-three, share files over the
|
|
home network, and open apps from a launcher or a shared link. Backups run on their own to three
|
|
places — the machine's drive, a second drive, and an encrypted off-site copy — and a customer can
|
|
restore files and app data from the drive alone. Proven end to end on real hardware.
|
|
|
|
**Apps come back after a power cut.** The machine tells an app the customer switched off from one
|
|
that simply did not come back, and waits for the system to finish starting before deciding instead of
|
|
glancing once, five seconds in. Hard-reset the demo box six times in a row: everything came back every
|
|
time, and an app switched off deliberately stayed off every time.
|
|
|
|
## What's broken
|
|
|
|
**The off-site copy can be erased by the machine that made it** — the credential that writes it can
|
|
also delete it. A daily snapshot is armed as a stopgap, and we have never restored from that copy.
|
|
*(R-95, R-87)*
|
|
|
|
**Three apps out of fifty-three kept their data where backups never looked.** They reported healthy;
|
|
the data would vanish on the next update. Two are fixed, the third is now clear to fix because it is
|
|
installed nowhere. *(R-156)*
|
|
|
|
**Local backups get 20 GB while apps get 50 GB.** An app that outgrows the smaller space stops being
|
|
backed up locally — and the off-site copy is made from the local one, so that stops too. Nothing is
|
|
lost: the last good copy is kept intact. *(R-163)*
|
|
|
|
**When that happens, only one page says so** — no email, no alert. The page that answers "is this app
|
|
backed up?" is the one that stays silent. *(R-158)*
|
|
|
|
**The checks now have two nets, and the second one emails you.** Every repository has one command
|
|
that runs all of its checks; it runs by itself before every push and refuses a push that fails. That
|
|
one lives on the workstation and can be skipped. So the build server now runs the same checks again,
|
|
on a machine that does not care who pushed or what they typed — and **when they fail it sends you an
|
|
email**, because a red mark on a page nobody watches is not a warning. Proven with a real broken
|
|
change, not assumed. The one thing it still cannot do is *stop* the change: every change here goes
|
|
straight to the main copy with no review step, so there is no point in the road for it to stand at.
|
|
It notices, quickly, and tells you. *(R-29, R-161, R-168, R-169)*
|
|
|
|
**A filling disk now warns the customer before anything breaks, and a failed backup now reaches you.**
|
|
Until today the first sign that a disk was filling up was a backup that did not happen — nothing said
|
|
anything beforehand. Two things changed. The customer is now warned while there is still room to act,
|
|
naming the drive and how much space is left, in plain Hungarian that says what to do about it. And
|
|
when one app's backup fails for any reason, **you** are told which app and why, with the disk figures
|
|
attached — the page that answers "is this app backed up?" was, until now, the one page that never
|
|
said. The customer is deliberately *not* told about that second one: they can free up space, but they
|
|
can do nothing about a backup that failed, so telling them would only alarm them.
|
|
|
|
Both were proven on the demo machine by actually filling a disk. One detail is worth knowing because
|
|
it is why there are two rules and not one: the serious warning fired when free space dropped below a
|
|
fixed amount while the disk was only 91% full — a percentage on its own would have missed it.
|
|
|
|
**These went in *before* the partition change deliberately.** The partition being removed is also a
|
|
barrier against a runaway backup filling the space the machine needs to run; putting the warnings in
|
|
first means that when it comes down, the thing watching is already working and already tested.
|
|
*(R-167, R-158)*
|
|
|
|
## What we're working on
|
|
|
|
- **Now:** the last app whose data was never saved; today's decisions written down.
|
|
- **Next:** merging the small backup partition into the large one — **the warnings for it are already
|
|
done and working**, so this step is now only the partition change. It needs one decision from you
|
|
first (below).
|
|
- **After:** rebuilding how the machine records whether an app is meant to be running.
|
|
|
|
## Waiting on you
|
|
|
|
- **How a new version reaches a machine.** Pushing the installer publishes it — half a minute later
|
|
every new machine downloads it, with no staging and no way back but another push. And publishing is
|
|
a step we remember rather than one the release performs, forgotten twice: a fix can be live here
|
|
and still not reach a new machine. Nothing is installing today, so this is the cheapest moment to
|
|
settle both. *(R-110, R-115)*
|
|
- **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a
|
|
session log; nothing suggests anyone else saw it. *(R-132)*
|
|
- **The partition merge: one decision, now measured.** Removing the backup partition also removes a
|
|
barrier — today a runaway backup is refused on its own and cannot touch the space the machine needs
|
|
to run; afterwards it can, and a machine out of that space is stopped, not slow. So: do we add a
|
|
hard stop that refuses a backup before it eats the last of the room, or do we rely on the new
|
|
warnings? The recommendation is the hard stop, because it keeps exactly what the barrier gave us.
|
|
**Second question, which only you can answer:** does the tester's box need converting in place, or
|
|
can it be reinstalled? It does not report to the hub, so nothing here can tell. *(R-165, R-176)*
|
|
|
|
## Changed since last update
|
|
|
|
- **2026-08-02** — The false "host offline" warning is fixed, and the cause was not what it looked
|
|
like. The hub's database was supposed to be in a mode where reading a page cannot block a machine's
|
|
status update — the code said so, but a one-word syntax difference meant the setting had **never
|
|
taken effect**, for the hub's whole life. So opening an operator page could make a machine's report
|
|
fail; two failures in a row crossed the half-hour threshold and sent you an alert about a machine
|
|
that was up and healthy. It had already done that twice that day. Now genuinely fixed and verified
|
|
live. **Also found while checking it: the hub's own database is not in any automatic backup** — it
|
|
holds every machine's emergency password and the escrow records. Filed, not yet fixed.
|
|
|
|
- **2026-08-02** — Boot recovery finished: the machine records what the customer asked for, and waits
|
|
for the system to finish starting before deciding what is missing. Six hard resets, everything back
|
|
every time. A hole the previous day's change had opened — starting an app whose external drive was
|
|
missing — was found by reading the code, reproduced on the demo box first, and fixed the same day.
|
|
**A second instance of the same hole was found and fixed today**, on the path that restarts an app
|
|
after an interrupted backup.
|
|
|
|
- **2026-08-02** — Thirteen mechanical checks had built up across the four repositories and nothing
|
|
ran most of them; two were failing quietly, one since 14 July. Both fixed, and the arrangement that
|
|
replaced them is described above.
|
|
- **2026-08-02** — Decided: the 20 GB backup partition goes away and shares space with app data. That
|
|
changes the disk layout, so it happens before any machine is installed outside the house. **Measured
|
|
since:** no machine outside the house is registered yet, so this is as cheap now as it will ever be;
|
|
and the two demo machines can simply be reinstalled rather than converted.
|
|
- **2026-08-02** — Decided: only this machine and the tester's box are protected; every other box,
|
|
demo boxes included, may be broken or reinstalled freely. Two of the three apps that never saved
|
|
their data are fixed; this page created.
|