Files
felhom.eu/STATUS.md
T
admin 2584dfb938
gates / gates (push) Successful in 8s
R-193: a guest rebuild silently drops the offsite tier; R-192 cause established
Operator confirms no hub-side offsite config change, so the regression was not an
action. Evidence: demo-hp's controller went 0.187.0 -> 0.192.0 at 06:12:18 with a
new config hash and the agent re-keyed its leaf three minutes earlier — a guest
rebuild. The last pre-rebuild report shows the tier fully healthy: escrowed, last
success 02:16:39Z, 15 snapshots, 40.9 MB. No offsite object in the 108 reports
since.

Mechanism: the restic credential is delivered once. demo-hp consumed its secret on
2026-07-23; the rebuilt controller has no copy and no way to request another.
demo-felhom survived the SAME rebuild only because its secret was still unconsumed
— it consumed it four seconds after its config hash changed and was reporting
offsite again 76 seconds later. That difference was luck, not design.

Also sharpens R-192: the self-heal's guard refuses when any report since the
consume carried an offbox target, but that query reads the OLDEST 500 reports —
all of which predate the rebuild. Healthy history before a rebuild is not evidence
the credential still works, which is why the automation that exists for this case
declined to act.
2026-08-04 09:05:53 +02:00

90 lines
5.8 KiB
Markdown

# STATUS — what works, what's broken, what's next
**Updated 2026-08-03.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority on open work; this
> page restates part of it in plain words, and **nothing may exist only here**. **Not `CONTEXT.md`**,
> which is technical state written for Claude Code — keep the two separate. **Maintenance:** update
> at the end of every session in which something shipped, broke, or was decided. One screen; cut
> items rather than extend it.
## What works right now
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer,
who sets their own password. They install apps from a catalogue of fifty-three, share files over the
home network, and open apps from a launcher or a shared link. Backups run on their own to three
places — the machine's drive, a second drive, and an encrypted off-site copy — and a customer can
restore files and app data from the drive alone. Apps come back after a power cut: hard-reset the demo
box six times, everything returned every time, and an app switched off deliberately stayed off.
Proven end to end on real hardware.
## What's broken
- **One demo machine has no off-site copy of its app data, and rebuilding it is what took it away.**
`demo-hp` was rebuilt on 3 August; before that its off-site backup was healthy and had run
successfully at 04:16 that morning (15 snapshots). The rebuilt machine came up without it and has
not had it in 108 check-ins since. **The cause is that the off-site password is delivered exactly
once and a rebuilt machine cannot ask for another** — the other demo machine survived the same
rebuild only because it happened to have an unused password waiting for it, and recovered in 76
seconds. Nothing about that difference was designed. *(R-193)*
- **The daily email about it tells you the wrong story**, and the automatic repair that exists for
this declines without saying why. The message says the password was never applied; it was, on
23 July, and worked for eleven days. *(R-192)*
- **The weekly off-site backup reports FAILED although it worked.** It uploads correctly and then
trips on a tidy-up step it is deliberately not allowed to perform, so the job ends in an error and
you get an email. The backup itself is safe and on the endpoint. Both demo machines do it; one
setting per machine fixes it. *(R-191)*
- **The off-site copy can be erased by the machine that made it.** The credential that writes it can
also delete it. A daily snapshot is armed as a stopgap.
*(R-95, R-87)*
## What shipped recently
- **The on-machine backup copy has now been proved to restore — by the machines themselves.** Both
demo machines restored their own on-machine backup into a throwaway machine overnight, booted it,
checked it and destroyed it, without being asked: 84 and 109 seconds each. Every restore proof we
had before this was of the *off-site* copy; the copy an ordinary recovery would actually use had
never been tested on either machine. Both also proved their off-site copy on the same night, one
after the other rather than at once, which is the machine deciding for itself what to do first.
*(closes the last open half of R-86/R-185)*
- **A backup copy the machine was never allowed to read — and could not tell you about**, on both
demo machines. The permission was one command; the silence was the real fault, and the machine now
checks whether it may read each copy it depends on and says so when it may not. *(R-185)*
- **Three ways the alarm system was misreporting its own work — all fixed.** None of them ever risked
data. **(1)** When the machine proved a backup restores, that result could vanish if the agent was
restarted in the following quarter-hour — and yesterday's change made the gap a week rather than a
day, because the machine correctly refuses to re-prove an archive it has already proven. It is now
written to disk with the result and survives. This was caught happening, not predicted: a real
14.5 GB off-site restore passed and left no record at all. **(2)** Every release had about a
fifty-fifty chance of emailing you a failure for a release that worked; the version tag is now
published after the binary, and a new check catches the opposite mistake so nothing is traded away.
**(3)** A released binary can now be rebuilt by anyone and checked against the fingerprint you
approve — until today, rebuilding produced different bytes. *(R-189, R-188, R-186)*
## What we're working on
- **Now:** nothing outstanding.
- **Next:** proving the off-site *app-data* copy can actually be restored — the one tier nothing
tests unattended. Most of the machinery it needed arrived with the restore-test change below.
*(R-87)*
- **After:** the off-site copy that the machine making it can still erase. *(R-95)*
## Waiting on you
- **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a
session log; nothing suggests anyone else saw it. *(R-132)*
- **One small question, not urgent.** The automatic check cannot see which version you have told
machines to install, only which ones exist. Closing that needs either a password given to the build
server or a check inside the hub itself. *(R-184)*
- **Nothing else.**
## Changed since last update
- **2026-08-04** — Both demo machines proved their on-machine backup restores, on their own,
overnight — the copy an ordinary recovery uses, never tested until now. Found while checking: the
weekly off-site backup reports failure after a successful upload. *(R-185, R-191)*
- **2026-08-03** — Fixed three ways the alarm system misreported itself: a proof of a working backup
that could vanish on a restart (seen happening), a release that emailed a failure for a release
that worked, and a released binary nobody could rebuild and check. *(R-189, R-188, R-186)*